Geotags:

SwitzerlandTanzania

With MeditronFO, EPFL brings transparency to medical LLMs.

The framework developed in Lausanne documents data, code and evaluations to make the use of language models in healthcare more inspectable.

MeditronFO: Open Medical AI Illustration, with Researchers, Clinicians, and Healthcare Data Used to Make Clinical LLMs More Transparent and Auditable
Clinicians and researchers analyze MeditronFO in a sober editorial composition, including medical diagrams, a European network, and Swiss partner logos: the image highlights transparency, scientific collaboration, and verifiability of clinical LLMs. (Illustration: Innovando.News)

Healthcare is one of the fields where the adoption of theartificial intelligence and big data It is proceeding with greater caution, not out of lack of interest, but because error has direct consequences for people. Large language models are already used, or tested, to support doctors' work: they can help guide a diagnosis, summarize clinical documentation, suggest priorities, or assist in emergency room decisions. The critical point, however, remains verifiability. If a system produces a recommendation without making its data, methods, and training criteria accessible, independent control becomes fragile.

It is on this point that he intervenes MeditronFO, announced by theFederal Institute of Technology in LausanneAccording to the University of Vaud, this is the first completely open framework for building large language model doctorsThe emphasis is not only on the availability of the final model, but on the entire development chain: data, code, training procedures, documentation, and evaluations. In a regulated, sensitive, and professionally accountable industry, this difference is not insignificant.

The project was born in Laboratory for Intelligent Global Health and Humanitarian Response Technologies, known as LiGHT, within the School of Computer and Communication Sciences at EPFL. The technical basis is the open one already used by Meditron, released in 2023, but the evolution presented now aims to make the entire construction method more transparent. The stated goal is to enable the "medicalization" of basic open models, that is, their adaptation to the healthcare domain through clinical knowledge, documented datasets, and specific validation procedures.

Why the opening weighs more than just the model release

In AI parlance, the word "open" is often used inconsistently. A model may make weights available, but keep the datasets, filtering steps, training choices, or evaluation protocols opaque. For clinical use, this partial openness creates an auditing problem: hospitals, researchers, authorities, and the medical community can see the results, but not always reconstruct how they were obtained.

MeditronFO attempts to bridge this gap. The framework has been applied to completely open base models, including Elm, EuroLLM e Apertus, a Swiss model developed by EPFL and ETH Zurich as part of the Swiss AI Initiative. The move is significant because it shifts the focus from isolated performance to development traceabilityIn medicine, knowing the path a system takes to a response can be as important as measuring the accuracy of the response itself.

Xavier Theimer-Lienhard, a doctoral student who leads Meditron at LiGHT, summarizes the principle with a direct comparison between clinical training and model training.

“We would never trust a doctor whose training couldn't be verified: the same standard should apply to AI in healthcare.”

The statement clarifies an aspect often overlooked in the debate on generative systems: reliability does not coincide with the availability of an effective interface. Instead, it requires documentation, reproducibility, and the ability to be controlled. The issue also concerns the security and privacy, because healthcare models process knowledge and data that may border on sensitive information, even when they do not directly include identifiable medical records.

The open component does not by itself eliminate the risk of bias, errors, or misuse. However, it reduces one of the main barriers to independent analysis. In other words, it does not automatically guarantee clinical safety, but it makes the work of verification, comparison, and improvement more concrete.

MeditronFO: AI editorial visualization applied to medicine, including clinical language models, healthcare data, auditable processes, and collaboration between research and hospitals.
The combination of AI, health data, and medical practice raises questions of trust, privacy, and accountability, making open training and assessment processes particularly important in clinical LLMs. (Illustration: Envato)

Clinicians involved in the construction, not just the final test

A distinctive feature of the project is the involvement of healthcare professionals from the earliest stages. According to EPFL, MeditronFO was developed with clinicians involved in data selection, output validation, and the identification of potential critical issues. This approach reduces the risk of developing formally sophisticated tools that are poorly aligned with daily practice.

Validation also passes through MOOVE, an acronym for Massive Open Online Validation and Evaluations, an environment through which clinicians participate in the evaluation and continuous improvement of models. For a technology designed to interact with real-world medicine, this is a crucial point: the challenge isn't simply to correctly answer a question, but to do so under conditions similar to those in which a physician works, with time constraints, incomplete information, and operational responsibilities.

The framework combines public medical datasets with clinician-reviewed synthetic data derived from medical examinations, guidelines, and real-world patient cases. The EPFL source also indicates the use of a set of expert-curated clinical datasets drawn from over 46.000 clinical practice guidelinesThis is important data because it shows the scale of the documentation used to specialize the models, but it must be read in conjunction with the issue of quality: in healthcare, the quantity of incorporated knowledge is not enough if it is not accompanied by selection, review, and contextualization.

Mary-Anne Hartley, doctor and director of ,LiGHT connects the competitiveness of models to the active role of clinicians and communities.

“Our findings show that competitive medical models can be built with the active involvement of clinicians and communities.”

This shift has organizational implications. Medical AI isn't presented as a tool brought to bear on the healthcare system from the outside, but as a software architecture to be developed with those familiar with procedures, recurring errors, institutional constraints, and differences between contexts. For healthtech companies, university centers, and hospitals, the lesson is clear: the transfer from the center of research and development Clinical practice requires governance, not just computational power.

MeditronFO: Image dedicated to fully open medical LLMs, with references to transparency, Swiss research, clinical collaboration, and responsible use of artificial intelligence.
The comparison published by EPFL shows the performance on HealthBench of proprietary, open-weights and fully open models: MeditronFO is placed in the range of fully open models, with results superior to the respective base models (Graph: EPFL)

Technical results and the role of Switzerland's high computing power

In terms of performance, EPFL reports that each MeditronFO model outperformed its respective base model. The best result indicated in the press release concerns Apertus-70B-MeditronFO, which improved performance on medical tests 6,6 percentage points compared to the underlying model. The data should be interpreted with caution: benchmarks are useful for comparing technical configurations, but they are not a substitute for clinical evaluations in real-world environments.

The industrial relevance of the case also lies in the infrastructure that supports it. The development of MeditronFO was supported by Swiss AI Initiative, collaboration between EPFL, ETH Zurich and CSCS, the Swiss National Supercomputing Center. According to the initiative's website, the program was launched in December 2023 with over 10 million GPU hours on Alps, the supercomputer operated by CSCS, and with funding from 20 million Swiss francs by the Federal Polytechnics Domain.

These numbers show why the competition in healthcare AI isn't just about algorithms and datasets. It requires computing power, distributed expertise, evaluation infrastructures, and a critical mass of research. The Swiss AI Initiative itself claims to benefit from the contributions of over 800 researchers, included 70 professors focused on AI, from more than ten academic institutions in SwitzerlandIn questo quadro, il supercomputer it is not a simple technical accelerator, but a component of scientific sovereignty.

The connection with Apertus also reinforces this interpretation. The project website describes it as a fully open foundation model for sovereign AI, developed by the Swiss AI Initiative with EPFL, ETH Zurich, and CSCS. Apertus claims openness in its weights, data, code, methods, and alignment principles. MeditronFO uses this foundation as one of the models to specialize in the medical field, demonstrating how an upstream approach to openness can generate more controllable vertical applications.

MeditronFO: AI editorial visualization applied to medicine, including clinical language models, healthcare data, auditable processes, and collaboration between research and hospitals.
AI applied to healthcare is not just about algorithms, but also about training, governance, and control of the tools used by doctors, hospitals, and researchers: the difference between closed models and auditable systems becomes crucial in clinical contexts (Illustration: Envato)

From experimentation to the wards, the testing ground in Tanzania

The next step will not be just a new model comparison. According to EPFL, the team is preparing clinical trials at multiple sites, from Switzerland to the Tanzania, to evaluate how doctors use AI in real-world healthcare settings. The trials will observe whether clinicians follow or reject the generated recommendations and how these decisions impact patient care.

The multi-year project mentioned by EPFL, called MED.USE, also aims to understand whether AI can help improve the quality of care by reducing unnecessary treatments and interventions. This is a delicate area: a decision support system can be useful if it improves the quality of available information, but it can become problematic if it introduces automation or encourages excessive trust in algorithmic recommendations.

Hartley specifically calls for the need to measure the effect in treatment pathways, not just in technical tests.

“It is important to get real-world feedback based on patient outcomes.”

Clinical AI will be more credible when it can be inspected

For the healthcare sector, MeditronFO suggests a clear direction: clinical AI will be more credible when it can be inspected, adapted, and evaluated by independent communities. For businesses, this means that proprietary models will have to contend with growing expectations of transparency. For institutions, the issue concerns the ability to define rules that distinguish between formal and actual openness.

It is unrealistic to think that all medical systems based on machine learning algorithm e deep learning become completely open. It's plausible, however, that healthcare will demand higher standards of documentation, auditability, and accountability, especially when AI enters decision-making areas. In this sense, MeditronFO shouldn't be seen simply as a new model, but as an experiment in technical governance: it makes visible what often remains hidden behind a generated response.

The crucial question, therefore, isn't the abstract promise of more powerful AI. It's about the type of infrastructure that healthcare systems, universities, and regulators intend to accept. If the models become part of clinical practice, trust must be based on verifiable elements: data provenance, independent audits, transparently measured performance, and identifiable human responsibility. MeditronFO places these elements at the center of the discussion, leaving the clinical trial to assess their operational viability.

Meditron Open Sources: Auditable AI Workflows for Clinical LLMs

Here are three insights that might interest you:

From EPFL a new multimodal model for more flexible AI
Apertus is the Swiss "open AI" for a more transparent future.
A cutting-edge healthcare hub based entirely on AI in Geneva

MeditronFO: Open Medical AI Illustration, with Researchers, Clinicians, and Healthcare Data Used to Make Clinical LLMs More Transparent and Auditable
MeditronFO is presented as an open medical AI system supported by EPFL, LiGHT, ETH Zurich, CSCS and the Swiss AI Initiative: transparency becomes a design choice, not an afterthought (Illustration: Innovando.News)

Location

COMMENTS

Leave a comment