Google’s open MedGemma clears peer review: what Nature Medicine says it can and can’t do
A peer-reviewed study of Google’s open medical AI models finds small, downloadable models rivalling far larger ones on X-rays, eye scans and pathology — and it’s clear about where they still fall short.

On October 6, 2026, Nature Medicine published a peer-reviewed study of MedGemma, Google’s open medical vision-language models built on Gemma 3. The 4B and 27B models beat their general-purpose base models on medical text, X-ray, pathology and eye-scan tasks and rival much larger systems, but the authors say benchmarks are only a first step towards clinical use.
Most AI models that can “read” a chest X-ray are either closed systems you rent through an API or narrow tools trained for one job. On October 6, Nature Medicine published a peer-reviewed study of a different kind of model: MedGemma, a family of open medical AI models from Google that anyone can download, run on a single GPU and fine-tune. The headline finding is that small, specialised open models can match — and sometimes beat — far larger general-purpose systems on medical images. The fine print is just as important: this is evidence about benchmarks, not about patients.
The paper, “An open vision-language model for diverse medical applications”, has 88 authors and was led by Andrew Sellergren, David F. Steiner, Daniel Golden and Lin Yang. It is the peer-reviewed version of a technical report Google first posted in July 2025, when most of the models were released.
What is MedGemma?
MedGemma is built on Gemma 3, Google’s open-weight model family. The collection has four parts:
- MedSigLIP — a 400-million-parameter image encoder, the part of the system that turns a scan into something a language model can reason about. Google fine-tuned it on millions of medical image–text pairs from radiology, dermatology, pathology and ophthalmology.
- MedGemma 4B — a small model that takes text, images or both.
- MedGemma 27B — a larger multimodal model that can also work with electronic health records (EHRs).
- MedGemma 27B Text — tuned for text-only tasks such as clinical reasoning and summarisation.
How the collection fits together: imaging and text data (left) train the MedSigLIP encoder and the MedGemma models (right). Fig. 1 from Sellergren et al., Nature Medicine (2026), CC BY 4.0, unchanged.
The recipe is deliberately light-touch. Medical data made up just 2% of the training mix when Google tuned the image encoder, and about 12% when it re-trained the language model, so the models would keep Gemma 3’s general skills. Text knowledge came partly from distilling a larger teacher model, including roughly 200,000 synthetic exam-style questions generated with Gemini. All image-plus-text post-training used reinforcement learning, which the authors found generalised better than standard supervised fine-tuning.
How well does it read medical images?
This is where the results are strongest. On data the models never saw in training, MedGemma improved on its Gemma 3 base by 15.5–18.1% on chest X-ray finding classification, 2.6–10% on medical image question answering and 10.8% on agentic evaluations, the paper reports.
The size of the jump varies by specialty. In zero-shot tests — no task-specific training — MedGemma 4B scored:
- 64.9% on classifying retinal images (EyePACS), versus 14.4% for Gemma 3 4B and 27.7% for Gemini 2.5 Pro;
- 69.8% on a pathology multiple-choice test, versus 37.1% for Gemma 3 4B and 42.7% for Gemini 2.5 Pro;
- 50.1 macro F1 on the CXR14 chest X-ray set, versus 32.0 for Gemma 3 4B and 39.2 for Gemini 2.5 Pro.
Not every result goes MedGemma’s way. On a dermatology test, Gemini 2.5 Pro (81.0%) and Flash (78.4%) beat MedGemma 4B (71.8%). And the authors note that the best pathology-specific encoders still outperform MedSigLIP on several tasks.
Can it write a radiology report?
One of the most practical tests was chest X-ray report generation. A single US board-certified thoracic radiologist compared 306 MedGemma 4B reports against the original reports from the MIMIC-CXR dataset, viewing the actual images on a clinical viewer.
A radiologist’s comparison of 306 MedGemma-written chest X-ray reports with the originals. Blue: AI report better; grey: similar; red: original better; black: both missed findings. Extended Data Fig. 1 from Sellergren et al., Nature Medicine (2026), CC BY 4.0, unchanged.
The radiologist rated 68% of AI reports for normal scans and 49% for abnormal scans as equal to or better than the original. Overall, 81% were judged likely to lead to the same or better patient management — higher than the 73% reported for Google’s much larger Med-Gemini. With extra reinforcement-learning fine-tuning, MedGemma 4B reached a RadGraph F1 score of 30.3 on MIMIC-CXR, which the authors call state of the art for that metric.
Two caveats matter here. It was one reviewer, and the radiologist knew which reports were AI-generated. The chart also shows the trade-off clearly: on abnormal scans, a sizeable red slice marks cases where the AI missed key findings the original caught.
How does it compare with GPT and Gemini on text?
On MedQA, a US medical-licensing-style exam, MedGemma 27B Text scored 86.2% — the best open model at release apart from the 671-billion-parameter DeepSeek-R1 (91.0%), and well ahead of OpenBioLLM 70B (78.2%). Frontier models still lead: OpenAI’s o3 scored 93.3% and Gemini 2.5 Pro 93.9%.
The gap is widest on harder, unseen questions. On the text-only MedXpertQA benchmark, MedGemma 27B Text scored 25.7%, against 54.6% for o3. The authors are candid that “frontier models often still outperform smaller open models.”
Their argument is about efficiency, not supremacy: they cite a 500-fold difference in compute cost between MedGemma 4B and the most expensive model they compared. MedGemma also got better at spotting plausible-but-wrong answers. On the hard set of the MedHallu benchmark, MedGemma 4B scored 0.50 versus 0.09 for Gemma 3 4B. When three physicians rated open-ended answers, 15.4% of MedGemma 4B’s responses were judged inaccurate, versus 18.5% for Gemma 3 4B. That is an improvement, but still roughly one wrong answer in six.
In AgentClinic, a simulation where the model plays a doctor taking a history and making a diagnosis, MedGemma 27B beat human physicians on the MedQA-based scenarios and approached much larger models. The 4B models struggled to follow the simulation’s instructions at all.
Why “open” is the real story
The paper’s strongest case is for MedGemma as a starting point. When Google fine-tuned both MedGemma 4B and plain Gemma 3 4B on three new datasets, MedGemma came out ahead — especially when only 10% of the training data was used.
Fine-tuning MedGemma 4B (blue) versus Gemma 3 4B (orange) on 0%, 10% and 100% of the training data for pathology, skin-lesion and pneumothorax tasks. Extended Data Fig. 2 from Sellergren et al., Nature Medicine (2026), CC BY 4.0, unchanged.
On the pneumothorax (collapsed lung) task, Gemma 3 collapsed into always predicting the most common answer; MedGemma did not. For hospitals and startups with small, private datasets, that data efficiency is the point.
The authors list where an open model beats an API: a frozen model for documentation and reliability, lower cost, running locally or offline, and full control over adaptation. Google says MedGemma is free for research and commercial use on Hugging Face and Vertex AI, that every model runs on a single GPU, and that the 4B model and MedSigLIP can be adapted for mobile hardware.
Adoption is already real. Google reports “millions of downloads and hundreds of community-built variants” on Hugging Face. Malaysia’s Qmed Asia built a chat tool over more than 150 national clinical practice guidelines, and Taiwan’s National Health Insurance Administration used MedGemma to review preoperative assessments across more than 30,000 pathology reports. In January, Google shipped MedGemma 1.5, adding support for 3D CT and MRI volumes and whole pathology slides, and lifting EHR question-answering accuracy from 68% to 90%.
What are the limits?
The paper is unusually direct about them:
- Benchmarks aren’t patients. Benchmark scores are “only the first step towards validating real-world utility,” the authors write, and some older benchmarks may have leaked into training data.
- Coverage gaps. This version handles 2D images only; 3D scans and genomics came later or not at all. Anatomical localisation on chest X-rays was weak (an overlap score of 0.16 for the 27B model).
- Errors and bias. The authors flag “model errors or subtle forms of distribution bias in training data” as unresolved.
- Who did the work. All but two of the 88 authors are current or former Google employees who may own Alphabet stock, according to the paper’s competing-interests statement.
Google’s own guidance is that MedGemma is “not intended to be used without appropriate validation, adaptation and/or making meaningful modification by developers for their specific use case.”
What it means for builders
If you are building health AI, MedGemma is now a peer-reviewed, well-documented base model you can own. Benchmark it against your own data before trusting any number above, fine-tune on your task, and plan for clinical validation and regulatory review from day one. For other ways AI labs are releasing capable models to specialists first, see how Google rolled out Gemini 4 Argon to cyber defenders.
Cover image: MedGemma 4B answering questions about a chest X-ray, a skin lesion and a breast-tissue slide, with specialists’ impressions below. Fig. 2 from Sellergren et al., Nature Medicine (2026), CC BY 4.0, unchanged. Source images: NIH Clinical Center chest X-ray dataset, Google’s Skin Condition Image Network (SCIN) dataset and The Cancer Genome Atlas.
Questions readers ask
- What is MedGemma?
- MedGemma is a family of open medical AI models from Google, built on Gemma 3. It includes MedGemma 4B and 27B, which read both images and text, a text-only MedGemma 27B, and MedSigLIP, a 400-million-parameter encoder for medical images. Developers can download and fine-tune them.
- What did the Nature Medicine MedGemma study find?
- Published on October 6, 2026, the study found MedGemma beat its Gemma 3 base models on every medical benchmark tested and matched or beat much larger models on several imaging tasks. For example, MedGemma 4B scored 64.9% on retinal image classification, versus 14.4% for Gemma 3 4B and 27.7% for Gemini 2.5 Pro.
- Can MedGemma diagnose patients?
- No. Google says MedGemma is not intended to be used without appropriate validation, adaptation or meaningful modification by developers for a specific use case. The paper evaluates benchmarks and simulated tasks, not real-world patient outcomes.
- Is MedGemma free to use?
- Yes. Google says MedGemma is free for research and commercial use and is available on Hugging Face and Google Cloud's Vertex AI, under Google's Health AI Developer Foundations terms. The models can run on a single GPU, and the 4B model can be adapted for mobile hardware.
- How does MedGemma compare with GPT and Gemini?
- Frontier models still win on the hardest tests: on the out-of-distribution MedXpertQA text benchmark, MedGemma 27B Text scored 25.7% versus 54.6% for OpenAI's o3. But MedGemma 4B beat Gemini 2.5 Pro on several image classification tasks, and the authors cite a 500-fold compute-cost gap between MedGemma 4B and the most expensive model they compared.
+ Sources
- Nature Medicine — An open vision-language model for diverse medical applications (Sellergren et al., 2026)
- Google Research — MedGemma: our most capable open models for health AI development
- Google Research — Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR
- arXiv — MedGemma Technical Report
- Streamlinefeed — Nature Medicine tests Google's MedGemma for healthcare
- ReXrank — Chest X-ray report generation leaderboard
The Weekly Diff launches soon.
Follow via RSSKeep reading



