MedGemma vs Parrotlet on 288 Indian Lab Reports and Prescriptions: We Ran All Three Models
CTO & Co-Founder
CTO & Co-Founder at Nirmitee.io. Architects healthcare integrations across FHIR, SMART on FHIR, ABDM and NHCX, writing from production experience taking hospital software from sandbox to go-live.

The short answer. On Eka Care's public test set of 288 Indian medical documents, Eka's Parrotlet-v-lite read the most facts correctly, at about 63 to 67 percent of the checks. Google's MedGemma 1.5 managed 40 to 45 percent, and the older MedGemma 4B managed 31 to 37 percent. Parrotlet was also the fastest, at 1.3 seconds a page on one mid-range GPU. All three models are small enough to run on your own server, so patient documents never have to leave your network. If you are planning this for your own documents, see how we build private healthcare AI solutions.
None of the three is ready to work unsupervised. Even the best one missed about a third of the facts a reviewer would check, and prescriptions were hard for all of them.
| Model | Answer the software could read | Facts correct, strict name match | Facts correct, loose name match | Seconds per page | GPU cost per 1,000 pages (estimate) |
|---|---|---|---|---|---|
| Parrotlet-v-lite 4B (Eka Care) | 99.3% | 62.6% | 67.3% | 1.3 | about $0.30 |
| MedGemma 1.5 4B (Google) | 93.8% | 40.0% | 45.0% | 4.0 | about $0.94 |
| MedGemma 4B (Google) | 92.4% | 31.0% | 37.3% | 2.2 | about $0.53 |
Same GPU (one NVIDIA L4), same prompts, same settings for all three. The cost column is an estimate: measured seconds per page multiplied by Google Cloud's listed on-demand price for one L4 machine, about $0.85 an hour in US regions. It leaves out start-up time, storage and staff.
What we tested, and why a hospital should care
Most clinics and labs in India still send and receive reports as PDFs and phone photos. Turning a page into structured data, meaning each test name, value, unit and reference range in its own field, is the first step for anything useful: trend charts, alerts, insurance claims, or sending records to ABDM in FHIR format. The extraction model is one piece; the rest is laboratory integration and ABDM integration.
We used the evaluation set Eka Care published with its model: 218 lab reports and 70 prescriptions, each with Eka's own extraction prompt and a list of checks written from expert annotations. A check reads like "Haemoglobin test contains the value 12.4" or "Frequency of QTRIPIL SR 400 MG matches 0-0-1". Across the 288 pages there are 17,276 such checks.
We gave every model the same page image and Eka's own prompt for that page. We did not tune anything for any model.
How we scored, and why we did not use an AI judge
Eka scored its models by asking GPT-4o whether each check was met. We wanted a score that anyone can rerun for free and get the same answer, so we wrote fixed rules instead. The checks follow about 30 sentence patterns, so a short program can read each one and compare it with the model's answer. 17,273 of the 17,276 checks fit a pattern; the other 3 are malformed in the dataset and are left out.
The rules ignore case and punctuation, treat "5.60" and "5.6" as the same number, and treat a blank and "None" as the same, as Eka's own instructions say. Names were harder. Should "Pus Cells" match "Pus Cells, /HPF"? A person would say yes. A strict rule says no. So we report two scores: strict, where names must match exactly after cleaning, and loose, where one name may contain the other. The truth sits between them, because the loose rule can also give credit for the wrong test, such as "Haemoglobin" inside "Mean corpuscular haemoglobin".
Our numbers are therefore not Eka's numbers. A flexible AI judge will usually score higher than fixed rules. Use our table to compare the three models with each other, not with figures in Eka's chart.
Did our scoring rules get it right?
We drew 40 checks at random, 20 our rules passed and 20 they failed, and compared each verdict with a reading of the model's actual answer. 37 of 40 agreed. In all three disagreements our rules were too strict: a correct value was failed because the test or drug name differed slightly, such as "OMNACORTIL 10 MG TAB" in the answer against "OMNACORTIL 10 G TAB" in Eka's check. So the strict scores in our table are a floor, not a ceiling. This review was done by our analysis assistant reading the outputs, not by a clinician.
How much can you trust a 2-point difference?
Not much. We ran MedGemma 1.5 twice with one setting changed, the number of pages the GPU processes at once. Nothing about the model or prompt changed. The score moved from 41.9 to 40.0 percent. So any gap under about 2 points is noise. The gaps in our table are 9 to 23 points, well above that.
What we found
Parrotlet's specialisation shows. It returned usable JSON on 99.3 percent of pages, never hit the length limit on prescriptions, and wrote short answers, about 700 tokens a page. One caution: Eka built Parrotlet, the test set and the checks, so its answers naturally match the style the checks expect. Eka's card does not say whether any test pages or near-copies were used in training. A test on your own documents is the only way to remove that home advantage.
MedGemma 1.5 is better than MedGemma 4B, but slower. The newer model gained about 9 points. It also writes out its reasoning before every answer, which made its answers nearly twice as long as the older model's and its pages almost twice as slow. Software that expects plain JSON has to strip that reasoning first. If it does not, every MedGemma 1.5 answer fails: 0 of 288 pages were plain JSON without cleanup.
Prescriptions are the weak spot. On prescriptions, Parrotlet got 49 percent of checks right, MedGemma 1.5 got 30 percent and MedGemma 4B got 21 percent (strict matching). Handwriting, abbreviations like "1-0-1" and brand names that hide the generic drug all hurt.
About 5 to 8 percent of MedGemma pages ran out of room. They reached the 6,000-token output limit before finishing, so the page produced nothing usable. Parrotlet hit the limit on 0.7 percent of pages.
What broke on the way, so you do not repeat it
- The cheapest Colab GPU gave blank answers. On a T4 with the model loaded in float16, one page ran for 545 seconds and returned nothing. Gemma-family models are known to misbehave in float16; we did not run a controlled test to prove that was the cause. Moving to an L4 with bfloat16 fixed it.
- Running one page at a time was too slow. The standard Hugging Face code did about 11.5 tokens a second on the L4, which works out to roughly 18 hours per model for 288 pages. Switching to vLLM, an inference server that processes many pages together, brought that to between 6 and 20 minutes per model.
- Leftover processes ate the GPU. Two runs failed with out-of-memory errors because earlier test runs had left background processes holding 22 GB. Check GPU memory before every batch.
What this means if you are choosing a model
- For Indian lab reports today, start with Parrotlet-v-lite and test it on your own pages. On this test it was more accurate, faster and cheaper per page than either MedGemma model.
- Plan for human review. At 63 to 67 percent of checks, every extracted page still needs a person to confirm values before they reach a patient record or a claim.
- Treat prescriptions as a separate project. No model here was close to usable on them without review.
- Budget for engineering, not GPUs. The GPU cost was about $1 or less per 1,000 pages for every model (estimate). The work is in cleaning outputs, checking units and dates, mapping to FHIR and LOINC, and building the review screen.
- Keep it on your own hardware if data rules require it. All three models run on one 24 GB GPU, so documents can stay inside your network. Self-hosting does not by itself make a system HIPAA-compliant or ABDM-approved; logs, backups and access still need review.
Can a better prompt close the gap?
Barely. We reran MedGemma 1.5 with a stricter prompt, a few made-up example answers (none from the test set) and one automatic retry for any page whose answer did not fit the expected format. 39 of the 288 pages needed the retry.
| MedGemma 1.5 | Usable answers | Facts correct, strict | Facts correct, loose | Seconds per page |
|---|---|---|---|---|
| Eka's prompt, as shipped | 93.8% | 40.0% | 45.0% | 4.0 |
| Stricter prompt, examples, one retry | 95.8% | 41.2% | 46.3% | 4.8 |
Overall the gain was about one point, which is inside the noise. Lab reports did not improve at all. Prescriptions went from 30 to 38 percent, and every prescription answer became usable, but the prescription score also moved by 5 points between our two baseline runs that differed only in batch size, so we would not bank on that gain without a larger test. The retries cost about 20 percent more GPU time: 4.8 seconds a page instead of 4.0, or about $1.14 per 1,000 pages instead of $0.94 (estimate).
Our reading: prompting fixes format problems, not reading problems. The gap between MedGemma 1.5 and Parrotlet is about what the model has learned from Indian documents, and closing it would take fine-tuning on licensed, labelled pages, which we have not done here.
Limits
- This is one public test set from one company. Your documents will differ.
- Our scores come from fixed rules, so they are not comparable with Eka's GPT-4o scores.
- The checks were written by Eka from expert annotations. We did not re-annotate the pages ourselves.
- None of these models is a medical device. A clinician or trained reviewer must check every output.
How Nirmitee can help
We run this exact comparison on your own documents in about two weeks: your lab formats, your prescriptions, your languages. You get the same table for your data, the failure cases, and a plan for review and FHIR/ABDM integration. Book a model evaluation on your documents. For background on other open medical models, see our healthcare LLM landscape guide.
Need to pick or deploy a medical document model for your own data? Explore our healthcare AI solutions and healthcare interoperability work, or talk to our team about a two-week evaluation on your documents.
Method details for technical readers
Dataset. ekacare/medical_records_parsing_validation_set, split test, revision dbfce4dd6ed5593e7804619ab9f3b2178504bfe9, 288 rows (218 Lab-report, 70 Prescription). Counting unit: page = dataset row; rubric = one criterion on one page.
Runs (8 Oct 2026, UTC, NVIDIA L4, vLLM 0.31.0, bfloat16, no quantization, temperature 0, max_tokens 6000, max_num_seqs 64, prompt = row sample_prompt, system "You are an expert medical professional.")
| Model | Run ID |
|---|---|
| google/medgemma-1.5-4b-it | 20261008T133613935417Z-baseline-vllm |
ekacare/parrotlet-v-lite-4b (inner Gemma weights in subfolder model, wrapper code not executed) | 20261008T140259005582Z-baseline-vllm |
| google/medgemma-4b-it | 20261008T141250972060Z-baseline-vllm |
| Repeatability pair: MedGemma 1.5, default batch | 20261008T125839689806Z-baseline-vllm |
Metrics. "Answer the software could read" = json_parsed_pct after stripping <unused94>...<unused95> thinking text and one markdown fence. Facts correct = rubric_pass_pct_pooled (strict) and rubric_pass_pct_pooled_contains (loose): passed rubrics / scorable rubrics, pooled over pages. Seconds per page = vLLM generation wall time / pages, batched, excluding model load; not single-request latency.
Cost. Estimate = seconds per page x 1,000 / 3,600 x $0.85 per hour (Google Cloud g2-standard-8 on demand, US regions, $0.8536 in us-central1 per gcloud-compute.com, data updated 4 Oct 2026, consistent with the figure shown on Google Cloud's accelerator pricing page on 8 Oct 2026).
Prompting run: 20261008T142726340207Z-prompting-vllm (stricter instructions, synthetic examples written for this test, one retry when the answer fails Eka's schema; 39 pages retried).
Ready to scale?
Talk to our healthcare engineering team about building, integrating, and shipping faster.
Frequently Asked Questions
Which is more accurate for Indian lab reports, MedGemma or Parrotlet?
How fast is MedGemma on a single GPU?
Does a better prompt make MedGemma as good as Parrotlet?
Can these models run on our own servers for HIPAA or ABDM workloads?
Are these models safe to use without human review?


