Test Hospital Integrations Without PHI: What We Found in Synthea Data
Nirmitee.io Engineering
Author

Synthetic patients, a safety gate, and the three interfaces a hospital integration has to pass.
The short answer. Use Synthea, the free synthetic patient generator, for your test data. Do not use it as it comes out of the box. Raw Synthea names real hospitals, fails the US Core standard check, and is cleaner than any hospital feed you will receive. Each of those gaps comes back later as a security review question, a failed conformance test or a go-live surprise. Closing them takes days, and the tools are free. We built them, measured the results, and published everything here.
Watch: Test Hospital Integrations Without PHI: What We Found in Synthea Data (6:01). The findings and the checklist in six minutes.
Who this is for
You sell software into hospitals, health systems or payers. Your product reads and writes their data: FHIR, HL7 v2, X12 claims, or all three. Your team needs realistic patients to build and test against long before any customer hands over a real record. You want data that is safe to put in a ticket, a demo or a CI pipeline, and realistic enough that go-live is not the first time your code meets a messy feed.
Synthea in one paragraph
Synthea is a free, open-source synthetic patient generator from MITRE. It simulates patients from birth to death using public health statistics and writes their full records as FHIR R4, C-CDA and CSV, with no real patient data involved. CMS and ONC use it for their own sandboxes and test kits. For installation, modules and ready-made datasets, see our complete Synthea guide. This article is about one job: using it to test hospital and payer integrations.
The real cost of testing on the wrong data
PHI in test environments widens everything
The moment a real patient record lands on a developer laptop, a staging database or a shared test file, that system holds electronic PHI. HIPAA requires a risk analysis of all ePHI a business associate holds (45 CFR 164.308) and limits its use to the minimum necessary (164.502(b)) and to what the business associate agreement permits (164.504(e)). Every copy in a test system is one more thing to defend in your customer's security review.
The stakes are not abstract. IBM's 2026 study puts the average healthcare breach at $6.64 million (HIPAA Journal summary). In 2025, 35.8% of large healthcare breaches happened at business associates, the vendors hospitals share data with (HHS OCR data via HIPAA Journal). A test environment with no real patient in it has nothing to breach.
Clean test data hides the bugs that delay go-live
Hospital feeds carry local lab codes instead of LOINC, the same patient registered twice, weights in pounds, half-filled birth dates and names typed in capitals. A product tested only on tidy samples meets these for the first time in go-live week, in front of the customer. Each one becomes an escalation, a hotfix and a slipped date. Synthea's own documentation says its data has "no missing or incorrect information", so it will not find these bugs for you unless you add them.
The same patient as your tests see it and as a hospital sends it. Every amber tag is a defect we inject in the lab described below.
Integration is already the buyer's top worry
In the Peterson Health Technology Institute's 2026 purchasing survey, 45% of health plans, employers and health systems named difficulty integrating vendor data as a top barrier to buying digital health (PHTI via HIT Consultant). A go-live that slips because of untested data confirms that worry in front of your customer.
Synthetic patient data vs other test data options
Six ways to get patient test data. For integration work, aim for the highlighted row.
De-identified production data looks like the realistic, safe choice. It is neither cheap nor guaranteed. HIPAA's Safe Harbor method removes all dates except the year, which are exactly the values interface and claims tests depend on, and shifting dates instead moves you to Expert Determination with a qualified statistician. The tools also miss things. We measured one below.
EHR vendor sandboxes are the right place to prove you can connect to Epic, Oracle Health or athenahealth. They hold a handful of shared test patients and their terms forbid real PHI, so they do not replace your own dataset.
Synthea is what the US government uses. CMS's Blue Button 2.0, BCDA, DPC and AB2D sandboxes share a synthetic population produced by MITRE (CMS BFD synthetic data guide), the SMART Health IT sandbox's R4 data is Synthea, and ONC's Inferno US Core test data was selected from a Synthea run.
What we found when we tested Synthea's FHIR output like a hospital would
We generated 50 patients, validated them against US Core, loaded them into a FHIR server, sent them through an interface engine as HL7 v2, tried to bill them as X12 claims, and used them to test a de-identification tool and a duplicate-patient check.
Every number comes from a saved script, run on 3 Oct 2026.
It names real hospitals, with real IDs and phone numbers
Synthea places its synthetic patients in real facilities taken from public CMS directories. We checked every facility ID (NPI) in its export against the national NPPES registry: 158 of 160 belong to real, registered organizations, and the names match. The FHIR files carry those facilities' real names, street addresses and 480 real phone numbers.
Real run: each facility NPI looked up in NPPES.
That is public directory data, not patient data, and it is still the wrong thing to ship. A test claim from a fake patient now names a real hospital as the billing provider. Put that in a demo, a clearinghouse test or a payer sandbox and you have attached invented care to a real organization. We hit this ourselves: a sample ID in one of our open-source Da Vinci servers turned out to belong to a real provider, and only the final pre-release review caught it.
What it means for you: replace facility details before test data leaves your team, and check every NPI against the registry. A valid check digit only proves an NPI is well formed.
It fails US Core validation, for one fixable reason
US Core is the FHIR standard US regulators and EHRs expect (our US Core implementation guide explains it). Synthea labels its data as US Core, and a label is a claim, not proof. We ran the official HL7 FHIR validator on three patients: 2,412 errors. Nearly all came from one cause. Synthea adds two fields of its own to every patient that the standard does not define, so every patient fails, and every record pointing at that patient fails too. Remove those two fields and the count drops to 0.
How two extra fields become 2,412 errors. Validator run without a terminology server, so codes were not checked against value sets.
What it means for you: if you use Synthea to prepare for conformance testing or a customer's FHIR validation, fix this first, or you will chase thousands of false alarms.
It leaves out data your workflow may depend on
Synthea fills 37 of the 49 US Core 6.1.0 record types. The other 12 never appeared in our run, including insurance coverage, orders and lab specimens. Synthea's issue tracker has had requests for orders and appointments open since 2021.
US Core profiles with no data in a 56-patient run.
What it means for you: payer APIs under CMS-0057-F need coverage records, ordering workflows need orders, and lab workflows need specimens. Budget time to add them.
It is cleaner than any real feed
Every patient name has numbers glued to it ("Dian810"), including inside all 3,147 clinical notes. Ten of 56 patients have the placeholder ZIP code 00000. Nothing is ever wrong: no local codes, no duplicates, no odd units.
What it means for you: if your tests only see clean data, your error handling is untested, and that is the code that runs on go-live day.
Healthcare test data readiness checklist
Hand this to whoever owns your test environments.
Seven checks for integration test data.
The open-source toolkit we built, and what it proved
We turned the checklist into a free, open-source toolkit on GitHub. One command generates patients, cleans them, runs the safety gate, adds realistic defects with a record of each, produces HL7 v2 messages, loads a FHIR server and checks the server against that record. On 50 patients it runs in about 90 seconds.
The pipeline. Every decision runs on each build, and nothing is shared unless it passes.
- Real-world identifiers: 480 real phone numbers before, 0 after. The gate blocks the raw data and passes the cleaned data.
- Standard check: 2,412 validator errors before, 0 after.
- Realism: 994 local lab codes and 68 weights in pounds injected, and the FHIR server held exactly what the record said.
- Server behavior: the HAPI FHIR server rejected a record pointing at a missing patient, but stored a timestamp with no time zone, which FHIR R4 does not allow. If your product trusts the server to catch bad data, test that assumption.
Server counts checked against the defect manifest.
Use cases for synthetic patient data in integration testing
Before your first hospital go-live
Rehearse go-live on data that misbehaves the way the hospital's will. Inject the defects their feed is known to carry and prove your product handles each one before anyone is watching.
During the hospital's security review
"Do you use PHI in development or test?" gets a short answer: no, and here is the automated check that keeps it that way.
Sample HL7 messages and interface engine testing
Synthea does not produce HL7 v2; its maintainer answered a request for it with one word, "No" (issue #1011). Our toolkit writes ADT^A04 registrations and ORU^R01 lab results from the same patients, so you get sample HL7 messages that match your FHIR test data person for person.
We sent all 1,064 messages (56 registrations, 1,008 lab results) over MLLP into Open Integration Engine 4.5.2, the open-source fork of Mirth Connect, on a strict channel and a lenient one. The engine accepted every message, and its transformers saw every injected defect that reaches v2: 929 local lab codes, 6 partial birth dates, 6 record numbers without an assigning authority and 3 padded upper-case names.
Then the finding that matters for go-live. We also sent control messages with a malformed date and a text value in a numeric lab field. The strict channel acknowledged both as accepted (AA). It only rejected an unknown message type. An interface engine acknowledgement means "parsed", not "valid". The HAPI validator, the library inside the engine called directly, caught both, and it also caught two bugs in our own generator: order numbers longer than the 22 characters v2.5.1 allows, and a value-type code that was too long. We fixed both, and the validator now reports 0 issues on all 1,064 messages.
What it means for you: do not treat an AA acknowledgement as proof your messages are right. Add a validation step to your channels, and test them with data that is wrong on purpose. Our Mirth Connect Docker guide gets a local engine running in minutes.
Synthea claims and X12 837 testing
Synthea writes a claim and an explanation of benefit for every encounter, so it looks like a ready source of 837 test claims. It is not, yet. Across 4,379 claims in our run:
- Every diagnosis is SNOMED CT. An 837 requires ICD-10-CM. Synthea can map to ICD-10-CM, but only from a mapping file you supply; the release ships one narrow example.
- No procedure carries a CPT or HCPCS code, the codes payers price claims on. Services are SNOMED, LOINC, RxNorm and CVX.
- Only 1,007 of 2,922 professional claims have any diagnosis, and an 837P cannot be sent without one.
- No member ID, no payer ID, and no NPI on the billing provider. Synthea's CSV and CPCDS exports have the same gaps, and in a separate run 112 of 114 facility NPIs in the CSV export were again real organizations.
We built a five-claim 837P from Synthea data with an explicit crosswalk (SNOMED to ICD-10-CM, office visits to CPT 99213) and placeholder member and payer IDs. It passes the pyx12 X12 validator. pyx12 also passed a version with a raw SNOMED code in the diagnosis field, so a syntax check proves the file is well formed, not that a payer will accept it. It did catch a claim with no diagnosis.
What it means for you: Synthea gives you realistic encounters to bill, not billable claims. For clearinghouse onboarding you need a reviewed code crosswalk, valid test IDs, and the clearinghouse's own test mode. See our guide to EDI testing and clearinghouse integration.
Payer APIs under CMS-0057-F
Patient Access, Provider Access and prior authorization APIs need coverage and claims data. Synthea supplies claims and explanation-of-benefit records but no Coverage, so plan to add it. Our CMS interoperability rules guide covers what each API needs. CMS's own synthetic beneficiaries include a "golden" test member and a deliberate ID-collision case, a good pattern to copy.
Patient matching (EMPI) testing with synthetic duplicates
Nobody can tell you which real records are the same person, so matching logic is hard to test. With injected duplicates you have the answer key. A simple exact-match rule found 7 of 126 injected duplicates. A rule that allowed a typo and a partial birth date and checked the address found 117.
Recall of three matching rules on 126 injected duplicates among 715 records.
Synthea people are generated independently: no twins, no Jr and Sr, no relatives at one address. Those cause wrong merges in real registries, so add them before you trust a "no false matches" result. Synthea's Fixed Records and split-record features help here, and our guide to patient matching beyond demographics covers the matching side.
Microsoft Presidio on clinical notes: a de-identification benchmark
Because Synthea knows every patient detail behind each note, its notes make a free benchmark. We ran Microsoft Presidio, a widely used open-source de-identification tool, on 3,147 notes. It flagged every date and every age, and left the patient's first name in 1,576 notes. Only half the notes came out fully cleaned. Synthea notes are short and templated, so real notes will score differently. The lesson holds: measure a de-identification tool on your own notes before you trust it with production data. Our de-identification guide covers Safe Harbor and Expert Determination.
Presidio 2.2.364 with spaCy en_core_web_lg, measured 3 Oct 2026.
Sales demos
A demo full of believable patients, with no real hospital named on screen.
Resilience and failure testing
Injected defects pair well with deliberate outages. See our chaos engineering guide for FHIR and Mirth.
The bottom line
Synthea answers the question every integration team asks first: where do we get patient data without touching PHI? It is free, it is what CMS and ONC use, and it produces coherent records in minutes. Out of the box it is not test-ready. It names real hospitals, fails US Core validation until two fields are removed, skips record types that payer and order workflows need, and never sends the messy data hospitals do.
Each of those gaps is cheap to close before go-live and expensive to discover during it. The teams that go live smoothly treat test data as part of the integration, with the same care as the interface itself: synthetic patients, a safety gate, the exact standard version their customers use, and the defects their customers' feeds really carry.
Going live with a hospital or payer soon? Talk to an integration engineer.
Nirmitee builds healthcare integrations for companies that sell into hospitals and payers: FHIR and US Core APIs, HL7 v2 through Mirth Connect and Open Integration Engine, X12 claims and clearinghouse connections, EHR integrations, and CMS-0057-F payer APIs. We publish our work in the open: our Da Vinci CRD, DTR and PAS servers are tested against ONC's Inferno suites, with CRD and PAS passing at zero failures, and the toolkit in this article is open source.
Book a 30-minute test data review. Bring your next go-live: the customer, the formats and the feeds you expect. You leave with:
- the defects that kind of feed typically carries, and which ones your tests do not cover yet
- a gap check of your test data against the seven-point checklist above
- a plan for a gated, PHI-free dataset and a go-live rehearsal, sized in days
How synthetic patient data came about
Healthcare test data moved in three steps. First, rules for stripping identifiers from real records: HIPAA's Privacy Rule (2000, modified 2002) and HHS's de-identification guidance (2012). Second, perturbed real data released to the public: CMS's DE-SynPUF in 2013, synthetic Medicare claims for about 2.33 million beneficiaries. Third, patients simulated from public statistics: Synthea, started at MITRE in 2016, described in JAMIA in 2017 (Walonoski et al.), and now the source of CMS's sandbox beneficiaries.
Sources: HHS OCR, NORC, Synthea release notes, JAMIA, ASPE, CMS.
For whoever builds this
Everything above, as commands. Measured 3 Oct 2026 with the Synthea master-branch-latest build, seed 42, Massachusetts, unless stated.
Run Synthea in five minutes, no Java install
curl -L -o synthea-with-dependencies.jar \ https://github.com/synthetichealth/synthea/releases/download/master-branch-latest/synthea-with-dependencies.jar docker run --rm -v "$PWD":/w -w /w eclipse-temurin:17-jre \ java -jar synthea-with-dependencies.jar -s 42 -p 50 \ --exporter.baseDirectory=/w/output Massachusetts
A real Synthea run in Docker.
-psets how many living patients you get; Synthea also writes those who died during simulation, so 50 gave 56 records.-sfixes the seed. Pin the jar version too.- Load the hospitals and practitioners bundles before patient bundles; patient bundles reference them conditionally. Or set
exporter.fhir.transaction_bundle=false.
Settings that matter
| You want | Setting | Measured (200 requested, 228 records) |
|---|---|---|
| Smaller, faster | --exporter.years_of_history=2 | 206 MB, 19 s |
| Default | --exporter.years_of_history=10 | 756 MB, 27 s |
| Full lifetime | --exporter.years_of_history=0 | 2,058 MB, 57 s |
| Bulk loading | --exporter.fhir.bulk_data=true | 502 MB in 21 NDJSON files |
| One condition only | -k keep_diabetes.json | 22 of 22 diabetic; about 20 times slower per patient |
About 3.3 MB per patient at default settings; plan disk before a load test.
The toolkit
Clone it from github.com/Nirmitee-tech/healthcare-test-data-lab. Every number in this article is regenerated by a script in that repository.
docker run -d --name synthlab-hapi -p 127.0.0.1:8090:8080 hapiproject/hapi:latest python3 -m lab.pipeline --patients 50 --seed 42 --rate 0.1 \ --out out/run --fhir-base http://localhost:8090/fhir
The safety gate fails the raw data and passes the scrubbed data. It checks identifiers only; it does not detect real names or pasted notes.
HL7 validator_cli, US Core 6.1.0, before and after.
Every injected defect is written to a manifest your tests can check.
Sample HL7 v2 messages from a synthetic patient. The L25718 code is an injected local lab code.
A synthetic patient on HAPI FHIR 8.12.0.
Send the HL7 v2 messages through an interface engine
docker run -d --name oie -m 2g -p 8443:8443 -p 6661:6661 openintegrationengine/engine:latest # Import a channel: TCP Listener in MLLP mode on 6661, HL7v2 inbound, strict parser on. # Then stream every message with MLLP framing (0x0B ... 0x1C 0x0D) and record each ACK code.
Our harness builds the channel from the engine's own default objects and imports it over the REST API, because hand-written channel XML fails silently when a field is missing. Validate separately with the HAPI v2 validator; the engine's ACK alone is not validation.
Build and check an 837P from Synthea claims
- Map diagnosis codes from SNOMED CT to ICD-10-CM with a reviewed crosswalk, or supply
exporter.code_map.icd10-cmto Synthea. - Map services to CPT or HCPCS, and drop or fix claims with no diagnosis.
- Add test member IDs, payer IDs and a billing NPI that are checked against NPPES and known to be unissued.
- Validate syntax with pyx12, then send through your clearinghouse's test mode for payer edits.
Ready to scale?
Talk to our healthcare engineering team about building, integrating, and shipping faster.
Frequently Asked Questions
How do you test hospital integrations without PHI?
Can we use de-identified production data for integration testing instead?
Is Synthea data PHI?
Is Synthea good enough for US Core or certification testing?
Can Synthea produce HL7 v2 or X12 claims?
Does the same seed always give the same patients?
Can I train AI models on Synthea data?
Does Synthea work outside the US?


