Synthetic healthcare data · Medicare Advantage available now

Synthetic healthcare data, realistic enough to build on.

I'm an actuary, and getting real data to build with was always the hard part. The good vendor files cost six figures and take months of legal review, and the public files are too thin for real work. So I built this. It generates synthetic members and their full record (eligibility, claims, pharmacy, labs, encounters, revenue), calibrated to public benchmarks so the totals come out where you'd expect. Medicare Advantage is live today, with more lines of business coming on the same engine. It's early, and mostly a solo project.

1
Synthetic population
claims, labs, revenue agree
8
Data domains
incl. labs · ADT · journeys
MA
Available now
more lines of business coming
94/100
v3 fidelity score
vs. published benchmarks

The problem

Most teams can't get good data to build with.

To build a risk model, benchmark a population, or demo a product, you need realistic healthcare data. Real claims sit behind HIPAA and data-use agreements, and the vendor files run $50k to $250k a year with months of legal review before you touch a row.

The public CMS files are free but aggregated too far to model on. So startups stall before a deal closes, and analytics teams can't test a method until they've already paid for the data.

Synthetic data fills that gap, if you trust the numbers. So every dataset ships with the benchmark comparison, and you can check it yourself.

The archetype engine

I generate people first, then their records.

I generate synthetic people and their clinical histories first, then produce every data domain from that one source. A member's claims, labs, and revenue all describe the same person, instead of being stitched together after the fact.

1

Generate people

A language model writes a library of patient archetypes: realistic bundles of conditions, comorbidities, and demographics, grounded in clinical knowledge and published prevalence.

2

Generate journeys

Each member lives 36 months: enrollment, conditions progressing, acute episodes, medications starting and stopping. A diabetic with heart failure behaves like one over time.

3

Render any domain

The timeline becomes the files: claims, eligibility, encounters, labs, revenue. Same person, consistent across all of them.

Calibrated to public benchmarks today. Not trained on real claims yet; that's the roadmap. See the full methodology →

Data domains

The eight data domains.

Everything joins on member_id. Take the domains you need; more are coming from the same people.

Full product details →

Eligibility & enrollment

Available now

Member-month enrollment, demographics, plan/program, and benefit status — the spine every other domain links to.

Medical claims

Available now

Line-level institutional + professional claims: diagnoses, procedures, settings, allowed/paid, and realistic adjustment chains.

Pharmacy

Available now

NCPDP-grade drug fills with refill chains, benefit phases, formulary tiers, and net-of-rebate economics.

Revenue & payment

Available now

Payer-side revenue the way a plan receives it — capitation, risk scores, and the factors behind every dollar.

Encounters

Available now

Visit-level utilization rolled up from the journey: office, ER, inpatient, SNF, outpatient, and lab encounters with length of stay, primary diagnosis, and DRG.

Labs & results

Available now

Ordered tests with result values trended to each member's conditions. The A1c tracks the diabetic, the eGFR tracks the CKD stage.

ADT feed

Available now

HL7-style admit, discharge, and register events derived from facility and ER encounters, with patient class, facility, and discharge disposition.

Member journeys

Available now

One pseudo-chart per member: the archetype's true problem list, the HCCs actually coded this year, and a chart-note narrative. Ground truth paired with the observed record.

Quality measures

Expanding

Measure-ready numerators, denominators, and gaps (HEDIS-style / Stars) rendered from each member's actual care.

Lines of business

Medicare Advantage today · the rest expanding
Medicare AdvantageliveMedicare FFSsoonCommercial / employersoonMedicaidsoonACA / exchangesoon

How I check fidelity

I check the totals against published benchmarks.

Synthetic data is only useful if the numbers match reality. Every dataset ships with a credibility audit: dozens of metrics against published benchmarks, each cited. If something is off, it's in there.

Sample credibility auditv3 “Asclepius” · MA
MetricOursBenchmarkStatus
Avg risk score (V24/V28)1.000.90–1.10in band
Medical loss ratio86%85–92%in band
PMPM medical$983$850–1,150in band
Cost concentration (top 5%)52%~50%in band
Actuarial value (paid/allowed)82%80–85%in band
ER visits / 1,000571550–650in band
IP admits / 1,000236250–300below

Illustrative numbers. The one metric below band (inpatient admits) is shown, not hidden. Each dataset's real audit ships with it.

Versions

Versions, like model releases.

Each version raised fidelity. v3 “Asclepius” lands at 94/100 against published benchmarks, which is solid enough to work with. The next gains come from new capabilities, not a higher score.

v3 · AsclepiusCurrent

The AI engine: coherent comorbidity, real-HCC calibration, eight domains.

94
fidelity / 100
8
data domains

What every dataset ships as today. The highest-fidelity release, calibrated to the real CMS risk model.

v2 · HippocratesAvailable

Persistence, seasonality, and the social-determinants MLR fix.

88
fidelity / 100
4
data domains

The prior statistical release, for cost-sensitive work or simpler dynamics.

v1 · GalenAvailable

The first calibrated release. Solid control totals, simpler dynamics.

79
fidelity / 100
4
data domains

Schema validation, pipeline development, and the lineage starting point.

v4 · PanaceaRoadmap

Real-claims pattern learning, SNP cohorts, provider continuity.

The next leap in fidelity: patterns learned directly from licensed real claims, not just calibrated to published aggregates.

Compare versions in detail →

Where we're going

Where this is headed.

Right now it's standard panels. Next, you describe the population or benchmark you need, the engine generates it to match, and the tooling helps you build on top.

Next

Custom data on demand

Describe the cohort or benchmark you need in plain language, and the engine generates a calibrated dataset to match. Custom panels, not just the standard ones.

Building

A tooling layer

Models, functions, and scripts to load, validate, and build on the data, so you get a starting point and not just files.

Roadmap

Learned from real claims

Move from calibrated-to-benchmarks toward patterns learned from licensed real claims. The honest gap today, and the next big step.

Need a custom cohort or a specific benchmark now? Tell us what you need →

Who it's for

What people use it for.

Payers & risk-bearing teams

Build before the data is in hand

Stand up risk, revenue, and quality work on realistic data while your own extracts are still in the queue.

Health tech & digital health

Build and demo before the deal closes

Build and demo your product on realistic data with no PHI, then swap in the customer's real data once the contract closes.

Data science & actuarial

Train, validate, benchmark

A labeled, reproducible dataset for risk models, pricing, and pipelines, with ground truth you don't get from real claims.

See detailed use cases →

Our vision

I spent about a decade as an actuary in and around health plans and providers. The same problem kept coming up: the people trying to build something useful couldn't get data to build with. This is my attempt at a fix. It's early, and I'd rather hear it's not useful than not hear at all.

See if it holds up.

Pull the free 5,000-member sample, all 8 domains, with the report and the lag triangles. No signup, no card. Open it, check a member's journey, and see if the numbers hold.

Free 5,000-member sample · Complete 100k panel, all 8 domains, $5,000