Synthetic healthcare data, realistic enough to build on.
I'm an actuary, and getting real data to build with was always the hard part. The good vendor files cost six figures and take months of legal review, and the public files are too thin for real work. So I built this. It generates synthetic members and their full record (eligibility, claims, pharmacy, labs, encounters, revenue), calibrated to public benchmarks so the totals come out where you'd expect. Medicare Advantage is live today, with more lines of business coming on the same engine. It's early, and mostly a solo project.
The problem
Most teams can't get good data to build with.
To build a risk model, benchmark a population, or demo a product, you need realistic healthcare data. Real claims sit behind HIPAA and data-use agreements, and the vendor files run $50k to $250k a year with months of legal review before you touch a row.
The public CMS files are free but aggregated too far to model on. So startups stall before a deal closes, and analytics teams can't test a method until they've already paid for the data.
Synthetic data fills that gap, if you trust the numbers. So every dataset ships with the benchmark comparison, and you can check it yourself.
The archetype engine
I generate people first, then their records.
I generate synthetic people and their clinical histories first, then produce every data domain from that one source. A member's claims, labs, and revenue all describe the same person, instead of being stitched together after the fact.
Generate people
A language model writes a library of patient archetypes: realistic bundles of conditions, comorbidities, and demographics, grounded in clinical knowledge and published prevalence.
Generate journeys
Each member lives 36 months: enrollment, conditions progressing, acute episodes, medications starting and stopping. A diabetic with heart failure behaves like one over time.
Render any domain
The timeline becomes the files: claims, eligibility, encounters, labs, revenue. Same person, consistent across all of them.
Calibrated to public benchmarks today. Not trained on real claims yet; that's the roadmap. See the full methodology →
Data domains
The eight data domains.
Everything joins on member_id. Take the domains you need; more are coming from the same people.
Eligibility & enrollment
Available nowMember-month enrollment, demographics, plan/program, and benefit status — the spine every other domain links to.
Medical claims
Available nowLine-level institutional + professional claims: diagnoses, procedures, settings, allowed/paid, and realistic adjustment chains.
Pharmacy
Available nowNCPDP-grade drug fills with refill chains, benefit phases, formulary tiers, and net-of-rebate economics.
Revenue & payment
Available nowPayer-side revenue the way a plan receives it — capitation, risk scores, and the factors behind every dollar.
Encounters
Available nowVisit-level utilization rolled up from the journey: office, ER, inpatient, SNF, outpatient, and lab encounters with length of stay, primary diagnosis, and DRG.
Labs & results
Available nowOrdered tests with result values trended to each member's conditions. The A1c tracks the diabetic, the eGFR tracks the CKD stage.
ADT feed
Available nowHL7-style admit, discharge, and register events derived from facility and ER encounters, with patient class, facility, and discharge disposition.
Member journeys
Available nowOne pseudo-chart per member: the archetype's true problem list, the HCCs actually coded this year, and a chart-note narrative. Ground truth paired with the observed record.
Quality measures
ExpandingMeasure-ready numerators, denominators, and gaps (HEDIS-style / Stars) rendered from each member's actual care.
Lines of business
Medicare Advantage today · the rest expandingHow I check fidelity
I check the totals against published benchmarks.
Synthetic data is only useful if the numbers match reality. Every dataset ships with a credibility audit: dozens of metrics against published benchmarks, each cited. If something is off, it's in there.
Illustrative numbers. The one metric below band (inpatient admits) is shown, not hidden. Each dataset's real audit ships with it.
Versions
Versions, like model releases.
Each version raised fidelity. v3 “Asclepius” lands at 94/100 against published benchmarks, which is solid enough to work with. The next gains come from new capabilities, not a higher score.
The AI engine: coherent comorbidity, real-HCC calibration, eight domains.
What every dataset ships as today. The highest-fidelity release, calibrated to the real CMS risk model.
Persistence, seasonality, and the social-determinants MLR fix.
The prior statistical release, for cost-sensitive work or simpler dynamics.
The first calibrated release. Solid control totals, simpler dynamics.
Schema validation, pipeline development, and the lineage starting point.
Real-claims pattern learning, SNP cohorts, provider continuity.
The next leap in fidelity: patterns learned directly from licensed real claims, not just calibrated to published aggregates.
Where we're going
Where this is headed.
Right now it's standard panels. Next, you describe the population or benchmark you need, the engine generates it to match, and the tooling helps you build on top.
Custom data on demand
Describe the cohort or benchmark you need in plain language, and the engine generates a calibrated dataset to match. Custom panels, not just the standard ones.
A tooling layer
Models, functions, and scripts to load, validate, and build on the data, so you get a starting point and not just files.
Learned from real claims
Move from calibrated-to-benchmarks toward patterns learned from licensed real claims. The honest gap today, and the next big step.
Need a custom cohort or a specific benchmark now? Tell us what you need →
Who it's for
What people use it for.
Build before the data is in hand
Stand up risk, revenue, and quality work on realistic data while your own extracts are still in the queue.
Build and demo before the deal closes
Build and demo your product on realistic data with no PHI, then swap in the customer's real data once the contract closes.
Train, validate, benchmark
A labeled, reproducible dataset for risk models, pricing, and pipelines, with ground truth you don't get from real claims.
Our vision
I spent about a decade as an actuary in and around health plans and providers. The same problem kept coming up: the people trying to build something useful couldn't get data to build with. This is my attempt at a fix. It's early, and I'd rather hear it's not useful than not hear at all.
See if it holds up.
Pull the free 5,000-member sample, all 8 domains, with the report and the lag triangles. No signup, no card. Open it, check a member's journey, and see if the numbers hold.
Free 5,000-member sample · Complete 100k panel, all 8 domains, $5,000