Every dataset ships with something genuinely wrong in it — and a machine-readable record of exactly what, where, and how big.
Synthetic data is easy. Synthetic data with ground truth about what's broken in it is what actually lets you score a detector, test a pipeline, or benchmark a model against a known answer — instead of eyeballing whether the output looks plausible.
"kind": "spike", "table": "orders", "column": "shipping_cost", "description": "shipping_cost in orders spikes to roughly 4.2x its normal level for South between 2025-05-25 and 2025-07-10, then returns to baseline.", "severity": "obvious", "rows_affected": 330, "window": { "start": "2025-05-25", "end": "2025-07-10", "column": "placed_at" }, "segment": { "region": "South" }, "magnitude": 4.174, "detect_hint": "Aggregate shipping_cost by week over placed_at and compare each week to the trailing median."
Every dataset is seeded and deterministic. Regenerate it from its seed and the bytes match, checksum for checksum — checked automatically, including whether the generator itself has changed since the dataset was built. MySandboxData doesn't just say "looks right." It says exactly why, if it isn't.
Eight packs. Four relational business domains, one MDM/entity-resolution pack, three flat single-table archetypes. Every row count below is exact — not estimated, not sampled down for this page.
--scale, not with seed. Different seeds vary values and anomaly placement, not table shape or volume.| pack | tables | rows @ scale 1.0 | anomalies | description |
|---|---|---|---|---|
| retail | 5 | 1,298,400 | 6 | Omnichannel retail: customers, catalog, orders, line items and returns. |
| fintech | 4 | 1,469,800 | 6 | Card payments: accounts, merchants, transactions and chargeback disputes. |
| healthcare | 5 | 981,200 | 5 | Clinical encounters, procedures, and insurance claims across patients and providers. |
| saas | 4 | 219,500 | 6 | Subscription business: accounts, plans, MRR, product usage and support load. |
| enrichment | 3 | 99,500 | 2 | Clean customer/product masters and order facts for MDM and entity-resolution work. |
| measurements_iris | 1 | 6,000 | 2 | Flat table of flower measurements and species — no joins. |
| measurements_wine | 1 | 8,000 | 2 | Flat table of wine chemistry measurements and quality class. |
| measurements_sensor | 1 | 50,000 | 3 | Flat table of IoT sensor readings and operating state. |
Seven kinds of injected anomaly. Every instance is machine-labeled: which rows, which time window, which segment, how severe, how big. That label — not the data itself — is what's for sale.
A short, sharp jump in a metric confined to a time window.
A permanent step change — the metric never comes back.
A handful of individually extreme rows — fraud, fat fingers, bad ETL.
A pipeline outage: one column goes null for a stretch of time.
One category stops appearing partway through — a churned segment.
Two columns that normally move together stop doing so.
Rows duplicated by a retried job — the classic silent double-count.
outlier_burst. Full seven-kind variety exists in the four relational business packs (retail, fintech, healthcare, saas).| spike | level_ shift | outlier_ burst | missing- ness | segment_ collapse | correlation_ break | duplicate_ run | |
|---|---|---|---|---|---|---|---|
| retail | obvious | — | 2 | subtle | obvious | — | moderate |
| fintech | moderate | moderate | obvious | subtle | — | subtle | obvious |
| healthcare | obvious | moderate | moderate | subtle | obvious | — | — |
| saas | obvious | moderate | obvious | subtle | moderate | subtle | — |
| enrichment | — | — | 2 | — | — | — | — |
| measurements_* | — | — | only kind | — | — | — | — |
"kind": "spike", "table": "orders", "column": "shipping_cost", "description": "shipping_cost in orders spikes to roughly 4.2x its normal level for South between 2025-05-25 and 2025-07-10, then returns to baseline.", "severity": "obvious", "rows_affected": 330, "window": { "start": "2025-05-25", "end": "2025-07-10", "column": "placed_at" }, "segment": { "region": "South" }, "magnitude": 4.174, "detect_hint": "Aggregate shipping_cost by week over placed_at and compare each week to the trailing median."
rows_affected, window, and segment are what make a detector's output scoreable against ground truth. detect_hint doubles as a spec for a baseline detector, if you want to sanity-check your own before running it against ours.Three real datasets, generated the same way every dataset in the catalog is generated — download and inspect the ground truth yourself before signing up for anything.
| table | rows | cols | parquet |
|---|---|---|---|
| customers | 4,000 | 9 | 148 KB |
| products | 240 | 8 | 19 KB |
| orders | 32,000 | 8 | 662 KB |
| line_items | 89,000 | 6 | 1.4 MB |
| returns | 4,646 | 6 | 127 KB |
returns intentionally contains 46 duplicated rows (the duplicate_run anomaly) — this violates the table's own primary key, so a strict load will reject it unless you drop the constraint or use the label to find the rows first.| table | rows | cols | parquet |
|---|---|---|---|
| accounts | 2,500 | 8 | 80 KB |
| merchants | 680 | 6 | 24 KB |
| transactions | 140,060 | 10 | 3.2 MB |
| disputes | 3,800 | 7 | 127 KB |
correlation_break between disputed_amount and resolution_days is the hardest of the seven kinds to detect; worth trying your own method against it.transactions intentionally contains 60 duplicated rows (duplicate_run) — same primary-key caveat as the retail sample above.outlier_burst models fat-finger entries and bad ETL writes, and an obviously-wrong value is still a value your pipeline has to catch. If your detector can't flag this one, that's worth knowing.| table | rows | cols | parquet |
|---|---|---|---|
| sensor_readings | 50,000 | 11 | 1.1 MB |
outlier_burst — see the disclosure on the anomaly catalog for which packs carry the full seven-kind range.One-time, no spam — this just unlocks the download buttons above, your files start immediately.
Submitting starts the retail, fintech and sensor downloads right away — no waiting on a reply.
Unlocked — downloads have started. The Parquet / CSV buttons above will now work directly for the rest of this visit.
Every tier ships the same thing the free samples do — real rows, real anomalies, byte-verified reproducibility. The difference is scale.
Best value if you want ongoing access — one subscription beats repurchasing the bundle every time we add a new pack.
Every purchase includes both parquet and gzipped-CSV formats. Checkout, tax handling (VAT/sales tax where applicable), and delivery are all handled by Lemon Squeezy, our merchant of record.