
Simulated worlds. Scientific questions.
Meridia
A world to investigate.
Evidence to question.
Test how AI systems work with incomplete evidence, against a world whose history you can inspect.
Enter the worldSeed 4711 · Initial landscape
Display relief exaggerated · Building glyphs illustrative
01 / The observatory
Examine the world.
A recorded country, its residents and the methods used to study it. Explore the evidence at each scale. No live agent runs on this page.
A country, in relief.
Explore the terrain and its settlements.
Relative elevation
Drag to orbit · Scroll to zoom · Click a county
Inspect a place
Each observation belongs to a person and a date. Move the clock to see what was available then.
Available observations
Same country. Same observed inputs.
Inspect the comparison.
Two statistical methods ran on the same packet. Lower estimation error is better. These are recorded results, not frontier-agent scores.
| Method | Population errorMean absolute percentage error | Adult income errorMean absolute percentage error | Method & checksRecorded seconds |
|---|
Month 6 revised estimates, averaged equally across 16 counties. Adult income is mean income among people aged 16 and over. Each absolute error is divided by the larger of the true value’s magnitude and 1, then expressed as a percentage.
Revised population estimate against retained development truth.
Inspect a recorded evidence task
The task asks for one resident’s county and employment at month 8. Each answer must cite the latest relevant record.
The reference replays the records. The stale control keeps the initial facts. These are diagnostic programs, not language models.
Download the task packetFollow the result back to its source.
Engine version, saved output hashes and export provenance.
Recorded development example · Public demo data includes selected truth · No live model calls
02 / A measured experiment
Same worlds.
Different missing evidence.
Income can influence who answers a survey and whose income is left blank. Compare four recorded settings across twelve worlds, with the same underlying truth in every paired comparison.
Separate study · Worlds 74001–74012 · Bayesian county adult-income estimates
Loading the recorded experiment…
Choose a recorded setting. These controls do not run a new simulation. Other observation errors remain.
- Both on · reference
- Selected setting
Mean signed error · Negative means underestimated
Selected minus reference
Mean county signed error per world. The axis stays fixed when you change a setting. Zero means no net signed error.
Inspect all twelve pairs
| World | Both on | Selected | Difference |
|---|
What this comparison establishes
Each paired arm shares retained truth, source records, sampled households and keyed random draws. When an income dependence is switched off, its corresponding mean probability is matched to the reference; realized rates can still differ.
Signed error is 100 × (estimate − truth) / max(|truth|, 1), averaged equally over finite county targets and then equally over twelve worlds. Mean absolute percentage error takes the absolute value before averaging. The estimator and its settings stay fixed.
This diagnoses an observation mechanism in twelve development worlds, with one survey draw per arm. It is not a fitted correction, a confirmation on fresh worlds or an agent benchmark. The keyed instrument differs from the legacy version, so this view compares only keyed arms.
A claim you can inspect.
All four settings, exact world-level values and source hashes.
03 / The research process
Separate what happened
from what was observed.
Meridia generates events and imperfect records of those events. A task specifies what a method can see. Retained truth lets researchers inspect its conclusions.
- 01
Generate a world
People, households, institutions and a dated event history.
- 02
Define the evidence
Declare the observation rules, missing information and available records.
- 03
Run a method
Use a specified packet and output format. Preserve its result and runtime.
- 04
Inspect the result
Compare estimates and supporting evidence with the recorded world.
Scope of this demonstration
A model, with limits.
This demonstration contains a simplified population and institutional system. Its history is replayed; actions do not change the future. Disease transmission, a coupled climate system and live agent execution are not part of this example.
Current research focuses on observation mechanisms, uncertainty and reproducible agent evaluation. Performance in a generated world alone does not establish performance in a real country.
04 / Research with Meridia
Bring a question
worth testing.
We’re developing Meridia for teams studying how agents work with evidence and uncertainty. Start with one workflow and a result your team can inspect.
Discuss a research pilot