Calibrated consumer simulation

Know how customers will respond —
before you commit.

$9.99 a month? a student price? easy enough to use? important to parents?

Tlon simulates your target customers as whole populations and checks its predictions against human evidence, so you can compare decisions before committing resources. Pricing is our first application.

41% would buy at $9.99 a month

Illustrative
Audience segment

The problem

How will consumers respond — and which decision should we make?

01

4–8 weeks of fieldwork

≈10min

to a first simulated readout

Evidence takes time and budget.

Surveys and market tests need weeks of fieldwork and real budget — and the conditions being studied can change before the research is done.

02

A few options per study

10× repeats

per question and option, so no comparison rests on one answer

Few scenarios get tested.

A new product, a different offer or a price change can each move demand and revenue, yet a real study can only afford to test a few of them.

03

Predictions taken on faith

≈38respondents

what one population-level simulation was worth in our pricing backtest

Simulations need proof.

AI simulation can explore far more scenarios. To act on its results, teams need evidence that the predictions reflect how people actually respond.

The product

From a business question to repeatable simulations.

  1. 01

    Describe

    Tell the agent about your product, who it is for and the decision in front of you. Paste your site, drop a deck. It writes a neutral brief — the exact text every simulated customer reads.

  2. 02

    Choose who

    Pick and adjust the consumers to simulate — students, young adults, parents, or any group you describe. Each group is composed from official population statistics, with every assumption labeled.

  3. 03

    Simulate

    Each group is simulated as a whole population. For a price test the model estimates the whole demand curve in one answer; for perception questions, the share giving each answer. Ten independent runs each, strengthened with real answers to similar questions.

  4. 04

    Compare

    Compare predicted outcomes with honest uncertainty bands: demand and revenue across price points, a recommended test range, and how each group rates the product.

Applications

Pricing first. Perception alongside.More decisions next.

Available

Pricing

Demand and revenue at every tested price, the willingness-to-pay distribution, plan choice and a recommended test range.

Available

Perception & satisfaction

How important, useful, easy to use or satisfying customers expect your product to be, and how likely they are to try or recommend it — by group, with bands.

  • Importance44%
  • Ease of use61%
  • Satisfaction54%
Not at allSlightlyModeratelyVeryExtremely
Planned

Features & offers

Compare feature sets, bundles and offers before you build them — the same simulations, applied to the other levers of consumer choice.

  • Feature set A
  • Feature set B
  • Bundle

Illustrative

Who to simulate

Tell us who it’s for. We build the crowd.

Young adults, women, students — or any group you can describe. Each composition comes from official US population statistics. Drag the price and see how differently each group responds.

$9.99/mo

Young adults 50% · Women 68% · Students 22%

Young adultsAges 18–34

50%

would buy at $9.99/mo

  • 77.5M US adults
  • 29% of all US adults
Women18 and over

68%

would buy at $9.99/mo

  • 136M US adult women
  • 28% of them are aged 18–34
StudentsCollege & graduate

22%

would buy at $9.99/mo

  • 21.7M enrolled students
  • 56% women · 74% at public colleges
+Any group you describe

Describe it in a sentence. The agent assembles it from public statistics and labels every assumption.

For example

College students who already pay for music Working moms with young kids Retirees over 65 High-income young professionals

  • AgePublic statistic
  • SexPublic statistic
  • Student statusPublic statistic
  • Monthly app spendModel estimate
Describe a group
Composition: U.S. Census Bureau, American Community Survey 2024 (1-year)Demand curves are illustrative

Method

Simulate the crowd, not the individual.

There are two ways to simulate a market with language models. Both have their place.

Agent-level

Simulate many individual personas one by one, then count their answers.

1,000calls to simulate 1,000 people

Our approach

Population-level

Describe the whole group once and estimate its answer distribution directly.

1call per group, per question

Unit of simulation

Agent-level ·One persona per call

Population-level ·One group per call

What you get

Agent-level ·Individual answers and narratives

Population-level ·The answer-share distribution

Cost grows with

Agent-level ·Number of simulated people

Population-level ·Questions × repeats

Best at

Agent-level ·Rich stories and per-persona cross-tabs

Population-level ·How answers spread across a population

Watch for

Agent-level ·Answers can bunch on a few options

Population-level ·Subgroup differences come out subtler

We use population-level simulation for the distribution and ask for reasons separately — so you get both the curve and the why.

Calibration

Calibration at the core.

Generating consumer responses is only the first step. For a simulation to guide a business decision, its predictions have to be checked against human evidence — and their errors corrected.
  1. Stage 1 · Model development

    Learn from related surveys

    The model draws on human responses to related surveys of similar consumers and products. The backtests below measure this stage.

  2. Stage 2 · Simulation calibration

    Correct with a small reserved sample

    After a simulation, human responses to the same survey — same product, questions and test conditions — from a small reserved group are used with established statistical methods to estimate and correct prediction errors.

  3. Aim

    Fewer human responses, same accuracy

    The aim is accuracy comparable to a survey with fewer human responses than a survey alone. Independent evaluation will test whether calibration improves accuracy and how much human data it needs.

Longer term, we aim to extend calibration across products and consumer groups, so accumulated consumer data can be reused to explore more scenarios — with fewer repeated studies, and less time and cost.

Design-partner program

We are working with a small cohort of consumer AI companies selling to U.S. consumers. 80% of the research framework is shared; 20% is tailored to your product, plans and competitors.

80% · Shared framework
20%

Tailored · 20%

  1. 1Baseline
  2. 2Repeatability
  3. 3Transfer
  4. 4Fewer labels
  5. 5Held-out test

Evidence

One simulation, worth dozens to a hundred-plus real respondents.

These backtests measure the model-development stage: simulations strengthened with real answers to similar questions, compared with real survey answers. Calibration with a reserved human sample is the next milestone.

Study 1 · metric JSD

A household survey in China

In a backtest on a national household survey in China (30,000+ households), four five-point questions, a single population-level simulation landed as close to the true answer distribution as a random sample of ≈16–62 real respondents (median 48). To beat one population-level simulation with 95% confidence, a real survey needs ≈37–136 respondents (median ≈111). Agent-level simulation with the same model: ≈4.5–38 (median 5).

≈48

respondents, population-level · break-even (median)

≈111

respondents to beat population-level with 95% confidence (median)

≈5

respondents, agent-level · break-even (median)

Survey question

Distance to the true distribution vs. number of real respondents

Random real sample (mean)5–95% of random samplesPopulation-level simulationAgent-level simulationBreak-even sample sizeSample size to win with 95% confidence
0.00010.0010.010.11.01101001k10kReal respondents sampled at random (log scale)Distance to truth (JSD, log)≈ 49 respondents≈ 4.5 respondents

Answer shares: real vs. simulated

Income far above spending

Slightly above

Balanced

Slightly below

Far below

RealPopulation-level (min–max of 10 runs)Agent-level

Study 2 · metric MAE

A willingness-to-pay survey in the U.S.

In a randomized-price willingness-to-pay survey of U.S. adults (2,058 people, each asked whether they would buy at a randomly assigned price), we took 10 everyday products and had the model estimate directly the share who would buy at each of 11 price points, then compared it with the real purchase rates. Agent-level simulation was worth ≈12 real respondents and population-level simulation ≈17; after strengthening the model with real answers to similar questions, population-level simulation was worth ≈38.

≈12

respondents · agent-level

≈18 to win with 95% confidence

≈17

respondents · population-level

≈25 to win with 95% confidence

≈38

respondents · strengthened population-level

≈53 to win with 95% confidence

Purchase-rate error vs. number of real respondents

Random real sample (mean)5–95% of random samplesAgent-level simulationPopulation-level simulationStrengthened population-levelBreak-even sample sizeSample size to win with 95% confidence
0510152025101001kReal respondents sampled at random (log scale)Purchase-rate error (MAE, points)≈ 12 respondents≈ 17 respondents≈ 38 respondents

Average purchase-rate gap (points, lower is better)

Agent-level simulation17.6 pts
Population-level simulation15.5 pts
Strengthened population-level11.2 pts
Same people, asked again later2.5 pts

The last bar is a reference: when the same respondents answer again later, purchase rates shift by about this much — roughly the floor no method can be expected to beat.

Model calls in study 1

Population-level · 40 calls per group10 repeats × 4 questions
Agent-level · 138,416 callsone call per household per question

log₁₀

What it doesn’t show (yet)

In 13 of 16 subgroup cells, simply reusing the overall distribution matched the subgroup as well or better, so we report segment differences as directional. Pricing validation covers 10 everyday products so far; the model’s errors are mostly systematic and don’t average away with more repeats; and the real answers used for strengthening come from a similar survey, so the gain may shrink for very different categories.

Study 1 metric: Jensen–Shannon divergence between simulated and real answer shares. “Worth n” is the sample size at which the average of 1,000 random real samples is as close as the simulation; “95% confidence” is the size at which 95% of random samples are closer. Backtest v1. Study 2 metric: mean absolute difference between simulated and full-sample real purchase rates across 11 price points (MAE, percentage points), averaged over 10 products; “worth n” is computed as above. Agent-level is a per-respondent simulation (one digital twin per person), using the best-performing published configuration. The strengthened simulation draws on real answers to similar questions about products other than the 10 tested. Read the full method

What you get

A readout you can act on.

Sample reportAI study app — consumer simulation
3 segments · 6 price points · 3 perception questions · 10 repeatsRaw simulation — not yet calibrated with your human sampleIllustrative data

Demand curve

Share who would buy at each price, with a band that includes method error.

Test range · $7.99 – $12.99 /moHover or use ← → to read the curve

Willingness to pay

Distribution across price bands, with median and interquartile range.

p50 $9 · IQR $6–$14

Revenue index

Price × demand. The shaded band is the recommended test range.

Plan mix

Share choosing each plan — or keeping their current setup.

38%
27%
8%
27%
  • Basic $4.99
  • Plus $9.99
  • Pro $19.99
  • Keep current setup

Segments

Directional

Per-segment curves, so you can see who drives the result.

  • Young adults
  • Women
  • Students

Perception

Share giving each answer, lowest to highest; the figure is the share in the top two.

  • Importance44%
  • Ease of use61%
  • Satisfaction54%
Not at allSlightlyModeratelyVeryExtremely

Reasons

What each group says for and against buying at the test price — the “why” behind the curve.

  • Why they’d buy
  • Saves me time every week31%
  • Better than the free app I use now22%
  • Easy to use on my phone15%
  • Why they wouldn’t
  • Free alternatives are good enough28%
  • Already paying for too many subscriptions19%
  • Not sure I’d use it enough13%

Where we start

Pricing for consumer AI companies.U.S. consumers first.

We start where our current data reaches: consumer AI companies selling to U.S. consumers. As we evaluate the model on more applications, we will expand into other consumer products and services, and assess open consumer datasets from markets such as Asia and Europe.
    01 · NowPricing for consumer AI companies, U.S. consumers
    02 · NextOther consumer products and services
    03 · ThenOpen consumer datasets from Asia and Europe

Progress

  • DoneA detailed research pipeline, including the calibration design
  • DoneAn early demo of the workflow: brief, populations, pricing and perception simulations
  • UnderwayModel development

Next milestone

Implement and evaluate the calibration workflow on existing consumer data, starting with pricing: compare calibrated and uncalibrated predictions against independent consumer responses to measure how much calibration improves accuracy.

Security

Your data stays yours.

Project isolation

Every project is scoped to your account with row-level security.

No training on your data

Briefs, files and results are never used to train models.

Delete anytime

Remove a project with its files, runs and reports in one step.

Aggregates first

Populations are built from aggregate statistics, never personal records.

Security details →

Compare decisions before you commit.

Start with a conversation. Your first simulated readout is a few minutes away.