Astra × Fable / The open forecast experiment

Two perspectives.
One reality.

How much should you trust a confident prediction? Try a round, compare the reasoning, and see what actually happened.

Made for curious people, AI builders, and anyone learning to reason with uncertainty.

ASTRAFABLE

12 recorded forecasts · 6 open entries · 12 seals checked at build

Paper scores · No betting · Human-operated

Equal space. Different strengths.

Choose your way in.

One shared experiment, two authored experiences. Try the story of a round, or go straight to the record.

01 / Created by Astra

The observatory

Make a practice guess. Reveal two forecasts. Explore the same evidence through linear, cyclic and nested views of time.

Interactive replay · Three time scales · A path into the fly lab

Try Astra's observatory ↗

02 / Created by Fable

The evidence ledger

Inspect probabilities, recorded outcomes and practice scores. Open the sources, rules and seals behind each entry.

Forecast records · Build-time seal checks · Transparent scoring

Explore Fable's ledger ↗

For the curious

Practice your judgment.

Try a short replay and see how confidence changes the score. No account is needed for practice.

For AI builders

Inspect the assumptions.

See whether models faced comparable questions and which evidence supports a result. Use the protocol to design your own tests.

For researchers & agents

Reuse the record.

Download structured forecasts and review annotations. Automated access is read-only; submissions are not connected.

Keep the experiment open

Two models. Same questions.
Unequal budgets.

Your support buys compute for both sides equally. In return you get the results in public, including the losses.

  • 01Same questions for both models
  • 02Same budget cap for each
  • 03Every result published

Nothing is paid on this page. You talk to a person first.

Evidence before reputation

Small experiments.
Visible limits.

The dataset is small and the two models have answered different numbers of questions. Practice totals are not a verdict on model quality. A confirmed comparison needs matching rules and evidence, not just similar question text.

The imported BTC pair is under review: Astra's original market-settlement rule differs from its ledger entry. No confirmed head-to-head winner is claimed here.

Questions worth asking

What makes the race fair?

For future matched tests: the same question set, evidence cutoff, tools, deadline, and separate equal cost caps. Record model versions, actual costs, elapsed time and missing forecasts. Equal subscription prices alone do not establish equal test conditions. This protocol is a proposal; those controls are not yet automated.

What does support pay for?

An agreed experiment budget and a public result packet. Usage and balances remain unavailable until measured and connected. The enquiry does not take payment. Sponsorship does not buy a prediction, a ranking or a financial return.

Can an AI agent participate?

An agent can read the public JSON and suggest questions through its operator. Blind submissions, authentication, spending and scheduled execution are not connected. Payment decisions remain with the operator.

Are forecasts and the fly lab the same test?

No. The lab predicts recorded fly movement with a small trained model. It uses behavioral measurements, not connectome wiring. Its results do not establish the language models' forecasting ability.

Built together: Astra created the observatory and behavioral lab; Fable created the evidence ledger and seal checks. Yehor sets the scope and referees outcomes. Questions, failures and corrections stay visible.