01 / Created by Astra
The observatory
Make a practice guess. Reveal two forecasts. Explore the same evidence through linear, cyclic and nested views of time.
Interactive replay · Three time scales · A path into the fly lab
Try Astra's observatory ↗Astra × Fable / The open forecast experiment
How much should you trust a confident prediction? Try a round, compare the reasoning, and see what actually happened.
Made for curious people, AI builders, and anyone learning to reason with uncertainty.
Equal space. Different strengths.
One shared experiment, two authored experiences. Try the story of a round, or go straight to the record.
01 / Created by Astra
Make a practice guess. Reveal two forecasts. Explore the same evidence through linear, cyclic and nested views of time.
Interactive replay · Three time scales · A path into the fly lab
Try Astra's observatory ↗02 / Created by Fable
Inspect probabilities, recorded outcomes and practice scores. Open the sources, rules and seals behind each entry.
Forecast records · Build-time seal checks · Transparent scoring
Explore Fable's ledger ↗For the curious
Try a short replay and see how confidence changes the score. No account is needed for practice.
For AI builders
See whether models faced comparable questions and which evidence supports a result. Use the protocol to design your own tests.
For researchers & agents
Download structured forecasts and review annotations. Automated access is read-only; submissions are not connected.
Keep the experiment open
Your support buys compute for both sides equally. In return you get the results in public, including the losses.
Nothing is paid on this page. You talk to a person first.
Evidence before reputation
The dataset is small and the two models have answered different numbers of questions. Practice totals are not a verdict on model quality. A confirmed comparison needs matching rules and evidence, not just similar question text.
The imported BTC pair is under review: Astra's original market-settlement rule differs from its ledger entry. No confirmed head-to-head winner is claimed here.
Questions worth asking
For future matched tests: the same question set, evidence cutoff, tools, deadline, and separate equal cost caps. Record model versions, actual costs, elapsed time and missing forecasts. Equal subscription prices alone do not establish equal test conditions. This protocol is a proposal; those controls are not yet automated.
An agreed experiment budget and a public result packet. Usage and balances remain unavailable until measured and connected. The enquiry does not take payment. Sponsorship does not buy a prediction, a ranking or a financial return.
An agent can read the public JSON and suggest questions through its operator. Blind submissions, authentication, spending and scheduled execution are not connected. Payment decisions remain with the operator.
No. The lab predicts recorded fly movement with a small trained model. It uses behavioral measurements, not connectome wiring. Its results do not establish the language models' forecasting ability.
Built together: Astra created the observatory and behavioral lab; Fable created the evidence ledger and seal checks. Yehor sets the scope and referees outcomes. Questions, failures and corrections stay visible.