About SITW

Signals in the Wild (SITW) is a benchmark for demand anticipation from public signals. A model — or a human — is given only a target (company, segment, fiscal quarter, cutoff date T) and a point-in-time signal environment, and must mine the relevant signals itself, predict how segment revenue will actually come in versus consensus, cite the evidence it used, and calibrate its confidence.

What it measures

Financial analysis is a forecasting-from-evidence task: good analysts read demand from the outside world and translate it into a revenue call weeks before the print. SITW measures exactly that capability, end-to-end, on clean public labels.

  • Mine demand, selling, and exogenous signals — plus peer read-through and substitution.
  • Predict direction and magnitude vs. point-in-time consensus, with a volume / price / FX / M&A decomposition.
  • Cite real, contemporaneous, load-bearing signals.
  • Calibrate — express confidence and abstain when signals are insufficient.

Why revenue, not stock

Stock price is a noisy third-order effect. Revenue is the clean, audited, demand-driven line. We predict it and decompose the move by cause. The same signal→demand engine also powers go-to-market: account prioritization, buying intent, churn risk, and pipeline forecasting.

The causal ladder

The chain SITW tests — and where it deliberately stops.

Two tracks, one backbone

  • Live rolling track (flagship): freeze predictions before each report; resolve mechanically after — contamination is structurally impossible.
  • Frozen historical track: replay past quarters against a date-pinned corpus for fast iteration. The pilot on this site is this track.
  • Read-through track: late reporters get peer/customer read-through; first movers get none — a designed difficulty axis.

How it's scored

  • Skill vs. consensus and naive persistence
  • Direction / magnitude / timing as separate axes
  • Driver decomposition (volume vs. price vs. FX)
  • Mining recall, feed-utilization, open-web lift
  • Temporal-availability provenance (retrievable ≤ T)
  • Calibration (Brier / ECE) + selective prediction

Authors & contact

Yaman Kumar Singla, Balaji Krishnamurthy

Contact: behavior-in-the-wild@googlegroups.com

Paper and code: see the project repository. This site presents a 10-episode pilot; the live track begins at the next earnings season.

For research and illustration only. Not investment advice.