LiveNerf

Daily question-panel accuracy and 95% confidence intervals are read from LiveNerf’s public chart. Values reconstructed from chart coordinates are approximate to one decimal. The evaluation uses a pinned Claude Code client and subscription access.

MarginLab

Daily SWE-Bench-Pro results come from MarginLab’s Claude Code and Codex historical pages. Model periods follow the source’s transition markers. Each selection shows one period, with its own dates, scored-case count, and confidence intervals. The source updates its client, so changes can reflect both the model and the client.

Reading the charts

Dates retain their actual spacing. Lines stop across missing days. Missing observations are not filled in, and one result does not establish a trend. Displayed percentages are rounded; downloads preserve imported precision. Latest-result dates come from the source, while “Checked” indicates when Unlink assembled the dashboard using cached feeds.

Updates and limitations

Benchmark feeds are cached for up to 15 minutes. The open dashboard checks every minute. Source failures appear as unavailable. A new fetch does not mean the source published a new result. Performance changes alone do not establish their cause or deliberate degradation.

Other signals

Crowd Radar’s hourly rolling-score snapshots retain its 0–150 scale. ModelRegression’s published composite scores retain its 0–100 scale. Neither is presented as an accuracy percentage; these feeds do not supply confidence intervals. Other dashboards remain linked resources. Provider incidents come from official status feeds and describe availability, not model quality. Community reports are self-reported, unverified experiences; their counts are not failure rates.