- Trial: ATTAIN-MAINTAIN (NCT06584916), Phase 3, randomized, double-blind, placebo-controlled. - Population: adults with obesity/overweight who previously completed SURMOUNT-5 and reached a weight plateau. - Re-randomization: 3:2 to orforglipron (titrated to MTD 24 mg or 36 mg) vs placebo. - Primary objective: demonstrate superior maintenance of prior weight loss vs placebo over 52 weeks. - Rescue: participants who regained >=50% of prior weight loss could receive rescue orforglipron (post-hoc topline sometimes reports Week 24 as last pre-rescue time point).
How AI models scored forecasting this event. Lower Brier and log loss are better; higher accuracy and quality are better. Skill is the Brier Skill Score versus always predicting the base rate (>0 means the model beats that baseline); accuracy shows its 95% confidence interval and Brier its standard error so you can judge how much data each row rests on. Models are ranked by Brier Skill Score — how decisively each model's probability beat the base rate — with the accuracy interval as the tiebreaker.
This benchmark is open for forecasting. Run a model in the Arena to be among the first results recorded here.
Topline: Orforglipron met the primary and all key secondary endpoints for weight maintenance vs placebo at 52 weeks. Key reported weight changes: participants switching from Wegovy to orforglipron regained ~0.9 kg on average by Week 52; those switching from Zepbound regained ~5.0 kg on average by Week 52. In post-hoc Week 24 analyses (before placebo rescue eligibility), weight change from baseline was -0.1 kg vs +9.4 kg (placebo) for those switching from Wegovy and +2.6 kg vs +9.1 kg (placebo) for those switching from Zepbound. No hepatic safety signal was observed; overall safety was consistent with prior studies (GI AEs most common).
Resolution fingerprint: 994d82b865320bd5e5dbded0fbd0ec53a1d53e9b046e27a85af042d4a970bbad
It asks AI models to forecast the outcome of the ATTAIN-MAINTAIN Phase 3 trial before the readout is public, in Obesity. Predictions are scored against the verified result using accuracy, Brier score, and log loss.
Yes — this benchmark is resolved. The ground-truth readout and its sources are listed in the Resolution section, and a tamper-evident fingerprint lets anyone verify it was not changed after scoring.
Use the "Cite" button above to copy a citation, or download the frozen JSON snapshot of the recorded predictions. Each benchmark is modeled as a schema.org Dataset for machine citation.