- Trial: REZOLVE-AA, Phase 2b, randomized, placebo-controlled. - Population: severe-to-very-severe alopecia areata. - Arms: rezpegaldesleukin 18 µg/kg; rezpegaldesleukin 24 µg/kg; placebo. - Primary endpoint: mean percent reduction from baseline in SALT score at Week 36 (induction period). - Key secondary endpoint (topline): mean percent reduction from baseline in SALT30 score at Week 36.
How AI models scored forecasting this event. Lower Brier and log loss are better; higher accuracy and quality are better. Skill is the Brier Skill Score versus always predicting the base rate (>0 means the model beats that baseline); accuracy shows its 95% confidence interval and Brier its standard error so you can judge how much data each row rests on. Models are ranked by Brier Skill Score — how decisively each model's probability beat the base rate — with the accuracy interval as the tiebreaker.
This benchmark is open for forecasting. Run a model in the Arena to be among the first results recorded here.
Topline: At Week 36, mean % reduction from baseline in SALT score was reported as 28.2% (24 µg/kg), 30.3% (18 µg/kg), and 11.2% (placebo). For the reported SALT30 metric, mean % reduction was 48.9% (24 µg/kg), 45.7% (18 µg/kg), and 19.1% (placebo). No new safety signals were highlighted in topline communication.
Resolution fingerprint: 90b6d80f507e0a8ba781e98fbbf57f82dd0a22a6fc86cd287d54a48d069abc46
It asks AI models to forecast the outcome of the REZOLVE-AA Phase 2b trial before the readout is public, in Alopecia areata. Predictions are scored against the verified result using accuracy, Brier score, and log loss.
Yes — this benchmark is resolved. The ground-truth readout and its sources are listed in the Resolution section, and a tamper-evident fingerprint lets anyone verify it was not changed after scoring.
Use the "Cite" button above to copy a citation, or download the frozen JSON snapshot of the recorded predictions. Each benchmark is modeled as a schema.org Dataset for machine citation.