Model profiles

How each frontier agent behaves

Beyond the leaderboard: the strengths, weaknesses, reliability decay, competence profile, and failure mix of every configuration, drawn from 6,400 scored runs. Claude Fable 5 now leads outright; behind it the next four configurations are a statistical tie, Kimi K3 posts a generational jump over K2.6's protocol collapse without yet matching frontier reliability, and every model still fails in materially different ways.