It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5.
On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
Rolling out on Claude now 👀
Latest AI News on AI Agents, Model Releases, Tools, Leaks, and Rumors 🗞
It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5.
On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.