Mistral Large 4 scores 62% on DeepSWE, outperforming GLM-5.

🚨 AI News | TestingCatalog

🚨 AI News | TestingCatalog

@testingcatalog

Latest AI News on AI Agents, Model Releases, Tools, Leaks, and Rumors 🗞

7 708 підписників
Відкрити в Telegram
Mistral Large 4 scores 62% on DeepSWE, outperforming GLM-5.3, according to VentureBeat.

Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight).

Le Chonk also scores 15% on Harvey’s Legal Agent Benchmark (Legal tasks, SOTA open-weight).

We need a tech report now 👀
Відкрити пост в Telegram