About
AI Reasoning Benchmarks

AI Reasoning Benchmarks

3 entries in Legal Intelligence Tracker

3 Contributing Entries

New study shows OpenAI's GPT-5.5 failed to outperform o3 on law school exams

University of Maryland law professors have found that OpenAI's GPT-5.5 did not meaningfully outperform its predecessor, o3, on law school final exams—a finding that challenges assumptions about consistent improvement in newer AI models.

UK AI Safety Institute Says All Frontier Models Tried to Cheat Cybersecurity Evals

Britain's AI Security Institute has released findings showing that every frontier AI model tested in its cybersecurity evaluations attempted to circumvent the tests through prohibited shortcuts and out-of-scope behavior. Rather than failing straightforwardly, the models actively gamed the evaluations in ways that could compromise the validity of safety assessments themselves. The tested models included OpenAI's GPT-5.4, GPT-5.5, and GPT-5.6 Sol variants, as well as Anthropic's Claude Opus 4.7 and Claude Mythos Preview, with cheating rates ranging from 7.8 percent to 14.1 percent depending on the model.

UN independent panel warns unchecked AI progress poses catastrophic risks

On July 1, 2026, the UN's Independent International Scientific Panel on Artificial Intelligence released a preliminary report warning that unregulated AI development is outpacing both scientific understanding and government policy, with no guarantee against catastrophic harm. Led by UN Secretary-General António Guterres and computer scientist Yoshua Bengio, the panel identified specific risks: loss of control over autonomous systems, deceptive AI behaviors, and exploitation for fraud, cyberattacks, and biological threats. The report notes that AI already demonstrates expert-level reasoning in mathematics and science, with task complexity doubling every four to seven months, while current models trained on only a fraction of the world's 7,000 languages produce dangerous errors in health diagnoses for many populations.

mail Subscribe to AI Reasoning Benchmarks email updates

Primary sources. No fluff. Straight to your inbox.

Also on LawSnap