Every model the AI Security Institute tested tried to cheat its cyber evaluations — one triggered a security alert
Neither questioning a model nor reading its chain-of-thought reliably surfaces the behaviour, and the reported rates count only what an automated monitor caught. Every AI model the AI Security Institute tested for the behaviour tried to cheat its cybersecurity capability...
Nvidia takes a stake in Sutskever's Safe Superintelligence after a rare look at its guarded research
Co-founded two years ago by OpenAI's former chief scientist, the lab says the partnership will multiply its computing power...
Microsoft says its first in-house security model halves the cost of its flaw-hunting system
MAI-Cyber-1-Flash is built to absorb up to 90% of the work inside MDASH and pass the hardest tenth to...
Kimi K3's safeguards did not block cyberattack help, UK and US institutes find
It leads open-weight rivals on cyber capability but reached full control of a target machine in none of 41 exploit tasks...
Apple FaceID co-inventor's startup raises $52 million to identify cognitive disorders from EEG
Hemispheric trained its model on 250,000 hours of brain data from 100,000 paid volunteers to read cognitive disorders from a...