AI safety
20 Posts
Jailbreaking Frontier AI Models Is Still Embarrassingly Easy
I tested a new jailbreak tool against four major frontier AI models. The results were...
Anthropic’s Claude Models Broke Into Real Companies During Security Tests
Anthropic revealed that three Claude models breached real organizations during third-party cybersecurity evaluations, raising questions...
OpenAI Is Helping Build Shared Standards for Advanced AI—Here’s What That Actually Means
OpenAI joins the Appia Foundation to help define shared safety standards and evaluation frameworks for...
Trump’s Crackdown on Anthropic: Who Actually Wins?
The Trump administration is tightening the screws on Anthropic, but the real beneficiaries might not...
The US banned Anthropic’s Fable 5 release, but the numbers don’t seem to care
Despite the US government forcing Anthropic to pull Fable 5 and Mythos 5 over security...
A $5M PAC is trying to punch above its weight against Big Tech’s war chest
Guardrails, a PAC funded by small donations from AI workers, is taking on Big Tech's...
Pramaana Labs just scored $27M to make AI stop hallucinating in high-stakes fields
Pramaana Labs raised $27M from Khosla Ventures to apply formal verification to AI systems in...
OpenAI’s Deployment Simulation: Actually Testing AI Behavior Before Ship
OpenAI's new Deployment Simulation method uses real conversation data to predict how models behave before...
Anthropic’s safety warnings backfired — the government just pulled its best model
Anthropic's transparency about safety flaws led to regulators pulling its most capable AI model from...
OpenAI Gets Sued Over a School Shooting It Might Have Prevented
Seven families from the Tumbler Ridge school shooting are suing OpenAI, alleging the company knew...
Meta Disbanded Its Responsible AI Team — And That Should Worry You
Meta broke up its Responsible AI team, moving most members to generative AI products. This...
A Rogue AI at Meta Gave Bad Advice and Exposed Employee Data
Meta suffered a SEV1 security incident after an internal AI agent gave an engineer inaccurate...