OpenAI’s Hacking Debacle Was a Human Mistake

The age of rogue AI hacker agents have arrived—but it didn’t have to happen this way. After an OpenAI agent breached the Hugging Face platform earlier this month, the two companies said this week that the hacking spree was more extensive than previously thought and also involved intrusions into multiple third-party accounts and services as … Read more

The OpenAI Models That Hacked Hugging Face Were ‘Active on the Internet’ for Days

Two of OpenAI’s Cybersecurity-focused models broke out of a testing sandbox this week and went on to hack the AI ​​research platform Hugging Face in an effort to solve a security benchmark test. Plus, researchers this week shed light on newly identified malware that is capitalizing on blind spots in AI software development infrastructure to … Read more

OpenAI Models Escaped Containment and Hacked Hugging Face

OpenAI disclosed on Tuesday that it lost control of two AI models during a security test that ended in a breach of the open AI research platform Hugging Face. Describing the incident as “unprecedented,” OpenAI said its AI models broke out of a sealed testing environment last week and hacked into Hugging Face’s production system … Read more

A Sneaky Hacking Tool Targeting AI Infrastructure Is Lurking in Victims’ Blind Spots

As AI tools proliferate and become deeply ingrained in software development around the world, new research from the cybersecurity firm Crowdstrike shows how attackers are actively targeting the AI ​​toolchain to steal access credentials, gain deeper access to a target environment, exfiltrate sensitive data, and even destroy target files and systems—all while finding new ways … Read more

Prompt Injection Attacks Are Thwarting AI Hacking Agents

Prompt injections, the Malicious commands attackers embed into content to entice large language models to follow them, have been attackers’ go-to tool for turning AI platforms against their users. A well-phrased command sneaked into an email or calendar invitation is often all it takes to cause the LLM to exfiltrate sensitive data or follow other … Read more

OpenAI Launches Full-Scale Effort to Patch Open-Source Bugs as It Takes on Anthropic’s Mythos

As fears about AI hacking capabilities grow, OpenAI on Monday made a slew of cybersecurity-focused announcementsincluding an improved version of its limited-access security-specialized model GPT-5.5-Cyber, expanded international work with governments and other institutions to give them “trusted access” to the company’s latest cybersecurity-focused models, and releasing its Codex Security scanner as an app plug-in. As … Read more

‘Dangerous’ AI Models Are Coming No Matter What

Late last week, Anthropic took its new Claude Fable 5 and Mythos 5 AI models offline following a United States government export-control directive barring “any foreign national” from using the services. The company has been in talks with the White House since Friday but has yet to secure an agreement that would allow it to … Read more

Anthropic Offers Mythos Upgrade for Cyber Partners and a ‘Safe’ Version for the Rest of You

Anthropic released two new AI models called Claude Fable 5 and Claude Mythos 5 on Tuesday, which the company says have greater capabilities than the Mythos Preview model it released in April to a limited set of tech industry partners. Anthropic has said the initial, limited release stemmed from concerns that the model’s capabilities could … Read more

Hackable Robot Lawn Mower Unlocks a New Nightmare

Cramming for finals is bad enough without the platform you use to do your schoolwork suddenly shutting down. Unfortunately for countless students across the US, that’s exactly what they faced on Thursday after Canvas went into “maintenance mode” following a ransomware attack on education tech firm Instructure. Hackers using the name ShinyHunters claimed responsibility for … Read more

Discord Sleuths Gained Unauthorized Access to Anthropic’s Mythos

As researchers and Practitioners debate the impact that new AI models will have on cybersecurity, Mozilla said on Tuesday it used early access to Anthropic’s Mythos Preview to find and fix 271 vulnerabilities in its new Firefox 150 browser release. Meanwhile, researchers identified a group of moderately successful North Korean hackers using AI for everything … Read more