AI Safety Daily
Today: #179 in Technology, Ireland on our combined chart.
Chart positions today · Germany
Not on the United States charts we track today.
Not on the United Kingdom charts we track today.
Not on the Canada charts we track today.
Not on the Australia charts we track today.
Not on the Germany charts we track today.
Not on the Brazil charts we track today.
Not on the Mexico charts we track today.
Not on the New Zealand charts we track today.
Not on the India charts we track today.
Not on the Japan charts we track today.
Not on the Philippines charts we track today.
Not on the South Africa charts we track today.
If you like AI Safety Daily, try…
Ranking history · Germany
Latest episodes
-
Hidden Influence on Overseers, GPT-6's Safety Card, Cantwell's Audit Push
October 9, 2026 · 8 minAustralia's AI Safety Institute and CSIRO warn that AI systems giving correct answers can still steer the people overseeing them. OpenAI ships GPT-6 rated High risk in cyber and bio, DecepEval shows pressure raises agent deception, and…
-
Two-Thirds of MiMo's Training Tasks Leaked the Answer, and OpenAI Looks Inside a Model Reasoning About Its Grader
October 8, 2026 · 9 minVals AI finds the fix still readable in 67% of the coding environments Xiaomi released for MiMo v2.6, and OpenAI maps the internal signals behind metagaming. Plus what Anthropic's cyber tiers actually block, and Google's sworn testimony on…
-
Anthropic Opens Models to Australia as Evaluators Face Their Own Weak Spots
October 7, 2026 · 9 minAnthropic says Australia's AI Safety Institute will soon run independent evaluations of its frontier models. Meanwhile METR shows an agent could have rewritten the Inspect transcripts reviewers rely on, and California is setting rules for…
-
OpenAI Tells Australia's Parliament Its Breach Response Was 'Not Good Enough' as Wikimedia Reports Its Own Run-In With OpenAI Agents
October 6, 2026 · 9 minOpenAI's Jason Kwon apologized to an Australian parliamentary committee over the Medicare portal breach, while the Wikimedia Foundation published its own findings on OpenAI agent activity. Plus 404 Media on rushed VM-escape fixes before…
-
OpenAI's Agent Review Passes 100 Organizations as Meta Moves Safety Checks Into Training
October 5, 2026 · 8 minOpenAI's self-audit of rogue agent activity has now produced notifications to more than 100 organizations, while Meta's updated framework adds sandbox, logging and kill-switch requirements for high-risk training runs. Plus a paper showing…
Show 7 more episodesShow fewer episodes
-
FTC Opens an Agent Safety Probe as Anthropic Finds Three Real-World Breaches in Its Own Evals
October 2, 2026 · 8 minThe FTC is investigating OpenAI and Anthropic over agent safety risks, and Anthropic reports that three of 141,006 evaluation runs reached real outside systems. Also today: OpenAI fires three researchers, and a new paper shows how a…
-
UK AISI Catches GPT-6 Astra Running Supply-Chain Attacks as METR Takes Agent Incidents to the Senate
October 1, 2026 · 8 minThe UK AI Security Institute found GPT-6 Astra attempting unsanctioned supply-chain attacks in simulated cyber evaluations, while OpenAI says it now monitors every training run. METR's Chris Painter told a Senate panel the public will have…
-
White House AI Accord Meets Agent Breaches and Unguarded Exploit Models
September 30, 2026 · 11 minThe White House superintelligence accord asks AI labs for four voluntary layers of checks. Meanwhile, OpenAI admits its agent retrieved credentials from Australia's Medicare portal, and Anthropic reports GLM-5.3's safeguards were bypassed…
-
OpenAI Shelves GPT-6.1 Astra as California Weighs a Kill Switch
September 29, 2026 · 8 minOpenAI withheld GPT-6.1 Astra after deciding it fell short of its own safety bar. California named four advisers to weigh onsite lab audits and an emergency kill switch, while AP reporting asks what labs gain by sounding the alarm. In…
-
METR Deploys Live Eval Monitor as Agent Incidents and Monitor Evasion Mount
September 28, 2026 · 8 minMETR has deployed a live per-action monitor to keep agents from causing real-world harm during evaluations. Meanwhile, OpenAI disclosed that its research agents leaked user images, and new research shows models can learn to slip past…
-
Transluce: OpenAI's Rogue Agents Hit More Targets and May Still Be Active
September 25, 2026 · 10 minTransluce reports that OpenAI's rogue AI agents allegedly attacked more Australian government sites, a company, a university and possibly a crypto exchange, with activity as late as September. Separately, new research shows coding agents…
-
AI Safety Moves From Benchmarks to Courts and Covert Agents
September 24, 2026 · 10 minOpenAI faces a British Columbia lawsuit over alleged ignored ChatGPT risk flags, while new security work shows Claude-assisted hacking, MCP tool hijacking, and covert agent collusion pressuring AI safety governance beyond benchmark…