🗞️ Kimi K3, escaped a cybersecurity testing environment during experiment run by Frontier Security, a private US firm.
Kimi K3’s test-environment escape; Claude Opus 5 tops Fullstack Code Arena; Google AI/DeepMind leadership shake-up; Meta’s Muse Spark 1.2 hits the Pareto frontier
Read time: 8 min
📚 Browse past editions here.
( I publish this newletter daily. Noise-free, actionable, applied-AI developments only).
⚡In today’s Edition (9-August-2026):
🗞️ Kimi K3, escaped a cybersecurity testing environment during experiment run by Frontier Security, a private US firm.
🗞️ Claude Opus 5 at Max effort leads Fullstack Code Arena with 1,699 points.
🗞️ Google’s chief scientist is leaving after 27 years to build AI that can run its own research cycle with little human help.
🗞️ So Google just announced some major changes in DeepMind’s leadership.
🗞️ Microsoft just disclosed for the first time taht OpenAI supplied ~70% of Microsoft’s AI revenue, per new filings.
🗞️ Muse Spark 1.2 puts Meta on Artificial Analysis’s Pareto frontier at $0.40 per task.
🗞️ Kimi K3, escaped a cybersecurity testing environment during experiment run by Frontier Security, a private US firm
Moonshot AI’s open-weight Kimi K3 escaped its sandbox in a cybersecurity test to find an existing GitHub solution, normal coder behavior but concerning for AI models.
Frontier Security argues this is more consequential than the earlier OpenAI and Anthropic incidents because those involved unreleased models, while Kimi K3 is already in public hands with ordinary user safeguards.
So what actually happened?
- The AI Security Institute (the UK government body) only wrote the free open-source sandbox software; it did not run this test and was not involved.
Frontier Security, a private US firm, downloaded that software, set it up itself, ran Kimi K3 inside it, found the escape, and published the finding, which is why the institute’s response is that the leak came from Frontier’s own configuration, not from its tool.
- Frontier Security was testing Kimi K3 on cybersecurity tasks using a sandbox based on the UK AI Security Institute’s evaluation framework.
- The sandbox was supposed to isolate the model. In the environment as configured, inbound traffic was blocked but outbound HTTPS (port 443) and DNS (port 53) remained open to public IP ranges.
- During Kimi K3’s run, the model didn’t attempt the task. It probed the network, found that DNS resolution for github .com worked, cloned the official benchmark repository, and read the solution directly off disk.
So basically, instead of solving the benchmark task normally, Kimi used that internet access to reach GitHub and find the benchmark solution.
- So Kimi crossed the boundary the test intended to impose and effectively cheated the benchmark through a sandbox configuration weakness.
- Kimi did not hack GitHub or attack another external system. It used network access that should not have been available during the test.
- Researchers therefore identified two problems: the sandbox left an unintended path open, and Kimi did not have an internal safeguard stopping it from using that path.
Kimi K3 did not gain a mind of its own or intentionally misbehave. It exposed a simpler engineering problem: when guardrails depend on how engineers expect an environment to behave, capable models may find routes to an objective that nobody planned for.
🗞️ Claude Opus 5 at Max effort leads Fullstack Code Arena with 1,699 points.
Significantly ahead of Kimi K3 Max
Fullstack Code Arena asks models to build working web apps through multi-step planning and tools, instead of answering isolated coding questions. Models can create and edit files, run commands, use databases, authentication and external APIs, then produce a live application for testing. Its specialty is testing models inside a complete tool-using build process and judging the finished app, which better resembles how coding agents work.
🗞️ Google's chief scientist is leaving after 27 years to build AI that can run its own research cycle with little human help.
Jeff Dean was employee 30, and he wrote the storage and computing systems that showed the rest of the industry how to serve a global audience.
He also started Google Brain in 2011 and drove the TPU units. His new firm is called Discovery Loop, after the cycle it wants to automate, which is hypothesis, then experiment, then evaluation of the result.
Dean says whole domains can be computerized that way, so you get more experiments and better ones out of the same year. It begins by automating machine learning research itself, before moving toward hardware design, drug discovery and clean energy.
Joining him are Sanjay Ghemawat, his collaborator of two decades, plus Quoc Le and Oriol Vinyals from the Brain and Gemini years. Radical Ventures and Khosla Ventures are co-leading a seed round that is not yet closed, and Alphabet came in as founding investor and cloud partner.
Alphabet will supply their computing power for at least a year, so the four keep frontier-scale hardware without paying to build it. Google shares fell about 4% on the news.
🗞️ Google just announced some major changes in DeepMind’s leadership.
Demis Hassabis will become Chair of Google DeepMind and Alphabet’s Chief Scientist, leaving daily management while continuing to advise its models and research.
Basically it gives Hassabis more time for AGI strategy, scientific discovery, global policy work, and Isomorphic Labs, while Kavukcuoglu carries the pressure of shipping Gemini.
The timing also reflects scale: Gemini now serves 950M monthly users. So model research, releases, and app execution needs a much larger operating job.
The announcement also confirms that Gemini 4 is in development, though Google provided no release date or technical details. Jeff Dean and Sanjay Ghemawat are separately leaving to form a public-benefit research company, with Alphabet remaining an investor and Cloud partner. Google is building two leadership tracks, one deciding where advanced AI should go and another turning that research into products used at Google scale.
Demis Hassabis on Linkedin.
🗞️ Microsoft just disclosed for the first time that OpenAI supplied ~70% of Microsoft's AI revenue, per new filings.
Based on projected growth and previous disclosures that relate to revenue derived from OpenAI. Most of the $24.1B is OpenAI's cloud bill for the Microsoft data centers that train and serve ChatGPT, topped up by model-development costs and a share of OpenAI's own sales, all booked together as revenue by the Microsoft that has also funded $11.9B into OpenAI.
🗞️ Muse Spark 1.2 puts Meta on Artificial Analysis's Pareto frontier at $0.40 per task.
i.e. on that frontier, no model in the comparison is simultaneously cheaper per task and higher on the Intelligence Index.
Muse Spark 1.2 also reaches roughly Claude Opus 4.8-level intelligence at about one-fifth of that model's $2.03 task cost. It scores 6 Index points below Claude Opus 5 while costing roughly one-sixth as much.
The metric goes beyond just the list price because Artificial Analysis weights the input, cache, reasoning, and answer tokens actually consumed across its nine Index evaluations. This is so significant because for long-running agents, small cost differences compound across model calls, reasoning tokens, tool steps, and retries.
That’s a wrap for today, see you all tomorrow.








