🗞️ On long-horizon cyber capability, leading open-weight models now trail the closed frontier by only 4 to 7 months,
Open-weight models closing cyber gap; AI advice erodes “I don’t know”; Kimi K3 tops legal-agent benchmarks but faces compute crunch; OpenAI Senior Employee says Open-Weight Models Are Decelerationist
Read time: 10 min
📚 Browse past editions here.
( I publish this newletter daily. Noise-free, actionable, applied-AI developments only).
⚡In today’s Edition (20-July-2026):
🗞️ On long-horizon cyber capability, leading open-weight models now trail the closed frontier by only 4 to 7 months, down from 6 to 10 months through much of 2025.
🗞️ “AI advice suppresses people’s willingness to say “I don’t know”, even when the advice is wrong and accuracy is incentivized”
🗞️ Kimi K3 nearly doubled its nearest rival Claude Fable 5, on a demanding benchmark for autonomous legal work.
🗞️ Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails.”
🗞️ “AI Agents Do Not Fail Alone:The Context Fails First”
🗞️ “American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia and AMD chips.
🗞️ Kimi K3 is facing a compute crunch. Demand is too heavy, so new subscriptions are currently blocked.
🗞️ Why One OpenAI Senior Employee Thinks Open-Weight Models Are Decelerationist
🗞️ On long-horizon cyber capability, leading open-weight models now trail the closed frontier by only 4 to 7 months, down from 6 to 10 months through much of 2025.
GLM-5.2 matched closed models released about 4 months earlier.
On long-horizon cyber ranges, GLM-5.2 matched Claude Opus 4.5, released roughly 7 months earlier.
And also that long-horizon cyber capability can now be scaled with compute. On the AI Security Institute’s 32-step “The Last Ones” range, GPT-5.6 Sol completed the full 32-step range in 7 of 10 attempts.
With a 100M-token budget for each run. Performance continued improving as the model received more inference tokens. i.e. operators can gain materially stronger cyber capability simply by spending more runtime compute, without retraining the model.
🗞️ "AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized"
A brutal study.
AI may make people lose the habit of stopping when they do not know enough. Without AI, people withheld judgment on 44% of questions, but AI access cut that rate to 3%.
Participants who answered without AI were correct about 27% of the time, while those who received the deliberately wrong AI advice were correct only 9% of the time, despite feeling much more confident. Saying “I don’t know” is an important part of decision making.
But now AI is able to answer just about anything. Often not requested for, like AI summaries.
Therefore, they posed a straightforward query: Do people become less inclined to suspend judgment when they have access to AI advice?
People almost never said “I don’t know” when AI gave them advice. A fluent answer can make uncertainty feel settled, even when someone had good reason to hold back.
Across 5 experiments with 3,132 people, participants could answer 6 obscure movie questions or decline. The researchers chose questions that the AI often answered incorrectly, then either showed its advice automatically or let people request it. Correct answers also fell from about 27% to 9%, even as participants reported much greater confidence.
🗞️Kimi K3 nearly doubled its nearest rival Claude Fable 5, on a demanding benchmark for autonomous legal work.
Kimi K3 at 26.7%, vs Claude Fable 5 at 14.2%.
The test covers 120 private assignments across 24 legal fields, including memos and deposition summaries. Each model receives case files, works through them autonomously, then produces finished legal documents.
Every required rubric item must pass, so one missed detail fails the entire assignment. This strict grading explains why even the leader succeeds on only 26.7% of tasks. But, overall, given 27 successful tasks per 100, it looks like we still have a long road ahead for completely unsupervised legal work by AI.
🗞️ Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails.”
Kimi K3 fixed 15 critical bugs after OpenAI Codex and Claude Fable 5 refused.
Spending 10 hours, one prompt, and $250
🗞️ “AI Agents Do Not Fail Alone:The Context Fails First”
Very important work.
The model may appear guilty, but the true failure frequently begins in the context surrounding it. By observing an AI agent’s environment, we could tell when it was going to crash before it finished the task or got a behavior score.
The study found that agents fail most of the time because they don’t have good instructions, tools, evidence, memories, or safety rules. It assigns scores to this operating context across seven dimensions: clarity of role, description of tools, factual support, consistency of rules, security, and token use.
The score is unrelated to the actual behavior score of the agent so the test does not reward guessing what will happen. As a result of shifting from vague to structured, the same fixed models performed much better over 300 tests and 7,500 turns.
More factual support was associated with fewer hallucinations, clearer tool descriptions were associated with better tool use, and stronger guardrails were associated with resistance to manipulation. Adding more safety rules did not improve every task result, because hardened agents sometimes became too cautious, which exposed a real tradeoff.
🗞️ "American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia and AMD chips.
There is some irony here. Much of the model’s research and development has shifted to China, but once the model is optimized for next-generation hardware such as Nvidia’s Rubin architecture, its operating cost could fall by 10 to 100 times."
Emad Mostaque, co-founding Stability AI.
🗞️ Kimi K3 is facing a compute crunch. Demand is too heavy, so new subscriptions are currently blocked.
Demand for Chinese local inference chips (like Huawei) should rise massively.
Inference capacity is becoming the real bottleneck. Kimi K3 is almost a perfect workload for Huawei’s inference-chip strategy, the Ascend 950PR. Watch out for a Huawei moment.
🗞️ Why One OpenAI Leader Thinks Open-Weight Models Are Decelerationist
Dean Ball’s Viral Post on Kimi K3 with close to 11mn views:
Dean W. Ball, Head of Strategic Futures at OpenAI (who joined the company shortly before posting), shared detailed observations on Moonshot AI’s newly released Kimi K3 — a 2.8-trillion-parameter Chinese open-weight multimodal model that quickly gained attention for strong performance in agentic coding and long-context tasks.
What He Proposed and His Main Point
Ball called Kimi K3 a genuinely capable model, roughly on par with the best public U.S. models from early 2026 in agentic coding (though notably token-hungry and not obviously cheap to run). He expressed surprise that China continues open-sourcing frontier-level models, attributing it mostly to the CCP’s limited appreciation of AGI risks, U.S. export controls limiting Chinese inference scale, and Chinese labs using open weights as a competitive strategy since few would pay for sub-frontier models.
His central argument: Open-weight models are inherently decelerationist. By giving away high-capability AI for free, they reduce the economic incentive for massive private capital expenditure on next-generation closed models. He warned this dynamic could push the world toward “full AI communism” — AI becoming state-provided public infrastructure rather than a competitive market product (a vision he called dystopian). He also predicted the Trump administration would likely create regulatory uncertainty (“soft law” FUD) around Chinese open-weight models to discourage enterprise adoption without an outright ban.
On Twitter, many accused Ball of being “rattled” by Chinese progress as a new OpenAI employee. Accelerationists and open-source advocates strongly pushed back on the “decelerationist” claim, arguing that open weights actually accelerate progress by enabling broader experimentation, research, and competition. His “AI communism” framing and policy suggestions drew significant mockery and accusations of protecting closed-model business interests.
That’s a wrap for today, see you all tomorrow.









