The theme this week is trust—and how little of it is warranted.
On the security side: tool specifications in AI agents are a bigger attack surface than most teams realize. Schema-formatted tool definitions change model behavior in ways that degrade safety, and KV cache reuse opens a new class of injection vectors. These aren't theoretical. If you're shipping agentic systems, this cluster deserves a read before your next deploy.
On the measurement side: a validity audit of four major agent-safety benchmarks found their scores are being quoted interchangeably as if they measure the same thing. They don't. R-Judge, InjecAgent, AgentHarm, AgentDojo—all measuring different behaviors, all getting cited as 'safety.' The EvalSafetyGap survey reinforces this: benchmark scores can improve while actual alignment properties stay flat or get worse.
The one genuinely useful shipping item: Google's Gemini 3.6 Flash with managed agents and lifecycle hooks. Not revolutionary, but it's production infrastructure you can use today.
Everywhere else, the pattern is researchers identifying real problems—LLMs failing in clinical reasoning, code agents avoiding deletion, long-context models copying instead of reasoning—and offering partial fixes. Good signal, but none of it is plug-and-play.
The EU AI content labeling mandate went live. Start your compliance checklist now if you haven't.