cybersecurity

Evaluating GLM-5.3 Cybersecurity Capabilities

Romulo Santos

Last week, Anthropic published an interesting post analyzing the cybersecurity capabilities of Z.ai’s open-weight GLM-5.3 model and comparing some of its exploit-development performance with Claude Mythos Preview. Anthropic’s testing found GLM-5.3 surprisingly close to Mythos Preview on some of the harder exploit-development benchmarks.

A couple of weeks earlier, NIST’s Center for AI Standards and Innovation (CAISI) had published its own assessment. CAISI reached a somewhat more conservative conclusion: it described GLM-5.3 as the most cyber-capable open-weight model released so far, while estimating that its overall cyber capability still trails the current U.S. frontier by roughly four months.

The Abstraction Gap Between Security and Software Engineering

Romulo Santos

The other day, I was talking with some former teammates who now work as software engineers at different companies, and the conversation turned to their relationships with their security teams. A familiar theme quickly emerged around the disconnect between security recommendations and the realities of their day-to-day engineering work. Sometimes it feels like we have come a long way as an industry, but many security teams still operate at a considerable distance from the engineers building and maintaining the systems they are trying to protect.