artificial intelligence

Evaluating GLM-5.3 Cybersecurity Capabilities

Romulo Santos

Last week, Anthropic published an interesting post analyzing the cybersecurity capabilities of Z.ai’s open-weight GLM-5.3 model and comparing some of its exploit-development performance with Claude Mythos Preview. Anthropic’s testing found GLM-5.3 surprisingly close to Mythos Preview on some of the harder exploit-development benchmarks.

A couple of weeks earlier, NIST’s Center for AI Standards and Innovation (CAISI) had published its own assessment. CAISI reached a somewhat more conservative conclusion: it described GLM-5.3 as the most cyber-capable open-weight model released so far, while estimating that its overall cyber capability still trails the current U.S. frontier by roughly four months.

Autonomous Agentic Coding With Deterministic Verification Using Claude Code

Romulo Santos

After I wrote How AI Is Changing Software Engineering Work, I kept returning to one part of the argument: whether deterministic verification steps would help increase trust in AI-generated output. When I considered which languages and platforms would favor that, Go and Rust came to mind first. While other languages can, of course, also offer good verification controls, Go and Rust came to mind as good candidates because both ship standardized toolchains with strong support for automated, deterministic verification. Tests, formatting, vetting, and compilation can all be expressed as repeatable commands with clear pass/fail outcomes — exactly the kind of signal an agent can be held to.

How AI Is Changing Software Engineering Work

Romulo Santos

Over the past few years, we have watched AI move from completing lines of code to taking on work we might once have handed to another engineer. Coding agents can inspect a repository, propose an implementation, modify several files, write tests, and return something that looks remarkably close to finished. What has not become equally easy is deciding whether that work belongs in a production system.

Someone still has to understand what the change does, whether it solves the right problem, how it can fail, what it exposes, how it behaves under load, and whether the surrounding system can absorb it. As of today, the code may arrive faster, but the confidence required to move it into production still has to be established separately.

This changes where engineering value is concentrated: less in producing each line of code, and more in understanding systems, verifying outcomes, and operating software safely. It also means AI will amplify the strengths and weaknesses of the engineering environment around it.