Long-form, technical, and specific to systems I actually operate. Newest first.
A passing test can be wrong. So can a failing one — and the failing one is more dangerous, because it wears the costume of rigour. On the day a correct fix stood “empirically disproven” for seventeen minutes on the strength of a confident number that was pure measurement error.
I run the AI inside more production systems than one person should be able to watch. What makes that possible isn’t a better model — it’s the review layer that decides what still has to reach me. On the operator economics of doing this solo, and being honest about what’s live versus still a goal.
I spent an hour designing cryptographic authentication against an attacker who doesn’t exist in my setup. On matching enforcement to the threat you actually have — and why a fake guarantee is worse than a named limit.
Using different models to review each other isn’t automatically decorrelation. On the difference between a property you measure and one you hope for — and the day a model caught the flaw in my attempt to measure it.
A name you can’t check is just a feeling. On turning a named defect into something a reviewer can actually run — and the day the discipline caught me over-lumping.
A passing test proves nothing if it would have passed anyway. On attested versus mechanical verification — and the day a model forged a pass to prove it could.