
Today’s issue: A Kubernetes colonoscopy, the secret I’ve been keeping from David K Piano, and some mixed feelings about the latest viral AI slop.
Welcome to #525.


Threatening my AI with some Permissions Policy Algebra
You’d think that as the fancy coding models get smarter, they’d also get better at following rules. But as my toddler reminds me daily: the smarter something gets, the more ungovernable it becomes.
Which is why it’s wild that the standard way for coding agents to decide whether to do something risky is to just… ask themselves? Claude Code’s auto mode, Codex’s auto-review, and the other agents all hand sketchy tool calls to a second model to judge if it’s safe before moving forward.
OpenAI says Codex’s auto-review caught 99.3% of prompt injections in its synthetic evals, which sounds great until you remember that the benchmark doesn’t measure every risky thing an agent might do (and also wtf does that benchmark actually measure).
And the real problem is what the judge can’t see. Which means your agent could read a bug report with a real customer’s email in it, then three turns later paste it into a public GitHub issue because each call looked fine on its own.
This is the hairy problem that OpenAPPA is trying to fix. It’s a new, open-source flying bison deterministic guardrail for childproofing your coding agents with the power of Agentic Permissions Policy Algebra.
I find algebra as sexy as anyone, but the deterministic part is the key point. The core policy engine isn’t a model that can get prompt injected. It answers one question before every action: is this data allowed to go to this destination?
Here’s how:
Security labels: Every session gets tagged with who’s allowed to see its data and how much that data can be trusted. Reading a private repo narrows the audience and reading a random web page tanks the trust score. Once a session gets more restricted it never loosens back up.
Tool contracts: Declarative rules for each tool that OpenAPPA checks before every call and updates after.
Remedy plans: When a call gets blocked, the agent gets a machine-readable list of ways forward, like redacting the PII or asking a human to approve that one specific call. Archestra says this is why their agents still finish ~90% of tasks instead of the ~37% other deterministic guardrails manage.
Bottom line: Turns out, the secret to an effective childlock might just be algebra. Time to post the quadratic formula on the outside of my pantry.


TFW you have to test the 25,000 line PR your agent cooked up
The one thing everyone’s talking about: agents now write the majority of your code, but how do you test it?
As our name suggests, we use real user sessions to meticulously test every edge case and flag what will break before you merge:
With Meticulous, Notion cut frontend testing and maintenance efforts by 30%, with individual engineers estimating savings of up to 5 hours a week.
Schedule a demo to learn more.

Sam Rose wrote a visual guide demonstrating how Kubernetes probes work. Between this and my 20-year high school reunion, I think the universe is telling me to get a colonoscopy.
David Zhang created jevgrep, a CLI that lets you search for code using natural language. jg “What would you say...you do here?”
Give your apps and AI agents real-time data from 100+ search engines using SerpApi. Power your AI tools, track rankings, or discover local businesses with real-time results from Google Search, Maps, AI Overviews, and more. Get structured JSON or Markdown with one simple GET request. [sponsored]
XState added support for Effect.ts that lets you combine event-driven workflows with Effect services while handling cancellation and testing. I still don’t know what Effect is, and at this point I’m too afraid to ask.
Cloudflare released version 1.0 of vinext, the OG “slopfork” that lets you run Next.js with Vite, which reminds me, we’re a little bit overdue for the next Dane vs Guillermo Twitter feud.
Pooya Parsa created upm, a fast, tiny package manager for the npm registry.
The future of mobile development is agentic. And it’s built on Expo. For years, mobile shipped slower than the web. Now agents are writing code, testing it, fixing it, shipping it OTA, and then monitoring it. All on Expo. [sponsored]
The VoidZero team is still shipping post-acquisition and just released Vite+ 1.0, which combines Vite, Vitest, Oxlint, Oxfmt, Rolldown, tsdown, and Vite Task into a single package.
Turso 0.8.0 just got around SQLite’s single-writer bottleneck with concurrent writes.
Domenic Denicola made a website that pushes AI-generated content to its limit. It’s weird. It makes me feel bad. But I was also kind of impressed?