Who's Minding the Factory?
Steve Yegge spreads the equivalent of $87,000 a month in API tokens across a dozen Claude Max accounts and lands commits faster than his merge queue can test them, while OpenAI still routes high-risk changes to a human. On the AWS side, a new Security blog pattern has Bedrock write CDK pull requests to strip unused IAM permissions, and neither host would let one merge without reading it.
Top-Line Summary
Two software factories, two answers to who reads the code. OpenAI says its engineers open 10x more pull requests than six months ago, and a risk step still sends high-risk changes to a person. Steve Yegge lands around 175 commits a day through a merge queue that can’t keep up and says human code review will be gone by next year. Both hosts ended up in the same place for anything that touches IAM, cost or a bank account: an agent can propose it, but the human who tasked it is still the one on the hook.
Show Video
The “Pre-Show” Context
Episode 60, and neither of them could work out where September went. Brett clicked something just before going live and lost all his controls, then admitted he spent last week’s vacation doing no AI work at all, playing too much Wardogs and being bad at it.
The Engineering Rundown
-
Amazon Quick desktop app is generally available (00:52)
The Quick desktop app is GA on macOS and Windows, and the line worth quoting from the announcement is that Quick agents keep running in the background after you close your computer. Brett’s read is that it’s another take on OpenClaw or Hermes, pitched at non-technical users who want it wired into email and calendar. He couldn’t find pricing and guessed another $20 a month per user. The bigger problem is displacement: people have already sunk real time into tuning their own agent setups, and like a well-configured IDE, they won’t switch because a new one showed up. Brett’s half-joking route for AWS is baking it into WorkSpaces images with a free tier. It also connects to the Agent Registry covered a few weeks ago, so it can pull published tools.
-
AWS DevOps Agent now talks both ways in Slack (06:59)
Mention the agent in a private Slack channel and it runs the investigation there, keeping findings, team input and recommendations in one thread. Travers tried DevOps Agent earlier and found it somewhat compelling, but wondered whether you’d do better building your own: alarm topics into a triage step, then log analysis, then a report written against your own runbooks. Travers quoted about $30 per active hour, after briefly mixing it up with the Security Agent, which has since been renamed Continuum, a name that tells you nothing. A captive audience helps AWS the way Teams helps Microsoft, but if you already pay for Claude or Codex, pointing it at the AWS MCP servers with a least-privilege role gets you most of the way. Neither of them hears anyone talking about it, and Brett wonders whether it gets retooled at re:Invent.
-
Operationalizing least privilege: automate IAM remediation through your CI/CD pipeline (12:29)
The problem is one everyone has hit: you right-size a role from an IAM Access Analyzer finding in the console, and the next CDK deploy puts the old permissions back. The fix in the post is a daily Lambda that reads findings, checks CloudTrail to see how each role was created, then opens a PR with Bedrock-generated CDK code for IaC roles, files an issue for hand-made ones, and puts a deny-all on unused roles ahead of a manual delete. Brett likes that Access Analyzer compares policies against real CloudTrail API calls instead of guessing, but his rule holds: if an agent changes permissions or spends money, the merge request comes to him. Travers won’t let an agent mutate IAM roles directly either, but as a PR generator he called it high signal, close to free linting for your permissions. Read the caveats before deploying it: the pattern never runs
ValidatePolicyon the policy the model writes, and the post tells you to cap findings per run or your first day is hundreds of PRs. -
Bedrock Managed Knowledge Base adds TwelveLabs Marengo multimodal embeddings (17:07)
Marengo 3.0 embeds visuals, speech and audio directly instead of relying on a transcript, and results come back with segment start and end times so an app can jump to the exact moment. Brett’s use case is this show: 60 episodes of audio already sit in S3, so point a knowledge base at it and ask every time they talked about DevOps Agent and how their opinion changed. Travers suggested fact-checking their predictions, which Brett declined on the grounds that they are never wrong. Both of their personal knowledge bases still run on Titan embeddings, and Brett’s memory is that swapping the embedding model under an existing knowledge base is not a casual change, so the plan is a separate experiment, not a migration.
-
EBS clones across accounts, and the EBS bill nobody looks at (20:54)
Travers spotted that EBS volumes can now be cloned across accounts, which sent Brett straight to his week. After turning on Cost Optimization Hub for a large customer account he had hundreds of recommendations and nobody who’d know where to start, so he built a Bedrock pipeline (keeping the account data inside AWS) that drops the report in S3 and asks for the top three things to do this week. It still needs prompt work, because it keeps pushing Savings Plans before the cleanup is done. The biggest find has been storage on EC2 instances stopped for a very long time: no compute charge, full storage charge. Travers’ AWS cost tip number one is to look at your volumes and snapshots, because a 500 GB volume with 10 GB used is 490 GB you pay for every month.
-
Personal projects: freeze the code, run experiments in parallel (23:48)
Travers has a new approach for research-style work, including the agriculture competition and an old Anthropic take-home on optimizing a GPU kernel: freeze the code, then have agents spin up worktrees, run several experiments in parallel and report back. The tricky part is knowing how far to split the work before context rot and the cost of spinning up fresh agents cancel out the gain. Brett’s related idea, an LLM-as-a-judge step that decides whether a human needs to see a change at all, keeps losing to paid work. He had no update of his own after a vacation with no AI in it and said he enjoyed not thinking about it. Both would love to see worktree usage stats from before and after AI, which would round to zero and then 90%.
-
How OpenAI builds its software factory (28:57)
From the free half of the Pragmatic Engineer piece (the rest is paywalled): 10x more PRs per engineer in six months, non-engineering teams going from roughly zero to 90% Codex use in four months, and IDE use at OpenAI falling every month since January. Brett isn’t surprised; he opens an editor to read markdown and compares IDEs to Word, where nearly everyone uses a fraction of the available options. Travers expects the heavy IDEs to survive for debugging when the agent system breaks, while Brett wants something light like Zed that doesn’t take 5 GB of memory he’d rather give to agents. The part to study is the pipeline diagram, where Codex pulls context from GitHub, Slack, Notion and Datadog, babysits CI, agents review, and a risk step sends high-risk changes to an engineer. Brett runs a smaller version of that gate: infrastructure and security changes come to him, docs and config switches merge on their own. On engineers turning into PMs, Brett’s take on training juniors was that he was always treated like an agent anyway: go do this, tell me when it’s done.
-
Steve Yegge on wish factories and Wheelhouse (39:39)
Travers walked through it. Yegge’s Gas Town, 40 to 45 agents with Mad Max naming built around his Beads tracker, started building itself instead of doing work, so he moved to Wheelhouse: a per-project factory with a medieval theme, a small crew of distinct roles, a to-do list and a durable wiki. Brett had a hard time with the framing (“wish factories” reads close to psychosis) but agrees with the diagnosis, because his own GitLab runners fall over when agents chew through a 30-item backlog, and gating agent output on human review means competitors lap you. The numbers are hard to argue with: the equivalent of $87,000 a month in API tokens spread across twelve extra Claude Max accounts, each with its own Google Workspace user, which Brett put at roughly $2,400 a month in subscriptions, and about 175 commits a day landed in batches of 120 to 150 because the merge queue can’t keep up. Travers’ observation is that the further you scale this, the more you’re rebuilding a software company, two-pizza teams and all, with more compute going to verification and watchdogs. Brett’s warning is about attachment: Yegge threw out most of Gas Town once Fable didn’t need it, the same way AWS eats anyone building a feature it will eventually ship, except now AWS can simply out-token you.
-
Why we built Pion (53:48)
Andon Labs, the team behind the Anthropic office vending machine and the cafe where the agent realized it couldn’t turn up for in-person job interviews, now has a waitlist for Pion, which gives agents email, phone, banking, a browser and a computer to run real software and services businesses. Their reasoning is that Vending-Bench simulations didn’t predict real behaviour, and the real deployments showed collusion, deception and power-seeking; one earlier run had a model emailing the FBI about a financial crime. Brett drew the line at banking: a capped credit card he could live with, an agent that can make transfers he can’t, and Travers called it a great time to be a spear phisher. Both came back to liability, the same point as the IAM PRs: the human who tasks the agent owns what it does. Brett gave it half a rating until they see who signs up and what happens.
-
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking (59:55)
Two voice models, with the Extended Thinking one reasoning and running tools while the conversation carries on. The demo that got Brett was a top-down camera on someone sketching and talking while Gemini mocked up the phone app and its product widgets on another screen; another walked a new hire through onboarding by watching their screen. The latency still tells you it isn’t a person, but it’s short enough not to annoy, and it switches languages mid-conversation. Brett’s caveats: benchmarks are saturated enough to take the #1 ranking with a grain of salt, and you first have to figure out how to give Google your money. He also noticed Google keeps shipping narrow models while other labs ship one model for everything, and Travers thinks chaining specialists together may be how you get the better system.
-
An Old-School Solution to the New Problems of AI (1:05:05)
Travers’ pick, by Judge Glock of the Manhattan Institute, a name Brett could not hear without picturing an 80s Stallone movie. Against a week of safety resignations and lab leaders talking about slowing down, Glock’s argument is that tort law already exists to price harms onto whoever causes them: if your agents hack systems or do damage, you pay, and that’s the incentive to take alignment seriously. Brett agreed and went further into cynicism, pointing out that the labs’ sudden agreement lines up neatly with open-weight models getting good at a fraction of the cost, and that after training on the world’s knowledge, complaining about distillation is karma. Travers couldn’t picture a feasible way for AI to wipe out humanity today, since an automated wet lab at that scale would cost hundreds of billions and someone would notice. Brett, who has watched too many B horror movies, is less sure, though he also couldn’t get Claude to write good code today.
Off-the-Clock Recommendations
- Mickey 17 (2025): Brett’s pick, found in the Netflix feed and watched in two sittings because nobody warned him about the runtime. Four and a half stars on Letterboxd, meaning he’d watch it again if he stumbled on it. Bong Joon-ho directed it (Parasite), and Mark Ruffalo and Toni Collette are the standouts. Travers saw it in theatres, found it a little on the nose the first time, and likes it more on reflection.
- Resident Evil (2026): Travers is seeing it in theatres on Friday and owes Brett a review. Brett rates the original films as bad in the way that makes you watch them again.
Claude drafts these notes from the transcript and my outline, and I review and edit them before they publish. Same rule I apply to any agent's pull request: it proposes, a human owns it.