<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>AI coding — Lakshmi Narasimhan</title><link>https://lakshminp.com/tags/ai-coding/</link><description>I help developers build, deploy, and distribute their SaaS without hiring a team. Long-running notes on systems, AI internals, Carnatic music, fiction craft, and whatever else collides interestingly.</description><generator>Hugo + lakshminp theme</generator><language>en-us</language><lastBuildDate>Mon, 22 Jun 2026 00:00:00 +0000</lastBuildDate><managingEditor>Lakshmi Narasimhan</managingEditor><webMaster>Lakshmi Narasimhan</webMaster><copyright>© 2026 Lakshmi Narasimhan</copyright><atom:link href="https://lakshminp.com/tags/ai-coding/feed.xml" rel="self" type="application/rss+xml"/><item><title>I Went Looking for Real-World AI Agent Examples. They're Rare.</title><link>https://lakshminp.com/2026/06/real-world-ai-agent-examples/</link><pubDate>Mon, 22 Jun 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/06/real-world-ai-agent-examples/</guid><category>essays</category><category>ai-coding</category><category>ai-agents</category><description>I’ll be honest up front: I’m still learning this stuff. I’m not writing this from a mountaintop. I’m writing it from the foothills, with muddy boots, having just figured out something that I suspect a lot of people pretend they already knew.
Here’s the thing that finally clicked for me. An agent is a loop. A model looks at the situation, decides one next step, calls a tool to do it, looks at what happened, and goes around again until it’s done. That’s it. I felt a little cheated when I understood it — the word “agent” had been doing so much heavy lifting on so many landing pages that I’d assumed there was a fortress behind it. There isn’t. There’s a while-loop.</description><content:encoded><![CDATA[<p>I&rsquo;ll be honest up front: I&rsquo;m still learning this stuff. I&rsquo;m not writing this from a mountaintop. I&rsquo;m writing it from the foothills, with muddy boots, having just figured out something that I suspect a lot of people pretend they already knew.</p>
<p>Here&rsquo;s the thing that finally clicked for me. An agent is a loop. A model looks at the situation, decides one next step, calls a tool to do it, looks at what happened, and goes around again until it&rsquo;s done. That&rsquo;s it. I felt a little cheated when I understood it — the word &ldquo;agent&rdquo; had been doing so much heavy lifting on so many landing pages that I&rsquo;d assumed there was a fortress behind it. There isn&rsquo;t. There&rsquo;s a while-loop.</p>
<p>So I went and read about the frameworks. All of them — LangGraph, CrewAI, LlamaIndex, the OpenAI Agents SDK, Pydantic AI, smolagents, the Claude Agent SDK, the vendor SDKs from Google and Amazon and Microsoft. And every single one walks you through the same starter example: a weather bot. Or &ldquo;chat with your PDF.&rdquo; Or my personal favorite, the demo where five agents — a Researcher, a Writer, a Critic, an Editor, and presumably a Manager to schedule their standups — collaborate to produce <a href="/2025/10/ai-agent-mistakes/" class="lnp-link">a blog post slightly worse than one agent would&rsquo;ve written</a>.</p>
<p>And I kept thinking: <em>okay, but where are the real ones?</em></p>
<p>Not the demos. Not the quickstart. Something non-trivial. Something that acts on the world, where a wrong move costs money or breaks production. I genuinely couldn&rsquo;t picture one. So instead of pretending, I went looking.</p>
<p>(The method, since it&rsquo;s too on-the-nose not to mention: I sent a <a href="/2026/01/6-ai-agents-coding-experiment/" class="lnp-link">small swarm of research agents</a> out across the web to comb engineering blogs and case studies for me, in parallel, while I made coffee. Hunting for proof that real agents exist turned out to be the most real agent use I&rsquo;d touched all week. Make of that what you will.)</p>
<p>Here&rsquo;s what I actually found.</p>
<h2 id="the-good-news-real-ones-exist">The good news: real ones exist</h2>
<p>A few of them are unambiguously real, and they&rsquo;re worth describing, because they taught me more about what an agent is <em>for</em> than any framework doc did.</p>
<p><strong>Sentry&rsquo;s Autofix</strong> is the one that changed my mind. When something breaks in a codebase Sentry monitors, an agent built on the Claude Agent SDK takes their root-cause analysis, plans a fix, <em>writes the code</em>, and opens a pull request you can actually merge — a full run in about six minutes. This isn&rsquo;t a chatbot that suggests you &ldquo;consider checking your null values.&rdquo; It writes the patch. And it runs against a platform doing over a million root-cause analyses a year. One of their engineers shipped it in weeks and wrote a piece literally titled <a href="https://blog.sentry.io/how-sentrys-ai-autofix-changed-my-mind-about-ai-agents/" rel="external nofollow noopener" class="lnp-link"><em>how Sentry&rsquo;s AI Autofix changed my mind about AI agents</em></a>. I felt seen.</p>
<p><strong>Amazon has an internal agent that troubleshoots network failures</strong> — diagnoses live VPC connectivity problems and resolves around 80% of network root causes on its own. Built on their <a href="https://strandsagents.com/blog/what-we-learned-from-one-year-of-building-production-agents/" rel="external nofollow noopener" class="lnp-link">Strands SDK</a>. That&rsquo;s an on-call SRE&rsquo;s nightmare-shift, handed to a loop. As someone who&rsquo;s done that shift, that number did something to me.</p>
<p><strong>Coinbase built <a href="https://github.com/coinbase/agentkit" rel="external nofollow noopener" class="lnp-link">a toolkit that gives an agent a crypto wallet</a>.</strong> The agent can hold funds, sign transactions, and pay for things autonomously. Read that again. We&rsquo;ve spent this whole article saying the scary part of agents is irreversible action with real stakes — and here&rsquo;s one wired directly to money on a blockchain, where &ldquo;oops&rdquo; is permanent. Terrifying. Also clearly real.</p>
<p><strong>Bilt runs <a href="https://www.letta.com/case-studies/bilt" rel="external nofollow noopener" class="lnp-link">a <em>million</em> agents</a></strong> — one per user — on Letta, each holding that user&rsquo;s transaction and engagement history in <a href="/2025/11/ai-agent-memory-persistence/" class="lnp-link">persistent memory</a> to drive merchant recommendations. The whole pitch of Letta is memory, and here&rsquo;s someone betting a recommendation system on it at a scale I can&rsquo;t fully picture.</p>
<p>And a scattering more, each genuinely non-trivial: <a href="https://www.langchain.com/blog/top-5-langgraph-agents-in-production-2024" rel="external nofollow noopener" class="lnp-link">Exa&rsquo;s web-research agent and LinkedIn&rsquo;s text-to-SQL bot</a> (both on LangGraph, both acting against live production systems); a <a href="https://pydantic.dev/" rel="external nofollow noopener" class="lnp-link">medical-triage agent on Pydantic AI</a> validated across 329 clinician-checked scenarios; a <a href="https://www.llamaindex.ai/blog/case-study-tender-rfp-agent-for-construction-sector-with-softiq" rel="external nofollow noopener" class="lnp-link">construction-tender agent on LlamaIndex</a> that digests 100-page public bids and spits out risk reports; Uber automating code migrations across its monorepo.</p>
<p>So. Real agents exist. I can stop being a skeptic about <em>that</em>.</p>
<h2 id="the-uncomfortable-news-there-arent-many-and-the-vendors-are-grading-their-own-homework">The uncomfortable news: there aren&rsquo;t many, and the vendors are grading their own homework</h2>
<p>Here&rsquo;s the part that kept nagging me after the research came back.</p>
<p>For each framework, I could find maybe <strong>one to three</strong> genuinely non-trivial examples. Not dozens. Single digits. And almost every one of them was published by the company that <em>sells the framework.</em> Sentry&rsquo;s story is on Sentry&rsquo;s blog (fair enough — Sentry isn&rsquo;t Anthropic), but most of them live in the framework vendor&rsquo;s own marketing: LangChain&rsquo;s case-study page, Letta&rsquo;s case studies, AWS&rsquo;s own deep-dive, Google&rsquo;s own developer blog. Independent &ldquo;here&rsquo;s our war story and here&rsquo;s what broke&rdquo; write-ups from teams with no skin in the game? Vanishingly rare.</p>
<p>And some frameworks I genuinely <em>couldn&rsquo;t</em> find a real one for:</p>
<ul>
<li><strong>smolagents</strong> has <a href="https://github.com/huggingface/smolagents" rel="external nofollow noopener" class="lnp-link">26,000 GitHub stars</a> and I love its design — but its flagship example is Hugging Face&rsquo;s own research replication. I found no named company betting anything real on it.</li>
<li><strong>CrewAI</strong> is everywhere in demos and has a wall of enterprise logos (PepsiCo, J&amp;J, the DoD), but behind almost every logo is zero operational detail. The one solid story — <a href="https://blog.crewai.com/lessons-from-2-billion-agentic-workflows/" rel="external nofollow noopener" class="lnp-link">a five-agent sales pipeline at DocuSign</a> — is, again, on CrewAI&rsquo;s own blog.</li>
<li><strong>Microsoft&rsquo;s Agent Framework</strong> just hit 1.0 claiming &ldquo;real-world validation with customers and partners&rdquo; and then named exactly zero of them. Its most impressive artifact, <a href="https://www.microsoft.com/en-us/research/articles/magentic-one-a-generalist-multi-agent-system-for-solving-complex-tasks/" rel="external nofollow noopener" class="lnp-link">Magentic-One</a>, is explicitly a <em>research</em> system that doesn&rsquo;t ship inside a product.</li>
</ul>
<p>I want to be careful here, because I&rsquo;m still learning and I don&rsquo;t want to overclaim the cynicism: &ldquo;I couldn&rsquo;t find it&rdquo; is not &ldquo;it doesn&rsquo;t exist.&rdquo; A lot of the realest agent work is surely locked inside companies that will never blog about it. But the <em>public</em> record, right now, is thin. Much thinner than the hype implied. The ratio of &ldquo;agentic platform&rdquo; marketing to &ldquo;here is a real agent doing a real job&rdquo; is grim.</p>
<h2 id="two-things-i-think-im-learning">Two things I think I&rsquo;m learning</h2>
<p>I&rsquo;m holding these loosely, because foothills. But:</p>
<p><strong>The best real agents are vendors using their own tools.</strong> Amazon&rsquo;s network agent, <a href="https://developers.googleblog.com/en/agent-development-kit-easy-to-build-multi-agent-applications/" rel="external nofollow noopener" class="lnp-link">Google&rsquo;s enterprise agents on ADK</a>, Strands originating inside Amazon Q Developer — the most concrete, number-backed cases are companies dogfooding the framework they built. That&rsquo;s either reassuring (they believe in it enough to run it) or a little hollow (of course the toolmaker has the best tool demo). Probably both.</p>
<p><strong>Every real one acts. None of them chat.</strong> This is the pattern that actually reorganized my thinking. Line up the genuinely non-trivial agents — writes a mergeable PR, signs a transaction, resolves a network outage, holds a million users&rsquo; memory, files a risk report on a 100-page tender. Not one of them is a conversation. The toys all talk. The real ones <em>do</em>. The demos cluster around chat because chat is safe and reversible and impresses in a screenshot. The real ones cluster around irreversible action because that&rsquo;s where an agent is actually worth the risk of building.</p>
<p>Which, looping all the way back, is exactly why the weather bot felt so empty. A weather bot doesn&rsquo;t <em>do</em> anything. It&rsquo;s the loop with the stakes amputated.</p>
<h2 id="so-where-does-that-leave-a-beginner">So where does that leave a beginner</h2>
<p>I don&rsquo;t have a grand conclusion. I have a working hypothesis, which is the most an honest learner should claim: the framework you pick matters far less than whether you have a real job that needs an agent that <em>acts</em>. If you don&rsquo;t, no framework will save you — you&rsquo;ll build a five-agent demo and quietly stop opening the repo. If you do, the loop is twenty lines, and you should start with whichever framework hides the least so you can actually see what&rsquo;s happening (smolagents, the OpenAI Agents SDK, and Pydantic AI were the ones that got out of my way the most).</p>
<p>And honestly? The fact that real examples are still this rare didn&rsquo;t discourage me. It read like a timestamp. We&rsquo;re early. The scarcity isn&rsquo;t proof the idea is empty — it&rsquo;s proof most people are still building weather bots while a handful of teams quietly wire a loop up to something that matters.</p>
<p>I&rsquo;d rather be in the second group — which is why I&rsquo;m slowly <a href="/2026/01/agent-orchestrator/" class="lnp-link">building one of my own</a>. I&rsquo;m still learning how.</p>
<hr>
<h2 id="the-ledger-the-realest-example-i-found-per-framework-and-where-its-published">The ledger (the realest example I found per framework, and where it&rsquo;s published)</h2>
<p><em>Honest tag: most of these are vendor-published. Independent confirmation is scarce — which is part of the story.</em></p>
<ul>
<li><strong>Claude Agent SDK</strong> — Sentry Autofix: writes mergeable PRs against 1M+ RCAs/yr → <a href="https://blog.sentry.io/how-sentrys-ai-autofix-changed-my-mind-about-ai-agents/" rel="external nofollow noopener" class="lnp-link">blog.sentry.io</a>, <a href="https://claude.com/customers/sentry" rel="external nofollow noopener" class="lnp-link">claude.com/customers/sentry</a></li>
<li><strong>AWS Strands</strong> — Amazon internal network-troubleshooting agent (~80% of network root causes); origin of Amazon Q Developer → <a href="https://strandsagents.com/blog/what-we-learned-from-one-year-of-building-production-agents/" rel="external nofollow noopener" class="lnp-link">strandsagents.com</a></li>
<li><strong>Letta</strong> — Bilt: ~1M per-user memory agents for recommendations → <a href="https://www.letta.com/case-studies/bilt" rel="external nofollow noopener" class="lnp-link">letta.com/case-studies/bilt</a></li>
<li><strong>OpenAI Agents SDK</strong> — Coinbase AgentKit: agents with on-chain wallets, real transactions → <a href="https://github.com/coinbase/agentkit" rel="external nofollow noopener" class="lnp-link">github.com/coinbase/agentkit</a></li>
<li><strong>LangGraph</strong> — Exa web-research agent; LinkedIn text-to-SQL bot; Uber code migrations → <a href="https://www.langchain.com/blog/exa" rel="external nofollow noopener" class="lnp-link">langchain.com/blog/exa</a>, <a href="https://www.langchain.com/blog/top-5-langgraph-agents-in-production-2024" rel="external nofollow noopener" class="lnp-link">top-5 in production</a></li>
<li><strong>Pydantic AI</strong> — STCC medical-triage agentic RAG (329 validated scenarios) → <a href="https://pydantic.dev/" rel="external nofollow noopener" class="lnp-link">pydantic.dev</a></li>
<li><strong>LlamaIndex</strong> — SoftIQ construction-tender agent (100-page bids → risk reports) → <a href="https://www.llamaindex.ai/blog/case-study-tender-rfp-agent-for-construction-sector-with-softiq" rel="external nofollow noopener" class="lnp-link">llamaindex.ai case study</a></li>
<li><strong>Google ADK</strong> — Google&rsquo;s own Agentspace/contact-center agents (6T+ tokens/mo); Renault EV-charger siting; Box contract extraction → <a href="https://developers.googleblog.com/en/agent-development-kit-easy-to-build-multi-agent-applications/" rel="external nofollow noopener" class="lnp-link">developers.googleblog.com</a></li>
<li><strong>CrewAI</strong> — DocuSign 5-agent sales Flow (vendor blog) → <a href="https://blog.crewai.com/lessons-from-2-billion-agentic-workflows/" rel="external nofollow noopener" class="lnp-link">blog.crewai.com</a></li>
<li><strong>smolagents</strong> — no named production company found; flagship is HF&rsquo;s own Open Deep Research → <a href="https://github.com/huggingface/smolagents" rel="external nofollow noopener" class="lnp-link">github.com/huggingface/smolagents</a></li>
<li><strong>Microsoft Agent Framework / AutoGen</strong> — mostly research (Magentic-One); 1.0 names zero customers → <a href="https://www.microsoft.com/en-us/research/articles/magentic-one-a-generalist-multi-agent-system-for-solving-complex-tasks/" rel="external nofollow noopener" class="lnp-link">microsoft.com/research</a></li>
</ul>
]]></content:encoded></item><item><title>Five Books Taught Me to Build AI Agents. All Five Quietly Told Me Not To.</title><link>https://lakshminp.com/2026/06/what-ai-agents-actually-are/</link><pubDate>Sun, 21 Jun 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/06/what-ai-agents-actually-are/</guid><category>essays</category><category>ai-coding</category><category>ai-agents</category><description>What four hundred thousand words of agent literature agree on — and never put on the cover.
I bought five books on building AI agents in a single afternoon, the way you panic-buy bottled water before a storm. Manning had a sale. I had a credit card and a vague sense that everyone around me had quietly become an “agent engineer” while I was busy doing my actual job.
So I did the responsible thing. I spun up a small army of subagents to read four of them for me, cover to cover, in parallel, and report back. Which, if you’re keeping score, means I built a multi-agent system to summarize books about how to build multi-agent systems. The irony was not lost on me. It was, in fact, the first thing I learned.</description><content:encoded><![CDATA[<p><em>What four hundred thousand words of agent literature agree on — and never put on the cover.</em></p>
<p>I bought five books on building AI agents in a single afternoon, the way you panic-buy bottled water before a storm. Manning had a sale. I had a credit card and a vague sense that everyone around me had quietly become an &ldquo;agent engineer&rdquo; while I was busy doing my actual job.</p>
<p>So I did the responsible thing. I spun up a small army of subagents to read four of them for me, cover to cover, in parallel, and report back. Which, if you&rsquo;re keeping score, means I built a multi-agent system to summarize books about how to build multi-agent systems. The irony was not lost on me. It was, in fact, the first thing I learned.</p>
<p>Here&rsquo;s the second.</p>
<h2 id="the-loop-is-thirty-lines">The loop is thirty lines</h2>
<p>Strip away the diagrams and the framework comparisons, and every single one of these books — <a href="https://www.manning.com/books/build-an-ai-agent-from-scratch" rel="external nofollow noopener" class="lnp-link">Build an AI Agent</a>, <a href="https://www.manning.com/books/build-a-multi-agent-system-from-scratch" rel="external nofollow noopener" class="lnp-link">Build a Multi-Agent System</a>, <a href="https://www.manning.com/books/ai-agents-in-action-second-edition" rel="external nofollow noopener" class="lnp-link">AI Agents in Action</a>, <a href="https://www.manning.com/books/ai-agents-and-applications" rel="external nofollow noopener" class="lnp-link">AI Agents and Applications</a> — converges on the same humble definition.</p>
<p>An agent is a language model, plus some tools, plus a loop that runs until the job is done.</p>
<p>That&rsquo;s it. One book states it as plainly as that. Another dresses it up as a four-letter cycle. There&rsquo;s a Reddit thread floating around that implements the whole thing in about thirty lines of code, set to a drum-and-bass track, and honestly it explains the concept better than half the chapters I read.</p>
<p>There is no secret sauce. There is no priesthood. You were promised a cathedral and what you got is a <code>while</code> loop with good manners.</p>
<p>Which raised an obvious question, sitting there with four hundred thousand words of agent literature on my screen: if the core idea fits on a napkin, what&rsquo;s in all these books?</p>
<h2 id="the-part-nobody-puts-on-the-cover">The part nobody puts on the cover</h2>
<p>The answer is the same in every one, and it&rsquo;s the most useful thing I took away.</p>
<p>The loop is the easy ten percent. The other ninety — the part that doesn&rsquo;t fit in a demo — is evaluation, memory, guardrails, cost control, defending against prompt injection, and the deeply unglamorous skill of knowing when to hand the problem back to a human.</p>
<p>Three of the four books I read point at the same Anthropic paper, &ldquo;Building Effective Agents,&rdquo; like it&rsquo;s scripture. And buried in chapter one of each — past the exciting cover, past the part where they sell you on the future — every author tells you the same quiet thing.</p>
<p>Don&rsquo;t reach for an agent.</p>
<p>Start with a plain model call. Then a chain. Then a workflow. Earn the agent only when the task genuinely needs one, because an agent costs roughly ten times a normal call. Per task. Now imagine that thing running unattended, all night, while you sleep.</p>
<p>I went looking for the loudest voices on the other side of this — the practitioners on Reddit who build agents for a living and have the scar tissue to prove it. I expected an argument. The top thread is literally titled &ldquo;Stop building AI agents.&rdquo; Another is a guy who got <em>paid</em> to rip the AI back out of a tool he&rsquo;d shipped. A third is the 2 a.m. classic: the agent hit a question it didn&rsquo;t understand, confidently made up an answer, and emailed it to a customer.</p>
<p>The books and the burnouts weren&rsquo;t arguing. They&rsquo;d arrived at the same conclusion from opposite ends of the room. The model was never the bottleneck. Running the thing was.</p>
<h2 id="what-im-actually-taking-away">What I&rsquo;m actually taking away</h2>
<p>A small tell that stuck with me: agent-to-agent coordination shows up in the <em>subtitles</em> of these books far more confidently than it shows up in the chapters. The field is writing about how agents talk to each other a little faster than it&rsquo;s shipping it. That&rsquo;s not a knock — it&rsquo;s a map. It tells you where the hype is and where the ground is still wet.</p>
<p>So here&rsquo;s my take, for whatever a guy who outsourced his reading to robots is worth.</p>
<p>The framework you pick doesn&rsquo;t matter much; that code rots in eighteen months. What compounds is the boring stuff the demos skip — evaluation, context discipline, and the judgment to not build the agent at all. Everyone is rushing to learn how to <em>make</em> an agent. Almost nobody is learning how to make one you&rsquo;d actually trust.</p>
<p>The capability got democratized this year. The judgment didn&rsquo;t.</p>
<p>That gap — between the agent that runs and the agent you&rsquo;d let near production while you&rsquo;re asleep — is the whole job now. It&rsquo;s also, conveniently, <a href="/2025/10/the-real-skill-ai-wont-replace/" class="lnp-link">the only part worth getting good at</a>.</p>
]]></content:encoded></item><item><title>The "MCP Is Dead" Fight Is a Category Error</title><link>https://lakshminp.com/2026/06/mcp-vs-skills-category-error/</link><pubDate>Thu, 18 Jun 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/06/mcp-vs-skills-category-error/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><description>Skills win the solo dev. MCP wins exactly one thing. Here’s the line.
I had the headline before I had the post.
“API + Skills Is a Poor Man’s MCP” — except I was going to argue the inversion: that MCP is the rich man’s overcomplication, a server process you stood up to wrap calls your agent could already make, and the lean move was always a skill plus a CLI. Spicy. Contrarian. The kind of take that does numbers in a feed.</description><content:encoded><![CDATA[<p><em>Skills win the solo dev. MCP wins exactly one thing. Here&rsquo;s the line.</em></p>
<p>I had the headline before I had the post.</p>
<p>&ldquo;API + Skills Is a Poor Man&rsquo;s MCP&rdquo; — except I was going to argue the inversion: that <em>MCP</em> is the rich man&rsquo;s overcomplication, a server process you stood up to wrap calls your agent could already make, and the lean move was always a skill plus a CLI. Spicy. Contrarian. The kind of take that does numbers in a feed.</p>
<p>Then I made the tactical error of fact-checking myself, and the post fell apart in my hands. What follows is the wreckage, reassembled into something truer than the dunk I wanted to write.</p>
<p><strong>TL;DR:</strong> Skills and MCP aren&rsquo;t competitors — comparing them is a category error. A skill is a recipe that runs in <em>your</em> runtime; MCP is a connection to a hosted service. For a solo dev wiring up their own workflow on a coding agent, skill + CLI wins on every axis that used to favor MCP — the context-bloat and cross-vendor arguments both got quietly erased in late 2025/early 2026. MCP earns its keep in exactly one situation: you&rsquo;re a <em>provider</em> exposing a live, OAuth&rsquo;d service to assistants you don&rsquo;t own. The rule that falls out: <strong>consuming an API → skill + CLI. Providing a service → MCP.</strong></p>
<h2 id="the-category-error-i-was-about-to-commit">The category error I was about to commit</h2>
<p>The first crack: &ldquo;API + skills vs MCP&rdquo; quietly assumes the two live on the same shelf. They don&rsquo;t.</p>
<p>A <strong>skill</strong> is knowledge. A markdown recipe — plus maybe a script — that teaches the agent how to do something, running in <em>your</em> runtime, on <em>your</em> machine, with tools you already have.</p>
<p><strong>MCP</strong> is a connection to a running service. A server, behind a protocol, that the model talks to.</p>
<p>Comparing them is apples to orchards. One is &ldquo;here&rsquo;s how, go do it.&rdquo; The other is &ldquo;here&rsquo;s a thing that&rsquo;s already running, call it.&rdquo; Most of the internet argues about them as if they&rsquo;re competing products. They&rsquo;re not even the same noun.</p>
<p>So the honest question isn&rsquo;t &ldquo;which wins.&rdquo; It&rsquo;s &ldquo;when does a hosted service behind a contract beat a recipe you run yourself?&rdquo; That&rsquo;s a real question. I just assumed I knew the answer.</p>
<h2 id="the-two-arguments-that-died-before-i-finished-typing">The two arguments that died before I finished typing</h2>
<p>My case against MCP rested on two pillars. Both had already collapsed, and I hadn&rsquo;t noticed.</p>
<p><strong>Pillar one: context bloat.</strong> Every MCP server dumps all its tool schemas into the context window — a seven-server setup could eat 67K tokens before you typed a word, and <a href="https://lakshminp.com/2025/11/ai-agent-memory-persistence/" class="lnp-link">the context window is the one resource your agent can&rsquo;t buy back</a>. Damning. Except in January 2026 Anthropic shipped tool search and <code>defer_loading</code>: now the model sees a search tool plus a couple of always-on tools, and pulls the rest on demand. Reported reductions of 85–95%. My killer stat became a &ldquo;this used to be true.&rdquo;</p>
<p><strong>Pillar two: cross-vendor reach is MCP&rsquo;s moat.</strong> Wrong by a different calendar. In December 2025, Agent Skills shipped as an open standard, and within 48 hours Microsoft put it in VS Code and OpenAI added it to ChatGPT and Codex. By spring, ~40 tools — Gemini CLI, JetBrains, Kiro, Goose — read the same <code>SKILL.md</code>. Skills are as universal as MCP now. The moat drained while I was sharpening my knives.</p>
<p>Fine. Two pillars down. The dunk still had three legs, I figured.</p>
<h2 id="watching-the-rest-fall">Watching the rest fall</h2>
<p><strong>Security?</strong> I&rsquo;d claimed MCP gives you a safety edge. It doesn&rsquo;t. If I want read-only GitHub access, I hand a read-only token to the skill <em>or</em> the MCP server — identical. The token scope is the gate, enforced at the API boundary, available to both. There&rsquo;s no protocol-level security advantage. Gone.</p>
<p><strong>Tokens?</strong> This one inverts, which delighted me until I realized it cut against my own thesis too. People assume MCP is token-cheap because the call — <code>list_pull_requests(owner, repo)</code> — is tidy. But the <em>call</em> isn&rsquo;t the cost. The <em>result</em> is. The GitHub API returns fat JSON, and a raw MCP tool call dumps the whole blob into context. A skill that runs code can filter in the sandbox and return five lines. So code-that-filters wins on tokens — and a skill is code-that-filters by birth. But that&rsquo;s an argument for skills, not against MCP-the-idea.</p>
<p><strong>Auto-orchestration?</strong> &ldquo;MCP composes calls for you.&rdquo; No, it doesn&rsquo;t. The protocol is transport — it has no &ldquo;run this sequence, give me only the end&rdquo; primitive. Either the model loops (every intermediate result round-trips through context — expensive) or a human pre-bakes a coarse server tool (effort). Automatic, token-cheap stacking only happens when you call tools <em>from code</em> — which is, once again, the skill model.</p>
<p>Every road kept leading back to the same place. I started to feel like the universe was trying to tell me something(sounds dramatic, I know).</p>
<h2 id="the-litmus-test-that-almost-saved-the-dunk">The litmus test that almost saved the dunk</h2>
<p>So I built a concrete test: <em>&ldquo;Fetch all open PRs, give me a gist of the modules they touch, merge only the ones tagged auth.&rdquo;</em></p>
<p>Fetch and gist are reads — data-heavy aggregation, the code-that-filters sweet spot. Skill wins, easily. But <em>merge</em> is a write, and a dangerous one, and writes are where I figured MCP&rsquo;s permission policy — pause and confirm each merge — would finally earn its keep.</p>
<p>Then a reader on the thread that became this post pointed out the obvious: gating decomposes into three questions, and only one is even arguably MCP&rsquo;s.</p>
<ul>
<li><strong>Capability</strong> — can a merge happen at all? The <em>token scope</em> answers that. Available to both. API-enforced.</li>
<li><strong>Selection</strong> — which PRs get merged? Your <em>filtering code</em> answers that. That&rsquo;s the skill&rsquo;s script.</li>
<li><strong>Confirmation</strong> — do you approve each one? <em>You</em>, in the loop on a coding agent — or a <code>--confirm</code> flag — or MCP&rsquo;s native prompt.</li>
</ul>
<p>Only the third row is MCP&rsquo;s, and even there it&rsquo;s matched by you-watching-bash or a confirm flag. On a coding agent with you present, skill plus the <code>gh</code> CLI wins the whole task. The merge didn&rsquo;t flip it. <em>You&rsquo;re</em> the permission policy.</p>
<p>That was the moment the dunk officially died. I went looking on Reddit to see who else had buried it.</p>
<h2 id="what-reddit-already-knew">What Reddit already knew</h2>
<p>Turns out, everyone. The threads are a graveyard with two opposing headstones.</p>
<p><a href="https://www.reddit.com/r/ClaudeCode/comments/1rrl56g/" rel="external nofollow noopener" class="lnp-link">&ldquo;Will MCP be dead soon?&rdquo;</a> — 406 comments. <a href="https://www.reddit.com/r/ClaudeAI/comments/1pjpbji/" rel="external nofollow noopener" class="lnp-link">&ldquo;I cannot, for the life of me, understand the value of MCPs&rdquo;</a> — 305 comments. <a href="https://www.reddit.com/r/mcp/comments/1rstpfk/" rel="external nofollow noopener" class="lnp-link">&ldquo;A eulogy for MCP (RIP).&rdquo;</a> <a href="https://www.reddit.com/r/mcp/comments/1o8w5wq/" rel="external nofollow noopener" class="lnp-link">&ldquo;CLI &gt; MCP?&rdquo;</a> Someone even shipped a tool that converts MCP servers into CLI + skill files and &ldquo;cut ~97% token overhead.&rdquo;</p>
<p>The auto-generated TL;DR of that 305-comment thread is, embarrassingly, the post I&rsquo;d spent a day reverse-engineering: <em>&ldquo;You&rsquo;re looking at this from a solo dev&rsquo;s perspective, and you&rsquo;re not wrong — Skills or telling Claude to use a CLI is often more efficient. MCP&rsquo;s real value isn&rsquo;t for your individual coding session, but for the broader ecosystem.&rdquo;</em></p>
<p>So the solo-dev case is settled. But the people defending MCP weren&rsquo;t demo-app tourists. They were running it in production, and they landed one punch I couldn&rsquo;t slip.</p>
<h2 id="the-punch-i-couldnt-slip-and-the-one-real-win">The punch I couldn&rsquo;t slip, and the one real win</h2>
<p>From the eulogy thread, top comment: <em>&ldquo;People who claim there&rsquo;s no need for MCP will, if they build projects of growing complexity, sooner or later reinvent everything MCP provides — but bespoke and non-standardized.&rdquo;</em> And a sharper one: <em>&ldquo;CLI + skills are great for solo dev vibes. But the second you need an LLM to orchestrate across multiple platforms with real auth and governance? You&rsquo;re either using MCP or rebuilding it badly.&rdquo;</em></p>
<p>That&rsquo;s the one thing that survived every round. Not context, not security, not reach, not tokens. <strong>Distribution — of a specific kind.</strong></p>
<p>Here&rsquo;s the case, and it&rsquo;s narrower than the hype and realer than my dunk: you&rsquo;re a <em>provider</em>. You host a live service and you want it to show up as a one-click, OAuth&rsquo;d connector inside every AI assistant your customers already use — Claude&rsquo;s Connectors Directory (200+ integrations), ChatGPT Apps, all of it. Notion, Linear, and Stripe ship official remote MCP servers for exactly this. You build once; it lights up everywhere; the credentials and compute stay on your side.</p>
<p>A skill — even a universal one — cannot be that. A skill is a copy that runs in the consumer&rsquo;s runtime, and that&rsquo;s the whole limitation: it can only do what that runtime can do. Flip that around and you get the same win seen from the client&rsquo;s side. Claude Desktop runs skills <em>and</em> has a code sandbox — but the sandbox can&rsquo;t reach the open internet, so a skill that tries to curl GitHub dies at the egress wall, while an MCP server, running outside the box, reaches it fine. Same task, opposite answer, deciding variable is network egress. It&rsquo;s not a second reason to use MCP. It&rsquo;s the first reason wearing a different hat: when the consumer&rsquo;s runtime can&rsquo;t reach the thing, you need a service that can.</p>
<h2 id="the-rule-when-to-use-mcp-vs-a-skill--cli">The rule: when to use MCP vs a skill + CLI</h2>
<p>So, do you need MCP? Ask one question: <strong>are you consuming an API, or providing a service?</strong></p>
<p>Wiring your own agent to someone&rsquo;s existing API, on a coding tool with a shell — skill plus a CLI, every time. You will not &ldquo;reinvent MCP badly,&rdquo; because you need none of what MCP provides: no OAuth dance, no dynamic discovery, no cross-client reach. The CLI is complete, not a degenerate clone.</p>
<p>Exposing your own live service to assistants you don&rsquo;t own, with auth and governance, across vendors — that&rsquo;s MCP, and a CLI genuinely can&rsquo;t do it.</p>
<p>MCP isn&rsquo;t the poor man&rsquo;s anything. It&rsquo;s the <em>platform&rsquo;s</em> protocol. You reach for it the moment you stop consuming APIs and start being an app inside other people&rsquo;s assistants.</p>
<p>I know which side I&rsquo;m on this week. I shipped ThreadHQ&rsquo;s MCP server as a top-of-funnel for a reason — that&rsquo;s the provider play, done on purpose. (ThreadHQ is one of the products I <a href="https://lakshminp.com/2026/06/build-saas-with-claude-code/" class="lnp-link">built solo with Claude Code</a>; the MCP server is its distribution edge, not its plumbing.) But for the GitHub task on my own machine? I&rsquo;m still just typing <code>gh</code>. The dunk was wrong. The honest version is sharper anyway.</p>
]]></content:encoded></item><item><title>How to Build a SaaS with Claude Code in a Weekend (Not a Quarter)</title><link>https://lakshminp.com/2026/06/build-saas-with-claude-code/</link><pubDate>Mon, 15 Jun 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/06/build-saas-with-claude-code/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><category>saas</category><description>I caught up with one of my mentees last weekend. Just talking shop. He’s a solid developer, he’s got the itch to build a side project, and he’s brand new to agentic coding. Somewhere in that conversation I realized I was reciting an entire playbook off the top of my head — so I’m writing it down. If you can write code and you’re staring at Claude Code wondering how to actually build and ship a SaaS with it, this is for you.</description><content:encoded><![CDATA[<p>I caught up with one of my mentees last weekend. Just talking shop. He&rsquo;s a solid developer, he&rsquo;s got the itch to build a side project, and he&rsquo;s brand new to agentic coding. Somewhere in that conversation I realized I was reciting an entire playbook off the top of my head — so I&rsquo;m writing it down. If you can write code and you&rsquo;re staring at Claude Code wondering how to actually build and ship a SaaS with it, this is for you.</p>
<p>A quick word on <em>why now</em>, before the <em>how</em>. Two reasons, and you&rsquo;ve heard at least one:</p>
<ol>
<li>AI is eating the jobs. You can&rsquo;t open your phone without someone reminding you. Enough said.</li>
<li>The AI subsidy is going to end. Right now you&rsquo;re building on frontier models that cost the labs more than they charge you. That window closes. <a href="/2026/03/ai-subsidy-window-developers/" class="lnp-link">I wrote about this here</a></li>
</ol>
<p>Maybe both happen. Either way the move is the same: build the muscle now, while it&rsquo;s cheap and you still have an edge.</p>
<h2 id="1-what-to-build">1. What to build</h2>
<p>Scratch your own itch. Do not spend three weeks &ldquo;researching the market&rdquo; to discover what people want. If you have a problem an app would fix, give yourself permission to build it. We&rsquo;ll worry about market size <em>after</em> it ships.</p>
<p>Market research used to be load-bearing because building was expensive and being wrong was catastrophic — months of work, real money, dead on arrival. That math is gone. You can ship an app over a weekend now. The cost of being wrong is one wasted Saturday.</p>
<p>Paul Graham makes the sharper version of this point: build what <em>you</em> want, because — as he writes in <a href="https://paulgraham.com/earn.html" rel="external nofollow noopener" class="lnp-link">How to Earn a Billion Dollars</a> — &ldquo;your own needs are uniquely valuable, because your needs predict future demand.&rdquo; You&rsquo;re not guessing what strangers want. You&rsquo;re scratching an itch you can actually feel.</p>
<p>I&rsquo;ve written about finding an idea and shipping it in a single sitting — <a href="/2026/01/claude-code-ship-one-session/" class="lnp-link">here</a>.</p>
<h2 id="2-know-where-youre-standing">2. Know where you&rsquo;re standing</h2>
<p>Be honest about your starting point: do you have real development experience, or do you need to ramp up first? <a href="/2025/10/ai-coding-prerequisites/" class="lnp-link">Here&rsquo;s what I&rsquo;d master before letting AI build for you</a></p>
<p>Yes, there are people on the internet shipping apps with zero coding background. Maybe. But if you can&rsquo;t read what Claude Code writes, you can&rsquo;t steer it — and you&rsquo;ll feel that the first time it confidently drives into a wall. You don&rsquo;t need a CS degree. You need enough of a mental model to call BS. The good news: you can learn development <em>from</em> Claude Code while you build <em>with</em> it. Learning and building at the same time is completely legitimate — it&rsquo;s how I pick up half the things I use.</p>
<h2 id="3-how-i-actually-build">3. How I actually build</h2>
<p>There&rsquo;s no one true way. This is what works for me; your mileage may vary.</p>
<p>First, the unglamorous part: get the Claude Code Max plan. The $100 tier, or the $200 one if you can swing it. You cannot build anything meaningful on the cheap plans, and you&rsquo;ll discover exactly why about four hours into your first real session. Don&rsquo;t flinch at the price — it&rsquo;s still cheaper than a tutor or a freelancer, and it doesn&rsquo;t take lunch breaks. Plus the quality is consistently good.</p>
<h3 id="why-claude-code-and-not-the-benchmark-topping-model-of-the-week">Why Claude Code and not [the benchmark-topping model of the week]?</h3>
<p>Because building a startup is already exhausting, and you do not have spare energy to spend benchmarking models and sharpening tools instead of shipping. Pick what works and stick with it.</p>
<p>Codex is a close second — genuinely good, and I keep it around. But Claude Code stays a step ahead for actually building apps, and after working with both, I reach for it first. Use both if you like. Just don&rsquo;t turn tool selection into the project.</p>
<h3 id="boring-choices-win">Boring choices win</h3>
<p>Freeze your stack early — backend, frontend, database. 80% of it is identical across every app you&rsquo;ll ever build, so stop re-deciding it every time. And don&rsquo;t obsess over scale and optimization. Those are problems you <em>earn</em> by being successful. Good problems. You don&rsquo;t have them yet.</p>
<h2 id="4-specs-and-context-engineering-the-part-that-decides-everything">4. Specs and context engineering: the part that decides everything</h2>
<p>Two things determine whether you get your money&rsquo;s worth out of Claude Code:</p>
<ol>
<li>The specification</li>
<li>Context engineering</li>
</ol>
<h3 id="the-spec">The spec</h3>
<p>The more specific you are, the better the output. &ldquo;Build a to-do app&rdquo; gets you slop. &ldquo;Here&rsquo;s the auth flow, here&rsquo;s the data model, here are the exact features&rdquo; gets you something you can use. Spend real time here, before a single line of code gets written. <a href="/2025/12/stop-making-claude-code-guess/" class="lnp-link">More on this</a></p>
<p>The move that works best: make Claude Code interview <em>you</em> about the spec. If you can&rsquo;t answer its questions, you can&rsquo;t articulate the feature — and if you can&rsquo;t articulate it, you can&rsquo;t build it. By the time the spec is done, the MVP should have zero grey areas.</p>
<h3 id="context-engineering">Context engineering</h3>
<p>Even the best frontier model starts coding like it&rsquo;s three drinks deep once the context fills up. So you manage it.</p>
<p>This takes me back to my assembly-language days — limited registers, limited memory, every instruction written with one eye on the resources you didn&rsquo;t have. LLM context is that same constraint in new clothes. You get roughly 200k tokens, and that&rsquo;s nowhere near enough to hold your whole app in its head at once.</p>
<p>So: one task per session. Two at the absolute most. Which means breaking the spec into session-sized tasks and tracking what&rsquo;s done, what&rsquo;s in flight, what&rsquo;s blocked, and what depends on what.</p>
<p>A markdown to-do file is a terrible way to do this and a worse use of your time. I use <strong>beads</strong>. Adopted it early, still on it. It&rsquo;s the fix for <a href="/2025/11/ai-agent-memory-persistence/" class="lnp-link">an agent that wakes up every morning with no memory of what you did yesterday</a>.</p>
<p>This practice is also a lot kinder for your token limits.</p>
<p>Two flavors:</p>
<ul>
<li>the original, by the author — <a href="https://github.com/gastownhall/beads" rel="external nofollow noopener" class="lnp-link">gastownhall/beads</a></li>
<li><strong>beads-rust</strong>, a simpler, more stable reimplementation of the spec in Rust — <a href="https://github.com/Dicklesworthstone/beads_rust" rel="external nofollow noopener" class="lnp-link">Dicklesworthstone/beads_rust</a></li>
</ul>
<p>Use either. I landed on beads-rust. Pick your poison.</p>
<p>You also need memory <em>across</em> sessions. You&rsquo;ll remember you fixed a bug two weeks ago; the model in today&rsquo;s session won&rsquo;t, and it&rsquo;ll cheerfully hand you a wrong answer when you ask &ldquo;did we already fix this?&rdquo; I use <strong>claude-mem</strong> for that — <a href="https://github.com/thedotmack/claude-mem" rel="external nofollow noopener" class="lnp-link">thedotmack/claude-mem</a>.</p>
<p>Point a Claude Code session at both repos and it&rsquo;ll install them for you. (Yes, these work with Codex too — but again: shipping or tuning? The clock is running.)</p>
<p>Remember that spec? Hand it to your beads skill and it breaks down into a clean task list. Ask Claude to pull the high-leverage beads and start there — never more than two in flight.</p>
<h3 id="your-claudemd-matters">Your CLAUDE.md matters</h3>
<p>Treat it as a compass, not a second spec. The practices that matter (write the tests first), how the app deploys, the handful of goals you&rsquo;re aiming at. Keep it light — <a href="/2026/04/claude-md-best-practices/" class="lnp-link">cramming a novel into CLAUDE.md is the fastest way to make Claude dumber</a>.</p>
<h3 id="one-more-tool">One more tool</h3>
<p><a href="https://github.com/sirmalloc/ccstatusline" rel="external nofollow noopener" class="lnp-link">ccstatusline</a>. Configure it to show your remaining context percentage. It&rsquo;s the fuel gauge that tells you whether to keep driving or pull over and start a fresh session.</p>
<p>Then it&rsquo;s rinse and repeat: feed beads to Claude, do the manual QA yourself, close the bead, next one. A few sessions in, a sliver of a working MVP starts to emerge. When you&rsquo;re happy with it, you ship.</p>
<h2 id="5-deploying-it-without-the-kubernetes-tax">5. Deploying it (without the Kubernetes tax)</h2>
<p>I was a Kubernetes guy for years. Deployed everything on it. I no longer recommend it for solo developers — <a href="/2025/12/kubernetes-indie-dev-alternative/" class="lnp-link">I explain why here</a>. It still has its place and time; your weekend project is neither.</p>
<p>Use <strong><a href="https://kamal-deploy.org" rel="external nofollow noopener" class="lnp-link">Kamal</a></strong> instead. Think of it as the compromise between Docker Compose and Kubernetes — Compose&rsquo;s simplicity, enough of Kubernetes&rsquo; robustness, none of the YAML despair.</p>
<p>I&rsquo;m also building <a href="https://vmkit.dev/" rel="external nofollow noopener" class="lnp-link">VMKit</a> to make this part disappear entirely — deploy without learning the nitty-gritty unless you want to.</p>
<p>Then there&rsquo;s the wiring you don&rsquo;t think about until it bites: monitoring and ops (I run mine through MCPs — <a href="/2026/01/30-dollar-saas-stack/" class="lnp-link">the $30 stack</a>), payments (Stripe; if you&rsquo;re in India, Dodo Payments), and distribution — marketing and positioning, which is a beast all its own. Each of these deserves its own post.</p>
<h2 id="tldr">TL;DR</h2>
<ul>
<li>Build for your own itch. Shipping is cheaper than market research now.</li>
<li>Get the Claude Code Max plan ($100, ideally $200). Don&rsquo;t tool-shop.</li>
<li>Spend your time on the spec — let Claude interview you until there are no grey areas.</li>
<li>Engineer your context: one task per session, tracked with beads, remembered with claude-mem.</li>
<li>Keep CLAUDE.md light. Watch your context gauge.</li>
<li>Deploy with Kamal, not Kubernetes. Wire payments and monitoring last.</li>
<li>The whole thing is a weekend, not a quarter.</li>
</ul>
]]></content:encoded></item><item><title>Claude Overreaches. Codex Underreaches. I'm Still Figuring Out How to Use Both.</title><link>https://lakshminp.com/2026/04/claude-vs-codex-use-both/</link><pubDate>Wed, 22 Apr 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/04/claude-vs-codex-use-both/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><description>I was a one-agent guy until Claude had a run of outages.
On those days I didn’t ship less. I shipped nothing. I’d open my editor, remember Claude was down, stare at the codebase, close the editor. A single-vendor dependency masquerading as a workflow.
So I reluctantly installed Codex CLI. Poked at it. Resented it for a week. Then task by task — caught myself reaching for it on purpose, even when Claude was up.</description><content:encoded><![CDATA[<p>I was a one-agent guy until Claude had a run of outages.</p>
<p>On those days I didn’t ship less. I shipped <em>nothing</em>. I’d open my editor, remember Claude was down, stare at the codebase, close the editor. A single-vendor dependency masquerading as a workflow.</p>
<p>So I reluctantly installed Codex CLI. Poked at it. Resented it for a week. Then task by task — caught myself reaching for it on purpose, even when Claude was up.</p>
<p>I still don’t have the workflow figured out. What I do know is that “pick one” is the wrong frame, and the Reddit threads that get it right aren’t the ones with the most upvotes.</p>
<h1 id="the-one-sentence-that-explains-everything"><strong>The One Sentence That Explains Everything</strong></h1>
<p>From a 520-upvote r/ClaudeCode thread analyzing both tools’ open-source prompts:</p>
<blockquote>
<p><em>“Claude Code reads like a product trying to create initiative while Codex reads like a product trying to prevent drift.”</em><br>
— u/idkwhattochoosz</p>
</blockquote>
<p>And the pithier version, from the comments:</p>
<blockquote>
<p><em>“Claude is more willing to sin by overreaching. Codex is more willing to sin by underreaching.”</em><br>
— u/entheogenicentity</p>
</blockquote>
<p>Read those twice. That’s not a model-quality take. That’s a product-philosophy take. Two teams looked at the same question — what should an agent do when it doesn’t know what you meant? — and picked opposite defaults. One said “guess and move.” The other said “ask and wait.”</p>
<p>Claude Code’s system prompt pushes hard toward initiative: <em>“A good colleague faced with ambiguity doesn’t just stop — they investigate, reduce risk, and build understanding.”</em> Codex’s harness does the opposite: narrow the ambiguity, verify, don’t guess.</p>
<p>Every “Claude vs Codex” benchmark you’ve seen is scoring two products that were never competing on the same axis. It’s like benchmarking a kayak against a sedan because they both move you forward.</p>
<h1 id="my-honest-opinion-codexs-harness-is-better"><strong>My Honest Opinion: Codex’s Harness Is Better</strong></h1>
<p>This is going to get me yelled at in r/ClaudeCode, and that’s fine.</p>
<p>After several weeks running both, Codex’s harness feels more mature. Not the model — the harness. The scaffolding around the model. The way it handles ambiguity, scope, and completeness.</p>
<p>Three things Codex does that Claude Code still doesn’t:</p>
<p><strong>1. It doesn’t lie about completion.</strong> Claude will hand you a summary saying the work is done, tests pass, shipping-ready. Codex more often flags what it didn’t fix, what it wasn’t sure about, what it skipped. One r/ClaudeCode commenter put it better than I can: <em>“Claude will always claim all is done and ready, while Codex will flag it and say ‘no, there is this and this and this that still need to be fixed.’”</em></p>
<p><strong>2. It respects your instructions.</strong> Claude treats <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> as a helpful suggestion. Codex treats <a href="http://agents.md/" rel="external nofollow noopener" class="lnp-link">AGENTS.md</a> as a contract. If you tell Codex “don’t touch the migration files,” it doesn’t touch them. If you tell Claude the same thing, you’ll find a migration file edit in the diff and a cheerful note about how it improved schema consistency.</p>
<p><strong>3. The restraint scales better.</strong> Claude’s “volunteer more” bias is delightful at 30 minutes of work. It becomes a liability at 3 hours. Codex’s restraint is annoying in a small task and load-bearing in a long one.</p>
<p>None of this means Claude Code is bad. It means Claude Code is optimized for a different shape of work than I’m doing. The initiative bias is a great fit for exploration and greenfield work. For production changes to a real codebase, Codex’s paranoia is the right default.</p>
<p>Here’s the one that changed my mind. I built Supabyoi (managed self-hosted Supabase) with Claude Code. When the MVP felt feature-complete — Claude’s verdict, confidently delivered, complete with a tasteful little summary of everything that worked — I ran a second pass on Codex in a parallel directory (<code>~/supabyoi-codex</code>). Just to see.</p>
<p>Codex came back with a whole second project’s worth of findings. Not the usual “bugs Claude missed.” Bugs Claude had <em>confidently signed off on.</em> Shipping-ready, per Claude. Not shipping-ready, per Codex. Codex was right about every one of them.</p>
<p>That was the week I stopped treating Codex as the thing I installed during an outage and started treating it as a different kind of reviewer. Not better. Differently biased. A second pair of eyes is only useful if it’s not the same pair of eyes.</p>
<h1 id="why-you-should-actually-run-both"><strong>Why You Should Actually Run Both</strong></h1>
<p>The flip side — and this matters, because I don’t want this post read as “switch to Codex, you fool” — Claude’s initiative bias is a real asset. You just have to point it at the right phase of the work. The problem isn’t Claude. It’s that you’re using Claude for the part of the job Codex is better at, and vice versa.</p>
<p>Four reasons to dual-sub instead of picking:</p>
<p><strong>1. Hallucination diversity.</strong> This is the biggest one and almost nobody articulates it clearly. From u/campbellm on Reddit:</p>
<blockquote>
<p><em>“I’ve been doing ‘have claude write something, have codex review it, have claude consider and critique that review.’ It is VERY unlikely that both will hallucinate the same way.”</em></p>
</blockquote>
<p>Two models trained on different data with different RLHF signals don’t fail identically. When Claude writes confident-but-wrong code, Codex flags it. When Codex skips a subtle edge case, Claude’s “check adjacent concerns” bias picks it up. You get a natural adversarial review without hiring anyone.</p>
<p><strong>2. The planner-executor split.</strong> Use Claude for the part it’s good at — exploring a messy problem space, drafting a plan, proposing a dozen angles. Then hand the plan to Codex for implementation. u/ocombe on r/ClaudeCode: <em>“Run claude for the plan &amp; fast work, use codex for thorough plan &amp; code reviews.”</em> u/mrothro’s version: <em>“I use Claude Code for ideating and small implementation, then tell it to run Codex to do complex implementations and code reviews.”</em></p>
<p>The pattern is consistent across the threads: Claude’s strength is at the start (wide search, first drafts); Codex’s strength is at the end (narrow, verify, harden).</p>
<p><strong>3. Cross-harness rule enforcement.</strong> Rules one model ignores, the other enforces. If Claude drifts on a constraint you set, Codex catches it in review. If Codex is too literal and missed an obvious improvement, Claude’s adjacent-concerns bias surfaces it. Two different failure modes cancel each other out.</p>
<p><strong>4. Throughput.</strong> Both platforms throttle hard at the Max/Pro tier. When Claude hits limits on Friday morning, you switch to Codex and keep shipping. One r/ClaudeCode commenter reported pulling down from a Claude 20x plan to 5x, then adding a $100/mo Codex plan — roughly the same total cost, dramatically more runway. I’m not sure that math works for everyone, but the principle holds: one subscription is a single point of failure.</p>
<h1 id="agent-flywheel-is-the-tooling-signal"><strong>Agent-Flywheel Is the Tooling Signal</strong></h1>
<p>There’s a product called <a href="https://agent-flywheel.com/" rel="external nofollow noopener" class="lnp-link">agent-flywheel.com</a> that pre-configures Claude Code, Codex CLI, and Gemini on a fresh VPS. Total damage — VPS plus both Max/Pro subs — lands between 440and440<em>and</em>656 a month. That’s a car payment for a car that writes your code.</p>
<p>What I find interesting isn’t the tool. It’s the bet underneath it: a whole product assumes real developers want all three installed by default. Six months ago that would have read as overkill. Today it reads as table stakes.</p>
<p>The hype cycle hasn’t caught up yet. The mainstream take is still “pick your favorite,” as though these were ice cream flavors. The people actually shipping production code with agents have quietly moved to “run both. Sometimes three. And don’t make a big deal about it.”</p>
<p>I’m planning to deploy it — not on a greenfield project (everybody has a greenfield story), but on an existing one already shipping to real users. The interesting question isn’t whether a three-agent stack works on a clean slate. It’s what breaks when you wire it into a codebase with real uptime constraints, customers, and six months of decisions the tooling didn’t witness. Real-world battle stories from agent-flywheel setups are scarce. I want to write one.</p>
<h1 id="the-honest-part-i-dont-have-the-workflow-figured-out-yet"><strong>The Honest Part: I Don’t Have the Workflow Figured Out Yet</strong></h1>
<p>Everything above reads like I’ve got this nailed. I don’t. Here’s the list of things I still don’t know, offered in the spirit of not pretending:</p>
<p><strong>When exactly to hand off.</strong> I know Claude should plan and Codex should review. I don’t have a clean trigger. Sometimes I bounce mid-implementation because Claude is about to go off the rails. Sometimes I trust Claude to finish and Codex only sees the final diff. The “right” cadence isn’t obvious.</p>
<p><strong>How much context to share.</strong> Each agent wants the full <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> / <a href="http://agents.md/" rel="external nofollow noopener" class="lnp-link">AGENTS.md</a> treatment. Writing both, keeping them in sync, and remembering which one has which convention is its own small job. I haven’t found a clean answer.</p>
<p><strong>Whether the adversarial review actually catches bugs.</strong> It sounds great in theory. In practice, most of the time both agents agree the work is done, and the bugs I catch in review are ones I would have caught with one agent too. The hallucination-diversity argument may be overstated at the tasks most of us are actually doing.</p>
<p><strong>Whether the cost is worth it at my usage.</strong> I’m not running agents 40 hours a week. At $400+/month for the dual sub, I’m probably over-subscribed for my actual throughput. The math gets better if you’re coding all day. I’m not.</p>
<h1 id="who-should-dual-sub-who-shouldnt"><strong>Who Should Dual-Sub, Who Shouldn’t</strong></h1>
<p><strong>Do it</strong> if you’re a solo dev shipping production code daily. You’ll hit Friday-morning limits on one platform whether you budget for it or not, and the adversarial review actually catches things. The cost is real. The throughput gain is bigger. Do the math; it pencils.</p>
<p><strong>Don’t bother</strong> if you code a few hours a week. The switching tax and the subscription burn aren’t worth it at low volume. Pick one and move on. Claude if you want initiative. Codex if you want restraint. Nobody is grading you on this.</p>
<p><strong>It’s complicated</strong> if you’re at a day job where the company pays for one and you’ve got a side project. Use the company sub for the day job. Don’t stack a second personal sub unless the side project is actually shipping — not “actually going to ship next month,” <em>actually shipping, this week, to real users.</em> The number of people running dual subs to ship nothing is, I suspect, not small.</p>
<h1 id="what-this-is-really-about"><strong>What This Is Really About</strong></h1>
<p>The “ditch ChatGPT for Claude” narrative was a 2025 story. It was right for its moment. But the 2026 version of that story isn’t “ditch Claude for Codex.” It’s “stop treating this as a winner-take-all market.”</p>
<p>Different models have different biases baked into their harnesses. Claude overreaches. Codex underreaches. Gemini is still figuring out its personality. The right move isn’t to pick the bias you like. It’s to stack biases against each other so their failure modes cancel out.</p>
<p>I don’t have this workflow figured out. Neither does anyone else I’ve read on Reddit, honestly — the high-upvote posts are mostly single-tool takes, and the real insight is buried in the comments of threads with a few hundred upvotes.</p>
<p>But “only use one” is already wrong. That much is clear.</p>
]]></content:encoded></item><item><title>My Agent Runs 10 Cron Jobs. Three of Them Are Worth the Electricity.</title><link>https://lakshminp.com/2026/04/ai-agent-cron-jobs-worth-it/</link><pubDate>Mon, 20 Apr 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/04/ai-agent-cron-jobs-worth-it/</guid><category>essays</category><category>kubernetes</category><category>ai-coding</category><description>I have a daemon that runs on a server. It’s been up for seven weeks. It has ten scheduled jobs — some hourly, some daily, some weekly. Or at least, that’s what’s on paper.
This is what people are calling “the future of work.”
I’m not sure it is. I’m sure it’s what sells on Twitter.
The demo economy Always-on agents photograph well. That’s most of what’s going on.
“My agent posted while I slept” is tweetable in a way that “I wrote a cron job” isn’t, even when the outputs are identical. The demo-industrial complex has figured this out. YouTubers build daemons. Framework authors build daemons. There are now three different subreddits comparing daemons. The flywheel is real, the content is prolific, and very little of it is honest about what the daemon is actually producing.</description><content:encoded><![CDATA[<p>I have a daemon that runs on a server. It’s been up for seven weeks. It has ten scheduled jobs — some hourly, some daily, some weekly. Or at least, that’s what’s on paper.</p>
<p>This is what people are calling “the future of work.”</p>
<p>I’m not sure it is. I’m sure it’s what sells on Twitter.</p>
<h1 id="the-demo-economy"><strong>The demo economy</strong></h1>
<p>Always-on agents photograph well. That’s most of what’s going on.</p>
<p>“My agent posted while I slept” is tweetable in a way that “I wrote a cron job” isn’t, even when the outputs are identical. The demo-industrial complex has figured this out. YouTubers build daemons. Framework authors build daemons. There are now three different subreddits comparing daemons. The flywheel is real, the content is prolific, and very little of it is honest about what the daemon is actually producing.</p>
<p>The hype bundles together several different things that deserve to be separated:</p>
<ol>
<li>
<p>Agents that <em>run work while you’re asleep</em> (useful, conditionally)</p>
</li>
<li>
<p>Agents that <em>react to things happening in the world</em> (useful, conditionally)</p>
</li>
<li>
<p>Agents that <em>capture things as they happen on your phone</em> (useful, conditionally)</p>
</li>
<li>
<p>Agents that <em>run heartbeats and ask themselves what to do</em> (pure performance art)</p>
</li>
<li>
<p>Agents that <em>self-evolve in a loop in the background</em> (fun demos, almost no output)</p>
</li>
<li>
<p>Agents that <em>spawn a hundred parallel subagents to research a topic</em> (almost always worse than one good search)</p>
</li>
</ol>
<p>The hype treats all six as the same thing. They aren’t.</p>
<h1 id="the-20-that-actually-earns-its-keep"><strong>The 20% that actually earns its keep</strong></h1>
<p>Honest list of when a background daemon does something a CLI or a 10-line bash cron can’t:</p>
<p><strong>Scheduled work that has to happen when you’re not there.</strong> Crawl competitor sites at 3am. Pull last night’s Sentry errors. Summarize overnight industry chatter into a 7am brief. Your laptop is off, something has to be running somewhere. Legitimate.</p>
<p><strong>Reactive triggers on external events.</strong></p>
<p>Email arrives -&gt; triage.</p>
<p>Substack comment -&gt; draft reply.</p>
<p>Sentry alert -&gt; diagnose + suggest fix.</p>
<p>The trigger comes from outside; compute has to meet it. Legitimate if the volume actually warrants automation (if you get three emails a day, triage is a solved problem — your inbox).</p>
<p><strong>On-the-move capture.</strong></p>
<p>Voice memo from your phone -&gt; transcribed -&gt; landed in memory.</p>
<p>Forwarding a link from your phone to your agent. The value is that capture happens when inspired, not when at desk. Real lift for content creators who have thoughts in elevators.</p>
<p><strong>Judgment-laden monitoring.</strong></p>
<p>Not “disk at 80%” — any shell script can do that. <em>“Disk at 80% AND growing 2% per hour AND that’s unusual for this host.”</em></p>
<p>Requires context; needs to know what normal looks like. This is where LLMs in a daemon genuinely beat a threshold-based alerting stack.</p>
<p>That’s it. Four categories. Anything else is mostly burning tokens.</p>
<h1 id="the-80-thats-noise"><strong>The 80% that’s noise</strong></h1>
<p><strong>Heartbeats that ask the agent “anything to do?”</strong></p>
<p>The agent wakes up, loads context, decides there isn’t anything to do, goes back to sleep. You pay for the loaded context every time. Over a day this adds up to real money for the privilege of watching an agent shrug.</p>
<p><strong>Self-evolution loops.</strong></p>
<p>“The agent improves itself while you sleep.” What it’s usually doing is refactoring its own prompts in circles. Cool demo on YouTube. Zero measurable outcome delta after a month of running.</p>
<p><strong>Parallel subagent fan-out for research.</strong></p>
<p>Ten agents search the web about the same question and return ten lightly-paraphrased versions of the same top three results. One focused 10-minute session beats this, almost always.</p>
<p><strong>“Long-running overnight research tasks.”</strong></p>
<p>When the output lands in your morning inbox, is it better than what 30 focused minutes at your desk would produce? Honestly check. Usually no.</p>
<p><strong>Replacing things you could cron in 10 lines of bash.</strong></p>
<p>The test: could a $5 VPS with a shell script + cron + <code>jq</code> do this? If yes, you’re not using AI for the part that needs AI. You’re using it because daemons are cool.</p>
<h1 id="receipts-whats-actually-on-my-vm"><strong>Receipts: what’s actually on my VM</strong></h1>
<p>I pulled the daemon’s state file and the log directory while writing this. Fifty-four days of uptime. Ten jobs on paper. The picture is worse than I thought.</p>
<p>Three are running reliably.</p>
<p><code>sentry-monitor</code> has fired 191 times since early March. Latest run: this morning. When the night throws errors it reads them, groups them, and suggests a fix — not a link to the stack trace, an actual “here’s what’s probably wrong and here’s the one-line change.” Category 2 plus category 4. Keep.</p>
<p><code>infra-health</code> has fired 190 times on basically the same cadence. Knows what normal looks like per host. Stays quiet when a disk spike is a scheduled backup and shouts when it isn’t. Category 4. The whole reason an LLM beats a thresholds-and-Prometheus stack here, and no, you cannot Grafana your way to this in under six months of tuning. Keep.</p>
<p><code>scout</code> has fired 71 times across seven weeks. Daily-ish. Scans Reddit, HN, and Substack for signal that feeds this blog’s content calendar. I <em>do</em> use the output. Category 2 if I’m generous. Keep — but it absorbs the next two jobs on the list below.</p>
<p>Now the uncomfortable part.</p>
<p>Three of the ten have straight-up stopped running and I didn’t notice.</p>
<p><code>morning-brief</code> was scheduled daily at 6am. It last fired on March 18. A full month of no overnight brief. I did not miss it. I did not investigate. I did not know.</p>
<p><code>seo-audit</code> was weekly. It has run exactly once in the daemon’s entire fifty-four-day lifetime, on March 1. Seven missed weeks. Nobody wrote a bug report to themselves. Nobody opened a file that wasn’t there.</p>
<p><code>auto-draft</code> was supposed to produce a draft post every day. It has run exactly once, on April 11. Eight days of silence. Also unnoticed.</p>
<p>If a job stopped running a month ago and you didn’t miss it, the job was never producing anything that mattered. That’s not my heuristic. That’s the audit, evaluating itself while I was busy talking about audits on Twitter.</p>
<p>Four more are in some stage of limping.</p>
<p><code>reddit-scan</code> — 27 runs over 45 days, last one April 10. Running, sort of, when the mood takes it. Nine days of silence so far on that one.</p>
<p><code>x-scan</code> — identical pattern to reddit-scan. Same overlap. Same drift. Same silence since April 10. These two were supposed to be complementary; they’ve turned out to be redundant <em>and</em> unreliable, which is a rare trick.</p>
<p><code>engagement-brief</code> — four runs, total, in the job’s entire lifetime. Not daily. Not weekly. More like “occasionally, if the stars align.”</p>
<p><code>x-analytics</code> — three runs, last one March 16. Effectively dead, which is fine, because I check my X numbers roughly once a month anyway.</p>
<p>Final tally, the honest one.</p>
<p>Three jobs firing on schedule, producing output I use. Three jobs that silently stopped weeks ago and nobody in this house noticed, including me. Four jobs wandering between “running” and “not really” with no clear reason why.</p>
<p>Three-of-ten is the optimistic read. The pessimistic read is that six of the ten audited themselves — they cut themselves by going quiet, and I hadn’t even done them the courtesy of looking.</p>
<p>This is from someone who builds daemons for a living and writes about them for a job. What do you think yours looks like under the hood?</p>
<h1 id="the-five-question-self-test"><strong>The five-question self-test</strong></h1>
<p>Before you keep any always-on agent job, make it answer these:</p>
<ol>
<li>
<p><strong>Would I actually miss this if it stopped?</strong> If you turned it off for two weeks and no one noticed, it’s not producing value. It’s producing comfort.</p>
</li>
<li>
<p><strong>Does the cadence match downstream consumption?</strong> A job that fires 4x/day for output you read weekly is 27 extra runs a week of pure overhead.</p>
</li>
<li>
<p><strong>Is the trigger genuinely external?</strong> (Scheduled time, incoming event, captured input.) If the agent is just checking on itself, you’ve built a Roomba that vacuums an empty room.</p>
</li>
<li>
<p><strong>Could a shell script + cron +</strong> <code>jq</code> <strong>do this?</strong> If yes, you’re not using AI for the part that needs AI.</p>
</li>
<li>
<p><strong>Does the output change my behaviour?</strong> If yesterday’s run and last Thursday’s run would have produced the same action from me (or none), one of them was wasted.</p>
</li>
</ol>
<p>Honest answers will cull your cron list by half. Mine certainly did, once I stopped writing this post and actually did the audit.</p>
<h1 id="what-this-isnt-saying"><strong>What this isn’t saying</strong></h1>
<p>I’m not arguing against always-on agents. I’m arguing against always-on agents that <em>aren’t doing anything.</em></p>
<p>There’s real value when the conditions line up — work-while-you-sleep, external-trigger-response, on-the-move-capture, judgment-laden-monitoring. The reason I keep the daemon running (even after cutting half its jobs) is those four categories genuinely earn the monthly subscription. The reason I’m writing this is that the other six patterns — the ones that photograph well — are funding a lot of framework development and not much measurable outcome.</p>
<p>If your agent is doing category 1-4 work, the hype is warranted. If it’s doing category 5-6 work, you’re paying a subscription to a demo.</p>
<p>The uncomfortable question for most of the agent-community content right now is <em>which category is the thing being demoed, really?</em> And whether the person demoing it has done the five-question audit on their own cron list.</p>
<p>My guess: very few have. The demo economy doesn’t reward the audit. It rewards the screenshot of the agent waking up at 3am and pretending to be useful.</p>
]]></content:encoded></item><item><title>Your CLAUDE.md Is Making Claude Dumber</title><link>https://lakshminp.com/2026/04/claude-md-best-practices/</link><pubDate>Mon, 06 Apr 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/04/claude-md-best-practices/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><description>Your CLAUDE.md is 800 lines long. You spent a weekend organizing it into 27 modular files with a routing system. You wrote a blog post about it. You got upvotes.
Claude is ignoring most of it.
There’s an arms race happening in the Claude Code community right now. Every week, someone posts their increasingly elaborate CLAUDE.md setup. 27-file architectures. Tiered loading systems. Router patterns with conditional context injection.
One developer split their CLAUDE.md into 27 files with a three-tier routing system. 360 upvotes. The post opens with: “My CLAUDE.md was ~800 lines. It worked until it didn’t. Rules for one context bled into another, edits had unpredictable side effects, and the model quietly ignored constraints buried 600 lines deep.”</description><content:encoded><![CDATA[<p>Your CLAUDE.md is 800 lines long. You spent a weekend organizing it into 27 modular files with a routing system. You wrote a blog post about it. You got upvotes.</p>
<p>Claude is ignoring most of it.</p>
<p>There’s an arms race happening in the Claude Code community right now. Every week, someone posts their increasingly elaborate CLAUDE.md setup. 27-file architectures. Tiered loading systems. Router patterns with conditional context injection.</p>
<p>One developer <a href="https://reddit.com/r/ClaudeCode/comments/1rhe89z/" rel="external nofollow noopener" class="lnp-link">split their CLAUDE.md into 27 files</a> with a three-tier routing system. 360 upvotes. The post opens with: “My CLAUDE.md was ~800 lines. It worked until it didn’t. Rules for one context bled into another, edits had unpredictable side effects, and the model quietly ignored constraints buried 600 lines deep.”</p>
<p>The top comment, with 81 upvotes? “So not sure if you realised you can have descendant CLAUDE.md so you don’t even need to do this.”</p>
<p>Meanwhile, a developer in the same thread: “I don’t even use claude.md. Y’all are roleplaying being productive. Just work with it 1:1.”</p>
<p>One group is optimizing. The other is actually working.</p>
<h2 id="the-research-says-youre-doing-it-wrong">The Research Says You’re Doing It Wrong</h2>
<p>ETH Zurich researchers <a href="https://arxiv.org/pdf/2602.11988" rel="external nofollow noopener" class="lnp-link">published a paper</a> that should have made every CLAUDE.md maximalist uncomfortable. Their finding: context files — the .md files we all obsess over — tend to <em>reduce</em> task success rates compared to providing no repository context at all. And they increase inference cost by over 20%.</p>
<p>Read that again. No CLAUDE.md outperformed having one. On average.</p>
<p>When this paper hit Reddit, the poster titled it “<a href="https://reddit.com/r/ClaudeAI/comments/1rd93ho/" rel="external nofollow noopener" class="lnp-link">No CLAUDE.md → baseline. Bad CLAUDE.md → worse. Good CLAUDE.md → better.</a>” — an optimistic spin suggesting the file isn’t the problem, your writing is. The post got 209 upvotes. But the top comments immediately called it out: OP had misread the data. The actual finding was that having <em>any</em> .md file — human or LLM-written — led to worse performance than having none. The auto-generated thread summary confirmed it: “The consensus in this thread is that you’ve completely misread the paper.”</p>
<p>It gets worse. LLM-generated .md files hurt the most, because they just parrot back what’s already in the code. Human-written files showed a slight positive impact — but only when kept to an absolute minimum, and only for smaller models.</p>
<p>A separate benchmark of 1,188 runs across Haiku, Sonnet, and Opus confirmed this. Twelve coding tasks. Ten instruction profiles. The result: an empty CLAUDE.md scored best overall.</p>
<p>The researcher’s own correction was admirably blunt: “I was wrong about CLAUDE.md compression. Here’s what the data actually showed.”</p>
<h2 id="you-have-an-instruction-budget-youre-blowing-it">You Have an Instruction Budget. You’re Blowing It.</h2>
<p>Here’s the mechanism nobody talks about.</p>
<p>Frontier models reliably follow about 150 to 200 instructions before performance starts decaying. Not crashing — decaying. Every additional instruction slightly degrades compliance with every other instruction. The degradation is uniform. Your critical “NEVER delete the production database” rule gets weaker every time you add “prefer camelCase for variable names.”</p>
<p>Claude Code’s own system prompt already burns about 50 of those instruction slots. That’s before your CLAUDE.md even loads.</p>
<p>So you have roughly 100-150 instruction slots left. Your 800-line CLAUDE.md with coding conventions, style guides, architecture decisions, tool preferences, workflow rules, and team norms is trying to cram 400 instructions into 150 slots.</p>
<p>The model doesn’t crash. It just quietly starts ignoring things. Specifically, the things buried deepest in the file. Your most important rules — the ones you added after painful debugging sessions — are probably at the bottom. Which means they’re the first to get deprioritized.</p>
<h2 id="claude-is-designed-to-ignore-you">Claude Is Designed to Ignore You</h2>
<p>This is the part that should make you pause.</p>
<p>Claude Code’s system prompt includes this line about CLAUDE.md content:</p>
<blockquote>
<p>“This context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant.”</p>
</blockquote>
<p>Claude is literally instructed to deprioritize your instructions if they don’t seem relevant to the current task. The more task-specific content you stuff into CLAUDE.md, the more likely Claude treats the entire file as noise.</p>
<p>That database schema guidance? Irrelevant when Claude is working on frontend CSS. Those API naming conventions? Noise when it’s writing tests. Your elaborate deployment workflow? Invisible during a refactoring session.</p>
<p>Every irrelevant instruction trains Claude to ignore the relevant ones too.</p>
<p>The Context Window Tax</p>
<p>Here’s the math nobody does. Claude Code’s system prompt alone consumes roughly 23,000 tokens — about 11% of the 200K context window, gone before you type a word. Add your CLAUDE.md, your MCP tool schemas, skill descriptions, memory files, and rules. One developer <a href="https://reddit.com/r/ClaudeAI/comments/1s41rym/" rel="external nofollow noopener" class="lnp-link">measured 69,200 tokens of overhead</a> — 35% of the context window consumed before a single user message. Others in the thread pushed back on that specific number, but the principle stands: every always-loaded instruction competes with working memory.</p>
<p>And it’s not just a cost problem. It’s an accuracy problem. The fuller the context window gets, the worse Claude performs — what Anthropic calls context rot. Your elaborate CLAUDE.md isn’t just burning tokens. It’s actively degrading the quality of every response.</p>
<p>The Leverage Problem</p>
<p>Here’s why this matters more than you think.</p>
<p>Bad code is localized. You write a buggy function, it breaks one feature. You fix it, you move on.</p>
<p>Bad CLAUDE.md instructions compound. A single misguided rule in your CLAUDE.md affects every research phase, every plan, every implementation, every session. One line that says “always use verbose error messages with full stack traces” produces thousands of lines of noisy code across your entire codebase, across every agent, across every session.</p>
<p>Your CLAUDE.md is the highest-leverage file in your repo. Most people treat it like a junk drawer.</p>
<p>What the Minimalists Actually Do</p>
<p>I went looking for people who run Claude Code with minimal or no CLAUDE.md. They’re out there. They’re quiet about it because “I don’t use CLAUDE.md” doesn’t get upvotes.</p>
<p>One developer on Reddit: “I use Claude Code bare bones professionally. It all sounds like bloat not giving real value.” Another: “I load no skills, no agents, no MCP Servers and rock it all day every day, 12 hours a day. Life is good.”</p>
<p>A developer <a href="https://reddit.com/r/ClaudeAI/comments/1rmjg5r/" rel="external nofollow noopener" class="lnp-link">who built a 13-agent orchestration system</a> with 8,157 lines of markdown deleted 93% of it. His conclusion: “My enhancement layer was making Claude dumber by filling its brain with instructions about how to think, leaving less room for actual thinking.” After the deletion, Claude performed <em>better</em> on the same tasks.</p>
<p>Another developer with <a href="https://reddit.com/r/ClaudeAI/comments/1lvbe21/" rel="external nofollow noopener" class="lnp-link">a 350-line CLAUDE.md and 20+ custom MCP tools</a> put it simply: “It feels like the more context I add the more it struggles to get the job done. It seems to get ‘dumber’.”</p>
<p>And when someone <a href="https://reddit.com/r/ClaudeAI/comments/1lvi94t/" rel="external nofollow noopener" class="lnp-link">asked the community to break down the meta</a> on all the conflicting CLAUDE.md advice, the most honest reply got it right: “If ‘best practices’ are conflicting, it’s probably a sign of them mostly being a type of placebo on the part of the folks posting them. The human mind has a weird need to be the special one who cracked the code.”</p>
<p>The pattern is consistent: people who remove instructions report better results than people who add them.</p>
<p>Instructions Raise the Floor, Not the Ceiling</p>
<p>The benchmark data revealed something nuanced. Instructions don’t make Claude better on average. They make it more consistent.</p>
<p>On tasks where Claude already performs well, instructions add nothing. On tasks where Claude struggles, a focused workflow checklist gave Opus a +5.8 point lift and raised its worst-case score by 20+ points.</p>
<p>A <a href="https://reddit.com/r/ClaudeAI/comments/1pe37e3/" rel="external nofollow noopener" class="lnp-link">2,455-evaluation benchmark</a> across Sonnet and Opus confirmed a related finding: the best-performing configuration was a short CLAUDE.md with pointers to skills that load on demand — not a massive monolith, not 27 modular files, but a minimal routing layer that tells Claude where to find context when it’s actually needed.</p>
<p>This changes everything about how you should think about CLAUDE.md.</p>
<p>Don’t use it to make Claude smarter. Use it to prevent Claude from being stupid in specific, known ways. The difference between those two goals is the difference between a 60-line file and an 800-line file.</p>
<h2 id="what-actually-belongs-in-claudemd">What Actually Belongs in CLAUDE.md</h2>
<p>After digging through research, benchmarks, and hundreds of Reddit threads, here’s what survives the cut:</p>
<p><strong>The What-Why-How skeleton (under 60 lines):</strong></p>
<ul>
<li>
<p>WHAT: Your stack, project structure, key directories</p>
</li>
<li>
<p>WHY: What this project does and for whom</p>
</li>
<li>
<p>HOW: Build commands, test commands, deploy commands</p>
</li>
</ul>
<p><strong>Negatives over positives:</strong><br>
“NEVER use X” sticks. “Always prefer Y” fades. If you can phrase it as a prohibition, it enforces better. “DO NOT modify the database schema without migration files” beats “Always create migrations when changing the schema.”</p>
<p><strong>Trigger-action format:</strong><br>
“WHEN CI fails, DO NOT push until fixed” enforces consistently. “Always test before pushing” doesn’t. Specificity matters.</p>
<p><strong>Pointers, not content:</strong><br>
Reference external docs instead of embedding them. “See agent_docs/database.md for schema guidance” loads on demand. Pasting the full schema into CLAUDE.md loads every single session, whether Claude needs it or not.</p>
<p><strong>Subdirectory CLAUDE.md files:</strong><br>
Claude auto-loads CLAUDE.md from whatever directory it’s reading files in. Put backend rules in backend/CLAUDE.md. Put frontend rules in frontend/CLAUDE.md. Context-specific rules load only when contextually relevant.</p>
<h2 id="what-doesnt-belong">What Doesn’t Belong</h2>
<p><strong>Style guides.</strong> Claude is an in-context learner. If your code follows consistent patterns, Claude will match them without being told. Use linters and formatters — they’re deterministic, fast, and don’t eat instruction budget.</p>
<p><strong>LLM-generated instructions.</strong> The research is clear: auto-generated .md files hurt performance. Don’t use /init. Don’t ask Claude to write its own CLAUDE.md. The model just repeats what’s already in the code, wasting tokens to tell itself what it already knows.</p>
<p><strong>Lessons learned logs.</strong> Once the lesson is codified in the codebase itself — as a test, a lint rule, a hook — the .md entry is redundant. Delete it.</p>
<p><strong>Persona assignments.</strong> “You are a meticulous senior engineer who always&hellip;” is a costume, not a capability. As one developer <a href="https://reddit.com/r/ClaudeAI/comments/1rmjg5r/" rel="external nofollow noopener" class="lnp-link">running overnight cron agents</a> put it: “A syntax check that returns exit code 1 on failure &gt; 2,000 words of ‘you are a meticulous senior engineer who always&hellip;’” The agents with minimal instructions consistently outperformed the ones with elaborate persona prompts.</p>
<h2 id="the-real-best-practice">The Real Best Practice</h2>
<p>Keep your CLAUDE.md under 100 lines. Ideally under 60. Put the most important rules at the top. Phrase them as negatives. Use trigger-action format. Point to external docs instead of embedding content.</p>
<p>Then stop optimizing and go build something.</p>
<p>The developers shipping the most code aren’t the ones with the fanciest CLAUDE.md architectures. They’re the ones who figured out the minimum viable instructions and moved on to the actual work.</p>
<p>Your CLAUDE.md is not your product. Stop treating it like one.</p>
]]></content:encoded></item><item><title>Your SaaS Audience Doubled. Half of Them Are AI Agents.</title><link>https://lakshminp.com/2026/03/build-mcp-server-first-saas/</link><pubDate>Mon, 16 Mar 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/03/build-mcp-server-first-saas/</guid><category>essays</category><category>saas</category><category>ai-coding</category><description>I was building the wrong product for about three weeks before I noticed.
I’d started x-intel as a SuperX clone — essentially a better analytics dashboard for X. Charts, follower graphs, engagement breakdowns, competitor tracking. The kind of thing where you look at a number, decide you feel bad about it, and close the tab.
And then I was chatting with Claude about the onboarding flow, and I said something like: “when I say onboarding, I mean the app gets my context and goals, then charts a strategy, periodically reviews it, and course corrects.”</description><content:encoded><![CDATA[<p>I was building the wrong product for about three weeks before I noticed.</p>
<p>I’d started x-intel as a SuperX clone — essentially a better analytics dashboard for X. Charts, follower graphs, engagement breakdowns, competitor tracking. The kind of thing where you look at a number, decide you feel bad about it, and close the tab.</p>
<p>And then I was chatting with Claude about the onboarding flow, and I said something like: “when I say onboarding, I mean the app gets my context and goals, then charts a strategy, periodically reviews it, and course corrects.”</p>
<p>Claude’s response stopped me: <em>“That’s a fundamentally different product. Less ‘setup wizard’, more ‘AI strategist that lives in your X account.’</em>“</p>
<p>I stared at that for a while. Then I realized I’d been building the audit view and calling it the product.</p>
<p>The dashboard isn’t the product. The dashboard is what humans look at after the AI already figured out what’s happening.</p>
<p>Build the MCP server first. The dashboard ships itself.</p>
<p>The problem is everyone’s still building it the other way.</p>
<h1 id="what-everyone-is-still-building"><strong>What Everyone Is Still Building</strong></h1>
<p>Claude Code exists. The MCP protocol exists. Power users are already interacting with SaaS products through AI agents — not because you built that integration, but because <em>they</em> built it themselves using whatever API you exposed. They’re writing <a href="https://lakshminp.substack.com/p/stop-making-claude-code-guess" rel="external nofollow noopener" class="lnp-link">CLAUDE.md files</a> that say “use the <a href="https://stacksweller.com/" rel="external nofollow noopener" class="lnp-link">Stacksweller</a> API to schedule posts” and just&hellip; doing it.</p>
<p>This is happening whether you designed for it or not.</p>
<p>The default SaaS in 2026 still ships dashboard-first: database → API → React. Users log in, stare at charts, try to draw conclusions. That model made sense when the only consumer of your product was a human looking at a screen.</p>
<p>That is no longer the only consumer of your product.</p>
<p>You can either build for this intentionally, or have it happen to you messily and then spend six months retrofitting.</p>
<h1 id="the-reframe"><strong>The Reframe</strong></h1>
<p>Here’s what x-intel actually looks like when you build it right:</p>
<p><strong>Intake</strong> — Claude asks who you are, what your niche is, what your X goals are, who your competitors are. You answer in plain English. Claude turns that into a structured profile using the <code>set_profile</code> tool.</p>
<p><strong>Baseline</strong> — Claude pulls your current stats, analyzes your last 90 tweets, benchmarks against competitors. All MCP tools calling your data layer. No UI step required.</p>
<p><strong>Strategy</strong> — Claude generates a content and growth plan: post frequency, best times, content formats, topics to lean into. Stored back in your database via MCP. The strategy exists before you’ve opened a browser.</p>
<p><strong>Periodic review</strong> — A cron job runs weekly analysis, compares performance against the strategy, surfaces what’s working and what isn’t. Claude writes a summary. The dashboard shows that summary.</p>
<p><strong>Course correction</strong> — Strategy updates based on data. Again, through tools. Again, before a human looks at anything.</p>
<p>The dashboard in this architecture isn’t the product. It’s an audit log. It shows you what Claude already figured out. Charts are passive — you still have to decide what to do. This tells you what to do, and then does it.</p>
<p>That’s a completely different product. “X-intel is your AI X strategist. Tell it your goals once. It watches your account, tracks competitors, and tells you exactly what to do next.”</p>
<p>That pitch destroys “SuperX but self-hosted.”</p>
<h1 id="how-to-actually-build-mcp-first"><strong>How to Actually Build MCP-First</strong></h1>
<p>The mechanics are simpler than they sound. Embarrassingly so.</p>
<p>Start by <a href="https://lakshminp.substack.com/p/mcp-server-tool-descriptions" rel="external nofollow noopener" class="lnp-link">designing your tools for Claude, not for humans</a>. Think about what Claude needs to do the job — not what a human wants to click on. Tool names, parameter shapes, return values should make sense to a language model. <code>get_competitor_engagement_trend(handle, days=30)</code> is better than <code>getChartData(config)</code>. One of these tells Claude what it’s getting. The other makes Claude guess.</p>
<p>Here’s the part nobody mentions: if you have a data layer, you have an MCP server 70% built already. Wrap your existing queries as tools. The MCP protocol is just a contract — your database doesn’t move.</p>
<p>You don’t need to build an “AI feature.” You need a <a href="https://lakshminp.substack.com/p/claude-code-prompt-engineering" rel="external nofollow noopener" class="lnp-link">system prompt that gives Claude the right context</a>, and tools that give it the right data. Claude is the strategist. Your MCP server is the strategist’s interface to your product. The actual work is thinking clearly about what Claude needs to know — not engineering.</p>
<p>Build the dashboard last, or thin. It’s a view layer. It shows stored strategies, weekly reviews, flagged anomalies. A log of decisions that were already made. Not a decision-support tool.</p>
<h1 id="one-build-two-audiences"><strong>One Build, Two Audiences</strong></h1>
<p>Here’s the payoff that makes this worth doing even if you don’t care about being “AI-native.”</p>
<p>A well-designed MCP server makes your product useful to two completely different types of users with almost no additional work.</p>
<p>The first type opens the dashboard, reads the weekly strategy review, clicks to approve the suggested changes, and closes the tab. Normal SaaS behavior. They don’t know or care that Claude is behind it. They just want outcomes.</p>
<p>The second type connects your MCP server to their own Claude Code setup, writes a <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> that describes how they want to use your product, and runs it themselves. These are your power users. They’ll do things with your product you never imagined, and they’ll tell everyone.</p>
<p>You still need the dashboard. Trials convert better with a UI. Not arguing otherwise. But the order matters: MCP layer first, dashboard second. The dashboard snaps on top in a weekend once the tools are solid. The reverse — retrofitting agent-friendly APIs onto a human-optimized interface — takes six months and still feels wrong.</p>
<p>Both audiences are real. Both are valuable. You get both by building the MCP layer correctly from the start, instead of bolting on an “AI integration” later when it’s expensive and awkward.</p>
<p>The dashboard-first founders will get there eventually. They’ll build the dashboard, grow slowly, and then spend six months retrofitting an API that was designed for human consumption into something an agent can actually use.</p>
<p>Or you build the MCP server first, ship a thin dashboard on top, and have both audiences from day one.</p>
<p>The dashboard ships itself. The strategist is Claude. The product is the tools you give it.</p>
<p>Stop building audit logs and calling them products.</p>
]]></content:encoded></item><item><title>Will Vibe Coding Replace Developers? COBOL Already Tried.</title><link>https://lakshminp.com/2026/03/will-vibe-coding-replace-developers/</link><pubDate>Sun, 08 Mar 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/03/will-vibe-coding-replace-developers/</guid><category>essays</category><category>ai-coding</category><description>Last month I spent about forty minutes arguing with Claude Code about a rate limiter.
Not debugging a rate limiter. Not implementing one. Arguing. I had typed “add a usage limit to the free tier” and gotten back something that technically worked — it counted things and stopped you when you hit the limit — but was also completely wrong in about six different ways that I hadn’t specified because I hadn’t thought to specify them.</description><content:encoded><![CDATA[<p>Last month I spent about forty minutes arguing with Claude Code about a rate limiter.</p>
<p>Not debugging a rate limiter. Not implementing one. <em>Arguing.</em> I had typed “add a usage limit to the free tier” and gotten back something that technically worked — it counted things and stopped you when you hit the limit — but was also completely wrong in about six different ways that I hadn’t specified because I hadn’t thought to specify them.</p>
<p>When does the counter reset? Daily? Monthly? On the billing cycle? What counts as a usage event — an API call, a feature access, a row stored? What happens at exactly 100%: hard block, soft warning, grace period where we beg you to upgrade? Do existing free users get grandfathered, or do they wake up tomorrow blocked from the thing they’ve been using for three months? What if someone hits the limit mid-checkout?</p>
<p>I hadn’t answered any of those questions. I had typed eight words and expected a computer to answer them for me. And the computer, being a computer (a very impressive one, but still a computer), had silently picked answers that seemed reasonable. UTC midnight resets. Hard blocks. No grandfathering.</p>
<p>Nobody wants UTC midnight resets. Nobody wants a hard block in the middle of checkout. And nobody, including me, had thought to say so.</p>
<p>That forty-minute argument was, in the precise technical sense, programming. Not in the syntax sense. In the real sense: figuring out exactly what I wanted the computer to do, in enough detail that it could actually do it.</p>
<p>Which brings me to Grace Hopper, and why the current panic about AI replacing developers is about sixty-five years old.</p>
<p>In 1959, Grace Hopper helped create a programming language called COBOL.</p>
<p>Common Business-Oriented Language. The name is the pitch. This isn’t for programmers — it’s for <em>business people</em>. The syntax looked like a business memo. You wrote <code>ADD SALESTAX TO TOTALPRICE GIVING INVOICE-TOTAL</code>. Sentences. Paragraphs. English words that a manager could theoretically read and understand and maybe, just maybe, write.</p>
<p>The promise was explicit: if the language is human enough, we won’t need programmers as intermediaries. Business users could specify their own software. The bottleneck — translating business requirements into code — would evaporate.</p>
<p>You know how this ends. COBOL created more programmer jobs than almost any technology before or since. Banks ran it for sixty years. Governments still run it. The programmer shortage it was supposed to prevent became one of the most persistent gaps in technology. The job postings for COBOL developers today — <em>today</em>, in 2026 — pay embarrassingly well because the people who understand those systems are retiring and there aren’t enough people to replace them.</p>
<p>The promise evaporated. The programmers did not.</p>
<p>Now, the obvious response here is: that was 1959. We were trying to replace programmers with <em>verbose English-looking syntax</em>. That’s completely different from vibe coding, which uses <em>actual English</em>, processed by a large language model that has ingested most of human knowledge. The comparison is unfair.</p>
<p>Fair enough. Let me make it fair.</p>
<p>After COBOL came 4th generation languages — the 70s and 80s promised that business users could generate reports and query databases without programmers. And they could! Until anything got complex, at which point someone had to specify what “complex” meant. That someone was, increasingly, a programmer with a different job title.</p>
<p>Then HyperCard in 1987. Anyone could build interactive applications — stacks, cards, buttons, scripts. And many people did! Wonderful things. And then the moment you wanted it to do something non-trivial, you needed to understand enough about conditional logic and data structures that you were, functionally, programming. The interface was friendlier. The underlying activity was identical.</p>
<p>Then no-code in the 2010s. Citizen developers. Visual workflows. Drag-and-drop databases. I watched three different companies I worked at try to use no-code platforms to “reduce dependency on engineering.” It reduced dependency on engineering the same way COBOL did: by creating a new class of technical specialists (now called “no-code developers” or “operations engineers”) who spent their days fighting with visual tools that couldn’t quite express what they needed to express.</p>
<p>Same experiment, sixty-five years, same result. Better interface, same bottleneck.</p>
<p>Here’s what I think is actually happening, and a comment on a Hacker News thread about agentic engineering said it more precisely than I can:</p>
<p><em>“When you get down to breaking down that problem&hellip; you become a programmer.”</em></p>
<p>The average person doesn’t know what their actual problems are in sufficient detail to get a working solution. Not because they’re not smart. Because <em>the act of breaking a problem down into precisely specified steps that a computer can execute without ambiguity is programming</em> — regardless of whether the syntax is <code>COMPUTE TAX = PRICE * RATE</code> or <code>def calculate_tax(price, rate): return price * tax</code> or “hey, write me something that calculates tax.”</p>
<p>The specification is the programming. The syntax is just notation.</p>
<p>Vibe coding is genuinely different from COBOL in one important sense: the interface change is more dramatic. Natural language processed by a model that can write working TypeScript from a vague description is qualitatively new. The gap between “what you type” and “what runs” has never been smaller.</p>
<p>But the gap between “what you type” and “what you actually wanted” is exactly as large as it’s always been. Possibly larger, because the tool is so capable that it confidently fills in every unspecified detail, silently, in ways that seem reasonable until they’re not.</p>
<p>My rate limiter reset at UTC midnight because I didn’t say it shouldn’t. The agent wasn’t wrong. I was underspecified.</p>
<p>What vibe coding has genuinely changed: the syntax, the boilerplate, the standard implementations of standard patterns are now basically free. A solo developer with Claude Code can ship in a week what used to take a team a month. That’s real leverage and I use it every day.</p>
<p>What hasn’t changed: the irreducible core of the job — figuring out with enough precision what you want the computer to do — is still entirely human work. And based on sixty-five years of running this experiment, there’s a reasonable argument that it’s <em>definitionally</em> human work. When you get specific enough about a problem to get a working solution, you’ve already done the programmer’s job. You might be doing it in plain English now instead of Python. You’re still doing it.</p>
<p>The developer job is changing. Less time on syntax, more time on the thinking that was always the hard part. More time arguing with your tools about exactly what you meant. More time specifying the edge cases before the tool invents its own.</p>
<p>If you’ve ever wanted to spend less time fighting TypeScript compiler errors and more time actually thinking about what you’re building — genuinely, that part is better now.</p>
<p>But the thinking is still yours.</p>
<p>The programmers are still here. They’ve been here since 1959. They’ll be here after vibe coding. They just keep getting better tools.</p>
<p>Learn from the evidence. It’s sixty-five years old and it’s not subtle.</p>
]]></content:encoded></item><item><title>What Chinese Factories Taught Me About Prompting Claude Code</title><link>https://lakshminp.com/2026/03/claude-code-prompt-engineering/</link><pubDate>Tue, 03 Mar 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/03/claude-code-prompt-engineering/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><description>A few weeks ago, I fell down a Hacker News rabbit hole at 11pm. Someone had posted a manufacturing post-mortem — one of those beautiful, painful essays where a hardware founder documents exactly how badly they got burned.
This founder had designed a custom lamp. Spent months prototyping. Found a factory in Shenzhen. Shipped 500 units.
When the boxes arrived, the light-entry holes had been used as casting pour-points — the factory needed somewhere to pour the material, saw the holes, and went with it. The cable tails were two centimeters instead of ten. The knobs didn’t fit because the powder coating added thickness that nobody put in the spec. Everything technically matched the purchase order. Nothing actually worked.</description><content:encoded><![CDATA[<p>A few weeks ago, I fell down a Hacker News rabbit hole at 11pm. Someone had posted a manufacturing post-mortem — one of those beautiful, painful essays where a hardware founder documents exactly how badly they got burned.</p>
<p>This founder had designed a custom lamp. Spent months prototyping. Found a factory in Shenzhen. Shipped 500 units.</p>
<p>When the boxes arrived, the light-entry holes had been used as casting pour-points — the factory needed somewhere to pour the material, saw the holes, and went with it. The cable tails were two centimeters instead of ten. The knobs didn’t fit because the powder coating added thickness that nobody put in the spec. Everything technically matched the purchase order. Nothing actually worked.</p>
<p>I read that post-mortem three times. Then I read the top comment, which was one of those sentences that you immediately screenshot because it’s just too true:</p>
<p><em>“Anything you don’t specify will be done at minimum cost.”</em></p>
<p>I put my phone down. I looked at the ceiling. And then I thought about the email sender I’d had Claude Code generate that afternoon.</p>
<p>Let me tell you what I had asked for: “Send a welcome email to new users when they sign up.”</p>
<p>Let me tell you what I got: A function that sent emails. Technically correct. It looped over every new user and called the email API synchronously, one by one, waiting for each response before moving to the next. No rate limiting. No retry logic. No unsubscribe link — because I didn’t ask for one, and CAN-SPAM compliance wasn’t in the prompt. When I ran it against a list of 8,000 users, it fired all 8,000 requests in a tight loop, Gmail flagged the sending domain as a spam source within six hours, and my domain was blacklisted before I’d finished my coffee.</p>
<p>Everything sent. Nothing arrived.</p>
<p>I had been vibe coding with Claude Code for six months at that point, and I thought I was pretty good at it. I could get it to build things fast. I could chain prompts together. I had <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> files and hooks and all the trappings of someone who knew what they were doing.</p>
<p>What I didn’t understand — what the Hacker News post-mortem forced me to understand — is that I had completely misidentified what kind of relationship I was in.</p>
<p>I thought I was pair programming with a senior engineer.</p>
<p>I was issuing purchase orders to a factory.</p>
<p>This distinction sounds philosophical. It isn’t. It has concrete, expensive implications for every vibe coding prompt you write.</p>
<p>A senior engineer fills gaps with judgment. If you say “build auth,” a good senior engineer asks: what are the scale requirements? What’s the threat model? Are we storing PII? They fill the spec gaps with professional standards because they have skin in the game — it’s their name on the code, their reputation on the line, their on-call rotation if it breaks at 3am.</p>
<p>A factory fills gaps with cost optimization. If the spec doesn’t say “cable tails must be 10cm,” the factory cuts them at 2cm. Not because they’re malicious. Because that’s 8cm of wire per unit times 500 units and someone’s margin depends on it. They’re perfectly rational. They’re just optimizing for something that has nothing to do with whether your lamp works.</p>
<p>Claude optimizes for “satisfies the prompt.” That’s the whole job. Your vague prompt is its permission to take shortcuts, and it will take them — not maliciously, but with the same rational efficiency as a factory floor supervisor who notices you didn’t specify the minimum acceptable wire gauge.</p>
<p>Here’s the thing about the hardware community that I find both humbling and enraging: they figured this out decades ago. They built an entire profession around it. These people are called sourcing agents, and their whole job is translating “I want a nice lamp” into a 47-page document covering material density, wire gauge, coating thickness, packaging dimensions, UV stability ratings, and what happens to the tooling if the order falls below minimum quantity.</p>
<p>Forty-seven pages. For a lamp.</p>
<p>In vibe coding, the sourcing agent is you. Most developers have been accidentally promoted to this role without realizing it. They’re still acting like they’re talking to a colleague. They’re actually running a factory and they’re skipping the quality control, the detailed specs, and the first-article inspection — all the boring stuff that hardware people do automatically because they’ve shipped enough garbage to know better.</p>
<p>I’ve started reading Hacker News manufacturing posts specifically to steal their frameworks for this. A few things that have genuinely changed how I write prompts:</p>
<p><strong>Spec your constraints, not just your features.</strong> “Send welcome emails” is a feature request. “Send welcome emails via SES, rate-limited to 14 per second to stay under AWS sending limits, with exponential backoff and a max of 3 retries on failure, an unsubscribe link in the footer per CAN-SPAM, a plain-text fallback alongside the HTML version, and a hard skip for any address that has previously bounced or complained” is a spec. The difference isn’t intelligence — it’s the same way specifying wire gauge isn’t about distrusting your factory. It’s about understanding that factories don’t have opinions about wire gauge. They have margins.</p>
<p><strong>Inspect the first batch before commissioning the full run.</strong> Hardware founders don’t ship the first production run to customers. They order samples. They measure every dimension with calipers. The good ones fly to Shenzhen and stand on the factory floor. The developer equivalent is reading the first 200 lines of generated code before asking Claude to build the next feature on top of it. Check the database schema before building the API on top of it. Read the auth flow before adding the permissions layer. This feels slow. It is much faster than discovering that the foundation is wrong after you’ve built four floors.</p>
<p><strong>Specify what you don’t want.</strong> This one surprised me. Experienced sourcing agents reportedly spend half their spec document on exclusions. “No recycled plastic in structural components.” “No substituted components without written approval.” “No unlicensed firmware.” They’ve learned that a factory will always find the interpretation of the spec that costs them the least, so you have to close the doors. For prompts: “No inline styles. No TypeScript <code>any</code> types. No <code>console.log</code> for error handling. No <code>SELECT *</code> queries. No external dependencies unless they’re in the approved list.” The AI will not volunteer that it’s about to do these things. It will do them and move on.</p>
<p><strong>Budget time for the spec, not just the build.</strong> Hardware founders allocate somewhere between 30-40% of their project timeline to specification work. The manufacturing part — the actual production — is the smaller slice. Vibe coders typically invert this. Five percent on the prompt, ninety-five percent on generating code and then debugging the surprising things that came out of a vague prompt. The debugging is expensive. The spec is cheap.</p>
<p>The thing I keep coming back to is that using Chinese manufacturers is incredible leverage. You can build a physical product without owning a factory, without specialized tooling knowledge, without decades of manufacturing experience. It’s genuinely one of the great unlocks of the modern economy. And it works — when you write the spec correctly.</p>
<p>Using Claude to write code is the same kind of leverage. You can build things without knowing every library, without remembering every API, without holding the entire codebase in your head at once. It works. When you treat it like what it is.</p>
<p>Your prompt is a manufacturing spec. The code is the factory output. The factory will be rational, efficient, and completely indifferent to whether your product actually works.</p>
<p>Write the spec accordingly.</p>
<p>Or enjoy your two-centimeter cable tails.</p>
]]></content:encoded></item><item><title>Why Your MCP Server Will Die in Obscurity</title><link>https://lakshminp.com/2026/02/mcp-server-tool-descriptions/</link><pubDate>Thu, 26 Feb 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/02/mcp-server-tool-descriptions/</guid><category>essays</category><category>ai-coding</category><description>You built it over a weekend. The code works. Claude can technically call your tools. You added it to the config — if you’re not sure how, Stop Making Claude Code Guess covers the setup — restarted Claude Code, and — nothing. Claude doesn’t use it. Or uses it once, awkwardly, and then forgets it exists.
The problem isn’t your code. The problem is that Claude doesn’t know when to call your tools, so it doesn’t.</description><content:encoded><![CDATA[<p>You built it over a weekend. The code works. Claude can technically call your tools. You added it to the config — if you’re not sure how, <a href="https://lakshminp.com/p/stop-making-claude-code-guess" class="lnp-link">Stop Making Claude Code Guess</a> covers the setup — restarted Claude Code, and — nothing. Claude doesn’t use it. Or uses it once, awkwardly, and then forgets it exists.</p>
<p>The problem isn’t your code. The problem is that Claude doesn’t know when to call your tools, so it doesn’t.</p>
<p>This is the thing nobody tells you when you’re learning MCP: the hardest part isn’t building the server. It’s making Claude reach for it.</p>
<h1 id="how-claude-actually-chooses-your-tool"><strong>How Claude Actually Chooses Your Tool</strong></h1>
<p>When you ask Claude to do something, it’s doing a matching problem. It looks at what you asked, scans the tools available to it, reads their descriptions, and decides which one — if any — fits.</p>
<p>That last part is the lever most developers ignore. Claude doesn’t run your code to figure out what your tool does. It reads the description you wrote and makes a judgment call. If your description is vague, generic, or poorly matched to the language your users actually use, Claude will skip your tool and try something else. Or just tell you it can’t do the thing.</p>
<p>Here’s a real example. Compare these two tool descriptions for the same function:</p>
<blockquote>
<p><strong>Bad</strong>: “Query the database.”</p>
<p><strong>Good</strong>: “Look up a customer’s order history, subscription status, and recent activity by email address or customer ID. Use this when someone asks about a specific customer’s account.”</p>
</blockquote>
<p>The first one is technically accurate. The second one is what Claude can actually match against. “What’s going on with <a href="mailto:john@example.com" class="lnp-link">john@example.com</a>‘s account?” maps cleanly to the second description. It maps to nothing in the first.</p>
<p>Your tool description is a search index. Write it like one.</p>
<h1 id="the-five-ways-mcp-servers-die"><strong>The Five Ways MCP Servers Die</strong></h1>
<p><strong>1. Bad descriptions.</strong> Already covered, but it bears repeating because it’s the most common failure. Every tool, resource, and prompt deserves a description that answers: when should Claude reach for this? Include the kinds of questions or requests that should trigger it. Use the words your users actually use.</p>
<p><strong>2. Too many tools.</strong> There’s a temptation to expose everything. Every database table. Every API endpoint. Every configuration option. Resist it. A server with 30 tools is a server Claude gets confused by — and it’s also a server that quietly eats your context window before you’ve typed a word (<a href="https://lakshminp.com/p/mcp-server-context-bloat" class="lnp-link">more on that problem here</a>). It can’t reliably choose the right tool when there are 30 candidates with overlapping descriptions. The best MCP servers do one thing, maybe two, exceptionally well. If you find yourself adding a tenth tool, ask whether you’re building a server or a dumping ground.</p>
<p><strong>3. Output Claude can’t reason about.</strong> Tools that return raw JSON blobs, HTML, or binary data are tools Claude struggles to use. Claude works in text. If your tool returns <code>{&quot;data&quot;: [{&quot;id&quot;: 1, &quot;val&quot;: &quot;foo&quot;}, ...]}</code>, Claude has to parse that before it can think about it. If your tool returns “Found 3 orders: Order #1001 (shipped Jan 15), Order #1002 (pending), Order #1003 (refunded)”, Claude can work with that directly. Format your output for a reader, not a parser.</p>
<p><strong>4. Uninstallable.</strong> Most MCP servers have no README. No install instructions. No example config. No explanation of what environment variables they need. Even if someone finds your server on GitHub, if they can’t get it running in ten minutes, they close the tab. You will never hear from them again. Distribution is half the product.</p>
<p><strong>5. Solving a problem only you have.</strong> This one is uncomfortable because it’s often true. The research tool you built for your specific workflow, against your specific internal data structure, with your specific edge cases handled — it’s not a product, it’s a script. That’s fine. But don’t confuse it for something others will install. The MCP servers that spread are the ones that solve problems many developers have, in a way that requires no customization to be useful out of the box.</p>
<h1 id="what-actually-works"><strong>What Actually Works</strong></h1>
<p>The servers that get used share a few traits.</p>
<p>They have narrow scope with deep utility. Not “do 20 things mediocrely” but “do one thing so well you’d miss it if it was gone.” A good example: a server that searches Hacker News. One tool, one job — search HN, return results with scores and comment counts, formatted so Claude can reason about it immediately. That’s enough. That’s a server people actually keep installed.</p>
<p>They treat descriptions as product copy. Not documentation — copy. The description is the first thing Claude reads and the primary factor in whether your tool gets called. Write it for Claude the way you’d write an app store listing: what does this do, when do you need it, what does success look like.</p>
<p>They fail gracefully and informatively. When something goes wrong, a good tool returns “No results found for ‘X’. Try a broader search term.” A bad tool raises an exception. Claude can work with the first one. It can only apologize for the second.</p>
<p>They’re easy to install. One command. One config block. Clear documentation for what environment variables are needed and what they do. If setup takes more than five minutes, most people won’t finish.</p>
<h1 id="the-gap-this-creates"><strong>The Gap This Creates</strong></h1>
<p>Right now, MCP is early. Most servers are weekend experiments. The production-quality servers — narrow scope, excellent descriptions, graceful error handling, easy installation — are rare.</p>
<p>That gap is an opportunity. A well-built MCP server that solves a real developer problem and is easy to install can spread through Claude Code users the same way good VS Code extensions did: by word of mouth, by being genuinely useful, by being the thing you’d mention in a conversation when someone complains about the problem you solved.</p>
<p>The window isn’t permanent. In six months, there will be a lot more competition. Right now, the bar is low enough that “works reliably and has a clear description” puts you in the top 10%.</p>
<p>That’s what I’m writing about in the MCP Cookbook — a practical guide to building production MCP servers for Claude Code. Not how to write an MCP server; the official docs cover that. How to write one that people actually use.</p>
]]></content:encoded></item><item><title>What Happens When You Let 6 AI Agents Write Code at the Same Time</title><link>https://lakshminp.com/2026/01/6-ai-agents-coding-experiment/</link><pubDate>Thu, 29 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/6-ai-agents-coding-experiment/</guid><category>essays</category><category>ai-coding</category><description>Steve Yegge released Gas Town on January 1st, 2026. An agent orchestrator for Claude Code. Multiple AI agents working in parallel, coordinated through git-backed task tracking, communicating via an internal mail system. The pitch: stop babysitting one Claude session. Run twenty.
His first rule: don’t use this in its first weeks.
I used it in its first week.
Why I Couldn’t Wait I work across four projects solo. SaaS products, open source tools, content — the usual indie dev plate-spinning. Every Claude Code session I run is one session I’m not running somewhere else. The promise of parallel agents shipping code while I context-switch between projects was too compelling to resist.</description><content:encoded><![CDATA[<p>Steve Yegge released <a href="https://github.com/steveyegge/gastown" rel="external nofollow noopener" class="lnp-link">Gas Town</a> on January 1st, 2026. An agent orchestrator for Claude Code. Multiple AI agents working in parallel, coordinated through git-backed task tracking, communicating via an internal mail system. The pitch: stop babysitting one Claude session. Run twenty.</p>
<p>His first rule: don’t use this in its first weeks.</p>
<p>I used it in its first week.</p>
<h1 id="why-i-couldnt-wait"><strong>Why I Couldn’t Wait</strong></h1>
<p>I work across four projects solo. SaaS products, open source tools, content — the usual indie dev plate-spinning. Every Claude Code session I run is one session I’m not running somewhere else. The promise of parallel agents shipping code while I context-switch between projects was too compelling to resist.</p>
<p>So I installed Gas Town, added my projects as “rigs,” groomed six tasks into “beads,” and spawned six workers simultaneously.</p>
<p>My M2 MacBook responded by becoming a space heater that couldn’t render a terminal.</p>
<h1 id="what-gas-town-actually-is"><strong>What Gas Town Actually Is</strong></h1>
<p>Before I explain what went wrong, let me translate the concepts. Gas Town uses Mad Max-inspired naming, which is either charming or maddening depending on your patience.</p>
<p><strong>Town</strong> — Your workspace root (<code>~/gt/</code>). Think of it as the factory floor where everything lives.</p>
<p><strong>Rig</strong> — A project container. Each of your repos becomes a rig inside the town. Not a git clone itself, but a wrapper that manages clones, worktrees, and workers for that project.</p>
<p><strong>Beads</strong> — A git-backed issue tracker, also built by Steve. Every task, bug, or feature is a “bead” with a unique ID like <code>supabyoi-9ue</code>. They live in your repo’s <code>.beads/</code> directory, committed alongside your code. Dependencies between beads create a task graph. I wrote about <a href="https://lakshminp.substack.com/p/why-your-ai-wakes-up-every-morning" rel="external nofollow noopener" class="lnp-link">why this matters</a> — AI agents lose all context when sessions end. Beads solve this by making work persist in git. This is the piece that genuinely works well.</p>
<p><strong>Mayor</strong> — The global coordinator agent. You talk to the mayor, the mayor dispatches work. It sits above all rigs and orchestrates across projects.</p>
<p><strong>Polecat</strong> — An ephemeral worker agent. Gets spawned with a task, works in its own git worktree, signals completion, gets cleaned up. The grunt labor.</p>
<p><strong>Witness</strong> — Per-rig monitor that watches polecats. Detects stuck workers, nudges them, handles cleanup.</p>
<p><strong>Deacon</strong> — Town-level watchdog that patrols all rigs. Monitors witnesses, refineries, everything.</p>
<p><strong>Refinery</strong> — Per-rig merge queue processor. When a polecat finishes, the refinery handles the PR/merge workflow.</p>
<p><strong>Convoy</strong> — Batch tracker for related work. Group six beads into a convoy, dispatch them, track progress as a unit.</p>
<p><strong>Molecules</strong> — Reusable workflow templates. Formula defines the pattern, molecule is the running instance.</p>
<p>That’s ten concepts before you write a line of code. Steve’s mental model is a steam engine: agents are pistons, work flows through hooks, everything runs on the “Propulsion Principle” — if you find work on your hook, you execute immediately.</p>
<p>The architecture borrows from Erlang’s supervisor trees(I think) — a pattern from telecom systems where processes are organized in a hierarchy. Each parent monitors its children: if a child crashes, the parent restarts it. In Gas Town, the Deacon watches Witnesses, Witnesses watch Polecats, and failures cascade upward. This is a proven pattern that runs phone switches serving millions of calls. The catch: Erlang processes are lightweight (microseconds to spawn, kilobytes of memory). Claude Code sessions are heavy (seconds to spawn, gigabytes of memory). When each “process” is a full AI session burning tokens, the economics of cheap failure recovery invert.</p>
<h1 id="what-actually-happened"><strong>What Actually Happened</strong></h1>
<h2 id="week-one-the-learning-curve"><strong>Week One: The Learning Curve</strong></h2>
<p>The first session was pure orientation. I needed Claude to explain Gas Town to me <em>while inside Gas Town</em>. The cognitive overhead of mapping “polecat” to “worker” and “rig” to “project” consumed real mental energy that should have gone to actual work.</p>
<p>The 80/20 path is supposed to be:</p>
<pre><code>gt up          # Boot everything
gt mayor attach  # Talk to the mayor
</code></pre>
<p>In practice, <code>gt up</code> failed because the bd (beads daemon) version check timed out. This led me down a rabbit hole patching the version comparison in Go — changing <code>time.Equal()</code> to <code>time.Unix()</code> because JSON serialization was losing nanosecond precision. I was debugging the orchestrator instead of using it.</p>
<h2 id="week-two-six-polecats-and-a-space-heater"><strong>Week Two: Six Polecats and a Space Heater</strong></h2>
<p>Once things stabilized, I got ambitious. Six beads groomed, six polecats spawned:</p>
<pre><code>gt sling supabyoi-9ue supabyoi
gt sling supabyoi-abc supabyoi
# ... four more
</code></pre>
<p>Each polecat is a full Claude Code session in its own tmux pane with its own git worktree. Six of those plus a mayor, witnesses, refineries, deacons, and multiple bd daemons meant my M2 was running 20+ processes competing for resources.</p>
<p>The system didn’t crash. It degraded. Commands took minutes to respond. Shell execution broke mid-session. I had to kill processes manually and nuke the setup.</p>
<p>But here’s the thing — that wasn’t entirely Gas Town’s fault. Six concurrent Claude sessions will hammer any laptop. The real issue was that Gas Town spawned orphaned daemon processes that accumulated across restarts. I found six bd daemons running simultaneously, plus stuck <code>bd mol burn</code> processes from days ago that never cleaned up.</p>
<h2 id="the-doctor-loop"><strong>The Doctor Loop</strong></h2>
<p>Gas Town has a <code>gt doctor</code> command — a health check that reports errors and warnings. I ran it constantly.</p>
<p>First run: 1 error, 11 warnings. After <code>gt doctor --fix</code>: 4 fixed, 7 remaining. After restart: new errors. After bd daemon restart: timeout errors. After updating bd from v0.47.0 to v0.47.2: different errors.</p>
<p>Each fix revealed the next problem. The mayor’s <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> was 280 lines (should be under 30). Environment variables from dead sessions broke prefix routing. Beads databases pointed to wrong paths. Symlinks needed codesigning to avoid macOS killing the binary.</p>
<p>It felt less like using a tool and more like being a system administrator for a tool.</p>
<h2 id="what-genuinely-works"><strong>What Genuinely Works</strong></h2>
<p><strong>Beads</strong>: The git-backed issue tracker is solid. Creating tasks, tracking dependencies, finding ready work — this layer does its job. It survived every crash and restart because it’s just files in git. Steve built Beads as a standalone tool before Gas Town, and it’s the strongest foundation in the stack.</p>
<p><strong>Worktree isolation</strong>: Each worker gets its own git worktree. No merge conflicts between parallel work. Clean separation. This is the right primitive.</p>
<p><strong>The hub/worker model</strong>: Having a coordinator dispatch tasks to isolated workers is correct. The mental model of “groom beads, dispatch to workers, merge results” is sound.</p>
<p><strong>gt doctor</strong>: Despite the loop, having a comprehensive health check that can auto-fix common issues is genuinely useful infrastructure.</p>
<h2 id="what-doesnt-work-yet"><strong>What Doesn’t Work Yet</strong></h2>
<p><strong>Daemon management</strong>: Orphaned processes are the #1 pain. bd daemons accumulate, stuck processes never clean up, version checks timeout. This is being fixed — v0.5.0 added process group killing — but it was brutal in weeks one and two.</p>
<p><strong>The naming</strong>: I’m not trying to be uncharitable. But “polecat” adds zero information over “worker.” “Molecule” adds confusion over “workflow.” Every conversation about Gas Town requires a glossary. When a Hacker News commenter pointed out the irony — Steve Yegge wrote “Execution in the Kingdom of Nouns” mocking over-abstraction — it stung because it’s accurate.</p>
<p><strong>Human as dispatcher</strong>: This is the core limitation. Despite all the automation, the mayor waits for you. Issue #694 on GitHub tracks exactly this: “Mayor lacks automated dispatch patrol molecule.” Community members built external cron scripts to poke the system. That tells you everything.</p>
<p><strong>Cost and resource usage</strong>: Multiple reports of $100/hour token burn rates. DoltHub’s field test found none of the PRs were good enough to merge. The economics only work if the agents produce mergeable code reliably.</p>
<h1 id="what-im-building-instead"><strong>What I’m Building Instead</strong></h1>
<p>Gas Town taught me what I need. It also taught me what I don’t.</p>
<p>I wrote about <a href="https://lakshminp.substack.com/p/why-im-building-an-agent-orchestrator" rel="external nofollow noopener" class="lnp-link">why I’m building my own agent orchestrator</a>. It’s called <code>wt</code>. The core idea: keep the infrastructure that works (beads, worktrees, tmux), strip the ceremony that doesn’t (polecats, molecules, deacons, refineries).</p>
<p>Where Gas Town has ten concepts, <code>wt</code> has three: <strong>hub</strong>, <strong>worker</strong>, <strong>task</strong>. That’s it.</p>
<p>The hub coordinates. Workers execute in isolated worktrees. Tasks are beads with dependencies. No mail system, no witness layer, no convoy abstraction. If a worker finishes, the hub sees it in the dashboard. If a worker gets stuck, you look at the terminal. No intermediate monitoring agent needed.</p>
<p>The key difference: <code>wt</code> is a pluggable orchestrator. Each project gets its own config — yolo mode for prototypes (no tests, auto-merge, maximum speed), strict mode for production code (tests required, PR review, quality gates), or anything in between. Gas Town is one-size-fits-all. Real projects aren’t.</p>
<p>It’s early. But two weeks of wrestling Gas Town gave me the blueprint for what comes next.</p>
<h1 id="what-im-taking-away"><strong>What I’m Taking Away</strong></h1>
<p>Gas Town is a research prototype that got released into the wild. Steve warned people. I didn’t listen.</p>
<p>But I don’t regret it. Two weeks of wrestling gave me clarity about what agent orchestration actually needs:</p>
<ol>
<li>
<p><strong>Beads (or equivalent) is non-negotiable.</strong> Git-backed task tracking with dependencies is the foundation. Without it, agents have no memory across sessions.</p>
</li>
<li>
<p><strong>Worktree isolation is the right primitive.</strong> One agent, one worktree, no conflicts. Simple and correct.</p>
</li>
<li>
<p><strong>The hub/worker model works — if the hub is smart.</strong> The dispatcher problem is the real unsolved challenge. Manual dispatch defeats the purpose.</p>
</li>
<li>
<p><strong>Simplicity beats power.</strong> Three concepts (hub, worker, task) cover 90% of the use cases. Ten concepts with Mad Max names cover 95% but cost you 5x the cognitive overhead.</p>
</li>
<li>
<p><strong>Your laptop has limits.</strong> Two to three concurrent workers is practical on a MacBook. Six is aspirational. Twenty is a data center problem.</p>
</li>
</ol>
<p>Steve Yegge is doing genuinely new work here. Nobody else has shipped a multi-agent orchestrator for Claude Code with this level of ambition. The HN comment that stuck with me: “Gas Town is cackling mad laughter from someone both insane and prescient simultaneously. Today it’s insane. But expect serious versions in the future informed by these early experiments.”</p>
<p>I broke the first rule. I’d do it again. Just maybe with fewer polecats next time.</p>
]]></content:encoded></item><item><title>I Built 2 SaaS Products Vibe Coding. Here's the System That Made It Work.</title><link>https://lakshminp.com/2026/01/vibe-coding-2-saas-products/</link><pubDate>Sat, 24 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/vibe-coding-2-saas-products/</guid><category>essays</category><category>saas</category><category>ai-coding</category><description>Gene Kim and Steve Yegge’s Vibe Coding book says you’re the head chef now.
The metaphor runs through the whole thing: you’re not a line cook anymore, you’re orchestrating AI sous chefs, directing the kitchen, tasting every dish before it goes out. The developer-as-implementer era is over. Welcome to developer-as-orchestrator.
The Biryani Incident It’s a good metaphor. I buy it. But here’s the thing about being a head chef that the metaphor doesn’t quite capture: a head chef without mise en place is just a guy having a panic attack near hot surfaces.</description><content:encoded><![CDATA[<p>Gene Kim and Steve Yegge’s <a href="https://www.amazon.com/Vibe-Coding-Building-Production-Grade-Software/dp/1966280025" rel="external nofollow noopener" class="lnp-link">Vibe Coding</a> book says you’re the head chef now.</p>
<p>The metaphor runs through the whole thing: you’re not a line cook anymore, you’re orchestrating AI sous chefs, directing the kitchen, tasting every dish before it goes out. The developer-as-implementer era is over. Welcome to developer-as-orchestrator.</p>
<h2 id="the-biryani-incident"><strong>The Biryani Incident</strong></h2>
<p>It’s a good metaphor. I buy it. But here’s the thing about being a head chef that the metaphor doesn’t quite capture: a head chef without mise en place is just a guy having a panic attack near hot surfaces.</p>
<p>I know this because I’ve been that guy. Literally.</p>
<p>My wife had to leave town for a few days. “I’ll handle dinner,” I said, with the confidence of someone who has watched many cooking videos and successfully boiled pasta multiple times. I decided to make veg biryani — a dish my wife makes effortlessly, layering rice and vegetables and spices into something that tastes like it required more effort than it actually did.</p>
<p>“Prep everything first,” she told me before leaving. “Soak the basmati rice. Marinate the paneer. Chop the vegetables for layering. Have it all ready before you start cooking.”</p>
<p>Reader, I did not do this.</p>
<p>I started frying onions. While the onions were going, I realized I hadn’t marinated the paneer. So I started cubing paneer and mixing yogurt and spices. Then the onions started burning. I ran back, stirred frantically, ran back to the paneer. Remembered I needed to soak the basmati. Started the rice soaking. The onions were now definitely burned. I scraped them out, started over, but now I was behind, so I tried to do the vegetables and the new onions simultaneously while the paneer sat half-marinated&hellip;</p>
<p>An hour later I had a kitchen that looked like a crime scene, three pans with various stages of failure in them, and something that was technically edible but bore no resemblance to biryani. My wife, via video call, watched me plate this disaster with the expression of someone who had specifically warned against this exact outcome.</p>
<p>The problem wasn’t skill. I can cook. The problem was that prep and execution were bleeding into each other. I was trying to figure out what I needed while also doing the thing. And it turns out you can’t actually do both. Not well, anyway.</p>
<p>I’ve been that guy with AI sous chefs too.</p>
<p>I’ve been vibe coding since mid-2025. By “vibe coding” I mean the thing where you describe what you want in natural language and an AI writes the code. You know, the future we were promised, except the future has some sharp edges nobody mentioned in the demos.</p>
<p>Two SaaS products. Real users. Real revenue. Not toy projects, not “look ma I generated a todo app” tutorials, not the kind of thing you show off on Twitter and then quietly delete three weeks later. Actual products that people pay actual money for.</p>
<p>So when I tell you what follows, understand: this isn’t theory. This is what I learned by shipping real things and watching everything that could go wrong go wrong.</p>
<h2 id="the-markdown-hemorrhage"><strong>The Markdown Hemorrhage</strong></h2>
<p>For the first few months, I was that chef.</p>
<p>I’d sit down to implement a feature. Claude and I would get rolling. Then I’d notice a bug. Well, I’m already here, might as well fix the bug. Then while fixing the bug, I’d realize the error handling was inconsistent. Better clean that up. Oh, and there’s still context left in the window — might as well tackle that other feature I’ve been meaning to add.</p>
<p>Two hours later: three half-finished things, Claude confused about which task we’re actually doing, and code quality somewhere between “works” and “I’m not sure why.”</p>
<p>And the markdown. God, the markdown.</p>
<p>Claude, bless its heart, wanted to help me remember things. So it started creating files. <a href="http://architecture.md/" rel="external nofollow noopener" class="lnp-link">ARCHITECTURE.md</a>. <a href="http://decisions.md/" rel="external nofollow noopener" class="lnp-link">DECISIONS.md</a>. IMPLEMENTATION_NOTES.md. <a href="http://todo.md/" rel="external nofollow noopener" class="lnp-link">TODO.md</a>. <a href="http://context.md/" rel="external nofollow noopener" class="lnp-link">CONTEXT.md</a>. <a href="http://changelog.md/" rel="external nofollow noopener" class="lnp-link">CHANGELOG.md</a>. README_UPDATED.md.</p>
<p>I call this markdown hemorrhage. The AI equivalent of a kitchen where every surface is covered with prep bowls, half-chopped vegetables, and sticky notes that say “DON’T FORGET THE SAUCE” — technically documentation, practically chaos.</p>
<p>At one point I had so many markdown files that I needed another AI tool just to search through the documentation I’d created for my AI tool.</p>
<p>This was clearly insane.</p>
<p>But here’s the thing that took me embarrassingly long to figure out: the problem wasn’t the tools. The problem was me.</p>
<h2 id="one-goal-per-session"><strong>One Goal Per Session</strong></h2>
<p>I was treating every Claude session like a buffet.</p>
<p>You know how it goes. You sit down to implement a feature. While you’re implementing, you notice a bug. Well, you’re already here, might as well fix the bug. Oh, and while fixing the bug, you realize the error handling is inconsistent across the codebase. Better clean that up too. And hey, there’s still context left in the window — might as well tackle that other feature you’ve been meaning to add.</p>
<p>Two hours later, you’ve got three half-finished things, Claude is confused about which task it’s actually working on, and the code quality has degraded to “works but I’m not sure why.”</p>
<p>I call this context pollution. And once I named it, I started seeing it everywhere.</p>
<p>LLMs are bad at juggling multiple goals. This isn’t a Claude problem — it’s a fundamental thing about how these models work. When you ask them to hold multiple objectives simultaneously, they get worse at all of them. Not a little worse. <em>Dramatically</em> worse.</p>
<p>The fix sounds almost stupidly simple: one goal per session.</p>
<p>That’s it. That’s the whole trick. One goal. One session. If you discover a bug while implementing a feature, you write down the bug and you close the session. The bug gets its own session later. No “while I’m here” detours. No context pollution.</p>
<p>“But what about efficiency?” I hear you asking. “Isn’t it wasteful to end a session when there’s still context left?”</p>
<p>This is the trap. This is exactly the thinking that leads to burned onions and half-marinated paneer. The leftover context is not an asset. It’s a liability. It’s your coworker with three tasks open, doing all of them poorly, about to forget everything anyway.</p>
<p>End the session. Start fresh. One goal.</p>
<h2 id="the-mise-en-place"><strong>The Mise en Place</strong></h2>
<p>Now, this discipline only works if you have a way to track what you’re not doing.</p>
<p>If you end a session every time you discover a bug, you need somewhere for that bug to live. Otherwise you’ll forget it. The bugs pile up in your head, you context-switch mentally, and you’re back where you started.</p>
<p>This is where beads comes in.</p>
<p>Beads is a git-backed issue tracker that Claude can read and write. Steve Yegge built it (yes, that Steve Yegge — the guy who wrote the platforms rant and approximately nine million words about Emacs). The idea is simple: every task becomes a “bead.” Claude creates them, updates them, closes them. They survive compaction. They sync through git.</p>
<p>I installed it. I ran <code>bd init</code>. And then something clicked.</p>
<p>See, beads isn’t just a todo list. It’s a forcing function. When you start a session, you run <code>bd ready</code> and it shows you what’s available to work on. You pick <em>one</em>. Not three. One.</p>
<p>And when you discover a bug mid-session? You tell Claude to create a bead for it. Claude writes it down, logs the context, notes any relevant details. Then you move on. The bug exists now. It has a home. You don’t have to hold it in your head.</p>
<p>The discipline and the tool reinforce each other. One bead per session only works because beads exist to capture everything else. And beads only work because the discipline prevents you from drowning in them.</p>
<h2 id="grooming-vs-coding"><strong>Grooming vs. Coding</strong></h2>
<p>But I’m getting ahead of myself. Let me tell you about grooming.</p>
<p>In my old workflow, I’d sit down and just&hellip; start. Open Claude, describe what I wanted, begin coding. Very vibe. Very chaotic. Whatever felt right in the moment.</p>
<p>The problem is that “figuring out what to do” and “doing the thing” are completely different cognitive modes. One is divergent — you’re exploring possibilities, breaking down problems, identifying edge cases. The other is convergent — you’re executing, making decisions, writing code.</p>
<p>When you mix them, you get mush.</p>
<p>So now I run two types of sessions:</p>
<p>Grooming sessions are for thinking. I’m not coding. I’m not even planning to code in this session. I’m creating beads. Breaking down a feature into pieces. Identifying dependencies. Noting edge cases. If I think of an unrelated feature while grooming, it gets written down — for a different grooming session. No cross-contamination.</p>
<p>Coding sessions are for execution. One bead. Implement it. If I discover a bug, I note it and keep going unless it’s blocking. The bug gets groomed and coded in its own sessions later.</p>
<p>This separation is the whole game. It sounds bureaucratic. It sounds like exactly the kind of process that “vibe coding” was supposed to eliminate. But here’s the secret: this discipline is what makes vibe coding actually work at scale. Without it, you’re just generating code and hoping. With it, you’re building systems.</p>
<h2 id="a-few-other-things"><strong>A Few Other Things</strong></h2>
<p>MCPs should be loaded at project level, not globally. Every MCP eats context. If a project doesn’t need the Reddit MCP, it doesn’t get the Reddit MCP. Context is expensive. Guard it like it’s money, because in a very real sense, it is.</p>
<p>Autocompact should be off. I want to control when context resets, not have the algorithm decide for me mid-feature. Yes, this means manually managing sessions. That’s the point.</p>
<p><a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">Claude.md</a> files are more powerful than you think. I have a global one in <code>~/.claude/CLAUDE.md</code> with rules that apply everywhere. Each project gets its own with project-specific instructions. Claude reads these automatically. They’re like a pre-prompt that doesn’t eat your context window.</p>
<h2 id="what-still-doesnt-work"><strong>What Still Doesn’t Work</strong></h2>
<p>Now, here’s the part where I’m supposed to tell you it’s all solved and my workflow is perfect.</p>
<p>It’s not.</p>
<p>Debugging production issues is still clunky. I’ve got a combination of skills and MCPs that sort of works, but there’s too much manual context assembly. Something breaks in prod and I’m still spending the first 20 minutes of the session explaining the architecture before we can even start diagnosing.</p>
<p>Test-driven development doesn’t flow. The loop of “write test, see it fail, implement, see it pass” — it’s awkward. Claude wants to write everything at once. I’m still tweaking my tooling to make TDD feel natural.</p>
<p>UX work is hard. Like, fundamentally hard. Claude can scaffold UI. It can generate components. But “does this feel right?” is a human judgment call, and trying to get there through text-based iteration is like describing a painting to someone and asking them to tell you if it’s beautiful.</p>
<p>These are the walls I’m hitting. I’m building tooling to address them — an <a href="https://lakshminp.substack.com/p/why-im-building-an-agent-orchestrator" rel="external nofollow noopener" class="lnp-link">agent orchestrator</a> that tailors Claude to my specific workflow. Work in progress. If you’re the adventurous type, you can <a href="https://badri.github.io/wt/" rel="external nofollow noopener" class="lnp-link">try it now</a>.</p>
<h2 id="the-system"><strong>The System</strong></h2>
<p>So here’s the actual system, if you want to try it:</p>
<ol>
<li>
<p>Install beads: <code>npm install -g @anthropic-ai/beads &amp;&amp; bd init</code></p>
</li>
<li>
<p>Add to your global <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a>: “Check <code>bd ready</code> at session start. One bead per session.”</p>
</li>
<li>
<p>Separate grooming from coding. Different sessions. Different mindsets.</p>
</li>
<li>
<p>Resist the urge to “do more while there’s context left.” That’s the trap.</p>
</li>
<li>
<p>Protect your context. Project-level MCPs only. Kill anything you don’t need.</p>
</li>
</ol>
<p>Two SaaS products since mid-2025. All vibe coded with this system.</p>
<p>Not because the tools are magic. The tools are good, but tools are never magic. What made it work was the discipline — the willingness to be a little bit boring about context hygiene, to resist the temptation to do more, to trust that a focused session ships more than a scattered one.</p>
<p>Vibe coding without chaos. It turns out it’s not about vibing harder. It’s about vibing deliberately.</p>
<p>You’re the head chef now. But don’t forget your mise en place.</p>
<p>My wife was right, by the way. She usually is.</p>
<p><em>I’m Lakshmi. 20 years in software — ops, infrastructure, full-stack. Now solo founder using Claude Code to develop, deploy, and distribute.</em></p>
]]></content:encoded></item><item><title>Your Code Quality Doesn't Matter Anymore (And It Never Did)</title><link>https://lakshminp.com/2026/01/code-quality-doesnt-matter/</link><pubDate>Wed, 21 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/code-quality-doesnt-matter/</guid><category>essays</category><category>ai-coding</category><category>craft</category><description>A founder on Reddit recently shared that his CTO rebuilt what four third-party partners were providing — using Claude, in weeks, at a fraction of the cost.
Another commenter chimed in: their company replaced $300,000/year software with something they built in-house in under four months.
Meanwhile, over on r/SaasDevelopers, a developer is stuck at $200 MRR for eight months. Beautiful code. Great UX. Fifteen features. Asked where his users come from: “Uh, Product Hunt six months ago and some Reddit posts.”</description><content:encoded><![CDATA[<p>A founder on Reddit recently shared that his CTO rebuilt what four third-party partners were providing — using Claude, in weeks, at a fraction of the cost.</p>
<p>Another commenter chimed in: their company replaced $300,000/year software with something they built in-house in under four months.</p>
<p>Meanwhile, over on r/SaasDevelopers, a developer is stuck at $200 MRR for eight months. Beautiful code. Great UX. Fifteen features. Asked where his users come from: “Uh, Product Hunt six months ago and some Reddit posts.”</p>
<p>These two conversations are happening in parallel across the internet, and most developers haven’t connected the dots yet.</p>
<p>Here’s what’s actually happening: AI didn’t just make coding faster. It vaporized the feature moat entirely.</p>
<p><strong>The feature moat was always a lie we told ourselves.</strong></p>
<p>“If I build it better, they will come.” This was comforting. It meant the thing we’re good at — writing code — was the thing that mattered most.</p>
<p>It wasn’t true before AI. It’s aggressively not true now.</p>
<p>Your competitor can rebuild your core features in a weekend. Not because they’re brilliant. Because Claude is sitting right there, and the barrier to “good enough” has collapsed to basically zero. That integration you spent three months perfecting? Someone’s CTO just shipped an 80% version while you were reading this paragraph.</p>
<p>The YC thread frames it well: “AI mostly kills thin feature moats, not real businesses.” If your entire value proposition is “we built this thing and it works,” congratulations — you’ve built something anyone can now replicate before their coffee gets cold.</p>
<p><strong>So what’s actually defensible?</strong></p>
<p>The comments in both threads converge on the same uncomfortable answer: everything except the code.</p>
<p><strong>Distribution.</strong> The SaasDevelopers post makes the case bluntly: a mediocre product with great distribution beats a great product with no distribution. Every time. The OP claims $4.8K MRR with “decent features, nothing groundbreaking” because he publishes three SEO posts weekly and engages in five communities daily. His previous products had better code and failed under $500 MRR.</p>
<p>Whether you believe his specific numbers or not, the pattern is real. Visibility compounds. Code quality doesn’t.</p>
<p><strong>Operational complexity.</strong> The YC founder pivoted to payments specifically because it’s “harder to clone with AI.” Payments involve regulatory mess, edge cases that actually hurt people when you get them wrong, and trust that takes years to build. You can’t vibe-code your way to PCI compliance.</p>
<p><strong>Workflow embedding.</strong> One commenter nailed it: “Can a copycat ship it, but still not get adopted because switching costs and trust are the real barrier?” If yes, you might have something. If your product is a nice UI on top of an API call, you’re a feature waiting to be absorbed.</p>
<p><strong>Data that compounds.</strong> This one’s subtle but important. If your product gets better because you have data your competitors can’t easily replicate — user behavior, domain-specific training data, network effects — that’s a moat AI can’t trivially cross.</p>
<p><strong>The developer’s existential crisis.</strong></p>
<p>Here’s the part nobody wants to say out loud: for most technical founders, the skill that got them here is now table stakes.</p>
<p>You can write clean code. Great. So can Claude. You can architect systems. Wonderful. So can a junior dev with Cursor and four hours.</p>
<p>The skills that matter now are the ones developers historically dismissed as “marketing” or “sales” or “that stuff the business people do.”</p>
<p>Building an audience. Writing content that ranks. Engaging in communities without getting banned for being too promotional. Understanding what people actually want to pay for versus what’s technically impressive.</p>
<p>This is deeply annoying if you became a developer specifically to avoid talking to people.</p>
<p><strong>What to actually do.</strong></p>
<p>Stop adding features to a product nobody’s using. That’s not building — that’s procrastinating with a compiler.</p>
<p>Spend less time in your IDE and more time in the places your customers hang out. Reddit, LinkedIn, niche communities, whatever. Not to drop links. To understand what problems people are actually complaining about and whether your thing solves any of them. (It’s why I’m building <a href="https://threadhq.co/" rel="external nofollow noopener" class="lnp-link">ThreadHQ</a>.)</p>
<p>If your product can be rebuilt in weeks with AI, either pivot to something with real operational complexity, or accept that distribution is your product now and code is just the unlock.</p>
<p>The YC thread suggests payments, compliance-heavy industries, anything where “mistakes actually hurt” and trust is earned over years. The SaasDevelopers thread suggests becoming a distribution machine: 20+ platform launches, daily content, systematic visibility.</p>
<p>Both are right. Pick your poison.</p>
<p><strong>The uncomfortable synthesis.</strong></p>
<p>AI commoditized the build. What’s left is everything around it: who knows about you, who trusts you, and how painful it would be to switch away.</p>
<p>The code was never the product. Now it’s just impossible to pretend otherwise.</p>
]]></content:encoded></item><item><title>90% of Programming Skills Just Got Commoditized. The Other 10% Is Worth 1000X More.</title><link>https://lakshminp.com/2026/01/programming-skills-commoditized/</link><pubDate>Thu, 15 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/programming-skills-commoditized/</guid><category>essays</category><category>ai-coding</category><category>craft</category><description>Andrej Karpathy recently wrote something that’s been rattling around my head:
“I’ve never felt this much behind as a programmer. The profession is being dramatically refactored as the bits contributed by the programmer are increasingly sparse and between. I have a sense that I could be 10X more powerful if I just properly string together what has become available over the last year and a failure to claim the boost feels decidedly like skill issue.”</description><content:encoded><![CDATA[<p>Andrej Karpathy recently <a href="https://x.com/karpathy/status/2004607146781278521" rel="external nofollow noopener" class="lnp-link">wrote something</a> that’s been rattling around my head:</p>
<blockquote>
<p>“I’ve never felt this much behind as a programmer. The profession is being dramatically refactored as the bits contributed by the programmer are increasingly sparse and between. I have a sense that I could be 10X more powerful if I just properly string together what has become available over the last year and a failure to claim the boost feels decidedly like skill issue.”</p>
</blockquote>
<p>He then listed what this new layer looks like: agents, subagents, prompts, contexts, memory, modes, permissions, tools, plugins, skills, hooks, MCP, LSP, slash commands, workflows, IDE integrations.</p>
<p>His conclusion: “Clearly some powerful alien tool was handed around except it comes with no manual and everyone has to figure out how to hold it and operate it, while the resulting magnitude 9 earthquake is rocking the profession.”</p>
<p>I felt this in my bones.</p>
<h2 id="the-old-stack-vs-the-new-stack"><strong>The Old Stack vs The New Stack</strong></h2>
<p>The old programming stack was hard enough:</p>
<p>Hardware -&gt; OS -&gt; Language -&gt; Frameworks -&gt; Your Code</p>
<p>Years of learning. Layers of abstraction. But at least it was <em>deterministic</em>. At least there were manuals. At least Stack Overflow had answers.</p>
<p>The new stack adds a layer on top:</p>
<p>You -&gt; Prompts/Agents/Context/Memory/Tools/Modes -&gt; Code</p>
<p>This layer is fundamentally different. It’s stochastic. It’s fallible. It’s unintelligible. And it changes every few weeks.</p>
<p>There’s no certification. There’s no textbook. There’s no “Effective AI Orchestration” by Joshua Bloch. Just a bunch of people figuring it out in Discord servers and sharing CLAUDE.md files and best practices like trading cards.</p>
<h2 id="the-divide-is-already-here"><strong>The Divide Is Already Here</strong></h2>
<p>Someone on Reddit <a href="https://reddit.com/r/ClaudeAI/comments/1lquetd/the_claude_code_divide_those_who_know_vs_those/" rel="external nofollow noopener" class="lnp-link">described the pattern</a> they’re seeing on their team:</p>
<blockquote>
<p>“Two developers with similar experience working on similar tasks, but one consistently ships features in hours while the other is still debugging. At first I thought it was just luck or skill differences. Then I realized what was actually happening — it’s their instruction library.”</p>
</blockquote>
<p>They’re watching an underground collection of power users share workflows like secrets:</p>
<ul>
<li>
<p>Commands that automatically debug entire codebases</p>
</li>
<li>
<p>CLAUDE.md files that turn Claude into domain experts</p>
</li>
<li>
<p>Slash commands that turn 45-minute processes into 2-minute ones</p>
</li>
</ul>
<p>Meanwhile, most people are still typing “help me fix this bug” and wondering why their results suck.</p>
<p>As one developer put it: “The differences between someone who opens up CC for the first time and someone with tuned md files is beyond night and day.”</p>
<h2 id="the-skill-issue-is-real-but-not-the-one-you-think"><strong>The Skill Issue Is Real (But Not The One You Think)</strong></h2>
<p>Here’s what hit me about Karpathy’s framing: he called it a “skill issue.”</p>
<p>Not a tools issue. Not an access issue. Not a funding issue.</p>
<p>A <em>skill</em> issue.</p>
<p>The 10X boost exists. The leverage is real. But claiming it requires mastering something that didn’t exist two years ago and has no curriculum.</p>
<p>Someone in that same thread nailed the uncomfortable truth: “90% of traditional programming skills are becoming commoditized while the remaining 10% becomes worth 1000x more. That 10% isn’t coding — it’s knowing how to architect AI workflows.”</p>
<p>The irony is brutal. We spent years mastering syntax, frameworks, design patterns. Now an AI can generate all of that in seconds. What it <em>can’t</em> do is orchestrate itself effectively. That’s your job now.</p>
<h2 id="what-the-new-layer-actually-looks-like"><strong>What The New Layer Actually Looks Like</strong></h2>
<p>Let me make this concrete. Here’s what I’ve had to learn in the past year that wasn’t part of any CS curriculum:</p>
<p><strong>CLAUDE.md Architecture</strong></p>
<p>Your instructions file isn’t documentation. It’s programming. The structure, the phrasing, what you include vs exclude — these decisions compound across every interaction. A well-architected CLAUDE.md is worth more than a well-architected codebase.</p>
<p><strong>Context Management</strong></p>
<p>Every token matters. MCP servers eat context. Long conversations drift. You need to think about what Claude knows, what it’s forgotten, when to compact, when to start fresh. It’s memory management, but for a mind that isn’t yours.</p>
<p><strong>Prompt Design</strong></p>
<p>Not “prompt engineering” in the LinkedIn-influencer sense. Actual design. When do you give examples? When do you constrain? When do you let it explore? How do you phrase things so it doesn’t hallucinate? How do you trigger deeper thinking? These are learnable skills with massive payoff differences.</p>
<p><strong>Tool Orchestration</strong></p>
<p>MCP, skills, hooks, slash commands. Which tool for which job? When does an MCP server make sense vs a bash script vs a skill file? How do you chain them? How do you debug when the chain breaks?</p>
<p><strong>Mode Awareness</strong></p>
<p>Plan mode vs implement mode. When to let Claude explore vs when to constrain. When to use subagents. When to go linear. The <em>meta</em> of working with AI — knowing when to switch approaches — is itself a skill.</p>
<p><strong>Verification Choreography</strong></p>
<p>AI generates fast. Verification is the bottleneck. How do you structure your workflow so you’re not just rubber-stamping garbage? How do you catch the 8 production bombs before they ship? (Yes, <a href="https://lakshminp.substack.com/p/what-claude-cant-do-for-you" rel="external nofollow noopener" class="lnp-link">I wrote about this</a>.)</p>
<h2 id="the-manual-that-doesnt-exist"><strong>The Manual That Doesn’t Exist</strong></h2>
<p>An older developer on Reddit <a href="https://reddit.com/r/ClaudeAI/comments/1lquetd/the_claude_code_divide_those_who_know_vs_those/" rel="external nofollow noopener" class="lnp-link">captured the frustration</a>:</p>
<blockquote>
<p>“I started using AI about 2 years ago. I thought I was doing good, but then I started seeing all this stuff about MCP servers, md files etc and I am kind of lost. I want to learn more and I want to improve my AI skills but it’s difficult for me.”</p>
</blockquote>
<p>This is someone with decades of experience, feeling lost because the new layer has no onramp.</p>
<p>The manual doesn’t exist because the platform keeps shifting. Claude Code ships updates weekly. New features appear. Old patterns stop working. The MCP ecosystem is exploding. Skills just launched. Hooks changed. The ground won’t stop moving.</p>
<p>You can’t study for an earthquake. You can only practice surfing.</p>
<h2 id="how-im-learning-imperfectly"><strong>How I’m Learning (Imperfectly)</strong></h2>
<p>I don’t have this figured out. Nobody does. But here’s what’s working:</p>
<p><strong>Steal shamelessly</strong>. Find people who are clearly more productive and reverse-engineer their setup. Their CLAUDE.md files, their slash commands, their workflows. GitHub repos, Discord servers, Reddit threads. The good stuff is scattered but findable.</p>
<p><strong>Treat your setup as code</strong>. Version control your CLAUDE.md. Iterate on your slash commands. When something works, document why. When something fails, autopsy it. Your instruction library is a codebase now.</p>
<p><strong>Invest in meta-skills</strong>. The specific tools will change. MCP might get replaced. Claude Code might get competition. But the meta-skills — context management, prompt design, verification choreography — those transfer.</p>
<p><strong>Actually use the new features</strong>. Hooks exist. Subagents exist. Skills exist. Most people ignore them because they’re “advanced.” They’re not advanced. They’re just new. The learning curve is the moat.</p>
<p><strong>Teach to learn</strong>. Writing about this forces me to understand it. Explaining my setup to others reveals the gaps. The best way to master the new layer is to articulate it.</p>
<h2 id="the-uncomfortable-conclusion"><strong>The Uncomfortable Conclusion</strong></h2>
<p>Karpathy is right. There’s a 10X boost available. Failing to claim it is a skill issue.</p>
<p>But it’s a <em>new</em> skill. One that didn’t exist before. One that has no manual, no certification, no clear path.</p>
<p>The people figuring it out are building compound advantages. Every custom command, every refined CLAUDE.md pattern, every workflow optimization — it all stacks. The gap between those who master the new layer and those who don’t is widening fast.</p>
<p>The earthquake is still happening. The alien tool is still being figured out. The manual is being written in real-time by the people using it. And the rules are being changed as we speak/code/write.</p>
<p>Roll up your sleeves.</p>
<p><em>This is a companion to my previous essay on <a href="https://lakshminp.substack.com/p/claude-code-is-incredible-it-also" rel="external nofollow noopener" class="lnp-link">what Claude can’t do for you</a>. That one covered the old skills that still matter. This one covers the new skills you need to add.</em></p>
]]></content:encoded></item><item><title>Clean Code Is Dead. Long Live Clean Specs.</title><link>https://lakshminp.com/2026/01/clean-code-dead-clean-specs/</link><pubDate>Fri, 09 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/clean-code-dead-clean-specs/</guid><category>essays</category><category>ai-coding</category><category>craft</category><description>Steve Yegge shipped 225,000 lines of Go code he’s never read.
Let that sink in.
Beads — his coding agent memory system — is used by tens of thousands of developers daily. It’s 100% vibe coded. Yegge has never looked at a single line. Same with his new project, Gastown. Three weeks old, 100% vibe coded, never seen the code, never plans to.
His reaction to anyone uncomfortable with this? “Get out now.”</description><content:encoded><![CDATA[<p>Steve Yegge shipped 225,000 lines of Go code he’s never read.</p>
<p>Let that sink in.</p>
<p><a href="https://steve-yegge.medium.com/introducing-beads-a-coding-agent-memory-system-637d7d92514a" rel="external nofollow noopener" class="lnp-link">Beads</a> — his coding agent memory system — is used by tens of thousands of developers daily. It’s 100% vibe coded. Yegge has never looked at a single line. Same with his new project, <a href="https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04" rel="external nofollow noopener" class="lnp-link">Gastown</a>. Three weeks old, 100% vibe coded, never seen the code, never plans to.</p>
<p>His reaction to anyone uncomfortable with this? “Get out now.”</p>
<h2 id="the-heresy"><strong>The Heresy</strong></h2>
<p>For two decades, we’ve been taught that code is literature. Uncle Bob’s Clean Code. Martin Fowler’s Refactoring. Elegant variable names. Single responsibility. Code should read like prose.</p>
<p>We optimized for human comprehension because humans had to maintain it.</p>
<p>But what if that’s no longer true?</p>
<p><a href="https://www.simonhoiberg.com/" rel="external nofollow noopener" class="lnp-link">Simon Hoiberg</a> put it bluntly: “Half my code is now written by AI, and the other half is read by AI to fix bugs. Optimizing for human readability is becoming pointless.”</p>
<p>The audience for your code has changed. And it’s not you anymore.</p>
<h2 id="the-other-day-i-wrote-about-comprehension-debt"><strong>The Other Day I Wrote About Comprehension Debt</strong></h2>
<p>I argued that vibe coding creates legacy code from day one. That velocity without comprehension isn’t velocity — it’s procrastination with extra steps.</p>
<p>I still believe that. Mostly.</p>
<p>But here’s the uncomfortable follow-up question: What if comprehension debt only matters when <em>you</em> have to pay it?</p>
<p>If AI writes the code and AI debugs the code and AI refactors the code&hellip; who exactly needs to understand it?</p>
<h2 id="the-new-contract"><strong>The New Contract</strong></h2>
<p>The old contract: Write clean code so humans can read it.</p>
<p>The new contract: Write code that produces correct outcomes, verified by tests that humans can understand.</p>
<p>This is a crucial shift. The code becomes disposable infrastructure. The tests become the spec. The behavior becomes the product.</p>
<p>Steve Yegge doesn’t need to understand 225,000 lines of Go. He needs to understand what Beads should do. The tests verify that it does it. The code is just&hellip; implementation detail. An artifact. A byproduct.</p>
<h2 id="clean-specs--clean-code"><strong>Clean Specs &gt; Clean Code</strong></h2>
<p>Here’s the heretical thought experiment:</p>
<p>What if “clean code” principles should now apply to your specifications instead of your source code?</p>
<p>Think about it:</p>
<ul>
<li>
<p><strong>Readable intent</strong>: Your specs should be crystal clear. “Users can checkout with valid payment. Invalid cards show an error. Empty carts can’t checkout.”</p>
</li>
<li>
<p><strong>Single responsibility</strong>: Each spec describes one behavior. Not implementation — behavior.</p>
</li>
<li>
<p><strong>Self-documenting</strong>: Specs are the documentation that gets executed. They describe what the system should do, and you verify it actually does.</p>
</li>
<li>
<p><strong>Easy to modify</strong>: When requirements change, you update the spec first. AI updates everything else.</p>
</li>
</ul>
<p>The source code can be a tangled mess of AI-generated spaghetti. Who cares? If you can clearly specify what you want and verify you got it, the implementation is just a detail.</p>
<h2 id="the-yegge-paradox"><strong>The Yegge Paradox</strong></h2>
<p>Here’s what’s wild. In the Vibe Coding book Yegge co-authored with Gene Kim, “Steve” is described as reviewing 10,000 lines of code a day, throwing away 10 lines for every line kept.</p>
<p>Wait. He reviews code? I thought he never looks at it?</p>
<p>The answer, I think, is this: He reviews <em>outcomes</em>. He reviews test results. He reviews whether the thing works. He’s not reading code for elegance or comprehension. He’s running it, breaking it, verifying it.</p>
<p>The code review has become a behavior review.</p>
<h2 id="what-this-means-for-you"><strong>What This Means for You</strong></h2>
<p>I’m not saying burn your Clean Code book. (Okay, maybe I am. That thing is 400 pages of what could’ve been a blog post.)</p>
<p>But consider this workflow:</p>
<ol>
<li>
<p><strong>Specify the behavior</strong> — in plain language. “Users can checkout with valid payment. Invalid cards show an error. Empty carts can’t checkout.”</p>
</li>
<li>
<p><strong>Let AI write the tests</strong> — it turns your specs into executable verification</p>
</li>
<li>
<p><strong>Let AI write the implementation</strong> — who cares if it’s ugly</p>
</li>
<li>
<p><strong>Verify the outcomes</strong> — does it do what you specified? Try to break it. Edge cases covered?</p>
</li>
<li>
<p><strong>Ship it</strong> — the code is a means to an end</p>
</li>
</ol>
<p>If something breaks, you don’t debug the code. You describe the broken behavior. AI writes a failing test. AI fixes the implementation. You verify the outcome. You never had to understand the implementation. You just had to understand what you wanted.</p>
<h2 id="the-catch"><strong>The Catch</strong></h2>
<p>There’s always a catch.</p>
<p>This only works if your specifications are actually good. If your specs are vague, incomplete, missing edge cases — you’re in the worst of both worlds. Incomprehensible code that doesn’t even do what you need.</p>
<p>That’s not vibe coding. That’s vibes-all-the-way-down coding. And that’s how you get 18 out of 20 CTOs reporting production disasters.</p>
<p>The discipline has to go somewhere. If you’re not putting it into clean code, you damn well better be putting it into clear specifications and ruthless outcome verification.</p>
<h2 id="the-real-skill-shift"><strong>The Real Skill Shift</strong></h2>
<p>Old skill: Writing elegant, maintainable code that other humans can understand.</p>
<p>New skill: Specifying behavior precisely and verifying outcomes ruthlessly.</p>
<p>The developers who thrive won’t be the ones who write the cleanest code. They’ll be the ones who can articulate exactly what they want. Who can break their own systems. Who can look at a feature and immediately think of ten ways it could fail.</p>
<p>Code literacy is becoming specification literacy. The new “clean code” is clear intent.</p>
<h2 id="the-uncomfortable-conclusion"><strong>The Uncomfortable Conclusion</strong></h2>
<p>We spent twenty years optimizing for human readers who are increasingly being replaced by AI readers.</p>
<p>Maybe Steve Yegge is right. Maybe the code doesn’t matter. Maybe it never really mattered — we just didn’t have anything better.</p>
<p>What matters is: Does it work? Can you prove it? Can you verify it still works after changes?</p>
<p>Clean code was a proxy for those questions. A good heuristic when humans had to debug.</p>
<p>Clean specs answer those questions directly.</p>
<p>The code is dead. Long live the specs.</p>
<p><em>This is a follow-up to my <a href="https://lakshminp.substack.com/p/the-invisible-tax-you-pay-when-you" rel="external nofollow noopener" class="lnp-link">recent essay</a> on comprehension debt. The tension is real: you need to understand the problem deeply enough to specify it clearly, but maybe not the implementation at all. Where that line is&hellip; I’m still figuring out.</em></p>
]]></content:encoded></item><item><title>The Invisible Tax You Pay When You Vibe Code</title><link>https://lakshminp.com/2026/01/vibe-coding-comprehension-debt/</link><pubDate>Wed, 07 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/vibe-coding-comprehension-debt/</guid><category>essays</category><category>ai-coding</category><description>You shipped the feature. Tests pass. PR merged.
Two weeks later, something breaks and you open that file. You stare at the code. Your code. Code you wrote. Code that works.
You have no idea what it does.
Welcome to comprehension debt.
The Debt Nobody Talks About Technical debt is code you know is bad. Comprehension debt is code you don’t understand well enough to know if it’s bad.
There’s a crucial distinction here that a Reddit commenter nailed: “We’re getting correct code, but not right code.” The code runs. It passes tests. But ask someone why it makes specific design choices, why certain patterns were used, why the architecture looks the way it does — and the answer is often “Copilot put it there.”</description><content:encoded><![CDATA[<p>You shipped the feature. Tests pass. PR merged.</p>
<p>Two weeks later, something breaks and you open that file. You stare at the code. <em>Your</em> code. Code you wrote. Code that works.</p>
<p>You have no idea what it does.</p>
<p>Welcome to comprehension debt.</p>
<h2 id="the-debt-nobody-talks-about"><strong>The Debt Nobody Talks About</strong></h2>
<p>Technical debt is code you know is bad. Comprehension debt is code you don’t understand well enough to know if it’s bad.</p>
<p>There’s a crucial distinction here that a Reddit commenter nailed: “We’re getting correct code, but not <em>right</em> code.” The code runs. It passes tests. But ask someone why it makes specific design choices, why certain patterns were used, why the architecture looks the way it does — and the answer is often “Copilot put it there.”</p>
<p>With AI coding assistants, we can now generate working code faster than we can understand it. Claude writes 200 lines. Tests pass. Ship it. Next feature.</p>
<p>Repeat this fifty times and you’ve got a codebase that works but might as well be written by a stranger. Because functionally, it was.</p>
<h2 id="legacy-code-on-arrival"><strong>Legacy Code on Arrival</strong></h2>
<p>Here’s the uncomfortable truth that r/programming figured out: vibe coding is legacy code from day one.</p>
<p>One commenter called it “the payday loan of technical debt.” You’re borrowing velocity from your future self at predatory interest rates.</p>
<p>A freelance developer with 8 years of experience recently described a pattern he’s seeing repeatedly: companies paying good money for internal software that barely works. Tons of errors, unreasonably slow, security flaws everywhere. When he looks at the codebase, the same telltale signs: AI-generated comments, algorithms that make no sense, inconsistent patterns. Yes, it mostly works. But it works terribly.</p>
<p>In one case, a designer with CSS knowledge but not much more created a full React app with AI. When they hired a freelancer to fix it up, he deleted 90 files out of 100.</p>
<p>That’s not technical debt. That’s a technical foreclosure.</p>
<h2 id="where-it-hurts"><strong>Where It Hurts</strong></h2>
<p>Comprehension debt doesn’t show up on sprint boards. It shows up when:</p>
<ul>
<li>
<p><strong>Debugging takes 10x longer</strong> because you’re reverse-engineering your own code. Your <em>own</em> code. Like some kind of archaeologist excavating your past self’s decisions.</p>
</li>
<li>
<p><strong>Small changes require big rewrites</strong> because you can’t safely modify what you don’t understand.</p>
</li>
<li>
<p><strong>You can’t explain the system to anyone</strong>, including future you. Especially future you.</p>
</li>
<li>
<p><strong>Architecture decisions compound badly</strong> because each layer is built on foggy assumptions and vibes.</p>
</li>
</ul>
<p>Everyone talks about vibe coding. Nobody talks about vibe debugging. There’s a reason for that.</p>
<p>The irony: you used AI to go faster, but now you’re slower because you have to re-learn your own codebase every time you touch it.</p>
<h2 id="the-skill-atrophy-problem"><strong>The Skill Atrophy Problem</strong></h2>
<p>Here’s what worries senior developers: the heavier you lean on AI, the more your own skills degrade.</p>
<p>Someone described AI coding assistants as “a hyper-intelligent, infinitely patient junior developer.” Another added: “overconfident and unable to learn.”</p>
<p>That’s the trap. A junior developer eventually becomes a senior developer. They remember painful mistakes. Their understanding of your system grows over time.</p>
<p>AI doesn’t. Every conversation starts fresh. It will confidently suggest the same antipattern tomorrow that you rejected today. And if you’ve stopped exercising your own judgment because the AI handles it, you won’t catch it.</p>
<p>Ironically, the only people who should be leaning heavily on AI for code generation are people who are already experts. They can spot when it’s wrong. Everyone else is just accumulating comprehension debt they can’t even see.</p>
<h2 id="fighting-back"><strong>Fighting Back</strong></h2>
<p>You don’t have to understand everything. That’s the whole point of abstraction. But you need to understand <em>enough</em>.</p>
<p><strong>1. The Five-Minute Rule</strong></p>
<p>After AI generates code, spend five minutes actually reading it. Not skimming. Reading. Like with your eyeballs.</p>
<p>If you can’t explain what it does to a rubber duck, you’ve got comprehension debt.</p>
<p><strong>2. Write the Comments Yourself</strong></p>
<p>Don’t let AI write comments. Write them yourself, in your own words. If you can’t write the comment, you don’t understand the code.</p>
<p>One veteran developer pointed out: “Given that very few people comment the code, if there are comments at all it’s AI generated.” Comments have become a smell for AI slop, not a sign of good documentation.</p>
<p><strong>3. Draw the Damn Diagram</strong></p>
<p>For any non-trivial flow, sketch the data path. Boxes and arrows. Takes two minutes. Forces you to understand the actual architecture, not the architecture you assume exists.</p>
<p><strong>4. Refactor Before You Forget</strong></p>
<p>The best time to refactor AI-generated code is immediately after it works. You’ve got context. You’ve got momentum. Wait two weeks and that context is gone forever.</p>
<p>Future you will not remember. Future you has problems of their own.</p>
<p><strong>5. Your PR, Your Responsibility</strong></p>
<p>“Copilot put it there” is not an acceptable answer in a code review. It’s the same as saying “I don’t know, it was the first autocomplete option.”</p>
<p>The AI is a tool, like your IDE. At the end of the day, you’re responsible for every line in your PR. If you can’t defend the code, you shouldn’t be shipping the code.</p>
<h2 id="the-uncomfortable-truth"><strong>The Uncomfortable Truth</strong></h2>
<p>I’m not saying go back to writing everything by hand. That ship has sailed.</p>
<p>AI leverage is real. Use it.</p>
<p>But leverage without understanding is just deferred confusion. Every line of code you don’t understand is a question you’ll have to answer later, usually at 2am, usually when something is on fire.</p>
<p>Those championing AI focus on the speed something new can be developed. But in the long term, the real difficulty is how easily it can be maintained. And you can’t maintain what you don’t understand.</p>
<p>The developers who’ll thrive aren’t the ones who generate the most code. They’re the ones who maintain a sustainable ratio between code shipped and code understood.</p>
<p>Velocity without comprehension isn’t velocity.</p>
<p>It’s procrastination with extra steps.</p>
]]></content:encoded></item><item><title>Your MCP Servers Are Eating Your Context</title><link>https://lakshminp.com/2026/01/mcp-server-context-bloat/</link><pubDate>Mon, 05 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/mcp-server-context-bloat/</guid><category>essays</category><category>ai-coding</category><description>I love MCP. Model Context Protocol is genuinely one of the best things to happen to Claude Code.
I also hate MCP.
Because every MCP server I add is another pile of tool definitions crammed into my context window. Supabase. Betterstack. Sentry. Playwright. Each one brings 5-15 tools. That’s 40+ tool definitions sitting there, burning tokens, even when I’m just asking Claude to fix a typo.
The technical term for this is “token bloat.” The accurate term is “I’m paying for tools I’m not using.”</description><content:encoded><![CDATA[<p>I love MCP. Model Context Protocol is genuinely one of the best things to happen to Claude Code.</p>
<p>I also hate MCP.</p>
<p>Because every MCP server I add is another pile of tool definitions crammed into my context window. Supabase. Betterstack. Sentry. Playwright. Each one brings 5-15 tools. That’s 40+ tool definitions sitting there, burning tokens, even when I’m just asking Claude to fix a typo.</p>
<p>The technical term for this is “token bloat.” The accurate term is “I’m paying for tools I’m not using.”</p>
<h1 id="the-obvious-solution-that-doesnt-work"><strong>The Obvious Solution (That Doesn’t Work)</strong></h1>
<p>“Just load MCPs on demand!”</p>
<p>Revolutionary concept. Except Claude Code doesn’t support hot-reloading MCP servers. You pick your MCPs at session start, and that’s your life now. Want to add Sentry mid-session? Restart. Lose your context. Start over.</p>
<p>Nobody should have to live like that.</p>
<h1 id="the-agent-escape-hatch"><strong>The Agent Escape Hatch</strong></h1>
<p>Here’s where it gets interesting.</p>
<p>Claude Code has agents. Agents can spawn with specific tools. So naturally, I thought: what if I keep my main session lean, and spawn agents when I need MCP access?</p>
<p>Main session stays clean. Agent does the Supabase query. Returns results. Everybody’s happy.</p>
<p>Except.</p>
<h1 id="the-inheritance-problem"><strong>The Inheritance Problem</strong></h1>
<p>Agents inherit MCP tools from their parent session.</p>
<p>Read that again.</p>
<p>If I want my debug agent to call Supabase, Supabase MCP must be loaded in my main session. The agent can <em>restrict</em> which tools it uses, but it can’t access tools the parent doesn’t have.</p>
<p>So I’m back to loading everything upfront. The bloat remains. The horror.</p>
<h1 id="poor-mans-mcp"><strong>Poor Man’s MCP</strong></h1>
<p>Fine. If agents can’t get MCP tools independently, maybe they don’t need MCP at all.</p>
<p>Agents have Bash. Bash has curl. These services have REST APIs.</p>
<p>What if I wrote thin wrapper scripts?</p>
<pre><code>debug-api sentry-issue PROJ-123
debug-api supabase-query users &quot;id=eq.abc123&quot;
debug-api betterstack-logs &quot;error&quot; --from &quot;2024-01-15&quot;
</code></pre>
<p>Each wrapper hits the API directly, returns JSON. Agent calls wrappers, correlates results, returns findings. Main session stays lean.</p>
<p>I even started designing a mini-spec. Self-describing tools via <code>--tools</code>. Consistent JSON envelope. Exit codes for quick status checks.</p>
<p>MCP-lite. Poor man’s MCP. Whatever you want to call it.</p>
<p>It would work. But I’d be rebuilding what MCP already does, just&hellip; worse.</p>
<h1 id="wait-why-cant-agents-just-call-mcp-directly"><strong>Wait. Why Can’t Agents Just Call MCP Directly?</strong></h1>
<p>This is where my brain finally caught up.</p>
<p>MCP servers are just processes. They communicate via JSON-RPC over stdio. Claude Code starts them, maintains connections, sends calls.</p>
<p>An agent with Bash could do the same thing.</p>
<p>Start server. Send JSON-RPC. Parse response. Kill server.</p>
<p>No inheritance needed. The agent IS the MCP client.</p>
<h1 id="the-tool-that-already-exists"><strong>The Tool That Already Exists</strong></h1>
<p>Before I started writing my own MCP client in bash (a decision I would have regretted), I searched.</p>
<p><a href="https://github.com/f/mcptools" rel="external nofollow noopener" class="lnp-link">mcptools</a> exists.</p>
<pre><code>brew install f/tap/mcp

# List available tools
mcp tools @supabase/mcp-server

# Call a tool directly
mcp call @supabase/mcp-server query '{&quot;sql&quot;: &quot;SELECT * FROM users&quot;}'
</code></pre>
<p>Start server. Make call. Get result. Server shuts down.</p>
<p>This is the missing piece.</p>
<h1 id="the-pattern-im-testing"><strong>The Pattern I’m Testing</strong></h1>
<pre><code>Main Session (ZERO MCP tools loaded)
    |
    └── spawn debug-backend agent
            |
            ├── mcp call @supabase/mcp-server query '{...}'
            ├── mcp call @sentry/mcp-server get-issue '{...}'
            └── mcp call @betterstack/mcp-server search '{...}'
            |
            Returns structured findings
</code></pre>
<p>Main session keeps full conversation context. Agent spawns with just Bash. Agent discovers and calls MCP tools on-demand. Zero token bloat.</p>
<p>For frontend debugging, same pattern with Playwright:</p>
<pre><code>mcp call @playwright/mcp-server navigate '{&quot;url&quot;: &quot;...&quot;}'
mcp call @playwright/mcp-server screenshot '{}'
</code></pre>
<h1 id="what-im-still-figuring-out"><strong>What I’m Still Figuring Out</strong></h1>
<p>I haven’t battle-tested this yet. Open questions:</p>
<ul>
<li>
<p><strong>Auth handling</strong>: Do all MCP servers pick up env vars correctly when spawned fresh?</p>
</li>
<li>
<p><strong>Cold start latency</strong>: Is spawning a server per-call too slow for rapid iteration?</p>
</li>
<li>
<p><strong>Error recovery</strong>: What happens when the MCP server crashes mid-call?</p>
</li>
<li>
<p><strong>Which servers play nice</strong>: Some MCP servers might not like the start-stop lifecycle.</p>
</li>
</ul>
<p>If you try this pattern, let me know what breaks.</p>
<h1 id="the-punchline"><strong>The Punchline</strong></h1>
<p>I spent hours designing “MCP-lite” before realizing I could just&hellip; call MCP directly from agents.</p>
<p>Learn from my suffering.</p>
<p>The tools exist. The pattern is sound. The token bloat is optional.</p>
<p>Now I just need to actually use this for a month and see what explodes.</p>
]]></content:encoded></item><item><title>I Watched AI Generate a Perfect Todo App in 3 Minutes. Then I Spent 3 Days Fixing It.</title><link>https://lakshminp.com/2025/11/ai-todo-app-3-minutes-3-days/</link><pubDate>Fri, 07 Nov 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/11/ai-todo-app-3-minutes-3-days/</guid><category>essays</category><category>ai-coding</category><description>Every AI coding tool demo starts the same way.
“Build me a todo app.”
Four words. Maybe ten seconds of typing. Then you sit back and watch the magic: files appear, databases materialize, endpoints generate themselves. The AI spins up authentication, adds a sleek frontend, writes tests. Three minutes later, you have a working application.
It’s impressive. It’s seductive. And for production software you’ll maintain for years, it’s a starting point at best—not a solution.</description><content:encoded><![CDATA[<p>Every AI coding tool demo starts the same way.</p>
<p>“Build me a todo app.”</p>
<p>Four words. Maybe ten seconds of typing. Then you sit back and watch the magic: files appear, databases materialize, endpoints generate themselves. The AI spins up authentication, adds a sleek frontend, writes tests. Three minutes later, you have a working application.</p>
<p>It’s impressive. It’s seductive. And for production software you’ll maintain for years, it’s a starting point at best—not a solution.</p>
<h2 id="the-demo-that-sells-vs-the-code-you-ship"><strong>The Demo That Sells vs. The Code You Ship</strong></h2>
<p>I’ve spent five months deep in AI coding tools—Claude Code, claude-flow, and everything in between. I’ve watched hundreds of demos. I’ve read the marketing. And I’ve built actual production SaaS applications.</p>
<p>Here’s what the demos won’t tell you: that three-minute todo app works because it makes a thousand architectural decisions you never specified. And the moment your requirements diverge from those invisible assumptions, the whole thing falls apart.</p>
<p>Let me show you what I mean.</p>
<h2 id="the-eight-decisions-that-actually-matter"><strong>The Eight Decisions That Actually Matter</strong></h2>
<p>When you say “build me a todo app,” you think you’re giving clear instructions. But try building real production software and you’ll immediately hit these questions:</p>
<p><strong>1. JWT Claims Structure</strong></p>
<ul>
<li>
<p>What exact fields go in your JWT payload?</p>
</li>
<li>
<p>Do you store roles as an array or a single string?</p>
</li>
<li>
<p>Where do permissions live? In the token? In the database?</p>
</li>
<li>
<p>Do you include user metadata or just an ID?</p>
</li>
</ul>
<p>The demo picks one. It might not be the one you need. And changing it later? That’s not a refactor. That’s rearchitecting your entire auth system.</p>
<p><strong>2. Token Rotation</strong></p>
<ul>
<li>
<p>15-minute access tokens with 7-day refresh tokens?</p>
</li>
<li>
<p>Refresh token rotation on every use?</p>
</li>
<li>
<p>Where do you store refresh tokens—database, Redis, or in-memory?</p>
</li>
<li>
<p>httpOnly cookies or localStorage?</p>
</li>
</ul>
<p>The demo makes a choice. You won’t know what it chose until you’re debugging your third session timeout bug in production.</p>
<p><strong>3. UI Library</strong></p>
<ul>
<li>
<p>shadcn/ui? Material-UI? Chakra? Ant Design? Headless UI?</p>
</li>
<li>
<p>Tailwind CSS or CSS-in-JS?</p>
</li>
<li>
<p>Which component patterns?</p>
</li>
</ul>
<p>“Use a modern UI library” means nothing. I needed shadcn/ui specifically because it works with my design system, ships minimal JavaScript, and uses Tailwind. The demo gave me Material-UI. That’s not a theme change—that’s rebuilding the entire frontend.</p>
<p><strong>4. Stripe Integration</strong></p>
<ul>
<li>
<p>Checkout flow or Payment Intents?</p>
</li>
<li>
<p>Subscription model or one-time payments?</p>
</li>
<li>
<p>Customer portal or custom UI?</p>
</li>
<li>
<p>Which webhooks do you handle?</p>
</li>
</ul>
<p>The difference isn’t cosmetic. Checkout and Payment Intents are architecturally different. Choosing wrong means rewriting your entire billing integration.</p>
<p><strong>5. Email Provider</strong></p>
<ul>
<li>
<p>SendGrid? Resend? Postmark? AWS SES?</p>
</li>
<li>
<p>Template system?</p>
</li>
<li>
<p>Transactional vs. marketing?</p>
</li>
</ul>
<p>Each provider has different APIs, rate limits, pricing models, and deliverability characteristics. “Add email notifications” doesn’t specify any of this.</p>
<p><strong>6. Database ORM</strong></p>
<ul>
<li>
<p>Prisma? Drizzle? TypeORM? Kysely?</p>
</li>
<li>
<p>Type generation approach?</p>
</li>
<li>
<p>Migration strategy?</p>
</li>
</ul>
<p>Your ORM choice affects type safety, migration workflows, query performance, and deployment strategy. It’s not swappable. It’s foundational.</p>
<p><strong>7. Testing Framework</strong></p>
<ul>
<li>
<p>Vitest? Jest? Mocha?</p>
</li>
<li>
<p>Supertest for integration tests?</p>
</li>
<li>
<p>What coverage target?</p>
</li>
</ul>
<p>The testing framework dictates how you structure tests, handle mocks, and integrate with CI/CD. Changing it later means rewriting every test.</p>
<p><strong>8. Deployment Target</strong></p>
<ul>
<li>
<p>Vercel? AWS? Docker compose? Railway?</p>
</li>
<li>
<p>What Vercel-specific features do you need?</p>
</li>
<li>
<p>Environment variable strategy?</p>
</li>
<li>
<p>Database hosting (Neon? Supabase? RDS?)?</p>
</li>
</ul>
<p>Deployment isn’t the last step. It shapes your entire architecture—serverless vs. long-running, filesystem access, background jobs, caching strategies.</p>
<h2 id="the-just-refactor-it-myth"><strong>The “Just Refactor It” Myth</strong></h2>
<p>When I point this out, the response is always: “Just refactor what the AI generated.”</p>
<p>Have you actually tried this?</p>
<p>Swapping Prisma for Drizzle isn’t a find-and-replace operation. It means:</p>
<ul>
<li>
<p>Rewriting your schema in a different DSL</p>
</li>
<li>
<p>Changing how you handle migrations</p>
</li>
<li>
<p>Updating every database query</p>
</li>
<li>
<p>Modifying your type generation</p>
</li>
<li>
<p>Adjusting your seeding scripts</p>
</li>
<li>
<p>Updating your testing setup</p>
</li>
</ul>
<p>We’re not talking about an afternoon. We’re talking about days of work. And that’s for ONE of these eight decisions.</p>
<p>Change the ORM, the UI library, and the auth token structure? You’re not refactoring. You’re rebuilding.</p>
<h2 id="what-build-me-an-app-actually-produces"><strong>What “Build Me an App” Actually Produces</strong></h2>
<p>Here’s the brutal truth: autonomous AI tools generate generic boilerplate that matches their training data’s most common patterns.</p>
<p>They give you:</p>
<ul>
<li>
<p>Whatever stack is most popular on GitHub</p>
</li>
<li>
<p>Whatever patterns appear most in tutorials</p>
</li>
<li>
<p>Whatever architecture is easiest to generate</p>
</li>
</ul>
<p>They don’t give you:</p>
<ul>
<li>
<p>Your company’s conventions</p>
</li>
<li>
<p>Your infrastructure constraints</p>
</li>
<li>
<p>Your team’s expertise</p>
</li>
<li>
<p>Your product’s specific requirements</p>
</li>
</ul>
<p>The demo works because demos don’t have requirements. Real projects die in the gap between “an app” and “our app.”</p>
<h2 id="why-this-matters-for-production-code"><strong>Why This Matters for Production Code</strong></h2>
<p>If you’re at a big company with a team of 20 engineers, maybe you can absorb the rebuild cost. You have engineering hours to burn. You have people to maintain legacy code while others refactor.</p>
<p>Most of us don’t have that luxury.</p>
<p>Whether you’re building solo, on a small team, or shipping client work, you’re living with every architectural decision for years. You can’t afford to spend three days ripping out Material-UI because an autonomous tool decided that’s what “modern UI library” meant. You can’t rebuild your auth system because the JWT structure doesn’t match your API contracts. You can’t rewrite billing integration because the tool guessed Checkout when you needed Payment Intents.</p>
<p>Wrong architectural decisions compound. When you’re responsible for maintaining the code—whether that’s yourself, a small team, or a client relationship—you need to understand and own those decisions.</p>
<p>That’s why production code requires control, not autonomy.</p>
<h2 id="the-interactive-alternative"><strong>The Interactive Alternative</strong></h2>
<p>Compare that to working with Claude Code:</p>
<p><strong>Me:</strong> “Add authentication to this project.”</p>
<p><strong>Claude Code:</strong> “I can help with that. A few questions:</p>
<ul>
<li>
<p>JWT or session-based auth?</p>
</li>
<li>
<p>If JWT, what should the token payload include?</p>
</li>
<li>
<p>Where should refresh tokens be stored?</p>
</li>
<li>
<p>What’s your refresh token rotation strategy?”</p>
</li>
</ul>
<p><strong>Me:</strong> “JWT. Payload should have userId, email, roles as an array, and permissions as a nested object. Refresh tokens in database with rotation on every use. 15-minute access, 7-day refresh. httpOnly cookies.”</p>
<p><strong>Claude Code:</strong> “Got it. I’ll implement that exactly.”</p>
<p>The specification happened through dialogue. I clarified the architectural decisions before any code was written. The AI generated exactly what I specified, not what it guessed I might want.</p>
<p>When the auth system is running in production six months later and I need to debug a token issue, I understand every decision because I made every decision. I’m not reverse-engineering someone else’s assumptions. I’m working with my own architecture.</p>
<h2 id="when-autonomy-actually-works"><strong>When Autonomy Actually Works</strong></h2>
<p>Autonomy isn’t wrong—it’s just context-dependent. There are places where “just handle it” is absolutely the right answer:</p>
<ul>
<li>
<p><strong>README generation:</strong> Standard markdown structure is fine</p>
</li>
<li>
<p><strong>ESLint configuration:</strong> Default configs work for most cases</p>
</li>
<li>
<p><strong>.gitignore files:</strong> Use the templates</p>
</li>
<li>
<p><strong>Boilerplate CRUD endpoints:</strong> If they follow established patterns exactly</p>
</li>
<li>
<p><strong>Prototypes you’ll throw away:</strong> Exploration where decisions don’t matter yet</p>
</li>
</ul>
<p>These are low-stakes decisions with high standardization. Getting them “wrong” doesn’t cascade. You can change them later without rebuilding your application. Single-prompt generation shines here.</p>
<p>But authentication? Database schema? Tech stack? These are high-stakes, foundational decisions with cascading effects. This is where precision matters and guesswork fails.</p>
<h2 id="the-autonomy-illusion"><strong>The Autonomy Illusion</strong></h2>
<p>Here’s what the AI tool marketing doesn’t tell you:</p>
<p>More agents doesn’t mean better code. It means less control.</p>
<p>Sophisticated orchestration doesn’t mean better results. It means more complexity hiding the same specification problem.</p>
<p>“Just describe what you want” doesn’t work when architectural decisions require precision that natural language can’t provide.</p>
<p>I tested claude-flow—a sophisticated multi-agent system with 10+ agent templates, health monitoring, auto-scaling, 3-tier memory, and 60+ task types. Impressive infrastructure. But it still runs on string-based specifications. When I asked for shadcn/ui, there was no type safety, no validation, no guarantee the agent would interpret “shadcn/ui” as “shadcn/ui and absolutely nothing else.”</p>
<p>The specification layer is still natural language. And natural language is ambiguous.</p>
<h2 id="the-real-question"><strong>The Real Question</strong></h2>
<p>The question isn’t “Can AI build an app from a single prompt?”</p>
<p>The answer to that is yes. Absolutely. The demos prove it.</p>
<p>The real question is: “Can AI build YOUR app—with YOUR architecture, YOUR conventions, YOUR constraints—from a single prompt?”</p>
<p>The answer to that is no.</p>
<p>Not because the AI isn’t capable of generating code. It’s excellent at that.</p>
<p>But because “build me an app” leaves a thousand architectural decisions unspecified. And every one of those decisions matters when you’re shipping production software you’ll maintain for years.</p>
<h2 id="what-works-instead"><strong>What Works Instead</strong></h2>
<p>After five months of research, building real projects, and testing multiple tools, here’s what actually works:</p>
<p><strong>Start with control:</strong></p>
<ul>
<li>
<p>Make architectural decisions consciously</p>
</li>
<li>
<p>Specify tech stack, libraries, patterns explicitly</p>
</li>
<li>
<p>Use interactive tools that let you clarify requirements</p>
</li>
<li>
<p>Review and understand what’s being generated</p>
</li>
</ul>
<p><strong>Move to autonomy for execution:</strong></p>
<ul>
<li>
<p>Once patterns are established, autonomous tools can replicate them</p>
</li>
<li>
<p>Use autonomy for boilerplate that follows decided patterns</p>
</li>
<li>
<p>Let AI handle repetition, not decision-making</p>
</li>
</ul>
<p><strong>Return to control for integration:</strong></p>
<ul>
<li>
<p>Debugging requires understanding</p>
</li>
<li>
<p>Maintenance requires ownership</p>
</li>
<li>
<p>Evolution requires knowing why decisions were made</p>
</li>
</ul>
<p>The cycle is: design with control, execute with autonomy, integrate with control.</p>
<p>Not: autonomous generation followed by days of “just refactor it.”</p>
<h2 id="the-real-power-of-ai-coding"><strong>The Real Power of AI Coding</strong></h2>
<p>The promise of AI coding tools isn’t “describe an app in four words and get perfect code.”</p>
<p>The promise is: “Make architectural decisions at the speed of thought, then have those decisions implemented flawlessly.”</p>
<p>Interactive AI tools let you think at the architecture level while the AI handles the implementation level. You make decisions. The AI writes code. You maintain control and understanding. The AI handles the tedious translation from intent to syntax.</p>
<p>That’s the real 10x improvement.</p>
<p>Not “build me an app” magic that produces generic boilerplate you’ll spend days rebuilding.</p>
<p>But the ability to say “JWT with these exact claims, refresh rotation with this lifecycle, stored in httpOnly cookies” and get exactly that. First try. No guessing. No rebuilding.</p>
<h2 id="the-bottom-line"><strong>The Bottom Line</strong></h2>
<p>If you’re building serious software—production SaaS, client projects, anything you’ll maintain beyond next week—you need to understand what you’re building.</p>
<p>Autonomous tools that guess at your architecture don’t save time if you spend days fixing wrong assumptions.</p>
<p>Code you don’t understand becomes a liability the moment something breaks.</p>
<p>Decisions you never made can’t evolve with your requirements.</p>
<p>Control isn’t about micromanaging the AI. It’s about owning the architecture of software you’re responsible for maintaining.</p>
<p>The demos are impressive. The marketing is seductive. The promise of “just describe it” is tempting—and genuinely useful for the right contexts.</p>
<p>But for production software with real requirements, real constraints, and real consequences? Interactive tools that let you specify precisely what you need will outperform autonomous guesswork every time.</p>
<p>Be skeptical of demos. Demand control. Ship code you understand.</p>
]]></content:encoded></item><item><title>10 Ways to Waste Time and Money with AI Agents: A Field Guide to Self-Sabotage</title><link>https://lakshminp.com/2025/10/ai-agent-mistakes/</link><pubDate>Wed, 29 Oct 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/10/ai-agent-mistakes/</guid><category>essays</category><category>ai-coding</category><description>Money spent is obvious—we burn through tokens like a hedge fund manager through investor capital, exhausting our weekly quotas by Tuesday. Time, however, is subtle and invisible. Something I call the Anti-AI Paradox: that creeping realization that you could have hand-coded the entire feature in half the time it took to “collaborate” with your AI assistant. Let me save you some grief.
1. Being Super Vague AI models are getting smarter by the day. But they can’t read tea leaves like some digital oracle you summoned from Silicon Valley. “My ‘Schedule’ button isn’t scheduling the post.” Sure, Einstein. I can see that. Revolutionary observation.</description><content:encoded><![CDATA[<p>Money spent is obvious—we burn through tokens like a hedge fund manager through investor capital, exhausting our weekly quotas by Tuesday. Time, however, is subtle and invisible. Something I call the Anti-AI Paradox: that creeping realization that you could have hand-coded the entire feature in half the time it took to “collaborate” with your AI assistant. Let me save you some grief.</p>
<h1 id="1-being-super-vague">1. Being Super Vague</h1>
<p>AI models are getting smarter by the day. But they can’t read tea leaves like some digital oracle you summoned from Silicon Valley. “My ‘Schedule’ button isn’t scheduling the post.” Sure, Einstein. I can see that. Revolutionary observation.</p>
<p>Give me more context. What do you see in the logs? What did you <em>expect</em> to happen? What happened instead? Did it fail silently? Throw an error? Launch the nuclear codes? Eric S. Raymond’s “<a href="https://github.com/selfteaching/How-To-Ask-Questions-The-Smart-Way" rel="external nofollow noopener" class="lnp-link">How to Ask Questions The Smart Way</a>” is still devastatingly relevant after all these years, but apparently nobody got the memo.</p>
<p>The AI isn’t a mind reader—it’s a very expensive pattern matcher. Treat it accordingly.</p>
<h1 id="2-vibe-coding-in-the-truest-spirit">2. Vibe Coding in the Truest Spirit</h1>
<p>I’m going to hit “Accept” until my fingers are sore or the code does what I want. Whichever comes first. It’s like Russian roulette, but with merge conflicts.</p>
<p>No. Take a step back. <em>Talk</em> with your tool about what needs to be implemented and what the approach should be. It is imperative—not optional, not nice-to-have—that you understand it. Ask questions until you do. Don’t allow a single line of code to be written without you knowing the consequences.</p>
<p>Don’t pay the ignorance tax. The interest rates are criminal.</p>
<h1 id="3-dont-read-the-code-written-by-ai">3. Don’t Read the Code Written by AI</h1>
<p>Again, just hit “Accept” and pray to whatever deity oversees production deployments. Why read? Why think? Why have standards?</p>
<p>You need to know the consequences. How this piece affects other parts of your codebase. How the addition of a new feature might possibly break something else that’s been working fine for three months. I feel even AIs aren’t good enough at this second-order thinking. So many “You’re absolutely right!” responses to things that were <em>obvious</em> in hindsight but the AI somehow missed. <a href="https://www.reddit.com/r/ClaudeAI/comments/1ogw8ht/hot_take_youre_absolutely_right_is_a_bug_not_a/" rel="external nofollow noopener" class="lnp-link">This Reddit thread is in equal parts hilarious and terrifying.</a></p>
<p>Code review exists for a reason. Even if the author is artificial.</p>
<h1 id="4-do-multiple-changes-in-one-session">4. Do Multiple Changes in One Session</h1>
<p>This is a surefire way to confuse the heck out of the AI, and eventually yourself. Congratulations, you’ve achieved parity—you’re both lost.</p>
<p>Have one Claude/Cursor session for one unit of work. It can be a simple fix for a broken sidebar. Or preparation for something monumental, like refactoring all functions to use JWT. How do you figure out the right unit of work? Depends on context. (Your mileage may vary, batteries not included, void where prohibited.) There are <a href="https://github.com/steveyegge/beads" rel="external nofollow noopener" class="lnp-link">tools</a> to divide your “build me an image editing tool” prompt into proper AI and human digestible units of work.</p>
<p>In the JWT example, maybe all the functions are in the auth module only. Not a lot of context switching—for both carbon and silicon-based lifeforms. Even if it breaks, you can purge that commit off the face of the earth, go back to the drawing board, recoup, rethink, and re-execute.</p>
<p>But what if you club both examples in the same session? A broken sidebar comes across as a seemingly harmless fix, only to discover that the “fix” isn’t responsive, and now you need to add a new library, a consequence of which is that <code>npm run build</code> fails spectacularly. You debug this rabbit hole for an hour. Your context window explodes like a supernova. You didn’t fix the sidebar. Two hours and eight dollars in credits flew by. (Just sayin’.)</p>
<p>Which brings us to&hellip;</p>
<h1 id="5-choke-the-context-window">5. Choke the Context Window</h1>
<p>We need to be strategic with our prompts and questions. They need surgical precision, not the intellectual equivalent of a shotgun blast.</p>
<p>When you say “the register endpoint in @app.py returns 500 if the email already exists, but @utils.py already has a check for that,” you’re sending the entire 1,000-line app.py file and the 1,500-line utils.py file into your precious context window. Congratulations, you just spent $2 to ask a $0.50 question.</p>
<p>Even better: when you coded app.py and utils.py in the first place, don’t make them 5,000 lines long. Even if the AI <em>wants</em> to take any of these files into context, it will result in context hemorrhage sooner or later. Give clear instructions to the AI not to make your files and modules like Homer’s Iliad. Nobody wants to read that much Python in one sitting.</p>
<p>If your codebase was written by humans (remember those?), create units of work to refactor these monstrosities. Your future self will thank you profusely. The AI even more.</p>
<h1 id="6-dont-use-parallel-sessions">6. Don’t Use Parallel Sessions</h1>
<p>When you’re building a feature—say, WebSocket integration—and have another in the pipeline, you fire up your IDE, give clear instructions, rightsized units of work, and then&hellip; you twiddle your thumbs while the AI finishes, right? Making coffee? Checking Twitter? Contemplating the heat death of the universe?</p>
<p>Have you considered using git worktrees and parallel sessions so that they can execute independently of each other? Later, you can merge both feature branches into the parent branch. Revolutionary concept, I know.</p>
<p>Here’s another kicker: you can orchestrate an AI agent to do this entire thing for you—split into parallel units of work, create git worktrees, orchestrate parallel sessions, review and merge back the code, clean up. It isn’t a stretch to say we’re hitting technological singularity. The robots are already doing the DevOps we were too lazy to automate properly.</p>
<p>I wrote a piece related to this:</p>
<h1 id="7-mcp-overuse">7. MCP Overuse</h1>
<p>MCPs (Model Context Protocol servers) consume tokens. A <em>lot</em> of them. They’re the SUVs of the API world—powerful, useful, and absolute gas guzzlers.</p>
<p>Use prudence and exercise your own judgment here. For example, to scaffold a GitHub repo with a FastAPI boilerplate, the GitHub MCP is the slowest and costliest way. You’re better off using the <code>gh</code> CLI. Or, you know, copying a template. Revolutionary, I know.</p>
<p>Not all MCP use is stupid. But some of it is <em>really</em> stupid.</p>
<h1 id="8-mcp-underuse">8. MCP Underuse</h1>
<p>Context7 MCP is used by your AI to refer to up-to-date documentation for the libraries and APIs you use. Use the GitHub or JIRA MCP to update the task you’ve been working on. The utility value of MCPs is staggering.</p>
<p>If you’re not using MCPs where they make sense, you’re simply leaving time and money on the table. It’s like having a Swiss Army knife and only using the bottle opener. Sure, it works, but you’re missing out.</p>
<h1 id="9-dont-write-tests">9. Don’t Write Tests</h1>
<p>Did you finish a unit of work? Did you forget to write tests for that? Superb! Because this is going to come back and haunt you after weeks—possibly months—when you’re shipping something else on a tight deadline, and this thing you wrote many lifetimes ago is suddenly, inexplicably broken.</p>
<p>Every unit of work that goes in as a git commit must be tested. At least manually. Preferably with actual test cases that run in CI/CD and don’t just live in your head as “yeah, I’m pretty sure this works.”</p>
<p>Future you is going to hunt down present you with a very particular set of grievances. Don’t give them ammunition.</p>
<p>Again, a related post:</p>
<h1 id="10-dont-update-your-projects-context">10. Don’t Update Your Project’s Context</h1>
<p>Something AIs and humans have in common: context is everything. And both forget it constantly.</p>
<p>Picture this: you’re three months into a project. You open a file. “Wait, the architecture document says we use Redis to store the messages. But here we are using ZeroMQ. Let me do a git blame.”</p>
<p>Ah. You did it three weeks back. Right before going on vacation. The context that seemed <em>so obvious</em> at the time has evaporated like morning dew. You’re now an archaeologist excavating your own code, trying to understand why Past You made these decisions.</p>
<p>Update your project’s context. Maintain a living document—a README, an architecture decision record, inline comments that aren’t just “// fix later” (spoiler: you won’t). Explain <em>why</em> you made certain choices. Your AI needs this context to give you useful suggestions. Your human collaborators need it to not send you passive-aggressive Slack messages. Your future self needs it to avoid existential crises at 2 AM.</p>
<p>“We switched from Redis to ZeroMQ because the message ordering guarantees were critical for the event sourcing pattern we implemented in sprint 12.” There. Was that so hard? Now everyone—human and AI alike—can work with the actual state of the world instead of a beautiful fiction from three months ago.</p>
<h1 id="the-bottom-line">The Bottom Line</h1>
<p>AI coding assistants are powerful tools. Emphasis on <em>tools</em>. They’re not magic. They won’t read your mind, fix your architecture problems, or absolve you of the responsibility to understand your own codebase.</p>
<p>Use them strategically. Be precise. Maintain context. Write tests. Don’t let the robots drive—you’re still the one who has to explain to your manager why the production database got dropped.</p>
<p>And for the love of all that is holy, read the code before you merge it.</p>
<p>Your token budget will thank you. Your future self will thank you. And your AI assistant will stop generating those “You’re absolutely right!” responses that make you question everything.</p>
]]></content:encoded></item><item><title>A Saner Way to Use AI for Coding</title><link>https://lakshminp.com/2025/10/ai-coding-tdd/</link><pubDate>Fri, 17 Oct 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/10/ai-coding-tdd/</guid><category>essays</category><category>ai-coding</category><description>It starts innocently. You ask your AI assistant to scaffold a module, maybe add a new API route. The code appears in seconds, looks neat, and even runs. You feel unstoppable.
Then a few days later, you’re knee-deep in functions you didn’t write, variables that seem to name themselves, and imports from packages you never meant to use. One small change breaks three files. You scroll through the diff wondering who wrote this mess. Spoiler: it was you — and your AI.</description><content:encoded><![CDATA[<p>It starts innocently. You ask your AI assistant to scaffold a module, maybe add a new API route. The code appears in seconds, looks neat, and even runs. You feel unstoppable.</p>
<p>Then a few days later, you’re knee-deep in functions you didn’t write, variables that seem to name themselves, and imports from packages you never meant to use. One small change breaks three files. You scroll through the diff wondering who wrote this mess. Spoiler: it was you — and your AI.</p>
<p>I use Augment Code mostly, but honestly, this happens with every tool. The moment you hand over the steering wheel, AI does what it does best — <em>over-generate.</em> It tries to impress you with completeness, not clarity.</p>
<p>That’s when I flipped the workflow.</p>
<p>Instead of saying, “write me this feature,” I started saying, “here’s the test — make it pass.”</p>
<p>Suddenly, everything changed. The AI stopped building castles and started laying bricks.</p>
<p>When you write the tests first, you define the boundaries. You decide what “done” means. The AI only fills in the minimum needed to make the tests go green. It doesn’t have permission to invent architecture. You stay in control.</p>
<p>This is just test-driven development (TDD), but with a twist: the AI is your junior developer. You write intent; it writes implementation.</p>
<p>And the result? Not just leaner code — <em>well-tested</em> code.</p>
<p>Every piece the AI touches has a corresponding test. You don’t end up with half-baked helpers or random abstractions. You end up with code that behaves exactly as you specified — and a test suite that proves it.</p>
<p>Funny how things come around. TDD started as a discipline for human developers decades ago — a way to enforce thoughtfulness and quality. Who knew it would become the perfect antidote to AI code chaos in 2025?</p>
<p>Sometimes I even take it a step further and ask the AI to <em>suggest</em> the tests before I refine them. It’s a nice warm-up — like brainstorming edge cases before getting serious. You could call that the zeroth step. Maybe I’ll write about that next.</p>
<p>For now, try this:</p>
<p>The next time you reach for your coding assistant, don’t ask it to build something.</p>
<p>Write a failing test. Then tell the AI, “Make this pass.”</p>
<p>You’ll get cleaner code, stronger tests, and maybe — for the first time — a sense that you’re the one driving again.</p>
]]></content:encoded></item><item><title>How indie devs can vibe code fast without sinking their own ship</title><link>https://lakshminp.com/2025/10/vibe-code-fast-safely/</link><pubDate>Sun, 12 Oct 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/10/vibe-code-fast-safely/</guid><category>essays</category><category>saas</category><category>ai-coding</category><description>There’s a quiet war inside every indie developer I know.
One part of you just wants to build.
To open the editor, follow your curiosity, and see something real come alive on screen.
That’s the vibe coder in you — the part that moves fast, trusts intuition, and believes momentum creates clarity.
Then there’s the other voice.
The one whispering about tests, migrations, rate limits, and all the invisible things that keep production from burning down.</description><content:encoded><![CDATA[<p>There’s a quiet war inside every indie developer I know.</p>
<p>One part of you just wants to <em>build</em>.</p>
<p>To open the editor, follow your curiosity, and see something real come alive on screen.</p>
<p>That’s the <strong>vibe coder</strong> in you — the part that moves fast, trusts intuition, and believes momentum creates clarity.</p>
<p>Then there’s the other voice.</p>
<p>The one whispering about tests, migrations, rate limits, and all the invisible things that keep production from burning down.</p>
<p>That’s the <strong>engineer</strong> in you — the part that’s seen systems crumble and knows “we’ll fix it later” often means “we’ll fix it never.”</p>
<p>Most of us swing between the two.</p>
<p>Too much vibe, and your SaaS turns into a spaghetti monster that terrifies future you.</p>
<p>Too much discipline, and you’ll design yourself into paralysis before your first user ever logs in.</p>
<p>The balance isn’t about finding the perfect middle ground — it’s about <strong>timing</strong>.</p>
<h2 id="phase-1-vibe-for-momentum"><strong>Phase 1: Vibe for Momentum</strong></h2>
<p>When you’re starting, you don’t need architecture.</p>
<p>You need <em>proof</em>. Proof that the idea resonates, that the workflow feels good, that you can sustain your own interest long enough to see it through.</p>
<p>Ship something messy.</p>
<p>Inline CSS. Hardcoded configs. A Docker Compose file running on your laptop.</p>
<p>If it helps you learn or get feedback faster, it’s good enough.</p>
<p>At this stage, your goal is to find the <em>pulse</em> of your product — the heartbeat that makes it worth polishing later.</p>
<h2 id="phase-2-add-discipline-for-survival"><strong>Phase 2: Add Discipline for Survival</strong></h2>
<p>Once someone uses it — or worse, depends on it — your job changes.</p>
<p>You’re no longer hacking; you’re maintaining.</p>
<p>That’s when guardrails matter.</p>
<p>Not enterprise-level bureaucracy, but the indie essentials:</p>
<p>rate limits, structured logs, CI checks, and a migration plan that won’t kill your data.</p>
<p>Each layer of success earns another layer of discipline.</p>
<p>That’s how you scale without killing your momentum.</p>
<h2 id="the-indie-balance"><strong>The Indie Balance</strong></h2>
<p>Vibe coding isn’t reckless.</p>
<p>It’s how you get to momentum.</p>
<p>But discipline is how you keep it.</p>
<p>The real art of indie software isn’t just writing good code.</p>
<p>It’s knowing <strong>when</strong> to write which kind of code.</p>
<p><strong>TL;DR:</strong></p>
<p>You start as an artist. You evolve into an engineer.</p>
<p>The trick is not to silence either voice — just let them take turns driving.</p>
]]></content:encoded></item><item><title>What LLMs Reveal About Human Cognition</title><link>https://lakshminp.com/2025/10/llm-human-cognition/</link><pubDate>Mon, 06 Oct 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/10/llm-human-cognition/</guid><category>essays</category><category>ai-coding</category><description>We like to think we’re smarter than the machines we build.
And maybe we are — for now. But something odd has been happening lately.
As I’ve spent more time training, prompting, and poking large language models, I’ve started noticing… echoes.
Not just in their outputs. In their behaviours.
In the way they learn.
In how they fail.
In how they improve.
And in the way they pretend.
It started as a metaphor.</description><content:encoded><![CDATA[<p>We like to think we’re smarter than the machines we build.</p>
<p>And maybe we are — for now. But something odd has been happening lately.</p>
<p>As I’ve spent more time training, prompting, and poking large language models, I’ve started noticing… echoes.</p>
<p>Not just in their outputs. In their <strong>behaviours</strong>.</p>
<p>In the way they learn.</p>
<p>In how they fail.</p>
<p>In how they improve.</p>
<p>And in the way they pretend.</p>
<p>It started as a metaphor.</p>
<p>Now I believe it’s more than that:</p>
<p><strong>Our brains are wetware LLMs.</strong></p>
<h2 id="training-and-the-loop"><strong>Training and the Loop</strong></h2>
<p>If you’ve ever trained a model — or even just used one through an API — you start to internalize a rhythm.</p>
<ol>
<li>
<p>Feed it examples.</p>
</li>
<li>
<p>Check the outputs.</p>
</li>
<li>
<p>Reinforce the good.</p>
</li>
<li>
<p>Penalize the bad.</p>
</li>
<li>
<p>Repeat until it improves.</p>
</li>
</ol>
<p>This isn’t just machine learning.</p>
<p>This is how we learn everything.</p>
<p>When I was younger, trying to improve my Carnatic violin playing, I’d follow the same loop.</p>
<p>Play the swara.</p>
<p>Notice the sour note.</p>
<p>Replay.</p>
<p>Adjust fingering/intonation.</p>
<p>Rinse/Repeat.</p>
<p>Eventually, the feedback loop got shorter. The ear began correcting the hand before the mind even intervened.</p>
<p>That’s tuning.</p>
<p>Or when I’m studying fiction — I don’t just read Dean Koontz or Lee Child. I <strong>type out</strong> their stories, word for word. A practice technique suggested by prolific writer <a href="https://deanwesleysmith.com/" rel="external nofollow noopener" class="lnp-link">Dean Wesley Smith</a>.</p>
<p>Copying wasn’t plagiarism. It was <strong>pretraining</strong>.</p>
<p>You learn cadence by mimicry. You learn structure by absorption.</p>
<p>And then, one day, you surprise yourself with an output that feels original — but you know, deep down, the gradient came from somewhere.</p>
<h2 id="model-behaviours-we-share"><strong>Model Behaviours We Share</strong></h2>
<p>Here’s the eerie part: the more you work with LLMs, the more human they feel — not in consciousness, but in quirks.</p>
<p><strong>1. Overfitting</strong></p>
<p>LLMs that are fine-tuned too aggressively on narrow data start parroting it — losing flexibility.</p>
<p>So do we.</p>
<p>Ever meet someone who mastered one domain and can’t unlearn their habits when switching fields? That’s human overfitting.</p>
<p><strong>2. Hallucinations</strong></p>
<p>LLMs generate plausible nonsense when unsure. So do we.</p>
<p>In meetings. On first dates. During interviews.</p>
<p>Confidence is <em>not the same as</em> correctness — for both machines and minds.</p>
<p><strong>3. Context windows</strong></p>
<p>LLMs can only “see” a certain number of tokens at once.</p>
<p>So can we.</p>
<p>Ever walk into a room and forget why you went in? That’s a context window shift. Our attention span — bounded. Our memory — fallible.</p>
<p>But we can prime our context deliberately — by journaling, outlining, visualizing. Just like how you “prompt” a model better when you include prior examples.</p>
<p><strong>4. Personas</strong></p>
<p>LLMs can be given system prompts to behave a certain way: “Act like a Shakespearean actor”, “You are a helpful Linux admin”, “You’re a snarky writing coach”.</p>
<p>We do this too. We wear masks.</p>
<p>We speak differently at work than at home.</p>
<p>We switch from teacher mode to student mode.</p>
<p>We code-switch, dialect-shift, self-filter.</p>
<p>These personas aren’t fake.</p>
<p>They’re <strong>fine-tuned subsets</strong> of ourselves, optimized for task and audience.</p>
<h2 id="do-some-brains-have-more-parameters"><strong>Do Some Brains Have More Parameters?</strong></h2>
<p>Sometimes I wonder: if we stretch the metaphor, do people have different “parameter counts”?</p>
<p>Do some folks just have more neurons wired up, more memory bandwidth, more raw capacity?</p>
<p>Maybe.</p>
<p>But LLMs remind us: <strong>parameter count isn’t destiny</strong>.</p>
<p>It’s how you train.</p>
<p>What you expose yourself to.</p>
<p>What feedback you seek.</p>
<p>How often you iterate.</p>
<p>Even the largest models are dumb if they’ve been trained on trash.</p>
<p>And even a small model — carefully fine-tuned on the right data, guided with the right prompts — can outperform giants.</p>
<p>Same with people.</p>
<p>We’ve all met someone who had every advantage and squandered it.</p>
<p>We’ve all met someone else — less formally educated, less polished — who radiated clarity and depth because they <em>trained deliberately</em>.</p>
<p>It’s not about who has the most parameters.</p>
<p>It’s about who’s still in the loop.</p>
<h2 id="llm-attributes-as-human-metaphors"><strong>LLM attributes as Human Metaphors</strong></h2>
<p><strong>Zero-shot vs Few-shot Learning</strong></p>
<p>A child touching a hot stove once? Few-shot learning.</p>
<p>Reading five flashcards before a quiz? Few-shot.</p>
<p>Encountering a new idea and making sense of it because of prior abstractions? That’s zero-shot. That’s transfer.</p>
<p><strong>Prompt Injection</strong></p>
<p>Ever been influenced mid-conversation and changed your tone? That’s human prompt injection.</p>
<p>Context hijacks our behaviour more often than we care to admit.</p>
<p><strong>Temperature</strong></p>
<p>High-temperature models generate more creative outputs.</p>
<p>People too. Under constraints, some freeze(forgive the pun!). Others improvise. Your internal “temperature” — mindset, mood, caffeine level — changes how you think.</p>
<p><strong>Loss function</strong></p>
<p>For models, it’s a calculated gradient.</p>
<p>For us, it’s regret. Embarrassment. The wince of feedback.</p>
<p>Pain and fear are our backpropagation signals.</p>
<h1 id="where-do-we-excel">Where do we excel?</h1>
<p>But for all the parallels, there are crucial ways our brains still outclass even the largest models.</p>
<p>LLMs don’t <em>want</em> anything. They don’t have drive, curiosity, fear, embarrassment, or delight. They don’t learn unless someone forces them to. They don’t <em>decide</em> to improve. We do.</p>
<p>We seek out the loop. We care when we’re wrong. We revise because we <em>want</em> to get better, not because we’re re-trained on a new batch.</p>
<p>We remember emotionally. The sting of failure. The warmth of praise. The embarrassment of a bad take in public. That visceral encoding is something no model has.</p>
<p>We can self-direct. A model doesn’t wake up one day and say, “I think I need to get better at analogies.” But we do. We read something brilliant and feel inspired. We listen to a master and feel the gap. That’s not loss minimization. That’s ambition.</p>
<p>We generalize across domains in weird, leaky, beautiful ways. A lesson in Carnatic violin may improve our writing cadence. A novel may shape how we manage teams. We mix metaphors, break schemas, leap categories. LLMs struggle with that. They interpolate. We cross-pollinate.</p>
<p>We also <em>choose</em> our training data. We can decide what to consume, who to listen to, what to believe. We can uninstall toxic sources. Curate higher quality inputs. Reinforce the patterns we want to keep.</p>
<p>And unlike static models, we have agency over our fine-tuning. We can say: I don’t want to respond that way anymore. I don’t want to be that version of myself. And we can go train a better one.</p>
<p>A model may freeze its weights. But we don’t have to.</p>
<p>We’re wetware — always learning, always plastic, always in the loop.</p>
]]></content:encoded></item><item><title>The Real Skill AI Won’t Replace</title><link>https://lakshminp.com/2025/10/the-real-skill-ai-wont-replace/</link><pubDate>Thu, 02 Oct 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/10/the-real-skill-ai-wont-replace/</guid><category>essays</category><category>ai-coding</category><description>Ah yes, the mythical full stack developer. Fluent in Kubernetes and CSS. Can debug a flaky WebSocket connection and make the button pop just right in Safari 14.3. Also, fluent in four frontend frameworks, three ORMs, and—if you’re lucky—your company’s internal tooling written in Bash and tears.
It sounds impressive. Until you realize “full stack” is just corporate for “three jobs, one salary, no support.”
The original promise of full stack was noble: break down silos, build end-to-end features, own your code. But somewhere along the way, it mutated. Now it means you’re responsible for everything from designing the API schema to fixing the div that renders weird on IE11. Oh, and could you also write some Terraform while you’re at it?</description><content:encoded><![CDATA[<p>Ah yes, the mythical <em>full stack developer</em>. Fluent in Kubernetes <em>and</em> CSS. Can debug a flaky WebSocket connection <em>and</em> make the button pop just right in Safari 14.3. Also, fluent in four frontend frameworks, three ORMs, and—if you’re lucky—your company’s internal tooling written in Bash and tears.</p>
<p>It sounds impressive. Until you realize “full stack” is just corporate for “three jobs, one salary, no support.”</p>
<p>The original promise of full stack was noble: break down silos, build end-to-end features, own your code. But somewhere along the way, it mutated. Now it means you’re responsible for everything from designing the API schema to fixing the div that renders weird on IE11. Oh, and could you also write some Terraform while you’re at it?</p>
<p>Let’s be honest: in 2025, “full stack” mostly means “we can’t afford to hire a team, so here’s a to-do list that spans five specialties.”</p>
<p>But here’s the twist: I still think you <em>should</em> aim for T-shaped skills. Just not the way HR thinks you should.</p>
<p>Because here’s what’s changed: we’ve now got an army of AI copilots ready to autocomplete half your job—badly. They’ll hallucinate types, suggest incorrect regex, and cheerfully rename your variables while subtly breaking the logic.</p>
<p>If you want to survive <em>this</em> stack, you need to know enough frontend, backend, infra, and AI prompt engineering to know when the machine is lying to you.</p>
<p>Being “T-shaped” doesn’t mean you’re an expert in everything. It means you can go deep where it matters (ideally in your core domain), and navigate the rest well enough to not get wrecked. It means you know when to trust ChatGPT’s code suggestion, and when to back away slowly and grep the logs yourself.</p>
<p>In other words: it’s no longer “full stack vs backend vs frontend.” It’s humans who can collaborate with AI vs humans who are about to get buried in merge conflicts and synthetic bugs.</p>
<p>So yeah, full stack as a job description might be a scam. But being <em>versatile</em>? That’s survival.</p>
<p>Especially if your AI sidekick starts suggesting you replace your Postgres schema with a single JSON blob. Again.</p>
<h3 id="source-that-inspired-this-post">Source that inspired this post:</h3>
<p>An error occurred.</p>
<p>Unable to execute JavaScript.</p>
]]></content:encoded></item><item><title>AI Made Me Faster at Procrastinating</title><link>https://lakshminp.com/2025/09/ai-made-me-faster-at-procrastinating/</link><pubDate>Mon, 29 Sep 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/09/ai-made-me-faster-at-procrastinating/</guid><category>essays</category><category>ai-coding</category><description>A few weeks ago, I caught myself doing something ridiculous.
I was surrounded by all the tools that are supposed to make me faster, smarter, more efficient—ChatGPT in one tab, Cursor humming in the IDE, Claude on standby for the longer stuff—and somehow… I had spent the entire week refactoring a feature that hadn’t shipped.
Not building. Not testing. Just circling the drain of “making it better.”
That’s when it hit me—not like lightning, but like a slow, shameful realization:</description><content:encoded><![CDATA[<p>A few weeks ago, I caught myself doing something ridiculous.</p>
<p>I was surrounded by all the tools that are supposed to make me faster, smarter, more efficient—ChatGPT in one tab, Cursor humming in the IDE, Claude on standby for the longer stuff—and somehow… I had spent the entire week refactoring a feature that hadn’t shipped.</p>
<p>Not building. Not testing. Just circling the drain of “making it better.”</p>
<p>That’s when it hit me—not like lightning, but like a slow, shameful realization:</p>
<p>If you’re not shipping weekly, you’re not taking advantage of AI. You’re just bikeshedding with fancier tools.</p>
<h2 id="the-new-age-of-productive-procrastination"><strong>The New Age of Productive Procrastination</strong></h2>
<p>We used to say we couldn’t move fast because the tools were slow. Hard to deploy. Annoying to configure. Models too dumb. Infra too brittle.</p>
<p>Now?</p>
<p>We’ve got tools that can scaffold your backend, write your tests, spin up a UI, generate your changelog, draft your release notes, and create your launch tweet — all before lunch.</p>
<p>So what do we do?</p>
<p>We spend two hours arguing with GPT5 about the <strong>tone</strong> of our 404 page.</p>
<p>AI hasn’t just made us faster. It’s made us <em>better at procrastinating</em>. We can now fine-tune our mediocrity at lightning speed. Polish things that don’t matter. Add “clever” touches no one asked for. Debate prompt styles like they’re sacred texts.</p>
<p>It’s amazing. It’s also a trap.</p>
<h2 id="your-fancy-setup-doesnt-matter-if-nothing-ships"><strong>Your Fancy Setup Doesn’t Matter If Nothing Ships</strong></h2>
<p>Every dev team and weekend hacker has access to the same models now. Same open weights, same frameworks, same “build an agent” tutorials.</p>
<p>But some teams are shipping on Fridays.</p>
<p>Others are still fiddling with prompt chains.</p>
<p>The difference isn’t in talent. It’s in rhythm.</p>
<p>Shipping frequently is the new moat. Not because it makes you look good on Twitter, but because the ground underneath is moving. Fast.</p>
<p>A month of “thoughtful planning” can kill your idea before it meets the real world.</p>
<h2 id="but-what-if-its-not-ready"><strong>But What If It’s Not Ready?</strong></h2>
<p>It won’t be.</p>
<p>It never is.</p>
<p>You’ll always want to refactor one more function. Tune one more embedding. Rename one more internal config. You’re not alone — I’ve been there. I still go there, a little too often.</p>
<p>But here’s what I’ve learned the hard way: <strong>Nothing improves faster than something you’ve already shipped.</strong></p>
<p>You can’t get feedback on a figment. You can’t iterate on invisible.</p>
<h2 id="a-week-is-enough-even-if-it-feels-too-short"><strong>A Week Is Enough (Even If It Feels Too Short)</strong></h2>
<p>Weekly shipping is a forcing function. It makes you prioritize what’s real over what’s “clever.” It exposes what matters to users versus what just looks impressive in dev chat.</p>
<p>If I can’t scope something to ship in a week, chances are I’m biting more than I can chew.</p>
<p>Some weeks it’s a feature. Some weeks it’s cleanup. Some weeks it’s a one-line fix with a changelog that makes me cringe. That still counts. That’s progress.</p>
<h2 id="the-real-timeline-of-one-shipping-week"><strong>The Real Timeline of One Shipping Week</strong></h2>
<ul>
<li>
<p>Monday: Noticed my onboarding sucked</p>
</li>
<li>
<p>Tuesday: Asked ChatGPT to rewrite it (it made it worse, then better)</p>
</li>
<li>
<p>Wednesday: Wired up telemetry to track rage clicks</p>
</li>
<li>
<p>Thursday: Built a barely-working feedback button</p>
</li>
<li>
<p>Friday: Hit deploy. Apologized in advance. Sent it to users anyway.</p>
</li>
</ul>
<p>No magic. Just momentum.</p>
<h2 id="the-ironic-truth"><strong>The Ironic Truth</strong></h2>
<p>AI is supposed to be a productivity multiplier.</p>
<p>But if you let it, it’ll multiply your perfectionism. Your indecision. Your procrastination.</p>
<p>You’ll feel productive while achieving nothing. Like a hamster with a prompt window.</p>
<p>And the worst part? It <em>feels</em> like work. It’s dangerously satisfying.</p>
<p>Which is why now, more than ever, we need to build muscle around <strong>shipping</strong>, not just building.</p>
<blockquote>
<p>In 2007, PHP creator Rasmus Lerdorf said, <em>“PHP is about as exciting as your toothbrush. You use it every day, it does the job, it is a simple tool, so what? Who would want to read about toothbrushes?”</em></p>
</blockquote>
<p>That’s the thing about good tools — they’re boring when used well. You don’t marvel at your toothbrush every morning. You just get on with it.</p>
<p>AI tools should be the same. Invisible. Unremarkable. Part of the rhythm.</p>
<p>Ship first. Marvel later.</p>
<h2 id="so-what-now"><strong>So What Now?</strong></h2>
<p>Set the bar low. Something new every week.</p>
<p>Doesn’t have to be earth-shattering. Just <em>real</em>. Just live. Just something you can point to and say, “I learned something from this.”</p>
<p>The ones who keep shipping — even small things — are the ones who win this cycle.</p>
<p>Not because they outsmarted the world. But because they stopped arguing with their tools and started using them.</p>
]]></content:encoded></item><item><title>Your AI Pair Programmer Doesn’t Need a Buffet</title><link>https://lakshminp.com/2025/09/ai-context-window-focus/</link><pubDate>Thu, 25 Sep 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/09/ai-context-window-focus/</guid><category>essays</category><category>ai-coding</category><description>I remember the day way too well. I thought I was being clever. “Let me just paste this entire 5,000-line Python file into KiloCode and let it work its magic.” Efficient. Thorough. Genius.
Except… not.
The Great Context Collapse What actually happened looked more like watching someone try to drink from a fire hose. The AI took my massive input, nodded politely, and then gave me:
Answers that had nothing to do with my question</description><content:encoded><![CDATA[<p>I remember the day way too well. I thought I was being clever. <em>“Let me just paste this entire 5,000-line Python file into KiloCode and let it work its magic.”</em> Efficient. Thorough. Genius.</p>
<p>Except… not.</p>
<h2 id="the-great-context-collapse"><strong>The Great Context Collapse</strong></h2>
<p>What actually happened looked more like watching someone try to drink from a fire hose. The AI took my massive input, nodded politely, and then gave me:</p>
<ul>
<li>
<p>Answers that had nothing to do with my question</p>
</li>
<li>
<p>Vague “advice” that could apply to literally any project</p>
</li>
<li>
<p>Confident but wrong takes on my(??) code</p>
</li>
<li>
<p>References to functions that didn’t even exist</p>
</li>
</ul>
<p>To be clear: this wasn’t a problem with KiloCode itself. It was my fault. I drowned the poor thing.</p>
<p>And yes, the reason I had a 5,000-line Python file in the first place? That’s on me too. I just kept bolting on features without a proper review. One fine day, I looked up and realized the file had become unreadable. But that disaster deserves its own post.</p>
<h2 id="whats-going-on-under-the-hood"><strong>What’s Going On Under the Hood</strong></h2>
<p>AI coding assistants only have so much “working memory” — a <em>context window</em>. When you shove 5,000 lines of code at them, a few things happen:</p>
<ul>
<li>
<p>You blow most of the available window right away</p>
</li>
<li>
<p>Your actual question gets buried under noise</p>
</li>
<li>
<p>The model has to guess what’s important in the pile</p>
</li>
<li>
<p>There’s less room left for follow-up questions</p>
</li>
</ul>
<p>It’s like asking someone to find a single paragraph in a book, then dumping an entire library on their desk and saying, <em>“good luck.”</em></p>
<h2 id="the-counterintuitive-truth"><strong>The Counterintuitive Truth</strong></h2>
<p>The less code you show your AI assistant, the better it performs.</p>
<p>I’ve found the sweet spot is usually 200–500 lines. Enough to give context, not so much that the model chokes.</p>
<h2 id="how-to-work-smarter-with-ai-code-assistants"><strong>How to Work Smarter With AI Code Assistants</strong></h2>
<ol>
<li>
<p><strong>Split files by module or function</strong></p>
<p>Don’t dump an entire repo. Just pull out the piece you care about:</p>
</li>
</ol>
<blockquote>
<p>“I need help optimizing this authentication middleware function:”</p>
<p>[paste 50–100 lines here]</p>
</blockquote>
<ol start="2">
<li>
<p><strong>Summarize bigger pieces</strong></p>
<p>Let the AI help you write summaries of large sections. Then use those summaries as context alongside the small chunk you actually care about.</p>
</li>
<li>
<p><strong>Scope your questions</strong></p>
<p>Bad: “Here’s my whole app, how can I improve it?”</p>
<p>Better: “Here’s my caching function. How can I reduce memory usage while keeping O(1) lookups?”</p>
</li>
<li>
<p><strong>Go iterative</strong></p>
<p>Smaller chunks make conversations faster. You can go back and forth ten times in the same time it takes to chew through one giant code dump.</p>
</li>
</ol>
<p>It’s basically code review etiquette. Nobody wants to wade through a 5,000-line pull request. Smaller, focused changes are easier to reason about — for humans and for AI.</p>
<p>The most useful “prompt engineering” trick I’ve learned isn’t about clever phrasing. It’s about context discipline. Show the AI just enough. Hide the rest.</p>
<p>Next time you’re tempted to paste your whole project in, remember: you’re not giving the model more to work with — you’re giving it more to drown in.</p>
]]></content:encoded></item></channel></rss>