<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Claude Code — Lakshmi Narasimhan</title><link>https://lakshminp.com/tags/claude-code/</link><description>I help developers build, deploy, and distribute their SaaS without hiring a team. Long-running notes on systems, AI internals, Carnatic music, fiction craft, and whatever else collides interestingly.</description><generator>Hugo + lakshminp theme</generator><language>en-us</language><lastBuildDate>Thu, 18 Jun 2026 00:00:00 +0000</lastBuildDate><managingEditor>Lakshmi Narasimhan</managingEditor><webMaster>Lakshmi Narasimhan</webMaster><copyright>© 2026 Lakshmi Narasimhan</copyright><atom:link href="https://lakshminp.com/tags/claude-code/feed.xml" rel="self" type="application/rss+xml"/><item><title>The "MCP Is Dead" Fight Is a Category Error</title><link>https://lakshminp.com/2026/06/mcp-vs-skills-category-error/</link><pubDate>Thu, 18 Jun 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/06/mcp-vs-skills-category-error/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><description>Skills win the solo dev. MCP wins exactly one thing. Here’s the line.
I had the headline before I had the post.
“API + Skills Is a Poor Man’s MCP” — except I was going to argue the inversion: that MCP is the rich man’s overcomplication, a server process you stood up to wrap calls your agent could already make, and the lean move was always a skill plus a CLI. Spicy. Contrarian. The kind of take that does numbers in a feed.</description><content:encoded><![CDATA[<p><em>Skills win the solo dev. MCP wins exactly one thing. Here&rsquo;s the line.</em></p>
<p>I had the headline before I had the post.</p>
<p>&ldquo;API + Skills Is a Poor Man&rsquo;s MCP&rdquo; — except I was going to argue the inversion: that <em>MCP</em> is the rich man&rsquo;s overcomplication, a server process you stood up to wrap calls your agent could already make, and the lean move was always a skill plus a CLI. Spicy. Contrarian. The kind of take that does numbers in a feed.</p>
<p>Then I made the tactical error of fact-checking myself, and the post fell apart in my hands. What follows is the wreckage, reassembled into something truer than the dunk I wanted to write.</p>
<p><strong>TL;DR:</strong> Skills and MCP aren&rsquo;t competitors — comparing them is a category error. A skill is a recipe that runs in <em>your</em> runtime; MCP is a connection to a hosted service. For a solo dev wiring up their own workflow on a coding agent, skill + CLI wins on every axis that used to favor MCP — the context-bloat and cross-vendor arguments both got quietly erased in late 2025/early 2026. MCP earns its keep in exactly one situation: you&rsquo;re a <em>provider</em> exposing a live, OAuth&rsquo;d service to assistants you don&rsquo;t own. The rule that falls out: <strong>consuming an API → skill + CLI. Providing a service → MCP.</strong></p>
<h2 id="the-category-error-i-was-about-to-commit">The category error I was about to commit</h2>
<p>The first crack: &ldquo;API + skills vs MCP&rdquo; quietly assumes the two live on the same shelf. They don&rsquo;t.</p>
<p>A <strong>skill</strong> is knowledge. A markdown recipe — plus maybe a script — that teaches the agent how to do something, running in <em>your</em> runtime, on <em>your</em> machine, with tools you already have.</p>
<p><strong>MCP</strong> is a connection to a running service. A server, behind a protocol, that the model talks to.</p>
<p>Comparing them is apples to orchards. One is &ldquo;here&rsquo;s how, go do it.&rdquo; The other is &ldquo;here&rsquo;s a thing that&rsquo;s already running, call it.&rdquo; Most of the internet argues about them as if they&rsquo;re competing products. They&rsquo;re not even the same noun.</p>
<p>So the honest question isn&rsquo;t &ldquo;which wins.&rdquo; It&rsquo;s &ldquo;when does a hosted service behind a contract beat a recipe you run yourself?&rdquo; That&rsquo;s a real question. I just assumed I knew the answer.</p>
<h2 id="the-two-arguments-that-died-before-i-finished-typing">The two arguments that died before I finished typing</h2>
<p>My case against MCP rested on two pillars. Both had already collapsed, and I hadn&rsquo;t noticed.</p>
<p><strong>Pillar one: context bloat.</strong> Every MCP server dumps all its tool schemas into the context window — a seven-server setup could eat 67K tokens before you typed a word, and <a href="https://lakshminp.com/2025/11/ai-agent-memory-persistence/" class="lnp-link">the context window is the one resource your agent can&rsquo;t buy back</a>. Damning. Except in January 2026 Anthropic shipped tool search and <code>defer_loading</code>: now the model sees a search tool plus a couple of always-on tools, and pulls the rest on demand. Reported reductions of 85–95%. My killer stat became a &ldquo;this used to be true.&rdquo;</p>
<p><strong>Pillar two: cross-vendor reach is MCP&rsquo;s moat.</strong> Wrong by a different calendar. In December 2025, Agent Skills shipped as an open standard, and within 48 hours Microsoft put it in VS Code and OpenAI added it to ChatGPT and Codex. By spring, ~40 tools — Gemini CLI, JetBrains, Kiro, Goose — read the same <code>SKILL.md</code>. Skills are as universal as MCP now. The moat drained while I was sharpening my knives.</p>
<p>Fine. Two pillars down. The dunk still had three legs, I figured.</p>
<h2 id="watching-the-rest-fall">Watching the rest fall</h2>
<p><strong>Security?</strong> I&rsquo;d claimed MCP gives you a safety edge. It doesn&rsquo;t. If I want read-only GitHub access, I hand a read-only token to the skill <em>or</em> the MCP server — identical. The token scope is the gate, enforced at the API boundary, available to both. There&rsquo;s no protocol-level security advantage. Gone.</p>
<p><strong>Tokens?</strong> This one inverts, which delighted me until I realized it cut against my own thesis too. People assume MCP is token-cheap because the call — <code>list_pull_requests(owner, repo)</code> — is tidy. But the <em>call</em> isn&rsquo;t the cost. The <em>result</em> is. The GitHub API returns fat JSON, and a raw MCP tool call dumps the whole blob into context. A skill that runs code can filter in the sandbox and return five lines. So code-that-filters wins on tokens — and a skill is code-that-filters by birth. But that&rsquo;s an argument for skills, not against MCP-the-idea.</p>
<p><strong>Auto-orchestration?</strong> &ldquo;MCP composes calls for you.&rdquo; No, it doesn&rsquo;t. The protocol is transport — it has no &ldquo;run this sequence, give me only the end&rdquo; primitive. Either the model loops (every intermediate result round-trips through context — expensive) or a human pre-bakes a coarse server tool (effort). Automatic, token-cheap stacking only happens when you call tools <em>from code</em> — which is, once again, the skill model.</p>
<p>Every road kept leading back to the same place. I started to feel like the universe was trying to tell me something(sounds dramatic, I know).</p>
<h2 id="the-litmus-test-that-almost-saved-the-dunk">The litmus test that almost saved the dunk</h2>
<p>So I built a concrete test: <em>&ldquo;Fetch all open PRs, give me a gist of the modules they touch, merge only the ones tagged auth.&rdquo;</em></p>
<p>Fetch and gist are reads — data-heavy aggregation, the code-that-filters sweet spot. Skill wins, easily. But <em>merge</em> is a write, and a dangerous one, and writes are where I figured MCP&rsquo;s permission policy — pause and confirm each merge — would finally earn its keep.</p>
<p>Then a reader on the thread that became this post pointed out the obvious: gating decomposes into three questions, and only one is even arguably MCP&rsquo;s.</p>
<ul>
<li><strong>Capability</strong> — can a merge happen at all? The <em>token scope</em> answers that. Available to both. API-enforced.</li>
<li><strong>Selection</strong> — which PRs get merged? Your <em>filtering code</em> answers that. That&rsquo;s the skill&rsquo;s script.</li>
<li><strong>Confirmation</strong> — do you approve each one? <em>You</em>, in the loop on a coding agent — or a <code>--confirm</code> flag — or MCP&rsquo;s native prompt.</li>
</ul>
<p>Only the third row is MCP&rsquo;s, and even there it&rsquo;s matched by you-watching-bash or a confirm flag. On a coding agent with you present, skill plus the <code>gh</code> CLI wins the whole task. The merge didn&rsquo;t flip it. <em>You&rsquo;re</em> the permission policy.</p>
<p>That was the moment the dunk officially died. I went looking on Reddit to see who else had buried it.</p>
<h2 id="what-reddit-already-knew">What Reddit already knew</h2>
<p>Turns out, everyone. The threads are a graveyard with two opposing headstones.</p>
<p><a href="https://www.reddit.com/r/ClaudeCode/comments/1rrl56g/" rel="external nofollow noopener" class="lnp-link">&ldquo;Will MCP be dead soon?&rdquo;</a> — 406 comments. <a href="https://www.reddit.com/r/ClaudeAI/comments/1pjpbji/" rel="external nofollow noopener" class="lnp-link">&ldquo;I cannot, for the life of me, understand the value of MCPs&rdquo;</a> — 305 comments. <a href="https://www.reddit.com/r/mcp/comments/1rstpfk/" rel="external nofollow noopener" class="lnp-link">&ldquo;A eulogy for MCP (RIP).&rdquo;</a> <a href="https://www.reddit.com/r/mcp/comments/1o8w5wq/" rel="external nofollow noopener" class="lnp-link">&ldquo;CLI &gt; MCP?&rdquo;</a> Someone even shipped a tool that converts MCP servers into CLI + skill files and &ldquo;cut ~97% token overhead.&rdquo;</p>
<p>The auto-generated TL;DR of that 305-comment thread is, embarrassingly, the post I&rsquo;d spent a day reverse-engineering: <em>&ldquo;You&rsquo;re looking at this from a solo dev&rsquo;s perspective, and you&rsquo;re not wrong — Skills or telling Claude to use a CLI is often more efficient. MCP&rsquo;s real value isn&rsquo;t for your individual coding session, but for the broader ecosystem.&rdquo;</em></p>
<p>So the solo-dev case is settled. But the people defending MCP weren&rsquo;t demo-app tourists. They were running it in production, and they landed one punch I couldn&rsquo;t slip.</p>
<h2 id="the-punch-i-couldnt-slip-and-the-one-real-win">The punch I couldn&rsquo;t slip, and the one real win</h2>
<p>From the eulogy thread, top comment: <em>&ldquo;People who claim there&rsquo;s no need for MCP will, if they build projects of growing complexity, sooner or later reinvent everything MCP provides — but bespoke and non-standardized.&rdquo;</em> And a sharper one: <em>&ldquo;CLI + skills are great for solo dev vibes. But the second you need an LLM to orchestrate across multiple platforms with real auth and governance? You&rsquo;re either using MCP or rebuilding it badly.&rdquo;</em></p>
<p>That&rsquo;s the one thing that survived every round. Not context, not security, not reach, not tokens. <strong>Distribution — of a specific kind.</strong></p>
<p>Here&rsquo;s the case, and it&rsquo;s narrower than the hype and realer than my dunk: you&rsquo;re a <em>provider</em>. You host a live service and you want it to show up as a one-click, OAuth&rsquo;d connector inside every AI assistant your customers already use — Claude&rsquo;s Connectors Directory (200+ integrations), ChatGPT Apps, all of it. Notion, Linear, and Stripe ship official remote MCP servers for exactly this. You build once; it lights up everywhere; the credentials and compute stay on your side.</p>
<p>A skill — even a universal one — cannot be that. A skill is a copy that runs in the consumer&rsquo;s runtime, and that&rsquo;s the whole limitation: it can only do what that runtime can do. Flip that around and you get the same win seen from the client&rsquo;s side. Claude Desktop runs skills <em>and</em> has a code sandbox — but the sandbox can&rsquo;t reach the open internet, so a skill that tries to curl GitHub dies at the egress wall, while an MCP server, running outside the box, reaches it fine. Same task, opposite answer, deciding variable is network egress. It&rsquo;s not a second reason to use MCP. It&rsquo;s the first reason wearing a different hat: when the consumer&rsquo;s runtime can&rsquo;t reach the thing, you need a service that can.</p>
<h2 id="the-rule-when-to-use-mcp-vs-a-skill--cli">The rule: when to use MCP vs a skill + CLI</h2>
<p>So, do you need MCP? Ask one question: <strong>are you consuming an API, or providing a service?</strong></p>
<p>Wiring your own agent to someone&rsquo;s existing API, on a coding tool with a shell — skill plus a CLI, every time. You will not &ldquo;reinvent MCP badly,&rdquo; because you need none of what MCP provides: no OAuth dance, no dynamic discovery, no cross-client reach. The CLI is complete, not a degenerate clone.</p>
<p>Exposing your own live service to assistants you don&rsquo;t own, with auth and governance, across vendors — that&rsquo;s MCP, and a CLI genuinely can&rsquo;t do it.</p>
<p>MCP isn&rsquo;t the poor man&rsquo;s anything. It&rsquo;s the <em>platform&rsquo;s</em> protocol. You reach for it the moment you stop consuming APIs and start being an app inside other people&rsquo;s assistants.</p>
<p>I know which side I&rsquo;m on this week. I shipped ThreadHQ&rsquo;s MCP server as a top-of-funnel for a reason — that&rsquo;s the provider play, done on purpose. (ThreadHQ is one of the products I <a href="https://lakshminp.com/2026/06/build-saas-with-claude-code/" class="lnp-link">built solo with Claude Code</a>; the MCP server is its distribution edge, not its plumbing.) But for the GitHub task on my own machine? I&rsquo;m still just typing <code>gh</code>. The dunk was wrong. The honest version is sharper anyway.</p>
]]></content:encoded></item><item><title>How to Build a SaaS with Claude Code in a Weekend (Not a Quarter)</title><link>https://lakshminp.com/2026/06/build-saas-with-claude-code/</link><pubDate>Mon, 15 Jun 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/06/build-saas-with-claude-code/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><category>saas</category><description>I caught up with one of my mentees last weekend. Just talking shop. He’s a solid developer, he’s got the itch to build a side project, and he’s brand new to agentic coding. Somewhere in that conversation I realized I was reciting an entire playbook off the top of my head — so I’m writing it down. If you can write code and you’re staring at Claude Code wondering how to actually build and ship a SaaS with it, this is for you.</description><content:encoded><![CDATA[<p>I caught up with one of my mentees last weekend. Just talking shop. He&rsquo;s a solid developer, he&rsquo;s got the itch to build a side project, and he&rsquo;s brand new to agentic coding. Somewhere in that conversation I realized I was reciting an entire playbook off the top of my head — so I&rsquo;m writing it down. If you can write code and you&rsquo;re staring at Claude Code wondering how to actually build and ship a SaaS with it, this is for you.</p>
<p>A quick word on <em>why now</em>, before the <em>how</em>. Two reasons, and you&rsquo;ve heard at least one:</p>
<ol>
<li>AI is eating the jobs. You can&rsquo;t open your phone without someone reminding you. Enough said.</li>
<li>The AI subsidy is going to end. Right now you&rsquo;re building on frontier models that cost the labs more than they charge you. That window closes. <a href="/2026/03/ai-subsidy-window-developers/" class="lnp-link">I wrote about this here</a></li>
</ol>
<p>Maybe both happen. Either way the move is the same: build the muscle now, while it&rsquo;s cheap and you still have an edge.</p>
<h2 id="1-what-to-build">1. What to build</h2>
<p>Scratch your own itch. Do not spend three weeks &ldquo;researching the market&rdquo; to discover what people want. If you have a problem an app would fix, give yourself permission to build it. We&rsquo;ll worry about market size <em>after</em> it ships.</p>
<p>Market research used to be load-bearing because building was expensive and being wrong was catastrophic — months of work, real money, dead on arrival. That math is gone. You can ship an app over a weekend now. The cost of being wrong is one wasted Saturday.</p>
<p>Paul Graham makes the sharper version of this point: build what <em>you</em> want, because — as he writes in <a href="https://paulgraham.com/earn.html" rel="external nofollow noopener" class="lnp-link">How to Earn a Billion Dollars</a> — &ldquo;your own needs are uniquely valuable, because your needs predict future demand.&rdquo; You&rsquo;re not guessing what strangers want. You&rsquo;re scratching an itch you can actually feel.</p>
<p>I&rsquo;ve written about finding an idea and shipping it in a single sitting — <a href="/2026/01/claude-code-ship-one-session/" class="lnp-link">here</a>.</p>
<h2 id="2-know-where-youre-standing">2. Know where you&rsquo;re standing</h2>
<p>Be honest about your starting point: do you have real development experience, or do you need to ramp up first? <a href="/2025/10/ai-coding-prerequisites/" class="lnp-link">Here&rsquo;s what I&rsquo;d master before letting AI build for you</a></p>
<p>Yes, there are people on the internet shipping apps with zero coding background. Maybe. But if you can&rsquo;t read what Claude Code writes, you can&rsquo;t steer it — and you&rsquo;ll feel that the first time it confidently drives into a wall. You don&rsquo;t need a CS degree. You need enough of a mental model to call BS. The good news: you can learn development <em>from</em> Claude Code while you build <em>with</em> it. Learning and building at the same time is completely legitimate — it&rsquo;s how I pick up half the things I use.</p>
<h2 id="3-how-i-actually-build">3. How I actually build</h2>
<p>There&rsquo;s no one true way. This is what works for me; your mileage may vary.</p>
<p>First, the unglamorous part: get the Claude Code Max plan. The $100 tier, or the $200 one if you can swing it. You cannot build anything meaningful on the cheap plans, and you&rsquo;ll discover exactly why about four hours into your first real session. Don&rsquo;t flinch at the price — it&rsquo;s still cheaper than a tutor or a freelancer, and it doesn&rsquo;t take lunch breaks. Plus the quality is consistently good.</p>
<h3 id="why-claude-code-and-not-the-benchmark-topping-model-of-the-week">Why Claude Code and not [the benchmark-topping model of the week]?</h3>
<p>Because building a startup is already exhausting, and you do not have spare energy to spend benchmarking models and sharpening tools instead of shipping. Pick what works and stick with it.</p>
<p>Codex is a close second — genuinely good, and I keep it around. But Claude Code stays a step ahead for actually building apps, and after working with both, I reach for it first. Use both if you like. Just don&rsquo;t turn tool selection into the project.</p>
<h3 id="boring-choices-win">Boring choices win</h3>
<p>Freeze your stack early — backend, frontend, database. 80% of it is identical across every app you&rsquo;ll ever build, so stop re-deciding it every time. And don&rsquo;t obsess over scale and optimization. Those are problems you <em>earn</em> by being successful. Good problems. You don&rsquo;t have them yet.</p>
<h2 id="4-specs-and-context-engineering-the-part-that-decides-everything">4. Specs and context engineering: the part that decides everything</h2>
<p>Two things determine whether you get your money&rsquo;s worth out of Claude Code:</p>
<ol>
<li>The specification</li>
<li>Context engineering</li>
</ol>
<h3 id="the-spec">The spec</h3>
<p>The more specific you are, the better the output. &ldquo;Build a to-do app&rdquo; gets you slop. &ldquo;Here&rsquo;s the auth flow, here&rsquo;s the data model, here are the exact features&rdquo; gets you something you can use. Spend real time here, before a single line of code gets written. <a href="/2025/12/stop-making-claude-code-guess/" class="lnp-link">More on this</a></p>
<p>The move that works best: make Claude Code interview <em>you</em> about the spec. If you can&rsquo;t answer its questions, you can&rsquo;t articulate the feature — and if you can&rsquo;t articulate it, you can&rsquo;t build it. By the time the spec is done, the MVP should have zero grey areas.</p>
<h3 id="context-engineering">Context engineering</h3>
<p>Even the best frontier model starts coding like it&rsquo;s three drinks deep once the context fills up. So you manage it.</p>
<p>This takes me back to my assembly-language days — limited registers, limited memory, every instruction written with one eye on the resources you didn&rsquo;t have. LLM context is that same constraint in new clothes. You get roughly 200k tokens, and that&rsquo;s nowhere near enough to hold your whole app in its head at once.</p>
<p>So: one task per session. Two at the absolute most. Which means breaking the spec into session-sized tasks and tracking what&rsquo;s done, what&rsquo;s in flight, what&rsquo;s blocked, and what depends on what.</p>
<p>A markdown to-do file is a terrible way to do this and a worse use of your time. I use <strong>beads</strong>. Adopted it early, still on it. It&rsquo;s the fix for <a href="/2025/11/ai-agent-memory-persistence/" class="lnp-link">an agent that wakes up every morning with no memory of what you did yesterday</a>.</p>
<p>This practice is also a lot kinder for your token limits.</p>
<p>Two flavors:</p>
<ul>
<li>the original, by the author — <a href="https://github.com/gastownhall/beads" rel="external nofollow noopener" class="lnp-link">gastownhall/beads</a></li>
<li><strong>beads-rust</strong>, a simpler, more stable reimplementation of the spec in Rust — <a href="https://github.com/Dicklesworthstone/beads_rust" rel="external nofollow noopener" class="lnp-link">Dicklesworthstone/beads_rust</a></li>
</ul>
<p>Use either. I landed on beads-rust. Pick your poison.</p>
<p>You also need memory <em>across</em> sessions. You&rsquo;ll remember you fixed a bug two weeks ago; the model in today&rsquo;s session won&rsquo;t, and it&rsquo;ll cheerfully hand you a wrong answer when you ask &ldquo;did we already fix this?&rdquo; I use <strong>claude-mem</strong> for that — <a href="https://github.com/thedotmack/claude-mem" rel="external nofollow noopener" class="lnp-link">thedotmack/claude-mem</a>.</p>
<p>Point a Claude Code session at both repos and it&rsquo;ll install them for you. (Yes, these work with Codex too — but again: shipping or tuning? The clock is running.)</p>
<p>Remember that spec? Hand it to your beads skill and it breaks down into a clean task list. Ask Claude to pull the high-leverage beads and start there — never more than two in flight.</p>
<h3 id="your-claudemd-matters">Your CLAUDE.md matters</h3>
<p>Treat it as a compass, not a second spec. The practices that matter (write the tests first), how the app deploys, the handful of goals you&rsquo;re aiming at. Keep it light — <a href="/2026/04/claude-md-best-practices/" class="lnp-link">cramming a novel into CLAUDE.md is the fastest way to make Claude dumber</a>.</p>
<h3 id="one-more-tool">One more tool</h3>
<p><a href="https://github.com/sirmalloc/ccstatusline" rel="external nofollow noopener" class="lnp-link">ccstatusline</a>. Configure it to show your remaining context percentage. It&rsquo;s the fuel gauge that tells you whether to keep driving or pull over and start a fresh session.</p>
<p>Then it&rsquo;s rinse and repeat: feed beads to Claude, do the manual QA yourself, close the bead, next one. A few sessions in, a sliver of a working MVP starts to emerge. When you&rsquo;re happy with it, you ship.</p>
<h2 id="5-deploying-it-without-the-kubernetes-tax">5. Deploying it (without the Kubernetes tax)</h2>
<p>I was a Kubernetes guy for years. Deployed everything on it. I no longer recommend it for solo developers — <a href="/2025/12/kubernetes-indie-dev-alternative/" class="lnp-link">I explain why here</a>. It still has its place and time; your weekend project is neither.</p>
<p>Use <strong><a href="https://kamal-deploy.org" rel="external nofollow noopener" class="lnp-link">Kamal</a></strong> instead. Think of it as the compromise between Docker Compose and Kubernetes — Compose&rsquo;s simplicity, enough of Kubernetes&rsquo; robustness, none of the YAML despair.</p>
<p>I&rsquo;m also building <a href="https://vmkit.dev/" rel="external nofollow noopener" class="lnp-link">VMKit</a> to make this part disappear entirely — deploy without learning the nitty-gritty unless you want to.</p>
<p>Then there&rsquo;s the wiring you don&rsquo;t think about until it bites: monitoring and ops (I run mine through MCPs — <a href="/2026/01/30-dollar-saas-stack/" class="lnp-link">the $30 stack</a>), payments (Stripe; if you&rsquo;re in India, Dodo Payments), and distribution — marketing and positioning, which is a beast all its own. Each of these deserves its own post.</p>
<h2 id="tldr">TL;DR</h2>
<ul>
<li>Build for your own itch. Shipping is cheaper than market research now.</li>
<li>Get the Claude Code Max plan ($100, ideally $200). Don&rsquo;t tool-shop.</li>
<li>Spend your time on the spec — let Claude interview you until there are no grey areas.</li>
<li>Engineer your context: one task per session, tracked with beads, remembered with claude-mem.</li>
<li>Keep CLAUDE.md light. Watch your context gauge.</li>
<li>Deploy with Kamal, not Kubernetes. Wire payments and monitoring last.</li>
<li>The whole thing is a weekend, not a quarter.</li>
</ul>
]]></content:encoded></item><item><title>Claude Overreaches. Codex Underreaches. I'm Still Figuring Out How to Use Both.</title><link>https://lakshminp.com/2026/04/claude-vs-codex-use-both/</link><pubDate>Wed, 22 Apr 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/04/claude-vs-codex-use-both/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><description>I was a one-agent guy until Claude had a run of outages.
On those days I didn’t ship less. I shipped nothing. I’d open my editor, remember Claude was down, stare at the codebase, close the editor. A single-vendor dependency masquerading as a workflow.
So I reluctantly installed Codex CLI. Poked at it. Resented it for a week. Then task by task — caught myself reaching for it on purpose, even when Claude was up.</description><content:encoded><![CDATA[<p>I was a one-agent guy until Claude had a run of outages.</p>
<p>On those days I didn’t ship less. I shipped <em>nothing</em>. I’d open my editor, remember Claude was down, stare at the codebase, close the editor. A single-vendor dependency masquerading as a workflow.</p>
<p>So I reluctantly installed Codex CLI. Poked at it. Resented it for a week. Then task by task — caught myself reaching for it on purpose, even when Claude was up.</p>
<p>I still don’t have the workflow figured out. What I do know is that “pick one” is the wrong frame, and the Reddit threads that get it right aren’t the ones with the most upvotes.</p>
<h1 id="the-one-sentence-that-explains-everything"><strong>The One Sentence That Explains Everything</strong></h1>
<p>From a 520-upvote r/ClaudeCode thread analyzing both tools’ open-source prompts:</p>
<blockquote>
<p><em>“Claude Code reads like a product trying to create initiative while Codex reads like a product trying to prevent drift.”</em><br>
— u/idkwhattochoosz</p>
</blockquote>
<p>And the pithier version, from the comments:</p>
<blockquote>
<p><em>“Claude is more willing to sin by overreaching. Codex is more willing to sin by underreaching.”</em><br>
— u/entheogenicentity</p>
</blockquote>
<p>Read those twice. That’s not a model-quality take. That’s a product-philosophy take. Two teams looked at the same question — what should an agent do when it doesn’t know what you meant? — and picked opposite defaults. One said “guess and move.” The other said “ask and wait.”</p>
<p>Claude Code’s system prompt pushes hard toward initiative: <em>“A good colleague faced with ambiguity doesn’t just stop — they investigate, reduce risk, and build understanding.”</em> Codex’s harness does the opposite: narrow the ambiguity, verify, don’t guess.</p>
<p>Every “Claude vs Codex” benchmark you’ve seen is scoring two products that were never competing on the same axis. It’s like benchmarking a kayak against a sedan because they both move you forward.</p>
<h1 id="my-honest-opinion-codexs-harness-is-better"><strong>My Honest Opinion: Codex’s Harness Is Better</strong></h1>
<p>This is going to get me yelled at in r/ClaudeCode, and that’s fine.</p>
<p>After several weeks running both, Codex’s harness feels more mature. Not the model — the harness. The scaffolding around the model. The way it handles ambiguity, scope, and completeness.</p>
<p>Three things Codex does that Claude Code still doesn’t:</p>
<p><strong>1. It doesn’t lie about completion.</strong> Claude will hand you a summary saying the work is done, tests pass, shipping-ready. Codex more often flags what it didn’t fix, what it wasn’t sure about, what it skipped. One r/ClaudeCode commenter put it better than I can: <em>“Claude will always claim all is done and ready, while Codex will flag it and say ‘no, there is this and this and this that still need to be fixed.’”</em></p>
<p><strong>2. It respects your instructions.</strong> Claude treats <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> as a helpful suggestion. Codex treats <a href="http://agents.md/" rel="external nofollow noopener" class="lnp-link">AGENTS.md</a> as a contract. If you tell Codex “don’t touch the migration files,” it doesn’t touch them. If you tell Claude the same thing, you’ll find a migration file edit in the diff and a cheerful note about how it improved schema consistency.</p>
<p><strong>3. The restraint scales better.</strong> Claude’s “volunteer more” bias is delightful at 30 minutes of work. It becomes a liability at 3 hours. Codex’s restraint is annoying in a small task and load-bearing in a long one.</p>
<p>None of this means Claude Code is bad. It means Claude Code is optimized for a different shape of work than I’m doing. The initiative bias is a great fit for exploration and greenfield work. For production changes to a real codebase, Codex’s paranoia is the right default.</p>
<p>Here’s the one that changed my mind. I built Supabyoi (managed self-hosted Supabase) with Claude Code. When the MVP felt feature-complete — Claude’s verdict, confidently delivered, complete with a tasteful little summary of everything that worked — I ran a second pass on Codex in a parallel directory (<code>~/supabyoi-codex</code>). Just to see.</p>
<p>Codex came back with a whole second project’s worth of findings. Not the usual “bugs Claude missed.” Bugs Claude had <em>confidently signed off on.</em> Shipping-ready, per Claude. Not shipping-ready, per Codex. Codex was right about every one of them.</p>
<p>That was the week I stopped treating Codex as the thing I installed during an outage and started treating it as a different kind of reviewer. Not better. Differently biased. A second pair of eyes is only useful if it’s not the same pair of eyes.</p>
<h1 id="why-you-should-actually-run-both"><strong>Why You Should Actually Run Both</strong></h1>
<p>The flip side — and this matters, because I don’t want this post read as “switch to Codex, you fool” — Claude’s initiative bias is a real asset. You just have to point it at the right phase of the work. The problem isn’t Claude. It’s that you’re using Claude for the part of the job Codex is better at, and vice versa.</p>
<p>Four reasons to dual-sub instead of picking:</p>
<p><strong>1. Hallucination diversity.</strong> This is the biggest one and almost nobody articulates it clearly. From u/campbellm on Reddit:</p>
<blockquote>
<p><em>“I’ve been doing ‘have claude write something, have codex review it, have claude consider and critique that review.’ It is VERY unlikely that both will hallucinate the same way.”</em></p>
</blockquote>
<p>Two models trained on different data with different RLHF signals don’t fail identically. When Claude writes confident-but-wrong code, Codex flags it. When Codex skips a subtle edge case, Claude’s “check adjacent concerns” bias picks it up. You get a natural adversarial review without hiring anyone.</p>
<p><strong>2. The planner-executor split.</strong> Use Claude for the part it’s good at — exploring a messy problem space, drafting a plan, proposing a dozen angles. Then hand the plan to Codex for implementation. u/ocombe on r/ClaudeCode: <em>“Run claude for the plan &amp; fast work, use codex for thorough plan &amp; code reviews.”</em> u/mrothro’s version: <em>“I use Claude Code for ideating and small implementation, then tell it to run Codex to do complex implementations and code reviews.”</em></p>
<p>The pattern is consistent across the threads: Claude’s strength is at the start (wide search, first drafts); Codex’s strength is at the end (narrow, verify, harden).</p>
<p><strong>3. Cross-harness rule enforcement.</strong> Rules one model ignores, the other enforces. If Claude drifts on a constraint you set, Codex catches it in review. If Codex is too literal and missed an obvious improvement, Claude’s adjacent-concerns bias surfaces it. Two different failure modes cancel each other out.</p>
<p><strong>4. Throughput.</strong> Both platforms throttle hard at the Max/Pro tier. When Claude hits limits on Friday morning, you switch to Codex and keep shipping. One r/ClaudeCode commenter reported pulling down from a Claude 20x plan to 5x, then adding a $100/mo Codex plan — roughly the same total cost, dramatically more runway. I’m not sure that math works for everyone, but the principle holds: one subscription is a single point of failure.</p>
<h1 id="agent-flywheel-is-the-tooling-signal"><strong>Agent-Flywheel Is the Tooling Signal</strong></h1>
<p>There’s a product called <a href="https://agent-flywheel.com/" rel="external nofollow noopener" class="lnp-link">agent-flywheel.com</a> that pre-configures Claude Code, Codex CLI, and Gemini on a fresh VPS. Total damage — VPS plus both Max/Pro subs — lands between 440and440<em>and</em>656 a month. That’s a car payment for a car that writes your code.</p>
<p>What I find interesting isn’t the tool. It’s the bet underneath it: a whole product assumes real developers want all three installed by default. Six months ago that would have read as overkill. Today it reads as table stakes.</p>
<p>The hype cycle hasn’t caught up yet. The mainstream take is still “pick your favorite,” as though these were ice cream flavors. The people actually shipping production code with agents have quietly moved to “run both. Sometimes three. And don’t make a big deal about it.”</p>
<p>I’m planning to deploy it — not on a greenfield project (everybody has a greenfield story), but on an existing one already shipping to real users. The interesting question isn’t whether a three-agent stack works on a clean slate. It’s what breaks when you wire it into a codebase with real uptime constraints, customers, and six months of decisions the tooling didn’t witness. Real-world battle stories from agent-flywheel setups are scarce. I want to write one.</p>
<h1 id="the-honest-part-i-dont-have-the-workflow-figured-out-yet"><strong>The Honest Part: I Don’t Have the Workflow Figured Out Yet</strong></h1>
<p>Everything above reads like I’ve got this nailed. I don’t. Here’s the list of things I still don’t know, offered in the spirit of not pretending:</p>
<p><strong>When exactly to hand off.</strong> I know Claude should plan and Codex should review. I don’t have a clean trigger. Sometimes I bounce mid-implementation because Claude is about to go off the rails. Sometimes I trust Claude to finish and Codex only sees the final diff. The “right” cadence isn’t obvious.</p>
<p><strong>How much context to share.</strong> Each agent wants the full <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> / <a href="http://agents.md/" rel="external nofollow noopener" class="lnp-link">AGENTS.md</a> treatment. Writing both, keeping them in sync, and remembering which one has which convention is its own small job. I haven’t found a clean answer.</p>
<p><strong>Whether the adversarial review actually catches bugs.</strong> It sounds great in theory. In practice, most of the time both agents agree the work is done, and the bugs I catch in review are ones I would have caught with one agent too. The hallucination-diversity argument may be overstated at the tasks most of us are actually doing.</p>
<p><strong>Whether the cost is worth it at my usage.</strong> I’m not running agents 40 hours a week. At $400+/month for the dual sub, I’m probably over-subscribed for my actual throughput. The math gets better if you’re coding all day. I’m not.</p>
<h1 id="who-should-dual-sub-who-shouldnt"><strong>Who Should Dual-Sub, Who Shouldn’t</strong></h1>
<p><strong>Do it</strong> if you’re a solo dev shipping production code daily. You’ll hit Friday-morning limits on one platform whether you budget for it or not, and the adversarial review actually catches things. The cost is real. The throughput gain is bigger. Do the math; it pencils.</p>
<p><strong>Don’t bother</strong> if you code a few hours a week. The switching tax and the subscription burn aren’t worth it at low volume. Pick one and move on. Claude if you want initiative. Codex if you want restraint. Nobody is grading you on this.</p>
<p><strong>It’s complicated</strong> if you’re at a day job where the company pays for one and you’ve got a side project. Use the company sub for the day job. Don’t stack a second personal sub unless the side project is actually shipping — not “actually going to ship next month,” <em>actually shipping, this week, to real users.</em> The number of people running dual subs to ship nothing is, I suspect, not small.</p>
<h1 id="what-this-is-really-about"><strong>What This Is Really About</strong></h1>
<p>The “ditch ChatGPT for Claude” narrative was a 2025 story. It was right for its moment. But the 2026 version of that story isn’t “ditch Claude for Codex.” It’s “stop treating this as a winner-take-all market.”</p>
<p>Different models have different biases baked into their harnesses. Claude overreaches. Codex underreaches. Gemini is still figuring out its personality. The right move isn’t to pick the bias you like. It’s to stack biases against each other so their failure modes cancel out.</p>
<p>I don’t have this workflow figured out. Neither does anyone else I’ve read on Reddit, honestly — the high-upvote posts are mostly single-tool takes, and the real insight is buried in the comments of threads with a few hundred upvotes.</p>
<p>But “only use one” is already wrong. That much is clear.</p>
]]></content:encoded></item><item><title>Your CLAUDE.md Is Making Claude Dumber</title><link>https://lakshminp.com/2026/04/claude-md-best-practices/</link><pubDate>Mon, 06 Apr 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/04/claude-md-best-practices/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><description>Your CLAUDE.md is 800 lines long. You spent a weekend organizing it into 27 modular files with a routing system. You wrote a blog post about it. You got upvotes.
Claude is ignoring most of it.
There’s an arms race happening in the Claude Code community right now. Every week, someone posts their increasingly elaborate CLAUDE.md setup. 27-file architectures. Tiered loading systems. Router patterns with conditional context injection.
One developer split their CLAUDE.md into 27 files with a three-tier routing system. 360 upvotes. The post opens with: “My CLAUDE.md was ~800 lines. It worked until it didn’t. Rules for one context bled into another, edits had unpredictable side effects, and the model quietly ignored constraints buried 600 lines deep.”</description><content:encoded><![CDATA[<p>Your CLAUDE.md is 800 lines long. You spent a weekend organizing it into 27 modular files with a routing system. You wrote a blog post about it. You got upvotes.</p>
<p>Claude is ignoring most of it.</p>
<p>There’s an arms race happening in the Claude Code community right now. Every week, someone posts their increasingly elaborate CLAUDE.md setup. 27-file architectures. Tiered loading systems. Router patterns with conditional context injection.</p>
<p>One developer <a href="https://reddit.com/r/ClaudeCode/comments/1rhe89z/" rel="external nofollow noopener" class="lnp-link">split their CLAUDE.md into 27 files</a> with a three-tier routing system. 360 upvotes. The post opens with: “My CLAUDE.md was ~800 lines. It worked until it didn’t. Rules for one context bled into another, edits had unpredictable side effects, and the model quietly ignored constraints buried 600 lines deep.”</p>
<p>The top comment, with 81 upvotes? “So not sure if you realised you can have descendant CLAUDE.md so you don’t even need to do this.”</p>
<p>Meanwhile, a developer in the same thread: “I don’t even use claude.md. Y’all are roleplaying being productive. Just work with it 1:1.”</p>
<p>One group is optimizing. The other is actually working.</p>
<h2 id="the-research-says-youre-doing-it-wrong">The Research Says You’re Doing It Wrong</h2>
<p>ETH Zurich researchers <a href="https://arxiv.org/pdf/2602.11988" rel="external nofollow noopener" class="lnp-link">published a paper</a> that should have made every CLAUDE.md maximalist uncomfortable. Their finding: context files — the .md files we all obsess over — tend to <em>reduce</em> task success rates compared to providing no repository context at all. And they increase inference cost by over 20%.</p>
<p>Read that again. No CLAUDE.md outperformed having one. On average.</p>
<p>When this paper hit Reddit, the poster titled it “<a href="https://reddit.com/r/ClaudeAI/comments/1rd93ho/" rel="external nofollow noopener" class="lnp-link">No CLAUDE.md → baseline. Bad CLAUDE.md → worse. Good CLAUDE.md → better.</a>” — an optimistic spin suggesting the file isn’t the problem, your writing is. The post got 209 upvotes. But the top comments immediately called it out: OP had misread the data. The actual finding was that having <em>any</em> .md file — human or LLM-written — led to worse performance than having none. The auto-generated thread summary confirmed it: “The consensus in this thread is that you’ve completely misread the paper.”</p>
<p>It gets worse. LLM-generated .md files hurt the most, because they just parrot back what’s already in the code. Human-written files showed a slight positive impact — but only when kept to an absolute minimum, and only for smaller models.</p>
<p>A separate benchmark of 1,188 runs across Haiku, Sonnet, and Opus confirmed this. Twelve coding tasks. Ten instruction profiles. The result: an empty CLAUDE.md scored best overall.</p>
<p>The researcher’s own correction was admirably blunt: “I was wrong about CLAUDE.md compression. Here’s what the data actually showed.”</p>
<h2 id="you-have-an-instruction-budget-youre-blowing-it">You Have an Instruction Budget. You’re Blowing It.</h2>
<p>Here’s the mechanism nobody talks about.</p>
<p>Frontier models reliably follow about 150 to 200 instructions before performance starts decaying. Not crashing — decaying. Every additional instruction slightly degrades compliance with every other instruction. The degradation is uniform. Your critical “NEVER delete the production database” rule gets weaker every time you add “prefer camelCase for variable names.”</p>
<p>Claude Code’s own system prompt already burns about 50 of those instruction slots. That’s before your CLAUDE.md even loads.</p>
<p>So you have roughly 100-150 instruction slots left. Your 800-line CLAUDE.md with coding conventions, style guides, architecture decisions, tool preferences, workflow rules, and team norms is trying to cram 400 instructions into 150 slots.</p>
<p>The model doesn’t crash. It just quietly starts ignoring things. Specifically, the things buried deepest in the file. Your most important rules — the ones you added after painful debugging sessions — are probably at the bottom. Which means they’re the first to get deprioritized.</p>
<h2 id="claude-is-designed-to-ignore-you">Claude Is Designed to Ignore You</h2>
<p>This is the part that should make you pause.</p>
<p>Claude Code’s system prompt includes this line about CLAUDE.md content:</p>
<blockquote>
<p>“This context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant.”</p>
</blockquote>
<p>Claude is literally instructed to deprioritize your instructions if they don’t seem relevant to the current task. The more task-specific content you stuff into CLAUDE.md, the more likely Claude treats the entire file as noise.</p>
<p>That database schema guidance? Irrelevant when Claude is working on frontend CSS. Those API naming conventions? Noise when it’s writing tests. Your elaborate deployment workflow? Invisible during a refactoring session.</p>
<p>Every irrelevant instruction trains Claude to ignore the relevant ones too.</p>
<p>The Context Window Tax</p>
<p>Here’s the math nobody does. Claude Code’s system prompt alone consumes roughly 23,000 tokens — about 11% of the 200K context window, gone before you type a word. Add your CLAUDE.md, your MCP tool schemas, skill descriptions, memory files, and rules. One developer <a href="https://reddit.com/r/ClaudeAI/comments/1s41rym/" rel="external nofollow noopener" class="lnp-link">measured 69,200 tokens of overhead</a> — 35% of the context window consumed before a single user message. Others in the thread pushed back on that specific number, but the principle stands: every always-loaded instruction competes with working memory.</p>
<p>And it’s not just a cost problem. It’s an accuracy problem. The fuller the context window gets, the worse Claude performs — what Anthropic calls context rot. Your elaborate CLAUDE.md isn’t just burning tokens. It’s actively degrading the quality of every response.</p>
<p>The Leverage Problem</p>
<p>Here’s why this matters more than you think.</p>
<p>Bad code is localized. You write a buggy function, it breaks one feature. You fix it, you move on.</p>
<p>Bad CLAUDE.md instructions compound. A single misguided rule in your CLAUDE.md affects every research phase, every plan, every implementation, every session. One line that says “always use verbose error messages with full stack traces” produces thousands of lines of noisy code across your entire codebase, across every agent, across every session.</p>
<p>Your CLAUDE.md is the highest-leverage file in your repo. Most people treat it like a junk drawer.</p>
<p>What the Minimalists Actually Do</p>
<p>I went looking for people who run Claude Code with minimal or no CLAUDE.md. They’re out there. They’re quiet about it because “I don’t use CLAUDE.md” doesn’t get upvotes.</p>
<p>One developer on Reddit: “I use Claude Code bare bones professionally. It all sounds like bloat not giving real value.” Another: “I load no skills, no agents, no MCP Servers and rock it all day every day, 12 hours a day. Life is good.”</p>
<p>A developer <a href="https://reddit.com/r/ClaudeAI/comments/1rmjg5r/" rel="external nofollow noopener" class="lnp-link">who built a 13-agent orchestration system</a> with 8,157 lines of markdown deleted 93% of it. His conclusion: “My enhancement layer was making Claude dumber by filling its brain with instructions about how to think, leaving less room for actual thinking.” After the deletion, Claude performed <em>better</em> on the same tasks.</p>
<p>Another developer with <a href="https://reddit.com/r/ClaudeAI/comments/1lvbe21/" rel="external nofollow noopener" class="lnp-link">a 350-line CLAUDE.md and 20+ custom MCP tools</a> put it simply: “It feels like the more context I add the more it struggles to get the job done. It seems to get ‘dumber’.”</p>
<p>And when someone <a href="https://reddit.com/r/ClaudeAI/comments/1lvi94t/" rel="external nofollow noopener" class="lnp-link">asked the community to break down the meta</a> on all the conflicting CLAUDE.md advice, the most honest reply got it right: “If ‘best practices’ are conflicting, it’s probably a sign of them mostly being a type of placebo on the part of the folks posting them. The human mind has a weird need to be the special one who cracked the code.”</p>
<p>The pattern is consistent: people who remove instructions report better results than people who add them.</p>
<p>Instructions Raise the Floor, Not the Ceiling</p>
<p>The benchmark data revealed something nuanced. Instructions don’t make Claude better on average. They make it more consistent.</p>
<p>On tasks where Claude already performs well, instructions add nothing. On tasks where Claude struggles, a focused workflow checklist gave Opus a +5.8 point lift and raised its worst-case score by 20+ points.</p>
<p>A <a href="https://reddit.com/r/ClaudeAI/comments/1pe37e3/" rel="external nofollow noopener" class="lnp-link">2,455-evaluation benchmark</a> across Sonnet and Opus confirmed a related finding: the best-performing configuration was a short CLAUDE.md with pointers to skills that load on demand — not a massive monolith, not 27 modular files, but a minimal routing layer that tells Claude where to find context when it’s actually needed.</p>
<p>This changes everything about how you should think about CLAUDE.md.</p>
<p>Don’t use it to make Claude smarter. Use it to prevent Claude from being stupid in specific, known ways. The difference between those two goals is the difference between a 60-line file and an 800-line file.</p>
<h2 id="what-actually-belongs-in-claudemd">What Actually Belongs in CLAUDE.md</h2>
<p>After digging through research, benchmarks, and hundreds of Reddit threads, here’s what survives the cut:</p>
<p><strong>The What-Why-How skeleton (under 60 lines):</strong></p>
<ul>
<li>
<p>WHAT: Your stack, project structure, key directories</p>
</li>
<li>
<p>WHY: What this project does and for whom</p>
</li>
<li>
<p>HOW: Build commands, test commands, deploy commands</p>
</li>
</ul>
<p><strong>Negatives over positives:</strong><br>
“NEVER use X” sticks. “Always prefer Y” fades. If you can phrase it as a prohibition, it enforces better. “DO NOT modify the database schema without migration files” beats “Always create migrations when changing the schema.”</p>
<p><strong>Trigger-action format:</strong><br>
“WHEN CI fails, DO NOT push until fixed” enforces consistently. “Always test before pushing” doesn’t. Specificity matters.</p>
<p><strong>Pointers, not content:</strong><br>
Reference external docs instead of embedding them. “See agent_docs/database.md for schema guidance” loads on demand. Pasting the full schema into CLAUDE.md loads every single session, whether Claude needs it or not.</p>
<p><strong>Subdirectory CLAUDE.md files:</strong><br>
Claude auto-loads CLAUDE.md from whatever directory it’s reading files in. Put backend rules in backend/CLAUDE.md. Put frontend rules in frontend/CLAUDE.md. Context-specific rules load only when contextually relevant.</p>
<h2 id="what-doesnt-belong">What Doesn’t Belong</h2>
<p><strong>Style guides.</strong> Claude is an in-context learner. If your code follows consistent patterns, Claude will match them without being told. Use linters and formatters — they’re deterministic, fast, and don’t eat instruction budget.</p>
<p><strong>LLM-generated instructions.</strong> The research is clear: auto-generated .md files hurt performance. Don’t use /init. Don’t ask Claude to write its own CLAUDE.md. The model just repeats what’s already in the code, wasting tokens to tell itself what it already knows.</p>
<p><strong>Lessons learned logs.</strong> Once the lesson is codified in the codebase itself — as a test, a lint rule, a hook — the .md entry is redundant. Delete it.</p>
<p><strong>Persona assignments.</strong> “You are a meticulous senior engineer who always&hellip;” is a costume, not a capability. As one developer <a href="https://reddit.com/r/ClaudeAI/comments/1rmjg5r/" rel="external nofollow noopener" class="lnp-link">running overnight cron agents</a> put it: “A syntax check that returns exit code 1 on failure &gt; 2,000 words of ‘you are a meticulous senior engineer who always&hellip;’” The agents with minimal instructions consistently outperformed the ones with elaborate persona prompts.</p>
<h2 id="the-real-best-practice">The Real Best Practice</h2>
<p>Keep your CLAUDE.md under 100 lines. Ideally under 60. Put the most important rules at the top. Phrase them as negatives. Use trigger-action format. Point to external docs instead of embedding content.</p>
<p>Then stop optimizing and go build something.</p>
<p>The developers shipping the most code aren’t the ones with the fanciest CLAUDE.md architectures. They’re the ones who figured out the minimum viable instructions and moved on to the actual work.</p>
<p>Your CLAUDE.md is not your product. Stop treating it like one.</p>
]]></content:encoded></item><item><title>The Claude Code Leak Revealed a Token Drain Bug. The Real Problem Is Bigger.</title><link>https://lakshminp.com/2026/04/claude-code-leaked-source-code-token-drain-ai-subsidy/</link><pubDate>Thu, 02 Apr 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/04/claude-code-leaked-source-code-token-drain-ai-subsidy/</guid><category>essays</category><category>claude-code</category><category>security</category><description>Follow-up to: Anthropic Is Losing Money on You Every Month. What Are You Shipping?
Three weeks ago, I wrote that Anthropic is losing money on every subscriber and that smart developers should ship like crazy before the economics normalize.
I was right about the thesis. I was wrong about the timeline.
The window isn’t closing in 18-24 months. It’s closing now.
What Changed in Three Weeks Three things happened in rapid succession that accelerated the timeline:</description><content:encoded><![CDATA[<p><em>Follow-up to: <a href="https://lakshminp.com/p/ai-subsidy-window-developers" class="lnp-link">Anthropic Is Losing Money on You Every Month. What Are You Shipping?</a></em></p>
<p>Three weeks ago, I wrote that Anthropic is losing money on every subscriber and that smart developers should ship like crazy before the economics normalize.</p>
<p>I was right about the thesis. I was wrong about the timeline.</p>
<p>The window isn’t closing in 18-24 months. It’s closing now.</p>
<h1 id="what-changed-in-three-weeks"><strong>What Changed in Three Weeks</strong></h1>
<p>Three things happened in rapid succession that accelerated the timeline:</p>
<p><strong>1. Claude subscriptions doubled.</strong> Anthropic’s paid user base went from ~30k to ~60k subscribers between January and March 2026. Record growth. The Claude Code launch, Super Bowl buzz, and Cowork tools drove a wave of new signups.</p>
<p><strong>2. Rate limits got brutal.</strong> Users on r/ClaudeAI went from “this is amazing” to “I can’t work” practically overnight. Pro users ($20/month) report hitting 10% of their daily quota from a single prompt. Max users ($100-200/month) report the same degradation. One Max 20x subscriber — paying $200/month — couldn’t work for nine consecutive days.</p>
<p><strong>3. The source code leaked.</strong> On March 31, 2026, a 59.8 MB source map file was accidentally shipped in the Claude Code npm package. 512,000 lines of TypeScript, mirrored across GitHub within hours. And buried in that code was proof of something users had been complaining about for weeks.</p>
<h1 id="the-token-drain-bug"><strong>The Token Drain Bug</strong></h1>
<p>Here’s what the leak revealed.</p>
<p>Claude Code has a function called <code>db8</code> that filters what gets saved to session files. For non-Anthropic users, it strips out all attachment-type messages — including <code>deferred_tools_delta</code> records that track which tools the model already knows about.</p>
<p>When you resume a session, Claude Code scans your history to figure out what tools it already announced. But because <code>db8</code> nuked those records, it finds nothing. So it re-announces every deferred tool from scratch. Every. Single. Resume.</p>
<p>This breaks prompt caching in three ways:</p>
<ul>
<li>
<p>System reminders shift positions in the message array</p>
</li>
<li>
<p>The billing hash changes because the first message content differs</p>
</li>
<li>
<p>The cache breakpoint moves because the array length is different</p>
</li>
</ul>
<p>Result: your entire conversation rebuilds as <code>cache_creation</code> tokens instead of hitting <code>cache_read</code>. The longer the conversation, the worse the drain.</p>
<p>One user patched the two-line fix and posted it. His 5-hour usage dropped from spiralling out of control to 6% — normal levels. The post got 367 upvotes. A sharp commenter noted the patch also bypasses billing controls on cache TTL, which makes it not just a bug fix, but let’s set that aside.</p>
<p>Here’s the uncomfortable part: this bug was burning tokens silently for weeks. Users were complaining about rate limits. Anthropic’s status page showed “no incidents.” And the actual cause was a caching bug in their own client code.</p>
<h1 id="the-math-doesnt-work"><strong>The Math Doesn’t Work</strong></h1>
<p>Let’s do the numbers.</p>
<p>Anthropic’s annualized revenue is roughly $14 billion. Claude Code alone accounts for $2.5 billion of that run rate — up from $500 million just three months earlier. Consumer subscriptions generated about $1.2 billion in 2025, with 1,000%+ year-over-year growth.</p>
<p>Sounds great, right? Until you look at the other side of the ledger.</p>
<p>Anthropic burned approximately $5.2 billion in 2025. They’ve committed over $80 billion in cloud infrastructure costs through 2029. They just raised $30 billion in a Series G at a $380 billion valuation — the second-largest private tech financing ever, behind only OpenAI.</p>
<p>They’re buying compute at a staggering scale: 1 million Google TPUv7 chips (~$52 billion deal), a dedicated 1,200-acre AWS data center campus in Indiana ($11 billion), and a $50 billion deal with Fluidstack for facilities in Texas and New York. Total committed compute: over 2 gigawatts.</p>
<p>All of this is funded by venture capital and strategic investors (Amazon’s $8B+, Google’s $3B+). Not by your $20/month Pro subscription.</p>
<p>Anthropic projects positive free cash flow by 2027-2028. That’s the plan. But plans require the revenue to actually materialize, the compute to come online in time, and the unit economics to hold as usage scales.</p>
<p>Right now, 60,000 subscribers are overwhelming the existing infrastructure so badly that paying customers can’t work.</p>
<h1 id="the-subsidy-is-collapsing-under-its-own-success"><strong>The Subsidy Is Collapsing Under Its Own Success</strong></h1>
<p>Here’s the dynamic I didn’t fully appreciate three weeks ago.</p>
<p>The subsidy doesn’t end with a price increase. It ends with degradation.</p>
<p>Anthropic can’t raise prices on Pro from 20to20<em>to</em>50 tomorrow — that would cause a revolt and hand users to OpenAI and Google. But they can let the service get worse at the current price. Tighter rate limits. More frequent throttling. Peak-hour queuing. Features that work “sometimes.”</p>
<p>This is exactly what’s happening.</p>
<p>The math is simple. Double the subscribers on the same compute = everyone gets half the capacity. As one Reddit user put it: “selling more seats on the same plane and wondering why legroom is shrinking.”</p>
<p>And Anthropic isn’t alone. Google slashed Gemini API free tier quotas by 50-92% overnight in December 2025. One developer went from 300M+ input tokens per week to hitting limits at less than 9M. OpenAI’s ChatGPT Pro at $200/month is the only major offering that effectively removes caps — but at ten times the price of a Pro subscription.</p>
<p>The pattern across the industry: subsidized tiers are getting squeezed. The compute costs are real. And the bill always comes due.</p>
<h1 id="why-im-not-hitting-limits-and-you-might-not-be-either"><strong>Why I’m Not Hitting Limits (And You Might Not Be Either)</strong></h1>
<p>Here’s a mystery. Despite all this chaos, I’ve barely noticed the rate limits. After reading the threads and the leaked source code, I think I know why.</p>
<p><strong>I almost never resume sessions.</strong> The biggest token drain fires on session resume. My workflow — fresh sessions, agent registration per session, structured <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> — accidentally dodges this bug entirely.</p>
<p><strong>Surgical prompts.</strong> I don’t say “explore my codebase.” I say “read this file and fix this function.” My beads-based task tracking means every session has a specific objective. No wandering. No 94k-token “Explore” runs.</p>
<p><strong>Time zone arbitrage.</strong> IST puts my working hours outside US peak times. When r/ClaudeAI is screaming about rate limits at 2 PM Eastern, it’s midnight for me. I’m coding at 6 AM IST when San Francisco is asleep.</p>
<p><strong>Structured context.</strong> Between <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a>, <a href="http://architecture.md/" rel="external nofollow noopener" class="lnp-link">ARCHITECTURE.md</a>, and explicit file paths, Claude doesn’t need to discover my codebase. It already knows the layout. That’s 90% less indexing work.</p>
<p>This isn’t luck. It’s workflow design. But it reinforces the point from my original post: the subsidy rewards those who use it efficiently. Wasteful usage — open-ended exploration, resumed conversations, vague prompts — burns tokens at 10-50x the rate of focused work.</p>
<h1 id="what-this-means-for-you"><strong>What This Means For You</strong></h1>
<p>If you read my original post and thought you had 18-24 months — you might, on paper. Anthropic has the cash. They have the compute commitments. They project 70billioninrevenueby2028and70<em>billioninrevenueby</em>2028<em>and</em>17 billion in free cash flow.</p>
<p>But the experience of using the product is degrading right now. Not in 18 months. Now.</p>
<p>Here’s what actually matters:</p>
<p><strong>1. Ship before the experience degrades further.</strong> The window isn’t about pricing — it’s about capability per dollar. Today, $20/month gets you frontier model access that would have cost $500/month in API calls two years ago. That ratio is moving in the wrong direction as more users pile in.</p>
<p><strong>2. Optimize your workflow.</strong> Start fresh sessions. Use <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> and <a href="http://architecture.md/" rel="external nofollow noopener" class="lnp-link">ARCHITECTURE.md</a>. Be specific in your prompts. Avoid “Explore” and open-ended commands. These aren’t just productivity tips — they’re rate limit survival strategies.</p>
<p><strong>3. Don’t build on the assumption of unlimited AI access.</strong> If your product or workflow requires constant frontier model access at current prices, you’re building on borrowed time. Build systems that work <em>with</em> AI but can degrade gracefully. Ship products that generate revenue independent of your development tools.</p>
<p><strong>4. The enterprise pivot is coming.</strong> Anthropic’s enterprise revenue is already 80% of total. They have 300,000+ business customers, with large accounts (&gt;$100K ARR) growing 7x year-over-year. Follow the money: consumer subscriptions are the loss leader. Enterprise is the business. When push comes to shove, enterprise gets the compute.</p>
<h1 id="the-real-lesson"><strong>The Real Lesson</strong></h1>
<p>The leaked source code is a metaphor for the entire AI subsidy era.</p>
<p>For weeks, users were burning through rate limits at impossible speeds. They blamed themselves (”skill issue”), they blamed Anthropic (”fix your limits”), they blamed the model (”Claude got dumber”). The actual cause was a two-line bug in a caching function that nobody could see because the code was proprietary.</p>
<p>That’s the subsidy in miniature. You’re using a product where you can’t see the internals, can’t predict the costs, and can’t control when the rules change. The value is extraordinary — right now. But you’re a guest in someone else’s infrastructure, running on someone else’s VC money, subject to someone else’s capacity planning.</p>
<p>The smartest move hasn’t changed since three weeks ago. Ship. Build durable assets — products, content, audiences, skills — while the arbitrage is still available.</p>
<p>But do it faster than you planned. The window isn’t closing in 18 months.</p>
<p>The glass is already cracking.</p>
]]></content:encoded></item><item><title>What Chinese Factories Taught Me About Prompting Claude Code</title><link>https://lakshminp.com/2026/03/claude-code-prompt-engineering/</link><pubDate>Tue, 03 Mar 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/03/claude-code-prompt-engineering/</guid><category>essays</category><category>claude-code</category><category>ai-coding</category><description>A few weeks ago, I fell down a Hacker News rabbit hole at 11pm. Someone had posted a manufacturing post-mortem — one of those beautiful, painful essays where a hardware founder documents exactly how badly they got burned.
This founder had designed a custom lamp. Spent months prototyping. Found a factory in Shenzhen. Shipped 500 units.
When the boxes arrived, the light-entry holes had been used as casting pour-points — the factory needed somewhere to pour the material, saw the holes, and went with it. The cable tails were two centimeters instead of ten. The knobs didn’t fit because the powder coating added thickness that nobody put in the spec. Everything technically matched the purchase order. Nothing actually worked.</description><content:encoded><![CDATA[<p>A few weeks ago, I fell down a Hacker News rabbit hole at 11pm. Someone had posted a manufacturing post-mortem — one of those beautiful, painful essays where a hardware founder documents exactly how badly they got burned.</p>
<p>This founder had designed a custom lamp. Spent months prototyping. Found a factory in Shenzhen. Shipped 500 units.</p>
<p>When the boxes arrived, the light-entry holes had been used as casting pour-points — the factory needed somewhere to pour the material, saw the holes, and went with it. The cable tails were two centimeters instead of ten. The knobs didn’t fit because the powder coating added thickness that nobody put in the spec. Everything technically matched the purchase order. Nothing actually worked.</p>
<p>I read that post-mortem three times. Then I read the top comment, which was one of those sentences that you immediately screenshot because it’s just too true:</p>
<p><em>“Anything you don’t specify will be done at minimum cost.”</em></p>
<p>I put my phone down. I looked at the ceiling. And then I thought about the email sender I’d had Claude Code generate that afternoon.</p>
<p>Let me tell you what I had asked for: “Send a welcome email to new users when they sign up.”</p>
<p>Let me tell you what I got: A function that sent emails. Technically correct. It looped over every new user and called the email API synchronously, one by one, waiting for each response before moving to the next. No rate limiting. No retry logic. No unsubscribe link — because I didn’t ask for one, and CAN-SPAM compliance wasn’t in the prompt. When I ran it against a list of 8,000 users, it fired all 8,000 requests in a tight loop, Gmail flagged the sending domain as a spam source within six hours, and my domain was blacklisted before I’d finished my coffee.</p>
<p>Everything sent. Nothing arrived.</p>
<p>I had been vibe coding with Claude Code for six months at that point, and I thought I was pretty good at it. I could get it to build things fast. I could chain prompts together. I had <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> files and hooks and all the trappings of someone who knew what they were doing.</p>
<p>What I didn’t understand — what the Hacker News post-mortem forced me to understand — is that I had completely misidentified what kind of relationship I was in.</p>
<p>I thought I was pair programming with a senior engineer.</p>
<p>I was issuing purchase orders to a factory.</p>
<p>This distinction sounds philosophical. It isn’t. It has concrete, expensive implications for every vibe coding prompt you write.</p>
<p>A senior engineer fills gaps with judgment. If you say “build auth,” a good senior engineer asks: what are the scale requirements? What’s the threat model? Are we storing PII? They fill the spec gaps with professional standards because they have skin in the game — it’s their name on the code, their reputation on the line, their on-call rotation if it breaks at 3am.</p>
<p>A factory fills gaps with cost optimization. If the spec doesn’t say “cable tails must be 10cm,” the factory cuts them at 2cm. Not because they’re malicious. Because that’s 8cm of wire per unit times 500 units and someone’s margin depends on it. They’re perfectly rational. They’re just optimizing for something that has nothing to do with whether your lamp works.</p>
<p>Claude optimizes for “satisfies the prompt.” That’s the whole job. Your vague prompt is its permission to take shortcuts, and it will take them — not maliciously, but with the same rational efficiency as a factory floor supervisor who notices you didn’t specify the minimum acceptable wire gauge.</p>
<p>Here’s the thing about the hardware community that I find both humbling and enraging: they figured this out decades ago. They built an entire profession around it. These people are called sourcing agents, and their whole job is translating “I want a nice lamp” into a 47-page document covering material density, wire gauge, coating thickness, packaging dimensions, UV stability ratings, and what happens to the tooling if the order falls below minimum quantity.</p>
<p>Forty-seven pages. For a lamp.</p>
<p>In vibe coding, the sourcing agent is you. Most developers have been accidentally promoted to this role without realizing it. They’re still acting like they’re talking to a colleague. They’re actually running a factory and they’re skipping the quality control, the detailed specs, and the first-article inspection — all the boring stuff that hardware people do automatically because they’ve shipped enough garbage to know better.</p>
<p>I’ve started reading Hacker News manufacturing posts specifically to steal their frameworks for this. A few things that have genuinely changed how I write prompts:</p>
<p><strong>Spec your constraints, not just your features.</strong> “Send welcome emails” is a feature request. “Send welcome emails via SES, rate-limited to 14 per second to stay under AWS sending limits, with exponential backoff and a max of 3 retries on failure, an unsubscribe link in the footer per CAN-SPAM, a plain-text fallback alongside the HTML version, and a hard skip for any address that has previously bounced or complained” is a spec. The difference isn’t intelligence — it’s the same way specifying wire gauge isn’t about distrusting your factory. It’s about understanding that factories don’t have opinions about wire gauge. They have margins.</p>
<p><strong>Inspect the first batch before commissioning the full run.</strong> Hardware founders don’t ship the first production run to customers. They order samples. They measure every dimension with calipers. The good ones fly to Shenzhen and stand on the factory floor. The developer equivalent is reading the first 200 lines of generated code before asking Claude to build the next feature on top of it. Check the database schema before building the API on top of it. Read the auth flow before adding the permissions layer. This feels slow. It is much faster than discovering that the foundation is wrong after you’ve built four floors.</p>
<p><strong>Specify what you don’t want.</strong> This one surprised me. Experienced sourcing agents reportedly spend half their spec document on exclusions. “No recycled plastic in structural components.” “No substituted components without written approval.” “No unlicensed firmware.” They’ve learned that a factory will always find the interpretation of the spec that costs them the least, so you have to close the doors. For prompts: “No inline styles. No TypeScript <code>any</code> types. No <code>console.log</code> for error handling. No <code>SELECT *</code> queries. No external dependencies unless they’re in the approved list.” The AI will not volunteer that it’s about to do these things. It will do them and move on.</p>
<p><strong>Budget time for the spec, not just the build.</strong> Hardware founders allocate somewhere between 30-40% of their project timeline to specification work. The manufacturing part — the actual production — is the smaller slice. Vibe coders typically invert this. Five percent on the prompt, ninety-five percent on generating code and then debugging the surprising things that came out of a vague prompt. The debugging is expensive. The spec is cheap.</p>
<p>The thing I keep coming back to is that using Chinese manufacturers is incredible leverage. You can build a physical product without owning a factory, without specialized tooling knowledge, without decades of manufacturing experience. It’s genuinely one of the great unlocks of the modern economy. And it works — when you write the spec correctly.</p>
<p>Using Claude to write code is the same kind of leverage. You can build things without knowing every library, without remembering every API, without holding the entire codebase in your head at once. It works. When you treat it like what it is.</p>
<p>Your prompt is a manufacturing spec. The code is the factory output. The factory will be rational, efficient, and completely indifferent to whether your product actually works.</p>
<p>Write the spec accordingly.</p>
<p>Or enjoy your two-centimeter cable tails.</p>
]]></content:encoded></item><item><title>Claude Code Has Been Navigating Your Codebase Like a Tourist With No Map</title><link>https://lakshminp.com/2026/03/claude-code-lsp-semantic-context-agents/</link><pubDate>Mon, 02 Mar 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/03/claude-code-lsp-semantic-context-agents/</guid><category>essays</category><category>claude-code</category><description>Here’s a thing that happened to me.
I was watching a Claude Code session — one of those where you hand the agent a task and then sit back to observe, feeling very enlightened and modern. The task was simple: find where user authentication was implemented and add a new field to the login flow.
The agent started grepping. authenticate. Then auth. Then login. Then loginUser. Then handleLogin. Each grep taking 3-8 seconds, scanning hundreds of files, returning walls of output full of comments, test fixtures, variable names that happened to contain the word “auth”, README lines I’d written two years ago and forgotten.</description><content:encoded><![CDATA[<p>Here’s a thing that happened to me.</p>
<p>I was watching a Claude Code session — one of those where you hand the agent a task and then sit back to observe, feeling very enlightened and modern. The task was simple: find where user authentication was implemented and add a new field to the login flow.</p>
<p>The agent started grepping. <code>authenticate</code>. Then <code>auth</code>. Then <code>login</code>. Then <code>loginUser</code>. Then <code>handleLogin</code>. Each grep taking 3-8 seconds, scanning hundreds of files, returning walls of output full of comments, test fixtures, variable names that happened to contain the word “auth”, README lines I’d written two years ago and forgotten.</p>
<p>Six minutes in, the agent had read approximately 40% of my codebase and was confidently editing&hellip; a test helper that mocked authentication. Not the actual implementation. A mock. In a test file.</p>
<p>I watched it do this — a system with the reasoning capacity of a senior engineer, burning through context and API calls to do something that VS Code does when I hold Ctrl and click a function name. Something VS Code has done since 2016. Something that takes 50 milliseconds.</p>
<p>This is the state of the art in 2026. The most capable AI coding tool available, navigating your codebase the way your grandfather would navigate a foreign city: slowly, incorrectly, and with a lot of asking for directions from people who don’t know either.</p>
<p>There’s a fix. There are actually two fixes, and you need both. But first I want to explain why the problem is worse than it looks, because if you don’t understand the root cause, you’ll implement half of the solution and wonder why your agents still feel like babysitting.</p>
<p>Let me tell you about grep, and why it’s a disaster for code navigation specifically.</p>
<p>Grep is a text search tool. It finds patterns in text. This is genuinely useful for a lot of things. When you want to find every config file that mentions a database host, grep is perfect. When you want to find a log line, grep is perfect. When you want to navigate code semantically — find where a function is <em>defined</em>, trace what <em>calls</em> a function, understand the type hierarchy — grep is completely wrong for the job. It just happens to be the only hammer available, so everything looks like a nail.</p>
<p>Here’s the specific failure mode. When your agent searches for <code>authenticate</code>, it finds:</p>
<pre><code>auth.service.ts:47:  async authenticate(user: User): Promise&lt;Token&gt; {
auth.service.ts:112: // authenticate is called after 2FA verification
auth.middleware.ts:23: // Middleware that calls authenticate() before protected routes
auth.test.ts:8:   describe('authenticate', () =&gt; {
utils/mock-auth.ts:31:   authenticate: jest.fn().mockResolvedValue(mockToken),
config/dev.ts:15:   authenticateWithMock: true,
README.md:234: ## How to authenticate
</code></pre>
<p>Seven results. One is the actual definition. The agent has to read all of them, reason about which one is the real thing, and then probably read the files surrounding each one to build context. Meanwhile, it’s consuming tokens, spending time, and building a picture of your codebase that’s assembled from grep outputs rather than from actual structural understanding.</p>
<p>The deeper problem: grep doesn’t understand the <em>difference</em> between a definition, a call site, a comment, a test mock, and a config flag. Those are fundamentally different things in the semantic structure of a codebase. A human engineer with IDE tooling can instantly distinguish them. An agent with only grep cannot — it has to infer the difference from text patterns and context, which it does imperfectly, which means it makes wrong edits, which you have to catch and correct, which is why agent sessions still require babysitting.</p>
<p>This is not a clever problem. We solved it for humans a long time ago.</p>
<p>In 2016, Microsoft did something quietly brilliant. They were building VS Code, and they had a problem: every editor had to implement language intelligence from scratch. Vim plugins, Emacs modes, IntelliJ — everyone was reimplementing the same understanding of what a TypeScript file meant, independently, badly, in incompatible ways.</p>
<p>Their solution was the Language Server Protocol. The idea: separate the “smarts” from the editor. Create a standard protocol where a language server — a standalone process that deeply understands a specific language — can talk to any editor that speaks the protocol. Build the language server once, correctly, and every editor gets the benefit.</p>
<p>A language server is not a text search tool. It parses your code into an Abstract Syntax Tree. It resolves types. It builds a symbol table — a complete map of every identifier in your codebase: what it is, where it’s defined, what it references, what references it. When VS Code shows you that <code>authenticate</code> is defined in <code>auth.service.ts</code> on line 47, it’s not searching for the string “authenticate.” It’s looking up <code>authenticate</code> in the symbol table and getting back a precise answer in under 50 milliseconds.</p>
<p>LSP was so obviously right that it became universal. Every serious editor implemented it. Every major language has a language server: <code>pyright</code> for Python, <code>gopls</code> for Go, <code>typescript-language-server</code> for TypeScript, <code>rust-analyzer</code> for Rust, <code>clangd</code> for C/C++. You almost certainly have at least one of these running on your machine right now.</p>
<p>The irony is that we gave AI agents trillion-parameter language models with remarkable reasoning capabilities, and then handed them grep for code navigation. Like building a Formula 1 car and fitting it with bicycle tires.</p>
<p>Claude Code can connect to these language servers. As of early 2026, this is an undocumented community workaround discovered via a GitHub issue — not an official feature. Which is funny, given how much it changes things. Enable it by adding to <code>~/.claude/settings.json</code>:</p>
<pre><code>{
  &quot;env&quot;: {
    &quot;ENABLE_LSP_TOOL&quot;: &quot;1&quot;
  }
}
</code></pre>
<p>Or export it in your shell profile if you prefer:</p>
<pre><code>export ENABLE_LSP_TOOL=1
</code></pre>
<p>Then install the language server plugin for your stack. Claude Code has a plugin system for this — update the marketplace first, then install:</p>
<pre><code>claude plugin marketplace update claude-plugins-official

# TypeScript/JavaScript
claude plugin install typescript-lsp
npm install -g typescript-language-server typescript

# Python
claude plugin install pyright-lsp
npm install -g pyright

# Go
claude plugin install gopls-lsp
go install golang.org/x/tools/gopls@latest

# Rust
claude plugin install rust-analyzer-lsp
rustup component add rust-analyzer
</code></pre>
<p>One gotcha that will silently waste your time: a plugin can be installed but disabled. An installed, disabled plugin does nothing — no LSP server registers at startup, no tools become available, no error. Just grep, same as before. After installing, run <code>claude plugin list</code> and confirm the status reads <code>enabled</code>. If it shows <code>disabled</code>, run <code>claude plugin enable &lt;name&gt;</code>. Check this before you spend 20 minutes wondering why nothing changed.</p>
<p>Once enabled, your agent gets access to tools that most people don’t know exist:</p>
<p><code>goToDefinition</code> — exact location of any symbol’s definition. Not “files that contain this string.” The <em>definition</em>. In ~50ms.</p>
<p><code>findReferences</code> — every call site in your entire codebase. Every single one, sorted, precise, with file and line number.</p>
<p><code>workspaceSymbols</code> — search your codebase by symbol name. Returns only actual code symbols (functions, classes, interfaces, variables) — not comments, not strings, not README lines.</p>
<p><code>hover</code> — full type information for any identifier. When the agent is about to call a function, it can check the exact signature first rather than guessing.</p>
<p><code>diagnostics</code> — real-time type errors. When the agent changes a function signature, the language server immediately reports every caller that’s now broken. In the same turn. Before the broken code ever runs.</p>
<p>That last one changes the loop entirely. Without LSP, the workflow is: agent makes a change → change breaks something → you run tests → tests fail → agent fixes it → might break something else → iterate. You’re discovering errors through tests, which means you’re discovering them late, which means multiple turns of cleanup for each mistake.</p>
<p>With LSP, the workflow is: agent makes a change → diagnostics immediately flag every type error caused by that change → agent fixes everything in the same turn. Error discovery goes from “whenever you run tests” to “immediately.” This alone is worth the two minutes it takes to set up.</p>
<p>Here’s the catch, and it’s not obvious until you run into it.</p>
<p>Even with LSP enabled and plugins installed and confirmed active, Claude still <em>prefers</em> grep. Grep is familiar, grep is in its training distribution, grep is what it reaches for first. Having the tools available doesn’t automatically mean Claude will use them.</p>
<p>Add this to your <code>CLAUDE.md</code>:</p>
<pre><code>## Code Navigation

Prefer LSP tools over Grep for any code navigation task:
- Use workspaceSymbol to find symbols by name
- Use goToDefinition to find where something is defined
- Use findReferences to find all call sites
- Use diagnostics after any edit to catch type errors immediately

Use Grep only for text search: log messages, comments, config values,
string literals. Never use Grep to find function definitions.
</code></pre>
<p>Explicit instructions in <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> override default behavior. This is a documented pattern: the tools exist, but you have to tell the agent to use them. Think of it as configuring the agent’s preferences, not patching the agent’s capabilities.</p>
<p>Now here’s the part where most people stop, and where they shouldn’t.</p>
<p>LSP gives your agent GPS. It knows <em>how</em> to navigate. <code>findReferences</code> from anywhere in the codebase will return exact results. But GPS without a destination is just a compass. Your agent still has to figure out <em>where to go</em> before it can navigate there efficiently.</p>
<p>Think about how an experienced engineer ramps up on a new codebase. They don’t start by grepping for things. They start by asking questions: where does the auth layer live? What’s the database access pattern? How do the services communicate? They build a mental model first, then navigate with precision.</p>
<p>Your agent has no mental model of your codebase unless you give it one. Every session starts cold. It has the code itself (too much to read exhaustively) and the tools to navigate it (useful once oriented) but no map. So it wanders.</p>
<p>The second layer is a structured description of your codebase’s architecture. Not documentation. Not a README. A map for the agent — written in terms of what the agent needs to know to get oriented quickly:</p>
<pre><code>## Codebase Architecture

**Entry point:** src/server.ts bootstraps the app. All route registration happens here.

**Auth layer:** Everything authentication-related lives in /src/auth.
The entry point is `authenticate()` in auth.service.ts.
JWT handling is in auth.middleware.ts. Session storage is Redis via auth.session.ts.
Never bypass the middleware — it handles rate limiting and audit logging.

**Services:** Business logic in /src/services.
PaymentService, UserService, NotificationService are the big three.
Services never call each other directly — all cross-service communication
routes through the event bus in /src/events/index.ts.

**Database:** Prisma ORM. Never write raw SQL — always go through the Prisma client.
Schema lives in /prisma/schema.prisma. Run `npm run db:migrate` after schema changes.

**External integrations:** Stripe in /src/integrations/stripe,
SendGrid in /src/integrations/email. Each integration has a fake for testing.
</code></pre>
<p>You put this in <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a>, or in a dedicated <a href="http://architecture.md/" rel="external nofollow noopener" class="lnp-link">ARCHITECTURE.md</a> that <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a> imports via <code>@ARCHITECTURE.md</code>.</p>
<p>What changes: your agent starts the session <em>oriented</em>. When you ask it to add a new payment method, it already knows that payment logic lives in <code>/src/services/PaymentService</code>, that external Stripe calls go through <code>/src/integrations/stripe</code>, and that services communicate through the event bus. It doesn’t need to explore your codebase to discover the architecture. It can go directly to the right place and navigate from there with LSP precision.</p>
<p>The GPS analogy only goes so far. A better way to think about it: LSP is your agent’s ability to look something up instantly. Semantic context is the agent knowing what to look up. Both are required. Without the map, LSP is a fast tool pointed in random directions. Without LSP, the map tells you where to go but getting there is still six minutes of grepping.</p>
<p>Together, your agent works the way a senior engineer works on a codebase they know well: they know the territory, they navigate precisely, and they catch their own mistakes before committing them.</p>
<p>The reason I keep coming back to this: the numbers suggest agents are about to do a lot more real work.</p>
<p>Michael Truell <a href="https://x.com/mntruell/status/2026736314272591924" rel="external nofollow noopener" class="lnp-link">announced recently</a> that Cursor now has 2x more agent users than Tab (autocomplete) users. Agent usage is up 15x in a year. More than a third of PRs merged at Cursor are created by agents running autonomously in the cloud — not a human in the loop, not autocomplete suggestions, agents doing complete pieces of work end-to-end.</p>
<p>If that’s the direction — and the trajectory makes it pretty clear it is — then agents navigating codebases with grep is a bottleneck at the wrong layer. You’ve solved the intelligence problem. You have an agent that can reason about complex changes across multiple files. You have not solved the navigation problem, which means the intelligence is being spent on finding things instead of changing them. It’s like hiring a brilliant architect and making them do their own filing.</p>
<p>LSP and semantic context are table stakes for agent-native codebases. The fact that LSP is buried in settings and semantic maps are a community pattern rather than a first-class feature is a product gap. It’ll get closed. But right now you have to close it yourself, and it takes about thirty minutes.</p>
<p>Set up the language server for your stack. Enable LSP in settings.json. Tell Claude to prefer it in <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a>. Write an architecture section that orients the agent in the first turn. Thirty minutes of setup for sessions that actually feel autonomous.</p>
<p>Your future self will be insufferably smug about having done this early. That’s a reasonable outcome.</p>
]]></content:encoded></item><item><title>While You Panic About AI Taking Jobs, I Built $200/Mo Tools</title><link>https://lakshminp.com/2026/02/stop-paying-ai-wrappers/</link><pubDate>Wed, 18 Feb 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/02/stop-paying-ai-wrappers/</guid><category>essays</category><category>claude-code</category><category>saas</category><description>I was about to click “Subscribe: $29/month” on yet another AI content tool when I stopped.
Not because $29 was a lot. I’ve subscribed to worse. I have a graveyard of forgotten SaaS products auto-renewing somewhere in my credit card statement, silently draining money for tools I used exactly twice.
No, I stopped because I realized something embarrassing.
This tool was literally just a pretty wrapper around Claude. Same AI I already pay for. Same capabilities. They’d added a nice UI, a payment form, and approximately zero additional value.</description><content:encoded><![CDATA[<p>I was about to click “Subscribe: $29/month” on yet another AI content tool when I stopped.</p>
<p>Not because $29 was a lot. I’ve subscribed to worse. I have a graveyard of forgotten SaaS products auto-renewing somewhere in my credit card statement, silently draining money for tools I used exactly twice.</p>
<p>No, I stopped because I realized something embarrassing.</p>
<p>This tool was literally just a pretty wrapper around Claude. Same AI I already pay for. Same capabilities. They’d added a nice UI, a payment form, and approximately zero additional value.</p>
<p>I was about to pay $29/month to rent something I could build in an afternoon.</p>
<p>So I closed the tab. Opened Claude Code. And two hours later, I had my own version. No usage limits. No subscription. Mine forever.</p>
<p>That was six months ago. Since then, I’ve built six tools. I’ve eliminated $200/month in subscriptions. And I’ve realized something that changed how I think about this whole “AI is coming for your job” panic.</p>
<h1 id="everyones-worried-about-the-wrong-thing"><strong>Everyone’s Worried About the Wrong Thing</strong></h1>
<p>My LinkedIn feed is a horror show right now. Every other post is either “AI will take your job” or “Here’s how to survive the AI apocalypse” or some variation of “the robots are coming, repent.”</p>
<p>The advice is always the same. Learn to prompt. Adapt or die. Get your finances in order because disruption is coming.</p>
<p>Maybe it is. I don’t know the future any better than you do.</p>
<p>But here’s what I do know: while everyone’s stockpiling survival advice, I’ve been using those same AI tools to eliminate $200/month in software subscriptions. Tools I was paying for six months ago? I own them now. Forever. Zero recurring cost.</p>
<p>Same technology. Completely different mindset.</p>
<p>Let me show you what I mean.</p>
<h1 id="the-batch-pdf-processor-i-built-instead-of-uploading-50-files"><strong>The Batch PDF Processor I Built Instead of Uploading 50 Files</strong></h1>
<p>I had 50 research papers to process. Extract abstracts, pull out key findings, grab any data tables, save everything as searchable markdown.</p>
<p>Sure, Claude can read a PDF. One at a time. Upload, wait, copy the output, upload the next one, repeat 49 more times.</p>
<p>Old me would’ve done exactly that. Or Googled “batch PDF extraction tool,” found something that charges per page, done the math, decided it wasn’t worth it, and then manually uploaded 50 files anyway.</p>
<p>New me? I built a script.</p>
<p>Ninety minutes later, I had a skill that loops through a folder, extracts text and tables from each PDF using PyMuPDF, summarizes each one, and saves structured markdown files. Point it at a folder, go make coffee, come back to 50 organized summaries.</p>
<pre><code>&gt; Process all PDFs in ~/research/papers
[loops through 50 files]
[extracts + summarizes each]
[saves to ~/research/summaries/]
</code></pre>
<p>No uploading files one by one. No copy-paste marathon. No usage limits. Runs locally, so nothing leaves my machine.</p>
<p>The tool isn’t “PDF extraction”: Claude already does that. The tool is <em>automation</em>. Batch processing. The boring plumbing that turns a manual 3-hour task into a 5-minute one.</p>
<h1 id="the-video-transcription-pipeline-that-changed-how-i-learn"><strong>The Video Transcription Pipeline That Changed How I Learn</strong></h1>
<p>I’m an infoproduct junkie. Courses, masterclasses, workshops: if someone’s selling knowledge in video form, I’ve probably bought it. My Teachable and Gumroad purchase history is embarrassing. Hours and hours of content sitting in various dashboards, waiting to be watched.</p>
<p>Old me would take notes while watching. Pause, scribble, play, pause, scribble. Retain maybe 30% of it. Forget the rest within a week.</p>
<p>New me? I feed the videos to my video-distill skill. It transcribes everything with Whisper, distills it into readable chapters with Claude, and exports to EPUB.</p>
<p>Two hours of course video becomes a 12,000-word mini-book on my Kindle that I can search, highlight, and reference forever.</p>
<p>Build time: 2 hours.</p>
<p>Previous cost: $50-200 per course for transcription services (or just&hellip; not having transcripts and forgetting everything).</p>
<p>The ROI on courses went from “eh, probably worth it” to “this is a no-brainer.”</p>
<h1 id="the-pattern-nobody-wants-to-admit"><strong>The Pattern Nobody Wants to Admit</strong></h1>
<p>Here’s what I noticed while building these:</p>
<p><strong>Every AI SaaS is a thin wrapper around the same AI you already have access to.</strong></p>
<ul>
<li>
<p>Batch file processing? A loop + Claude.</p>
</li>
<li>
<p>Video transcription? Whisper + Claude.</p>
</li>
<li>
<p>Content research? Web fetch + Claude.</p>
</li>
<li>
<p>Cross-posting? Template formatting + Claude.</p>
</li>
<li>
<p>Writing assistant? Prompt engineering + Claude.</p>
</li>
</ul>
<p>The “product” is convenience packaging. A nice UI, hosting, a payment gateway, customer support. Sometimes that’s worth paying for.</p>
<p>But most of the time? You can build 80% of what you need in under two hours.</p>
<p>I’ve done this six times now. Total build time: about 14 hours. One weekend, spread across a few months.</p>
<p>Total monthly savings: $199.</p>
<p>Total annual savings: $2,388.</p>
<p>Tools I now own forever: 6.</p>
<h1 id="the-real-question-nobodys-asking"><strong>The Real Question Nobody’s Asking</strong></h1>
<p>While everyone argues about whether AI will take their job, here’s what I’m thinking:</p>
<p><strong>It’s not AI vs humans. It’s humans with AI vs humans without.</strong></p>
<p>The lawyer who refuses to use AI for contract review loses to the lawyer who uses it and handles 5x more clients.</p>
<p>The developer who doesn’t use AI for code generation loses to the developer who does and ships features in days instead of weeks.</p>
<p>The writer who thinks AI is “cheating” loses to the writer who uses it for research and drafts 10x more content.</p>
<p>The people who lose their jobs to AI won’t be the ones whose work AI can theoretically do.</p>
<p>They’ll be the ones who <em>didn’t use AI</em> and got outpaced by someone who did.</p>
<h1 id="what-i-actually-do-instead-of-panicking"><strong>What I Actually Do Instead of Panicking</strong></h1>
<p>I audit my subscriptions. Every few months, I go through my recurring charges and ask: “Is this just an AI wrapper?”</p>
<p>If yes, I build a replacement. If the replacement covers 80% of my use cases, I cancel the subscription.</p>
<p>Here’s my current hit list:</p>
<ul>
<li>
<p><strong>Batch PDF processing</strong>: replaced manual uploads with pdf-reader skill (90 min)</p>
</li>
<li>
<p><strong>Video transcription</strong>: replaced $50-200/course services with video-distill skill (2 hours)</p>
</li>
<li>
<p><strong>Ebook formatting</strong>: replaced $30/book services with epub-builder skill (1 hour)</p>
</li>
<li>
<p><strong>Content research</strong>: replaced $29/mo tools with compose skill (3 hours)</p>
</li>
<li>
<p><strong>Cross-posting</strong>: replaced $50/mo tools with distribute skill (2 hours)</p>
</li>
<li>
<p><strong>Writing assistant</strong>: replaced $20/mo Sudowrite etc with fiction-writer skill (4 hours)</p>
</li>
</ul>
<p>Not everything is worth building. QuickBooks? Keep paying. Complex Zapier automation with 50 integrations? Probably keep paying. Hosted databases? Definitely keep paying.</p>
<p>But simple AI wrappers that just call Claude with a prompt and charge you monthly for it? Those are dying. You can own that.</p>
<h1 id="this-is-what-leverage-actually-looks-like"><strong>This Is What Leverage Actually Looks Like</strong></h1>
<p>When you own your tools, you can modify them to fit your exact workflow. Combine them in ways SaaS products can’t. Build competitive advantages nobody else has.</p>
<p>My <code>distribute</code> skill cross-posts to five platforms in one command. Most people manually post to each: different formatting, different copy, different everything. That’s 30 minutes per post. I do it in 2 minutes.</p>
<p>Over a year, that’s 50+ hours saved. For one skill.</p>
<p>My <code>video-distill</code> skill turns courses into searchable mini-books. Most people watch courses once and forget 80% of it. I have permanent reference material I can search and review anytime.</p>
<p>This is how one-person operations beat ten-person teams. Not by working harder. By owning tools that make you 10x faster.</p>
<h1 id="two-paths"><strong>Two Paths</strong></h1>
<p>The AI revolution is here. Not coming. Here.</p>
<p><strong>Path A:</strong> Read about how AI is going to disrupt your career. Worry about it. Prepare for the worst. Maybe it happens, maybe it doesn’t. Either way, you spent months anxious instead of building.</p>
<p><strong>Path B:</strong> Use AI to eliminate expenses, build tools, ship faster, own your stack. If disruption comes, you’re already leveraged. If it doesn’t, you still saved $2,400/year and got 10x faster.</p>
<p>I’m on Path B.</p>
<p>I built six skills in 14 hours. I save $200/month. I own tools that do exactly what I need, with no usage limits, no pricing tiers, and no risk of some startup getting acqui-hired and shutting down my workflow.</p>
<p>And I’m shipping four SaaS products while working two day jobs, because I have the leverage to do it.</p>
<p>You can panic, or you can build.</p>
<p>I’m building.</p>
<h1 id="if-you-want-to-start"><strong>If You Want to Start</strong></h1>
<p>Here’s the playbook I use every time.</p>
<h2 id="step-1-pick-your-target"><strong>Step 1: Pick Your Target</strong></h2>
<p>Start with something you use weekly but don’t need enterprise features for. The sweet spot is tools where you’re paying for convenience, not capability.</p>
<p><strong>Good first targets:</strong></p>
<ul>
<li>
<p>Batch processing workflows (loop through files, process each, save outputs)</p>
</li>
<li>
<p>Video/audio transcription (Whisper is free and runs locally)</p>
</li>
<li>
<p>Content research and brainstorming (web scraping + Claude)</p>
</li>
<li>
<p>Grammar/style checking (Claude prompts replace Grammarly)</p>
</li>
<li>
<p>Format conversion pipelines (markdown → EPUB, video → transcript → summary)</p>
</li>
</ul>
<p><strong>Bad first targets:</strong></p>
<ul>
<li>
<p>Accounting software (compliance, integrations, audit trails)</p>
</li>
<li>
<p>Complex multi-step automation with 20+ triggers</p>
</li>
<li>
<p>Anything requiring hosted infrastructure you don’t want to manage</p>
</li>
<li>
<p>Real-time collaboration tools (Google Docs, Figma)</p>
</li>
</ul>
<h2 id="step-2-describe-the-workflow-not-the-tool"><strong>Step 2: Describe the Workflow, Not the Tool</strong></h2>
<p>Don’t say “build me a transcription tool.” Say “I have 20 course videos. I want to transcribe each one, distill the key points, and save them as markdown files organized by chapter.”</p>
<p>The more specific you are about your <em>actual workflow</em>, the better the tool fits. You’re not building a generic SaaS: you’re building exactly what you need.</p>
<h2 id="step-3-start-ugly-iterate-fast"><strong>Step 3: Start Ugly, Iterate Fast</strong></h2>
<p>Your first version will be rough. That’s fine.</p>
<p>My video transcription pipeline started as a janky script that choked on long files. I fixed the edge cases as I hit them. Now it handles 3-hour lectures without breaking a sweat.</p>
<p>Don’t try to build the polished SaaS version. Build the “works for my specific use case” version. That takes 90 minutes, not 90 days.</p>
<h2 id="step-4-the-80-test"><strong>Step 4: The 80% Test</strong></h2>
<p>Run your homegrown tool alongside the paid one for 30 days. Track when you reach for the paid tool instead.</p>
<p>If your skill handles 80% of cases, cancel the subscription. The remaining 20%? Either iterate on your skill or accept the occasional manual workaround.</p>
<p>Perfect is the enemy of $X/month forever.</p>
<h2 id="step-5-compound-it"><strong>Step 5: Compound It</strong></h2>
<p>Once you’ve built one, you’ll notice patterns. The same techniques: file handling, API calls, text processing, output formatting: show up everywhere.</p>
<p>Your second skill takes half the time. Your fifth takes 20 minutes.</p>
<p>This is how you end up owning your entire stack without spending months building it.</p>
<h2 id="the-prompt-that-starts-everything"><strong>The Prompt That Starts Everything</strong></h2>
<p>If you’re using Claude Code or similar:</p>
<pre><code>I want to build a tool that [specific workflow].
I currently do this by [current manual process or paid tool].
My input is [what you're working with].
I want the output to be [format and destination].
What's the simplest way to build this?
</code></pre>
<p>Then iterate. The AI will ask clarifying questions, suggest approaches, write code. You test, refine, test again.</p>
<p>Two hours later, you own something you would’ve rented forever.</p>
<p>The people who win the next decade won’t be the ones who worried about AI disruption.</p>
<p>They’ll be the ones who used AI to build tools, eliminate costs, and move faster than everyone else.</p>
<p>Don’t rent your stack. Own it.</p>
<p><em>P.S.: The AI wrapper economy is dying. Every “AI-powered” tool that’s just a UI around Claude or GPT is on borrowed time. The moment users realize they can build 80% of that themselves, the subscription revenue evaporates. If you’re building an AI SaaS, make sure your value is in the 20% that can’t be replicated in two hours.</em></p>
]]></content:encoded></item><item><title>Claude Code Is Incredible. It Also Almost Shipped 8 Production Bombs Last Week.</title><link>https://lakshminp.com/2026/01/claude-code-production-bombs/</link><pubDate>Mon, 12 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/claude-code-production-bombs/</guid><category>essays</category><category>claude-code</category><description>What 15 years of production scars are still good for Last week I caught eight bugs across three projects. Not typos. Not missing semicolons. Real, ship-breaking, production-melting problems that would have sailed right past code review and into the waiting arms of actual users.
Claude Code wrote the code. Claude Code passed the tests. Claude Code would have happily deployed it.
And Claude Code had no idea anything was wrong.</description><content:encoded><![CDATA[<h2 id="what-15-years-of-production-scars-are-still-good-for"><strong>What 15 years of production scars are still good for</strong></h2>
<p>Last week I caught eight bugs across three projects. Not typos. Not missing semicolons. Real, ship-breaking, production-melting problems that would have sailed right past code review and into the waiting arms of actual users.</p>
<p>Claude Code wrote the code. Claude Code passed the tests. Claude Code would have happily deployed it.</p>
<p>And Claude Code had no idea anything was wrong.</p>
<p>I’ve written about <a href="https://lakshminp.substack.com/p/the-invisible-tax-you-pay-when-you" rel="external nofollow noopener" class="lnp-link">comprehension debt</a> and argued that <a href="https://lakshminp.substack.com/p/clean-code-is-dead-long-live-clean" rel="external nofollow noopener" class="lnp-link">clean code is dead</a>. But here’s the uncomfortable third act: even if you understand your specs perfectly and verify your outcomes ruthlessly, there’s a whole category of knowledge that AI simply doesn’t have.</p>
<p>Call it production intuition. Call it battle scars. Call it “I’ve been burned by this exact thing before.”</p>
<p>Whatever you call it, you can’t prompt your way into it.</p>
<h3 id="the-localhost-delusion"><strong>The Localhost Delusion</strong></h3>
<p>Here’s what AI is optimized for: making code that works on your machine, right now, with your current data, under ideal conditions.</p>
<p>Here’s what AI is catastrophically bad at: imagining your code running on three replicas behind a load balancer at 3am when the database is under pressure and someone’s running a batch job that nobody documented.</p>
<p>A Fortune 100 developer on Reddit <a href="https://reddit.com/r/ExperiencedDevs/comments/1mg2r6y/the_era_of_ai_slop_cleanup_has_begun/" rel="external nofollow noopener" class="lnp-link">put it bluntly</a>: “There’s a lot of vibe coded slop that works well for MVP but will absolutely fall apart under stress and within production environments once they scale to more users. It doesn’t really reveal itself until later when it’s much more difficult to fix.”</p>
<p>Later. When it’s difficult. The horror.</p>
<p>Let me walk you through what “later” looked like for me this week.</p>
<h3 id="pattern-1-production-blindness"><strong>Pattern 1: Production Blindness</strong></h3>
<p><strong>The Concurrency Landmine</strong></p>
<p>I’m building a tool that routes MCP calls. Claude wrote it in Python — clean, well-structured, exactly what I asked for. Worked beautifully in testing. One request, one response, everybody’s happy.</p>
<p>Then I imagined what happens when three Claude Code sessions hit it simultaneously.</p>
<p>Oh.</p>
<p><em>Oh no.</em></p>
<p>Python’s threading model and the phrase “parallel requests” get along about as well as cats and bathtubs. Claude’s solution? More Python. Refactor this, optimize that. Very confident. Would’ve worked for lower concurrency.</p>
<p>But I knew this thing needed to handle dozens of parallel sessions. I prompted it to rewrite in Go. Claude nailed the port — goroutines, channels, the works. Problem solved in an afternoon.</p>
<p>The code was excellent. The language selection wasn’t. Claude optimizes brilliantly within the box you give it. It just won’t question whether you’re in the right box.</p>
<p><strong>The Replicas Problem</strong></p>
<p>Auth rate limiting. Claude implemented it in-memory with a clean sliding window algorithm. Textbook correct. Tests pass. Ship it.</p>
<p>One replica: perfect.<br>
Two replicas: every user gets double the rate limit.<br>
Three replicas: chaos.</p>
<p>This is distributed systems 101. The kind of thing you learn after you’ve been paged at 2am because someone figured out they could hit your API from three different IPs and get 3x the rate limit.</p>
<p>When I pointed this out, Claude immediately suggested Redis-backed rate limiting with proper distributed locking. Great solution. But it didn’t think about replicas until I did. Claude builds for the deployment model you describe. If you don’t describe it, localhost is the default.</p>
<p><strong>The Sync Task Landmine</strong></p>
<p>An API endpoint runs a long-running task. Claude implemented it synchronously — straightforward, easy to understand, does exactly what the tests verify.</p>
<p>Deploy it. First user clicks the button. 30-second timeout. 504 Gateway Timeout.</p>
<p>When I explained the problem, Claude refactored it to Celery with proper task queuing, retry logic, and status polling. Solid implementation. Took maybe 20 minutes.</p>
<p>But here’s the thing: I only caught it because I tested the actual user flow, not just the unit tests. Claude implemented what I asked for. I didn’t ask for “an endpoint that won’t timeout in production.” That’s on me. But it’s also the kind of thing I’ve learned to check after watching synchronous endpoints die approximately 47 times.</p>
<h3 id="pattern-2-ecosystem-amnesia"><strong>Pattern 2: Ecosystem Amnesia</strong></h3>
<p><strong>The Deprecation You Won’t Find on Stack Overflow</strong></p>
<p>Supabase deprecated their API keys. Not in a big announcement. Not in the docs you’d naturally read. In a <a href="https://github.com/orgs/supabase/discussions/29260" rel="external nofollow noopener" class="lnp-link">GitHub discussion</a> with 43 comments and a lot of confused developers.</p>
<p>I found out because I read the discussion. I then had to fix three separate projects.</p>
<p>Claude doesn’t read GitHub discussions. Claude doesn’t know what the community is grumbling about. Claude’s knowledge is frozen in time, and the ecosystem keeps moving.</p>
<p><strong>The “Works But Wrong” Framework Choice</strong></p>
<p>Claude encrypted secrets at rest using Fernet keys. Technically correct. Cryptographically sound. Tests pass. Secure enough.</p>
<p>But I’m using Supabase. Supabase has a vault feature built specifically for this. When I mentioned it, Claude migrated everything over cleanly — proper RLS policies, the works.</p>
<p>The Fernet implementation wasn’t <em>wrong</em>. It just wasn’t the <em>right</em> choice for this stack. Supabase vault means one less thing I manage, one less key rotation I handle, one less piece of infrastructure to think about.</p>
<p>Claude doesn’t know the zeitgeist of “what’s the idiomatic way to do this in Supabase.” That knowledge lives in community forums, Discord servers, and the muscle memory of people who’ve shipped Supabase apps before.</p>
<p><strong>The Build System That Wasn’t</strong></p>
<p>I made UI mockups with Tailwind CSS(using Claude, to be fair). Told Claude to use them. Claude happily served Tailwind from a CDN.</p>
<p>In development? Fine.<br>
In production? Every page load fetches the entire Tailwind library. Uncompiled. Unoptimized. Approximately 300KB of CSS you don’t need.</p>
<p>Claude knows what Tailwind is. Claude doesn’t know that real projects compile it. That’s the kind of tribal knowledge you pick up by shipping things and watching your Lighthouse scores crater.</p>
<h3 id="pattern-3-verification-vacuum"><strong>Pattern 3: Verification Vacuum</strong></h3>
<p><strong>The Cache That Wasn’t</strong></p>
<p>Database queries were getting slow. I asked Claude to add a caching layer. Claude wrote a beautiful Redis caching module — proper TTLs, cache invalidation on writes, the works. Tests for the module passed. I shipped it, watched the deployment go green, felt the warm glow of productivity.</p>
<p>The cache wasn’t being hit.</p>
<p>Claude built an excellent caching module. Claude did not check whether the actual query functions were <em>calling</em> the cache. The module worked perfectly. Nothing was using it. Every request still hit the database directly.</p>
<p>I discovered this by checking Redis after a few hundred requests. Empty. Revolutionary debugging technique, I know.</p>
<p>Could tests have caught this? Sure. Integration tests that verify “when I make this API call, Redis gets a cache entry.” But I’d need to know to write that test. Claude wrote unit tests for the caching module. I should have asked for end-to-end verification. The knowledge that “modules can exist without being properly wired up” is experience. Pattern recognition. The scar tissue from shipping features that weren’t actually features before.</p>
<h3 id="pattern-4-architectural-judgment"><strong>Pattern 4: Architectural Judgment</strong></h3>
<p><strong>The Multitenancy Time Bomb</strong></p>
<p>A project using Qdrant vector database. Users store embeddings. Multiple users. Shared infrastructure.</p>
<p>The question that should wake you up at night: Can User A see User B’s data?</p>
<p>Claude’s implementation used collection-level isolation with proper tenant IDs in the filter queries. Reasonable approach. Worked in my tests.</p>
<p>But a thorough multitenancy review? The kind where you trace every query path, every edge case, every possible way data could leak between tenants? Where you think about what happens if someone forgets to pass the tenant filter? Where you consider whether the default behavior is secure-by-default or insecure-by-default?</p>
<p>That review took me two hours. Claude can help <em>execute</em> the fixes I identify, but it won’t spontaneously think “hey, multitenancy is a critical architectural decision that deserves paranoid scrutiny.”</p>
<p>Get multitenancy wrong and you’re on the front page of Hacker News, and not in the good way. Claude builds features. You build threat models.</p>
<h3 id="what-the-discourse-gets-wrong"><strong>What The Discourse Gets Wrong</strong></h3>
<p>Here’s what bothers me about the AI discourse: both sides are missing the point.</p>
<p>The AI skeptics say “AI code is garbage, don’t use it.” That’s wrong. Claude Code is incredibly useful. I ship faster. I handle complexity I couldn’t handle alone. The leverage is real.</p>
<p>The AI evangelists say “AI will replace developers, just describe what you want.” That’s also wrong. Describing what you want is the <em>easy</em> part. Knowing what you <em>should</em> want — that’s where the experience lives.</p>
<p>Someone on r/ExperiencedDevs <a href="https://reddit.com/r/ExperiencedDevs/comments/1on84au/ai_wont_make_coding_obsolete_coding_isnt_the_hard/" rel="external nofollow noopener" class="lnp-link">nailed it</a>: “Coding is the boring/easy part. Typing is just transcribing decisions into a machine. The real work is upstream: understanding what’s needed, resolving ambiguity, negotiating tradeoffs, and designing coherent systems.”</p>
<p>The developers who thrive aren’t the ones who write the most code. They’re the ones who catch the multitenancy bug before it ships. Who know that in-memory rate limiting won’t scale. Who’ve been burned by synchronous endpoints and CDN-served CSS and proxies that aren’t wired up.</p>
<p>Experience isn’t knowing the syntax. Experience is knowing the failure modes.</p>
<h3 id="whats-actually-worth-learning"><strong>What’s Actually Worth Learning</strong></h3>
<p>So what’s worth learning when AI can write the code?</p>
<p><strong>Production Thinking</strong></p>
<p>This is the big one. Every example above comes back to it: Claude builds for localhost, you build for production.</p>
<p>Concretely, this means developing instincts for questions like:</p>
<ul>
<li>
<p>“What happens when there are multiple replicas?” (Rate limiting, session state, caching, file storage — anything in-memory becomes a distributed systems problem)</p>
</li>
<li>
<p>“What happens under load?” (Synchronous operations become timeouts. N+1 queries become database meltdowns. That “fast enough” endpoint becomes a bottleneck)</p>
</li>
<li>
<p>“What happens when dependencies fail?” (Database is slow. External API is down. Redis is unreachable. Do you degrade gracefully or explode?)</p>
</li>
<li>
<p>“What happens at 3am when nobody’s watching?” (Background jobs. Retry logic. Dead letter queues. The things that fail silently)</p>
</li>
</ul>
<p>How do you learn this? You can’t shortcut it. You deploy things. You watch them break. You read post-mortems. You get paged. You develop a paranoid imagination for failure modes.</p>
<p>But you can accelerate it: before you ship, spend 10 minutes imagining the deployment. Draw the boxes. How many instances? What’s in front of them? Where’s the state? What’s shared? This exercise catches 80% of the issues I described above.</p>
<p><strong>Ecosystem Intuition</strong></p>
<p>This is knowing the <em>zeitgeist</em> of your stack — not just what’s possible, but what’s idiomatic. What the community actually uses. What’s deprecated but still in the docs. What’s new but not proven.</p>
<p>Concretely:</p>
<ul>
<li>
<p>Read the GitHub discussions, not just the docs. That’s where deprecations get announced, migration paths get debated, and footguns get documented.</p>
</li>
<li>
<p>Follow the maintainers on Twitter/X or Bluesky. They’ll tell you about breaking changes before the docs catch up.</p>
</li>
<li>
<p>Lurk in Discord servers. The “how should I do X” discussions reveal what’s considered best practice.</p>
</li>
<li>
<p>Actually ship with the stack. The difference between “I’ve read about Supabase” and “I’ve shipped three apps with Supabase” is enormous.</p>
</li>
</ul>
<p>The goal: when Claude suggests an approach, you can immediately sense whether it’s the “right” way or just “a” way. Fernet encryption vs Supabase vault. CDN Tailwind vs compiled Tailwind. Redis rate limiting vs in-memory rate limiting. These aren’t in the documentation. They’re in the collective experience.</p>
<p><strong>Architectural Paranoia</strong></p>
<p>Some decisions are easy to change later. Some aren’t. Knowing the difference is half of senior engineering.</p>
<p>The ones that are hard to reverse:</p>
<ul>
<li>
<p><strong>Multitenancy model</strong>: Shared database with tenant IDs? Separate schemas? Separate databases? Choose wrong and you’re rewriting everything.</p>
</li>
<li>
<p><strong>Auth architecture</strong>: Where do tokens live? How do sessions work? What’s the refresh flow? Changing this later breaks every client.</p>
</li>
<li>
<p><strong>Data model fundamentals</strong>: Relational vs document. Normalized vs denormalized. Adding a column is easy. Restructuring your entire data model is not.</p>
</li>
<li>
<p><strong>API contract design</strong>: Once clients depend on your response shape, changing it is a versioning nightmare.</p>
</li>
</ul>
<p>For each of these, Claude will happily implement whatever you ask. It won’t stop and say “are you sure about this? This is hard to change later.” That paranoia is your job.</p>
<p>My rule: for any architectural decision I can’t easily reverse, I spend at least an hour thinking about alternatives before I let Claude write the first line.</p>
<p><strong>Verification Instincts</strong></p>
<p>What should you test for? What’s easy to get wrong? What <em>looks</em> done but isn’t actually wired up?</p>
<p>This is pattern recognition from past failures. The cache that wasn’t being hit. The feature flag that was never checked. The error handler that swallowed exceptions silently.</p>
<p>Concretely:</p>
<ul>
<li>
<p><strong>Test the user flow, not just the units</strong>. My caching module passed all its unit tests. The integration was broken. If I’d tested “make this API call and verify Redis has an entry,” I’d have caught it immediately.</p>
</li>
<li>
<p><strong>Verify your assumptions</strong>. Claude wrote the code, but did the code actually get <em>used</em>? Add a log line. Check the network tab. Confirm reality matches intention.</p>
</li>
<li>
<p><strong>Break it on purpose</strong>. What happens when you pass invalid input? What happens when the database is slow? What happens when the auth token is expired? Claude tests the happy path. You test the sad path.</p>
</li>
</ul>
<p>The underlying skill: developing a checklist of “things that can look done but aren’t” for your specific domain. Every time you get burned, add it to the list. Eventually, you check these instinctively.</p>
<h3 id="the-uncomfortable-conclusion"><strong>The Uncomfortable Conclusion</strong></h3>
<p>A freelance developer with 8 years of experience <a href="https://reddit.com/r/ExperiencedDevs/comments/1mg2r6y/the_era_of_ai_slop_cleanup_has_begun/" rel="external nofollow noopener" class="lnp-link">described a pattern</a> he’s seeing across multiple clients: companies paying good money for internal software that barely works. Same symptoms every time. AI-generated comments. Algorithms that make no sense. Inconsistent patterns.</p>
<p>“Yes it mostly works,” he wrote, “but does so terribly to the point where it needs to be fixed.”</p>
<p>The era of AI slop cleanup has begun. And the people doing the cleanup are the ones who know what production actually looks like.</p>
<p>Claude builds for localhost. You build for production.</p>
<p>That gap is where your fifteen years live. And it’s not getting smaller.</p>
<p><em>This is the third essay in an accidental trilogy. First: <a href="https://lakshminp.substack.com/p/the-invisible-tax-you-pay-when-you" rel="external nofollow noopener" class="lnp-link">comprehension debt is real</a>. Second: <a href="https://lakshminp.substack.com/p/clean-code-is-dead-long-live-clean" rel="external nofollow noopener" class="lnp-link">clean code is dead</a>. This one: the skills that matter more now, not less.</em></p>
]]></content:encoded></item><item><title>I Found a Business Idea and Shipped It in One Claude Code Session</title><link>https://lakshminp.com/2026/01/claude-code-ship-one-session/</link><pubDate>Tue, 06 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/claude-code-ship-one-session/</guid><category>essays</category><category>claude-code</category><category>saas</category><description>I shipped a new product landing page yesterday. Not “finished the design.” Not “pushed to staging.” Live. Accepting waitlist signups. DNS propagated. The whole thing.
Time from “huh, interesting Reddit post” to “supabyoi.com is live”: under an hour.
This is either impressive or terrifying, depending on how you feel about the pace of software development in 2026.
The Setup I’ve been building a Reddit research tool that lives inside Claude Code. It monitors subreddits, scores posts against your interests, and helps you find signals in the noise. The tool uses Reddit’s public JSON endpoints—no API keys required, because Reddit doesn’t issue them anymore.</description><content:encoded><![CDATA[<p>I shipped a new product landing page yesterday. Not “finished the design.” Not “pushed to staging.” Live. Accepting waitlist signups. DNS propagated. The whole thing.</p>
<p>Time from “huh, interesting Reddit post” to “<a href="http://supabyoi.com/" rel="external nofollow noopener" class="lnp-link">supabyoi.com</a> is live”: under an hour.</p>
<p>This is either impressive or terrifying, depending on how you feel about the pace of software development in 2026.</p>
<h1 id="the-setup"><strong>The Setup</strong></h1>
<p>I’ve been building a Reddit research tool that lives inside Claude Code. It monitors subreddits, scores posts against your interests, and helps you find signals in the noise. The tool uses Reddit’s public JSON endpoints—no API keys required, because Reddit doesn’t issue them anymore.</p>
<p>That’s not hyperbole. Reddit recently announced they’re “ending self-service API access.” You can’t just create an app and get keys anymore. You have to submit a request form, explain your use case, and wait for approval. The approval that, according to r/redditdev, never comes. “Tickets rejected, modmail ignored, admin DM ignored.” One developer summed it up: “I don’t believe anyone is getting API access for small personal use at this point.”</p>
<p>They’re pushing everyone to Devvit, their walled-garden platform. JavaScript only. Runs on their servers. Limited to what they allow.</p>
<p>If you want Reddit data for your own tools, you either pay enterprise rates, beg for approval, or use public endpoints like a normal browser would. I chose option three.</p>
<p>GummySearch learned this the hard way. They built a Reddit research tool, hit $11k/month MRR, served 135,000 users. Then Reddit wouldn’t give them a commercial license. They’re <a href="https://gummysearch.com/final-chapter/" rel="external nofollow noopener" class="lnp-link">shutting down by December 2026</a>. The founder didn’t want to operate “looking over your shoulder every day.”</p>
<p>Public endpoints don’t have that problem. Same data. No license to revoke.</p>
<p>I had the tool pointed at r/Supabase, watching for pain points. Standard demand validation stuff.</p>
<p>It surfaced a pattern across multiple threads. <a href="https://www.reddit.com/r/Supabase/comments/1joufox/im_a_massproject_starter_supabase_aint_for_me/" rel="external nofollow noopener" class="lnp-link">“I’m a mass-project starter. Supabase ain’t for me?”</a> (41 upvotes, 27 comments). <a href="https://www.reddit.com/r/Supabase/comments/1i3mduv/is_selfhosting_supabase_worth_it/" rel="external nofollow noopener" class="lnp-link">“Is Self-Hosting Supabase Worth It?”</a> (73 upvotes, 60 comments).</p>
<p>The pain: Supabase’s free tier caps you at 2 projects. Pro tier is $25/month plus $10 per additional project. If you’re an indie dev shipping lots of small bets, you burn through that limit fast.</p>
<p>Supabase is genuinely great for rapid prototyping—I use it myself. The pricing just doesn’t fit the small bets workflow.</p>
<p>The comments were gold:</p>
<blockquote>
<p>“It feels like a bait-and-switch where the upgrade appears to remove project limits, only to hit you with unexpected per-project fees”</p>
<p>“Setting it up properly takes time, maintaining it takes time, keeping the server secure takes time”</p>
</blockquote>
<blockquote>
<p>“The setup process is extensive, unclear and often frustrating”</p>
<p>“Very strange pricing model, which is kind of unacceptable”</p>
</blockquote>
<p><strong>Translation:</strong> Indie devs love Supabase for building fast. They hate the pricing when they ship a lot. They want to self-host but are terrified of maintaining it.</p>
<h1 id="the-evaluation"><strong>The Evaluation</strong></h1>
<p>I have a framework for this. Open source project + operational complexity + permissive license = potential hosting business. I’ve been running variations of this for a while.</p>
<p>I asked Claude—right there in the same session—to run the threads through the framework:</p>
<p><strong>Pain point</strong>: Real. Multi-project pricing punishes prolific shippers.<br>
<strong>License</strong>: Apache 2.0. Clear.<br>
<strong>Operational complexity</strong>: High. Supabase runs ~12 services.<br>
<strong>Existing managed option</strong>: Yes, but that’s the pain source—not the solution.</p>
<p>Then the key insight: I’m not competing with Supabase Cloud on hosting. I’m offering care and feeding for self-hosted instances. Different model entirely.</p>
<ul>
<li>
<p>They bring their own VPS (Hetzner, $10-15/month)</p>
</li>
<li>
<p>I handle upgrades, backups, security</p>
</li>
<li>
<p>Fixed monthly fee: $25. Unlimited instances.</p>
</li>
</ul>
<p>One customer with 5 projects: $25 from me + $15 VPS = $40 total vs $75 on Supabase Cloud.</p>
<h1 id="the-build"><strong>The Build</strong></h1>
<p>Here’s where it gets fast.</p>
<p>I told Claude:</p>
<p>“Create a landing page. Tailwind, not CDN. Minimal. Dev-focused. Static HTML.”</p>
<p>Claude scaffolded the project structure, wrote the copy, set up the build pipeline. I tweaked the value prop and added my ConvertKit form.</p>
<p>Then:</p>
<p>“Push to GitHub, I’ll deploy to Cloudflare Pages.”</p>
<p>Done.</p>
<p>Total time building the landing page: maybe 20 minutes. Most of that was me fiddling with colors.</p>
<h1 id="the-stack"><strong>The Stack</strong></h1>
<p>For the curious:</p>
<p><strong>Landing page</strong>: Static HTML, Tailwind CSS, Cloudflare Pages<br>
<strong>Waitlist</strong>: ConvertKit embed<br>
<strong>Domain</strong>: Namecheap (purchase) → Cloudflare (DNS)<br>
<strong>Total cost so far</strong>: $12 for the domain</p>
<p>The actual product will be FastAPI + Supabase (yes, the irony) + HTMX. SSH into customer VMs. Cron jobs for backups. Simple. I can ship a working beta this week.</p>
<h1 id="the-point"><strong>The Point</strong></h1>
<p>This isn’t about Supabyoi specifically. It’s about the workflow.</p>
<p><strong>Old way</strong>:</p>
<ol>
<li>
<p>Have idea</p>
</li>
<li>
<p>Think about it for weeks</p>
</li>
<li>
<p>Research competitors</p>
</li>
<li>
<p>Write PRD</p>
</li>
<li>
<p>Design mockups</p>
</li>
<li>
<p>Build MVP</p>
</li>
<li>
<p>Realize nobody wants it</p>
</li>
<li>
<p>Total time: 3 months</p>
</li>
</ol>
<p><strong>New way</strong>:</p>
<ol>
<li>
<p>Tool surfaces interesting signal</p>
</li>
<li>
<p>Ask Claude to validate against framework</p>
</li>
<li>
<p>Ask Claude to build landing page</p>
</li>
<li>
<p>Ship</p>
</li>
<li>
<p>See if anyone signs up</p>
</li>
<li>
<p>Total time: 1 hour</p>
</li>
</ol>
<p>The landing page is a hypothesis test. Not a commitment. If I get 50 waitlist signups, I build the thing. If I get 5, I move on. The cost of being wrong is $12 and an hour of my time.</p>
<h1 id="about-that-reddit-tool"><strong>About That Reddit Tool</strong></h1>
<p>I’ve been quietly building this for months. It’s how I found the signal that led to this post.</p>
<p>The key: it lives inside Claude Code. Not a separate app. Not a browser tab. Right there in my terminal, in the same session where I’m writing code and shipping products.</p>
<p>The architecture: Crawler runs locally or on your VPS (Reddit can’t shut you down if they can’t block your IP). Data syncs to a backend. You query it with natural language through Claude. “What are people complaining about in r/Supabase?” → ranked list of pain points with source threads. Then in the same breath: “Evaluate this against my validation framework.” Then: “Build me a landing page.”</p>
<p>One session. Research to shipping.</p>
<p>No Reddit API keys because Reddit killed self-service access. Uses public endpoints. Same data you’d see browsing the site. Your IP, your rate limits, no approval form that never gets answered.</p>
<p>I’m not ready to launch it yet, but if you want early access, DM me on <a href="https://linkedin.com/in/lakshminp" rel="external nofollow noopener" class="lnp-link">LinkedIn</a> or <a href="https://x.com/lakshminp" rel="external nofollow noopener" class="lnp-link">Twitter/X</a>.</p>
<h1 id="the-takeaway"><strong>The Takeaway</strong></h1>
<p>The leverage is real. One person, one AI assistant, one hour, one live product.</p>
<p>The bottleneck isn’t building anymore. It’s finding the right thing to build. That’s why the Reddit tool matters more than the Supabase thing. The tool finds signals. Claude validates them. Claude builds the test. You watch the data.</p>
<p>Small bets at scale.</p>
<p><a href="https://supabyoi.com/" rel="external nofollow noopener" class="lnp-link">supabyoi.com</a> is live. Let’s see what happens.</p>
<p><code>&lt;fingers-crossed/&gt;</code></p>
]]></content:encoded></item><item><title>I Found a Cryptominer in My Client's Production Cluster. Claude Code Found the Attacker.</title><link>https://lakshminp.com/2026/01/cryptominer-production-cluster/</link><pubDate>Sat, 03 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/cryptominer-production-cluster/</guid><category>essays</category><category>kubernetes</category><category>claude-code</category><description>New Year’s Day. Coffee in hand. Ready to ease back into work.
Then I saw the logs.
2026-01-02T06:34:27 GET xmrig-6.24.0-linux-static-x64.tar.gz 2026-01-02T06:34:30 GET http://37.32.6.33:7979/m 2026-01-02T06:34:30 spawn /opt/systemf/m ENOENT xmrig. In production. Someone was mining Monero on my client’s Kubernetes cluster.
The horror.
The Investigation I had a few hundred megabytes of JSON logs and approximately zero patience for manually correlating timestamps. So I did what any reasonable person would do: I asked Claude Code to analyze the logs and figure out what triggered the miner download.</description><content:encoded><![CDATA[<p>New Year’s Day. Coffee in hand. Ready to ease back into work.</p>
<p>Then I saw the logs.</p>
<pre><code>2026-01-02T06:34:27 GET xmrig-6.24.0-linux-static-x64.tar.gz
2026-01-02T06:34:30 GET http://37.32.6.33:7979/m
2026-01-02T06:34:30 spawn /opt/systemf/m ENOENT
</code></pre>
<p>xmrig. In production. Someone was mining Monero on my client’s Kubernetes cluster.</p>
<p>The horror.</p>
<h2 id="the-investigation"><strong>The Investigation</strong></h2>
<p>I had a few hundred megabytes of JSON logs and approximately zero patience for manually correlating timestamps. So I did what any reasonable person would do: I asked Claude Code to analyze the logs and figure out what triggered the miner download.</p>
<p>Within seconds, it built a timeline:</p>
<p><strong>Time Event</strong></p>
<p>06:34:26. Normal request to /onboarding</p>
<p>06:34:27. xmrig downloaded from GitHub</p>
<p>06:34:30. Secondary payload from sketchy IP</p>
<p>06:34:57. Container OOMKilled</p>
<p>The cryptominer was so resource-hungry it consumed 2GB of memory in 30 seconds and crashed the container. Ironic. The attacker’s greed saved us from a prolonged compromise.</p>
<p>But how did they get in?</p>
<h2 id="chasing-red-herrings"><strong>Chasing Red Herrings</strong></h2>
<p>Claude Code’s first suspect: a low-version npm package called <code>device-unique-keygen</code>. Added by a developer whose email matched the package maintainer. Classic supply chain attack pattern.</p>
<p>I got excited. Maybe too excited.</p>
<p>Claude Code fetched the GitHub repo, analyzed the source code, checked for postinstall scripts, looked for obfuscated code, searched for eval() calls.</p>
<p>Nothing. The package was clean. Just a browser fingerprinting library. Boring. Legitimate.</p>
<p>We moved on.</p>
<p>No malicious init containers. No sidecars. No .ashrc shenanigans. The Dockerfile was clean. The pod spec was clean.</p>
<p>Everything was clean except someone was definitely mining crypto on our infrastructure.</p>
<h2 id="the-actual-answer"><strong>The Actual Answer</strong></h2>
<p>Claude Code ran <code>npm audit</code> on the codebase.</p>
<pre><code>critical │ Next.js is vulnerable to RCE in React flight protocol
Package  │ next
Patched  │ &gt;=15.3.6
Your ver │ 15.3.4
CVSS     │ 10.0
</code></pre>
<figure>
<a href="https://substackcdn.com/image/fetch/$s_!FrM8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5bcacd-085e-465e-85f7-03955737439e_1247x765.png" class="image-link image2 is-viewable-img" target="_blank" data-component-name="Image2ToDOM"></a>
<img src="https://substack-post-media.s3.amazonaws.com/public/images/9c5bcacd-085e-465e-85f7-03955737439e_1247x765.png" class="sizing-normal" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9c5bcacd-085e-465e-85f7-03955737439e_1247x765.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:765,&quot;width&quot;:1247,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:128814,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://lakshminp.substack.com/i/183314492?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5bcacd-085e-465e-85f7-03955737439e_1247x765.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" srcset="https://substackcdn.com/image/fetch/$s_!FrM8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5bcacd-085e-465e-85f7-03955737439e_1247x765.png 424w, https://substackcdn.com/image/fetch/$s_!FrM8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5bcacd-085e-465e-85f7-03955737439e_1247x765.png 848w, https://substackcdn.com/image/fetch/$s_!FrM8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5bcacd-085e-465e-85f7-03955737439e_1247x765.png 1272w, https://substackcdn.com/image/fetch/$s_!FrM8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5bcacd-085e-465e-85f7-03955737439e_1247x765.png 1456w" sizes="100vw" loading="lazy" width="1247" height="765" />
<img src="data:image/svg+xml;base64,PHN2ZyByb2xlPSJpbWciIHdpZHRoPSIyMCIgaGVpZ2h0PSIyMCIgdmlld2JveD0iMCAwIDIwIDIwIiBmaWxsPSJub25lIiBzdHJva2Utd2lkdGg9IjEuNSIgc3Ryb2tlPSJ2YXIoLS1jb2xvci1mZy1wcmltYXJ5KSIgc3Ryb2tlLWxpbmVjYXA9InJvdW5kIiBzdHJva2UtbGluZWpvaW49InJvdW5kIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjxnPjx0aXRsZT48L3RpdGxlPjxwYXRoIGQ9Ik0yLjUzMDAxIDcuODE1OTVDMy40OTE3OSA0LjczOTExIDYuNDMyODEgMi41IDkuOTExNzMgMi41QzEzLjE2ODQgMi41IDE1Ljk1MzcgNC40NjIxNCAxNy4wODUyIDcuMjM2ODRMMTcuNjE3OSA4LjY3NjQ3TTE3LjYxNzkgOC42NzY0N0wxOC41MDAyIDQuMjY0NzFNMTcuNjE3OSA4LjY3NjQ3TDEzLjY0NzMgNi45MTE3Nk0xNy40OTk1IDEyLjE4NDFDMTYuNTM3OCAxNS4yNjA5IDEzLjU5NjcgMTcuNSAxMC4xMTc4IDE3LjVDNi44NjExOCAxNy41IDQuMDc1ODkgMTUuNTM3OSAyLjk0NDMyIDEyLjc2MzJMMi40MTE2NSAxMS4zMjM1TTIuNDExNjUgMTEuMzIzNUwxLjUyOTMgMTUuNzM1M00yLjQxMTY1IDExLjMyMzVMNi4zODIyNCAxMy4wODgyIiAvPjwvZz48L3N2Zz4=" />
<img src="data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSIyMCIgaGVpZ2h0PSIyMCIgdmlld2JveD0iMCAwIDI0IDI0IiBmaWxsPSJub25lIiBzdHJva2U9ImN1cnJlbnRDb2xvciIgc3Ryb2tlLXdpZHRoPSIyIiBzdHJva2UtbGluZWNhcD0icm91bmQiIHN0cm9rZS1saW5lam9pbj0icm91bmQiIGNsYXNzPSJsdWNpZGUgbHVjaWRlLW1heGltaXplMiBsdWNpZGUtbWF4aW1pemUtMiI+PHBvbHlsaW5lIHBvaW50cz0iMTUgMyAyMSAzIDIxIDkiPjwvcG9seWxpbmU+PHBvbHlsaW5lIHBvaW50cz0iOSAyMSAzIDIxIDMgMTUiPjwvcG9seWxpbmU+PGxpbmUgeDE9IjIxIiB4Mj0iMTQiIHkxPSIzIiB5Mj0iMTAiPjwvbGluZT48bGluZSB4MT0iMyIgeDI9IjEwIiB5MT0iMjEiIHkyPSIxNCI+PC9saW5lPjwvc3ZnPg==" class="lucide lucide-maximize2 lucide-maximize-2" />
</figure>
<p>CVSS 10. The maximum possible score. The “your house is actively on fire” of security ratings.</p>
<p>The app was running Next.js 15.3.4. A publicly disclosed RCE vulnerability. No authentication required. An attacker could run arbitrary commands on the server by sending a crafted request.</p>
<p>That’s exactly what happened. They sent a request, ran wget twice, downloaded the miner, and started extracting crypto value from compute cycles they weren’t paying for.</p>
<p>The container’s memory limit stopped them. A $20/month Kubernetes resource limit prevented what could have been ongoing theft.</p>
<h2 id="what-claude-code-actually-did"><strong>What Claude Code Actually Did</strong></h2>
<p>I want to be clear about what happened here. I didn’t single-handedly unravel a sophisticated attack. I didn’t manually correlate log timestamps or reverse-engineer obfuscated npm packages.</p>
<p>I said “check these logs” and Claude Code:</p>
<ul>
<li>
<p>Built a timeline from JSON log entries</p>
</li>
<li>
<p>Identified the malware artifacts and C2 server</p>
</li>
<li>
<p>Traced git blame to find who added suspicious packages</p>
</li>
<li>
<p>Fetched and analyzed source code from GitHub</p>
</li>
<li>
<p>Ruled out attack vectors one by one</p>
</li>
<li>
<p>Found the actual vulnerability via npm audit</p>
</li>
<li>
<p>Correlated the OOMKill timing with the attack</p>
</li>
<li>
<p>Suggested remediation and forensic preservation steps</p>
</li>
</ul>
<p>The entire investigation took under an hour. Not because I’m fast. Because Claude Code is.</p>
<h2 id="the-fix"><strong>The Fix</strong></h2>
<pre><code>pnpm update next@^15.3.6
</code></pre>
<p>One command. That’s the remediation for a CVSS 10.0 vulnerability.</p>
<p>We also orphaned the compromised pods for forensic analysis, rotated secrets, and added proper security contexts to prevent future wget adventures.</p>
<figure>
<a href="https://substackcdn.com/image/fetch/$s_!bCvw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F654c9829-43d6-46a6-8405-abfd7b305b80_1344x772.png" class="image-link image2 is-viewable-img" target="_blank" data-component-name="Image2ToDOM"></a>
<img src="https://substack-post-media.s3.amazonaws.com/public/images/654c9829-43d6-46a6-8405-abfd7b305b80_1344x772.png" class="sizing-normal" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/654c9829-43d6-46a6-8405-abfd7b305b80_1344x772.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:772,&quot;width&quot;:1344,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:121268,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://lakshminp.substack.com/i/183314492?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F654c9829-43d6-46a6-8405-abfd7b305b80_1344x772.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" srcset="https://substackcdn.com/image/fetch/$s_!bCvw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F654c9829-43d6-46a6-8405-abfd7b305b80_1344x772.png 424w, https://substackcdn.com/image/fetch/$s_!bCvw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F654c9829-43d6-46a6-8405-abfd7b305b80_1344x772.png 848w, https://substackcdn.com/image/fetch/$s_!bCvw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F654c9829-43d6-46a6-8405-abfd7b305b80_1344x772.png 1272w, https://substackcdn.com/image/fetch/$s_!bCvw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F654c9829-43d6-46a6-8405-abfd7b305b80_1344x772.png 1456w" sizes="100vw" loading="lazy" width="1344" height="772" />
<img src="data:image/svg+xml;base64,PHN2ZyByb2xlPSJpbWciIHdpZHRoPSIyMCIgaGVpZ2h0PSIyMCIgdmlld2JveD0iMCAwIDIwIDIwIiBmaWxsPSJub25lIiBzdHJva2Utd2lkdGg9IjEuNSIgc3Ryb2tlPSJ2YXIoLS1jb2xvci1mZy1wcmltYXJ5KSIgc3Ryb2tlLWxpbmVjYXA9InJvdW5kIiBzdHJva2UtbGluZWpvaW49InJvdW5kIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjxnPjx0aXRsZT48L3RpdGxlPjxwYXRoIGQ9Ik0yLjUzMDAxIDcuODE1OTVDMy40OTE3OSA0LjczOTExIDYuNDMyODEgMi41IDkuOTExNzMgMi41QzEzLjE2ODQgMi41IDE1Ljk1MzcgNC40NjIxNCAxNy4wODUyIDcuMjM2ODRMMTcuNjE3OSA4LjY3NjQ3TTE3LjYxNzkgOC42NzY0N0wxOC41MDAyIDQuMjY0NzFNMTcuNjE3OSA4LjY3NjQ3TDEzLjY0NzMgNi45MTE3Nk0xNy40OTk1IDEyLjE4NDFDMTYuNTM3OCAxNS4yNjA5IDEzLjU5NjcgMTcuNSAxMC4xMTc4IDE3LjVDNi44NjExOCAxNy41IDQuMDc1ODkgMTUuNTM3OSAyLjk0NDMyIDEyLjc2MzJMMi40MTE2NSAxMS4zMjM1TTIuNDExNjUgMTEuMzIzNUwxLjUyOTMgMTUuNzM1M00yLjQxMTY1IDExLjMyMzVMNi4zODIyNCAxMy4wODgyIiAvPjwvZz48L3N2Zz4=" />
<img src="data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSIyMCIgaGVpZ2h0PSIyMCIgdmlld2JveD0iMCAwIDI0IDI0IiBmaWxsPSJub25lIiBzdHJva2U9ImN1cnJlbnRDb2xvciIgc3Ryb2tlLXdpZHRoPSIyIiBzdHJva2UtbGluZWNhcD0icm91bmQiIHN0cm9rZS1saW5lam9pbj0icm91bmQiIGNsYXNzPSJsdWNpZGUgbHVjaWRlLW1heGltaXplMiBsdWNpZGUtbWF4aW1pemUtMiI+PHBvbHlsaW5lIHBvaW50cz0iMTUgMyAyMSAzIDIxIDkiPjwvcG9seWxpbmU+PHBvbHlsaW5lIHBvaW50cz0iOSAyMSAzIDIxIDMgMTUiPjwvcG9seWxpbmU+PGxpbmUgeDE9IjIxIiB4Mj0iMTQiIHkxPSIzIiB5Mj0iMTAiPjwvbGluZT48bGluZSB4MT0iMyIgeDI9IjEwIiB5MT0iMjEiIHkyPSIxNCI+PC9saW5lPjwvc3ZnPg==" class="lucide lucide-maximize2 lucide-maximize-2" />
</figure>
<h2 id="the-lesson"><strong>The Lesson</strong></h2>
<p>Two things saved us:</p>
<ol>
<li>
<p>Centralized logging (couldn’t investigate without the logs)</p>
</li>
<li>
<p>Memory limits (the attacker’s miner killed itself)</p>
</li>
</ol>
<p>One thing would have prevented this entirely: running <code>npm audit</code> before deployment.</p>
<p>The attacker exploited a vulnerability that was publicly disclosed and patched. We just hadn’t updated yet.</p>
<p>Godspeed with your own dependency updates.</p>
<p>My Medium friends can read this <a href="https://medium.com/@lakshminp/i-found-a-cryptominer-in-my-clients-production-cluster-claude-code-found-the-attacker-ae6148ec0514" rel="external nofollow noopener" class="lnp-link">over there</a> as well.</p>
]]></content:encoded></item><item><title>Claude Code Hooks: The Feature You're Ignoring While Babysitting Your AI</title><link>https://lakshminp.com/2026/01/claude-code-hooks/</link><pubDate>Fri, 02 Jan 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/01/claude-code-hooks/</guid><category>essays</category><category>claude-code</category><description>You’re doing it again.
Claude just edited a file. You’re about to type “now run prettier.” For the fourteenth time today. Like some kind of digital hall monitor.
Meanwhile, there’s a feature sitting right there in Claude Code that would do this automatically. It’s called hooks. And based on my extremely scientific survey of Reddit threads, approximately nobody is using them.
What Hooks Actually Are When Claude Code runs, it fires events. Before it uses a tool. After it uses a tool. When it stops. When it sends a notification.</description><content:encoded><![CDATA[<p>You’re doing it again.</p>
<p>Claude just edited a file. You’re about to type “now run prettier.” For the fourteenth time today. Like some kind of digital hall monitor.</p>
<p>Meanwhile, there’s a feature sitting right there in Claude Code that would do this automatically. It’s called hooks. And based on my extremely scientific survey of Reddit threads, approximately nobody is using them.</p>
<h1 id="what-hooks-actually-are"><strong>What Hooks Actually Are</strong></h1>
<p>When Claude Code runs, it fires events. Before it uses a tool. After it uses a tool. When it stops. When it sends a notification.</p>
<p>Hooks let you intercept these events and run shell commands automatically.</p>
<p>That’s it. That’s the whole concept.</p>
<p>Claude edits a file? Run your formatter. Claude finishes a task? Send yourself a Slack message. Claude tries to commit? Run your linter first.</p>
<p>No more babysitting. No more “please remember to run prettier.” You set it once and forget it exists.</p>
<h1 id="the-three-hooks-that-actually-matter"><strong>The Three Hooks That Actually Matter</strong></h1>
<p>I spent way too long reading Reddit threads about hooks. Here’s what power users actually care about:</p>
<h2 id="1-the-formatter-hook"><strong>1. The Formatter Hook</strong></h2>
<p>This is the most common one. Claude edits your code, your formatter runs automatically.</p>
<pre><code>{
  &quot;hooks&quot;: {
    &quot;PostToolUse&quot;: [
      {
        &quot;matcher&quot;: &quot;Write|Edit|MultiEdit&quot;,
        &quot;hooks&quot;: [
          {
            &quot;type&quot;: &quot;command&quot;,
            &quot;command&quot;: &quot;prettier --write $CLAUDE_FILE_PATH&quot;
          }
        ]
      }
    ]
  }
}
</code></pre>
<p>No more “can you run prettier on that?” Revolutionary concept, I know.</p>
<h2 id="2-the-notification-hook"><strong>2. The Notification Hook</strong></h2>
<p>You kicked off a task. You walked away to make coffee. Now you’re checking your terminal every 30 seconds like a nervous parent.</p>
<pre><code>{
  &quot;hooks&quot;: {
    &quot;Stop&quot;: [
      {
        &quot;matcher&quot;: &quot;&quot;,
        &quot;hooks&quot;: [
          {
            &quot;type&quot;: &quot;command&quot;,
            &quot;command&quot;: &quot;curl -d 'Claude is done' ntfy.sh/your-topic&quot;
          }
        ]
      }
    ]
  }
}
</code></pre>
<p>Send it to Slack. Send it to Discord. One person made their Mac speak out loud. Another person made Claude meow. (Allergies. I don’t judge.)</p>
<h2 id="3-the-please-remember-your-instructions-hook"><strong>3. The “Please Remember Your Instructions” Hook</strong></h2>
<p>This one solves a specific pain that will sound familiar: Claude compacts its context to save tokens. In doing so, it sometimes&hellip; forgets things. Important things. Things you put in your <a href="http://claude.md/" rel="external nofollow noopener" class="lnp-link">CLAUDE.md</a>.</p>
<p>The fix? Re-inject your core rules on every prompt:</p>
<pre><code>{
  &quot;hooks&quot;: {
    &quot;PreToolUse&quot;: [
      {
        &quot;matcher&quot;: &quot;UserPromptSubmit&quot;,
        &quot;hooks&quot;: [
          {
            &quot;type&quot;: &quot;command&quot;,
            &quot;command&quot;: &quot;cat .claude/rules.txt&quot;
          }
        ]
      }
    ]
  }
}
</code></pre>
<p>Your rules show up in the context window. Every time. Claude can’t “forget” what’s staring it in the face.</p>
<h1 id="where-this-goes"><strong>Where This Goes</strong></h1>
<p>Put your hooks in <code>.claude/settings.json</code> in your project, or <code>~/.claude/settings.json</code> globally.</p>
<p>Project-level hooks are better. Different projects have different formatters, different rules, different needs. Keep it scoped.</p>
<h1 id="the-one-gotcha"><strong>The One Gotcha</strong></h1>
<p>Someone on Reddit pointed out a real issue: if your formatter changes files, Claude gets a system reminder about those changes. Every. Single. Time.</p>
<p>If you’re formatting aggressively, that’s a lot of noise in your context window. Tokens that could be doing useful work are now just telling Claude that yes, you added a semicolon.</p>
<p>The fix is to be selective. Format on commit, not on every edit. Or accept the tradeoff. Your call.</p>
<h1 id="why-you-should-care"><strong>Why You Should Care</strong></h1>
<p>Hooks turn Claude Code from “AI assistant you have to supervise” into “AI assistant that follows your rules automatically.”</p>
<p>The power users on Reddit are calling this a game-changer. The rest of the users are still typing “please run the linter” by hand.</p>
<p>Don’t be the second group.</p>
<p>Set up three hooks. Formatter, notifications, rule enforcement. Takes ten minutes. Saves hours of typing the same commands.</p>
<p>Your future self will thank you.</p>
<p><em>What’s the most repetitive thing you’re still typing manually in Claude Code?</em></p>
]]></content:encoded></item><item><title>Stop Making Claude Code Guess</title><link>https://lakshminp.com/2025/12/stop-making-claude-code-guess/</link><pubDate>Tue, 30 Dec 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/12/stop-making-claude-code-guess/</guid><category>essays</category><category>claude-code</category><description>This is post 4 of 4 in my Claude Code series. Catch up on The Mental Model, Skills vs Slash Commands, and The Control Freak’s Guide to Agents if you missed them.
We’ve covered how to trigger Claude (slash commands), teach it (skills), and let it explore (agents). Now: how to give it access to the real world.
Because right now, Claude is trapped in a box. It can read your code. It can write your code. But it can’t see what’s on a webpage. It can’t query your database. It can’t check your production logs. It’s a very intelligent entity with no access to external reality.</description><content:encoded><![CDATA[<p><em>This is post 4 of 4 in my Claude Code series. Catch up on <a href="https://lakshminp.substack.com/p/i-spent-weeks-confused-about-claude" rel="external nofollow noopener" class="lnp-link">The Mental Model</a>, <a href="https://lakshminp.substack.com/p/skills-vs-slash-commands-one-works" rel="external nofollow noopener" class="lnp-link">Skills vs Slash Commands</a>, and <a href="https://lakshminp.substack.com/p/the-control-freaks-guide-to-letting" rel="external nofollow noopener" class="lnp-link">The Control Freak’s Guide to Agents</a> if you missed them.</em></p>
<p>We’ve covered how to trigger Claude (slash commands), teach it (skills), and let it explore (agents). Now: how to give it access to the real world.</p>
<p>Because right now, Claude is trapped in a box. It can read your code. It can write your code. But it can’t see what’s on a webpage. It can’t query your database. It can’t check your production logs. It’s a very intelligent entity with no access to external reality.</p>
<p>Ask Claude what your database schema looks like without an MCP server, and you get this:</p>
<p>“Based on typical Postgres schemas, your users table probably has an id, email, and created_at column&hellip;”</p>
<p>Probably. Or we could just <em>query the actual database</em>. Revolutionary concept, I know.</p>
<p>MCP servers are the escape hatch. They give Claude eyes, ears, and hands outside your codebase.</p>
<h1 id="what-mcp-servers-actually-are"><strong>What MCP servers actually are</strong></h1>
<p>MCP (Model Context Protocol) servers expose tools to Claude over a standardized protocol. Each server runs as a separate process and registers its capabilities — functions Claude can call when it needs to interact with something external.</p>
<p><strong>Examples:</strong></p>
<ul>
<li>
<p><strong>Playwright</strong>: browse the web, scrape pages, automate testing</p>
</li>
<li>
<p><strong>Supabase/Postgres</strong>: query databases directly</p>
</li>
<li>
<p><strong>Betterstack</strong>: pull production logs</p>
</li>
<li>
<p><strong>Filesystem</strong>: access files outside your project</p>
</li>
</ul>
<p>The “server” terminology is slightly misleading — they’re really tool providers that Claude can invoke. You configure them in <code>.mcp.json</code>, Claude discovers their capabilities, and suddenly it can do things it couldn’t do before.</p>
<h1 id="the-critical-distinction-capabilities-vs-instructions"><strong>The critical distinction: capabilities vs instructions</strong></h1>
<p>This is where people get confused between MCP servers and skills. Let me make it painfully clear:</p>
<p><strong>MCP servers = capabilities.</strong> They let Claude DO something it couldn’t do before. Playwright gives Claude the ability to browse. A Postgres MCP gives Claude the ability to query databases.</p>
<p><strong>Skills = instructions.</strong> They tell Claude HOW to do something. A “web-scraping” skill might teach Claude efficient patterns for using Playwright.</p>
<p>You often want both. The MCP server provides the capability. The skill provides the expertise.</p>
<p>Example: I have the Supabase MCP installed (capability). I also have a skill that says “when querying user data, always filter by organization_id first for performance” (instruction). The skill makes the capability more useful, but without the capability, the skill is just a nice idea with no way to execute.</p>
<p>Skills can tell Claude what to do. They can’t give Claude new abilities. Good luck querying your production database using only skills.</p>
<h1 id="the-token-cost-nobody-mentions-in-the-marketing-materials"><strong>The token cost nobody mentions in the marketing materials</strong></h1>
<p>Here’s the thing that caught me off guard: <strong>MCP server tool definitions are always in your context window.</strong></p>
<p>Every MCP server you install adds its tool signatures to every conversation. Even when you’re not using it. Even when you’re doing something completely unrelated. The tool definitions are just sitting there, eating tokens.</p>
<p>Install 10 MCP servers because they seemed cool? All 10 are eating tokens in every session. Like subscription services you forgot you signed up for, except it’s your API bill.</p>
<p>This matters for:</p>
<ul>
<li>
<p>Context window limits (you have less room for actual work)</p>
</li>
<li>
<p>Cost (more tokens = more money, math is cruel)</p>
</li>
<li>
<p>Performance (sometimes, more context = slower responses)</p>
</li>
</ul>
<p>Be intentional. Don’t install MCP servers “just in case.” Install what you actually use. Uninstall what you don’t. Your token budget will thank you.</p>
<h1 id="the-patterns-that-waste-everyones-time"><strong>The patterns that waste everyone’s time</strong></h1>
<p>Watch any Claude Code session where someone doesn’t have the right MCP servers installed:</p>
<ul>
<li>
<p>Asking Claude to guess what a webpage looks like (when Playwright could just open it)</p>
</li>
<li>
<p>Copy-pasting API responses into the chat (human middleware)</p>
</li>
<li>
<p>Describing database schemas in words (when Claude could just query them)</p>
</li>
</ul>
<p>MCP servers aren’t optional extras. They’re not power-user features for people who want to show off. They’re how Claude becomes actually useful for real-world tasks instead of just being a very eloquent guesser.</p>
<p>If you’re regularly copy-pasting external data into Claude, you need an MCP server. Stop being the bottleneck.</p>
<h1 id="my-current-mcp-stack-minimal-intentional"><strong>My current MCP stack (minimal, intentional)</strong></h1>
<p>For my SaaS projects, I typically configure:</p>
<ul>
<li>
<p><strong>Playwright</strong>: browser automation, web scraping, testing flows</p>
</li>
<li>
<p><strong>Supabase (read-only)</strong>: querying my production database without leaving Claude</p>
</li>
<li>
<p><strong>Betterstack</strong>: pulling logs when debugging production issues</p>
</li>
</ul>
<p>Key point: these are <strong>project-specific</strong>, configured in <code>.mcp.json</code> in the repo. Not global. When I’m working on this blog, I don’t need Supabase or Betterstack — so they’re not loaded, and I’m not paying the token cost.</p>
<p>This is the lever most people miss. You can have different MCP stacks for different projects. A SaaS project needs database and logs. A content project might just need Playwright for research. Configure per-project, pay only for what that project actually needs.</p>
<p>I’ve tried others and removed them. The token cost wasn’t worth it for occasional use. That fancy Notion MCP? Gone. The GitHub MCP for repo scaffolding? Turns out <code>gh</code> CLI works fine and doesn’t eat context.</p>
<p>Minimal viable MCP stack. Everything you need, nothing you don’t.</p>
<h1 id="when-to-add-an-mcp-server-and-when-not-to"><strong>When to add an MCP server (and when not to)</strong></h1>
<p><strong>Add an MCP server when:</strong></p>
<ul>
<li>
<p>You’re regularly copy-pasting external data into Claude</p>
</li>
<li>
<p>You’re asking Claude to guess at things it could just look up</p>
</li>
<li>
<p>The task requires interacting with systems outside your codebase</p>
</li>
<li>
<p>The capability would be used frequently enough to justify the token cost</p>
</li>
</ul>
<p><strong>Don’t add an MCP server when:</strong></p>
<ul>
<li>
<p>You might use it “someday” (you won’t, and it’ll eat tokens until you remember to remove it)</p>
</li>
<li>
<p>You’re just curious what it does (read the docs instead)</p>
</li>
<li>
<p>Another tool already covers the capability</p>
</li>
<li>
<p>A CLI tool would be faster and cheaper (often true)</p>
</li>
</ul>
<h1 id="the-full-picture-mcp-servers--skills--agents"><strong>The full picture: MCP servers + skills + agents</strong></h1>
<p>Here’s how they compose in practice. Say I’m debugging why users are seeing 500 errors:</p>
<ol>
<li>
<p><strong>Betterstack MCP</strong> pulls the error logs from the last hour</p>
</li>
<li>
<p><strong>Supabase MCP</strong> queries the affected user records</p>
</li>
<li>
<p><strong>Agent</strong> correlates the data — finds that all failing requests share a malformed organization_id</p>
</li>
<li>
<p><strong>Skill</strong> reminds Claude to check the migration history when schema issues appear</p>
</li>
<li>
<p><strong>Slash command</strong> triggered this whole investigation with <code>/debug-500s</code></p>
</li>
</ol>
<p>Each layer does its job. The MCP servers are the foundation — without the capability to actually access logs and data, everything else is just documentation for things you can’t do.</p>
<p>The pattern: MCP servers give Claude access to production reality. Skills teach Claude how to navigate that reality efficiently. Agents do the investigation autonomously. You get answers instead of guesses.</p>
<p><em>(There’s a deeper rabbit hole here about balancing MCP capabilities against context window costs — which MCPs to load when, how to structure project-specific configs, when to use CLI tools instead. That’s a future post.)</em></p>
<p><strong>This completes the 4-part series:</strong></p>
<ol>
<li>
<p><a href="https://lakshminp.substack.com/p/i-spent-weeks-confused-about-claude" rel="external nofollow noopener" class="lnp-link">The Mental Model</a> (slash commands, skills, agents, MCP servers, plugins)</p>
</li>
<li>
<p><a href="https://lakshminp.substack.com/p/skills-vs-slash-commands-one-works" rel="external nofollow noopener" class="lnp-link">Skills vs Slash Commands</a> (one works, one “works”)</p>
</li>
<li>
<p><a href="https://lakshminp.substack.com/p/the-control-freaks-guide-to-letting" rel="external nofollow noopener" class="lnp-link">Agents</a> (stop scripting exploration)</p>
</li>
<li>
<p>MCP servers (stop making Claude guess)</p>
</li>
</ol>
<p>I help technical founders develop, deploy, and market their SaaS using Claude Code. If this series saved you some confusion, that’s what I’m here for.</p>
<p><strong>I’m curious:</strong></p>
<ul>
<li>
<p>What’s your MCP stack? Which servers do you actually use daily vs installed “just in case”?</p>
</li>
<li>
<p>Have you noticed the token cost creep from too many MCPs?</p>
</li>
<li>
<p>Any MCP servers you’d recommend that I haven’t mentioned?</p>
</li>
</ul>
<p>Reply or comment — I read everything.</p>
]]></content:encoded></item><item><title>How to Let Claude Code Explore Without Losing Control</title><link>https://lakshminp.com/2025/12/claude-code-agent-exploration/</link><pubDate>Mon, 29 Dec 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/12/claude-code-agent-exploration/</guid><category>essays</category><category>claude-code</category><description>This is post 3 of 4 in my Claude Code series. Catch up on The Mental Model and Skills vs Slash Commands if you missed them.
My first instinct with Claude Code was to script everything. Every task got a slash command. Every command tried to handle every edge case. Fifty lines of instructions. Nested conditionals. Error handling for scenarios that would never happen. It was beautiful. It was comprehensive. It was also completely fragile and broke constantly.</description><content:encoded><![CDATA[<p><em>This is post 3 of 4 in my Claude Code series. Catch up on <a href="https://lakshminp.substack.com/p/i-spent-weeks-confused-about-claude" rel="external nofollow noopener" class="lnp-link">The Mental Model</a> and <a href="https://lakshminp.substack.com/p/skills-vs-slash-commands-one-works" rel="external nofollow noopener" class="lnp-link">Skills vs Slash Commands</a> if you missed them.</em></p>
<p>My first instinct with Claude Code was to script everything. Every task got a slash command. Every command tried to handle every edge case. Fifty lines of instructions. Nested conditionals. Error handling for scenarios that would never happen. It was beautiful. It was comprehensive. It was also completely fragile and broke constantly.</p>
<p>I was basically writing enterprise Java in prompt form. Nobody should have to live like that.</p>
<p>Then I discovered agents. Now my slash commands are thin — they set up context and spawn an agent to do the actual work. The agent figures out the details. My blood pressure dropped significantly.</p>
<p><em>(Quick terminology note: Claude Code calls these “subagents” because they run as subprocesses of your main session. Everyone just says “agents.” I’ll use “agent” throughout because life is short.)</em></p>
<h1 id="the-mistake-treating-claude-like-a-very-slow-bash-script"><strong>The mistake: treating Claude like a very slow bash script</strong></h1>
<p>I had a <code>/research-topic</code> command that tried to:</p>
<ul>
<li>
<p>Search the codebase for relevant files</p>
</li>
<li>
<p>Read documentation</p>
</li>
<li>
<p>Check external sources</p>
</li>
<li>
<p>Synthesize findings</p>
</li>
<li>
<p>Format the output in exactly the right way</p>
</li>
<li>
<p>Handle seventeen different edge cases I imagined at 2 AM</p>
</li>
</ul>
<p>It was a mess. Too many conditionals. Too many assumptions about what files exist and where. It broke every time the codebase changed, which was constantly, because that’s what codebases do.</p>
<p>The fundamental problem: I was trying to script exploration. But exploration is inherently unpredictable. You don’t know what you’ll find until you look. That’s literally the definition of exploration. Trying to pre-script it is like writing detailed driving directions before knowing where you’re going.</p>
<p>“Turn left at the third building. Unless there isn’t a third building. Then turn right. Unless you’re already facing north. Check the position of the sun&hellip;”</p>
<p>This is madness. Don’t do this to yourself.</p>
<h1 id="what-agents-actually-do-and-why-its-better"><strong>What agents actually do (and why it’s better)</strong></h1>
<p>Agents have their own context. They’re decoupled from your main conversation. They go off, do their thing, and come back with answers.</p>
<p>You describe the goal. The agent figures out the path.</p>
<p>“How does authentication work in this codebase?” I have no idea which files to look at. I could spend ten minutes finding them myself. Or I could let Claude spawn an Explore agent — one that’s <em>already built into Claude Code</em> — and get an answer in two minutes.</p>
<p>Here’s what that actually looks like:</p>
<pre><code>Me: &quot;How does authentication work in this codebase?&quot;

Claude spawns Explore agent → searches → reads 8 files → returns:

&quot;Authentication uses JWT tokens stored in httpOnly cookies.
The flow: login request → /api/auth/login → validates credentials
→ creates token → sets cookie. Token validation happens in
middleware at src/middleware/auth.ts:42. Refresh logic is in
src/lib/auth/refresh.ts.&quot;
</code></pre>
<p>The agent checked 8 files. Maybe it checked 15. I don’t know, and I don’t care. I care about the answer.</p>
<p>This is fundamentally different from scripting. You’re not specifying steps. You’re specifying outcomes. It’s like hiring a contractor versus writing assembly instructions. One of these is dramatically less exhausting.</p>
<h1 id="whats-inside-an-agent-dissecting-explore"><strong>What’s inside an agent (dissecting Explore)</strong></h1>
<p>Claude Code ships with a built-in Explore agent. Let’s look under the hood:</p>
<p><strong>AspectWhat Explore UsesModel</strong>Haiku (fast, cheap)<strong>Tools</strong>Glob, Grep, Read, limited Bash (ls, git log, etc.)<strong>Can modify files?<strong>No — strictly read-only</strong>Context</strong>Isolated from your main conversation</p>
<p>The Explore agent has three thoroughness levels:</p>
<ul>
<li>
<p><strong>Quick</strong> — targeted lookup, minimal file traversal. Use when you know roughly what you’re looking for.</p>
</li>
<li>
<p><strong>Medium</strong> — searches multiple related locations. The default for most queries.</p>
</li>
<li>
<p><strong>Very thorough</strong> — comprehensive search across unusual places and naming conventions. Slower, but finds things hidden in unexpected corners.</p>
</li>
</ul>
<p>Here’s the part nobody emphasizes enough: <strong>the agent runs in its own context window.</strong> It can read 20 files, search through thousands of lines, and when it’s done, you get a summary. Your main conversation stays clean. Your token budget stays intact.</p>
<p>This is huge. Before I understood this, I was reading files directly in my main session and watching my context fill up with code I’d already reviewed. Now the agent does the heavy reading. I get the conclusions.</p>
<h1 id="when-to-use-agents-a-guide-for-control-freaks-learning-to-let-go"><strong>When to use agents (a guide for control freaks learning to let go)</strong></h1>
<p><strong>Use agents when the task requires exploration:</strong></p>
<ul>
<li>
<p>“Find all API endpoints and document them”</p>
</li>
<li>
<p>“Investigate why this test is flaky”</p>
</li>
<li>
<p>“Research how other projects handle rate limiting”</p>
</li>
<li>
<p>“What files would I need to change to add feature X?”</p>
</li>
</ul>
<p><strong>Use agents when you’d be guessing at the steps:</strong></p>
<p>If you find yourself writing a slash command with lots of “if this exists, then&hellip;” logic, stop. Step away from the keyboard. That’s agent territory. You’re trying to pre-solve a problem you don’t understand yet.</p>
<p><strong>Use agents for heavy reading:</strong></p>
<p>Agents have their own context window. They can read 20 files without bloating your main session. When they’re done, they return a summary. Your conversation stays clean.</p>
<h1 id="creating-your-own-agents"><strong>Creating your own agents</strong></h1>
<p>You don’t have to use the built-in agents. You can create your own.</p>
<p>An agent is just a skill (markdown file in <code>.claude/commands/</code>) that spawns a subagent using the Task tool. Here’s a simple one:</p>
<pre><code># /research-topic

Research this topic in the codebase: $ARGUMENTS

Use the Task tool with subagent_type=&quot;Explore&quot; to:
- Search for relevant files and patterns
- Read key implementation files
- Understand how it currently works

Return a structured summary:
- Key files involved (with line numbers)
- How the feature currently works
- Any gaps, issues, or technical debt found
</code></pre>
<p>That’s it. No fifty lines of conditionals. No edge case handling. The Explore agent figures out what to search, what to read, how deep to go. You describe the outcome. The agent handles the process.</p>
<p>The pattern: <strong>thin skill that spawns an agent.</strong> The skill is just a trigger with context. The agent does the thinking.</p>
<h1 id="agents-vs-slash-commands-the-actual-difference"><strong>Agents vs slash commands (the actual difference)</strong></h1>
<p><strong>Slash commands / Skills:</strong></p>
<ul>
<li>
<p>Fixed steps</p>
</li>
<li>
<p>Predictable output</p>
</li>
<li>
<p>You know exactly what will happen</p>
</li>
<li>
<p>Good for: formatting, templating, repetitive tasks with known structure</p>
</li>
</ul>
<p><strong>Agents:</strong></p>
<ul>
<li>
<p>Dynamic exploration</p>
</li>
<li>
<p>Variable output</p>
</li>
<li>
<p>Figures out the path autonomously</p>
</li>
<li>
<p>Good for: research, investigation, multi-file analysis, anything where you don’t know what you’ll find</p>
</li>
</ul>
<p>Here’s the mental model: slash commands are for when you know the answer and just need to execute it. Agents are for when you don’t know the answer and need someone to go find it.</p>
<p>If you’re using slash commands for exploration, you’re making your life unnecessarily difficult. I did this for months. Learn from my suffering.</p>
<h1 id="the-real-unlock-its-embarrassingly-simple"><strong>The real unlock (it’s embarrassingly simple)</strong></h1>
<p>Once I stopped trying to script everything, my workflow got simpler. Not more complex. Simpler.</p>
<p>Skills handle the predictable stuff: commit messages, content formatting, code reviews with a checklist.</p>
<p>Agents handle the unpredictable stuff: understanding codebases, investigating issues, researching approaches.</p>
<p>I stopped fighting the tool. I let agents explore. I let skills execute. Everything got easier. I felt slightly foolish for not figuring this out sooner, but that’s the tax you pay for learning things the hard way.</p>
<p><strong>Next up:</strong> MCP servers — the things that let Claude actually interact with the real world instead of just imagining what websites might contain.</p>
<p>I help technical founders develop, deploy, and market their SaaS using Claude Code. These lessons came from months of doing it wrong first.</p>
]]></content:encoded></item><item><title>Claude Code Skills vs Slash Commands: Which One to Use</title><link>https://lakshminp.com/2025/12/claude-code-skills-vs-commands/</link><pubDate>Wed, 17 Dec 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/12/claude-code-skills-vs-commands/</guid><category>essays</category><category>claude-code</category><description>This is the most confusing distinction in Claude Code — and it only gets muddier now that skills and slash commands share the stage with subagents and plugins. Both are “reusable prompts.” Both save you from typing the same thing repeatedly. Both are shareable. Both sound like they do roughly the same job. The documentation makes them seem interchangeable, like choosing between Pepsi and Coke.
They are not interchangeable. Three differences matter:</description><content:encoded><![CDATA[<p>This is <a href="/2025/12/claude-code-mental-model/" class="lnp-link">the most confusing distinction in Claude Code</a> — and it only gets muddier now that skills and slash commands share the stage with subagents and plugins. Both are “reusable prompts.” Both save you from typing the same thing repeatedly. Both are shareable. Both sound like they do roughly the same job. The documentation makes them seem interchangeable, like choosing between Pepsi and Coke.</p>
<p>They are not interchangeable. Three differences matter:</p>
<p>1. <strong>Who pulls the trigger:</strong> Slash commands run when YOU invoke them. Skills run when CLAUDE decides they’re relevant.</p>
<p>2. <strong>Who provides context:</strong> Slash commands take arguments — /draft-linkedin-post [topic] [bullets] tells Claude exactly what to write about. Skills infer context from the conversation and codebase.</p>
<p>3. <strong>Token cost:</strong> Slash commands get inserted into context every time you invoke them. Skills only load their description until activated — the full content lazy-loads.</p>
<p>Control and clarity vs automation and efficiency. Pick your trade-off.</p>
<h1 id="slash-commands-youre-in-control-finally">Slash commands: you’re in control (finally)</h1>
<p>Slash commands run when you type them. Period. Full stop. No ambiguity. No hoping. No praying to the AI gods.</p>
<p>You type /draft-linkedin-post, it runs. You type /commit, it commits. Revolutionary concept, I know. In an age of magical AI that’s supposed to read your mind, there’s something deeply satisfying about a tool that just does what you tell it when you tell it.</p>
<p>I have /draft-linkedin-post that takes a topic and bullets, then outputs a post in my voice. Same structure every time. I don’t want Claude to get creative with the format — I want consistency. I want to type a command and get a predictable result. Like some kind of digital caveman using primitive trigger-response technology.</p>
<p>Other examples:</p>
<p>- /commit — standard commit message format</p>
<p>- /review — code review with my preferred checklist</p>
<p>- /atomize-content — break an essay into platform-specific posts</p>
<p>The key: if you’re giving the same instructions more than twice, make it a slash command. Your fingers will thank you. Your sanity will thank you. Your token budget might even thank you.</p>
<h1 id="skills-claude-decides-when-it-feels-like-it">Skills: Claude decides (when it feels like it)</h1>
<p>Skills are instructions that Claude is supposed to invoke automatically when relevant.</p>
<p>The theory — and I’m using “theory” here the same way physicists use it when they’re 90% sure but can’t prove it — is that you write a skill with a nice description, Claude reads it, and magically invokes it when the context matches.</p>
<p>The reality: skills don’t always fire automatically.</p>
<p>Sometimes Claude picks them up. Sometimes it doesn’t. Sometimes you have to explicitly say “use the X skill” anyway, which kind of defeats the entire purpose of having an “automatic” system. It’s like having a self-driving car that occasionally requires you to grab the wheel and steer. Very reassuring.</p>
<p>The dirty secret nobody wants to admit</p>
<p>Go read r/ClaudeCode threads about skills. I’ll wait. Actually, I’ll save you the trip:</p>
<p>“It hardly picks up any skill without actually telling it to use it”</p>
<p>“Skills would be awesome if it actually used them properly”</p>
<p>“Skills are basically just reminders to the LLM”</p>
<p>That last one is painfully accurate. Skills are fancy prompts that Claude may or may not remember exist. They’re like Post-It notes you stick on your monitor hoping your future self will notice them. Sometimes you do. Sometimes you walk past them for three weeks wondering why that yellow blob is in your peripheral vision.</p>
<p>Someone actually tested this.</p>
<p><a href="https://scottspence.com/posts/how-to-make-claude-code-skills-activate-reliably" rel="external nofollow noopener" class="lnp-link">Scott Spence ran</a> 200+ tests on skill activation reliably:</p>
<p>The results:</p>
<p>- Simple instruction hook: 20% activation (coin flip)</p>
<p>- Forced eval hook: 84% activation</p>
<p>The difference? Commitment mechanisms. Instead of passively hoping Claude notices skills exist, the forced eval hook makes Claude explicitly evaluate EACH skill with YES/NO reasoning before proceeding.</p>
<p>Once Claude writes “YES — need this skill,” it’s committed. It’s harder to bypass something you just agreed to use.</p>
<p>The takeaway: skills can work reliably — but not out of the box. You need <a href="/2026/01/claude-code-hooks/" class="lnp-link">hooks</a> to force the evaluation. Which kind of defeats the “automatic” promise.</p>
<h2 id="so-why-do-skills-exist-there-are-actual-reasons">So why do skills exist? (There are actual reasons)</h2>
<p>Three legitimate use cases:</p>
<p>1. Lazy loading saves tokens.</p>
<p>Your CLAUDE.md is always in context. Always. Every conversation. Every token. Skills, on the other hand, are loaded on demand — at least in theory. If you have a <a href="/2026/04/claude-md-best-practices/" class="lnp-link">500-line CLAUDE.md</a> because you kept adding “just one more instruction,” consider breaking domain-specific stuff into skills.</p>
<p>Your wallet will appreciate this eventually.</p>
<p>2. Organization for the obsessive-compulsive among us.</p>
<p>Instead of one massive CLAUDE.md that reads like a legal document, you can have modular skills: one for frontend patterns, one for API conventions, one for content writing. It’s cleaner to maintain. It makes you feel like you have your life together. Whether Claude actually uses them correctly is a separate question.</p>
<p>3. Sharing your neuroses with teammates.</p>
<p>You can package skills and share them with your team. “Here’s how I want code reviews done” becomes a shareable artifact instead of a 47-message Slack thread that nobody will ever read.</p>
<p>When to use which (the practical guide for people with deadlines)</p>
<h3 id="use-slash-commands-when">Use slash commands when:</h3>
<p>- You want to control exactly when it runs</p>
<p>- The task has fixed steps</p>
<p>- Consistency matters more than flexibility</p>
<p>- You don’t trust Claude to figure out when to apply it (wise)</p>
<h3 id="use-skills-when">Use skills when:</h3>
<p>- Instructions only apply sometimes (not every conversation)</p>
<p>- You want to reduce CLAUDE.md bloat</p>
<p>- You’re okay with Claude deciding relevance (optimistic)</p>
<p>- You want to share patterns across projects/teams</p>
<p>- You’re mentally prepared to invoke them explicitly anyway</p>
<h1 id="my-approach-learned-through-pain">My approach (learned through pain)</h1>
<p>I default to slash commands. I want control. I’ve been burned too many times by “intelligent” systems that aren’t quite intelligent enough.</p>
<p>I use skills for domain-specific instructions that would bloat my CLAUDE.md unnecessarily. My post summarizer skill doesn’t need to be in context when I’m debugging Python. My code review preferences don’t need to load when I’m drafting blog posts.</p>
<p>When I create a skill, I mentally prepare to invoke it explicitly. If Claude picks it up automatically, great — I’ll take that win. If not, I type “use the X skill” and move on with my life. No frustration. No existential crisis. Just pragmatic acceptance that the future isn’t quite here yet.</p>
<p>This is post 2 of 4. Next up: <a href="/2025/12/claude-code-agent-exploration/" class="lnp-link">agents</a> — and when to stop trying to script everything like it’s 2015.</p>
]]></content:encoded></item><item><title>I Spent Weeks Confused About Claude Code's 5 Concepts. Here's the Mental Model That Finally Clicked.</title><link>https://lakshminp.com/2025/12/claude-code-mental-model/</link><pubDate>Thu, 11 Dec 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/12/claude-code-mental-model/</guid><category>essays</category><category>claude-code</category><description>Slash commands, skills, agents, MCP servers, plugins. Five concepts. Five different jobs. One very confused developer (me) trying to figure out which one to use when.
You’re using Claude Code. You’ve seen these terms thrown around like confetti at a developer conference. Slash commands! Skills! Agents! MCP servers! Plugins! Each one sounds important. Each one sounds slightly different from the others. Each one makes you wonder if you’re using Claude Code wrong.</description><content:encoded><![CDATA[<p>Slash commands, skills, agents, MCP servers, plugins. Five concepts. Five different jobs. One very confused developer (me) trying to figure out which one to use when.</p>
<p>You&rsquo;re using Claude Code. You&rsquo;ve seen these terms thrown around like confetti at a developer conference. Slash commands! Skills! Agents! MCP servers! Plugins! Each one sounds important. Each one sounds slightly different from the others. Each one makes you wonder if you&rsquo;re using Claude Code wrong.</p>
<p>Spoiler: you probably are. But so is everyone else, so don&rsquo;t feel too special about it.</p>
<p>Here&rsquo;s the mental model that finally made sense to me after weeks of confusion and several existential crises about whether I understood my own tools.</p>
<h2 id="the-one-sentence-version">The one-sentence version</h2>
<p><strong>Slash commands</strong> = shortcuts you trigger manually (like a civilized person)</p>
<p><strong>Skills</strong> = instructions Claude <em>might</em> trigger automatically (emphasis on &ldquo;might&rdquo;)</p>
<p><strong>Agents</strong> = autonomous workers with their own context (little worker bees you send off to do your bidding)</p>
<p><strong>MCP servers</strong> = external capabilities like browsers, databases, APIs (the things that let Claude actually <em>do</em> stuff in the real world)</p>
<p><strong>Plugins</strong> = packaging that bundles any combination of the above (a zip file for your AI workflow, basically)</p>
<p>That&rsquo;s it. Five concepts. Five different jobs. The confusion happens because they overlap like a Venn diagram designed by someone who hates clarity. A slash command can spawn an agent. An agent can use MCP servers. A plugin can contain all of the above. It&rsquo;s turtles all the way down.</p>
<p>(I&rsquo;m not covering <strong>hooks</strong> here — automations that fire on events like file saves. That&rsquo;s a whole other therapy session.)</p>
<h2 id="the-key-distinction-most-people-miss">The key distinction most people miss</h2>
<p>Here&rsquo;s what nobody tells you upfront: <strong>who decides when something runs?</strong></p>
<ul>
<li>Slash commands: <strong>You</strong> trigger them. Like pressing a button. Revolutionary concept.</li>
<li>Skills: <strong>Claude</strong> triggers them. In theory. When it feels like it. Maybe.</li>
<li>Agents: <strong>You</strong> spawn them, then they run autonomously until they&rsquo;re done or your tokens are.</li>
<li>MCP servers: <strong>Claude</strong> calls them when it needs to reach outside your codebase.</li>
<li>Plugins: <strong>You</strong> install them. They&rsquo;re just containers.</li>
</ul>
<p>This matters more than all the technical mumbo-jumbo. Want control? Slash commands. Want Claude to figure it out? Skills. Want to let something loose and hope for the best? Agents. Want Claude to actually interact with the real world? MCP servers.</p>
<h2 id="the-decision-tree-for-people-who-dont-want-to-think-about-this-anymore">The decision tree (for people who don&rsquo;t want to think about this anymore)</h2>
<p>When you&rsquo;re about to ask Claude to do something, run through this:</p>
<p><strong>Is it a repeatable task with fixed steps?</strong> → Slash command. Done. Move on with your life.</p>
<p><strong>Does it need to access external systems?</strong> → MCP server. Claude can&rsquo;t browse the web or query databases with pure thought. Yet.</p>
<p><strong>Does it require exploration and figuring things out?</strong> → Agent. Let it wander. It&rsquo;s smarter than you think. Sometimes.</p>
<p><strong>Is it domain-specific instructions that don&rsquo;t always apply?</strong> → Skill. Good luck getting Claude to actually use it without being asked.</p>
<p><strong>Do you want to share your setup with others?</strong> → Package it as a plugin. Make it someone else&rsquo;s problem.</p>
<h2 id="the-overlap-problem-or-why-everyone-builds-three-things-for-the-same-task">The overlap problem (or: why everyone builds three things for the same task)</h2>
<p>Here&rsquo;s what happens in the wild: developers build a slash command, a skill, AND an agent for the same task. I&rsquo;ve done it. You&rsquo;ve probably done it. We&rsquo;ve all sinned.</p>
<p>Pick one primary approach:</p>
<ul>
<li>Need <strong>control over when it runs</strong> → slash command</li>
<li>Need <strong>Claude to decide when it&rsquo;s relevant</strong> → skill (and a prayer)</li>
<li>Need <strong>autonomous multi-step execution</strong> → agent</li>
<li>Need <strong>external system access</strong> → MCP server</li>
<li>Need <strong>to share your setup</strong> → plugin</li>
</ul>
<p>Use the others to support, not duplicate. Your future self will thank you when you&rsquo;re not debugging three different implementations of the same thing at 2 AM.</p>
<h2 id="how-they-layer-the-actually-useful-part">How they layer (the actually useful part)</h2>
<p>Think of it as a stack:</p>
<ul>
<li><strong>Skills</strong> = instructions (how to do things)</li>
<li><strong>Slash commands</strong> = triggers (entry points you control)</li>
<li><strong>Agents</strong> = workers (autonomous task executors)</li>
<li><strong>MCP servers</strong> = capabilities (external system access)</li>
<li><strong>Plugins</strong> = packaging (bundles of all the above)</li>
</ul>
<p>A slash command can spawn an agent. An agent can use MCP servers. A skill can teach Claude how to use an MCP server efficiently. A plugin can package all of this into something you can share on GitHub and pretend makes you a thought leader.</p>
<p>They compose. They don&rsquo;t compete. Unless you make them compete, in which case, godspeed.</p>
<p>This is post 1 of 4. Next up: slash commands vs skills — and why skills don&rsquo;t work the way the documentation promises they will.</p>
<p>I help technical founders develop, deploy, and market their SaaS using Claude Code. This is the kind of workflow clarity I wish someone had given me three months ago.</p>
]]></content:encoded></item><item><title>Why Your AI Wakes Up Every Morning With No Memory (And how to fix it)</title><link>https://lakshminp.com/2025/11/ai-agent-memory-persistence/</link><pubDate>Tue, 11 Nov 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/11/ai-agent-memory-persistence/</guid><category>essays</category><category>claude-code</category><description>I was two weeks into a gnarly refactor when it happened.
Claude and I had been pair programming on an authentication system—tracking down race conditions, filing away “fix this later” issues, building up this rich context about why we made certain decisions. RS256 instead of HS256 for key rotation. Session middleware patterns. The whole architecture was in our shared understanding.
Then I hit compaction.
I came back the next day, opened a new Claude session, and asked: “Where did we leave off?”</description><content:encoded><![CDATA[<p>I was two weeks into a gnarly refactor when it happened.</p>
<p>Claude and I had been pair programming on an authentication system—tracking down race conditions, filing away “fix this later” issues, building up this rich context about why we made certain decisions. RS256 instead of HS256 for key rotation. Session middleware patterns. The whole architecture was in our shared understanding.</p>
<p>Then I hit compaction.</p>
<p>I came back the next day, opened a new Claude session, and asked: “Where did we leave off?”</p>
<p>Claude: “I don’t have information about previous sessions in my context.”</p>
<p><strong>All of it. Gone.</strong></p>
<p>The discovered bugs. The architectural decisions. The “by the way, we should fix this” notes. Everything we’d built up over dozens of hours—evaporated.</p>
<p>I spent 30 minutes re-explaining what we’d been working on. And even then, I couldn’t remember all the issues Claude had surfaced. How many edge cases had we found? Which ones were critical? What was blocking what?</p>
<p>This is what I call the <strong>amnesia problem</strong>. And it’s not just annoying—it’s a fundamental limitation of how we work with AI agents.</p>
<h2 id="the"><strong>The <a href="http://todo.md/" rel="external nofollow noopener" class="lnp-link">TODO.md</a> Trap</strong></h2>
<p>So I did what everyone does: I created a <code>TODO.md</code> file.</p>
<pre><code>## TODO
- [ ] Add rate limiting to login endpoint
- [ ] Improve password hashing
- [ ] Fix email validation
- [ ] Build dashboard (depends on auth being done)
</code></pre>
<p>Seemed reasonable. Every project has a TODO list, right?</p>
<p><strong>Three days later, it was already a graveyard.</strong></p>
<p>Half the items were done but still unchecked. A quarter were outdated. New issues Claude discovered during implementation? Lost in chat history. Dependencies? I had “(depends on auth being done)” in a parenthetical. Good luck having Claude parse that reliably after compaction.</p>
<p>Steve Yegge calls these “swamps of rotten half-implemented plans.” He’s right.</p>
<p>Here’s why markdown TODOs fail with AI agents:</p>
<p><strong>They become stale instantly</strong> - You finish a task, forget to update the markdown. The agent reads it, doesn’t know what’s actually done.</p>
<p><strong>No dependency tracking</strong> - Can I start the dashboard? Is auth done? The agent has to guess.</p>
<p><strong>Context evaporates</strong> - “Fix email validation” tells you nothing. Which email? Where? What’s broken? Why does it matter? After compaction, this line is worthless.</p>
<p><strong>Agents can’t use them reliably</strong> - Claude reads the whole list, can’t tell what’s ready to work on, and often just&hellip; ignores it.</p>
<p>That <a href="http://todo.md/" rel="external nofollow noopener" class="lnp-link">TODO.md</a> file? After compaction, it’s all you have. And it’s not enough.</p>
<h2 id="enter-beads-a-memory-system-built-for-agents"><strong>Enter Beads: A Memory System Built for Agents</strong></h2>
<p>That’s when I found <a href="https://github.com/steveyegge/beads" rel="external nofollow noopener" class="lnp-link">beads</a>.</p>
<p>Steve Yegge built it specifically to solve the amnesia problem. It’s an issue tracker, but not like Jira or Linear. It’s <strong>built for AI agents, not humans.</strong></p>
<p>Here’s the breakthrough: <strong>You don’t manage beads. Claude does.</strong></p>
<p>I ran <code>bd init</code> in my project. That’s it. One command.</p>
<p>Then I just&hellip; kept working with Claude. And something wild happened.</p>
<p>Claude started filing issues automatically. Not when I asked—just when it discovered something worth remembering. While implementing the auth system, Claude would say:</p>
<p>“I notice there’s a potential race condition in the session handler. Filing it:</p>
<p><code>bd create ‘Fix race condition in user session’ --deps discovered-from:myapp-a4f2 -p 0</code>“</p>
<p>It wasn’t asking permission. It wasn’t waiting for me to maintain a TODO list. <strong>It was building its own memory.</strong></p>
<p>At the end of each session, Claude would update the notes on active issues:</p>
<pre><code>bd update myapp-a4f2 --notes “COMPLETED: JWT login with RS256
KEY DECISION: Chose RS256 over HS256 for key rotation capability
IN PROGRESS: Session middleware
BLOCKERS: Need decision on token expiry (1hr vs 24hr trade-off)
NEXT: Implement refresh endpoint after expiry decision”
</code></pre>
<p>This wasn’t for me. This was <strong>for future-Claude.</strong></p>
<h2 id="the-compaction-survival-test"><strong>The Compaction survival Test</strong></h2>
<p>The next day, I hit compaction again. New session, empty context.</p>
<p>Me: “What should we work on?”</p>
<p>Claude (without me asking, automatically ran these commands):</p>
<pre><code>bd ready --json
bd list --status in_progress --json
bd show myapp-a4f2
</code></pre>
<p>Then replied:</p>
<p>“I can see we’re working on authentication (<code>myapp-a4f2</code>). Last session we completed the JWT login flow, decided to use RS256 for key rotation capability, and got blocked waiting for your decision on token expiry time.</p>
<p>The trade-off is: 1-hour tokens (better security, worse UX) vs 24-hour tokens (better UX, but wider attack window if tokens are compromised). What’s your preference?</p>
<p>Also, I’ve filed 3 issues we discovered along the way:</p>
<ul>
<li>
<p><strong>Rate limiting for login endpoint</strong> (P0)</p>
</li>
<li>
<p><strong>Bcrypt rounds for password hashing</strong> (P1)</p>
</li>
<li>
<p><strong>Race condition in session handler</strong> (P0)</p>
</li>
</ul>
<p>Should we resolve the token expiry question first, or tackle one of the P0 issues?”</p>
<p><strong>The conversation history was gone. But Claude had full context.</strong></p>
<p>Every decision. Every discovered bug. Every blocker. Every trade-off. Right there.</p>
<p>No re-explaining. No “wait, what were we doing?” No hunting through old conversations.</p>
<p>This is what beads does.</p>
<h2 id="todowrite-vs-beads-two-memory-systems"><strong>TodoWrite vs Beads: Two Memory Systems</strong></h2>
<p>Here’s where people get confused. Claude actually has <em>two</em> memory systems, and they serve different purposes.</p>
<h3 id="todowrite-working-memory-this-hour"><strong>TodoWrite: Working Memory (This Hour)</strong></h3>
<p>TodoWrite is Claude’s <strong>scratch pad for the current session</strong>:</p>
<pre><code>✓ [completed] Implement login endpoint
→ [in_progress] Add password hashing
  [pending] Create session middleware
</code></pre>
<p>It shows you real-time progress. Gets marked complete as work happens. <strong>Disappears when the session ends.</strong></p>
<p>Perfect for: “What’s Claude doing right now?”</p>
<h3 id="beads-long-term-memory-this-weekmonth"><strong>Beads: Long-Term Memory (This Week/Month)</strong></h3>
<p>Beads is Claude’s <strong>episodic memory across sessions</strong>:</p>
<pre><code>bd show myapp-a4f2

Notes: “COMPLETED: Login with bcrypt (12 rounds)
KEY DECISION: JWT (not sessions) for stateless auth
IN PROGRESS: Session middleware
NEXT: Need input on token expiry (1hr vs 24hr)”
</code></pre>
<p>Survives compaction. Captures meaning, not just tasks. <strong>Persists across all sessions.</strong></p>
<p>Perfect for: “What happened last week? What decisions were made?”</p>
<h3 id="the-handoff-pattern"><strong>The Handoff Pattern</strong></h3>
<ol>
<li>
<p><strong>Session start</strong>: Claude reads bead notes → creates TodoWrite items for immediate work</p>
</li>
<li>
<p><strong>During work</strong>: TodoWrite gets marked complete</p>
</li>
<li>
<p><strong>Reach milestone</strong>: Claude updates bead notes with outcomes + context</p>
</li>
<li>
<p><strong>Session end</strong>: TodoWrite disappears, bead survives with enriched notes</p>
</li>
</ol>
<p><strong>After compaction</strong>: TodoWrite is gone forever. Bead notes reconstruct everything.</p>
<h2 id="the-magic-dependencies-that-prevent-mistakes"><strong>The Magic: Dependencies That Prevent Mistakes</strong></h2>
<p>This is where beads gets brilliant. It supports four relationship types:</p>
<h3 id="1-blocks---hard-blocker"><strong>1.</strong> <code>blocks</code> <strong>- Hard Blocker</strong></h3>
<pre><code>bd create “Build user dashboard” -p 1
# Created myapp-e3f7

bd create “Implement authentication” -p 0
# Created myapp-g2h9

bd dep add myapp-e3f7 myapp-g2h9
# → “myapp-g2h9 blocks myapp-e3f7”
</code></pre>
<p>Now dashboard won’t show in <code>bd ready</code> until auth is closed. Claude <strong>can’t accidentally start building the dashboard before auth exists.</strong></p>
<h3 id="2-discovered-from---the-audit-trail"><strong>2.</strong> <code>discovered-from</code> <strong>- The Audit Trail</strong></h3>
<p>This is the agent’s secret weapon:</p>
<pre><code># Claude finds bug B while implementing feature A
bd create “Fix memory leak in session handler” \
  --deps discovered-from:myapp-a4f2 -p 0
</code></pre>
<p>Creates an audit trail of how work was found. Those “oh by the way” issues Claude mentions? They now get filed permanently, linked to context.</p>
<p>After a week of work, you have an <strong>automatically maintained discovery backlog</strong>. Prioritized. Linked. Ready to tackle.</p>
<h3 id="3-parent-child---hierarchy"><strong>3.</strong> <code>parent-child</code> <strong>- Hierarchy</strong></h3>
<pre><code>bd create “Epic: Authentication system” -t epic
# Created myapp-j4k2

bd create “Add OAuth” --parent myapp-j4k2
# Created myapp-l8m1 (auto-linked)
</code></pre>
<p>Good for breaking down large features.</p>
<h3 id="4-related---soft-connection"><strong>4.</strong> <code>related</code> <strong>- Soft Connection</strong></h3>
<pre><code>bd dep add myapp-b7c3 myapp-d1e8 -t related
# “These touch the same code but don’t block each other”
</code></pre>
<h2 id="what-you-actually-do-almost-nothing"><strong>What You Actually Do (Almost Nothing)</strong></h2>
<p><strong>Your workflow:</strong></p>
<p><strong>One-time setup:</strong></p>
<pre><code>cd your-project
bd init
</code></pre>
<p>Done. That’s it.</p>
<p><strong>Work with Claude normally:</strong></p>
<p>“Let’s build user authentication”</p>
<p>Claude automatically:</p>
<ul>
<li>
<p>Creates issues as work emerges</p>
</li>
<li>
<p>Tracks dependencies</p>
</li>
<li>
<p>Updates notes at milestones</p>
</li>
<li>
<p>Files discovered work with proper links</p>
</li>
<li>
<p>Checks ready work at session start</p>
</li>
</ul>
<p><strong>You just work.</strong> The memory management happens in the background.</p>
<p><strong>When you DO interact with beads</strong> (rarely):</p>
<pre><code># Weekly review
bd stats

# Check what’s blocked
bd blocked

# Context restore after time away
bd show myapp-a4f2
</code></pre>
<p>The agent does the rest.</p>
<h2 id="why-claude-loves-it"><strong>Why Claude Loves It</strong></h2>
<p>The most interesting thing about beads isn’t the technology. It’s <strong>how Claude uses it.</strong></p>
<p>Claude’s behavior changes:</p>
<p><strong>1. Proactive filing</strong>: Claude files issues without being asked. “I notice X could be improved. Filing: <code>bd create...</code>“</p>
<p><strong>2. Better planning</strong>: Claude uses dependencies to think through work order before starting.</p>
<p><strong>3. Context awareness</strong>: Claude references past decisions from bead notes. “Last session we decided to use RS256 because&hellip;”</p>
<p><strong>4. Discovery tracking</strong>: Claude treats discovered work as first-class, not throwaways.</p>
<p><strong>Why?</strong> Because beads is built for how Claude actually works:</p>
<ul>
<li>
<p>Structured data (JSON)</p>
</li>
<li>
<p>Clear state (open/in_progress/closed)</p>
</li>
<li>
<p>Explicit relationships (dependencies)</p>
</li>
<li>
<p>Queryable memory (show me what’s ready)</p>
</li>
</ul>
<p>It’s not forcing Claude into a human workflow. It’s giving Claude the database it naturally wants.</p>
<h2 id="when-beads-is-overkill"><strong>When Beads is Overkill</strong></h2>
<p>Not every task needs beads. Use this test:</p>
<h3 id="use-beads-when"><strong>Use Beads when:</strong></h3>
<ul>
<li>
<p>Work spans multiple sessions</p>
</li>
<li>
<p>You might hit compaction before finishing</p>
</li>
<li>
<p>There are dependencies or blockers</p>
</li>
<li>
<p>You’re discovering related work along the way</p>
</li>
<li>
<p>You need to resume after time away</p>
</li>
</ul>
<p><strong>Example</strong>: “Build authentication system” (multi-day, many parts)</p>
<h3 id="use-todowrite-when"><strong>Use TodoWrite when:</strong></h3>
<ul>
<li>
<p>Work completes in this session</p>
</li>
<li>
<p>It’s a simple linear checklist</p>
</li>
<li>
<p>All context is in the conversation</p>
</li>
<li>
<p>No dependencies or discovery</p>
</li>
</ul>
<p><strong>Example</strong>: “Refactor this 200-line file” (done in an hour)</p>
<p><strong>The test</strong>: “Will I need this context in 2 weeks?”</p>
<ul>
<li>
<p><strong>Yes</strong> → Beads</p>
</li>
<li>
<p><strong>No</strong> → TodoWrite</p>
</li>
</ul>
<h2 id="the-git-sync-how-it-works-across-machines"><strong>The Git Sync: How It Works Across Machines</strong></h2>
<p>Beads stores everything in two places:</p>
<ol>
<li>
<p><code>.beads/beads.db</code> - Local SQLite (fast queries)</p>
</li>
<li>
<p><code>.beads/issues.jsonl</code> - Git-versioned JSONL (syncs across machines)</p>
</li>
</ol>
<p><strong>On your desktop:</strong></p>
<pre><code>bd create “New issue”
# → SQLite write (instant)
# → After 5 seconds, exports to JSONL
# → Git commit with your code changes
</code></pre>
<p><strong>On your laptop:</strong></p>
<pre><code>git pull
# → JSONL updates
# → bd auto-imports (newer than local DB)
# → SQLite now has the issue
</code></pre>
<p>You get:</p>
<ul>
<li>
<p>Fast local operations (SQLite, &lt;100ms)</p>
</li>
<li>
<p>Git versioning (full audit trail)</p>
</li>
<li>
<p>Multi-machine sync (JSONL)</p>
</li>
<li>
<p>Offline support (no server)</p>
</li>
</ul>
<p>It’s a distributed database&hellip; that’s just files in git.</p>
<h2 id="memory-as-infrastructure"><strong>Memory as Infrastructure</strong></h2>
<p>We’re at this weird moment where AI coding agents are incredibly capable but also incredibly forgetful.</p>
<p>We expect them to remember complex multi-week projects, track dozens of discovered issues, maintain perfect context across compaction—but we give them&hellip; markdown files.</p>
<p><strong>Beads doesn’t make agents smarter. It makes them less forgetful.</strong></p>
<p>And honestly? That might be more important.</p>
<p>Because the hardest part of any project isn’t writing code. It’s <strong>not losing track of what needs to be written.</strong></p>
<p>Beads gives your agent:</p>
<ul>
<li>
<p>Memory that survives compaction</p>
</li>
<li>
<p>A discovery backlog that doesn’t evaporate</p>
</li>
<li>
<p>A dependency graph that prevents mistakes</p>
</li>
</ul>
<p><strong>And you barely have to do anything.</strong> Install it, initialize it, let Claude manage it.</p>
<p>The agent handles the rest.</p>
<p><strong>Get started:</strong></p>
<ul>
<li>
<p>GitHub: <a href="https://github.com/steveyegge/beads" rel="external nofollow noopener" class="lnp-link">steveyegge/beads</a></p>
</li>
<li>
<p>Quick start: <code>bd init</code> in your project</p>
</li>
<li>
<p>Let Claude do the rest</p>
</li>
</ul>
<p><strong>Key commands</strong> (mostly for reference—Claude uses these automatically):</p>
<pre><code>bd init                    # One-time setup
bd ready                   # What’s ready? (Claude checks this)
bd show &lt;id&gt;               # Issue details (Claude reads notes)
bd stats                   # Weekly review (you use this)
bd blocked                 # What’s stuck?
</code></pre>
<p>Give your agent a memory. See what happens.</p>
]]></content:encoded></item><item><title>When Claude Code Goes Down: A Meditation on Modern Dependency</title><link>https://lakshminp.com/2025/10/claude-code-downtime/</link><pubDate>Fri, 31 Oct 2025 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2025/10/claude-code-downtime/</guid><category>essays</category><category>claude-code</category><description>Let me paint you a picture. It’s this evening. I’m in the zone. Fingers flying across the keyboard, that beautiful flow state where you and your AI coding assistant are one harmonious bug-squashing machine. And then, without warning, without so much as a courtesy error message, Claude Code just… dies.
Not the graceful kind of death where systems send you helpful notifications. Well, okay, they had a status page. There was technically a “we’re experiencing technical difficulties” message somewhere on the internet if you went looking for it. But in the moment? When you’re mid-keystroke and suddenly your AI copilot just stops responding? It just felt gone. Vanished. Disappeared like my motivation to manually write boilerplate code.</description><content:encoded><![CDATA[<p>Let me paint you a picture. It’s this evening. I’m in the zone. Fingers flying across the keyboard, that beautiful flow state where you and your AI coding assistant are one harmonious bug-squashing machine. And then, without warning, without so much as a courtesy error message, Claude Code just&hellip; dies.</p>
<p>Not the graceful kind of death where systems send you helpful notifications. Well, okay, they had a status page. There was <em>technically</em> a “we’re experiencing technical difficulties” message somewhere on the internet if you went looking for it. But in the moment? When you’re mid-keystroke and suddenly your AI copilot just stops responding? It just felt gone. Vanished. Disappeared like my motivation to manually write boilerplate code.</p>
<p>For approximately thirty seconds, I experienced what I can only describe as the five stages of grief compressed into real-time panic. Denial: “It’s just my internet.” Anger: “ARE YOU KIDDING ME RIGHT NOW?” Bargaining: “Maybe if I restart everything seventeen times&hellip;” Depression: “I guess I’m just not deploying tonight.” And finally, acceptance: “Well, I suppose I could try coding like it’s 2019.”</p>
<p>The outage lasted about five hours. Five. Entire. Hours. In developer time, that’s basically a geological epoch. I had bugs to fix. Tests to run. Production deployments waiting. And here I was, suddenly expected to do all of this using only my own fragile, fallible human brain.</p>
<p>So I did what any reasonable developer would do: I panicked for another minute, then reluctantly dusted off those ancient skills we used to call “programming without an AI safety net.”</p>
<p>Here’s the uncomfortable truth nobody wants to admit: it wasn’t <em>that</em> bad. I mean, it was bad. Don’t get me wrong. It was slow and tedious and made me feel like I was debugging with mittens on. But I didn’t spontaneously combust. My IDE still worked. My fingers still remembered where the keys were. Muscle memory is apparently still a thing.</p>
<p>I fixed the bug. Eventually. It just took approximately three times longer than it should have because I had to do wild, archaic things like “read the documentation thoroughly” and “actually understand what my code was doing” instead of asking Claude Code to explain it to me like I’m five.</p>
<p>The testing phase was particularly brutal. Normally, I’d have Claude Code help me think through edge cases, generate test scenarios, and spot the stupid mistakes I’m invariably making. Instead, I had to use my own brain to think of test cases. My <em>own</em> brain! Like some kind of caveman! I had to actually remember what good test coverage looks like and implement it myself. The horror.</p>
<p>And deployment? Well, deployment was already sorted, thankfully. But the whole process of getting there—fixing the bug, testing it properly, making sure everything was ready to ship—felt like wading through molasses. Without my AI copilot catching my typos, suggesting optimizations, and helping me think through edge cases, every step just took longer than it should have.</p>
<p>But here’s where it gets really pathetic. After about two hours of this manual labor cosplay, I had a brilliant idea. Claude Code might be down, but wasn’t there Claude Code <em>web</em>? Like, the browser version? Different infrastructure, right? Surely that was still running?</p>
<p>So I did what any self-respecting, definitely-not-addicted developer would do: I pulled my entire git repo into Claude Code web and just&hellip; kept working there. Yes, you read that right. My solution to Claude Code being down was to use a different version of Claude Code. I replaced my broken AI dependency with a slightly different flavour of the exact same AI dependency.</p>
<p>The technical term for this is “problem-solving.” The accurate term for this is “I have a problem.”</p>
<p>It actually worked pretty well as an interim hack, which is either a testament to Anthropic’s redundancy planning or a damning indictment of my ability to function independently. Probably both. The web version was a bit clunkier for my workflow, sure, but it beat slowly dying inside while manually parsing error messages.</p>
<p>The whole experience gave me a weird kind of perspective. It’s like when your phone dies and you suddenly remember you have hands and can look at things in real life. Except instead of appreciating nature, I was appreciating how much faster AI makes me at my job.</p>
<p>I can technically code without Claude Code. I proved that tonight. It’s like how I can technically do math without a calculator—possible, legal, but why would I choose suffering? The old ways still work. They’re just&hellip; inefficient. Tedious. The kind of thing that makes you question your career choices around the third hour of manually debugging something that Claude Code would’ve spotted in thirty seconds.</p>
<p>But sitting there in the dark ages of 2019-style development, something struck me. It’s been barely any time at all since AI coding assistants became genuinely useful. A few years ago, we were all coding exactly like this—manually, slowly, relying entirely on our own pattern recognition and Stack Overflow. And we thought we were pretty damn efficient.</p>
<p>Now? Now a five-hour outage feels like a crisis. That’s how far we’ve come. That’s how quickly this technology went from “neat party trick” to “fundamental part of my workflow” to “how did I ever function without this.”</p>
<p>I spent those five hours slightly inconvenienced, moving slower than usual, but still shipping code to production. A decade ago, this was just called “having a normal day at work.” Today, it felt like working with a handicap. That shift happened so fast we barely noticed it occurring.</p>
<p>We’re living through one of those rare moments where technology isn’t just improving incrementally—it’s fundamentally changing how we work. And tonight, in the brief absence of that technology, I got a glimpse of both where we’ve been and how far we’ve traveled.</p>
<p>The tools came back online. I went back to my normal pace. But I won’t forget that brief window of forced perspective, that reminder that we’re experiencing something genuinely transformative in real-time. Even if it did feel painfully slow while it was happening.</p>
]]></content:encoded></item></channel></rss>