<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>AI agents — Lakshmi Narasimhan</title><link>https://lakshminp.com/tags/ai-agents/</link><description>I help developers build, deploy, and distribute their SaaS without hiring a team. Long-running notes on systems, AI internals, Carnatic music, fiction craft, and whatever else collides interestingly.</description><generator>Hugo + lakshminp theme</generator><language>en-us</language><lastBuildDate>Mon, 22 Jun 2026 00:00:00 +0000</lastBuildDate><managingEditor>Lakshmi Narasimhan</managingEditor><webMaster>Lakshmi Narasimhan</webMaster><copyright>© 2026 Lakshmi Narasimhan</copyright><atom:link href="https://lakshminp.com/tags/ai-agents/feed.xml" rel="self" type="application/rss+xml"/><item><title>I Went Looking for Real-World AI Agent Examples. They're Rare.</title><link>https://lakshminp.com/2026/06/real-world-ai-agent-examples/</link><pubDate>Mon, 22 Jun 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/06/real-world-ai-agent-examples/</guid><category>essays</category><category>ai-coding</category><category>ai-agents</category><description>I’ll be honest up front: I’m still learning this stuff. I’m not writing this from a mountaintop. I’m writing it from the foothills, with muddy boots, having just figured out something that I suspect a lot of people pretend they already knew.
Here’s the thing that finally clicked for me. An agent is a loop. A model looks at the situation, decides one next step, calls a tool to do it, looks at what happened, and goes around again until it’s done. That’s it. I felt a little cheated when I understood it — the word “agent” had been doing so much heavy lifting on so many landing pages that I’d assumed there was a fortress behind it. There isn’t. There’s a while-loop.</description><content:encoded><![CDATA[<p>I&rsquo;ll be honest up front: I&rsquo;m still learning this stuff. I&rsquo;m not writing this from a mountaintop. I&rsquo;m writing it from the foothills, with muddy boots, having just figured out something that I suspect a lot of people pretend they already knew.</p>
<p>Here&rsquo;s the thing that finally clicked for me. An agent is a loop. A model looks at the situation, decides one next step, calls a tool to do it, looks at what happened, and goes around again until it&rsquo;s done. That&rsquo;s it. I felt a little cheated when I understood it — the word &ldquo;agent&rdquo; had been doing so much heavy lifting on so many landing pages that I&rsquo;d assumed there was a fortress behind it. There isn&rsquo;t. There&rsquo;s a while-loop.</p>
<p>So I went and read about the frameworks. All of them — LangGraph, CrewAI, LlamaIndex, the OpenAI Agents SDK, Pydantic AI, smolagents, the Claude Agent SDK, the vendor SDKs from Google and Amazon and Microsoft. And every single one walks you through the same starter example: a weather bot. Or &ldquo;chat with your PDF.&rdquo; Or my personal favorite, the demo where five agents — a Researcher, a Writer, a Critic, an Editor, and presumably a Manager to schedule their standups — collaborate to produce <a href="/2025/10/ai-agent-mistakes/" class="lnp-link">a blog post slightly worse than one agent would&rsquo;ve written</a>.</p>
<p>And I kept thinking: <em>okay, but where are the real ones?</em></p>
<p>Not the demos. Not the quickstart. Something non-trivial. Something that acts on the world, where a wrong move costs money or breaks production. I genuinely couldn&rsquo;t picture one. So instead of pretending, I went looking.</p>
<p>(The method, since it&rsquo;s too on-the-nose not to mention: I sent a <a href="/2026/01/6-ai-agents-coding-experiment/" class="lnp-link">small swarm of research agents</a> out across the web to comb engineering blogs and case studies for me, in parallel, while I made coffee. Hunting for proof that real agents exist turned out to be the most real agent use I&rsquo;d touched all week. Make of that what you will.)</p>
<p>Here&rsquo;s what I actually found.</p>
<h2 id="the-good-news-real-ones-exist">The good news: real ones exist</h2>
<p>A few of them are unambiguously real, and they&rsquo;re worth describing, because they taught me more about what an agent is <em>for</em> than any framework doc did.</p>
<p><strong>Sentry&rsquo;s Autofix</strong> is the one that changed my mind. When something breaks in a codebase Sentry monitors, an agent built on the Claude Agent SDK takes their root-cause analysis, plans a fix, <em>writes the code</em>, and opens a pull request you can actually merge — a full run in about six minutes. This isn&rsquo;t a chatbot that suggests you &ldquo;consider checking your null values.&rdquo; It writes the patch. And it runs against a platform doing over a million root-cause analyses a year. One of their engineers shipped it in weeks and wrote a piece literally titled <a href="https://blog.sentry.io/how-sentrys-ai-autofix-changed-my-mind-about-ai-agents/" rel="external nofollow noopener" class="lnp-link"><em>how Sentry&rsquo;s AI Autofix changed my mind about AI agents</em></a>. I felt seen.</p>
<p><strong>Amazon has an internal agent that troubleshoots network failures</strong> — diagnoses live VPC connectivity problems and resolves around 80% of network root causes on its own. Built on their <a href="https://strandsagents.com/blog/what-we-learned-from-one-year-of-building-production-agents/" rel="external nofollow noopener" class="lnp-link">Strands SDK</a>. That&rsquo;s an on-call SRE&rsquo;s nightmare-shift, handed to a loop. As someone who&rsquo;s done that shift, that number did something to me.</p>
<p><strong>Coinbase built <a href="https://github.com/coinbase/agentkit" rel="external nofollow noopener" class="lnp-link">a toolkit that gives an agent a crypto wallet</a>.</strong> The agent can hold funds, sign transactions, and pay for things autonomously. Read that again. We&rsquo;ve spent this whole article saying the scary part of agents is irreversible action with real stakes — and here&rsquo;s one wired directly to money on a blockchain, where &ldquo;oops&rdquo; is permanent. Terrifying. Also clearly real.</p>
<p><strong>Bilt runs <a href="https://www.letta.com/case-studies/bilt" rel="external nofollow noopener" class="lnp-link">a <em>million</em> agents</a></strong> — one per user — on Letta, each holding that user&rsquo;s transaction and engagement history in <a href="/2025/11/ai-agent-memory-persistence/" class="lnp-link">persistent memory</a> to drive merchant recommendations. The whole pitch of Letta is memory, and here&rsquo;s someone betting a recommendation system on it at a scale I can&rsquo;t fully picture.</p>
<p>And a scattering more, each genuinely non-trivial: <a href="https://www.langchain.com/blog/top-5-langgraph-agents-in-production-2024" rel="external nofollow noopener" class="lnp-link">Exa&rsquo;s web-research agent and LinkedIn&rsquo;s text-to-SQL bot</a> (both on LangGraph, both acting against live production systems); a <a href="https://pydantic.dev/" rel="external nofollow noopener" class="lnp-link">medical-triage agent on Pydantic AI</a> validated across 329 clinician-checked scenarios; a <a href="https://www.llamaindex.ai/blog/case-study-tender-rfp-agent-for-construction-sector-with-softiq" rel="external nofollow noopener" class="lnp-link">construction-tender agent on LlamaIndex</a> that digests 100-page public bids and spits out risk reports; Uber automating code migrations across its monorepo.</p>
<p>So. Real agents exist. I can stop being a skeptic about <em>that</em>.</p>
<h2 id="the-uncomfortable-news-there-arent-many-and-the-vendors-are-grading-their-own-homework">The uncomfortable news: there aren&rsquo;t many, and the vendors are grading their own homework</h2>
<p>Here&rsquo;s the part that kept nagging me after the research came back.</p>
<p>For each framework, I could find maybe <strong>one to three</strong> genuinely non-trivial examples. Not dozens. Single digits. And almost every one of them was published by the company that <em>sells the framework.</em> Sentry&rsquo;s story is on Sentry&rsquo;s blog (fair enough — Sentry isn&rsquo;t Anthropic), but most of them live in the framework vendor&rsquo;s own marketing: LangChain&rsquo;s case-study page, Letta&rsquo;s case studies, AWS&rsquo;s own deep-dive, Google&rsquo;s own developer blog. Independent &ldquo;here&rsquo;s our war story and here&rsquo;s what broke&rdquo; write-ups from teams with no skin in the game? Vanishingly rare.</p>
<p>And some frameworks I genuinely <em>couldn&rsquo;t</em> find a real one for:</p>
<ul>
<li><strong>smolagents</strong> has <a href="https://github.com/huggingface/smolagents" rel="external nofollow noopener" class="lnp-link">26,000 GitHub stars</a> and I love its design — but its flagship example is Hugging Face&rsquo;s own research replication. I found no named company betting anything real on it.</li>
<li><strong>CrewAI</strong> is everywhere in demos and has a wall of enterprise logos (PepsiCo, J&amp;J, the DoD), but behind almost every logo is zero operational detail. The one solid story — <a href="https://blog.crewai.com/lessons-from-2-billion-agentic-workflows/" rel="external nofollow noopener" class="lnp-link">a five-agent sales pipeline at DocuSign</a> — is, again, on CrewAI&rsquo;s own blog.</li>
<li><strong>Microsoft&rsquo;s Agent Framework</strong> just hit 1.0 claiming &ldquo;real-world validation with customers and partners&rdquo; and then named exactly zero of them. Its most impressive artifact, <a href="https://www.microsoft.com/en-us/research/articles/magentic-one-a-generalist-multi-agent-system-for-solving-complex-tasks/" rel="external nofollow noopener" class="lnp-link">Magentic-One</a>, is explicitly a <em>research</em> system that doesn&rsquo;t ship inside a product.</li>
</ul>
<p>I want to be careful here, because I&rsquo;m still learning and I don&rsquo;t want to overclaim the cynicism: &ldquo;I couldn&rsquo;t find it&rdquo; is not &ldquo;it doesn&rsquo;t exist.&rdquo; A lot of the realest agent work is surely locked inside companies that will never blog about it. But the <em>public</em> record, right now, is thin. Much thinner than the hype implied. The ratio of &ldquo;agentic platform&rdquo; marketing to &ldquo;here is a real agent doing a real job&rdquo; is grim.</p>
<h2 id="two-things-i-think-im-learning">Two things I think I&rsquo;m learning</h2>
<p>I&rsquo;m holding these loosely, because foothills. But:</p>
<p><strong>The best real agents are vendors using their own tools.</strong> Amazon&rsquo;s network agent, <a href="https://developers.googleblog.com/en/agent-development-kit-easy-to-build-multi-agent-applications/" rel="external nofollow noopener" class="lnp-link">Google&rsquo;s enterprise agents on ADK</a>, Strands originating inside Amazon Q Developer — the most concrete, number-backed cases are companies dogfooding the framework they built. That&rsquo;s either reassuring (they believe in it enough to run it) or a little hollow (of course the toolmaker has the best tool demo). Probably both.</p>
<p><strong>Every real one acts. None of them chat.</strong> This is the pattern that actually reorganized my thinking. Line up the genuinely non-trivial agents — writes a mergeable PR, signs a transaction, resolves a network outage, holds a million users&rsquo; memory, files a risk report on a 100-page tender. Not one of them is a conversation. The toys all talk. The real ones <em>do</em>. The demos cluster around chat because chat is safe and reversible and impresses in a screenshot. The real ones cluster around irreversible action because that&rsquo;s where an agent is actually worth the risk of building.</p>
<p>Which, looping all the way back, is exactly why the weather bot felt so empty. A weather bot doesn&rsquo;t <em>do</em> anything. It&rsquo;s the loop with the stakes amputated.</p>
<h2 id="so-where-does-that-leave-a-beginner">So where does that leave a beginner</h2>
<p>I don&rsquo;t have a grand conclusion. I have a working hypothesis, which is the most an honest learner should claim: the framework you pick matters far less than whether you have a real job that needs an agent that <em>acts</em>. If you don&rsquo;t, no framework will save you — you&rsquo;ll build a five-agent demo and quietly stop opening the repo. If you do, the loop is twenty lines, and you should start with whichever framework hides the least so you can actually see what&rsquo;s happening (smolagents, the OpenAI Agents SDK, and Pydantic AI were the ones that got out of my way the most).</p>
<p>And honestly? The fact that real examples are still this rare didn&rsquo;t discourage me. It read like a timestamp. We&rsquo;re early. The scarcity isn&rsquo;t proof the idea is empty — it&rsquo;s proof most people are still building weather bots while a handful of teams quietly wire a loop up to something that matters.</p>
<p>I&rsquo;d rather be in the second group — which is why I&rsquo;m slowly <a href="/2026/01/agent-orchestrator/" class="lnp-link">building one of my own</a>. I&rsquo;m still learning how.</p>
<hr>
<h2 id="the-ledger-the-realest-example-i-found-per-framework-and-where-its-published">The ledger (the realest example I found per framework, and where it&rsquo;s published)</h2>
<p><em>Honest tag: most of these are vendor-published. Independent confirmation is scarce — which is part of the story.</em></p>
<ul>
<li><strong>Claude Agent SDK</strong> — Sentry Autofix: writes mergeable PRs against 1M+ RCAs/yr → <a href="https://blog.sentry.io/how-sentrys-ai-autofix-changed-my-mind-about-ai-agents/" rel="external nofollow noopener" class="lnp-link">blog.sentry.io</a>, <a href="https://claude.com/customers/sentry" rel="external nofollow noopener" class="lnp-link">claude.com/customers/sentry</a></li>
<li><strong>AWS Strands</strong> — Amazon internal network-troubleshooting agent (~80% of network root causes); origin of Amazon Q Developer → <a href="https://strandsagents.com/blog/what-we-learned-from-one-year-of-building-production-agents/" rel="external nofollow noopener" class="lnp-link">strandsagents.com</a></li>
<li><strong>Letta</strong> — Bilt: ~1M per-user memory agents for recommendations → <a href="https://www.letta.com/case-studies/bilt" rel="external nofollow noopener" class="lnp-link">letta.com/case-studies/bilt</a></li>
<li><strong>OpenAI Agents SDK</strong> — Coinbase AgentKit: agents with on-chain wallets, real transactions → <a href="https://github.com/coinbase/agentkit" rel="external nofollow noopener" class="lnp-link">github.com/coinbase/agentkit</a></li>
<li><strong>LangGraph</strong> — Exa web-research agent; LinkedIn text-to-SQL bot; Uber code migrations → <a href="https://www.langchain.com/blog/exa" rel="external nofollow noopener" class="lnp-link">langchain.com/blog/exa</a>, <a href="https://www.langchain.com/blog/top-5-langgraph-agents-in-production-2024" rel="external nofollow noopener" class="lnp-link">top-5 in production</a></li>
<li><strong>Pydantic AI</strong> — STCC medical-triage agentic RAG (329 validated scenarios) → <a href="https://pydantic.dev/" rel="external nofollow noopener" class="lnp-link">pydantic.dev</a></li>
<li><strong>LlamaIndex</strong> — SoftIQ construction-tender agent (100-page bids → risk reports) → <a href="https://www.llamaindex.ai/blog/case-study-tender-rfp-agent-for-construction-sector-with-softiq" rel="external nofollow noopener" class="lnp-link">llamaindex.ai case study</a></li>
<li><strong>Google ADK</strong> — Google&rsquo;s own Agentspace/contact-center agents (6T+ tokens/mo); Renault EV-charger siting; Box contract extraction → <a href="https://developers.googleblog.com/en/agent-development-kit-easy-to-build-multi-agent-applications/" rel="external nofollow noopener" class="lnp-link">developers.googleblog.com</a></li>
<li><strong>CrewAI</strong> — DocuSign 5-agent sales Flow (vendor blog) → <a href="https://blog.crewai.com/lessons-from-2-billion-agentic-workflows/" rel="external nofollow noopener" class="lnp-link">blog.crewai.com</a></li>
<li><strong>smolagents</strong> — no named production company found; flagship is HF&rsquo;s own Open Deep Research → <a href="https://github.com/huggingface/smolagents" rel="external nofollow noopener" class="lnp-link">github.com/huggingface/smolagents</a></li>
<li><strong>Microsoft Agent Framework / AutoGen</strong> — mostly research (Magentic-One); 1.0 names zero customers → <a href="https://www.microsoft.com/en-us/research/articles/magentic-one-a-generalist-multi-agent-system-for-solving-complex-tasks/" rel="external nofollow noopener" class="lnp-link">microsoft.com/research</a></li>
</ul>
]]></content:encoded></item><item><title>Five Books Taught Me to Build AI Agents. All Five Quietly Told Me Not To.</title><link>https://lakshminp.com/2026/06/what-ai-agents-actually-are/</link><pubDate>Sun, 21 Jun 2026 00:00:00 +0000</pubDate><author>Lakshmi Narasimhan</author><guid isPermaLink="true">https://lakshminp.com/2026/06/what-ai-agents-actually-are/</guid><category>essays</category><category>ai-coding</category><category>ai-agents</category><description>What four hundred thousand words of agent literature agree on — and never put on the cover.
I bought five books on building AI agents in a single afternoon, the way you panic-buy bottled water before a storm. Manning had a sale. I had a credit card and a vague sense that everyone around me had quietly become an “agent engineer” while I was busy doing my actual job.
So I did the responsible thing. I spun up a small army of subagents to read four of them for me, cover to cover, in parallel, and report back. Which, if you’re keeping score, means I built a multi-agent system to summarize books about how to build multi-agent systems. The irony was not lost on me. It was, in fact, the first thing I learned.</description><content:encoded><![CDATA[<p><em>What four hundred thousand words of agent literature agree on — and never put on the cover.</em></p>
<p>I bought five books on building AI agents in a single afternoon, the way you panic-buy bottled water before a storm. Manning had a sale. I had a credit card and a vague sense that everyone around me had quietly become an &ldquo;agent engineer&rdquo; while I was busy doing my actual job.</p>
<p>So I did the responsible thing. I spun up a small army of subagents to read four of them for me, cover to cover, in parallel, and report back. Which, if you&rsquo;re keeping score, means I built a multi-agent system to summarize books about how to build multi-agent systems. The irony was not lost on me. It was, in fact, the first thing I learned.</p>
<p>Here&rsquo;s the second.</p>
<h2 id="the-loop-is-thirty-lines">The loop is thirty lines</h2>
<p>Strip away the diagrams and the framework comparisons, and every single one of these books — <a href="https://www.manning.com/books/build-an-ai-agent-from-scratch" rel="external nofollow noopener" class="lnp-link">Build an AI Agent</a>, <a href="https://www.manning.com/books/build-a-multi-agent-system-from-scratch" rel="external nofollow noopener" class="lnp-link">Build a Multi-Agent System</a>, <a href="https://www.manning.com/books/ai-agents-in-action-second-edition" rel="external nofollow noopener" class="lnp-link">AI Agents in Action</a>, <a href="https://www.manning.com/books/ai-agents-and-applications" rel="external nofollow noopener" class="lnp-link">AI Agents and Applications</a> — converges on the same humble definition.</p>
<p>An agent is a language model, plus some tools, plus a loop that runs until the job is done.</p>
<p>That&rsquo;s it. One book states it as plainly as that. Another dresses it up as a four-letter cycle. There&rsquo;s a Reddit thread floating around that implements the whole thing in about thirty lines of code, set to a drum-and-bass track, and honestly it explains the concept better than half the chapters I read.</p>
<p>There is no secret sauce. There is no priesthood. You were promised a cathedral and what you got is a <code>while</code> loop with good manners.</p>
<p>Which raised an obvious question, sitting there with four hundred thousand words of agent literature on my screen: if the core idea fits on a napkin, what&rsquo;s in all these books?</p>
<h2 id="the-part-nobody-puts-on-the-cover">The part nobody puts on the cover</h2>
<p>The answer is the same in every one, and it&rsquo;s the most useful thing I took away.</p>
<p>The loop is the easy ten percent. The other ninety — the part that doesn&rsquo;t fit in a demo — is evaluation, memory, guardrails, cost control, defending against prompt injection, and the deeply unglamorous skill of knowing when to hand the problem back to a human.</p>
<p>Three of the four books I read point at the same Anthropic paper, &ldquo;Building Effective Agents,&rdquo; like it&rsquo;s scripture. And buried in chapter one of each — past the exciting cover, past the part where they sell you on the future — every author tells you the same quiet thing.</p>
<p>Don&rsquo;t reach for an agent.</p>
<p>Start with a plain model call. Then a chain. Then a workflow. Earn the agent only when the task genuinely needs one, because an agent costs roughly ten times a normal call. Per task. Now imagine that thing running unattended, all night, while you sleep.</p>
<p>I went looking for the loudest voices on the other side of this — the practitioners on Reddit who build agents for a living and have the scar tissue to prove it. I expected an argument. The top thread is literally titled &ldquo;Stop building AI agents.&rdquo; Another is a guy who got <em>paid</em> to rip the AI back out of a tool he&rsquo;d shipped. A third is the 2 a.m. classic: the agent hit a question it didn&rsquo;t understand, confidently made up an answer, and emailed it to a customer.</p>
<p>The books and the burnouts weren&rsquo;t arguing. They&rsquo;d arrived at the same conclusion from opposite ends of the room. The model was never the bottleneck. Running the thing was.</p>
<h2 id="what-im-actually-taking-away">What I&rsquo;m actually taking away</h2>
<p>A small tell that stuck with me: agent-to-agent coordination shows up in the <em>subtitles</em> of these books far more confidently than it shows up in the chapters. The field is writing about how agents talk to each other a little faster than it&rsquo;s shipping it. That&rsquo;s not a knock — it&rsquo;s a map. It tells you where the hype is and where the ground is still wet.</p>
<p>So here&rsquo;s my take, for whatever a guy who outsourced his reading to robots is worth.</p>
<p>The framework you pick doesn&rsquo;t matter much; that code rots in eighteen months. What compounds is the boring stuff the demos skip — evaluation, context discipline, and the judgment to not build the agent at all. Everyone is rushing to learn how to <em>make</em> an agent. Almost nobody is learning how to make one you&rsquo;d actually trust.</p>
<p>The capability got democratized this year. The judgment didn&rsquo;t.</p>
<p>That gap — between the agent that runs and the agent you&rsquo;d let near production while you&rsquo;re asleep — is the whole job now. It&rsquo;s also, conveniently, <a href="/2025/10/the-real-skill-ai-wont-replace/" class="lnp-link">the only part worth getting good at</a>.</p>
]]></content:encoded></item></channel></rss>