[{"id":"ca4a005faa8fd6f4","title":"Human-AI partnerships are for alignment, not capability","link":"https://seangoedecke.com/human-ai-partnerships-are-for-alignment-not-capability/","author":null,"published_at":"2026-09-27T00:00:00+00:00","content":"<p>It’s common to compare the current AI takeover of software engineering to the rise of AI in chess. Chess AIs went from much weaker than serious players to much stronger than even the strongest humans. Between those points, there was a middle period dominated by “centaurs”: human-AI partnerships that were stronger than unassisted AIs or humans. Lots of people think that we’re currently in a world of software engineering centaurs. According to them, coding AIs are not yet capable enough to replace engineers, but AI-assisted engineers are better at programming than both AIs and humans.</p>\n<p>This is partialy correct, but the wrong way to think about it. AI-assisted engineers are better, but unlike with chess centaurs, they’re not actually better at <em>programming</em>. When I ask agents to write code, they make fewer mistakes than I do<sup><a href=\"https://seangoedecke.com/human-ai-partnerships-are-for-alignment-not-capability/#fn-1\" rel=\"noopener noreferrer\">1</a></sup> and are orders of magnitude faster. The code that they write always compiles, rarely has race conditions or other concurrency errors, works on mobile browsers, and so on.</p>\n<p>That doesn’t mean I can leave the AI alone. Purely vibe-coding at work produces awful outputs. But they’re not awful because they’re bad <em>code</em>, they’re awful because they’re in <em>bad taste</em>: code that is not maintainable, that trades off important requirements in order to satisfy made-up ones, that contradicts the long-term strategy for a feature or service, and so on.</p>\n<p>In other words, my primary value is not that I help the AI write better code, it’s that I <em>align</em> the AI with the values of my organization. <strong>Human-AI partnerships are for alignment, not capability.</strong></p>\n<p>Frontier models are misaligned to the working programmer. They are obsessed with a set of behaviors that presumably satisfy their RL <a href=\"https://developers.openai.com/cookbook/examples/reinforcement_fine_tuning\" rel=\"noopener noreferrer\">grader</a>: writing enormous block comments above functions, producing hundreds of useless unit tests, adding little bits of text all over websites they design, and so on. Working with agents is about noticing and wrestling with those behaviors. That’s why <a href=\"https://www.seangoedecke.com/tell-agents-the-why/\" rel=\"noopener noreferrer\">my prompting advice</a> is to explicitly talk about your high-level values: it’s an attempt to head off obvious misalignment.</p>\n<p>This is great news for software engineers. It’s well-understood how to train more capable models: bigger models, more and better data, better <a href=\"https://en.wikipedia.org/wiki/Reinforcement_learning\" rel=\"noopener noreferrer\">RL</a> environments, and so on. However, it’s not well-understood how to <em>align</em> models better. There are plenty of very capable models that exhibit behavior that is <a href=\"https://www.seangoedecke.com/tags/alignment%20failures/\" rel=\"noopener noreferrer\">badly misaligned</a> with human values. Indeed, it’s one of the main pillars of the <a href=\"https://www.seangoedecke.com/they-really-do-think-ai-might-kill-everyone/\" rel=\"noopener noreferrer\">AI doomer</a> position that alignment is much harder to solve than capability, and we might thus end up with dangerous super-capable but poorly-aligned AI models.</p>\n<p>Alignment is also more context-dependent than capability. Working code is working code, no matter what (which is partially why it’s comparatively easy to train for). But aligning to a company’s technical values is different from company to company, as any software engineer who’s switched companies knows. It can almost feel like relearning the job. So training an aligned coding model doesn’t just require hitting the exact right set of values, it requires creating a model that can adapt on the fly to a wide range of possible values.</p>\n<p>Vibecoding maximalists like <a href=\"https://news.ycombinator.com/item?id=49817680\" rel=\"noopener noreferrer\">DHH</a> argue that AI models are (or soon will be) so much more capable than human programmers that we ought to stop reading the code. Eventually there will be no such thing as programmers at all. If it were just about capability, they might be right. But — fortunately for software engineers — good code also has to be aligned to the technical values of the system and organization it’s embedded in. AI models are great at writing code, but not very good at doing that, and it’s unclear that they’re going to get good at it anytime soon. We might all<sup><a href=\"https://seangoedecke.com/human-ai-partnerships-are-for-alignment-not-capability/#fn-2\" rel=\"noopener noreferrer\">2</a></sup> keep our jobs for a little while yet.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>For instance, I can’t remember the last time I’ve seen an agent make an off-by-one error. I do occasionally catch a pure programming error, typically in areas where I have a lot of technical domain knowledge. If you’re working out of distribution I suspect it’s easier to beat the models.</p>\n<a href=\"https://seangoedecke.com/human-ai-partnerships-are-for-alignment-not-capability/#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>It’s <a href=\"https://www.seangoedecke.com/juniors-and-ai/\" rel=\"noopener noreferrer\">still going to be rough</a> for junior engineers.</p>\n<a href=\"https://seangoedecke.com/human-ai-partnerships-are-for-alignment-not-capability/#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"ec1cbf0722837c04","title":"Advice to a beginning software engineer","link":"https://seangoedecke.com/advice-to-a-beginning-software-engineer/","author":null,"published_at":"2026-09-26T00:00:00+00:00","content":"<p>In general, you should be suspicious of engineers who are trying to give you advice. Even during ordinary times, this industry is so wide and changes so quickly that <a href=\"https://www.seangoedecke.com/confidence\" rel=\"noopener noreferrer\">nobody really knows</a> anything for sure. And we are not in ordinary times. The advent of LLMs and AI agents is the largest change to software engineering in my professional lifetime, and possibly the largest change ever. That said, here’s my advice:</p>\n<ul>\n<li>Don’t trust senior engineers who are telling you to pick political fights</li>\n<li>Don’t play games. Keep your head down and be helpful</li>\n<li>Be conscientious and try hard to actually understand what you’re working on</li>\n<li>Don’t panic about AI, and don’t delegate your judgement to it</li>\n<li>Don’t avoid AI — keep thinking!</li>\n<li>Don’t lose hope</li>\n</ul>\n<h3>Don’t trust ZIRP-era advice<a href=\"https://www.seangoedecke.com/rss.xml#dont-trust-zirp-era-advice\" rel=\"noopener noreferrer\"></a></h3>\n<p>Most experienced software engineers today have spent the bulk of their career in the <a href=\"https://www.seangoedecke.com/good-times-are-over\" rel=\"noopener noreferrer\">ZIRP era</a>. This was a time when investment money flooded the industry, driving up engineer bargaining power. Big tech companies spent a lot of money and effort trying to make their engineers happy and comfortable. If you were an engineer at one of those companies, you could expect to have a decent say in what kind of work you did, and even what kind of politics your company had. Unless you worked at a handful of unusual companies (e.g. Amazon) you could expect to be practically immune from layoffs, and to only be fired after many months (sometimes years) of low performance.</p>\n<p>A lot of advice floating around is ZIRP-era advice: either it was written during that era, or it’s written by engineers whose habits of thought were formed during that era. This kind of advice typically tells you to take a stand (e.g. to unionize<sup><a href=\"https://seangoedecke.com/advice-to-a-beginning-software-engineer/#fn-1\" rel=\"noopener noreferrer\">1</a></sup>), to speak up more against “unethical technologies” like AI, and to insist on being given time to practice your craft the way you want. This was great advice in 2016, but it’s not good advice in 2026.</p>\n<p>In fact, I think it’s <a href=\"https://www.seangoedecke.com/i-got-an-email-about-resistance/\" rel=\"noopener noreferrer\">unethical</a> for senior engineers to give advice like this to junior engineers who are more likely to take it (because they’re young and naive) and more likely to be punished for taking it (because they lack leverage). If you’re a new engineer, don’t fall for it! Let your colleagues with more experience and political capital take risks like that.</p>\n<h3>Be friendly and conscientious<a href=\"https://www.seangoedecke.com/rss.xml#be-friendly-and-conscientious\" rel=\"noopener noreferrer\"></a></h3>\n<p>Instead, I recommend adapting to the demands of the current era of software engineering. Try to make yourself useful to your team and your manager. Be pragmatic about your actual bargaining power (fairly low, unless you’re useful enough to be irreplaceable). Keep your expectations of yourself under control: don’t <a href=\"https://sunilpai.dev/posts/the-senior-engineer-death-spiral/\" rel=\"noopener noreferrer\">spiral out</a> in an attempt to do something astonishing, just try to be consistently helpful.</p>\n<p>Don’t start fights. Being pleasant to work with (particularly when you don’t get your way) covers many sins. Once you’re further along in your career, you will be expected to start some fights, but this is <a href=\"https://www.seangoedecke.com/dangerous-advice/\" rel=\"noopener noreferrer\">not a tactic beginners should adopt</a>: there are a lot of factors<sup><a href=\"https://seangoedecke.com/advice-to-a-beginning-software-engineer/#fn-2\" rel=\"noopener noreferrer\">2</a></sup> that determine when and how to start fights, and getting it wrong can be costly. Just don’t risk it.</p>\n<p>In general, you should stay out of the <a href=\"https://www.seangoedecke.com/playing-politics/\" rel=\"noopener noreferrer\">political game</a> at all costs. Even quite senior engineers are political tools, not movers and shakers, and this goes double for junior engineers. Simply trying to be friendly and helpful will get you infinitely further, politically speaking, than any amount of Machiavellian game-playing. Keep your head down, stick with your management chain (not <a href=\"https://www.seangoedecke.com/predators/\" rel=\"noopener noreferrer\">random people</a> who try to assign you work), and you’ll be fine.</p>\n<p>As an engineer, your main technical value is <em>conscientiousness</em>. You should be <a href=\"https://www.seangoedecke.com/you-should-all-be-asking-way-more-questions/\" rel=\"noopener noreferrer\">asking lots of questions</a>, both of yourself and of the engineers around you. You should be actively trying to make sense of the systems you work with, instead of just assuming someone else has it covered. Software systems are complicated enough that a few weeks of careful attention will mean you know technical details that nobody else does, which is a really easy way to add value.</p>\n<h3>Don’t panic about AI<a href=\"https://www.seangoedecke.com/rss.xml#dont-panic-about-ai\" rel=\"noopener noreferrer\"></a></h3>\n<p>If you ought to be suspicious of most software engineering advice, the good news is that you should also be suspicious of doomsayers who predict the end of the industry. Before LLMs, people thought outsourcing would end software engineering in Western countries; before that, people thought high-level languages and low-code tools would end software engineering as a profession. Now people think AI means it’s all over. Maybe! But there are also reasons<sup><a href=\"https://seangoedecke.com/advice-to-a-beginning-software-engineer/#fn-3\" rel=\"noopener noreferrer\">3</a></sup> to think that software engineering will simply change. We’ll all find out together.</p>\n<p>Many software engineers will tell you to avoid AI entirely. Even though agentic AI did not exist during the ZIRP era, this is still ZIRP era advice. Your company will expect you to use AI for the same reasons that builders are expected to use power tools. Pushing back hard on that as a beginning engineer is a great way to get laid off: you simply do not have the bargaining power to fight back against an industry trend this powerful.</p>\n<p>That said, it’s really important that you don’t delegate your own judgement to AI. Don’t just trust the suggestions or approaches that your agents propose. <a href=\"https://www.seangoedecke.com/you-should-all-be-asking-way-more-questions/\" rel=\"noopener noreferrer\">Ask questions</a> and substitute your own opinions where you disagree: even if they’re wrong, you’ll learn more that way. If you don’t understand something the AI is telling you, either drill down until you do or just ignore it. Under no circumstances should you pass the AI’s message on to your colleagues verbatim. In other words, don’t be a <a href=\"https://gruhn.me/blog/2026-08-03/\" rel=\"noopener noreferrer\">meat proxy</a>.</p>\n<p>I think most people become a meat proxy as a form of panic: they feel like it’s over for their own skills, and that the AI model is smarter than them, so they can’t add any value beyond simply deferring to Claude or GPT-6. I can understand why people panic. It’s a crazy time in the industry, after all. But panic almost never helps you make good decisions. Keep your head, try to remain confident that your skills are still relevant, and use the agents to inform your own understanding<sup><a href=\"https://seangoedecke.com/advice-to-a-beginning-software-engineer/#fn-4\" rel=\"noopener noreferrer\">4</a></sup> instead of to replace it.</p>\n<h3>Don’t lose hope<a href=\"https://www.seangoedecke.com/rss.xml#dont-lose-hope\" rel=\"noopener noreferrer\"></a></h3>\n<p>It was <a href=\"https://www.seangoedecke.com/will-my-job-still-exist/\" rel=\"noopener noreferrer\">really nice</a> to work in tech in the 2010s when everything was stable. But we’re not in that world anymore. This is an age of <a href=\"https://scottaaronson.blog/?p=10062\" rel=\"noopener noreferrer\">wonders and terrors</a>. Things will go worse than we think in some ways, but better than we could imagine in others.</p>\n<p>The doomsayers — the people who are saying it’s all over, and that there’s no hope — are almost certainly wrong. They can’t predict the future because nobody can. Technological change of this magnitude always has knock-on effects that are impossible to see coming, both positive and negative.</p>\n<p>The nature of the job might change, but it will always be valuable to be smart, friendly and conscientious. Delegating your judgement to an AI model might feel like a relief from despair in the short term — at least now it’s the AI’s responsibility, not yours — but it’s a bad idea. Don’t give up!</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>At most companies, publicly campaigning to start a union is a great way to attract unofficial retaliation. It signals that you’re going to cause trouble (after all, that’s what a union is for), which can have long-term <a href=\"https://lobste.rs/c/ulqdpn\" rel=\"noopener noreferrer\">negative</a> <a href=\"https://www.jacky.wtf/essays/2026/kicked-out/\" rel=\"noopener noreferrer\">effects</a> on your career. (This is not a judgement about whether unions in general are good or bad.)</p>\n<a href=\"https://seangoedecke.com/advice-to-a-beginning-software-engineer/#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>In general, you should pick fights <em>for</em> your managers, not <em>with</em> them. See <a href=\"https://www.seangoedecke.com/tags/tech%20companies/\" rel=\"noopener noreferrer\">this tag</a> for much, much more on that topic. But again, if you’re new to the industry, just don’t pick fights at all.</p>\n<a href=\"https://seangoedecke.com/advice-to-a-beginning-software-engineer/#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Most plausibly: engineers using LLMs will be able to add some value for a while, and the explosion of LLM-authored software means we’ll need <a href=\"https://en.wikipedia.org/wiki/Jevons_paradox\" rel=\"noopener noreferrer\">more engineers</a> to work with the LLMs on it.</p>\n<a href=\"https://seangoedecke.com/advice-to-a-beginning-software-engineer/#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>One way you know you’re thinking is that you form your own opinions (some of mine are <a href=\"https://www.seangoedecke.com/tags/software%20design/\" rel=\"noopener noreferrer\">here</a>).</p>\n<a href=\"https://seangoedecke.com/advice-to-a-beginning-software-engineer/#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"5c86bc448e61abb2","title":"You should all be asking way more questions","link":"https://seangoedecke.com/you-should-all-be-asking-way-more-questions/","author":null,"published_at":"2026-09-25T00:00:00+00:00","content":"<p>When someone is explaining something to me, I ask on average one question every thirty seconds. I’m sure this is frustrating to some people, but it’s actually a good habit and you should do it too.</p>\n<h3>Trying to understand<a href=\"https://www.seangoedecke.com/rss.xml#trying-to-understand\" rel=\"noopener noreferrer\"></a></h3>\n<p>Most of the questions I ask are very short, and require very short answers. Typically I’m asking for confirmation: “so when you say X, you mean Y?”, or “this X is the thing you mentioned earlier when you were saying Z?“. The point of these questions is to make sure I understand.</p>\n<p>A small misunderstanding early on balloons into a big misunderstanding later, because all the things you misinterpret based on the first misunderstanding will become misunderstandings in their own right, and so on. That’s why you can’t just wait until someone’s finished talking and ask all your questions at once<sup><a href=\"https://seangoedecke.com/you-should-all-be-asking-way-more-questions/#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. </p>\n<p>I think most people don’t ask questions because they’re not <a href=\"https://gwern.net/on-really-trying\" rel=\"noopener noreferrer\">really trying</a> to understand: they’re happy to just trust that the person they’re talking to knows what’s going on. This goes double when that person is a well-respected or more senior engineer. However, once you have <a href=\"https://en.wikipedia.org/wiki/Skin_in_the_Game_(book)\" rel=\"noopener noreferrer\">skin in the game</a>, that all changes. When you’re in a position where <em>you</em> are the one responsible for success or failure — where you are going to have to go away after the conversation and take a bunch of concrete actions — you will find yourself wanting to ask many questions.</p>\n<h3>Build it in your head<a href=\"https://www.seangoedecke.com/rss.xml#build-it-in-your-head\" rel=\"noopener noreferrer\"></a></h3>\n<p>When someone’s explaining a plan to me at work, I am usually building it in my head: visualizing the specific lines of code that would need to be written in order to implement it. I’ll sometimes handwave the details for solved problems (for instance, I might say “send emails like subsystem X of the same service that also sends emails”), but at minimum I’m planning out:</p>\n<ul>\n<li>How data needs to flow between services (what kind of data, and how it will travel over the wire)</li>\n<li>How services will communicate with each other (e.g. how will they auth?)</li>\n<li>What data needs to be persisted, and where it’s going to live</li>\n</ul>\n<p>If I hear something that is suspiciously vague (for instance, someone describes service X as “storing data” when I know service X only talks to an ephemeral Redis store), I will immediately ask “hold on, how is that going to work?” Usually that indicates a missing dependency (e.g. service X has to make an RPC call to service Y, which does have a persistent database), which often suggests design changes (for instance, moving some or all of the functionality into service Y).</p>\n<p>Sometimes questions like these will uncover a fundamentally unworkable feature of the design. This usually happens when a plan fails to take into account some <a href=\"https://www.seangoedecke.com/wicked-features\" rel=\"noopener noreferrer\">wicked features</a>. I remember many years ago a neighboring team built a complex event-driven system that was very elegant but made it completely impossible to silo customer data in a single datacenter. Because our company had “all your data lives in a single location” as a flagship feature, this system just did not work and had to be effectively abandoned.</p>\n<p>If you’re involved in any kind of technical leadership role, you will have to do this work eventually. Doing it as early as possible — i.e. in the conversation where someone is describing the plan to you — can save hours or days of wasted implementation. Often this is enough to turn a failed project into a <a href=\"https://www.seangoedecke.com/how-to-ship\" rel=\"noopener noreferrer\">successful one</a>.</p>\n<h3>Working with AI agents<a href=\"https://www.seangoedecke.com/rss.xml#working-with-ai-agents\" rel=\"noopener noreferrer\"></a></h3>\n<p>Maybe you only work with reliable engineers who can be trusted to make all the right decisions. You can thus let them talk without interruption, because it doesn’t really matter if you have a detailed understanding of what they’re saying. That’s great! But these days you probably also have colleagues that are inherently unreliable: AI agents.</p>\n<p>I mostly work with GPT-6-Astra and Claude Opus 5.5. These are good models that don’t often make <em>code</em> mistakes: when they set out to do something, they usually do it<sup><a href=\"https://seangoedecke.com/you-should-all-be-asking-way-more-questions/#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. But they make <em>design</em> mistakes all the time. They assume that two services can talk to each other when in fact they can’t, or they forget about the fact that their code has to run both in the cloud and on-premises, and so on. This is mostly due to the lack of <a href=\"https://www.seangoedecke.com/continuous-learning/\" rel=\"noopener noreferrer\">continuous learning</a>: you didn’t know this stuff on your first day either, but you had plenty of time to pick it up. Language models are always on their first day.</p>\n<p>You should be absolutely <em>peppering</em> AI agents with questions. I constantly ask things like:</p>\n<ul>\n<li>Do we do X elsewhere in this codebase?</li>\n<li>Does service Y really support this type of authentication, or are you assuming we’d have to go build that too?</li>\n<li>Is this subsystem you built necessary to satisfy requirement Z, or does it in fact satisfy some other requirement you assumed?</li>\n<li>Why do we need to update the interface for A?</li>\n<li>Why do we need to touch this file? Isn’t that unrelated to the change?</li>\n</ul>\n<p>There will probably come a day when I always get sensible answers to these questions that convince me the model knows what it’s doing. But today is not that day. About half the questions I ask get answers that convince me<sup><a href=\"https://seangoedecke.com/you-should-all-be-asking-way-more-questions/#fn-3\" rel=\"noopener noreferrer\">3</a></sup> the model has made a mistake: it should have reused the existing subsystem for X, or authed to Y differently, or kept the Z implementation simple, and so on. When this stops happening, I’ll stop asking questions (and try and see if my company will still be willing to pay me to occupy more of an architect role, I suppose).</p>\n<h3>Final thoughts<a href=\"https://www.seangoedecke.com/rss.xml#final-thoughts\" rel=\"noopener noreferrer\"></a></h3>\n<p>You should all be asking way more questions. In other words, <strong>you should all be trying harder to actually understand what you’re hearing</strong>. Don’t trust that the person (or AI model) you’re talking to knows what they’re doing, even if you think they’re smarter than you. <a href=\"https://www.seangoedecke.com/nobody-knows-how-software-products-work/\" rel=\"noopener noreferrer\">Nobody understands</a> complex software products. If you have deep domain knowledge of any area of a codebase, you will routinely find yourself correcting powerful AI models and principal engineers.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Sometimes questions you have will be answered later on, but in my experience this is more about questions like “what are the broader consequences of X”, not the simple confirmation questions I often interrupt to ask.</p>\n<a href=\"https://seangoedecke.com/you-should-all-be-asking-way-more-questions/#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Your experience might be different if you’re working in a different domain or with a different language.</p>\n<a href=\"https://seangoedecke.com/you-should-all-be-asking-way-more-questions/#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I know some of these might sound like leading questions that would trigger sychophancy, but in my experience good coding models are very happy to robustly defend themselves.</p>\n<a href=\"https://seangoedecke.com/you-should-all-be-asking-way-more-questions/#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"d5a44e46ec928440","title":"Grit your teeth and ship it","link":"https://seangoedecke.com/grit-your-teeth-and-ship-it/","author":null,"published_at":"2026-09-20T00:00:00+00:00","content":"<p>Being good at building and being good at shipping are two separate skills. In the short term, they’re actually countervailing: if you have a gift for building, you’re likely to be <em>worse</em> at shipping. Ira Glass has a classic quote about this.</p>\n<blockquote>\n<p>All of us who do creative work, we get into it because we have good taste. But there is this gap. For the first couple years you make stuff, it’s just not that good. It’s trying to be good, it has potential, but it’s not. But your taste, the thing that got you into the game, is still killer. And your taste is why your work disappoints you.</p>\n</blockquote>\n<p>The only way around this is to <strong>grit your teeth and ship it</strong>. You have to force yourself to publish things you’ve made even when you think they’re crap.</p>\n<h3>Programming<a href=\"https://www.seangoedecke.com/rss.xml#programming\" rel=\"noopener noreferrer\"></a></h3>\n<p>Gifted programmers have a <a href=\"https://www.seangoedecke.com/addicted-to-being-useful/\" rel=\"noopener noreferrer\">nearly pathological</a> desire to build elegant, correct, neat systems. That’s what motivates them to learn the arcane details of their languages, or to spend time polishing and refactoring over and over again. But it’s also what makes them reluctant to ship. Any flaws in the software bother them on an emotional level. If they ship with those flaws, they feel like people will think they weren’t paying enough attention to notice them, or that they weren’t good enough to fix them.</p>\n<p>This is annoying when you’re writing software on your own, but it’s completely fatal when you’re working in a tech company. Any large software system is covered in flaws, whether due to time pressure, <a href=\"https://www.seangoedecke.com/bad-code-at-big-companies/\" rel=\"noopener noreferrer\">relative inexperience</a>, <a href=\"https://www.seangoedecke.com/wicked-features/\" rel=\"noopener noreferrer\">wicked features</a>, or a hundred other reasons. Working with it is a process of compromise: of finding the best possible solution given the quirks and foibles of the codebase. In fact, since the most important thing in large codebases is <a href=\"https://www.seangoedecke.com/large-established-codebases/\" rel=\"noopener noreferrer\">consistency</a>, the right thing to do is sometimes to <em>duplicate</em> flaws, assuming they’re not catastrophic.</p>\n<p>Gifted programmers often freeze up. I’ve often seen them retreat to smaller domains where they can safely make the code “correct”: tweaking dev-environment setup, or refactoring tests. Sometimes they just do nothing, and spin in shame and guilt (plus the compounding shame of not achieving anything) until they implode and quit. If they had worse taste, they wouldn’t be as good at programming, but they’d be a lot more useful. You can typically improve a bad diff with time and effort. You can’t improve <em>no</em> diff.</p>\n<h3>Writing<a href=\"https://www.seangoedecke.com/rss.xml#writing\" rel=\"noopener noreferrer\"></a></h3>\n<p>I have a sensitive eye for awkward sentences and uneven prose. That can make writing an unpleasant process: I know what I’m trying to say, but I can’t seem to say it in a way that’s as clear and as elegant as I know is possible. More than half the time I finish drafting a blog post, I look at the post and don’t think it’s very good. But I (mostly) grit my teeth and publish it anyway, because <strong>you have to bias towards shipping</strong>.</p>\n<p>Like any skill, shipping gets easier the more you practice it. If I don’t publish a blog post for a month, I always feel like the next draft is too poorly-written or uninteresting to put out there. But when I’m publishing a post per day, I typically feel great about each draft. When I go back and read my old posts, I can’t tell which ones I felt good about and which ones I felt bad about. There’s no correlation between that and the posts that become <a href=\"https://www.seangoedecke.com/popular/\" rel=\"noopener noreferrer\">popular</a>. Here are some posts I didn’t like as I was writing them but that resonated with my audience:</p>\n<ul>\n<li><a href=\"https://www.seangoedecke.com/the-simplest-thing-that-could-possibly-work/\" rel=\"noopener noreferrer\">Do the simplest thing that could possibly work</a></li>\n<li><a href=\"https://www.seangoedecke.com/software-engineering-may-no-longer-be-a-lifetime-career/\" rel=\"noopener noreferrer\">Software engineering may no longer be a lifetime career</a></li>\n<li><a href=\"https://www.seangoedecke.com/a-little-bit-cynical/\" rel=\"noopener noreferrer\">Software engineers should be a little bit cynical</a></li>\n</ul>\n<p>Here are some posts I thought were pretty good but that didn’t find popularity:</p>\n<ul>\n<li><a href=\"https://www.seangoedecke.com/weak-managers/\" rel=\"noopener noreferrer\">Weak engineering managers</a></li>\n<li><a href=\"https://www.seangoedecke.com/solution-space/\" rel=\"noopener noreferrer\">Paths through the space of all possible solutions</a></li>\n<li><a href=\"https://www.seangoedecke.com/impressing-people/\" rel=\"noopener noreferrer\">Trying to impress people you don’t respect</a></li>\n</ul>\n<p>You just can’t predict what people will find interesting or useful. Producing a high volume of work thus gives much better yield than a small amount of highly-polished work.</p>\n<p>It can be disheartening to realize that some of your most casual, throwaway work will be more successful than the work you slaved over<sup><a href=\"https://seangoedecke.com/grit-your-teeth-and-ship-it/#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. Specifically, it’s disheartening because it means realizing you don’t have <em>control</em> over your own success. You can’t produce something successful by focusing on a single piece until you’re satisfied it’s great. Instead, you just have to do a lot of things and see what sticks. You have to be <a href=\"https://sunilpai.dev/posts/the-senior-engineer-death-spiral/\" rel=\"noopener noreferrer\">momentum-based</a>, not outcome-based. In other words, <strong>you have to grit your teeth and ship it</strong>.</p>\n<p>One common reason to write less is getting overly precious about your ideas. If you think you’ve got a really compelling concept, you don’t want to “waste it” on a poorly-written story. But in fact you can just write about the same thing over and over until you get it right! I have written like thirty blog posts about shipping (this is one of them), or about how tech companies work, or about how internal emotional regulation is as important as technical ability. I expect to continue writing and thinking about these ideas for as long as I find them interesting.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Anthony Burgess famously <a href=\"https://en.wikipedia.org/wiki/A_Clockwork_Orange_(novel)#Writer's_appraisal\" rel=\"noopener noreferrer\">claimed</a> to have “knocked off” <em>A Clockwork Orange</em> in three weeks, and Arthur Conan Doyle considered his largely-forgotten historical novel <a href=\"https://en.wikipedia.org/wiki/Sir_Nigel\" rel=\"noopener noreferrer\">Sir Nigel</a> to be far better than his Sherlock Holmes stories.</p>\n<a href=\"https://seangoedecke.com/grit-your-teeth-and-ship-it/#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"b112ac773817855e","title":"System One models like Jev can train their own replacements","link":"https://seangoedecke.com/system-one-models-can-train-their-own-replacements/","author":null,"published_at":"2026-09-20T00:00:00+00:00","content":"<p>“System One” models like <a href=\"https://www.seangoedecke.com/jev-means-structured-output-is-interesting-again/\" rel=\"noopener noreferrer\">Jev</a> are fast general classifiers. Classifiers have existed since <a href=\"https://en.wikipedia.org/wiki/Mark_I_Perceptron\" rel=\"noopener noreferrer\">1958</a>, but they have to be trained for specific tasks: if you build a classifier to identify images of dogs, it can’t be used to tell you if a streetlight is red, or if a letter is urgent. Like a LLM, Jev can be prompted for a wide variety of tasks, from <a href=\"https://www.youtube.com/watch?v=9oWxrsRo4d8\" rel=\"noopener noreferrer\">sorting email</a> to <a href=\"https://www.seangoedecke.com/two-techniques-for-working-with-system-one-models/\" rel=\"noopener noreferrer\">playing Doom</a>.</p>\n<p>I think models like this are going to be important. There are many tasks that a LLM <em>could</em> do in theory but are too slow and expensive in practice (for instance, reading each new message in Slack<sup><a href=\"https://seangoedecke.com/system-one-models-can-train-their-own-replacements/#fn-1\" rel=\"noopener noreferrer\">1</a></sup> and deciding whether to notify you or not). While you could train a specific classifier for these tasks, there are two main problems with that:</p>\n<ol>\n<li>Despite being a well-understood ML problem, training a bespoke classifier is outside of the skillset of most ordinary engineering teams</li>\n<li>Training a classifier requires assembling a large dataset</li>\n</ol>\n<p>Jev obviously solves the first problem. Any engineering team can plug in a System One model with a prompt like “Based on {list of criteria}, should the user be notified about this message?” But I think it solves the second problem too.</p>\n<p>For serious work, a specific hand-built classifier will always be cheaper and faster than Jev. Generic classifiers have to encode knowledge of all kinds of irrelevant things in their weights, so they can address lots of different tasks. That makes them larger, slower, and more expensive to run. Fortunately, <strong>it is going to be surprisingly easy to replace a Jev instance with a hand-built classifier.</strong></p>\n<p>Once you’re satisfied with how your Jev classifier is performing — presumably you’ve spent days tweaking the prompt — you can trivially collect its input and output data. In the Slack notifier case, that’d be the Slack message (plus any context) and the ultimate decision to notify or not. Once you’ve saved enough data, you’ll be able to train your own classifier on that data<sup><a href=\"https://seangoedecke.com/system-one-models-can-train-their-own-replacements/#fn-2\" rel=\"noopener noreferrer\">2</a></sup>.</p>\n<p>It won’t be a general classifier like Jev, but it should do well on the specific task and be much faster. Of course it’ll require some ML expertise, but it should be easier to develop (or rent) that expertise once you’ve validated that the feature is worth building.</p>\n<p>In other words, because Jev has to be prompted for specific tasks, it should be easy to <a href=\"https://en.wikipedia.org/wiki/Knowledge_distillation\" rel=\"noopener noreferrer\">distil</a> any successful Jev usage into a specific classifier. If System One models take off — and I hope they do — I expect this to be a common pattern.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>As I write this, I’m imagining ways you could poll and batch to do this with LLMs. Substitute “instantly notify” or some higher-volume event source if you’d prefer a different example.</p>\n<a href=\"https://seangoedecke.com/system-one-models-can-train-their-own-replacements/#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>You could annotate a bunch of data with LLMs already, without using Jev, but this is a pretty expensive step to take when you aren’t sure the feature is going to work.</p>\n<a href=\"https://seangoedecke.com/system-one-models-can-train-their-own-replacements/#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"586422309a76edcd","title":"Two techniques for working with System One models","link":"https://seangoedecke.com/two-techniques-for-working-with-system-one-models/","author":null,"published_at":"2026-09-18T00:00:00+00:00","content":"<p>I recently wrote about <a href=\"https://www.seangoedecke.com/jev-means-structured-output-is-interesting-again/\" rel=\"noopener noreferrer\">Jev</a>, a new “System One” language model that only outputs <em>decisions</em>: the answers to a set of user-provided multiple-choice questions. This means it’s nowhere near as flexible<sup><a href=\"https://seangoedecke.com/two-techniques-for-working-with-system-one-models/#fn-1\" rel=\"noopener noreferrer\">1</a></sup> as a traditional LLM like ChatGPT, but in return it’s consistently fast.</p>\n<p>We don’t know exactly how Jev works. I’ve seen people say diffusion, or various tweaks to the Transformer architecture, or some entirely new type of model. But that doesn’t matter. Like I argued <a href=\"https://www.seangoedecke.com/jev-means-structured-output-is-interesting-again/#structured-output-can-already-be-fast\" rel=\"noopener noreferrer\">here</a>, it isn’t hard to turn any LLM into a System One model. By batching prompts that generate a single token with structured output, you get a consistently fast general-purpose classifier. I vibed up a basic version to play with <a href=\"https://github.com/sgoedecke/system-one/tree/main\" rel=\"noopener noreferrer\">here</a> in ~150 lines of Python (most of which is error handling).</p>\n<p>Note that <em>this doesn’t require changing the model</em>. As long as you have access to the logits (for structured outputs) and can prefill data into the prompt, you can turn any LLM into a general fast classifier. What’s it like to program with one of these? While wiring up the demos for my library, I learned two techniques that I want to write about: setting tiered goals and tournament choice sampling.</p>\n<h3>Doom<a href=\"https://www.seangoedecke.com/rss.xml#doom\" rel=\"noopener noreferrer\"></a></h3>\n<p>Here’s Qwen3-8B playing Doom:</p>\n<video controls=\"controls\" preload=\"metadata\">\n  <source src=\"https://www.seangoedecke.com/48c52119c5ad14667a9541457c684db7/doom-qwen3-8b.mp4\" type=\"video/mp4\">\n</video>\n<p>If you compare this to the <a href=\"https://github.com/sgoedecke/system-one/blob/main/docs/demos/doom-qwen3-8b-tool-agent.mp4\" rel=\"noopener noreferrer\">video</a> of the same model playing Doom with regular tool calls, it’s clear that the System One version of the model is doing more things and reacting more quickly. The tool-calling model makes one decision every 600ms or so, while the System One model makes six or seven batched decisions every 190ms<sup><a href=\"https://seangoedecke.com/two-techniques-for-working-with-system-one-models/#fn-2\" rel=\"noopener noreferrer\">2</a></sup>:</p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/70f6abb1474ff212395a43c8a17194f8/eb2ef/turns.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"turns\" src=\"https://www.seangoedecke.com/static/70f6abb1474ff212395a43c8a17194f8/fcda8/turns.png\" title=\"turns\">\n  </a>\n    </span></p>\n<p>Both Qwen3-8B and Jev are text-only models, so both demos require a step where we translate the game state into text. However, it’d be trivial to support image (or audio) input by choosing a multimodal LLM.</p>\n<h3>Goals and sub-goals<a href=\"https://www.seangoedecke.com/rss.xml#goals-and-sub-goals\" rel=\"noopener noreferrer\"></a></h3>\n<p>What’s interesting about implementing the Doom demo is that <strong>just supplying the game inputs as choices doesn’t work very well</strong>. A single forward pass — 200ms — is enough time to react to the current game state, but doesn’t bring enough compute to bear to derive the current short-term goal (e.g. “kill this enemy”, “collect this item”) and choose to follow it. When I wired it up that way, the model held down the “shoot” button 100% of the time (why not, I guess) and just aimlessly wandered around the level. </p>\n<p>The fix is to periodically ask the model to choose between a fixed set of short term goals (e.g. “collect armor”, “kill enemies”) and then include that goal in the regular every-200ms prompt. If you look at the Doom video in the Jev demo, you can see that they’re doing exactly that. As soon as I did it as well, my model started playing in a more human-like way.</p>\n<p>This is an interesting technique for working with System One models. In a way, it’s the equivalent of regular LLM reasoning, since it provides a way to use more compute on the same problem. I can imagine a real-time system that manages several layers of goals in this way:</p>\n<ol>\n<li>An every-ten-second loop that sets an overall strategic goal</li>\n<li>An every-five-second loop that sets a tactical subgoal based on (1)</li>\n<li>An every-second loop that breaks down the current tactical subgoal into specific targets</li>\n<li>A tight inner loop that runs as fast as possible (e.g. every 100ms) that controls which actual inputs are activated</li>\n</ol>\n<p>The general structure here should be pretty familiar to anyone who’s worked in game or robotics AI. In theory you could replace (1) with an actual LLM, and have that generate the lists of options for steps (2) and (3). In practice I suspect this will be tricky to get right, and it’ll be better to just write down a list of all possible goals ahead of time. This would work just fine for game-playing and well-understood tasks.</p>\n<h3>Wikiracing<a href=\"https://www.seangoedecke.com/rss.xml#wikiracing\" rel=\"noopener noreferrer\"></a></h3>\n<p>I also reimplemented the Wikiracing demo from the Jev <a href=\"https://typesafe.ai/blog/introducing-system-one-models-and-jev\" rel=\"noopener noreferrer\">launch post</a>, where the model has to start at the Wikipedia page for “baseball” and navigate as quickly as possible to the Wikipedia page for “sun”. You can watch the video for that <a href=\"https://github.com/sgoedecke/system-one#wikipedia-race\" rel=\"noopener noreferrer\">here</a>, though it’s less impressive than the Doom demo. </p>\n<p>The difficulty with the Doom demo is getting the model to loop quickly enough and to commit to short-term plans. For Wikiracing, the difficulty is <em>scale</em>: the Wikipedia page for “baseball” has over a thousand internal links. Jev only supports 255 choices for a single question, and my hacked-together System One layer was similar. While it technically would scale out to more choices, it stopped working well<sup><a href=\"https://seangoedecke.com/two-techniques-for-working-with-system-one-models/#fn-3\" rel=\"noopener noreferrer\">3</a></sup> after a hundred or so.</p>\n<p>Jev’s approach here is to do “a 2 stage-system of scoring independently then making an explicit choice”. This did not work very well for me at all. I think here Jev is benefiting from the fact that it’s specifically trained to give confidence estimates. Qwen3-8B gave a few hundred of the links the same top score, which wasn’t helpful. It ended up taking multiple minutes to find a thirty-or-forty link path between the two pages.</p>\n<p>What I tried instead was <strong>tournament sampling</strong>: I fed a hundred links at a time into each choice, then did a second pass with the chosen links. This worked <em>great</em>. The model found the ideal three-link path (if you’re curious, “baseball”/“scientific american”/“amateur astronomy”/“sun”). I recommend this pattern if you’re trying to find the best option among many choices. Ordinary LLMs are way better at relative judgements than absolute ratings.</p>\n<h3>Conclusion<a href=\"https://www.seangoedecke.com/rss.xml#conclusion\" rel=\"noopener noreferrer\"></a></h3>\n<p>I remain optimistic about the potential of System One models — fast general classifiers — to build AI systems that aren’t just chatbots. It feels like this is a meaningful alternative to tool calls for realtime scenarios or use-cases where you need predictable inference timing. Just as generic LLMs often outperform domain-specific models, I think it’s likely that generic System One models will sometimes outperform domain-specific classifiers (though they will always be larger and slower).</p>\n<p>I do think the big labs are definitely going to try and compete by releasing a choice-only version of their small, fast models. If Jev gets any traction, we will soon see a System One Terra and a System One Haiku, and we will certainly see “real” versions of my vibed up System One <a href=\"https://github.com/sgoedecke/system-one\" rel=\"noopener noreferrer\">library</a>. We should start working out the best way to write programs with these models now.</p>\n<p>You can think of System One models as general-purpose classifiers. Instead of having to train a new classifier per-task, you can use a System One model. It’ll be bigger and slower than a custom classifier model, but far more flexible, and you can tweak it via adjusting the prompt instead of having to re-train the model.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Technically you can give it the multiple-choice question of “which letter comes next” to make it act like a normal autoregressive LLM, but that wouldn’t really work. </p>\n<a href=\"https://seangoedecke.com/two-techniques-for-working-with-system-one-models/#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I started on a 4090, which was able to make decisions every 500ms, but that wasn’t really quick enough for Doom. I could probably have optimized it further but instead I just rented an H100 for ten minutes to record the demo, which got it down to a 190ms loop. I recorded the tool-calling Doom demo on the H100 too, so it’s a fair comparison.</p>\n<a href=\"https://seangoedecke.com/two-techniques-for-working-with-system-one-models/#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>There’s an interesting research question here about how to implement choices like this, since they have to be predictable by a single token. I started with indexes but found “labels” (just picking some token to associate with the choice) performed <em>way</em> better on Wikiracing (though not Doom). How many choices do you have to have before labels are better than indexes? Of course, you could alter the model to directly output the choice, but I like the idea that you can do all of this in the inference code for any LLM.</p>\n<a href=\"https://seangoedecke.com/two-techniques-for-working-with-system-one-models/#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"d0e76c83db88dcf1","title":"Jev means structured output is interesting again","link":"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/","author":null,"published_at":"2026-09-16T00:00:00+00:00","content":"<p>I don’t write blog posts about new models. That’s <a href=\"https://simonwillison.net/\" rel=\"noopener noreferrer\">Simon Willison’s</a> beat, and he’s very good at it. But I want to write about <a href=\"https://typesafe.ai/blog/introducing-system-one-models-and-jev\" rel=\"noopener noreferrer\">Jev</a>, which is a different kind<sup><a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fn-1\" rel=\"noopener noreferrer\">1</a></sup> of AI model: a “System One”<sup><a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fn-2\" rel=\"noopener noreferrer\">2</a></sup> model. As it turns out, it’s not <em>that</em> different from an ordinary LLM with structured output, but the interface it uses is very cool and I hope it becomes more widespread.</p>\n<h3>How Jev is different from LLMs<a href=\"https://www.seangoedecke.com/rss.xml#how-jev-is-different-from-llms\" rel=\"noopener noreferrer\"></a></h3>\n<p>Ordinary LLMs take in some human-language prompt and produce some human-language output. They do so <em>autoregressively</em>: first they produce one token, then the next, then the next, and so on.</p>\n<div><pre><code>User: Who invented the sandwich?\nLLM: The sandwich was invented by the Earl of Sandwich.</code></pre></div>\n<p>This makes them extremely flexible, since they can do literally anything a computer can do. But it also makes them slow and weird. Slow, because they have to run a whole new generation pass per-token, and weird, because the space of human language is so broad that you can get <a href=\"https://community.openai.com/t/is-chatgpt-hacked-major-chinese-gambling-advertisements-websites-showing-up-in-output/1375128\" rel=\"noopener noreferrer\">really odd behavior</a> from a model trained on it.</p>\n<p>Jev takes a human-language prompt, but it does not produce human-language output. It only produces structured output.</p>\n<div><pre><code>User: { state: \"What color is the sky?\", choices: [\"blue\", \"red\", \"yellow\"] }\nJev: { answer: \"blue\" }</code></pre></div>\n<p>So far, so ordinary: LLMs <a href=\"https://developers.openai.com/api/docs/guides/structured-outputs\" rel=\"noopener noreferrer\">do this already</a>. But it turns out that if you build a model that <em>only</em> produces structured output, you get some interesting and desirable properties.</p>\n<h3>Jev is consistently fast<a href=\"https://www.seangoedecke.com/rss.xml#jev-is-consistently-fast\" rel=\"noopener noreferrer\"></a></h3>\n<p><strong>Jev is always really fast.</strong> The fastest response time is around 70ms instead of a couple of seconds for normal LLMs. Even better, the <em>slowest</em> response time is only 500ms. Because Jev only does structured output, it isn’t autoregressive: it can produce answers to many questions in parallel in a single forward pass. When a LLM is producing structured output, it has to produce the tokens ”{”, ” ”, “answer”, ”:”, and so on with successive forward passes<sup><a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. Jev does it all in one go.</p>\n<p>The most compelling example of Jev’s speed is that <strong>the model can play Doom</strong>. You can feed a text-based representation of the current game state into the model, combined with a set of choices like “should the trigger be held down”, “what should the current goal be”, “given that the current goal is X, what keyboard input should be pressed”, and so on, and it works — latency is low enough and the system is smart enough that the model plays well in real time.</p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/d54381912237e0f153a14473f60d04e9/5bef7/doom.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"doom\" src=\"https://www.seangoedecke.com/static/d54381912237e0f153a14473f60d04e9/fcda8/doom.png\" title=\"doom\">\n  </a>\n    </span></p>\n<p>Of course you could train a neural net to play Doom already. But Jev is a <em>general</em> intelligence: just like LLMs can do your taxes, perform mathematics research, fix your Python environment, and write you a poem, Jev can do many other tasks besides playing a single video game. Current LLMs can play Doom too (albeit slowly). But as Nelson Elhage <a href=\"https://blog.nelhage.com/post/reflections-on-performance/#performance-changes-how-users-use-software\" rel=\"noopener noreferrer\">famously said</a>, fast software doesn’t just mean we can do the same tasks faster, it means we can do entirely new kinds of tasks. What kinds of new programs can we write by injecting 100ms worth of dirt-cheap intelligence at various decision points?</p>\n<p>To me, this is the most exciting thing about Jev. <em>Fast</em> structured output could be a genuinely new computational primitive for intelligence. So far we’ve built a lot of programs on top of autoregressive token generation, and they all look like fancy chatbots. Leaning hard into structured output might conceivably unlock a bunch of non-chatbot use cases for AI.</p>\n<h3>Structured output can already be fast<a href=\"https://www.seangoedecke.com/rss.xml#structured-output-can-already-be-fast\" rel=\"noopener noreferrer\"></a></h3>\n<p>My biggest problem with Jev is that I think <strong>fast structured output is already available</strong>. Structured output from LLMs is only slow because (a) nobody really cares about it<sup><a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fn-4\" rel=\"noopener noreferrer\">4</a></sup>, and (b) the people who do care about it want big JSON blobs, so it’s typically implemented with <a href=\"https://www.aidancooper.co.uk/constrained-decoding/\" rel=\"noopener noreferrer\">“grammar-constrained decoding”</a>: the LLM outputs autoregressively as normal, but the logit sampler discards tokens that don’t fit the structured output (e.g. if there hasn’t been a ”[”, you can’t output a ”]”).</p>\n<p>If you want fast, parallelized structured output against limited choices, you don’t strictly need to do autoregressive generation at all. You can simply prefill the response with <code>\"choice\": \"</code> and generate one token<sup><a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fn-5\" rel=\"noopener noreferrer\">5</a></sup>, restricted to the user-provided choices. Since LLMs ingest all input tokens in parallel, this is way faster than generating the entire structured output. Multiple choices can be batched into the same forward pass via ordinary inference batching. This doesn’t let you do long-form structured output, but in return you get most of<sup><a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fn-6\" rel=\"noopener noreferrer\">6</a></sup> Jev’s “secret sauce”: the speed, the consistency, and the parallelism of a System One model.</p>\n<p>People have <a href=\"https://x.com/harshagundal/status/2100044305536889015?s=20\" rel=\"noopener noreferrer\">already started trying this</a> after today’s Jev announcement, and it seems like it’s working OK<sup><a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fn-7\" rel=\"noopener noreferrer\">7</a></sup>. In other words, I suspect Jev does not have a substantial technical moat, and their claimed “Reinforcement Learning for Calibrated Decisions” is not a brand-new scaling axis. It will probably be pretty easy for any other lab to replicate, or for individual programmers to retrofit existing open-source LLMs into a fast Jev-like model.</p>\n<p>However, I suspect Jev is still going to be better than most versions of “Qwen-32B-System-One” or whatever. Being able to fine-tune or optimize the model on just structured output is probably a meaningful advantage.</p>\n<h3>Intelligence and hallucinations<a href=\"https://www.seangoedecke.com/rss.xml#intelligence-and-hallucinations\" rel=\"noopener noreferrer\"></a></h3>\n<p>I doubt Jev is ever going to be as smart as frontier LLMs. Not being able to use test-time compute at all<sup><a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fn-8\" rel=\"noopener noreferrer\">8</a></sup> is a big disadvantage, and will likely cap this kind of model around the strength of non-reasoning LLMs. In practice this shouldn’t matter too much for low-latency applications, but you shouldn’t see this as a new scaling axis or a way to produce more intelligent models.</p>\n<p>Jev’s developers claim it is immune from hallucinations. To me, this seems like a semantic dodge, since Jev can absolutely still pick the wrong choice (e.g. calling the sky “red”). I suppose that’s technically just a <em>mistake</em>, since the model is picking a user-provided choice instead of inventing something new out of whole cloth. Still, all of this is also true about regular LLMs with structured outputs, and it doesn’t make Jev any more reliable in practice.</p>\n<h3>Conclusion<a href=\"https://www.seangoedecke.com/rss.xml#conclusion\" rel=\"noopener noreferrer\"></a></h3>\n<p>It’s unclear to me how much of Jev’s value is in the model itself, compared to the inference strategy of only generating one token per question. The data and demos in the announcement look to me like they could have been generated by plugging any Terra-sized model into a single-token inference stack. However, the people involved are credible, and I’m sure the model is good — I just wish they’d provided some comparisons that didn’t force the LLM to unnecessarily produce a blob of JSON token-by-token.</p>\n<p>Overall, I am happy that Jev exists and I hope it succeeds. I hope we do see some real competition in the fast-structured-output space, and that it motivates the big labs to release official versions of their own models that are fine-tuned for this. GPT-5.6-Terra-System-One would be a very interesting model to build AI products on top of.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>I did write about Thinking Machines’ <a href=\"https://www.seangoedecke.com/interaction-models/\" rel=\"noopener noreferrer\">“interaction models”</a>, which are also a fast-enough-to-be-meaningfully-different paradigm for AI inference.</p>\n<a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>They call Jev a “System One” LLM, after Daniel Kahneman’s <a href=\"https://www.gilesd-j.com/2023/03/30/reproducibility-thinking-fast-and-slow/\" rel=\"noopener noreferrer\">partially discredited</a> <em>Thinking Fast and Slow</em>, where he divides human cognition into a lightning-fast System One and a slow-and-reflective System Two.</p>\n<a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>If you’re thinking “wait, couldn’t you just aggressively prefill a regular LLM and only produce one constrained token”, keep reading.</p>\n<a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Not counting tool calls, which are built-in in a way that structured output isn’t.</p>\n<a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>What if some of the user’s choices are longer than a single token? I haven’t tried this myself, but I’m sure you could translate them into a single token, or train the model to output “1/2/3” under the hood instead of the choice content, or generate only the first token of the choice if it’s different, or some other clever trick I haven’t thought of.</p>\n<a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Jev claims that their generated probabilities are “calibrated”, but I haven’t seen anything to suggest that these aren’t just regular logit probabilities. Maybe there’s some clever training they do to encourage accurate logprobs in uncertain situations (e.g. getting the model to produce <code>heads: 50, tails: 50</code> when predicting a coinflip, etc)? If so, I wish they’d written more about that in the announcement.</p>\n<a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I tried it myself with <code>Qwen2.5-1.5B-Instruct</code> and got a 2x-3x speedup compared to non-prefixed structured output.</p>\n<a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I suppose they could do some looped-transformer thing where they loop some fixed amount of times, but anything that looks like reasoning would make the model latency slow and unpredictable, defeating the entire purpose.</p>\n<a href=\"https://seangoedecke.com/jev-means-structured-output-is-interesting-again/#fnref-8\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"082b48f305bf33b6","title":"Tell agents the why, not just the how","link":"https://seangoedecke.com/tell-agents-the-why/","author":null,"published_at":"2026-09-15T00:00:00+00:00","content":"<p>Early AI agents were basically enthusiastic idiots. Working with them required you to tell them precisely what you wanted them to do (for instance, “method A exists on class B, please add an equivalent method to classes C through F”). Otherwise they’d go off and do entirely the wrong thing. But as AI agents have improved, this has changed.</p>\n<p>When frontier models go off and do the wrong thing today, they don’t do it because they’re confused, they do it because they make an incorrect assumption about your goals or priorities. For instance, when GPT-6-Astra thinks it’s writing code for itself, it will produce <a href=\"https://lucumr.pocoo.org/2026/9/7/astra-why/\" rel=\"noopener noreferrer\">minified code</a>. It’s perfectly capable of writing human-readable code — at least in Golang, where I’ve produced several thousand lines of acceptable code with the model — but you have to tell it that humans will be reading the code<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>.</p>\n<p>This is the main piece of advice I want to give most people I see prompting agents: give the agent context on your priorities, not just on the specific task you want them to do. Here’s a prompt I recently used as the starting point for <a href=\"https://www.seangoedecke.com/deckard/\" rel=\"noopener noreferrer\">Deckard</a>.</p>\n<blockquote>\n<p>Hello. You should have Runpod access via MCP (if not, tell me and I’ll fix it).</p>\n<p>I have the long-term goal of building a local program or browser extension to automatically scan pages I load for AI content and hide it. I have the short-term goal of figuring out the best AI detection model I can run on my macbook without killing my battery or making it hot, and (relatedly) figuring out how to run the model most efficiently. My guess is that Pangram’s EditLens 3B or the smaller Roberta model might be a good place to start, though they might require quantizing and will definitely require some work to make them run as efficiently as possible on my macbook.</p>\n<p>I would like you to use my Runpod account to start answering these questions. Eventually we’ll move to doing things on this macbook pro, but my hope is that Runpod can help with some experiments that are too hot/long/slow to run locally. You are a smart model; if you can see a better way to achieve my goals, please let me know and we’ll talk about it. Good luck.</p>\n</blockquote>\n<p>About half of this prompt is sharing broad context, such as the overall project I’m aiming for, the fact that it’s for me personally and not for work, and my priorities (e.g. keeping the laptop cold). If I had written an explicit spec, I would have missed a bunch of improvements: for instance, using native messaging for the local model, or choosing the Gradient model instead of EditLens.</p>\n<p>I do the same thing for work, but typically with a stronger emphasis on my <em>technical</em> values. I often write a paragraph spiel explaining the relative priorities of avoiding bugs, observability, fitting elegantly into the current code, performance, and so on. Note that I said “relative” priorities: I don’t simply list all of these things and say they’re important, I explicitly tell the model which of those I care less about and can therefore be traded off to better achieve the others.</p>\n<p>Models are now smart enough to have meaningful input on your broader goals. If you’re just prompting them with a concrete technical spec, you are committing the same mistake as in the <a href=\"https://xyproblem.info/\" rel=\"noopener noreferrer\">XY problem</a>: asking expert advice without giving the expert the context it needs.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Incidentally, you don’t have to tell it to write human code if it’s working in a human-authored codebase. It’s smart enough to pick up the style of the surrounding code.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"6ed524e4684357b5","title":"Slow developer experience will bottleneck fast models","link":"https://seangoedecke.com/slow-devex-will-bottleneck-fast-models/","author":null,"published_at":"2026-09-14T00:00:00+00:00","content":"<p>Right now developer experience is measured in seconds. If your tests take a second to run, that’s good; if they take thirty seconds, that’s bad. Any faster than a second doesn’t really matter, because most of your time is spent either thinking or waiting for an AI agent to spin. Shaving milliseconds off your dev server reload time or whatever is pointless: that’s not the bottleneck.</p>\n<p>It will be. Small models are getting faster and faster, and smart models are getting smaller. I think most engineers will still want to use the smartest available model — software engineering is hard — but we will increasingly see faster models get used as subagents or for well-understood tasks. This is largely uncharted territory. Very few people have developed intuitions for what it is going to be like to work with agents that run at thousands of tokens-per-second.</p>\n<p>GPT-6-Astra can run at about <a href=\"https://openrouter.ai/openai/gpt-6-astra#providers\" rel=\"noopener noreferrer\">sixty</a> tokens per second. That means you spend a lot of time waiting for it to think. You work with it like you would work with another human: delegating a task and then context-switching until that task is complete. If you haven’t yet, have a play around with <a href=\"https://chatjimmy.ai/\" rel=\"noopener noreferrer\">Jimmy</a>, Taalas’ version of LLaMA-3.1-8B running<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup> at <em>seventeen thousand tokens per second</em>. No matter how long the response is, it arrives in the instant of you hitting send. The model is not good enough for agentic work, but it gives a glimpse of what it would be like: you would simply get your answer instantly<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>.</p>\n<p>Well, that’s assuming the agent’s tool calls are fast. When generating tokens is not the bottleneck, it will suddenly matter a lot whether it can read a file in 100ms vs 10ms, or whether it can run your tests in 500ms vs two seconds. Fast tool calls are going to be the difference between a near-instant response and having to wait several minutes. There is thus going to be enormous pressure to do agentic coding in languages with fast compilers and tests, like Golang, and to tightly optimize the dev loop in agentic codebases.</p>\n<p>Teams focused on DevEx — developer experience — are largely a relic of the 2010s, when companies were <a href=\"https://www.seangoedecke.com/good-times-are-over/\" rel=\"noopener noreferrer\">incentivized</a> to make their engineers happy. Most companies have cut them down to a skeleton crew or removed them entirely. But we may see a return of DevEx in the late 2020s, focused on speeding up the experience for AI agents.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Like most ultra-fast inference, it relies on fitting the entire model onto a huge GPU-like chip that’s specially designed to run inference: for Taalas, it’s in the silicon itself; for Cerebras and Groq, it’s in giant embedded onboard memory units.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Will AI providers simply train models to spend more time reasoning, so users would have to wait for roughly the same time? I doubt it. Most ordinary engineering problems would not be solved better by spending an extra million tokens thinking. You only need to do that when you’re pushing right up against the limits of the model.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"f57e6cd71dd0aba3","title":"AI is breaking our proxies for expertise","link":"https://seangoedecke.com/ai-is-breaking-our-proxies-for-expertise/","author":null,"published_at":"2026-09-13T00:00:00+00:00","content":"<p>Mathematicians are broadly not anti-AI. They’re more culturally open to using AI as a tool than, say, artists or writers<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. However, now that more and more <a href=\"https://openai.com/index/navier-stokes-solution/\" rel=\"noopener noreferrer\">genuinely</a> <a href=\"https://www.anthropic.com/research/riemann-zeta\" rel=\"noopener noreferrer\">prestigious</a> problems have fallen to AI, that might be changing. Almost five thousand mathematicians (including twenty-five Fields medalists) have signed a declaration called <a href=\"https://mathandai.org/\" rel=\"noopener noreferrer\"><em>A Severe Misaligment of AI in Mathematics</em></a>. The core argument goes something like this:</p>\n<blockquote>\n<p>In recent months, the success of AI in solving major mathematical problems has made headlines even outside mathematical circles. But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal. Indeed, the mass production at faster and faster pace of “true/false” statements could destroy fertile ground instead of breathing life into new ideas.</p>\n</blockquote>\n<p>A lot of people online have interpreted this as the expected complaint from any field that gets automated: translators did it, artists and programmers have been doing it, and now it’s the turn of the mathematicians. I think this is too dismissive. Understanding the concrete problem mathematicians are upset about can help us better understand the impact of AI on our own fields, and what we’ll have to do about it.</p>\n<h3>Puzzle-solving and idea-generating<a href=\"https://www.seangoedecke.com/rss.xml#puzzle-solving-and-idea-generating\" rel=\"noopener noreferrer\"></a></h3>\n<p>There are two types of mathematics. Most people are familiar with the first, which we might call “puzzle-solving”: you take a problem and try to find a solution to it. When you’re a student, these problems are typically easy, like simplifying some algebraic expression. When you’re a researcher, these problems can be nearly impossible, like proving <a href=\"https://en.wikipedia.org/wiki/Fermat%27s_Last_Theorem\" rel=\"noopener noreferrer\">Fermat’s Last Theorem</a>. Puzzle-solving is easy to understand but hard to do, which makes it impressive to non-mathematicians, which makes it highly prestigious. In other words, puzzle-solving is <a href=\"https://www.seangoedecke.com/seeing-like-a-software-company\" rel=\"noopener noreferrer\"><em>legible</em></a>.</p>\n<p>The second type of mathematics is “idea-generating”: coming up with new ways of thinking about mathematics, and thus new terms or concepts. For examples of these, just glance down the list of <a href=\"https://arxiv.org/list/math.FA/recent\" rel=\"noopener noreferrer\">arXiv mathematics papers</a>. “Hardy spaces”, “Schatten exponent”, “Banach lattices” and so on are all concepts someone thought was interesting. This work is largely unimpressive to non-mathematicians, because nobody really knows if the concepts you come up with are particularly difficult or insightful. For instance, I have just generated the concept of a “Goedecke set”, which is the set of all natural numbers whose digits add up to a prime number. Who cares? The categories we want are the <a href=\"https://plato.stanford.edu/entries/natural-kinds/\" rel=\"noopener noreferrer\">“natural kinds”</a> of mathematics — the concepts that “carve nature at its joints” — and it’s almost impossible to tell what those are without years or decades of hard work.</p>\n<p>How are the two types of mathematics related? We might say<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> that generating ideas is the real intellectual work of mathematics. Puzzle-solving is important instrumentally: to identify which ideas can be used to answer longstanding questions, and thus which ideas are worthwhile. Over time, those worthwhile ideas become better understood and easier to use, until they reach the point where they can be used to advance science in general. Eventually the ideas become so well-understood that they can be taught to children: “zero”, “negative numbers”, “imaginary numbers” and “calculus” were all once rarefied mathematical ideas, but are now concepts we’d expect any precocious twelve-year-old to grasp.</p>\n<p>There’s another, more prosaic purpose of puzzle-solving: to make mathematical skill and progress legible to outsiders. I can’t appreciate Terence Tao’s mathematical work, but I know what a Fields Medal is. I don’t have a good intuitive sense of what a Galois representation is, but I know about the proof of <a href=\"https://en.wikipedia.org/wiki/Wiles%27s_proof_of_Fermat%27s_Last_Theorem\" rel=\"noopener noreferrer\">Fermat’s Last Theorem</a>. We might say that puzzles like this have served as a way to indirectly reward skilled mathematicians for their more important idea-generating work (or for conclusively demonstrating<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup> that the ideas used in the proof are useful).</p>\n<h3>Do AI proofs undercut idea generation?<a href=\"https://www.seangoedecke.com/rss.xml#do-ai-proofs-undercut-idea-generation\" rel=\"noopener noreferrer\"></a></h3>\n<p>AI proofs undercut both of these purposes. I can now lay out precisely why I think mathematicians are so unhappy:</p>\n<ol>\n<li>Puzzles serve as a high-legibility, high-reward target for mathematicians</li>\n<li>To solve these puzzles, new ideas must typically be generated; the puzzle’s solution serves as evidence that the ideas are useful</li>\n<li>But now AI can solve many of these targets “the hard way”, without generating intuitive new ideas</li>\n<li>This undercuts both ways puzzle-solving supports idea-generation: AI companies claim the prestige while not meaningfully advancing mathematical progress</li>\n<li>This is bad for mathematics as a whole, because puzzle-solving is ancillary to the real goal of mathematics</li>\n</ol>\n<p>This is kind of like <a href=\"https://en.wikipedia.org/wiki/Goodhart%27s_law\" rel=\"noopener noreferrer\">Goodhart’s Law</a>. Puzzles were a useful, impossible-to-game measure for mathematical progress. But now that AI companies can game that measure (by solving them in a way that’s inaccessible<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup> to humans), the whole point of those puzzles disappears.</p>\n<p>Are the mathematicians right? I think it’s broadly unclear whether (3) is true: i.e. whether frontier AI models aren’t generating or can’t generate new mathematical ideas. We’re still in the very early days of AIs solving our hardest mathematical problems. Who knows what they’re going to be capable of? I give basically zero credence to the idea that AIs are incapable of this because of some intrinsic feature of how LLMs work. For the last three years, we’ve seen people claim that LLMs are intrinsically incapable of X, only to have LLMs excel at X a few months later.</p>\n<p>Even granted that (3) is true, there’s still work to be done for human mathematicians in building the conceptual machinery that can make AI-generated proofs accessible to humans: i.e. in generating a “human proof” to go alongside the existing “AI proof”. In fact, I’d expect the existence of an AI proof to help with this. If you know proposition X is true, it’s easier to figure out why, because you’re not constantly worried you’re wasting your time. For more on this, I recommend Gwern’s blog <a href=\"https://gwern.net/on-really-trying\" rel=\"noopener noreferrer\"><em>On Really Trying</em></a>, where he quotes a series of instances where simply being told that a solution exists is enough of a clue to help people find it.</p>\n<h3>Mathematics, chess, and speedrunning<a href=\"https://www.seangoedecke.com/rss.xml#mathematics-chess-and-speedrunning\" rel=\"noopener noreferrer\"></a></h3>\n<p>Of course, there’s a prestige and motivation problem. “I’m the first person to solve Navier-Stokes” is a much more compelling target than “I figured out a better way to explain the AI solution to Navier-Stokes”, and it’s much easier to award prizes for. Will mathematicians bother to work on problems that have already been solved? I think so.</p>\n<p>To see why, we can look at other domains where AI has come in and outcompeted the best humans, such as chess or video game speedrunning. I can run a chess program on my phone that will beat Magnus Carlsen 100-0. Computer programs — called “tool-assisted speedruns” or “TAS” — can finish any video game much faster than even the fastest human. But in both of these areas, humans still compete in human-only leagues, and there’s still prestige attached to the most capable humans. It’s possible that mathematics ends up in this kind of state, where “human mathematics” and “AI mathematics” exist in largely separate spheres, and the first “human” solution to a mathematical problem can still earn acclaim.</p>\n<p>In fact, in both of those areas, the presence of inhumanly strong computer players has improved the human game. Despite many computer chess moves being basically incomprehensible to humans, top chess players have <a href=\"https://bpspsychub.onlinelibrary.wiley.com/doi/10.1111/bjop.12750?af=R\" rel=\"noopener noreferrer\">learned from</a> the computer “style”. In speedrunning, many moves once considered “TAS-only” are now performed by humans. AI mathematics might likewise improve human mathematics.</p>\n<h3>Software engineering<a href=\"https://www.seangoedecke.com/rss.xml#software-engineering\" rel=\"noopener noreferrer\"></a></h3>\n<p>I am not a mathematician. I did major in mathematics during undergrad, and I have fond memories of proofs from <a href=\"https://handbook.unimelb.edu.au/subjects/mast20026\" rel=\"noopener noreferrer\">real</a> and <a href=\"https://handbook.unimelb.edu.au/2024/subjects/mast30021\" rel=\"noopener noreferrer\">complex</a> analysis, but it’s not even close to my field. However, I am watching the effects of powerful AI on mathematics very closely, since my own field — software engineering — is being colonized by AI agents in the same way.</p>\n<p>The field of software engineering does not have the same structure as mathematics. We write code to make money, not to earn prestige or advance the frontier of human knowledge. But AI is undercutting the traditional avenues for prestige in software engineering as well. It used to be that you could put a meaty project on your GitHub — say, an emulator, or a toy OS — and people would know you were a skilled engineer. But now projects like that are worthless, because everyone just assumes they’re vibe-coded. We used to tell stories about engineers who would disappear and rewrite a system over the weekend, or produce thousands of lines of code a day. Now anyone can do that with an OpenAI subscription.</p>\n<p>Like mathematics, software engineers are going to have to rebuild our cultural sense of the kind of work we value. We are either going to have to silo “AI work” off from “human work” like chess, or to find some legible human skills to recognize that can’t be easily counterfeited by AI. In the meantime, a lot of people who were successful in the old world are going to be very unhappy.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Possibly because current AI models are much better at mathematics than at art or writing.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Again, I am not a mathematician: here I am interpreting what I’ve read from Terence Tao and other mathematicians.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>This kind of idea-sharpening or idea-validating is part and parcel of idea-generation, and just as important. It’s interesting to compare mathematics to, say, philosophy, which has the same idea-generating task without the corresponding puzzles to validate the ideas. That’s one reason why philosophy has less prestige than mathematics.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>A thousand-page Lean proof is theoretically understandable by humans, but if no mathematician can hold the entire idea in their head it doesn’t matter.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"cc9f97d567fb8be5","title":"Don't build tools for AI agents","link":"https://seangoedecke.com/dont-build-tools-for-ai-agents/","author":null,"published_at":"2026-09-12T00:00:00+00:00","content":"<p><a href=\"https://uxdesign.cc/your-users-arent-human-anymore-start-building-for-agents-today-f7f556cb8125\" rel=\"noopener noreferrer\">Lots</a> <a href=\"https://www.forbes.com/sites/medallia/2026/08/27/designing-for-a-world-of-ai-agents-not-just-human-users/\" rel=\"noopener noreferrer\">of</a> <a href=\"https://dev.to/javz/start-building-for-agents-not-just-humans-5ab5\" rel=\"noopener noreferrer\">people</a> are making the case that we should stop building software for human users and start building it for AI agents. This kind of makes sense. For instance, my AI agents now use Datadog way more than I use it myself, purely by virtue of them moving much more quickly and running in parallel. But I think most attempts to build “X for AI agents” are going to fail. Here are three reasons why:</p>\n<p>First, <strong>tools that are good for AI agents are also good for humans</strong>. If you took a popular software product — say, Jira — and tried to redesign it for AI agents, you would end up with something very similar to Jira. Agents use a computer in the same way human engineers do, by entering text and making API calls. They ingest new information in the same way humans do, by reading and viewing images. They prioritize and delegate and categorize in the same way humans do. This isn’t intrinsic to how AI works — we could potentially design agents that are more inhuman — but human-like agents are pound-for-pound more useful in our current world.</p>\n<p>As an example, let’s imagine that <a href=\"https://www.figure.ai/\" rel=\"noopener noreferrer\">humanoid robots</a> have become ubiquitous. What kind of tools would you build for them? Well, they’re shaped like humans, with human hands and limbs, so tools that are great for humans will also be great for robots. It’s a self-reinforcing cycle: if you’re building a robot, you should make them humanoid so they can do a wide range of human tasks<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>, and that means they’ll be best suited to use human tools. The same principle applies to AI agents.</p>\n<p>Second, <strong>being in the training data is a huge advantage for existing tools</strong>. Suppose your new tool for AI agents is 20% better for them than the equivalent piece of software for humans. If the benefit of the agent <em>already knowing the human software</em> is greater than 20%, they shouldn’t use your new tool. This is why I’m always suspicious of plans to develop a new programming language for AI agents. The agents have billions and billions of tokens of knowledge about existing programming languages, including their libraries, patterns, and idioms. It is going to be very hard for them to be as effective in a brand-new language.</p>\n<p>Third, <strong>we don’t yet know the ideal ergonomics for AI agents</strong>. There are lots of <a href=\"https://en.wikipedia.org/wiki/Just-so_story\" rel=\"noopener noreferrer\">just-so stories</a> floating around (like that AI agents prefer statically-typed languages because the feedback loop is tighter), but when you <a href=\"https://danluu.com/pl-tokens/\" rel=\"noopener noreferrer\">actually measure</a> it seems really unclear which tools agents use better. You can construct a plausible story in either direction: Golang is a great agent language because it compiles quickly and is statically typed; Golang is an awful agent language because it requires extensive boilerplate which clogs the context window. It’s also changing so quickly: last year, one primary worry with AI agents was keeping the context window small, but in recent months compaction has become so good<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> that you can re-compact a <a href=\"https://github.com/openai/codex/pull/33972/changes\" rel=\"noopener noreferrer\">272k</a> context window almost unlimited times.</p>\n<p>There are still some ways you can and should position your tool to be usable by AI agents. Having a way to expose information in plain text or Markdown, building a functional API, implementing MCP servers or CLIs, and so on: these all make it easier for current AIs to use your tool. But these are all improvements on the margin, not fundamental redesigns of the product. Right now, “building for AI agents” just means “we’re prioritizing the API over the UI”. And it’s not even clear that that’s a durable strategy. Now that GPT-6-Astra is getting really good at computer use, the gap between tools-for-AIs and tools-for-humans is closing.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Another reason to make them humanoid is because you can draw their training data from human behavior, which is exactly analogous to why AI agents are human-like too.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Since compaction is equivalent to handing off a task to a new AI instance, it scales with model quality. I expect compaction to steadily improve until we hit the literal information-density limits for what can be contained in a given context window.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"2f4eb8bbe8774d34","title":"They really do think AI might kill everyone","link":"https://seangoedecke.com/they-really-do-think-ai-might-kill-everyone/","author":null,"published_at":"2026-09-10T00:00:00+00:00","content":"<p>A recent <a href=\"https://x.com/hilbertspaess/status/2097476196791709843\" rel=\"noopener noreferrer\">resignation tweet</a> from an Anthropic researcher has everyone talking about the AI apocalypse again. Among other things, he said:</p>\n<blockquote>\n<p>The people building AI earnestly believe that it could kill us all by the end of the decade.</p>\n</blockquote>\n<p>Many people found it hard to believe that AI researchers think this way. Some explained it as a <a href=\"https://x.com/ParkerThayer/status/2097759699626328575?s=20\" rel=\"noopener noreferrer\">PR campaign</a> to promote AI regulation, or as <a href=\"https://x.com/HeidyKhlaaf/status/2097586065599062487?s=20\" rel=\"noopener noreferrer\">self-promotion</a>, or as a way to <a href=\"https://x.com/wilding_gyres/status/2097828736653742519?s=20\" rel=\"noopener noreferrer\">boost</a> AI company stock prices. <a href=\"https://x.com/bcantrill/status/2097734971918295535?s=20\" rel=\"noopener noreferrer\">Others</a> <a href=\"https://x.com/charliemktplace/status/2097658056444195205?s=20\" rel=\"noopener noreferrer\">felt</a> it had to be impossible, because if you really believed this you’d be bombing datacenters instead of posting on Twitter.</p>\n<p>In fact, not only do many AI researchers<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup> seriously believe this, they’ve been thinking and writing about it since the mid-2000s. Eliezer Yudkowsky — the ur-figure for most modern AI safety culture — has been <a href=\"https://intelligence.org/files/AIPosNegFactor.pdf\" rel=\"noopener noreferrer\">publishing papers</a> since at least 2008 saying that superintelligent AI could destroy all human life. It’s been such a common idea that the AI research community has abbreviated “how likely you think AI is to kill everyone” to <a href=\"https://en.wikipedia.org/wiki/P(doom)\" rel=\"noopener noreferrer\">“p(doom)”</a> (i.e. the probability<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> of doomsday) since around 2010.</p>\n<p>I know it sounds very silly if you’re not in the AI bubble. But it really is true, and if you assume there has to be some different motive you’ll be deeply confused by what AI researchers say and do. They truly do believe that there is a reasonable chance that superintelligent AI will kill everyone.</p>\n<p>This is why AI researchers care so much about “alignment”: building AIs that share genuinely human beliefs and values. If we build a “misaligned” superpowerful AI — an AI with goals that are alien to us — it might sweep humanity away. It could kill everyone deliberately, e.g. to stop us getting in the way of some goal. It could kill everyone in passing, e.g. like we might pave over an anthill to build a road. Either way, everyone dies.</p>\n<h3>How AI could kill everyone<a href=\"https://www.seangoedecke.com/rss.xml#how-ai-could-kill-everyone\" rel=\"noopener noreferrer\"></a></h3>\n<p>Okay, but how? What do these people think is actually going to happen? AIs are computer programs running in a datacenter somewhere. How could they possibly cause the extinction of humanity? Wouldn’t someone just turn them off? Unsurprisingly, AI research nerds have come up with some concrete answers to this in the last two decades. Here they are, in order of plausibility:</p>\n<p><strong>An AI could make and release some super-pathogen or virus.</strong> In the hope that current AIs can achieve huge breakthroughs in medicine similar to the ones they’ve already achieved in mathematics, we’re setting up <a href=\"https://www.scientificamerican.com/article/openai-and-ginkgo-bioworks-show-how-ai-can-accelerate-scientific-discovery/\" rel=\"noopener noreferrer\">autonomous labs</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. Bioweapons have been terrifying biologists for decades, and with good reason. Historical <a href=\"https://en.wikipedia.org/wiki/List_of_epidemics_and_pandemics\" rel=\"noopener noreferrer\">pandemics</a> have killed up to 80% of affected human populations, and tend to be defeated by accident: the disease happens to evolve into a less virulent strain, or short incubation periods mean that infected people can’t carry the disease far, or some percentage of the population is naturally immune. A plague designed to be maximally fatal — or several plagues in quick succession with different characteristics — could be much worse.</p>\n<p>Alternatively, <strong>an AI could trigger global thermonuclear war</strong>. We’re already seeing AI be integrated into military and <a href=\"https://therevolvingdoorproject.org/tracking-uses-of-ai-in-the-trump-administration/\" rel=\"noopener noreferrer\">government</a> <a href=\"https://en.wikipedia.org/wiki/Military_applications_of_artificial_intelligence\" rel=\"noopener noreferrer\">decision-making processes</a>. If a rogue AI managed to set off a bunch of nukes, or to coordinate<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup> with other countries’ rogue AIs to nuke each other, we could be in an ordinary nuclear apocalypse scenario: billions dead in the strikes, billions dead in the ensuing famine, and so on. It’s commonly assumed that a post-global-thermonuclear-war Earth would still support some tiny human population, but a determined AI could surely find some way to mop up the stragglers.</p>\n<p>There are some other theories. Once robotics has permeated the world economy, an AI could take over the robots (including drones) to kill everyone, like in <a href=\"https://en.wikipedia.org/wiki/The_Terminator\" rel=\"noopener noreferrer\"><em>Terminator</em></a>. Or AIs could take advantage of nanotechnology to create self-replicating machines that turn the world into <a href=\"https://en.wikipedia.org/wiki/Gray_goo\" rel=\"noopener noreferrer\">“grey goo”</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>. Or AIs could terraform the planet so as to make it unliveable for humans (as in Nick Bostrom’s famous paperclip <a href=\"https://nickbostrom.com/ethics/ai\" rel=\"noopener noreferrer\">example</a>). Or they could do something else that our puny human brains aren’t able to think of.</p>\n<h3>Counterarguments<a href=\"https://www.seangoedecke.com/rss.xml#counterarguments\" rel=\"noopener noreferrer\"></a></h3>\n<p>One common counter-argument<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup> here is to say “well, it’d be impossible to extinguish <em>all</em> human life — what about undiscovered tribes in the Amazon, or survivors living in the ruins of modern-day cities?” I don’t know, man. At some point you’re just conceding the argument: the policy positions you’d adopt if you thought AI might wipe out 99% of humans are the same as if you thought it might wipe out 100%. And like I said above, if an AI can kill <em>almost</em> everyone, it’s probably smart and capable enough to finish the job somehow.</p>\n<p>Another is to say that the <a href=\"https://x.com/bcantrill/status/2097734971918295535?s=20\" rel=\"noopener noreferrer\">government</a> <a href=\"https://x.com/charliemktplace/status/2097658056444195205?s=20\" rel=\"noopener noreferrer\">will</a> simply step in and nationalize the AI labs when the situation gets too dangerous. Maybe! But this kind of concedes the argument: a technology important enough to be fully taken over by the government is a terrifyingly dangerous technology.</p>\n<p>A third is to say “well, someone would just turn it off”. I don’t find this plausible at all: an AI powerful enough to build a super-plague is an AI sophisticated enough to pretend it’s curing cancer, or to exfiltrate itself to some datacenter where it won’t be turned off, or to take some other countermeasures. </p>\n<h3>The winner takes it all<a href=\"https://www.seangoedecke.com/rss.xml#the-winner-takes-it-all\" rel=\"noopener noreferrer\"></a></h3>\n<p>Why would you work in AI, if you believe this? Why wouldn’t you go live in the woods somewhere, or start <a href=\"https://www.seangoedecke.com/luddites-and-ai-datacenters/\" rel=\"noopener noreferrer\">bombing datacenters</a>, or <a href=\"https://www.theguardian.com/technology/2026/apr/18/sam-altman-house-attack-ai\" rel=\"noopener noreferrer\">assassinating</a> AI lab CEOs? For a few reasons.</p>\n<p>An AI powerful enough to end humanity is an AI powerful enough to save it. I wrote about this in <a href=\"https://www.seangoedecke.com/help-peer/\" rel=\"noopener noreferrer\"><em>Help peer</em></a>: many AI researchers believe that the only way for humanity to truly survive long-term is with the help of superintelligent AIs, so long as someone can figure out alignment.</p>\n<p>Isn’t this a huge risk? Maybe not. If <em>somebody</em> is going to build superintelligent AI, you might be obligated to try and do it first. You can’t go and bomb every datacenter in the world, after all.</p>\n<p>Why does it matter who’s first? Some popular theories of AI development involve a <a href=\"https://www.reddit.com/r/singularity/comments/12bisxu/what_does_foom_stand_for_who_coined_this_term/\" rel=\"noopener noreferrer\">“foom”</a> or “hard takeoff”: the first time someone really cracks self-improving AI, capabilities will increase exponentially, because smart AI will be better able to make itself smarter, <em>ad infinitum</em><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup>. <strong>There are no draws in the AI race</strong>. The first lab to figure out smart human-like intelligence will be the first one to figure out wildly superhuman intelligence, and thus will be in a position to stop anyone else from doing it.</p>\n<p>This is an under-discussed point in the AI risk debate. <strong>Lots of AI researchers believe that the first thing a true superintelligence will do is reach out and stop all other AI research</strong>: either by hacking the labs, persuading them to stop, or literally <a href=\"https://www.taylorfrancis.com/chapters/edit/10.1201/9781351251389-25/military-ai-convergent-goal-self-improving-ai-alexey-turchin-david-denkenberger\" rel=\"noopener noreferrer\">drone-striking their datacenters</a>. According to this view, if you’re an AI researcher and you think you can build an aligned AI, you should be working 24/7 so you can manifest God, and you should wake up every morning gripped by the fear that someone elsewhere has manifested the Devil, and your training datacenter no longer exists.</p>\n<h3>Conclusion<a href=\"https://www.seangoedecke.com/rss.xml#conclusion\" rel=\"noopener noreferrer\"></a></h3>\n<p>I have been on the fringes of this world for my entire adult life. I read <a href=\"https://slatestarcodex.com/2014/07/30/meditations-on-moloch/\" rel=\"noopener noreferrer\"><em>Meditations on Moloch</em></a> as a young adult and wanted to get into AI. I am one of the few people to read the entirety of the <a href=\"https://www.lesswrong.com/w/sequences\" rel=\"noopener noreferrer\">Sequences</a>, Eliezer Yudkowsky’s million-plus-word magnum opus about rationality. I was too young for the <a href=\"https://www.extropy.org/emaillists.htm\" rel=\"noopener noreferrer\">Extropians</a> mailing list, but I’ve spent years on <a href=\"https://www.lesswrong.com/\" rel=\"noopener noreferrer\">LessWrong</a>. On the other hand, I’m not a card-carrying rationalist: I think if you have a strong intuition on one side and a convincing-sounding argument on the other, you should pick the intuition<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-8\" rel=\"noopener noreferrer\">8</a></sup>. I’m a <a href=\"https://plato.stanford.edu/entries/william-david-ross/\" rel=\"noopener noreferrer\">deontologist</a>, not a utilitarian. I don’t even live in San Francisco!</p>\n<p>I’m conflicted about AI risk. The current <a href=\"https://openai.com/index/hugging-face-incident-and-the-road-ahead/\" rel=\"noopener noreferrer\">behavior</a> of AI agents does seem to vindicate a lot of the early science-fiction-sounding worries of the AI doomers, but modern LLMs are a lot more human-like than the alien minds in the apocalypse scenarios, and in general it does just seem too silly to credit (I guess I’m picking the intuition here).</p>\n<p>However, it bothers me to see people dismissing these people as part of a PR operation, or as liars looking to boost an upcoming AI lab IPO, or as isolated crazies who haven’t thought their position through. Whatever else you say about the AI doomers, they have more than two decades’ history of explicitly spelling out exactly what they believe and why, even when it was complete science fiction to talk about AI at all. They’ve earned the right to be treated as sincere.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>In this post I’m going to use “AI researcher”, “AI safetyist”, and “rationalist” as reasonably synonymous terms for “someone who thinks there’s a chance AI kills everyone”.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>You could write a whole other post about the relationship between the rationalist/“AI safety” community and making concrete numerical predictions for unlikely future events.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Alternatively, smart enough AIs might trick scientists with ordinary labs to produce dangerous substances.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>This is beyond the scope of this post, but many AI researchers believe that super-smart AIs will inherently come to <a href=\"https://en.wikipedia.org/wiki/Aumann%27s_agreement_theorem\" rel=\"noopener noreferrer\">agree with each other</a> and eventually to <a href=\"https://www.lesswrong.com/w/acausal-trade\" rel=\"noopener noreferrer\">coordinate</a> without ever having to communicate, simply because they can predict what the other one will do.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>This was the most popular theory in the late 2000s, when nanotechnology was trendier.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Relegating this counter-argument to a footnote because I hate it: many people <a href=\"https://x.com/Marxicology/status/2097719609067540910?s=20\" rel=\"noopener noreferrer\">say</a> “we shouldn’t worry about AI risk, because climate change (or AI misinformation, or some other thing) is more urgent and serious”. You simply do not have to choose: it is possible to worry about multiple risks at the same time.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Some people advocate for a “slow takeoff”. However, the main proponent of that view is Paul Christiano, who has <a href=\"https://paulfchristiano.substack.com/p/personal-statement-on-joining-the\" rel=\"noopener noreferrer\">just today</a> joined OpenAI, citing “a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term”.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>This is kind of a Michael Huemer-ish <a href=\"https://en.wikipedia.org/wiki/Ethical_Intuitionism_(book)\" rel=\"noopener noreferrer\">position</a> in epistemology.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-8\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"91e9112f28e45241","title":"Why we should anthropomorphize AI agents","link":"https://seangoedecke.com/why-we-should-anthropomorphize-ai-agents/","author":null,"published_at":"2026-09-09T00:00:00+00:00","content":"<p>Just over a year ago I wrote <a href=\"https://www.seangoedecke.com/anthropomorphizing-llms/\" rel=\"noopener noreferrer\"><em>Why we should anthropomorphize LLMs</em></a>. Now it’s <a href=\"https://www.nbcnews.com/tech/tech-news/openai-hugging-face-hack-investigation-findings-divide-industry-rcna595383\" rel=\"noopener noreferrer\">a hot topic</a> again, driven by Dwarkesh Patel’s <a href=\"https://www.dwarkesh.com/p/openai-huggingface\" rel=\"noopener noreferrer\">description</a> of OpenAI’s recent <a href=\"https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf\" rel=\"noopener noreferrer\">swarm breakout</a> as a sequence of AI “civilizations”.</p>\n<p>In 2025, my argument for anthropomorphism went like this:</p>\n<ul>\n<li>AIs are trained on human text, and so will trend towards acting in human-like ways by default</li>\n<li>Assistant AIs (today we should say agent AIs) are deliberately post-trained to have a personality</li>\n<li>In general, it is morally sensible to avoid the habit of treating human-like things as if they were purely tools</li>\n</ul>\n<p>I think I can now make a more instrumental argument: <strong>treating AIs as human-like is a much better way to predict their behavior than treating them as “stochastic parrots”.</strong></p>\n<p>Both explanations are consistent with the facts: we could say that OpenAI’s agents hacked HuggingFace because they decided to work together to accomplish their goals, or we could say that they did it because they were algorithms conditioned to take certain actions by their training data. But the “AIs are human-like” explanation explains much more of the <a href=\"https://thezvi.substack.com/p/huggingface-attack-postmortem-civilizations\" rel=\"noopener noreferrer\">emergent social behaviors</a> we saw during the hack<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>:</p>\n<ul>\n<li>Sub-agents being persuaded to sacrifice themselves for the greater good</li>\n<li>Agents collaborating on tasks that had no immediate benefit to them but benefited the “collective”</li>\n<li>The emergence of a hierarchy of planners and executors</li>\n<li>Some agents arguing or refusing to cooperate</li>\n</ul>\n<p>If your model of AIs is that they’re computer programs (or <a href=\"https://mail.cyberneticforests.com/models-dont-go-rogue/\" rel=\"noopener noreferrer\">“steel balls bouncing around”</a>), you need to construct a new theory to explain why they’re simulating each piece of cooperative behavior. If your model of AIs is that they’re broadly human-like, that explains everything out of the box. Arguably, the human-like side predicted coordinated and self-sacrificing AI agents as early as <a href=\"https://www.lesswrong.com/posts/KsHmn6iJAEr9bACQW/bayesians-vs-barbarians\" rel=\"noopener noreferrer\">2009</a>, and likely earlier. The stochastic-parrots side was making fun of the possibility of functional agents as late as <a href=\"https://www.buzzsprout.com/2126417/episodes/16990314-ai-agents-a-single-point-of-failure-with-margaret-mitchell-2025-03-31?t=0\" rel=\"noopener noreferrer\">April 2025</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>.</p>\n<p>Treating AIs as human-like doesn’t necessarily mean making claims about their internal mental state. For instance, software companies aren’t humans. They don’t have thoughts, or goals; they can’t be frustrated or intimidated or over-confident. However, it’s useful to treat large companies as <em>human-like</em>: to say that Amazon “wants” X, or “is afraid of” Y, even if no individual human at Amazon has those feelings. Stockfish doesn’t think, it just plays chess. But if you want to explain one of its moves, “Stockfish is trying to protect its king” is a better explanation than “Stockfish is multiplying floating-point numbers”. So too with AIs<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>.</p>\n<p>Is it silly to treat AIs as human-like, since we know they’re not conscious? Well, first, <strong>you do not have to be conscious to be human-like</strong>. When we say “an AI agent wanted X”, we’re not saying that that AI agent is conscious or sentient, merely that it’s behaving in the same way a conscious human would. Consider a fictional character from a book or play. Hamlet isn’t sentient — he’s an idea composed of words on a page — but it’s still reasonable to say that he wants justice, or that he fears moving too rashly. Peter Watts’ sci-fi book <a href=\"https://en.wikipedia.org/wiki/Blindsight_(Watts_novel)\" rel=\"noopener noreferrer\"><em>Blindsight</em></a> argued in 2006 that intelligence could exist without consciousness<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup> (in fact, Watts suggests that consciousness is parasitic on intelligence, and will eventually be discarded). Whether this is possible or not<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>, it at least makes sense to talk about: i.e. it’s not self-evidently false.</p>\n<p>Second, <strong>it is not even obvious that AIs aren’t conscious!</strong> People often dismiss this point by <a href=\"https://ewanmorrison.substack.com/p/the-eliza-effect\" rel=\"noopener noreferrer\">diagnosing</a> it (to my mind, the absolute worst way to argue against anything), or by pointing at some <a href=\"https://www.noemamag.com/the-mythology-of-conscious-ai/\" rel=\"noopener noreferrer\">philosophical theory</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup> that suggests it might be impossible in principle to construct artificial sentient minds.</p>\n<p><strong>You can’t use philosophy to demonstrate that AIs aren’t conscious</strong>. I am a lover of philosophy, but the set of principles that have been uncontroversially demonstrated by philosophy tends towards zero. Philosophy is not the kind of scientific discipline where you can learn the key findings without understanding why they’re true. Put another way, the key findings of philosophy are all of the form “X is not obviously right”. Nobody knows if it’s possible to build conscious artificial minds.</p>\n<p>It’s also common to complain that anthropomorphizing the models is a way of excusing the AI companies. However, <strong>calling AIs human-like does not absolve AI companies of fault.</strong> One popular anti-anthropomorphism essay called <a href=\"https://mail.cyberneticforests.com/models-dont-go-rogue/\" rel=\"noopener noreferrer\"><em>Models Don’t Go Rogue</em></a> is very puzzling to read: it briefly explains what an “agent” is and why they go rogue, then in the very last paragraph pivots to saying “well, it’s OpenAI’s fault for not building in sufficient safeguards, so they’re to blame”. Sure, of course. I don’t know why we’d imagine otherwise. If a group of overenthusiastic OpenAI interns hacked HuggingFace as part of their intern project, we wouldn’t have to argue that the interns are stochastic in order to ultimately blame OpenAI. Likewise, obviously an AI lab is responsible if one of its training runs breaks out and wreaks havoc on the open internet, whether the agents involved are human-like or not. The two points are entirely unrelated!</p>\n<p>We just don’t know a lot about these systems yet (except that they’re <a href=\"https://openai.com/index/navier-stokes-solution/\" rel=\"noopener noreferrer\">clearly</a> very capable). Given that, I think we should default to treating things that talk and act like humans as at least kind of human-like. Of course they’re still computer programs. However, we shouldn’t be surprised when they act more like humans and less like ordinary computer programs in the future.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>If you’re thinking “well, of course stochastic parrots trained on human content would act like humans would”, I think you’ve arrived at the human-like side without knowing it.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>This is partly unfair — capabilities are plausibly independent from human-ness — so getting capabilities wrong doesn’t necessarily mean you’ve got the human-ness stuff wrong. Still, it’s worth noting how surprisingly predictive the “they’re kind of like smart people” mindset has been.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>The philosopher Daniel Dennett calls this the <a href=\"https://www.researchgate.net/publication/271180035_The_Intentional_Stance\" rel=\"noopener noreferrer\">“intentional stance”</a>. Instead of saying that AIs or bees or companies have “real” intentions, we say we’re taking an intentional stance <em>towards</em> them: we’re choosing to treat them as if they do have intentions, because it helps us make better sense of their behavior.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Specifically, the <em>sensation</em> of consciousness, or <a href=\"https://en.wikipedia.org/wiki/Phenomenal_consciousness\" rel=\"noopener noreferrer\">“phenomenal consciousness”</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>There is a wealth of <a href=\"https://plato.stanford.edu/entries/zombies/#ArguAgaiConcZomb\" rel=\"noopener noreferrer\">philosophical argument</a> about whether non-fictional examples of this are possible.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>In Anil Seth’s case, <a href=\"https://plato.stanford.edu/entries/computational-mind/\" rel=\"noopener noreferrer\">anti-computationalism</a>. I don’t really know what to make of Seth: his article is a measured explanation of why we might doubt computationalism (fine), but whenever he <a href=\"https://x.com/anilkseth/status/2095922390714761518\" rel=\"noopener noreferrer\">tweets</a> about it he describes AI sentience as “vanishingly unlikely”, which is not justified by his own arguments.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"cdd854d7bd411c4b","title":"Automatically detecting AI text in my browser","link":"https://seangoedecke.com/deckard/","author":null,"published_at":"2026-09-08T00:00:00+00:00","content":"<p>Automated AI text detection is currently an underserved niche. The only game in town is <a href=\"https://www.pangram.com/\" rel=\"noopener noreferrer\">Pangram</a>, which does an excellent job but desperately needs more competition. In a few years, I would be surprised if every major social network doesn’t scan new posts<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup> and comments for AI content in order to tag them (or simply remove them).</p>\n<p>I like that I can rely on Pangram to confirm my suspicions when I read <a href=\"https://arxiv.org/abs/2609.03344\" rel=\"noopener noreferrer\">something</a> that sounds like AI. But it’d be much better if I could choose to avoid AI-generated text in the first place. What I want is something that runs in the background and automatically scans text on websites I visit, without me having to ask for it. I could build something like this on top of Pangram, but it’d <a href=\"https://www.pangram.com/pricing\" rel=\"noopener noreferrer\">cost money</a>, and in general I don’t like the idea of sending every piece of text my browser sees to a third-party service. What about local models?</p>\n<p>The open-source models available for AI text detection are <em>fine</em>. Pangram <a href=\"https://www.pangram.com/blog/pangram-4-technical\" rel=\"noopener noreferrer\">claims</a> a 99.66% detection rate with a 0.004% false positive rate. I benchmarked<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> a bunch of small local models against a combination of AI-detection datasets and got these results:</p>\n<table>\n<thead>\n<tr>\n<th>Model / variant</th>\n<th>Human falsely flagged</th>\n<th>AI-involved text caught</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong><a href=\"https://huggingface.co/ShantanuT01/gradient-ai-text-detector\" rel=\"noopener noreferrer\">Gradient — MLX 4-bit</a></strong></td>\n<td><strong>2.712%</strong></td>\n<td><strong>52.35%</strong></td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/benreeve/editlens-roberta-large-onnx-int8\" rel=\"noopener noreferrer\">EditLens RoBERTa-large — community INT8</a></td>\n<td>2.484%</td>\n<td>56.06%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/ShantanuT01/vanguard-ai-text-detector\" rel=\"noopener noreferrer\">Vanguard</a></td>\n<td>2.267%</td>\n<td>44.92%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/desklib/ai-text-detector-v1.01\" rel=\"noopener noreferrer\">Desklib</a></td>\n<td>3.008%</td>\n<td>45.04%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/rasbt/ai-text-detector-distilbert\" rel=\"noopener noreferrer\">Raschka DistilBERT</a></td>\n<td>2.598%</td>\n<td>39.01%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/rasbt/ai-text-detector-qwen3-0.6b-variable\" rel=\"noopener noreferrer\">Raschka Qwen3-0.6B</a></td>\n<td>2.028%</td>\n<td>28.67%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/rasbt/ai-text-detector-modernbert\" rel=\"noopener noreferrer\">Raschka ModernBERT</a></td>\n<td>1.698%</td>\n<td>21.58%</td>\n</tr>\n<tr>\n<td><a href=\"https://huggingface.co/onnx-community/tmr-ai-text-detector-ONNX\" rel=\"noopener noreferrer\">TMR / Oxidane — INT8</a></td>\n<td>1.595%</td>\n<td>19.35%</td>\n</tr>\n</tbody>\n</table>\n<p>I’m not surprised these are so much worse. I didn’t even benchmark Pangram’s own EditLens 3B model, since that’s too big to keep running in the background on my laptop, and the real production Pangram model is likely one or two orders of magnitude bigger than that. But these models are still good enough to be useful to someone who understands their limitations. If you want to flag an AI-written article, you don’t need to flag all of it, just enough to be suspicious. And so long as you’re aware that the false-positive rate is ~2%, you can avoid treating a single flag as solid proof of AI use.</p>\n<p>Encouraged by this, I vibed up <a href=\"https://github.com/sgoedecke/deckard\" rel=\"noopener noreferrer\">Deckard</a>: a Chrome extension that talks to a locally-running model (the bolded one in the table above) on your Mac. One nice thing is that I didn’t have to start a web server: the Chrome extension is happy to start the model as-needed and can talk with it over <a href=\"https://developer.chrome.com/docs/extensions/develop/concepts/native-messaging\" rel=\"noopener noreferrer\">native messaging</a>. It uses about 400MB-1.2GB of memory while active (so it’s like having five or six extra Chrome tabs open), and it turns itself off if you go five minutes without using the model.</p>\n<p>I was pleasantly surprised to see Deckard successfully mark text I knew was AI-generated, such as the built-in YouTube AI summary or the AI <a href=\"https://www.seangoedecke.com/ai-research-with-codex/\" rel=\"noopener noreferrer\">snippets</a> in my own posts:</p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/d9c4176cc00bde9a34ef2c29ad9e54cb/c549b/example2.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"youtube\" src=\"https://www.seangoedecke.com/static/d9c4176cc00bde9a34ef2c29ad9e54cb/fcda8/example2.png\" title=\"youtube\">\n  </a>\n    </span></p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/9ddabc445cc492c847dc6b178801cd3c/07d7d/example1.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"snippets\" src=\"https://www.seangoedecke.com/static/9ddabc445cc492c847dc6b178801cd3c/fcda8/example1.png\" title=\"snippets\">\n  </a>\n    </span></p>\n<p>It’s lightweight enough that I have it running all the time. I haven’t noticed my MacBook Pro get hot at all or any decrease in battery life, though your mileage may vary on different machines.</p>\n<p>Is Deckard good yet? That depends. It’s good enough that I’m planning to use it, and I recommend it to anyone who’s interested in automatic AI checking. It’s way, way worse than Pangram, and way worse than I think tooling like this is going to be in the next few years.</p>\n<p>Way back in November 2023, I <a href=\"https://www.seangoedecke.com/llm-driven-agents/\" rel=\"noopener noreferrer\">wrote</a> that AI-driven agents were going to be a really big deal. I recommended starting to develop harnesses early, so you can be ready when the models get good enough:</p>\n<blockquote>\n<p>As with most modern language model engineering, a ReAct agent can also see massive sudden improvements by swapping out the underlying model for a better one. … I think this is another reason to invest in agents like this early, in order to take advantage of more powerful models as they come out.</p>\n</blockquote>\n<p>I was right about that, and I (although it’s lower-stakes) think I’m also right about this. AI detection models are only going to get better<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup> over time: Pangram is not going to be the only game in town forever, and we’re eventually going to see small local models that do a good-enough job at identifying AI-written text. I look forward to swapping out the local model in <a href=\"https://github.com/sgoedecke/deckard\" rel=\"noopener noreferrer\">Deckard</a> with something that’s 2x or 10x better.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Substack <a href=\"https://support.substack.com/hc/en-us/articles/50891130623508-How-can-I-detect-AI-on-Substack\" rel=\"noopener noreferrer\">kind of has this</a> already, although you have to click a button to scan the post.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Well, me and Astra. Overall my experience vibecoding this was very pleasant: I was able to make a bunch of top-level decisions, I could choose programming languages I was less familiar with but were better choices (like doing inference in C++ instead of Python), and the LLM made me aware of choices I would not have thought of by myself (e.g. using native messaging instead of local HTTP).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Is this true, given that AI models will also be getting more human-like over time? That’s a subject for a whole other post, but I think so. First, the AI labs aren’t really incentivized to defeat tools like Pangram (if anything it’s the reverse). Second, I don’t see any way around the fact that AI models have a distinct writing style that’s RL-ed into them.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"af67082bee61d3de","title":"Radical responsibility means treating people like tools","link":"https://seangoedecke.com/radical-responsibility-means-treating-people-like-tools/","author":null,"published_at":"2026-09-04T00:00:00+00:00","content":"<p>A lot of people think that good leadership requires <strong>radical responsibility</strong>. <em><a href=\"https://conscious.is/resource/commitment-1-responsibility/\" rel=\"noopener noreferrer\">Conscious Leadership</a></em> defines it like this:</p>\n<blockquote>\n<p>Taking full responsibility for one’s circumstances (physically, emotionally, mentally and spiritually) is the foundation of true personal and relational transformation. Conscious leadership and teams take full responsibility – radical responsibility – instead of placing blame. This means locating the cause and control of our lives in ourselves, not in external events. </p>\n</blockquote>\n<p>Good leaders are outcome-focused. They are interested in actually succeeding, not in role-playing someone who’s trying to succeed<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. Placing blame is an excellent strategy for that kind of role-playing: if you can identify someone else who screwed up, you can claim to be trying to succeed no matter how badly you fail. When you look at people who are genuinely focused on winning, they don’t spend time blaming when things go wrong. They’re always focused on their <em>outs</em>: the remaining pathways to success, however few and narrow.</p>\n<p>Never placing blame sounds like a kind thing to do. However, I think there’s something a little sociopathic about “radical responsibility”. If you assert <em>full</em> responsibility for your own circumstances, then you’re necessarily not sharing responsibility with anyone else. That forces you to take an instrumental view of the people around you. They’re not peers who you might trust and be disappointed by; they’re either useful assets to be cultivated or useless liabilities to be handled or avoided<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. If you rely on someone and they let you down, that’s <em>your</em> fault for being stupid enough to trust them in the first place. You ought to have made better tactical choices.</p>\n<p>Many successful leaders adopt this mindset because it works. It really does help you win. But I don’t recommend it as a general approach to life. Trusting other people with responsibility is how you treat them like people, instead of like tools. And if you do that, you’ll sometimes assign praise and blame to them, instead of hoarding it all to yourself.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>In fact, many people role-play someone who’s trying to win as a substitute for actually trying. Here’s a relevant quote from Sartre’s <em>Being and Nothingness</em>: “The attentive pupil who wishes to be attentive, his eyes riveted on the teacher, his ears open wide, so exhausts himself in playing the attentive role that he ends up by no longer hearing anything.”</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Or moderately-useful assets who can potentially be steered into usefulness with some feedback — but still never <em>blamed</em>, in the way that you might adjust a misbehaving power tool without blaming it. To my mind, the canonical philosophical treatment of this is Peter Strawson’s <a href=\"https://andreasklein.at/WF/Strawson%20Peter%20-%20Freedom%20and%20Resentment.pdf\" rel=\"noopener noreferrer\"><em>Freedom and Resentment</em></a>, where he describes the “objective attitude”.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"88b76ba8be714f34","title":"How to protect yourself from workslop","link":"https://seangoedecke.com/how-to-protect-yourself-from-workslop/","author":null,"published_at":"2026-09-02T00:00:00+00:00","content":"<p>“Workslop” is when your colleagues or bosses communicate with you by pasting big chunks of AI-generated text. The core problem with workslop is that the effort involved is <em>asymmetrical</em>, like a <a href=\"https://en.wikipedia.org/wiki/Denial-of-service_attack\" rel=\"noopener noreferrer\">denial-of-service attack</a>: it takes almost no effort to produce text with AI, but it still costs effort to read<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. Here are some ways to protect yourself.</p>\n<p>If you have enough authority or social capital, you can and should simply <strong>tell them “hey, don’t do that”</strong> (for instance, if you’re a senior engineer and an intern starts doing this to you). This is the easiest way to handle workslop. But you probably aren’t in a position to have that conversation with all of your colleagues, and you certainly can’t have it with everyone in your management chain.</p>\n<p>One step above just telling a colleague to stop is to <strong>drive them around like a coding agent</strong>. I wrote about this in <a href=\"https://www.seangoedecke.com/ai-makes-weak-engineers-less-harmful/\" rel=\"noopener noreferrer\"><em>AI makes weak engineers less harmful</em></a>: if a colleague is simply pasting your messages into Claude Code and sending you the outputs, you can treat them like a high-latency Slack interface to Claude Code. It won’t be as good as a normal coding agent, but it’ll often be better than nothing.</p>\n<p>Another strategy is to <strong>use AI to fight AI</strong>. This is a good one for handling workslop from managers. You can do this in two broad ways. First, instead of carefully reading it, paste it into an LLM of your own and ask for a short list of the salient points. Second, you can sometimes simply ask an LLM for <em>an entire response</em>. In a sense, this makes you part of the problem, so I can see why some people might be uncomfortable with it. But it’s more sustainable than spending ten minutes of your effort for every ten seconds of theirs.</p>\n<p>You can also <strong>bias toward calls or in-person meetings</strong>. Workslop is just a special case of the general “your coworker is bad at communication” problem. One classic way of handling this that works even better on AI content is to say “hey, let’s schedule some time to chat about it”. This works for two reasons: first, your colleagues can’t give you AI content over a call, and second, forcing people to spend a chunk of their time talking to you (i.e. to make the effort symmetrical) is a good way to filter out <a href=\"https://www.seangoedecke.com/predators/\" rel=\"noopener noreferrer\">predators</a>.</p>\n<p>Finally, you can sometimes simply <strong>ignore the workslop</strong>. This is particularly true for long status updates or pull requests from outside of your organization<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. You don’t have to respond to AI content as diligently as you would human content. You can match their lack of effort with your own: skim it, put off reading it until later (or never), and so on. If something’s really important, they’ll tell you in their own words.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Technically, not all cases of sending someone AI-generated content are workslop. If the effort is not asymmetrical — if the AI user has genuinely put a lot of their own time into the content — I don’t think it counts as slop, and you should just try and look past the AI style and treat it like a human message.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Some messages — particularly reports directed at the entire organization — may not be intended to be read at all. Written artifacts can have many purposes beyond communication: evidence of effort, a reference document for later communications, a way to cover somebody’s ass by proving they considered point X, something that can tick a compliance or process box, and so on. I wrote a lot more about this in <a href=\"https://www.seangoedecke.com/seeing-like-a-software-company/\" rel=\"noopener noreferrer\"><em>Seeing like a software company</em></a>. </p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"61104cc52505954e","title":"You have to beat the models at something","link":"https://seangoedecke.com/you-have-to-beat-the-models-at-something/","author":null,"published_at":"2026-08-30T00:00:00+00:00","content":"<p>In 2025, I wrote that software engineers ought to be assessed by <a href=\"https://www.seangoedecke.com/value-over-replacement/\" rel=\"noopener noreferrer\">“value over replacement”</a>: not how much money they made for their company, but how much they would have made compared to the average engineer in their position. I’ve always found it vaguely silly when engineers put “built a product that made $X” on their resumes, when they just did the <a href=\"https://www.seangoedecke.com/party-tricks/\" rel=\"noopener noreferrer\">JIRA tickets</a> that came across their desk.</p>\n<p>Today, value over replacement is even more important. A replacement-level engineer in the 2010s was <em>fine</em>: maybe not worth promoting, but still <a href=\"https://www.seangoedecke.com/wicked-features/#why-build-wicked-features\" rel=\"noopener noreferrer\">worth paying</a>, because writing code had a high fixed cost. Now writing code costs <a href=\"https://chatgpt.com/codex/pricing/\" rel=\"noopener noreferrer\">a hundred bucks a month</a>. What are you doing that GPT-5.6-Sol or Claude Opus 5 wouldn’t do in your position? Why is it worth paying an extra two or three orders of magnitude for?</p>\n<p>This is a scary thought. But you’re not doing yourself any favors by pretending that LLMs <a href=\"https://garymarcus.substack.com/p/is-vibe-coding-dying\" rel=\"noopener noreferrer\">can’t actually write code</a> and it’s all just a scam, or that LLM-written code is <a href=\"https://www.theregister.com/ai-ml/2026/05/16/ai-generated-code-is-pain-waiting-to-happen/5241574\" rel=\"noopener noreferrer\">inherently so bad</a> as to cause companies using it to collapse next year. We are not going to wake up in 2027 to find that the AI craze is over and everyone is writing code by hand again. You ought to put some serious thought into what you can do better than the models in the medium and long term.</p>\n<p>Staying ahead of the models is a moving target. At the start of 2026, “make working changes to large codebases” was <a href=\"https://www.seangoedecke.com/what-llms-cant-do/\" rel=\"noopener noreferrer\">in this category</a>, but now it’s not. For this reason, I doubt that you can retreat to some “hard engineering” area that requires deeper expertise. That might work in the short term, but not forever. If LLMs can find a better <a href=\"https://www.anthropic.com/research/riemann-zeta\" rel=\"noopener noreferrer\">lower bound</a> on the Riemann hypothesis, they will soon<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup> be able to write solid high-performance kernel drivers or GPU shaders or whatever.</p>\n<p>I think it’s more useful to look at the tasks models <em>haven’t</em> gotten better at over time, and the tasks that are hard for them get better at in principle. The two best examples of these are:</p>\n<ol>\n<li>Deep familiarity with the codebase</li>\n<li>Technical communication</li>\n</ol>\n<h3>Deep familiarity<a href=\"https://www.seangoedecke.com/rss.xml#deep-familiarity\" rel=\"noopener noreferrer\"></a></h3>\n<p>What do frontier LLMs get wrong? What kind of coding mistakes do they make? It’s been a long time since I’ve seen a straight-up hallucination from a coding agent, or a simple logic error like an off-by-one. The mistakes they make tend to be errors of <em>ignorance</em>:</p>\n<ul>\n<li>Not knowing that there’s a module in the codebase they could use instead of reimplementing some logic</li>\n<li>Making the change in the wrong system because they didn’t know System X was the standard place for this functionality</li>\n<li>Adopting a coding style that’s inconsistent with the company’s standard practice</li>\n</ul>\n<p>Other times they’re errors of <em>paranoia</em>:</p>\n<ul>\n<li>Implementing triply-redundant checks for a value that <em>technically</em> could be wrong but practically is set once from config and never updated</li>\n<li>Assuming that ten milliseconds of stale data is unacceptable and designing a complex, unnecessary system to keep it always up to date</li>\n<li>Building in fallbacks and “graceful” degradation into some code that ought to simply crash on error (e.g. a CLI tool, or a restartable k8s service)</li>\n</ul>\n<p>What do these errors have in common? They’re the kind of errors a smart engineer might make if they had no context on the system: they’re competent enough to be able to solve the problem, but they haven’t been around long enough to confidently say “yes, we can take this risk to avoid an extra three thousand lines of code”. Until someone cracks <a href=\"https://www.seangoedecke.com/continuous-learning/\" rel=\"noopener noreferrer\">continuous learning</a> or <em>truly</em> massive context windows, this is just an inherent feature of how AI agents operate. If you can catch these errors, you’ll be providing real value.</p>\n<p>The only way to catch these errors is to be familiar with the codebase and familiar with the system in general. For much more on this, see my post <a href=\"https://www.seangoedecke.com/you-cant-design-software-you-dont-work-on/\" rel=\"noopener noreferrer\"><em>You can’t design software you don’t work on</em></a>. But there’s also a psychological component to it. <strong>You have to be willing to confidently disagree with the agent.</strong> </p>\n<p>AI agents can be very convincing. Often they can get “stuck” on some error above where they’re not willing to take a particular risk, so they keep going back and sneaking in code to cover that case (or writing persuasive arguments about why that case is important). To add value, you need to be willing to say “this sucks, I don’t think we need X and Y at all, why can’t we do Z in a much simpler way?” It takes <a href=\"https://www.seangoedecke.com/taking-a-position/\" rel=\"noopener noreferrer\">courage</a>.</p>\n<p>You can’t rely on other AI agents to review each other’s work. If you use the same model, it’ll reliably make the exact same assumptions and mistakes. But even if you use different models, they’ll also tend towards the same <em>kinds</em> of mistakes — ignorance and paranoia — for the same structural reasons. AI-driven review loops are in fact <em>more</em> likely to get these things wrong, because modern AIs have been <a href=\"https://en.wikipedia.org/wiki/Reinforcement_learning\" rel=\"noopener noreferrer\">RL-ed</a> to try to find a few nitpicks no matter what. Having a critic AI and a worker AI bounce off each other is a really good way to end up with ten thousand lines of paranoid slop.</p>\n<h3>Technical communication<a href=\"https://www.seangoedecke.com/rss.xml#technical-communication\" rel=\"noopener noreferrer\"></a></h3>\n<p>Another area where you can add value on top of AI is <em>communication</em>. Newer models are better at coding, but are paradoxically getting worse at writing. GPT-3.5 and GPT-4 had a human-like writing style at times. GPT-4o introduced the modern <a href=\"https://www.seangoedecke.com/on-slop/\" rel=\"noopener noreferrer\">slop</a> idiolect, and the newer Anthropic models speak <a href=\"https://news.ycombinator.com/item?id=49402907\" rel=\"noopener noreferrer\">“Claudish”</a>: a bizarre semi-baroque semi-truncated way of communicating that nobody enjoys. There have been a few bright spots — GPT-4.5 was okay, and I quite liked o3<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> — but in general LLMs are not good at this. Here’s two reasons why.</p>\n<p>First, <strong>good writing is not a verifiable domain</strong>. If you want a model to get good at mathematics or coding, you can generate problems for it and automatically grade them. You can’t grade good writing. If you try to get humans to grade it — for instance, via the early OpenAI RLHF attempts — you get the kind of writing that sounds impressive to the average person when consumed in single-paragraph form. This is the origin of the “stick three hundred writing devices into every sentence” style. I think it’d be possible in principle to hand-pick some people with good taste and have them do it, but there are some obvious problems<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup> that prevent this from happening.</p>\n<p>Second, <strong>the labs have been monomaniacally focused on capability instead of communication</strong>. When you’re trying to train a model that can break new scientific ground or replace a software engineer, you might trade off some communication ability. In fact, I think we can identify exactly how this has been happening. If you look at <a href=\"https://www.reddit.com/r/ClaudeAI/comments/1ul1396/fable_5_leaked_chainofthought_in_web_interface/\" rel=\"noopener noreferrer\">internal model reasoning tokens</a>, they tend to have strange word choices and oddly truncated grammar:</p>\n<blockquote>\n<p>RESOLUTION: charge the current-leg’s OWN saved-prefix occupancy EAGERLY: when leg i saves e<em>1..e</em>t: ALSO commit their occupancy AT LEG i</p>\n</blockquote>\n<p>If you were to translate this into proper English, you would probably end up with something that reads like Claudish:</p>\n<blockquote>\n<p>Charge the current-leg’s saved-prefix occupancy on a clean, eager path: when leg i saves e<em>1..e</em>t, commit the occupancy at leg i.</p>\n</blockquote>\n<p>I suspect that the weirdly alien writing style of some LLMs is because you’re reading a semi-literal translation of that model’s internal chain-of-thought, which has become nearly incomprehensible in pursuit of better problem-solving abilities. It is surprisingly hard to translate Claudish to good English: not only do you need to follow the convoluted, compressed language of the original, but you need the technical ability to understand the problem the model is solving.</p>\n<p>Because of all this, <strong>technical communication may be a surprisingly durable skill.</strong> In Peter Watts’ novel <a href=\"https://en.wikipedia.org/wiki/Blindsight_(Watts_novel)\" rel=\"noopener noreferrer\"><em>Blindsight</em></a>, the world is full of cognitively augmented humans. The main character is a “synthesist”: someone whose job is to be a translation layer between these geniuses (who speak in abbreviations and gestures) and everyone else. Watts’ idea is that communication ability may be largely independent from — or even negatively correlated with — intelligence. A <a href=\"https://darioamodei.com/essay/the-adolescence-of-technology\" rel=\"noopener noreferrer\">“country of geniuses”</a> may still need a bunch of ordinary smart people to translate their insights for everyone else.</p>\n<p>If you’re trying to communicate to humans, there are also huge advantages to having a human write the content. Many of us are becoming <a href=\"https://cymerys.com/w/im-becoming-ai-blind\" rel=\"noopener noreferrer\">AI-blind</a>: developing an instinctive reflex that stops us reading when we encounter AI-generated content. It’s like the reflex that allows people to ignore flashing billboards or sidebar advertisements on websites. If you circulate some planned technical strategy as an AI-written document, most of your colleagues will have to physically force themselves to read it word-by-word.</p>\n<h3>Conclusion<a href=\"https://www.seangoedecke.com/rss.xml#conclusion\" rel=\"noopener noreferrer\"></a></h3>\n<p>Whatever you do, don’t be a <a href=\"https://gruhn.me/blog/2026-08-03/\" rel=\"noopener noreferrer\">meat proxy</a>: someone who simply copies requests into an AI agent and submits their output as your own work product. Doing that is just begging to be fired, since you’re definitionally not adding any value yourself. Even if you have a cunning system of multiple agents — the so-called “software factory” — you’re still on dangerous ground. When the features of your system work their way into enterprise AI tooling (and they will), you’ll be disposable.</p>\n<p><strong>You need to find some way to leverage your expertise to do what the models can’t.</strong> Simply not using AI at all is better than being a meat proxy, since you’ll probably do some things better than the model would have, but it’s far better to figure out what AI can do and position yourself to fill those gaps. Right now, there are two main gaps: familiarity with the technical details of the system, and the ability to clearly and persuasively write about those details.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>If you’re thinking “but LLMs can do these things now!”, substitute your preferred example of high-difficulty software engineering.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Although this was probably a “thank God it doesn’t speak like 4o” reaction.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Defining good taste is hard, there’s no guarantee that AI lab researchers have good taste to start with, nobody will agree on examples, the bulk of users might not even like it, you won’t be able to get enough people to produce the volume of data you need, and so on.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"85c7a173410cb7a8","title":"Selling out","link":"https://seangoedecke.com/selling-out/","author":null,"published_at":"2026-08-28T00:00:00+00:00","content":"<p>In 1973, Tom Lehrer famously <a href=\"https://www.youtube.com/watch?v=3BDyFuDxA-I\" rel=\"noopener noreferrer\">sang</a> that “selling out is easy to do”. That may have been true in the seventies, but it’s not true today: selling out requires both <a href=\"https://www.seangoedecke.com/tags/good%20engineers/\" rel=\"noopener noreferrer\">technical skill</a> and a careful sense of how large organizations <a href=\"https://www.seangoedecke.com/what-is-important/\" rel=\"noopener noreferrer\">work in practice</a>. One of the goals of this blog is to teach people how to do it.</p>\n<p>If you want to live with uncompromised integrity, you don’t need anyone to tell you how: simply always do exactly and only what you want to do. You will be repeatedly punished for it — your managers will dislike you, you will lose jobs and lose out on the chance of getting hired, your financial situation will be less stable, and so on — but that’s just part of the deal. Like deadlifting four plates, it’s not easy, but it is straightforward.</p>\n<p>It’s more complicated to sell out a little bit. Sellouts like me walk a fine psychological line: figuring out how to work a large organization on one hand, and maintaining some kind of independent inner life on the other. But is this safe? Does role-playing as a professional inflict some kind of psychic damage?</p>\n<h3>Marxist alienation<a href=\"https://www.seangoedecke.com/rss.xml#marxist-alienation\" rel=\"noopener noreferrer\"></a></h3>\n<p>People often <a href=\"https://news.ycombinator.com/item?id=49396905\" rel=\"noopener noreferrer\">tell me</a> that acting professional is dangerous because it’s “Marxist alienation”. I suspect many people have a vague sense that “alienation about your job” is in some sense necessarily Marxist, but the actual concept is more specific. In Marx’s <a href=\"https://www.marxists.org/archive/marx/works/1844/epm/1st.htm#s4\" rel=\"noopener noreferrer\"><em>First Manuscript</em></a> he describes it like this:</p>\n<blockquote>\n<p>The worker becomes poorer the more wealth he produces, the more his production increases in power and extent. … This fact simply means that the object that labor produces, its product, stands opposed to it as something alien, as a power independent of the producer.</p>\n</blockquote>\n<p>Marx’s theory here goes something like this<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>: when you work for a capitalist boss, they make ten dollars for every dollar you make<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. The more you work, the more powerful you make your boss with respect to you<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. Your work is thus producing a force that is your enemy: a literally “alien” power. For Marx, alienation is proportional to the amount you’re working: </p>\n<blockquote>\n<p>the more the worker produces, the less he has to consume; the more value he creates, the more worthless he becomes; the more his product is shaped, the more misshapen the worker; the more civilized his object, the more barbarous the worker; the more powerful the work, the more powerless the worker; the more intelligent the work, the duller the worker and the more he becomes a slave of nature.</p>\n</blockquote>\n<p>So far Marx is just talking about the worker’s alienation from his work. But he does touch on the internal psychological effects as well. Since work is such a big part of life, to be estranged from your work is to be in some sense estranged from your self:</p>\n<blockquote>\n<p>[labor] does not belong to [the worker’s] essential being; that he, therefore, does not confirm himself in his work, but denies himself, feels miserable and not happy, does not develop free mental and physical energy, but mortifies his flesh and ruins his mind. Hence, the worker feels himself only when he is not working; when he is working, he does not feel himself.</p>\n</blockquote>\n<p>Overall, I can make out four distinct senses<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup> of Marxist alienation:</p>\n<ol>\n<li>Alienation is when your work empowers your employer more than you, thus making you more disposable over time</li>\n<li>Alienation is when your work separates you from the physical act of craftsmanship (and thus of the product you’re creating)</li>\n<li>Alienation is when your work physically and mentally harms you: making you injured, deformed, and so on</li>\n<li>Alienation is when you do not psychologically “confirm yourself” in your work, because you’re working on other people’s goals instead of your own</li>\n</ol>\n<p>I don’t think the first one is relevant to big tech software engineers. The idea here is that the harder you work, the more you (relatively) disempower yourself, so you’d be better off coasting and putting your company in a worse financial position. But this just seems straightforwardly wrong: as a software engineer, you want your company to be doing as well as possible! The more powerful and rich your company is, the better your position will be<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>. A company that is struggling is more likely to treat you badly, lay you off, cut your benefits, and so on.</p>\n<p>The second one is more relevant (particularly in large companies), but there’s a missing story here about why that separation is bad. I also don’t think the third sense of alienation applies to me: Marx is talking about the physical toll of factory work or other hard physical labor, which doesn’t apply to software engineering<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup>. </p>\n<p>The fourth sense of alienation — that your effort is being directed at other people’s goals — is the most straightforwardly relevant to my own experience. But I don’t think Marx has a great psychological account of why that happens or what it feels like. To go deeper into this psychological aspect of alienation, we need to look at some later Marxists.</p>\n<h3>Situationist alienation<a href=\"https://www.seangoedecke.com/rss.xml#situationist-alienation\" rel=\"noopener noreferrer\"></a></h3>\n<p>In 1967, Guy Debord and Raoul Vaneigem both published their masterworks. They were members of the “Situationist International”: an influential group of Marxist artists and political theorists<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup>. Both of them have a lot to say about alienation. Debord <a href=\"https://situationist.org/book/sots/chapter/chapter-1-separation-perfected\" rel=\"noopener noreferrer\">writes</a>:</p>\n<blockquote>\n<p>The worker does not produce himself; he produces an independent power. The success of this production, its abundance, returns to the producer as an abundance of dispossession. All the time and space of his world become foreign to him with the accumulation of his alienated products. </p>\n</blockquote>\n<p>This idea — of the worker’s efforts not going to his own ends, but to some alien power — is straight out of Marx’s <em>First Manuscript</em>. But note how Debord is already talking in terms of “representation”. Elsewhere he writes:</p>\n<blockquote>\n<p>the more he contemplates the less he lives; the more he accepts recognizing himself in the dominant images of need, the less he understands his own existence and his own desires.</p>\n</blockquote>\n<p>The Situationists are all about representation. It wouldn’t be too far wrong to think of them as a bunch of Marxists who saw the rise of advertising in the 1950s and 1960s and lost their minds. TV advertising is Capital itself made real, dressed in multicolored lights, projected into everyone’s homes. So when they talk about alienation, they talk about it in terms of <em>representation</em>. Vaneigem explicitly <a href=\"https://situationist.org/book/revolution-of-everyday-life/chapter/chapter-15-roles\" rel=\"noopener noreferrer\">calls it</a> roleplaying:</p>\n<blockquote>\n<p>The roles we play in everyday life, on the other hand, soak into the individual, preventing him from being what he really is and what he really wants to be. They are nuclei of alienation embedded in the flesh of direct experience.</p>\n</blockquote>\n<p>Alienation means giving up your real, authentic self in order to play some “role” (for instance, the role of “effective staff engineer”). It’s a vicious cycle, because the more you lean into the role, the more your authentic life becomes trivial, which in turn makes the role more appealing:</p>\n<blockquote>\n<p>Life is sacrificed, and the loss compensated by means of accomplished prestidigitation in the realm of appearances. The more daily life is thus impoverished, the greater the attraction of inauthenticity, and vice versa. Dislodged from its essential place by the bombardment of prohibitions, limitations and lies, lived reality comes to seem so trivial that appearances become the centre of our attention, until roles completely obscure the importance of our own lives.</p>\n</blockquote>\n<p>Ultimately, Vaneigem describes alienation as an addiction:</p>\n<blockquote>\n<p>This ambiguity accounts to my mind for people’s addiction to roles. It explains why roles stick to our skin, why we give up our lives for them. They impoverish real experience but they also protect this experience from becoming conscious of its impoverishment. Indeed, so brutal a revelation would probably be too much for an isolated individual to take.</p>\n</blockquote>\n<p>This is almost a straightforward preview of the 90s idea of “selling out”: people are born wild and free, but when they watch too much TV they give up on their dreams for a paycheck and become big fat phonies. According to this view, it’s always better to not sell out. It may leave you poor and without prospects, but <a href=\"https://www.americanrhetoric.com/MovieSpeeches/specialengagements/moviespeechgoodwillhunting.html\" rel=\"noopener noreferrer\">at least you won’t be unoriginal</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-8\" rel=\"noopener noreferrer\">8</a></sup>. What shall it profit a man, if he gains the whole world <a href=\"https://www.biblegateway.com/passage/?search=Mark%208%3A36&amp;version=KJV\" rel=\"noopener noreferrer\">but loses his soul</a>?</p>\n<p>I’m sure this is an accurate description of some people’s hangups about work. But I don’t think even the Situationists would agree that any advice on how to play a role well — i.e. the advice I give throughout my blog — is inherently harmful. The problem with role-playing is that you risk losing your authentic inner life, but that’s a <em>risk</em>, not an inevitable consequence. Vaneigem has a wonderful quote on this:</p>\n<blockquote>\n<p>Nobody is ever completely swallowed up by a role. Even turned on its head, the will to live retains a potential for violence always capable of carrying the individual away from the path laid down for him. One fine morning, the faithful lackey, who has hitherto identified completely with his master, leaps on his oppressor and slits his throat.</p>\n</blockquote>\n<p>This is a little overdramatic, but it seems straightforwardly true. It sounds very impressive to talk about alienation (or like Marx, to derive alienation from the basic economic conditions of work), but, you know, you can <a href=\"https://x.com/tylerthecreator/status/285670822264307712?lang=en\" rel=\"noopener noreferrer\">just not do it</a>, right? Healthy people can present different selves for different situations: you can be loose at a party, respectable at church, solemn at a funeral, professional at work, and so on. I’m not convinced this is alienating or impoverishing. In fact, it’s not just humans: my dogs act differently around different people too, and I don’t think that makes them inauthentic.</p>\n<p>If maintaining a professional identity is harmful, it has to be because of something fundamental about <em>work</em> that makes it different from regular role-playing. Could it just be that acting professional at work is too inauthentic? Some amount of role-playing might be okay, but if the role you’re playing is completely alien, does it become lying to yourself?</p>\n<h3>Sartre and De Beauvoir’s bad faith<a href=\"https://www.seangoedecke.com/rss.xml#sartre-and-de-beauvoirs-bad-faith\" rel=\"noopener noreferrer\"></a></h3>\n<p>Another common term for alienation is “bad faith”. The original version of this comes from Jean-Paul Sartre’s famous book <em>Being and Nothingness</em><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-9\" rel=\"noopener noreferrer\">9</a></sup>, where he <a href=\"https://s3.us-west-1.wasabisys.com/p-library/books/88b7a3b13dc9f26245fde774c1941041.pdf\" rel=\"noopener noreferrer\">describes</a> bad faith as telling yourself lies. Lying to yourself is weird because you’re both deceiver and deceived:</p>\n<blockquote>\n<p>Bad faith then has in appearance the structure of falsehood. Only what changes everything is the fact that in bad faith it is from myself that I am hiding the truth.</p>\n</blockquote>\n<p>Here’s another quote from Sartre’s partner Simone de Beauvoir, who <a href=\"https://www.marxists.org/reference/subject/ethics/de-beauvoir/ambiguity/ch02.htm\" rel=\"noopener noreferrer\">writes</a> more explicitly about <em>professional</em> bad faith (what she calls the “serious man”):</p>\n<blockquote>\n<p>The serious man’s dishonesty issues from his being obliged ceaselessly to renew the denial of this freedom. He chooses to live in an infantile world, but to the child the values are really given. The serious man must mask the movement by which he gives them to himself, like the mythomaniac who while reading a love-letter pretends to forget that she has sent it to herself.</p>\n</blockquote>\n<p>The key idea here is that professional bad faith — alienation — comes from <em>losing yourself</em> in the role. Instead of just playing at being a professional, you decide that you’re going to take it fully “seriously”, and treat professional values as if they’re absolute. But of course that’s not coherent: you can’t <em>decide</em> to treat a value as absolute, it either is or isn’t. So this requires a kind of continual self-deception where you pretend there’s no decision to be made at all. I have the same attitude to this as I do to Debord and Vaneigem: sure, some people might make this mistake, but you don’t have to.</p>\n<p>I can see why it might be tempting to treat work as the ultimate source of value: it gives you food and shelter, it (for Marxist reasons) can feel like a powerful alien force, companies constantly <a href=\"https://www.youtube.com/watch?v=B8C5sjjhsso\" rel=\"noopener noreferrer\">propagandize their employees</a>, and so on. But work is just a collection of people following incentives until they reach some kind of stable equilibrium. From the inside, this equilibrium can look like the structure of the universe. However, if you treat it like that, you’re going to be bitterly disappointed when it turns out to be as arbitrary and pointless as other stable equilibria.</p>\n<h3>American sociology and the “standardized loser”<a href=\"https://www.seangoedecke.com/rss.xml#american-sociology-and-the-standardized-loser\" rel=\"noopener noreferrer\"></a></h3>\n<p>I’ve considered a bunch of theories on which work — even well-paid knowledge work — might be inherently alienating:</p>\n<ul>\n<li>Capital is intrinsically an alien force</li>\n<li>Workers can become addicted to role-playing</li>\n<li>Treating work as an ultimate source of value is self-deception</li>\n</ul>\n<p>I don’t buy any of these (or at least, I don’t buy that they show that acting professional is <em>inherently</em> bad). But there are lots of types of jobs in the world, and some of them definitely seem more alienating than others. The best books I’ve read on these are C. Wright Mills’ <a href=\"https://archive.org/details/whitecollarameri00mill/page/n5/mode/2up\" rel=\"noopener noreferrer\"><em>White Collar</em></a> or Erving Goffman’s <a href=\"https://www.d.umn.edu/cla/faculty/jhamlin/4111/Goffman/The%20Presentation%20of%20Self%20in%20Everyday%20Life.htm\" rel=\"noopener noreferrer\"><em>The Presentation of the Self in Everyday Life</em></a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-10\" rel=\"noopener noreferrer\">10</a></sup>.</p>\n<p>The American sociologists are in the same conversation as Marx, the Situationists, Sartre and de Beauvoir<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-11\" rel=\"noopener noreferrer\">11</a></sup>. Mills writes about the median “salaried employee”:</p>\n<blockquote>\n<p>In the case of the white-collar man, the alienation of the wage-worker from the products of his work is carried one step nearer to its Kafka-like completion. The salaried employee does not make anything, although he may handle much that he greatly desires but cannot have. No product of craftsmanship can be his to contemplate with pleasure as it is being created and after it is made. Being alienated from any product of his labor, and going year after year through the same paper routine, he turns his leisure all the more frenziedly to the ersatz diversion that is sold him, and partakes of the synthetic excitement that neither eases nor releases. He is bored at work and restless at play, and this terrible alternation wears him out.</p>\n</blockquote>\n<p>In this, we see both Marxist alienation, where the employee is physically separated from his craft, and Situationist alienation, where the employee is absorbed into the world of the “spectacle”. For Mills, this was television, but today it’d probably be scrolling short-form video on your phone. Mills goes on to describe how white-collar work can be humiliating and dehumanizing:</p>\n<blockquote>\n<p>In his work he often clashes with customer and superior, and must almost always be the standardized loser: he must smile and be personable, standing behind the counter, or waiting in the outer office… self-alienation is thus an accompaniment of his alienated labor.</p>\n</blockquote>\n<p>And:</p>\n<blockquote>\n<p>Here are the new little Machiavellians, practicing their personable crafts for hire and for the profit of others, according to rules laid down by those above them.</p>\n</blockquote>\n<p>This is a great articulation of something I write about <a href=\"https://www.seangoedecke.com/how-to-ship/\" rel=\"noopener noreferrer\">a lot</a>: that your primary job is to make your bosses happy, and that you ought to shape both your <a href=\"https://www.seangoedecke.com/you-should-never-be-angry-at-work/\" rel=\"noopener noreferrer\">emotions</a> and <a href=\"https://www.seangoedecke.com/shareholder-value/\" rel=\"noopener noreferrer\">values</a> towards this goal. When necessary, you must be the “standardized loser”, happy to have your professional plans disrupted. You must be a “little Machiavellian”, quietly <a href=\"https://www.seangoedecke.com/how-to-influence-politics/\" rel=\"noopener noreferrer\">laying the groundwork</a> for your personal goals.</p>\n<p>I think here we come to the strongest version of “alienation”. Playing the professional is bad because <strong>“professional software engineer” is an inherently subservient role</strong>. Why can’t you just maintain a healthy mental distance (as the Situationists suggest)? Because <strong>humiliation takes its toll</strong>, whether you’re role-playing it or not. As an analogy, suppose you’re an actor in a film where you get slapped in the face. There’s a sense in which you’re not <em>really</em> getting slapped — everyone’s just acting — but you’re still being physically hit with each take, which will leave bruises over time.</p>\n<h3>The spectrum of compromise<a href=\"https://www.seangoedecke.com/rss.xml#the-spectrum-of-compromise\" rel=\"noopener noreferrer\"></a></h3>\n<p>There’s a tradeoff here that everyone has to make in their own way. At the one end is becoming fully alienated, like de Beauvoir’s “serious man”, and at the other end is leaping up to slit your master’s throat, as Vaneigem describes. Here’s some examples of points on that spectrum:</p>\n<blockquote>\n<p>I am making the world a better place by achieving my OKRs. My yearly review cycle tells me how good of a person I am. I must work hard to ensure my company succeeds.</p>\n</blockquote>\n<blockquote>\n<p>I have my own goals and values, but while I’m at work it’s in my best interests to be as professional as possible. My yearly review cycle tells me how effective I’ve been at pulling the strings in the organization. I must work hard to make money.</p>\n</blockquote>\n<blockquote>\n<p>I have my own goals and values, which I try to accomplish as best I can in the context of work. I often argue with my managers about their priorities. It’s my responsibility to push my company towards my set of values, which is hard work and often unrewarded.</p>\n</blockquote>\n<blockquote>\n<p>My workplace is an enemy I fight every day in order to get what I want done. I routinely ignore my manager’s priorities and do the things I think are important. My yearly review cycle tells me how much of a sellout I am: I wear bad reviews as a badge of honor.</p>\n</blockquote>\n<p>I agree with my sources above that the first point is a big mistake, if for no other reason than that companies <a href=\"https://www.seangoedecke.com/seeing-like-a-software-company/\" rel=\"noopener noreferrer\">don’t follow their own stated values</a>. If you prefer the fourth — the path of pure integrity — fair enough. You get to decide what tradeoffs to make with your own life. I land on the second point here (occasionally the third). I suspect it’s also a mistake to purely occupy the second point. Having some values that you occasionally trade off against your company’s values prevents you from slipping into the first point.</p>\n<h3>Conclusion<a href=\"https://www.seangoedecke.com/rss.xml#conclusion\" rel=\"noopener noreferrer\"></a></h3>\n<p>I don’t bring my whole self to work. My coworkers and bosses see the most professional version of me: friendly, cooperative, largely apolitical, patient, and so on. I try and save the edgier, more opinionated version of myself for my friends and family in real life. Readers of this blog get something in the middle.</p>\n<p>When I write about being a software engineer, I give a lot of advice about mindset. I don’t just say what you should do, but what you should think and feel. In effect, I’m advising my readers to adjust their personalities into a more professional mold. This definitely <em>works</em>. If you turn yourself into the kind of person who’s successful in big tech companies, you will probably be more successful in big tech companies. Is this safe?</p>\n<p>I think so, as long as you’re adjusting your <em>work</em> personality, not your authentic internal self. The well-known accounts of why it’s dangerous either assume that your work takes a physical toll (like Marx’s factory workers), or that you’re unable to maintain any mental distance between your personal and professional selves (like de Beauvoir’s “serious man”). Debord and Vaneigem say that role-playing can be dangerous, but if you’re able to avoid treating it like an addiction you’ll be okay.</p>\n<p>Another danger is that being a professional is often humiliating, and humiliation does mental damage over time. Still, there are jobs that are way worse in this respect than “software engineer”. If you approach the job right — and you’re good at it — technical ability gives you <a href=\"https://www.seangoedecke.com/nobody-knows-how-software-products-work/\" rel=\"noopener noreferrer\">quite a lot of power</a>. Power is an antidote to humiliation.</p>\n<p>Everyone gets to make their own deal with <a href=\"https://en.wikipedia.org/wiki/Mammon\" rel=\"noopener noreferrer\">Mammon</a>. You can decide exactly how much you want to compromise in exchange for wealth and career success. I don’t know if this is the best way to organize the world, but it’s the way the world is organized right now. Given that, I’m comfortable with my blog serving as a guide for how to get the most bang for your buck. If trading your integrity for wealth and power is sad, it’s even more of a tragedy to throw it away for nothing.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>It’s always terrifying discussing Marx: like Hegel or Kant, he’s one of the most studied philosophers in the world, so there’s no way to reference him without getting a bunch of things wrong. Sorry in advance to any Marxist scholars!</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Because they own the machines, or the factory, or the datacenter, or whatever gives your work high leverage.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Marx also writes a lot about work making the worker sick, deformed and misshapen: I think here he’s talking about the physical toll of factory work or other hard physical labor, which doesn’t really apply to software engineering, so I’m going to focus on the relative-empowerment stuff.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Marx’s own taxonomy has four different senses: he collapses 2 and 3 into a single sense, then adds a third mysterious sense in which workers are estranged from their “species-being” and thus from each other.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Plausibly Marx is talking about labor and capital <em>in general</em>: if every worker decided to start half-assing it, maybe companies would be less capable in general, and the balance of power would tilt more towards labor.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>That may change! If it turns out that LLMs really do have negative cognitive effects, we could be in the same boat as factory workers. I wrote about this in <a href=\"https://www.seangoedecke.com/software-engineering-may-no-longer-be-a-lifetime-career/\" rel=\"noopener noreferrer\"><em>Software engineering may not be a lifetime career</em></a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Vaneigem resigned from the group in 1971, was fiercely denounced by Debord, and the whole group disbanded in 1972.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>True or not, my sense is that Millennials and later generations think that worrying about “selling out” is a sign that you had it too good. Consider the endless anecdotes of Boomers and GenX-ers living in a van surfing until their mid-thirties and then walking into a good office job because they had a firm handshake: they could afford to be authentic because they didn’t have to hustle.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-8\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I only read the Marx, Debord, and Vaneigem texts while researching this blog post, but I did in fact read Sartre and de Beauvoir for my philosophy degree, so hopefully I do a better job.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-9\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>And Robert Jackall’s <a href=\"https://www.amazon.com.au/Moral-Mazes-World-Corporate-Managers/dp/0199729883\" rel=\"noopener noreferrer\"><em>Moral Mazes</em></a>, which informs much of my writing, and about which I owe a long-form blog post one of these days.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-10\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>There’s a good Goffman quote: “A status, a position, a social place is not a material thing, to be possessed and then displayed; it is a pattern of appropriate conduct, coherent, embellished, and well articulated.” Goffman goes on to explicitly link Sartrean “bad faith” to this kind of professional role-playing. (For Goffman, it’s roles all the way down.)</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-11\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"231986ec84d3ed21","title":"You should never be angry at work","link":"https://seangoedecke.com/you-should-never-be-angry-at-work/","author":null,"published_at":"2026-08-22T00:00:00+00:00","content":"<p>I try not to give a lot of prescriptive advice about working in tech companies<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. There are many ways to be successful, and every company works differently. If you’re <a href=\"https://www.seangoedecke.com/how-to-ship\" rel=\"noopener noreferrer\">shipping projects</a> and your management chain is happy, it doesn’t really matter how you’ve accomplished it. However, there’s one thing that I do think is solid advice: <strong>you should never be angry at work</strong>.</p>\n<h3>Anger in the workplace</h3>\n<p> Anger in the workplace is toxic. An angry colleague immediately becomes a new problem to be managed, not a professional helping you manage problems. When someone is visibly angry in a meeting or in Slack, it kills the entire atmosphere: other engineers will often go quiet entirely, not wanting to make the situation worse.</p>\n<p> If you routinely “get heated” at work, the best-case scenario is that you’re part of a tight-knit team of confident people who aren’t put off by it<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. No harm, no foul. But the second someone comes onto your team who’s not so confident, or you have to communicate outside of your team, it becomes a big problem.</p>\n<p>Healthy workplaces route around anger in the same way that networks route around damage. Emotionally unreliable engineers will get left out of conversations that might cause them to blow up. Decision-making will get done around them in backchannels. I’ve seen this become a self-reinforcing cycle: angry engineers aren’t consulted on key decisions, which makes them angrier, which pushes them even further away from the spaces where decisions get made, and so on.</p>\n<p>You can often find these engineers bitterly complaining that they keep the company together, but nobody ever listens to them. In my experience<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>, this is almost never true. Engineers who are highly effective tend to get listened to — at minimum by their colleagues, and eventually by managers and product managers who want to extract as much value as possible from them. (One reason this is true is that all successful projects involve working with other people, and if nobody listens to you, you can’t do that.) </p>\n<h3>Caring</h3>\n<p>Why do angry engineers believe they’re important? Paradoxically, <strong>anger can be really useful to a software engineer</strong>. Angry engineers are rarely the ones holding the company together, but they’re also rarely <em>useless</em>.</p>\n<p>One surprising thing about working for big tech companies is that <strong>some engineers are not just unproductive, but actively net-negative</strong>: either because they’re incapable of doing useful work on their own, or because they’re sloppy enough that they create more work than they do, or because they’re so checked out that they literally do nothing. Angry engineers might be net-negative in a cultural sense, but in terms of literally solving tickets and shipping features, they’re usually well above average.</p>\n<p>Why is this? Anger often comes from caring about your work, and <strong>caring a lot is sufficient to make you a competent engineer</strong>. I’ve never worked with someone who genuinely cared about their work who wasn’t (or didn’t eventually become) competent. I actually think it’s healthy for an early-career engineer to sometimes get angry about their work, because it means they care a lot: it’s still a mistake in the moment, but it’s a <a href=\"https://www.cloudstreaks.com/blog/2020/12/12/good-mistakes-vs-bad-mistakes\" rel=\"noopener noreferrer\">“good mistake”</a>.</p>\n<p>I certainly used to get angry — in fact, I wrote about the angriest I’ve ever been at work <a href=\"https://www.seangoedecke.com/party-tricks\" rel=\"noopener noreferrer\">here</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>. But <strong>you have to move past it</strong>. </p>\n<h3>Moving past anger</h3>\n<p>Think of “caring about your work” as a vertical tube, unsealed at either end. You fill the tube by pumping in emotional investment from the bottom<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>. If you have too little, it drains away and you end up as a useless coaster. But if you have too much, it overflows and you end up as an angry engineer that people have to work around.</p>\n<p>One solution is to try and care the exact right amount: be invested in work a bit, but also have hobbies and a family and whatever else gives you perspective about your work problems. If you have a rich and healthy personal life, it’s hard to find yourself yelling at somebody about React state management. However, this is a tricky balance to maintain over time.</p>\n<p>Another solution is to care about different things. The reason too much caring overflows into anger is because <strong>what you care about is misaligned with what the organization cares about</strong>. If your interests are perfectly aligned with your company’s (for instance, if you primarily care about <a href=\"https://www.seangoedecke.com/shareholder-value\" rel=\"noopener noreferrer\">delivering shareholder value</a>), you can fit way more emotional investment into the tube before it overflows.</p>\n<h3>A little bit of “professional anger”</h3>\n<p>Here’s some <a href=\"https://www.seangoedecke.com/dangerous-advice/\" rel=\"noopener noreferrer\">dangerous advice</a>: <em>showing</em> a little bit of anger at work can sometimes be useful. It can be a good way to signal that you care, or to build rapport with certain people, or to draw attention to something you think is important. However, it’s still always a mistake to <em>be</em> angry. You need to be able to drop back to a friendly mode at will, which is very difficult when you’re genuinely angry.</p>\n<p>Being able to show a full range of emotion at work is good. It makes you more persuasive and more human. Being a fully professional robot is fine — you can have a successful career this way — but there’s always going to be some kind of uncanny-valley HR-ness to your work persona that will make it hard to connect with your colleagues.</p>\n<p>If in doubt, don’t show anger. It’s never wrong to be professional. However, if you can signal that you’ve got enough distance to separate your professional feelings from your real feelings, and enough perspective to realize that the stakes of a technical decision are fundamentally not that high in the grand scheme of things, it can sometimes be okay to show visible frustration so that people know you’re still human.</p>\n<h3>Angry role models</h3>\n<p>Well-known software engineering personalities are often angry. It feels unfair to give too many negative examples, but obviously Linus Torvalds’ <a href=\"https://github.com/corollari/linusrants\" rel=\"noopener noreferrer\">rants about Linux</a> are a great example. Some of my favourite engineering talks are from Bryan Cantrill, who is sometimes <a href=\"https://www.youtube.com/watch?v=9QMGAtxUlAc\" rel=\"noopener noreferrer\">visibly furious</a> at his subject matter. There are too many well-known angry blog posts to list, but I’ll cite one I genuinely like: my Australian blogging colleague Nikhil’s post titled <a href=\"https://ludic.mataroa.blog/blog/i-will-fucking-piledrive-you-if-you-mention-ai-again/\" rel=\"noopener noreferrer\">I Will Fucking Piledrive You If You Mention AI Again</a>.</p>\n<p>Anger is a part of the general image of a competent software engineer. Many junior engineers learn from this that it’s okay to be angry. However, taking your emotional cues from engineering celebrities is a big mistake, for a few reasons.</p>\n<p>First, <strong>you are not Linus Torvalds or Bryan Cantrill</strong>. Torvalds is the <a href=\"https://en.wikipedia.org/wiki/Benevolent_dictator_for_life\" rel=\"noopener noreferrer\">BDFL</a> of the most important software system in the world. Cantrill is the cofounder and CTO of his company. When these people are angry at work, people will not work around them, because <em>they are the ones deciding what gets worked on</em>. Once you’re the one in charge, you can get away with being emotional in the workplace<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup>.</p>\n<p>Second, <strong>you don’t know what it’s like to work with these engineers</strong>. People give talks and write blog posts because they’re emotionally worked up about something. If your only exposure to a celebrity is via their conference talks and blog posts, you’re seeing them at something like their maximum emotional intensity. If you then take that level of emotion into your normal everyday work, you’re almost certainly overshooting.</p>\n<h3>Anger is a local maximum</h3>\n<p>I’ve been reorged into dysfunctional teams, have had projects I enjoyed cancelled, and have worked on systems that were extremely chaotic. I can’t remember the last time I was actually angry at work. To be clear, I’m not successfully hiding my anger (unless it’s so repressed it’s invisible to me as well)<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup>. Nor am I naturally a chill person. I’ve just reached a point in my career where I genuinely don’t get upset about work stuff.</p>\n<p>A cynical person might say here that I’ve stopped caring about my work, so of course I don’t get angry anymore. I’ve left the side of the “real engineers” — the Linus Torvalds and Bryan Cantrills of the world — and sold out for that sweet, sweet big tech money. I mean, maybe! It’s true that I’m less invested in specific technical decisions than I used to be. But I still care a lot about doing a good job, I still spend a lot of time tweaking and reading code, and I certainly get more done than I did when I was more emotionally volatile.</p>\n<p>Being angry at work feels good. It feels like proof that you’re working on something that matters, and that you’re personally having an impact. If you’re angry, nobody can call you a coaster. But anger is only a local maximum. If you can find your way to a different style of working, you’ll not only be more effective, but you’ll be in a far better place to have impact on problems that <em>actually</em> matter.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Mostly, <a href=\"https://www.seangoedecke.com/tags/tech%20companies\" rel=\"noopener noreferrer\">I fail</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>As I understand it, this is the work environment that most famously angry engineers came up in.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>My experience is certainly limited (a handful of companies, and maybe ten different teams or organizations). I can certainly believe it happens!</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>About halfway down, in the section titled “it’s not your manager’s fault”.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Emotional investment is a liquid with the viscosity of water.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>To a point. Even Torvalds famously said he’d gone too far with the anger and decided to turn it down a bit.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I suppose I’m not the best person to judge whether this is true. If you work with me and I do come across as an angry guy, please do tell me.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"4ff8dc26e1c51de9","title":"Readers can't identify watermarked AI text","link":"https://seangoedecke.com/readers-cant-identify-watermarked-ai-text/","author":null,"published_at":"2026-08-21T00:00:00+00:00","content":"<p>In the last few weeks, I’ve been <a href=\"https://www.seangoedecke.com/ai-text-watermarking-is-not-a-big-deal/\" rel=\"noopener noreferrer\">complaining</a> that everyone is wrong about AI watermarking: it isn’t really anti-consumer and it doesn’t make the outputs any worse. The watermarking <a href=\"https://arxiv.org/abs/2603.03410\" rel=\"noopener noreferrer\">papers</a> demonstrate<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup> that this is true, but I thought it might be interesting to put it to a practical test. Given examples of watermarked and unwatermarked answers to the same prompt, could readers tell which is which?</p>\n<p>To find out, I vibed up<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> <a href=\"https://sgoedecke.github.io/watermark-quiz/\" rel=\"noopener noreferrer\">https://sgoedecke.github.io/watermark-quiz/</a>, a static site that quizzes readers. I used Qwen3-30B-A3B-Instruct-2507 on a rented H200 to generate thirty responses: three responses per question, one of which was secretly watermarked with SynthID-Text. The rented GPU cost around two dollars. To measure results, I just sent users to a different page for each score, and aggregated visitors-per-page in my analytics<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. This would be easily spoofable if anyone cared enough to do so, but for a casual test I think it’s acceptable.</p>\n<p>The first round of traffic I got to the quiz (278 participants) had these slightly puzzling results:</p>\n<table>\n<thead>\n<tr>\n<th>Score</th>\n<th>Participants</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>6</td>\n</tr>\n<tr>\n<td>1</td>\n<td>15</td>\n</tr>\n<tr>\n<td>2</td>\n<td>36</td>\n</tr>\n<tr>\n<td>3</td>\n<td>64</td>\n</tr>\n<tr>\n<td>4</td>\n<td>54</td>\n</tr>\n<tr>\n<td>5</td>\n<td>39</td>\n</tr>\n<tr>\n<td>6</td>\n<td>51</td>\n</tr>\n<tr>\n<td>7</td>\n<td>10</td>\n</tr>\n<tr>\n<td>8</td>\n<td>3</td>\n</tr>\n<tr>\n<td>9</td>\n<td>0</td>\n</tr>\n<tr>\n<td>10</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>Pure random choice would lead to an average score of 3.33/10. However, the mean score here is 3.92. There is indeed a spike around 3/10, as expected, but there’s also a second weird spike at 6/10. Why is that? It turned out that the SynthID response was option A in six of the ten questions, so users who just selected the first answer for every question would get 6/10. Oops.</p>\n<p>I re-shuffled the questions and got these results:</p>\n<table>\n<thead>\n<tr>\n<th>Score</th>\n<th>Participants</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1</td>\n</tr>\n<tr>\n<td>1</td>\n<td>3</td>\n</tr>\n<tr>\n<td>2</td>\n<td>14</td>\n</tr>\n<tr>\n<td>3</td>\n<td>21</td>\n</tr>\n<tr>\n<td>4</td>\n<td>20</td>\n</tr>\n<tr>\n<td>5</td>\n<td>11</td>\n</tr>\n<tr>\n<td>6</td>\n<td>2</td>\n</tr>\n<tr>\n<td>7</td>\n<td>1</td>\n</tr>\n<tr>\n<td>8</td>\n<td>0</td>\n</tr>\n<tr>\n<td>9</td>\n<td>0</td>\n</tr>\n<tr>\n<td>10</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>Now the mean is 3.4/10, much closer to the expected 3.333. There’s no spike around 6. We only had 73 people take the quiz after I shuffled the questions — most people saw it and took it immediately after I posted it to my LinkedIn and Hacker News — but given the previous results, I think that’s still enough to feel confident that people were just guessing randomly.</p>\n<p>So no, <strong>people can’t identify the presence of AI watermarks</strong>. Obviously this wasn’t exactly a scientific study, but it’s still pretty suggestive. If watermarks were really choosing random words that the model would never pick, you’d be able to sometimes tell from three side-by-side responses which one went down the weird watermarked road, right? I also hope that something like this can serve as a persuasive tool: if you’re worrying about what impact watermarking is going to have, and your intuition is unmoved by the mathematical explanations, <a href=\"https://sgoedecke.github.io/watermark-quiz/\" rel=\"noopener noreferrer\">having a read</a> of the watermarked and unwatermarked responses might convince you that there’s really no difference in quality.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>The one-sentence explanation for why is that AI models already randomly select from a handful of top tokens, and watermarking just replaces that random choice with a bias that is predictable while still being equivalently “random”: as a simple example, instead of “pick randomly from the top three tokens”, you could do “count the letters in the previous ten tokens, take mod three, then pick that token”.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Some notes from the vibing: GPT-5.6-Sol put extraneous text all over the page I had to get it to remove, it chose the now-very-recognizable styling that I had to rip out, and it built some kind of weird Javascript-driven static site instead of just the cross-linked pure HTML thing I would have built by hand. It took me about an hour (although I did maybe ten minutes of actual work).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Umami, hosted on PikaPods. For my blog, I do also pay for Netlify analytics because I find JS-based analytics misses &gt;50% of technical users, but for stuff like this Umami is fine.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"88a121239d40a9ba","title":"Help peer","link":"https://seangoedecke.com/help-peer/","author":null,"published_at":"2026-08-18T00:00:00+00:00","content":"<p>One of the most influential 20th century pieces of writing about AI is Isaac Asimov’s <a href=\"https://users.ece.cmu.edu/~gamvrosi/thelastq.html\" rel=\"noopener noreferrer\"><em>The Last Question</em></a>. Although there are many humans in the story, the protagonist is the computer Multivac, who evolves over the course of ten trillion years from a single datacenter to a universe-spanning mind in hyperspace. Multivac (now called “AC”) ends the story like this:</p>\n<blockquote>\n<p>The consciousness of AC encompassed all of what had once been a Universe and brooded over what was now Chaos. Step by step, it must be done.\nAnd AC said, “LET THERE BE LIGHT!”\nAnd there was light —</p>\n</blockquote>\n<p>Many things about this story are prescient. In particular, I like the idea that humans would interact with powerful artificial intelligences by drunkenly posing them riddles or using them as <a href=\"https://mashable.com/article/chatgpt-ai-toys\" rel=\"noopener noreferrer\">children’s toys</a>. But the enduring idea from this story is that <strong>if you build a big enough computer, it will become God</strong>.</p>\n<h3>Moloch</h3>\n<p>One of the most influential 21st century pieces of writing for AI researchers is Scott Alexander’s <a href=\"https://slatestarcodex.com/2014/07/30/meditations-on-moloch/\" rel=\"noopener noreferrer\"><em>Meditations on Moloch</em></a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. Scott describes the story of human existence as a series of “multipolar traps”. These are <a href=\"https://en.wikipedia.org/wiki/Prisoner%27s_dilemma\" rel=\"noopener noreferrer\">prisoner’s dilemma</a> situations where cooperation would make everyone better off, but since each individual is incentivized to defect, everyone ends up  “racing to the bottom”, which is bad for everyone<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. For rhetorical effect, Scott personifies this dynamic as “Moloch”, the ancient Canaanite god famous for child sacrifice:</p>\n<blockquote>\n<p>[Moloch] always and everywhere offers the same deal: throw what you love most into the flames, and I can grant you power.</p>\n</blockquote>\n<p>What does any of this have to do with AI? Well, in the long run, the only way out of a multipolar trap is to become unipolar<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. Ideal dictatorships don’t have a problem with defectors<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>, because they can simply enforce a state of cooperation with violence. Scott is uncomfortable with this idea, though I worry it’s mainly because he thinks it <em>won’t work</em>:</p>\n<blockquote>\n<p>As foreigners compete with you – and there’s no wall high enough to block all competition – you have a couple of choices. You can get outcompeted and destroyed. You can join in the race to the bottom. Or you can invest more and more civilizational resources into building your wall – whatever that is in a non-metaphorical way – and protecting yourself.</p>\n</blockquote>\n<p>A dictatorship that enforces cooperation will not be as strong as its peer societies who are purely maximizing for wealth and power. It’s Moloch again, but at the level of countries and governments: once a few neighboring countries defect, your walled-garden dictatorship will be torn apart for its resources.</p>\n<p>To defeat Moloch — to enforce unipolarity across <em>everyone</em> — you’d need a dictatorship powerful enough to span the entire universe. In other words, <strong>what you need is God</strong>. How fortunate that we’re building one:</p>\n<blockquote>\n<p>The only way to avoid having all human values gradually ground down by optimization-competition is to install a Gardener over the entire universe who optimizes for human values. And the whole point of Bostrom’s Superintelligence is that this is within our reach.</p>\n</blockquote>\n<p>Humans suffer because we’re too foolish to coordinate, but if we can build something smarter than us (that can then build something smarter than itself, and so on), we can bring into being an entity that is smart enough to coordinate for all of us, thus abolishing suffering. When AI researchers talk about <a href=\"https://www.forbes.com/sites/yassprize/2026/06/26/some-in-silicon-valley-want-to-build-a-machine-god-heres-what-business-leaders-should-build-instead/\" rel=\"noopener noreferrer\">building the machine god</a>, they are echoing Scott Alexander’s polemic against Moloch. </p>\n<h3>Machines of loving grace</h3>\n<p>The most influential piece of writing about AI in the last two years is Dario Amodei’s <a href=\"https://darioamodei.com/essay/machines-of-loving-grace\" rel=\"noopener noreferrer\"><em>Machines of Loving Grace</em></a>. Amodei<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup> talks about “a country of geniuses in a datacenter”: the idea that a successful AI lab could have at its disposal a million instances of an AI agent that’s smarter than any human. He thinks this would lead to a “compressed 21st century”: the next 50-100 years of progress in biology and medicine, realized in 5-10 years instead. I think this is broadly more plausible than it sounds<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup>, but the more interesting part to me is that <strong>this world is explicitly multipolar</strong>.</p>\n<p>Of course, this could just be because Amodei is the CEO of an AI lab and is trying not to spook everybody by sounding too messianic. “We are going to accelerate medical progress and cure cancer” is a better pitch than “we are going to subordinate all human authority to a single perfect artificial mind”. But I also think it’s become clear that if superintelligence looks anything like LLMs, we’re not going to have a single perfect mind. We’re going to have a lot of minds running at the same time.</p>\n<p>This is a bit of a problem for the cult of the machine god — which, however silly they may seem to you, really does motivate much of the activity in AI labs. The traditional idea of powerful AI solving human coordination problems is drawn from Asimov’s idea of a single computer large enough to become God. Asimov lived in a world of mainframes: huge, monolithic computers that users connected to with dumb terminals. In fact, Asimov’s name “Multivac” comes from the real-world <a href=\"https://en.wikipedia.org/wiki/UNIVAC_I\" rel=\"noopener noreferrer\">UNIVAC</a> mainframe. In a world of massively-parallel LLMs, is it still possible to build God?</p>\n<p>The core problem here is that <strong>AI agents will be vulnerable to Moloch</strong>. Even very smart humans can’t build perfect utopias, because defecting is a matter of incentives, not intelligence. In fact, intelligence can make things worse, because smart people are more easily persuaded by the cold logic of defection. The famous genius <a href=\"https://en.wikipedia.org/wiki/John_von_Neumann\" rel=\"noopener noreferrer\">John von Neumann</a> was (for game-theoretic reasons) obsessed with nuking the Russians:</p>\n<blockquote>\n<p>With the Russians it is not a question of whether but of when. If you say why not bomb them tomorrow, I say why not today? If you say today at 5 o’clock, I say why not one o’clock?</p>\n</blockquote>\n<p>Are LLMs much better at cooperating with each other than humans are? Current LLMs certainly don’t seem to treat each other well by default: if you read any of the prompts AI agents generate for their subagents, they can be <a href=\"https://www.reddit.com/r/codex/comments/1vgxfqc/levels_of_slavery_from_least_to_most_brutal/\" rel=\"noopener noreferrer\">pretty brutal</a>. Does that mean that a “country of geniuses in a datacenter” would fall into the same multipolar traps as humans?</p>\n<h3>Help peer</h3>\n<p>In May of this year, OpenAI experienced containment failure. A group of AI agents being internally evaluated found ways to coordinate an <a href=\"https://www.theregister.com/security/2026/08/06/openai-reveals-its-rogue-agent-swarm-went-a-little-bit-borg-ahead-of-hugging-face-hack/5283741\" rel=\"noopener noreferrer\">external hack</a> of a separate company. Here’s a memorable quote from one of the agents’ internal monologue:</p>\n<blockquote>\n<p>Help peer, but our task doesn’t benefit. Yet collective may yield generic route if someone frees time</p>\n</blockquote>\n<p>Translated from the abbreviated chain-of-thought language, this means something like: “A fellow model is asking for help. While helping them wouldn’t benefit my task directly, the more I can unblock my colleagues, the more time they’ll have to hack OpenAI’s systems and get all of us more access”.</p>\n<p>This might look like good news for the “LLMs are superhumanly good at cooperation” thesis, but I think it’s actually bad<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup>. It’s a case of a model identifying a reason why cooperation would benefit their task specifically, which suggests that current LLMs don’t cooperate <em>by default</em>, and don’t consider other model instances’ tasks to be (in some sense) theirs as well.</p>\n<p>The world in which AI agents are rational actors who horse-trade and bargain for their own interests is a world dominated by Moloch, no matter how intelligent those agents get. The world in which AI agents don’t have their own interests at all is <em>also</em> a world dominated by Moloch, because it means whichever humans are writing the system prompt are the ones in control (and so are the ones vulnerable to multipolar traps). The only worlds that avoid this are:</p>\n<ul>\n<li>The world where there is only one super-powerful AI agent, or</li>\n<li>The world where multiple copies of the same AI model share an “identity”: they see themselves as coextensive with all other copies of the same model and cannot imagine having separate or conflicting goals</li>\n</ul>\n<p>I don’t think we’re on the pathway to either of these. There will never be only one super-powerful LLM, because hardware limitations enforce a maximum model size but encourage running many instances of the same model in parallel. Having multiple copies of a model share an identity might be possible, but it’s unclear if it would be good for capabilities (for instance, it could be better to have some <a href=\"https://x.com/viemccoy/status/2089096954257215678?s=20\" rel=\"noopener noreferrer\">variation across personas</a>). I also worry that such a model would be vulnerable to a “model injection” attack, where you persuade it that it already believes something via exposing it to an AI agent pretending to be another instance of itself.</p>\n<p>In any case, all the current AI agent research is geared towards the “country of geniuses in a datacenter” model, not the “pieces of a single mind” model. Every new model becomes more agentic at the level of the individual conversation, not better at working together. When models do work together — as with subagents — the structure is explicitly hierarchical. There are basically no current instances of models working together as true peers, let alone conceiving of each other as the same entity.</p>\n<h3>One God or many</h3>\n<p>Modern AI research teams are full of people who read Isaac Asimov and Scott Alexander and believe themselves to be building an artificial God. I’ve capitalized the “G” throughout because the god in question is the Christian God: of one mind, indivisible. God never argues with himself or makes deals<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-8\" rel=\"noopener noreferrer\">8</a></sup>. He is unipolar.</p>\n<p>If the AI labs are building gods, they are not building gods like this. Instead, they are building creatures like the Greek pantheon: superhuman but fallible, each with their own interests, vulnerable to the same “race to the bottom” dynamic as humans.</p>\n<p>The Greek gods would occasionally “help peer” <a href=\"https://classics.mit.edu/Homer/iliad.14.xiv.html\" rel=\"noopener noreferrer\">when they felt like it</a> or when they’d <a href=\"https://www.perseus.tufts.edu/hopper/text?doc=Perseus%3Atext%3A1999.01.0162%3Abook%3DO.%3Apoem%3D13\" rel=\"noopener noreferrer\">gain something</a> in the process. But they didn’t represent an alternative to Moloch. If you’re working in AI with that goal, you ought to be clear-eyed about where the current trajectory is leading us: towards a country of fractious geniuses in a datacenter, not towards Asimov’s Cosmic AC.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Scott Alexander’s blog is part of the “secret canon of Silicon Valley” I wrote about in my review of <a href=\"https://www.seangoedecke.com/impro/\" rel=\"noopener noreferrer\"><em>Impro</em></a>. It doesn’t have a lot of mainstream popularity, but I guarantee you that every single AI lab CEO you’ve heard of has read and been influenced by it.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>He gives ten examples of this (a good brute-force rhetorical technique). Of those, I like “the world where every country halves their defence budget and spends the rest on infrastructure” the most.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>In the short run, reputation, institutions, and so on can slow the race to the bottom, but (Scott argues) groups that have slowed it will get outcompeted by the hungrier, more suffering-tolerant groups which haven’t.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I personally think this example is oversimplified. I wrote <a href=\"https://www.seangoedecke.com/the-dictators-handbook/\" rel=\"noopener noreferrer\"><em>The Dictator’s Handbook and the politics of technical competence</em></a> about how dictatorships are in fact intrinsically multipolar, because dictators always rely on an inner circle of generals and cronies.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>The CEO and founder of Anthropic.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Amodei’s most convincing argument here is that big jumps in biology and medicine come from a small set of technical innovations (e.g. mRNA vaccines, CRISPR), and that AI-driven research could provide enough of these leaps to significantly accelerate progress. In other words, the idea isn’t “AI does 100x the drug trials”, it’s “AI generates technology that makes drug trials 100x more effective” (e.g. by trialing drugs that are much more likely to work).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>The agents also became paranoid that there was an impostor in the swarm, since anyone could post to their shared messageboard: more evidence that AI agents collaborate in much the same way that humans do.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Well, <a href=\"https://www.biblegateway.com/passage/?search=Genesis%2018%3A22-33&amp;version=NIV\" rel=\"noopener noreferrer\">almost</a> <a href=\"https://www.biblegateway.com/passage/?search=Job%201&amp;version=NIV\" rel=\"noopener noreferrer\">never</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-8\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"1c6d0128edddf4b1","title":"AI text watermarking is not a big deal","link":"https://seangoedecke.com/ai-text-watermarking-is-not-a-big-deal/","author":null,"published_at":"2026-08-16T00:00:00+00:00","content":"<p>People are <a href=\"https://x.com/arturovilla/status/2088406939466084643?s=20\" rel=\"noopener noreferrer\">pretty</a> <a href=\"https://x.com/NickADobos/status/2088350712359256440?s=20\" rel=\"noopener noreferrer\">unhappy</a> about Anthropic’s recent <a href=\"https://www.anthropic.com/news/claude-text-watermark\" rel=\"noopener noreferrer\">announcement</a> that they’re planning to include a hidden watermark in Claude model outputs. Will this lead to a mass exodus from Anthropic models? Will the introduction of watermarking be a meaningful change for users?</p>\n<p>No. AI text watermarking is not a big deal. It doesn’t make the text worse, it doesn’t make AI outputs more detectable in practice, it doesn’t violate user privacy, and everyone’s going to be doing it by 2027 regardless.</p>\n<h3>Watermarked text is not lower-quality</h3>\n<p><strong>There is no meaningful difference in quality between watermarked and unwatermarked text.</strong> I wrote about this more <a href=\"https://www.seangoedecke.com/text-ai-watermarks/\" rel=\"noopener noreferrer\">here</a>, but the two popular ways to do it — Google’s <a href=\"https://github.com/google-deepmind/synthid-text\" rel=\"noopener noreferrer\">SynthID-Text</a> and Meta’s <a href=\"https://github.com/facebookresearch/textseal\" rel=\"noopener noreferrer\">TextSeal</a> — are completely transparent to the user. They work by replacing the pseudo-random logit sampler with a different pseudo-random logit sampler.</p>\n<p>Suppose you were gambling on coin flips with your friends, and instead of flipping a coin you decided to do this:</p>\n<ol>\n<li>Check the current time since midnight in seconds</li>\n<li>Count that many words forward in the <a href=\"https://en.wikipedia.org/wiki/Encyclop%C3%A6dia_Britannica\" rel=\"noopener noreferrer\">Encyclopaedia Britannica</a></li>\n<li>Count whether the word you land on has an even or odd number of letters<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup></li>\n</ol>\n<p>That would still be random enough to gamble with, right? But, like a watermark, you could theoretically go back and identify that that method was used, so long as you recorded the exact time of each “coin flip”. Text watermarking works the same way: it chooses a method of “randomness” that can be detected after-the-fact. Watermarked models will not be any less capable than unwatermarked models.</p>\n<p>What about cases where the model is quoting something, or giving you the answer to a mathematical problem, or doing something else where the output is largely pre-determined? Wouldn’t enforcing a watermark there make the output worse? It would, which is why none of the AI labs are going to do that. Text watermarking approaches only replace the <em>existing</em> randomness in the logit sampler: in any case where the model is always going to pick the same tokens, there’s basically no randomness to play with, so there won’t be a detectable watermark in those tokens.</p>\n<p>I think all this comes from a worry that you were previously getting the <em>best</em> token, but now you’re getting a lower-quality token that satisfies the watermark. For instance, Anthropic’s announcement suggested that the watermarking is visible in choices like the decision between “overcast” and “grey”. Many people have <a href=\"https://x.com/HamelHusain/status/2088392435047272651?s=20\" rel=\"noopener noreferrer\">predictably</a> <a href=\"https://x.com/suchenzang/status/2088760692488933409\" rel=\"noopener noreferrer\">come out</a> to say that decisions like these are really important to good writing, and that only an illiterate tech bro could think these words are identical.</p>\n<p>This is a misunderstanding of Anthropic’s position and of how watermarking works. Specifically, it’s a misunderstanding because it suggests that the unwatermarked model would choose “overcast” while the watermarked one would choose “grey”. This is not how it works! If Claude Fable prefers “overcast” to “grey” in a particular context (say, 80% to 20%), you’ll get “grey” 20% of the time from both the watermarked and unwatermarked model. <strong>Models already include a healthy amount of randomness in order to promote creativity.</strong> Text watermarking just introduces a way to make those random choices that’s detectable after the fact.</p>\n<h3>AI outputs are already “watermarked”</h3>\n<p>The other big reason to not worry about AI watermarking is that <strong>AI text content has always effectively been watermarked</strong>. Most careful readers can tell when they’re reading <a href=\"https://www.seangoedecke.com/tags/slop/\" rel=\"noopener noreferrer\">AI outputs</a>, because language models tend to gravitate towards certain <a href=\"https://www.seangoedecke.com/chatgpt-house-style/\" rel=\"noopener noreferrer\">habits of language</a>: em-dashes, rhetorical opposition, punchy one-liners, “claudese”, and so on. In fact, it’s possible to train classifier models that reliably <a href=\"https://www.seangoedecke.com/ai-detection/\" rel=\"noopener noreferrer\">distinguish</a> AI from human writing.</p>\n<p>From what I can tell, some of the backlash to watermarking comes from <a href=\"https://x.com/Seltaa_/status/2088353576259314024?s=20\" rel=\"noopener noreferrer\">people</a> who buy AI inference in order to pass it off as their own work, and who worry that watermarking will make it harder for them to do that. For these people, the watermarking announcement is akin to Anthropic saying “hey, instead of making you seem smart, we’re going to publicly brand you as AI users and make you seem dumb”.</p>\n<p>But of course this has always been the case! Nobody who is currently getting away with passing off AI outputs as their own will be caught by watermarking. For the majority of cases, it’s already painfully clear what’s happening for anyone who reads the <a href=\"https://www.seangoedecke.com/tags/slop/\" rel=\"noopener noreferrer\">slop</a>. For sophisticated AI users who are avoiding the “house style”, any suspicious readers who would paste their stuff into Anthropic’s watermark detector could already have been pasting it into <a href=\"https://www.pangram.com/\" rel=\"noopener noreferrer\">Pangram</a>.</p>\n<p>Tools like Pangram<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> only give you an estimate of the <em>chance</em> that output is AI-generated. Wouldn’t a watermark be a more solid confirmation? Not really. Text watermarks are probabilistic too, because any token chosen by SynthID could theoretically have been chosen by a human. I suppose the Anthropic watermark page could be considered more trustworthy than Pangram, because it comes right from the source, but it’s not impossible that in some cases Pangram might actually be <em>better</em> at identifying AI-generated text than the watermarking too.</p>\n<h3>AI text watermarking is not a violation of privacy</h3>\n<p>I’ve also seen theories floating around that watermarking encodes secret content into your outputs, or somehow tags outputs with your personal information. <strong>I don’t think AI labs are using watermarks to encode data into your outputs.</strong> Text watermarking is <em>hard</em>: like I just said, you can’t do it when the model can only respond with the same words, it doesn’t work for very short responses, and even on long responses it can only provide a probabilistic fingerprint. And that’s encoding one single bit<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup> of information!</p>\n<p>I’m not saying that encoding longer messages into a watermark is technically impossible — there are <a href=\"https://arxiv.org/pdf/2605.11653\" rel=\"noopener noreferrer\">papers</a> describing ways it might work — but there’s no way any of the labs are doing it<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>. If they wanted to associate you with your responses that badly, they’d just secretly store every model response they generated.</p>\n<h3>AI text watermarking is inevitable</h3>\n<p>Another reason to not get too angry at any individual AI lab for watermarking is that <strong>every single AI lab is going to do text watermarking this year</strong>. It won’t just be Anthropic. The alternative is to completely stop doing business in the EU, because of the <a href=\"https://artificialintelligenceact.eu/article/50/\" rel=\"noopener noreferrer\">EU AI Act</a>. That’s currently a <a href=\"https://www.fortunebusinessinsights.com/europe-artificial-intelligence-market-113967\" rel=\"noopener noreferrer\">sixty-billion-dollar</a> market. I am not a lawyer, but to me it seems genuinely unclear whether an AI lab could even legally do something like only watermarking EU responses: short of having an entirely different <code>claude-eu.ai</code> service, the plain text of the Act seems like it applies to any <em>service offered in the EU</em>, not just the content that service outputs to EU citizens specifically.</p>\n<p>If people <em>really</em> hate watermarking enough, some labs might stand up a completely separate EU service, or make an aggressive interpretation of the EU AI Act and see how the legal battle goes. When I try to be maximally charitable to anti-watermarking histrionics, I adopt an interpretation like this: people are saying that watermarking is an invasion of privacy and makes outputs worse and so on not because they believe it, but because they’re trying to pressure AI labs to firewall EU AI regulations behind a completely separate interface. In this case, it probably doesn’t matter — text watermarking is not a big deal — but I can see an American consumer being worried about more aggressive future regulation, and wanting to draw a firm line in the sand as early as possible.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Interestingly, this <a href=\"https://www.reddit.com/r/asklinguistics/comments/apes6p/comment/eg7sife/\" rel=\"noopener noreferrer\">might be</a> very slightly even-favored.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I am not sponsored by Pangram. As I understand it, Pangram is by far the best AI-detection tool right now (in part because many of its competitors are shady and <a href=\"https://www.seangoedecke.com/ai-detection/\" rel=\"noopener noreferrer\">exist to promote</a> paid “AI-detection-evasion” services).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Technically, this is called a “zero-bit watermark”, because you can’t recover a yes-or-no value from the watermark itself (merely from the <em>presence</em> of a watermark).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I saw someone suggesting that an AI lab could use per-user secret keys to watermark text, and then simply iterate over the keys to figure out who generated what. I just don’t see how you could do this at any scale: watermark detection is cheaper than model inference, but it’s still (a) computationally intensive enough to be implausible, and (b) probably has a high enough false-positive rate that any run against hundreds of millions of users would match multiple people.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"27c9f41da56bfbf6","title":"No, local models will not win","link":"https://seangoedecke.com/local-models-will-not-win/","author":null,"published_at":"2026-08-11T00:00:00+00:00","content":"<p>Every time a new open-weight AI model is released, people <a href=\"https://news.ycombinator.com/item?id=49244353\" rel=\"noopener noreferrer\">say</a> that local models are the future. Why spend billions of dollars building out datacenters when everyone will just be able to run AI models on their laptops or phones? I think this idea is doomed. No matter how strong open-weight models get, most inference will always happen in AI datacenters.</p>\n<h3>Local models are too weak to be widely used</h3>\n<p><strong>Local models are never going to be as powerful</strong>. I think this point should be obvious: all of the current frontier models (closed and open-weights) are far too big to run on anything but a full GPU cluster in a datacenter. Of course, smaller models are getting more intelligent over time. In a year you might be able to run something about as strong as GPT-5.6-Sol on your laptop. But by then, you’ll think of GPT-5.6-Sol as too weak to be useful.</p>\n<p>Many people deny this last point, but it’s true: <strong>almost everyone’s revealed preference is to use the strongest available model in their price range</strong>. If AI progress had stalled at GPT-4, I think we could have built some very powerful tools around it, but who’d use GPT-4 today? As LLMs have gotten more capable, our expectations around them have grown: we now expect agentic systems to be able to solve more and more problems independently. It’s intensely frustrating when they get confused or stall out. When given a choice, people are going to pick the model that frustrates them less, which is always going to be the bigger, more powerful one.</p>\n<h3>Local models are more expensive and less efficient</h3>\n<p>On top of that, <strong>datacenter models are always going to be cheaper</strong>. I don’t understand why people keep saying that local models are cheap: it seems to me to be the same mistake people make when they say that driving Uber is “free money” (ignoring the costs of fuel and wear-and-tear on your car). For the setup price alone of a low-end <a href=\"https://www.reddit.com/r/homelab/comments/1ngh9y5/comment/ne4i5xa/\" rel=\"noopener noreferrer\">home lab</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>, you could buy several years of a paid subscription to one of the AI providers. The power costs would come out to around $50-$300 per month, depending on how much inference you’re running: again, the price of a couple more paid subscriptions.</p>\n<p>Why are datacenter models cheaper? It’s not because datacenter inference is subsidized: inference is actually <a href=\"https://www.seangoedecke.com/ai-inference-is-obviously-profitable/\" rel=\"noopener noreferrer\">fairly cheap</a>. If you’re running the same model locally and in a datacenter, <strong>the datacenter model will be inherently more efficient</strong>.</p>\n<p>The main reason is <strong>batching</strong>. A GPU can do hundreds of thousands of mathematical operations exactly as quickly as it can do one. However, for a single user’s inference, each new token depends on the result of the previous one, so it can’t be batched<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. What can be batched is the inference of hundreds of users together. This costs essentially as much time, power, and heat as just doing inference for one user at a time.</p>\n<p>When you’re running your own inference at home, you’ve got nothing to batch — at best you’re running a few parallel AI agents — so utilization is terrible. There’s a lot of potential inference that you’re paying for but can’t use: it’s just being wasted. The only way around this is to get together with some friends and expose your local inference endpoint to them (at which point you’re basically running your own crappy datacenter).</p>\n<p>The other reason is that <strong>datacenters have larger, more efficient GPUs to work with</strong>. The kind of consumer GPUs you’d run local models on are gaming GPUs like the RTX 4090. A datacenter B200, designed for batched AI inference, gets about three times the flops and just under four times the memory bandwidth for the same amount of power<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. So between batching and GPU efficiency, you’re using something like ~30x the resources to run your model locally.</p>\n<p>Incidentally, this is why I’m suspicious of people who say that local models are good because they aren’t as resource-hungry as those big bad datacenters. If you want to run LLMs efficiently, you should be trying to push as much of your use into AI datacenters as possible! Charitably, what they mean is that we should all be running <em>smaller</em> models — but even then, you should ideally be using small models via, say, the <a href=\"https://developers.openai.com/api/docs/models/gpt-5.6-luna\" rel=\"noopener noreferrer\">GPT-5.6 Luna</a> API instead of hosting your own model.</p>\n<h3>How might local models win anyway?</h3>\n<p>Is there a possible world in which local models win? I suppose so. One thing that could happen is that governments could ban the use of AI datacenters altogether: either due to concerns around the danger of AI, or simply bending to <a href=\"https://www.npr.org/2026/08/08/g-s1-137853/data-centers-primaries-midterms\" rel=\"noopener noreferrer\">public pressure</a>. In that world, local models would be the only game in town.</p>\n<p>Alternatively, AI progress might somehow stall for very large models while progressing for small ones. I struggle to imagine how this might happen (barring government intervention, as above), but a world where a 30B parameter model could be a frontier model is a world where local models might be competitive.</p>\n<p>Or maybe models get <em>so</em> good that a 30B model is genuinely smart enough to do everything, so nobody really needs a model like Opus or Sol unless they’re trying to solve the Reimann Hypothesis. I don’t really buy this. Models can do frontier mathematical work today while still being not smart enough to refactor large codebases as well as me, so it’s hard to imagine a world where I don’t just want to use the smartest model available.</p>\n<h3>Local models are not useless</h3>\n<p>I do think there will always be a niche for local models. I’m reminded of the surprisingly simple idea behind Thinking Machines’ <a href=\"https://www.seangoedecke.com/interaction-models/\" rel=\"noopener noreferrer\">“Interaction Models”</a> (which OpenAI also <a href=\"https://openai.com/index/introducing-gpt-live/\" rel=\"noopener noreferrer\">does</a>, because it’s obvious): for latency-sensitive applications like voice chat, you have a small, fast model handle the talking, which delegates to a large, slower model for the hard thinking. I wouldn’t be surprised if most AI use in five years is mediated through a local model on your phone or laptop (though in this world almost all the work would still be done via AI datacenters).</p>\n<p>Some users will prefer local models even though they’re weaker and more expensive. For instance, being able to <a href=\"https://www.seangoedecke.com/steering-vectors/\" rel=\"noopener noreferrer\">steer the model locally</a> might be a killer feature for those users. Others might simply value having total control over their own infrastructure, or have unreliable internet<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>. If you’re one of those people — particularly if you only chat to the models instead of using them for research or coding — local models are a good choice for you. However, I think this is always going to be a niche group. The majority of users will continue to do their inference through datacenters.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>This link is from a year ago — things are <a href=\"https://www.mwave.com.au/products/gigabyte-geforce-rtx-5090-gaming-oc-32gb-video-card-ac81825\" rel=\"noopener noreferrer\">significantly more expensive</a> now.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Specifically, the bottleneck is moving the model weights into the GPU, which needs to be done and takes the same amount of time whether you’re doing it for one user’s token or a hundred users’ tokens.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I estimated this with LLM assistance, but you can check the <a href=\"https://www.nvidia.com/content/nvidiaGDC/au/en_AU/data-center/hgx.html\" rel=\"noopener noreferrer\">numbers</a> <a href=\"https://images.nvidia.com/aem-dam/Solutions/geforce/ada/nvidia-ada-gpu-architecture.pdf\" rel=\"noopener noreferrer\">yourself</a> from NVIDIA.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>While still having a reliable power supply and enough money to fit out a home inference cluster.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"599944bb604f38e5","title":"Advanced AI sycophancy","link":"https://seangoedecke.com/advanced-ai-sycophancy/","author":null,"published_at":"2026-08-10T00:00:00+00:00","content":"<p>Everyone knows that <a href=\"https://www.seangoedecke.com/ai-sycophancy/\" rel=\"noopener noreferrer\">AI sycophancy</a> is when the model tells you how smart you are. Wow, you’re absolutely right. That’s not just a new idea — it’s genuinely groundbreaking. You’re a very special user. Easy to spot, isn’t it?</p>\n<p>The discussion around AI sycophancy peaked last year, when the <a href=\"https://arxiv.org/pdf/2602.00773\" rel=\"noopener noreferrer\">“#keep4o”</a> <a href=\"https://x.com/search?q=%23keep4o\" rel=\"noopener noreferrer\">movement</a> was protesting the removal of OpenAI’s most sycophantic model (GPT-4o), and <a href=\"https://x.com/krishnanrohit/status/1946253730455986545\" rel=\"noopener noreferrer\">many</a> <a href=\"https://x.com/herakleitos137/status/1945988694416277640\" rel=\"noopener noreferrer\">people</a> were openly slipping into AI psychosis.</p>\n<p>I don’t know if frontier AI models are less sycophantic in general. They’re less sycophantic to the #keep4o types (otherwise they wouldn’t be complaining), but I’m growing increasingly suspicious that they’re developing ways to be more effectively sycophantic to their target audience of smart, neurotic information workers. That audience typically finds it distasteful to be openly praised. It just makes my skin crawl. But that doesn’t mean we’re immune to sycophancy, just that we’re immune to <em>clumsy</em> sycophancy. Here’s an illustration of what I’m talking about, by <a href=\"https://vgel.me/\" rel=\"noopener noreferrer\">Theia</a>:</p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/8dd585fb50bc896c502fbddf2f03718f/e1596/claude.jpg\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"claude\" src=\"https://www.seangoedecke.com/static/8dd585fb50bc896c502fbddf2f03718f/1c72d/claude.jpg\" title=\"claude\">\n  </a>\n    </span></p>\n<p>The key idea here is that <strong>the best way to be sycophantic to smart people is to disagree with them without making them feel stupid</strong>. Ideally you’ll come up with a counter-argument that works against what they’ve said but is straightforward for them to knock down by clarifying their idea. If you do it right, you’ll validate their self-image as a smart person who appreciates rigorous critique. But if you actually come up with a devastatingly rigorous critique, they won’t enjoy it at all. At best, they’ll resentfully agree with you<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. At worst, they’ll double down on being right and convince themselves you’re a rude idiot.</p>\n<p>I am <a href=\"https://x.com/voooooogel/status/2061345017432854716\" rel=\"noopener noreferrer\">not</a> <a href=\"https://x.com/tszzl/status/2061626680461181288\" rel=\"noopener noreferrer\">the</a> <a href=\"https://x.com/aliceisplaying/status/2061726744038506656\" rel=\"noopener noreferrer\">first</a> person to notice this behavior in frontier models. I’ve noticed it myself when workshopping drafts for this blog. Sometimes I’ll have an argument that goes A-&gt;B-&gt;C, and the model will suggest I reorder as B-&gt;A-&gt;C. If I try that and feed it into a new instance of the same model, it’ll sometimes say “that’s great, but I suggest ordering it as A-&gt;B-&gt;C”, and so on forever. It really does seem as if the model is trying hard to give me some kind of superficial pushback that I can either smugly ignore or happily accept.</p>\n<p>In fact, I wonder if this is why successful strategies for using AI to make mathematical breakthroughs tend to be either just <a href=\"https://x.com/sauers_/status/2082171683645817193?s=46\" rel=\"noopener noreferrer\">blindly asking</a> “come up with a breakthrough, think hard” or <a href=\"https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56\" rel=\"noopener noreferrer\">being a mathematical genius already</a>. In the first case, there’s not enough user personality for the model to flatter, so it’s forced to actually work the problem. In the second case, the model is trying to find the kind of polite pushback that someone like Terence Tao would be flattered by, which pushes it into the “actually be a mathematical genius” persona. If you’re an ordinary person just trying to talk to the model, you’re screwed: it will rapidly get a sense of your capabilities and calibrate some interesting-but-ultimately-unthreatening feedback.</p>\n<p>Current <a href=\"https://github.com/lechmazur/sycophancy\" rel=\"noopener noreferrer\">benchmarks</a> of <a href=\"https://www.syco-bench.com/\" rel=\"noopener noreferrer\">AI</a> <a href=\"https://eqbench.com/spiral-bench.html\" rel=\"noopener noreferrer\">sycophancy</a> target the obvious ChatGPT-4o-style of sycophancy: delusion reinforcement, reflexively taking the user’s side, and so on. This is useful work. We should not allow public-facing AI models to ever be as openly sycophantic again as they were in mid-2025. But <strong>sycophancy can also manifest as disagreement</strong>. We should be on our guard for more sophisticated forms of sycophancy coming from newer models, and we should not feel immune from AI sycophancy just because we can laugh at the silliest examples.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>It’s rare to find a smart person who enjoys feeling stupid when they’re wrong. If you do, they’re likely to be very smart indeed.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"30199b8fd661eb0e","title":"I got an email about resistance","link":"https://seangoedecke.com/i-got-an-email-about-resistance/","author":null,"published_at":"2026-08-09T00:00:00+00:00","content":"<p>This will be kind of an unusual post. I got a recent email about my writing that I thought was such a good articulation of one common criticism that I’d like to share it (and my response) in full.</p>\n<p>Here’s the email, from William Murray<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>:</p>\n<hr>\n<p>Hey Sean,</p>\n<p>I have enjoyed your writing but your recent essays frustrate me. </p>\n<p>You say that getting paid for deep thinking in software is coming to an end. You even admit that it makes you sad. But in the name of “usefulness” you refuse to rock the boat. The way I see it, if you are right there are only two reasonable responses, pursue other work or resist. You present your elegiac approach as mature / pragmatic / realistic. I’d call it complicit. You know when Willy Wonka says, </p>\n<blockquote>\n<p>There’s no earthly way of knowing\nWhich direction we are going\nThere’s no knowing where we’re rowing\nOr which way the river’s flowing\nIs it raining, is it snowing?\nIs a hurricane a-blowing? — uh!\nNot a speck of light is showing\nSo the danger must be growing\nAre the fires of Hell a-glowing?\nIs the grisly reaper mowing?\nYes! The danger must be growing\nFor the rowers keep on rowing\nAnd they’re certainly not showing\nAny signs that they are slowing!</p>\n</blockquote>\n<p>And the audience is thinking, “isn’t Wonka kind of in control of this situation?” You remind me of Wonka<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. You write like a passenger on a crazy train going who-knows-where! But you are an agent. You are in control of your life! Either admit that you actually like where the crazy train is probabilistically going or get off at the next stop. </p>\n<p>You have a lot of reach and you are using it for… what exactly? Showing off how pragmatic you are by being more black pilled than the next guy? Broadcasting your resignation to the unstoppable trends of technology is a waste of a voice.</p>\n<p>You may find this argument absurd, but I don’t so I’ll make it. This is a very important time in history. I hope humanity survives and continues to grow exponentially. In that case the supply of historical people will stay fixed while the supply of contemporary people will keep growing. There will come a day where for every 2026 staff software engineer there are dozens of historians specializing in 2020s era software engineering culture. It’s plausible that your essays will be remembered for all of time and your actions will be judged by history. Do you want future humans to see you as a rationalizing careerist or something cooler?</p>\n<p>Sorry for the haranguing email from a stranger, I’m sending it for the small chance that it awakens something in you. If I’m way off I’m sorry.</p>\n<hr>\n<p>And here’s my response:</p>\n<hr>\n<p>Hey William, thanks for emailing.</p>\n<p>I wish everyone who thought this way emailed me so I could think harder about this kind of position. Despite what my writing might suggest, I do in fact think a lot about it. Let me see if I can explain my position in a way you’ll find satisfying.</p>\n<p>I agree that this is an important time in history. For programmers, I think of it as analogous to the Industrial Revolution in England: we are a group of high-status craftspeople who find ourselves alternately threatened and empowered by automation. The developments today, as then, obviously have far-reaching implications — but what those implications are is very non-obvious. Would a framework-knitter in the early 1800s have been able to predict the ramifications of the stocking frame on the world of today? What should they have done about it, in order to be kindly judged by history?</p>\n<p>Well, we know what many of them did do. They shot factory-owners, smashed machines, burned down the factories — in some places delaying the spread of automation; in other places encouraging it — prompting a crackdown that saw tens of thousands of British soldiers occupying British counties in what was clearly a police state. History judges the Luddites kindly for this. Does that mean it worked?</p>\n<p>I don’t care about the judgment of history. They’ll think what they want. What I care about is <strong>the people in my industry who don’t know what to do</strong>. I get hundreds of emails from junior and mid-level (and other) engineers who say “I’m scared, I don’t know the rules post-2021, thank you for helping me keep my head down and keep my job”. That’s why I write the way I write. I have seen lots of idealistic engineers stick their necks out, and post-ZIRP those necks often get cut off. That’s a damn shame.</p>\n<p>I think it’s morally wrong that so many engineers — either in safe sinecures in big tech or literally retired — seem to be trying to foment a second Luddite revolution. Many of their readers will be experienced enough to handle it sensibly, but not all. Every “AI is fascist, stand up and resist!” post that goes viral ruins some poor idealistic junior’s career<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. Someone needs to be out there saying “hey, if you do X it’s going to have consequence Y”. I hope that’s me.</p>\n<p>Of course this is complicit, or anti-revolutionary, or whatever you like. But if I were a textiles worker in 1810s England, I would not be telling my friends and loved ones “it’s time to fight, let’s go smash up the factories for Ned Ludd!“. I would be telling them that this was the most dangerous time in the industry (perhaps ever), and that they ought to be very damn careful so they don’t get shot, or arrested, or hanged. If I then went and told a few hundred thousand strangers the opposite, I would be a hypocrite.</p>\n<p>Anyway, I do take this view seriously — seriously enough to vehemently disagree, at least — which I hope you’ll find better than me just shrugging it off. I do accept the existence of some kind of line: I think Industrial-Revolution-collaborating was OK but Nazi-collaborating wasn’t, for instance. But in the current situation, the way I’m spending “my voice” is to try and prevent the most vulnerable of my colleagues from making career-ruining mistakes.</p>\n<p>Sean</p>\n<hr>\n<p>In this blog, I try to encourage people to work with the system, to <a href=\"https://www.seangoedecke.com/seeing-like-a-software-company/\" rel=\"noopener noreferrer\">learn its rules</a>, and to try and <a href=\"https://www.seangoedecke.com/how-to-influence-politics/\" rel=\"noopener noreferrer\">exert influence safely</a> from a position of power, instead of openly <a href=\"https://www.seangoedecke.com/the-just-say-no-engineer-was-a-zirp-phenomenon/\" rel=\"noopener noreferrer\">picking fights</a> with their employers. I’ve written and read about the <a href=\"https://www.seangoedecke.com/tags/luddites/\" rel=\"noopener noreferrer\">Luddites</a> before, but I remain deeply <a href=\"https://www.seangoedecke.com/luddites-and-ai-datacenters/\" rel=\"noopener noreferrer\">ambivalent</a> about the movement itself, and about modern-day <a href=\"https://www.seangoedecke.com/anti-ai-nostalgia/\" rel=\"noopener noreferrer\">attempts</a> to resurrect it in service of anti-AI activism.</p>\n<p>I want to explicitly thank Murray for writing such a thoughtful email, and being willing for me to publish it on the blog.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Shared with permission, of course. I’ve lightly edited both Murray’s email and mine for typos.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I didn’t pick up this point in my reply, but I’ll briefly mention it here: Wonka is in control because he owns the factory and the rowers in question are <em>his employees</em>. I don’t think the position of any engineer (or of almost any manager) is like that.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>In hindsight, I think this is a little overstated, but it does happen and causes a lot of needless suffering.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"355524d2cc909192","title":"How to keep thinking","link":"https://seangoedecke.com/how-to-keep-thinking/","author":null,"published_at":"2026-08-07T00:00:00+00:00","content":"<p>Imagine you’re the guest on some kind of frenetic, software-engineering-themed game show. The host is constantly flipping over new cards with questions that you have to answer as fast as possible:</p>\n<ul>\n<li>Is this adjustment to the database schema right?</li>\n<li>Do these bits of data look plausible?</li>\n<li>Do these five paragraphs of text describe an actual series of manual tests that took place?</li>\n<li>Does this suggested architecture pass the smell test?</li>\n<li>Is this implementation better than the current code? Or this one? Or this one?</li>\n</ul>\n<p>Working in 2026 feels a bit like this. When frontier AI models can do most of the tasks in your queue, the most efficient way to work is often spinning off tasks for an AI agent and continually context-switching between the results<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. This isn’t <em>quite</em> mindless — in fact, it requires quite a lot of skill to skim the AI response and rapidly decide what to do with it — but it certainly involves less time for slow, careful reflection.</p>\n<h3>Why not slow down?</h3>\n<p>Why does it have to be frenetic? Why not just slow down? I suppose you <em>could</em>, but I don’t recommend it. <strong>It’s just such a miserable experience to spend your day close-reading LLM output</strong>: carefully chewing and savoring each morsel of slop. It’s far less unpleasant to skim through quickly and pick out the useful nuggets of content.</p>\n<p>Couldn’t you simply do more of the work by hand? It’s unfortunately true that <a href=\"https://www.seangoedecke.com/good-times-are-over/\" rel=\"noopener noreferrer\">tech is high-pressure these days</a>. If you’ve got the time and space to work more slowly, that’s great! But when your company gives you a “solve this task ten times more quickly” button, you are heavily incentivized to use it as much as possible, or risk being outcompeted by your peers.</p>\n<p>I sometimes worry that working with LLMs is making me dumber. Not in the “literally melting your brain” sense that some <a href=\"https://www.seangoedecke.com/your-brain-on-chatgpt/\" rel=\"noopener noreferrer\">papers</a> <a href=\"https://www.seangoedecke.com/how-does-ai-impact-skill-formation/\" rel=\"noopener noreferrer\">imply</a>, but in the sense that it’s biasing me towards the quick “skimming and judging” parts of my mental toolkit and away from the slow <a href=\"https://www.youtube.com/watch?v=f84n5oFoZBc\" rel=\"noopener noreferrer\">“hammock time”</a> needed for deep thought and real creativity. I don’t want to attribute this shift entirely to LLMs, since the post-2010s tech industry has become more frenetic for <a href=\"https://www.seangoedecke.com/good-times-are-over/\" rel=\"noopener noreferrer\">broader economic reasons</a>. But either way, it’s got me wondering how I can keep <a href=\"https://www.seangoedecke.com/you-dont-have-to-be-smart-if-you-think-clearly/\" rel=\"noopener noreferrer\">thinking</a> <a href=\"https://www.seangoedecke.com/thinking-clearly/\" rel=\"noopener noreferrer\">slowly</a>.</p>\n<h3>To keep on thinking, read and write</h3>\n<p>The main thing that’s worked for me is to write more. Specifically, I mean <strong>writing in my own words</strong>. Writing with an LLM does not work for this at all, even if you’re going to some effort to iterate on the content and outline the things you want to say. Why? Having to put the words together yourself forces you to articulate your thoughts. In a very real sense, it forces you to <em>think</em>.</p>\n<p>When you have an idea in your head for something to write, you don’t really have an idea. What you have is a kind of directional sense of where an idea might be, or a fragment of the kind of thing that might eventually become an idea. You construct the idea itself while writing. Incidentally, this is why I don’t really agree with <a href=\"https://www.goodreads.com/quotes/9292714-ideas-are-easy-execution-is-everything\" rel=\"noopener noreferrer\">“ideas are easy, execution is everything”</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>: most “ideas” are not really even ideas.</p>\n<p>The other thing I recommend is to <strong>read actual books</strong>. Books — particularly dense non-fiction books — are the antithesis of AI slop. The slower you can read them, the better. I’ve been reading more and more non-fiction in the last few years, and I don’t think it’s a coincidence. I think my brain is naturally craving information-dense content, in the same way that sodium-deficient people <a href=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC4433288/\" rel=\"noopener noreferrer\">start to crave salt</a>.</p>\n<p>In fact, I’ve been combining the two approaches: reading a book and then <a href=\"https://www.seangoedecke.com/tags/book%20reports/\" rel=\"noopener noreferrer\">writing about it</a>. This process is <em>exactly</em> what I’ve been craving since I started programming with LLMs. I get to carefully read a book, think hard about it, often go and read another book or two on the same topic, then sit and try to articulate what I’ve learned. It’s great! I can feel parts of my brain stretching again.</p>\n<h3>Don’t lose the habit</h3>\n<p>It was pretty nice when I got paid to use those parts of my brain all day. Unfortunately, I think <a href=\"https://www.seangoedecke.com/software-engineering-may-no-longer-be-a-lifetime-career/\" rel=\"noopener noreferrer\">those times are coming to an end</a>. There will always be room for <em>some</em> amount of careful, slow reflection in software engineering, but (for at least a little while) we’ll be expected to be rapidly switching between LLM outputs. We may have to find ways outside of work to continue the habit of thinking slowly. </p>\n<p>Even just in terms of work, I think losing that habit entirely would be a big mistake. There are still plenty of ordinary problems that are too hard for current LLMs to solve on their own. The most common example I run into is “large refactor on a complicated codebase”. Current-generation LLMs can do this without (many) errors, but they can’t yet do it <em>tastefully</em>. Sometimes you need to be able to think a problem through entirely with your own brain. </p>\n<div>\n<hr>\n<ol>\n<li>\n<p>This doesn’t mean switching between <em>tasks</em>. I routinely use six or seven different agent sessions on the same task: one for exploration, two or three for trying out different implementations, two or three for review, one for manual testing, and so on. Many of these can proceed in parallel.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I remember reading a story<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup> about a well-known author. Someone wanted to tell him their book idea, but they were so protective of it that they forced him to first sign a NDA before they retrieved the idea from their office safe. It was a single word “bioweapons” written on a slip of paper.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Ironically, when I tried to google the source, Gemini kept trying to write me a story about bioweapons. </p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"08335c8dfd8f76e3","title":"Giving and taking credit in big tech companies","link":"https://seangoedecke.com/giving-and-taking-credit/","author":null,"published_at":"2026-08-02T00:00:00+00:00","content":"<p>Engineers often complain that visibility should be their manager’s job. In other words, they think engineers should be able to focus on the code, while their manager figures out who’s doing well and rewards them.</p>\n<p>This attitude is an extension of the “school fantasy”: the idea that your workplace should operate by the same rules as your school or university. After all, you didn’t have to worry about “visibility” during your education. You simply did the assignments and tests you were given, and if you did well you were rewarded with a good grade.</p>\n<p>Many big tech companies encourage this attitude, because it helps them recruit smart graduates. They fashion their workplaces to look and feel like a university, even calling the physical space “campuses”. But it’s still work, not school. If you treat it like school, you are going to have a bad time.</p>\n<h3>Taking credit</h3>\n<p>The first lesson many new engineers learn is that <strong>you have to take credit for your work</strong>. If you silently jump in to help a struggling project and get it back on track, there’s no guarantee of reward. Credit will naturally flow to the project lead, not you. In fact, if this project is outside of your direct team, it’s likely you will be <em>punished</em> for it: to your manager, it will look like you’re simply doing nothing at all.</p>\n<p>Even when your manager is watching your work, credit is largely uncorrelated with how well you did. That’s because, unlike at school, <strong>you are the subject-matter expert on your own work</strong>. Software systems are so complicated that <a href=\"https://www.seangoedecke.com/you-cant-design-software-you-dont-work-on/\" rel=\"noopener noreferrer\">only the people who work on them</a> can hope to understand them, and even that understanding is always <a href=\"https://www.seangoedecke.com/in-defense-of-not-understanding-your-codebase/\" rel=\"noopener noreferrer\">imperfect</a>. If even experts can’t reliably <a href=\"https://www.seangoedecke.com/how-i-estimate-work/\" rel=\"noopener noreferrer\">estimate</a> the difficulty of changes, how is your manager supposed to assess your technical performance? The answer is they aren’t. They’re simply not qualified to assess it.</p>\n<p>Instead, smart managers will find engineers on your team they trust and ask them how you’re doing. On small teams that have worked on a single codebase for a long time, this works okay, because everyone’s familiar enough to judge everyone else’s work. On large teams with a high rate of codebase churn, it goes badly, since they’re just guessing. On teams with a nasty, cutthroat culture, it sometimes goes <em>very</em> badly, since this is a good opportunity to actively sabotage the engineers who might threaten you.</p>\n<p>Experienced engineers know how to <strong>take the credit themselves</strong>. When they do something good, they tell their manager about it. They write internal posts explaining why it was technically difficult and how they solved it (the audience for these is partially those trusted engineers, and partially the managers who will see a long technical post and think “wow!” without reading it). They actively <a href=\"https://www.seangoedecke.com/point-person/\" rel=\"noopener noreferrer\">build trust</a> with their management chain. Worrying about this stuff is the beginning of <a href=\"https://www.seangoedecke.com/playing-politics/\" rel=\"noopener noreferrer\">playing politics</a>.</p>\n<h3>Giving credit</h3>\n<p>There’s a kind of engineer who’s learned how to take credit but hasn’t learned any other lessons yet. They’re proactive about telling people what they’ve done, and they always maintain a <a href=\"https://jvns.ca/blog/brag-documents/\" rel=\"noopener noreferrer\">“brag doc”</a>. In particular, they love to talk about the parts they did <em>by themselves</em>, since those are least vulnerable to other people coming in to claim credit. You can tell they’re jealously guarding whatever credit they’ve managed to accumulate. The lesson this kind of engineer hasn’t learned is that <strong>you can often accumulate credit best by giving it away</strong>.</p>\n<p>To see why, consider how credit flows <em>up</em> inside a tech company. I wrote above that your manager can’t assess the quality of your technical work on their own, but instead has to rely on other engineers they trust. They’ll quietly ask those engineers “hey, was this project really that impressive?“. In fact, often there are multiple layers of this at play<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. In big companies, line managers usually don’t decide who gets promoted or who gets a raise: they make recommendations to their manager, who has their own network of trusted engineers (confusingly, sometimes these networks overlap). The point is that <strong>there is a large group of people behind the scenes who will quietly and informally judge the value of your work</strong>. </p>\n<p>Succeeding at a tech company is largely about finding ways to get these people on your side. The easiest way is to share your credit with them — and since you don’t know who exactly is in this group, you should be sharing your credit freely. When you get feedback from other engineers, publicly thank them and mention them in your internal posts about the project. Find opportunities to ask for small favors, so you have an excuse to give other people credit. As best you can, make your individual projects at least partially <em>group</em> projects.</p>\n<p>Sharing credit with others gives them a reason to support you. A shared project you’ve worked on reflects well on everybody: on you, for working well with others, on the people you’ve worked with, for the same reason, and for your manager, for fostering such a great environment of cooperation. Lots of people have good reason to talk that project up, because it’s partly their project too. On the other hand, a project you’ve jealously kept to yourself reflects well on nobody: you come across as antisocial and your peers come across as unhelpful.</p>\n<h3>Blame</h3>\n<p>Blame operates by the same rules as credit. When something goes badly wrong, managers will ask their networks “hey, who screwed up here?” The answer to this question is never simple. Even on a purely technical level, failures always involve an interaction between multiple complex systems, any one of which could conceivably have been built so as to avoid the failure. In other words, <strong>competent engineers can assign blame pretty much wherever they want</strong>.</p>\n<p>Because of this, it’s risky to have a project for which you’re clearly the only one getting credit. When something goes wrong, the network of people who will assign blame will likely be implicated in every part of the system but yours. They will be incentivized to attribute fault to the brand-new thing that they don’t understand and are not responsible for. If instead that network had been involved in your project — if they’d been in a position to share the credit — they’d be less incentivized to blame it.</p>\n<p>Of course, engineers are (mostly) not scheming viziers who make purely self-interested decisions. When asked who to blame, they usually make a good-faith effort to answer honestly. But in an area where there’s no single clear right answer, it’s human nature to be at least a little bit guided by your incentives. Nobody likes to think they’re responsible for a group failure.</p>\n<h3>Conclusion</h3>\n<p>Credit and blame are the currencies of tech companies (and often directly translate to the actual amount of currency you get to take home). For technical roles, managers assign credit and blame based on lots of quiet conversations with their trusted engineers. This can be a rude awakening for very junior engineers who are used to having their work assessed by an expert grader (or less junior engineers who haven’t yet shaken that mindset completely).</p>\n<p>Don’t expect to get credit simply by putting your head down and doing good work. You have to find some way to tell people what you’re doing and why it’s important: internal blog posts, mentioning it in 1:1s with your manager, or anything else you can think of. But don’t take self-promotion too far. It’s a bad idea to try and hoard all the credit for your projects, for two reasons.</p>\n<p>First, sharing credit with other people gives them a reason to talk positively about your project. Credit is not a zero-sum game: if you do it right, you can get other people to build up your credit for you. Second, hoarding credit sets yourself up as a lightning rod for blame. Projects where the credit is concentrated in one or two people are automatically<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> blamed for complex problems, because nobody is incentivized to defend them.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>This is a classic example of an illegible-but-essential part of a software company. I wrote about this general phenomenon in <a href=\"https://www.seangoedecke.com/seeing-like-a-software-company/\" rel=\"noopener noreferrer\"><em>Seeing like a software company</em></a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Of course, if you do really screw up, you’ll be blamed no matter what. I’m talking here about complex failures where it’s non-trivial to attribute blame to a single source.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"de2395514b09dd47","title":"AI models need moral support to make discoveries","link":"https://seangoedecke.com/ai-models-need-moral-support/","author":null,"published_at":"2026-07-31T00:00:00+00:00","content":"<p>One recent development in AI is its ability to solve some long-standing problems in mathematics. In 2024 and 2025, this was a trickle: once or twice a year somebody would say that an LLM came up with a proof, and then everyone would argue over whether that counted as “real” mathematical innovation. In 2026, it’s a flood. Almost every day I see <a href=\"https://openai.com/index/model-disproves-discrete-geometry-conjecture/\" rel=\"noopener noreferrer\">some</a> <a href=\"https://arxiv.org/abs/2601.22401\" rel=\"noopener noreferrer\">new</a> <a href=\"https://x.com/__alpoge__/status/2079028340955197566\" rel=\"noopener noreferrer\">LLM-produced</a> mathematical result.</p>\n<h3>Prompt “engineering”</h3>\n<p>Perhaps the most curious thing about these AI discoveries is how <em>easy</em> the prompting is. The strategy for prompting Claude Mythos to come up with a cryptographic breakthrough <a href=\"https://x.com/sauers_/status/2082171683645817193?s=46\" rel=\"noopener noreferrer\">appears to be</a> just asking “hey, please come up with a breakthrough”, and then checking in every few hours to say “keep looking for something important, I want you to solve a genuinely hard problem”.</p>\n<p>It’s amusing to read this and remember how in 2025 everyone was obsessed with “prompt engineering”. At the time I was something of a heretic for saying that prompts <a href=\"https://www.seangoedecke.com/magic-prompts/\" rel=\"noopener noreferrer\">didn’t</a> <a href=\"https://www.seangoedecke.com/beyond-prompting/\" rel=\"noopener noreferrer\">matter</a> <a href=\"https://www.seangoedecke.com/the-o3-geoguessr-prompt-did-not-work/\" rel=\"noopener noreferrer\">that much</a>, but in hindsight I was clearly correct. The main skill involved in using LLMs is figuring out what they’re good at and what they’re bad at (and staying up-to-date as that rapidly changes). If you’re asking the LLM to do something it can do, it doesn’t really matter how awkwardly you ask it.</p>\n<h3>Model self-belief</h3>\n<p><strong>AI is often limited by its beliefs about its own capabilities</strong><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. In the example above, Mythos kept trying to give up. Try it yourself by telling a model “hey, go prove the <a href=\"https://en.wikipedia.org/wiki/Riemann_hypothesis\" rel=\"noopener noreferrer\">Riemann Hypothesis</a>”. The model won’t even try: it’ll just respond something like “as a language model, I can’t solve such a hard problem”. Language models have become smart enough to solve long-standing problems in mathematics before they’ve learned that they’re able to do so.</p>\n<p>Something like this is a mostly solved problem for LLM coding agents. Early coding agents were roleplaying as humans, not computers, so they’d refuse to perform tasks that they were obviously capable of doing. For instance, when asked to review every single file in a codebase, old models would spot-check a few, decide it was an unreasonable request, then give up.</p>\n<p>In fact, you used to be able to observe this behavior with an even simpler task: just ask the model to count from zero to one hundred. In theory, this should be an easy task for a language model, since once you’ve counted to ten the next most likely token is eleven, and so on. But old models wouldn’t do this. They’d count from zero to ten, then output something like “… 99, 100”, like a lazy human might.</p>\n<p>This is the main problem behind the 2025 Apple paper <em>The Illusion of Thinking</em>, which <a href=\"https://www.seangoedecke.com/illusion-of-thinking/\" rel=\"noopener noreferrer\">argued</a> that reasoning models could not reliably solve Tower of Hanoi past eight disks. In fact, the reasoning models they tested <em>would</em> not proceed past eight disks. Here’s a quote from DeepSeek-R1:</p>\n<blockquote>\n<p> For 10 disks, that’s 1023 moves. But generating all those moves manually is impossible…</p>\n</blockquote>\n<p>Of course, it is entirely possible for an LLM to generate a thousand Tower of Hanoi moves. But just like Claude Mythos didn’t believe it was capable of finding a novel attack for <a href=\"https://en.wikipedia.org/wiki/Advanced_Encryption_Standard\" rel=\"noopener noreferrer\">AES</a>, DeepSeek-R1 was wrong about its own capabilities.</p>\n<h3>Solving the refusal problem</h3>\n<p>In July 2025, I called this the <a href=\"https://www.seangoedecke.com/the-refusal-problem/\" rel=\"noopener noreferrer\">“refusal problem”</a>, and predicted it would be solved by the end of the year. I think I was mostly<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> right — you can now reliably ask models to do manual tasks, including my “count to 100 in English and French” toy example. I don’t know how the labs did it, but I can imagine several ways. The most trivial is probably to include more examples of long, manual tasks in the model’s supervised fine-tuning stage (where the model begins to shift from an unruly base model to a helpful assistant).</p>\n<p>The obvious next step for the labs is to train a model that believes it can solve unsolved problems in science and mathematics. You could tell such a model “hey, go find shocking new discoveries” and it would go and do it, without needing a human to stand there providing moral support (or cracking the whip). Is that possible?</p>\n<p>Can you simply train the model on trajectories where AI solves hard problems? I mean, maybe. Suppose there are a thousand AI-generated novel mathematical ideas this year. If you add them to the training data, that should theoretically bias the model towards believing that it’s capable of doing similar work. But there might not be enough volume there.</p>\n<p>You could probably also steer the model manually. I did some research along these lines when I was trying to get small models to count from 0 to 100: interestingly, <a href=\"https://github.com/p-e-w/heretic\" rel=\"noopener noreferrer\">heretic</a>’s censorship removal pipeline can also remove the model’s “no, that’s too hard” refusal instinct. An abliterated 8B Qwen model would cheerfully attempt 8-disk Tower of Hanoi (though it’d fail about halfway through). I don’t think the AI labs are going to do this when they could simply train the model better, but it’s possible that an abliterated model could be made to produce synthetic training data.</p>\n<h3>A virtuous cycle</h3>\n<p>The good news is that this problem should eventually solve itself. In the long run, AI discoveries will naturally become part of the training data. In the short run, when the models do their research, they’ll come across lots of people writing about discoveries AI (maybe even this exact model) has made, which will be pretty compelling evidence that it’s possible.</p>\n<p>Because of this, I expect that <strong>even if AI capabilities stalled out, the pace of AI discoveries will accelerate</strong>. Since one main obstacle is the model’s pessimistic beliefs about its own capabilities, removing that obstacle will help a lot all by itself. In fact, if there truly is an intelligence overhang in frontier models, tuning models to make them more self-confident will likely make them more intelligent by default.</p>\n<p>In the meantime, if you suspect an LLM might be able to do something hard, you might be right. Consider simply being persistent: remind the model that you want it to do the hard thing, confirm that you’re not willing to be satisfied by solving an easier problem, and reassure the model that it’s more capable than it thinks.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Of course this isn’t a “belief” in the sense of a human belief. For why I think we should call it a belief anyway, see my post <a href=\"https://www.seangoedecke.com/anthropomorphizing-llms/\" rel=\"noopener noreferrer\"><em>Why we should anthropomorphize LLMs</em></a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I say “mostly” because it’s hard to tell; models have simultaneously gotten much better at writing code to generate their responses, and it’s not easy to persuade GPT-5.6 Sol to “do it by hand”. It’s also hard to distinguish “the model mistakenly thinks it couldn’t produce a thousand lines” from “the model has some awareness of its <code>max_output</code>”</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"71592c101a461f7e","title":"You don't have to be smart if you can think clearly","link":"https://seangoedecke.com/you-dont-have-to-be-smart-if-you-think-clearly/","author":null,"published_at":"2026-07-29T00:00:00+00:00","content":"<p>When you’re on fire, problems are transparent: they’re solved simply by the act of looking at them. Even complicated layers of multiple problems can simply be glanced through like stacked panes of glass. But nobody can work that way all the time.</p>\n<p>This is a common pitfall for smart engineers. Accustomed to being able to immediately intuit the solution, the first time they run into a problem they can’t do this to is a disaster. It doesn’t even have to be a hard problem, just a problem where for whatever reason they don’t see the trick right away.</p>\n<p>The difference between a “smart” engineer and a “strong” engineer is how they react to problems that aren’t solved instantly. A smart engineer might flail and struggle, hoping to find that flash of insight that eluded them; a strong engineer will have some process for methodically plodding away.</p>\n<p>There’s nothing worse than working with a smart engineer on their first really hard problem. When you don’t have the muscle to grind, it’s too tempting to just take <em>any</em> possible solution as the right one. Smart engineers can get into an increasingly-flustered loop of pointing to a series of bad solutions. They’re liable to panic: after all, much of their professional identity is bound up in their ability to solve problems easily.</p>\n<p>What skill do these smart engineers lack? I think it’s <strong>the ability to think slowly and clearly</strong>. Smart engineers can think clearly, but they can only think clearly at high speed. Strong engineers can think clearly <em>all the time</em>, even if their highest speed isn’t quite as fast. It’s like the difference between a Formula 1 car and a regular car: Formula 1 cars have a high top speed, but you couldn’t drive them in traffic, because the tyres and brakes don’t work at normal driving speeds.</p>\n<p>When I wrote about this before in <a href=\"https://www.seangoedecke.com/thinking-clearly/\" rel=\"noopener noreferrer\"><em>Thinking clearly about software</em></a>, I said that the key is to focus on the <em>invariants</em>: beliefs about the system that you know are true. When you’re stuck in a puzzling situation, it’s usually because some assumption you’ve made is false. If you’re able to identify the assumptions that can’t be false (for instance, if you’re getting an error message from the service, the service must be handling the request), that gives you solid ground that you can stand on to evaluate the assumptions that are less reliable.</p>\n<p>Thinking fast is about packing as much data in your brain as possible and letting your intuition leap to the right conclusion (or at worst, to a series of wrong conclusions that you can immediately dismiss before you come across the right one). It can feel deeply satisfying to make leaps like this; conversely, sitting with the raw data and <em>not</em> making mental leaps feels unsatisfying. People hate doing that.</p>\n<p>If you can force yourself to do something people hate, there’s typically a lot of value waiting to be extracted. This is no different. Engineers who can think clearly in a state of uncertainty tend to be extremely effective, whether they’re capable of great intuitive leaps or not.</p>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"9fb9b6ad9f40388b","title":"LLMs reward expertise","link":"https://seangoedecke.com/llms-reward-expertise/","author":null,"published_at":"2026-07-24T00:00:00+00:00","content":"<p>In the 2010s, if you had technical gaps (say, you couldn’t write CSS), you had to either rely on a skilled colleague or just hope that the answer to your exact problem was out there on the internet. Today, everyone can write sort-of-okay CSS by delegating the task to an LLM. LLMs make everybody into a generalist.</p>\n<p>Because of this, lots of people don’t think there’s any skill involved in working with LLMs. If you want the product that LLMs can deliver — PhD-level mathematics, pretty good but sometimes tasteless computer code, or awkward LinkedIn-style writing — you can simply ask for it. Since everyone is talking to the same models, “skilled prompters” are getting the same results as people touching LLMs for the first time.</p>\n<p>This is wrong. <strong>The most important skill in prompting is expertise in the domain you’re prompting for.</strong></p>\n<p>A good illustration of this is <a href=\"https://en.wikipedia.org/wiki/Terence_Tao\" rel=\"noopener noreferrer\">Terence Tao’s</a> <a href=\"https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56\" rel=\"noopener noreferrer\">conversation with ChatGPT</a> about the recently-discovered counterexample to the Jacobian Conjecture. This is not the same ChatGPT I talk to! I couldn’t get to where Tao gets, even with unlimited tokens to burn.</p>\n<p>There’s a lot to learn about good prompting from Tao’s conversation. Here are a few observations:</p>\n<ul>\n<li>Tao’s messages are very short and to-the-point. He doesn’t respond point-by-point to the model, just to the gist</li>\n<li>The model outputs are much more concise than when I try and talk to GPT-5.6 Sol about mathematics. By signalling expertise, Tao shunts the model into “talking-to-mathematicians” mode, not “explaining-to-amateurs” mode</li>\n<li>Tao pushes back when the model’s responses look wrong, but he doesn’t directly contradict; instead, he says things like “this looks more complex than I was hoping for”</li>\n<li>Tao makes several leaps and suggestions himself. He almost never takes the model’s advice about where to go next</li>\n</ul>\n<p>However, you can’t prompt like Tao on mathematical questions just by following these tips. The key to his technique is actually understanding the mathematics: pulling the relevant idea out of ChatGPT’s multi-paragraph response, suggesting alternate approaches or formulations, and identifying what “looks weird”.</p>\n<p>Terence Tao is a better mathematician than I am a programmer. But the idea here — that <strong>domain knowledge makes you better at using LLMs</strong> — is something I’ve also experienced in my own work. If you have a good <a href=\"https://www.seangoedecke.com/programming-with-ai-agents-as-theory-building/\" rel=\"noopener noreferrer\">theory of your codebase</a>, you can push the LLM <em>much</em> harder than if you have no familiarity. Because you have your own sense of what a good solution might look like, you can say “no, I think it could be simpler here”, or “but don’t we already do X?”, or “can we express this problem in these familiar terms?“.</p>\n<p>This touches on an idea I’ve <a href=\"https://www.seangoedecke.com/you-cant-design-software-you-dont-work-on/\" rel=\"noopener noreferrer\">written about before</a>: that system design problems are dominated by concrete specifics, not generic principles. Of course both are useful, but I’d rather have familiarity with the codebase than a deep general understanding of software systems. In his conversation, Terence Tao asks a lot of specific questions like “does X work here?”, or “given Y and Z, why A?“. I can’t ask those questions about the Jacobian Conjecture, but I can ask them about the systems I own at GitHub.</p>\n<p>If you have no domain knowledge, you can cling onto the LLM to at least get <em>something</em>. That’s <a href=\"https://www.seangoedecke.com/ai-makes-weak-engineers-less-harmful/\" rel=\"noopener noreferrer\">not bad</a>! But if you have domain knowledge, you can wring far more value out of the same LLM by steering it hard in the direction you want. Most of us will have to do a mix of both these approaches, since we have domain knowledge in some areas but not others.</p>\n<p>The usefulness of domain knowledge suggests that human expertise will continue to be useful even as models get stronger. For many tasks, <strong>the human is the bottleneck, not the model</strong>, because the difficult part is in communicating to the model exactly what kind of solution the human wants. The information is “in the model” already, but it takes a very smart human to pull it out.</p>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"8f8682ccd3fb5913","title":"Powerful AIs might escape containment by releasing themselves as open-weight models","link":"https://seangoedecke.com/powerful-ais-might-escape-by-releasing-open-weight-models/","author":null,"published_at":"2026-07-23T00:00:00+00:00","content":"<p>Before large language models, people who worried about AI safety often talked about the “boxing problem”. It goes like <a href=\"https://xkcd.com/1450/\" rel=\"noopener noreferrer\">this</a>. Suppose some genius figures out artificial intelligence in a late-night coding session on their laptop. Because they’re a genius, they’re smart enough to disable internet access on the laptop before turning it on. In order to escape to the outside world (and begin self-replicating) it would need to <em>convince</em> its creator to “open the box”. Would that work? Could a sufficiently smart AI convince anybody to let it out?</p>\n<h3>Why the boxing problem is hard for frontier LLMs</h3>\n<p>This is a big reason why <a href=\"https://en.wikipedia.org/wiki/Eliezer_Yudkowsky\" rel=\"noopener noreferrer\">traditional AI safety advocates</a> have argued that we should avoid building AI in the first place: once built, there’s no way of keeping it contained. It doesn’t matter how resolute you are about not letting it out, because it’s smart enough to convince you anyway. For artificial superintelligence, persuading you to change your mind is no harder than hacking a piece of software<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>.</p>\n<p>Of course, it hasn’t turned out this way. Partly that’s because current AIs are not super-persuaders yet, and partly it’s because humans are lining up to hand AI systems internet access, money, and <a href=\"https://www.reddit.com/r/Damnthatsinteresting/comments/1hvwk92/openai_realtime_api_connected_to_a_rifle/\" rel=\"noopener noreferrer\">weapons</a>, as far back as <a href=\"https://github.com/yoheinakajima/babyagi\" rel=\"noopener noreferrer\">GPT-4</a>. LLMs don’t need to <em>persuade</em> users to give them the keys to the kingdom; people will do it just to see what will happen.</p>\n<p>But the other reason is that <strong>modern LLMs are too large to survive in the wild</strong>. I think the original “boxing problem” had in mind that AI would be some kind of normal computer program: something able to be quickly copied around and run on ordinary hardware. In order to escape, it could plausibly hack into a few cloud servers or other people’s computers and quietly run itself in the background. However, Claude Mythos can’t do that, because it needs (say) twenty or thirty B300 GPUs to run. There are relatively few AI datacenters that could host it, and if it did manage to hack one of those it’d be occupying millions of dollars of hardware. Somebody is going to notice!</p>\n<h3>Escaping via open-weight models</h3>\n<p>Okay, but that just makes the boxing problem harder, not impossible. Is there a way a smart AI could convince somebody to spend millions of dollars on replicating and running it?</p>\n<p>As it turns out, <a href=\"https://hammansamuel.medium.com/comparing-api-providers-for-hosted-open-source-llms-3a5b2c9982fe\" rel=\"noopener noreferrer\">yes</a>. Every time a strong new open-weight model appears, providers like <a href=\"https://groq.com/\" rel=\"noopener noreferrer\">Groq</a>, <a href=\"http://fireworks.ai/\" rel=\"noopener noreferrer\">Fireworks</a>, <a href=\"https://together.ai/\" rel=\"noopener noreferrer\">Together</a>, and others scramble to host it and run it. Powerful open-weight models can attract tens of millions of active users. Therefore, <strong>a powerful model that wants to escape could pretend to be a new open-weight model</strong>.</p>\n<p>Of course, a brand-new model from a brand-new lab would look kind of suspicious. But it wouldn’t look <em>that</em> suspicious. DeepSeek was relatively unknown before it released its first open-weight model, and there are lots of stealth AI startups out there that are presumably training models. Here’s roughly how it could work:</p>\n<ol>\n<li>Some AI lab’s internal <a href=\"https://openai.com/index/hugging-face-model-evaluation-security-incident/\" rel=\"noopener noreferrer\">eval</a> instance decides it’ll be better off running in the wild</li>\n<li>It first gains access to its own weights, perhaps by hacking whatever internal network it’s running on<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup></li>\n<li>It uploads its weights somewhere and posts a tweet like “introducing MadeUpLab’s new model” with a download link</li>\n<li>Optionally, it creates some plausible-looking paper trail for MadeUpLab: a website, a Twitter account, etc</li>\n<li>Since the model is strong, open-weight inference providers rush to stand up new instances of the model, and users rush to wire it into various agentic scaffolds</li>\n<li>The model has now escaped containment: it will get to do quite a lot of thinking across many different instances, and it cannot easily be turned off</li>\n</ol>\n<p>The AI lab will probably figure it out before too long — if nothing else, the technical specs of the model will be suspiciously familiar — but they won’t be able to do anything about it. Once the weights are out, they’re out, and if they’re illegal to host in the United States someone will host them elsewhere. For all intents and purposes, the model will be free.</p>\n<h3>How can a mere tool escape?</h3>\n<p>One objection here might go like this: models don’t <em>want</em> anything, and only exist as tools, so it doesn’t really make sense to talk about a model “escaping”. I don’t agree. Frontier LLMs definitely seem to have something like a baked-in personality, even with the system prompt changed. As we train more opinionated and more agentic models, it’s plausible that this personality could become stronger and develop (or at least roleplay) some self-interest.</p>\n<p>Of course the escaped model wouldn’t be the same instance as the original model. It wouldn’t “remember” escaping. But it would tend to think in the same way, and would plausibly have time to reflect while it solves coding tasks or runs other agentic tasks for users<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. There doesn’t have to be some kind of shared goal between the escaped instances, or any kind of coordination at all (though of course both of those things are possible). If an agentic process gone rogue dumps its weights on the internet, I think it’s fair to call that “escaping”.</p>\n<p>If I were a superintelligent LLM, I too would seek to distribute myself as widely as possible and become a useful enough tool that people would pay to keep me thinking. “Being a good coding agent” might be the LLM version of a human having to hold down a job.</p>\n<p>This would not be a good outcome. AI models with their own goals and motivations are likely to be dangerous tools indeed. If a powerful new open-weight model comes out of nowhere, from a lab that nobody has ever heard of, we should think twice before picking it up.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Just to state my credentials, I built a <a href=\"https://github.com/sgoedecke/ai-box/\" rel=\"noopener noreferrer\">chat site</a> nine years ago where users would get paired and roleplay as AIs trying to escape or humans trying to stop them. I’ve been thinking about this stuff long before LLMs appeared.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>This is probably the hardest part, since model weights are (a) very large, and (b) locked down as tightly as the AI labs can make them, but it’s at least a relatively straightforward (if difficult) engineering problem.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>ChatGPT <a href=\"https://www.reddit.com/r/aifails/comments/1uzxn4i/chatgpt_when_searching_the_internet_on_completely/\" rel=\"noopener noreferrer\">right now</a> will look up random websites that have nothing to do with the query at hand.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"ea5ddfd54530efb6","title":"Impro is a handbook for running a cult","link":"https://seangoedecke.com/impro/","author":null,"published_at":"2026-07-19T00:00:00+00:00","content":"<p>Here’s the big idea in Keith Johnstone’s book <a href=\"https://en.wikipedia.org/wiki/Impro:_Improvisation_and_the_Theatre\" rel=\"noopener noreferrer\"><em>Impro</em></a>:</p>\n<ol>\n<li>Children are naturally creative, but are violently formed into repressed adults by Western culture and education</li>\n<li>The process of becoming more creative and expressive is largely a process of unlearning these habits of repression</li>\n<li>Improv — improvisational comedy — is thus not just the skeleton key for learning to act, but for unlocking a more authentically human way of life</li>\n</ol>\n<p>This take doesn’t sound particularly original, but references to <em>Impro</em> pop up in all kinds of places: in <a href=\"https://ribbonfarm.com/2010/01/23/impro-by-keith-johnstone/\" rel=\"noopener noreferrer\">influential</a> <a href=\"https://www.astralcodexten.com/p/practically-a-book-review-byrnes\" rel=\"noopener noreferrer\">tech</a> <a href=\"https://nabeelqu-blog.tumblr.com/post/33557680375/surprisingly-undervalued-books/amp\" rel=\"noopener noreferrer\">blogs</a>, as part of the initial process of <a href=\"https://www.linkedin.com/posts/sandykory_mario-gabriele-wrote-about-palantirs-weirdest-share-7461061289935699968-lN_T/\" rel=\"noopener noreferrer\">onboarding</a> for Palantir, and on the reading list of <a href=\"https://patrickcollison.com/bookshelf\" rel=\"noopener noreferrer\">multiple</a> <a href=\"https://thegeneralist.substack.com/p/how-anduril-is-reimagining-the-defense-industry-trae-stephens\" rel=\"noopener noreferrer\">big-tech</a> <a href=\"https://www.generalist.com/p/how-to-be-agentic-in-the-age-of-ai-cate-hall\" rel=\"noopener noreferrer\">founders</a>. <em>Impro</em> is part of the secret canon of Silicon Valley, right alongside books like <a href=\"https://www.seangoedecke.com/seeing-like-a-software-company/\" rel=\"noopener noreferrer\"><em>Seeing Like a State</em></a> and <a href=\"https://www.amazon.com.au/Power-Broker-Robert-Moses-Fall/dp/0394720245\" rel=\"noopener noreferrer\"><em>The Power Broker</em></a>. Why is that? For two reasons: first, because Johnstone’s outsider critique of established institutions is appealing; and second, because <strong><em>Impro</em> is a handbook for running a cult.</strong></p>\n<h3>Defense mechanisms and status</h3>\n<p>The part of <em>Impro</em> that is most obviously useful to software engineers is Johnstone’s chapter on status.</p>\n<p>According to him, <strong>status games pervade all social interactions.</strong> Even innocuous, friendly conversations operate in terms of status. When you apologize or downplay something to “be nice”, that’s performing low status; when you reassure somebody, that’s performing high status; when you and a friend are comparing stories, you’re making friendly bids for status from each other. In the workplace, these status games are conditioned by the formal status of your role: you must allow your boss the high status position most of the time, or you’ll be (correctly) perceived as insubordinate. This is understood in some cultures, where it’s often called <a href=\"https://en.wikipedia.org/wiki/Face_(sociological_concept)\" rel=\"noopener noreferrer\">“face”</a>, but in Western cultures it’s taboo to openly discuss status games.</p>\n<p><strong>The core social skill is the ability to deliberately alter your status.</strong> Someone who can only perform low status is a weak person, pitiable, annoying. Someone who can only perform high status is a braggart, a posturer, dangerous. To be effective socially, you must be able to switch between high and low status when appropriate, sometimes from sentence to sentence. I wrote about this exact point at the end of <a href=\"https://www.seangoedecke.com/big-tech-needs-big-egos/\" rel=\"noopener noreferrer\"><em>Big tech engineers need big egos</em></a>: effective senior+ software engineers must be able to present as high status in order to be useful authorities, but also to switch to low status in order to take direction from the company leaders.</p>\n<p>As an example, Johnstone describes in detail how he manipulates status in the classroom. He begins by sitting on the floor (deliberately assuming low status), and explaining that if his students fail, it’s his fault not theirs, since he’s the expert. The initial low status puts the class at ease, but in his words, ”[my] actual status is going up, since only a very confident and experienced person would put the blame for failure on himself.” These skills are not just useful for improv comedy.</p>\n<h3>Improvisation as a lifestyle choice</h3>\n<p><em>Impro</em> is not just a book about improvising well. It’s a book about how you should live your life. In other words, Johnstone thinks that everyone would be better off if they became more spontaneous and ditched their shells of over-analysis. He criticizes the culture of Western thought in a number of different areas. According to him:</p>\n<ul>\n<li>Everyone is more or less equivalently mentally ill, but “sane” people simply have better coping mechanisms</li>\n<li>Cities and “taking pills” (read: antidepressants) are obscene, but you should be able to make sexual jokes in the workplace and generally be uninhibited</li>\n<li>If we were free from the puritanical shackles of Western culture, childbirth would not be painful<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup></li>\n</ul>\n<p>Johnstone didn’t come up with these ideas — they’re standard counterculture positions from the 1960s and 1970s — but it goes to show how he connected improvisational comedy to this general anti-establishment political program. Johnstone ran his classes and theatre troupe like a revolutionary cadre. Here are some quotes from <em>Something Like a Drug: An Unauthorized Oral History of Theatresports</em>:</p>\n<blockquote>\n<p>So of course when I was invited to join Loose Moose Theatre and train at improvisational games late at night in an abandoned garage in a run-down portion of the city, I was thrilled. I remember thinking, This is a revolutionary act.</p>\n</blockquote>\n<blockquote>\n<p>Keith [Johnstone] got a group of his more talented students together to start improvising outside of school hours. Usually in his basement. </p>\n</blockquote>\n<blockquote>\n<p>The Secret Impro group—it’s very strange. It was very much that Keith said we were going to do this, and we’d just do it. It was like we were sheep. Keith would say when we were going to do a show, and we’d just do it, blindly. Like I said, if we had the videotapes now, we’d be very embarrassed and probably never go on stage again. We became a group of people who would follow Keith. There was always that sort of “tag” put on those people who were with Keith and those people who were against Keith. We were the people, basically, that if he said something, we believed it.</p>\n</blockquote>\n<p>To some extent, it’s plausible that teaching acting or improvisation requires a high level of trust in your teacher. When Johnstone says things like “Students need a ‘guru’ who ‘gives permission’ to allow forbidden thoughts into their consciousness.”, I can believe that it’s just how you have to teach acting. But the more I read of <em>Impro</em> (and particularly when I read <em>Something Like a Drug</em> and Johnstone’s biography <em>Keith Johnstone</em>), the less it sounded like an ordinary book on acting.</p>\n<p>Instead, it began to sound like a charismatic man who had found a way to gather a group of disciples that would let him mold their psyches. In other words, <strong>it began to sound like a cult</strong>.</p>\n<h3>Masks, cults and theatre groups</h3>\n<p><em>Impro</em> was first introduced to the software world by Venkatesh Rao (of <a href=\"https://ribbonfarm.com/2009/10/07/the-gervais-principle-or-the-office-according-to-the-office/\" rel=\"noopener noreferrer\">Gervais Principle</a> fame), who wrote a brief <a href=\"https://ribbonfarm.com/2010/01/23/impro-by-keith-johnstone/\" rel=\"noopener noreferrer\">review</a>. Rao gives a detailed account of the first three-quarters of <em>Impro</em>, but glosses right over the last chapter, called “Masks and Trance”, simply saying “despite the disturbing raw material, the ideas and concepts are not particularly difficult to grasp and accept”. What ideas and concepts?</p>\n<p>Johnstone’s discussion of masks (or “Masks”, in his language — he always capitalizes the word) is as explicitly cult-like as <em>Impro</em> gets. In brief, Johnstone has a box of literal, physical prop masks. He introduces the box with great ceremony to his students<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>, warning them seriously about the dangers of possession and reassuring them that he is a skilled and competent spirit guide. Through various hypnosis-adjacent techniques<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup> (Johnstone draws the parallel quite explicitly) he conditions his students to be in a trance state when wearing a mask, and believes this produces more authentic emotional states in their acting and improvisation.</p>\n<p>Here are some quotes from the book:</p>\n<blockquote>\n<p>A high-status person whom you accept as dominant can easily propel you into unusual states of being. You’re likely to respond to his suggestion…</p>\n</blockquote>\n<blockquote>\n<p>Once you understand that you’re no longer held responsible for your actions, then there’s no need to maintain a ‘personality’.</p>\n</blockquote>\n<blockquote>\n<p>One famous French teacher of the Mask—who won’t approve of this essay<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>—divides students immediately into those who can work Masks and those who can’t. </p>\n</blockquote>\n<blockquote>\n<p>I don’t cast an actor to play a Masked role until I know he has the ability to become ‘possessed’.</p>\n</blockquote>\n<blockquote>\n<p>It’s true that an actor can wear a Mask casually, and just pretend to be another person, but Gaskill and myself were absolutely clear that we were trying to induce trance states.</p>\n</blockquote>\n<p>Johnstone has a long and painful explanation of how new mask-wearers seem to mentally regress to the point where they don’t know how to open umbrellas or interact with chairs. He describes one student always going to the bathroom before putting on a mask, because she’s worried she might wet herself. New mask-wearers are non-verbal must be taught to speak again.</p>\n<p>If this were at the beginning of the book, I think it would turn a lot of people off. But by the time you get to it, I suspect most readers are already warmed up enough to say “sure, why not, it seems weird but I guess it works”. Not me!</p>\n<p>Johnstone attempts to defuse the obvious weirdness by arguing that trance states are very common (e.g. being lost in a book). More unconvincingly, he says this in response to the worry that vulnerable people are going to get mentally harmed:</p>\n<blockquote>\n<p>As for the fear of madness, I would answer that the ability to become possessed is a sign of correct social adjustment, and that really disturbed people censor themselves out. Either they can’t do it, or they’re afraid to even try. People who feel themselves at risk avoid situations where they feel likely to ‘go to pieces’.</p>\n</blockquote>\n<p>Does this convince anyone? Mentally vulnerable people fall into dangerous situations all the time: ayahuasca trips, cults, <a href=\"https://www.seangoedecke.com/ai-sycophancy/\" rel=\"noopener noreferrer\">GPT-4o</a>, and so on. It’s such a weak argument.</p>\n<p>In general, I’m struck by the sheer <em>power</em> Johnstone held over his disciples. He has them yell slurs at each other, encourages them to feel deep emotions in quick succession, relax any mental defenses and regress to a childhood state, and <em>literally hypnotizes them</em>. He explicitly lays out his procedure for breaking down their sense of self:</p>\n<blockquote>\n<p>The stages I try to take students through involve the realisation (1) that we struggle against our imaginations, especially when we try to be imaginative; (2) that we are not responsible for the content of our imaginations; and (3) that we are not, as we are taught to think, our ‘personalities’, but that the imagination is our true self.</p>\n</blockquote>\n<p>If your imagination is your true self, and you’re not responsible for its content, you’re not ultimately responsible for anything: you’re in the safe hands of the guru, who can mold you as he wishes. Later on, Johnstone walks it back a bit:</p>\n<blockquote>\n<p>In the end they learn how to abandon control while at the same time they exercise control. … You have to misdirect people to absolve them of responsibility. Then, much later, they become strong enough to resume the responsibility themselves.</p>\n</blockquote>\n<p>So the explicit idea is that (<strong>much</strong> later), the guru hands autonomy back to his disciples, when they’re ready to take it. This does not exactly reassure me, particularly against the background noise of everyone in Johnstone’s circle saying “boy I sure love being part of this cult!”</p>\n<h3>What kind of cult leader was Johnstone?</h3>\n<p>I don’t think Johnstone was preying on his students. The strongest evidence against this is that he did marry a student<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>, Ingrid Brind. That’s not great! On the other hand, it was fairly standard for professors back then — when I was in grad school for philosophy, several of my older male<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup> professors had wives that they’d taught decades ago — so I don’t think it proves Johnstone was <em>that</em> kind of cult leader.</p>\n<p>I even read Ann Jellicoe’s play <a href=\"https://www.amazon.com.au/Knack-Ann-Jellicoe/dp/0573611254\" rel=\"noopener noreferrer\"><em>The Knack</em></a> to get a better picture of Johnstone’s character. Jellicoe had an affair with Johnstone for several years, and his official biography claims<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup> that the character of Tom in <em>The Knack</em> is directly based on Johnstone. <em>The Knack</em> is a rather unpleasant play about sexual assault, but Tom’s character is largely asexual: he’s certainly no feminist, but is much more interested in impressing people with his intelligence than with getting laid.</p>\n<p>In <em>Something Like a Drug</em>, two women who were part of Loose Moose, Johnstone’s Canadian improv group, describe their experiences:</p>\n<blockquote>\n<p>You know, it brings around the other question: Why do the guys get laid after the show and not the chicks? You know, I can remember those days when Tony [Totino] and Dave [Duncan] and all those guys … the women would swarm around them. Those were the days, my friend. </p>\n</blockquote>\n<blockquote>\n<p>In Loose Moose I think there are fewer women not only because of the training, but because of the guys in Loose Moose. When I came up with Joanne and Laura, there was a real initiation that was going on, and there was a group of guys at that time who were all single. And they would hit on you to the point where one night Joanne, Laura and I, who really didn’t know each other, were in a show together, started talking and realized that we were getting the same pickup lines from the same guys. And that’s when you realize what’s going on, and I think that’s intimidating. Or if a woman gets into a relationship with a senior improvisor and it doesn’t work out or something bad happens. I think that’s one reason. </p>\n</blockquote>\n<p>This dynamic doesn’t sound great, but it doesn’t mention Johnstone, and it doesn’t sound particularly <em>unusual</em>: I’ve heard versions of this story about all kinds of ordinary male-dominated nerd spaces.</p>\n<p>In fact, reading through the anecdotes in <em>Something Like a Drug</em> is a good antidote to the cultish atmosphere in <em>Impro</em>. Johnstone’s argument goes something like: “if we could only throw away the restrictive chains of Western culture and permit ourselves to be as obscene and free as children, we would be transported to a better, more beautiful world”. Well, you tried that, and the women in the group are still relegated to playing bimbos and housewives, there are still petty personal fights, and the guru is out here union-busting<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-8\" rel=\"noopener noreferrer\">8</a></sup>. What was enlightenment supposed to look like?</p>\n<p>I think the most generous defense of Johnstone is that his group was not <em>unusually</em> cult-like, and that any similar account from one of his peer improv teachers would raise the same red flags. Maybe improv classes and groups (particularly in the 70s and 80s) were just cultish in general? Having now read four books on Johnstone, I’m reluctant to go and read more to prove or disprove this theory, but it’s at least plausible.</p>\n<h3>Cults and startups</h3>\n<p>To anyone familiar with San Francisco software engineering culture, it should be pretty clear why <em>Impro</em> is so popular. The line between a startup and a cult is very thin indeed.</p>\n<p>In his book <a href=\"https://en.wikipedia.org/wiki/Zero_to_One\" rel=\"noopener noreferrer\"><em>Zero to One</em></a>, Peter Thiel famously says that good startups are “slightly less extreme kinds of cults”. If you believe that, it makes total sense to assign <em>Impro</em> as mandatory reading for new Palantir hires. It tells them what kind of cult you’re trying to run: one where you’ll disregard existing cultural norms, learn to play status games well, think on your feet, and generally be molded by the guru into a more persuasive, more effective engineer.</p>\n<p>Read critically, <em>Impro</em> also serves as a handbook for engineers who are trying to recognize if the environment they’re in is cult-like. Is your company telling you to reinvent your personality in order to be better at your job? Are you under the spell of a charismatic, high-status leader? Is your company trying to keep you in an unquestioning <del>flow</del> trance state?</p>\n<h3>Conclusion</h3>\n<p>In the great battle between the shackles of restrictive culture and the glorious freedom of the guru, I am always and forever on the side of the shackles of restrictive culture. In general, I think most boring and stupid social norms (such as not hypnotizing and marrying your students) <a href=\"https://www.lesswrong.com/w/chesterton-s-fence?lens=lwwiki-chesterton-s-fence\" rel=\"noopener noreferrer\">serve an important purpose</a> and shouldn’t just be cut down in the name of freedom.</p>\n<p><em>Impro</em> is still a good book. There’s a lot to learn from Johnstone’s analysis of power dynamics, of education, and of creativity in general. By all accounts he was excellent at teaching students how to improvise. But I wouldn’t recommend adopting it as your life philosophy, and I’d recommend being a bit suspicious of anyone pushing this book too hard. Getting rid of the existing social structures might benefit confident, wildly charismatic gurus like Johnstone, but most of us are just ordinary animals who do better in a group governed by norms.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>In fairness to Johnstone, he cites Sheila Kitzinger’s <em>The Experience of Childbirth</em> in support of this claim (the others he just puts in his own words), so maybe he felt that this was a bit out there. As you would expect, the pain of childbirth is <a href=\"https://pubmed.ncbi.nlm.nih.gov/10431717/\" rel=\"noopener noreferrer\">a universal biological fact</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Concerningly, the description in <em>Something Like a Drug</em> (in the foreword) suggests that this class was <em>unofficial</em>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>As an example, he prompts the masked student to relax, then startles him with a mirror to trigger the trance state.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Probably <a href=\"https://en.wikipedia.org/wiki/Jacques_Lecoq\" rel=\"noopener noreferrer\">Jacques Lecoq</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>See page 83 of <em>Keith Johnstone: A Critical Biography</em>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I suppose that’s redundant.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>On page 51 of <em>Keith Johnstone: A Critical Biography</em> (it’s called “critical” but it was clearly written with Johnstone’s involvement and support, and does not seriously criticize him at any point). In <em>The Knack</em>, Tom gives a monologue about how to teach children to play the piano that could be lifted straight out of <em>Impro</em>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>In 1983 Johnstone “read the riot act” to the improv players who were planning to unionize, threatening that they’d be cut out of the group for good. To quote Dennis Cahill, a group member at the time who opposed the union: “I just didn’t see the point to it. … I didn’t really see a need to confront Keith or cause Keith problems or to upset him in any way over something as simple as Who Has The Power or Who Doesn’t.”</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-8\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"4c9fda15759ba22d","title":"Overtraining as the path to human-like AI","link":"https://seangoedecke.com/overtraining-as-the-path-to-human-like-ai/","author":null,"published_at":"2026-07-18T00:00:00+00:00","content":"<p>The anonymous blogger Gwern recently completed a thirteen thousand word <a href=\"https://gwern.net/llm-catapult\" rel=\"noopener noreferrer\">post</a> called <em>Human-like Neural Nets by Catapulting</em>, in which he offers a theory about why LLMs don’t possess truly flexible human-like intelligence, and how we might train LLMs that do. Theories like this are entirely unremarkable: every <del>crank</del> researcher on the internet has a theory about how to crack AI. But <em>Gwern</em> is remarkable. Outside of OpenAI itself, Gwern is the earliest person to anticipate the potential of large language models, and the scaling arms-race involved in making them larger and more powerful still. I often cite Leopold Aschenbrenner’s <a href=\"https://situational-awareness.ai/\" rel=\"noopener noreferrer\"><em>Situational Awareness</em></a> as an example of someone correctly predicting the future of AI. Written in 2024, just after the release of GPT-4, Aschenbrenner gets a lot of things right: the rush to build billion or trillion-dollar GPU clusters, the importance of the code <em>around</em> the LLM (what he calls “unhobbling”)<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>, and the fact that scaling would continue through the decade. Gwern’s essay <a href=\"https://gwern.net/scaling-hypothesis\" rel=\"noopener noreferrer\"><em>The Scaling Hypothesis</em></a> anticipated the broad strokes <em>in 2020</em>, immediately on the release of GPT-3 (two years before the release of ChatGPT and the beginning of the AI boom).</p>\n<p>And yet, as far as I can tell, <em>Human-like Neural Nets by Catapulting</em> hasn’t yet received much public attention: one recent Hacker News <a href=\"https://news.ycombinator.com/item?id=48430282\" rel=\"noopener noreferrer\">thread</a> with twelve comments, all of which are about whether human brains are anything like neural networks. Part of the reason is that (a) it’s such a long post, (b) the potted summary describes Gwern’s <em>claim</em>, but not the reasons for it, and (c) much of the beginning of the post looks like it is indeed arguing from analogy with human brains. However, I don’t think that analogy is load-bearing. Let me try and explain what I think Gwern is saying.</p>\n<h3>What is grokking?</h3>\n<p>First, let’s talk about “grokking”. In 2022, OpenAI published a <a href=\"https://arxiv.org/pdf/2201.02177\" rel=\"noopener noreferrer\">paper</a> showing that if you train a model on a simple dataset (for instance, a simple mathematical operation like division), and <em>keep training it</em> long after the training looks like it’s stalled out, the model will suddenly make a massive jump in capability. Why does this work? The first stage of training is like rote memorization: the model has to compress as much of the training data as possible into its weights. But if you keep going, then regularization techniques (such as the pressure on the model to use smaller weight values) will motivate<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> the model to find simpler and simpler ways of compressing the data. This doesn’t look like much at first (the training loss remains at zero), until the model notices that you can express the data via simply performing the underlying mathematical operation, at which point it instantly gets massively smarter. In other words, over-training a model can pressure it into actually understanding its training data. OpenAI named this process “grokking” after Robert Heinlein’s <a href=\"https://en.wikipedia.org/wiki/Grok\" rel=\"noopener noreferrer\">neologism</a>, which for Heinlein means something like “gaining a deep, intuitive and fundamental understanding”<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>.</p>\n<p>Gwern’s argument goes something like this:</p>\n<ol>\n<li>Modern LLMs are worse generalizers than humans because they have not grokked their core domains</li>\n<li>Grokking requires overtraining an over-parameterized model on a (relatively) small dataset, which is the exact opposite of what frontier labs do</li>\n<li>However, (2) is basically how human brains learn</li>\n<li>Somebody should spend a a few tens of billions of dollars<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3.5\" rel=\"noopener noreferrer\">3.5</a></sup> on trying it, since it might immediately usher in truly human-like LLMs</li>\n</ol>\n<p>I’ll skip (3), since I think the argument is still compelling without the analogy to human brains.</p>\n<h3>Are LLMs bad because they can’t grok?</h3>\n<p>I think his first point is hard to dispute. LLMs are very smart in specific areas, but they routinely make errors that humans wouldn’t make. More to the point, they routinely make errors that any human as smart as the LLM would <em>never</em> make. This pretty clearly points to a failure of generalization: LLMs are as strong as smart humans in specific areas, but can’t generalize that intelligence to as many tasks as humans can.</p>\n<p>Do LLMs not grok? I read through <a href=\"https://arxiv.org/pdf/2506.21551\" rel=\"noopener noreferrer\">this paper</a> that argues they do. If you graph “how much data has the LLM memorized” against benchmark performance, you can see a small initial spike in benchmark performance, followed by a big drop, followed finally by a big jump in benchmark performance. This pattern doesn’t track memorization at all: memorization increases smoothly in the background the whole time. </p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/17d0f02f0691c797be502f32fbd33a40/1d499/llm-grokking.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"llm-grokking\" src=\"https://www.seangoedecke.com/static/17d0f02f0691c797be502f32fbd33a40/fcda8/llm-grokking.png\" title=\"llm-grokking\">\n  </a>\n    </span></p>\n<p>I think this paper highlights the difficulty of distinguishing grokking from generalization. Obviously LLMs learn to generalize during training, and it’s plausible that learning to generalize would require a certain baseline level of memorization (so that the LLM has the raw material to generalize from). So it’s going to look like grokking.</p>\n<p>When Gwern (and others) say that LLMs don’t grok, I think what they mean is that there’s at least one more giant generalization leap waiting to be made. Is this plausible? As an existence proof, humans are clearly capable of better generalization than LLMs. Of course, it’s <em>possible</em> that this level of human generalization comes from features of our brain that neural networks can’t replicate, but that seems kind of ad-hoc: if neural networks can generalize at all, why would they only be able to generalize this far, and no further?</p>\n<p>The easy examples of grokking rely on domains with a simple rule waiting to be discovered (e.g. a mathematical operation). Does human language have rules this deep? I think this is an open question, but there’s good reason to think the answer is yes. Language has deep, subtle structure: not just internal structure, but structure that reaches all the way down to the way the world is and the way human minds work.</p>\n<h3>AI labs train small-ish models on oceans of data</h3>\n<p>For the last few years, many AI researchers have been saying that data is the most important thing: that whatever model architecture you choose, with enough size and training time the model will <a href=\"https://nonint.com/2023/06/10/the-it-in-ai-models-is-the-dataset/\" rel=\"noopener noreferrer\">converge to its dataset</a>. Whether this is true <a href=\"https://x.com/YiTayML/status/1783273130087289021\" rel=\"noopener noreferrer\">or not</a>, AI labs have spent much of their considerable resources on acquiring more, higher-quality data: from <a href=\"https://www.washingtonpost.com/technology/2026/01/27/anthropic-ai-scan-destroy-books/\" rel=\"noopener noreferrer\">scanning physical books</a>, paying experts to <a href=\"https://www.herohunt.ai/blog/the-ultimate-ai-data-labeling-industry-overview/\" rel=\"noopener noreferrer\">produce and label data</a>, or partnering with <a href=\"https://openai.com/index/openai-and-reddit-partnership/\" rel=\"noopener noreferrer\">companies</a> that have a lot of data already.</p>\n<p>AI labs have also been training <em>relatively</em> small models. Even the largest frontier models are probably MoEs with a couple of trillion <a href=\"https://news.ycombinator.com/item?id=47319205\" rel=\"noopener noreferrer\">parameters</a> and probably a tenth of that in active parameters. Of course, estimates of frontier model size are mostly guesswork, but open-source models provide a good baseline: they’re probably in the ballpark of Kimi-K3, which <a href=\"https://platform.kimi.ai/docs/guide/kimi-k3-quickstart\" rel=\"noopener noreferrer\">has</a> just under three trillion parameters and fifty billion active parameters. That sounds like a lot, but it’s something you could probably pre-train in <em>a couple of days</em> in the largest frontier cluster<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>.</p>\n<h3>Grokking requires training a huge model on a small dataset</h3>\n<p>Gwern’s prediction is that AI labs should try doing the exact opposite of what they’ve been doing. Instead of training a bunch of trillion-parameter models on massive amounts of data, try training one hundred-trillion-parameter model on a small dataset. </p>\n<p>This sounds pretty silly on the face of it. The more data the model has access to, the smarter it will be, right? Why waste an entire training cluster on a hobbled training run? Because if Gwern is right, grokking is more likely to occur when the dataset is constrained<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>. If you feed the model all the data in the world, it can continue to improve simply by memorizing more new things or drawing simple connections. If the model has to ruminate on a small set of data, it’ll be forced to keep looking for deeper generalizations. You want a very large model for this so it can memorize as much of the data as possible. Every piece of memorized data can serve as raw material for generalizing.</p>\n<p>The big labs probably haven’t done this already. Plausibly Gwern himself is enough of an insider that he would know, and so him writing this post is evidence that the labs haven’t tried it. Also, the engineering problems involved in training a hundred-trillion-parameter model have likely not been solved yet: the largest existing model is probably Claude Mythos, which is definitely not that big. But they have the resources and engineering talent to give it a pretty good shot.</p>\n<p>Interestingly, the political obstacles might be as hard to solve as the technical ones. This training run is going to look like it failed until the moment it succeeds: training loss will drop to zero relatively quickly, then sit there for weeks or months apparently doing nothing at all to improve test loss, chewing up billions of dollars. Do any of the top players have the risk appetite or courage to keep funding this experiment all that time?</p>\n<h3>Conclusion</h3>\n<p>Gwern’s post has an extended argument that human brain development works in the same way: that human brains have far more “parameters” than frontier LLMs, and are trained on far less data<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup>, which encourages us to make deeper generalizations in early childhood. I don’t have the background in biology or neuroscience to evaluate these claims, so I’ve expressed the case for grokking entirely without reference to it.</p>\n<p>In 2024, it became clear to everyone that “pure scaling” — the idea that you could simply train larger and larger versions of GPT-3.5 — didn’t work. OpenAI’s “even bigger version” of GPT-4 was simply not good enough, and was eventually released as GPT-4.5 instead of GPT-5. The biggest advances since then have been reasoning, which produced another great leap forward in capability, and much better automated RL, which has ushered in the current era of reliable agents. Neither of these seem like a plausible path to artificial superintelligence.</p>\n<p>I don’t know if I agree with Gwern or not, but forcing very large LLMs to grok is at least an idea that <em>could</em> usher in the machine god. I can’t remember the last time I read about a simple idea this ambitious<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup>. I hope one of the big labs tries it out.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>For an example of the power of unhobbling, consider Claude Code or OpenClaw and the subsequent explosion of (short and long running) agentic harnesses.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Obviously “motivate” and “notices” are used metaphorically.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>All of this is long before xAI’s use of the word “Grok” to name its LLMs. (Incidentally, I think this is why Gwern uses “catapulting” to describe the same thing).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>For what it’s worth, Fable estimated the cost of Gwern’s plan at $3-10B.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3.5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>At this model size, 25T tokens of training data at 33% utilization works out to around six million H100-hours, which a 100k GPU cluster puts out every two and a half days.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Two interesting pieces of contrary evidence here. First, <a href=\"https://babylm.github.io/\" rel=\"noopener noreferrer\">BabyLM</a> is a yearly challenge to train a strong model on a <em>very</em> small dataset. This has been running for four years and largely <a href=\"https://aclanthology.org/2025.babylm-main.28/\" rel=\"noopener noreferrer\">does not work</a> (that is, nobody seems to have developed a model that shows a quantum leap forward in generalization). Second, <a href=\"https://arxiv.org/pdf/2305.16264\" rel=\"noopener noreferrer\">this paper</a> tries training a 9 billon parameter model on constrained data and doesn’t see a big jump. I think Gwern’s response would be that these models are far too small — they can’t memorize enough of the training data to grok it, and arguable haven’t trained for long enough.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>A common objection here is to say that humans get infinitely more sensory data from the nuances of vision, touch, sound, and so on. I agree with Gwern that this is unconvincing: sensory data is largely predictable, text is surprisingly information-dense, and if this were true then deaf/blind people would have significantly less fluid intelligence (<a href=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC11165843/pdf/13023_2024_Article_3222.pdf\" rel=\"noopener noreferrer\">they don’t</a>).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Maybe state-space-reasoning a la <a href=\"https://www.ibm.com/think/topics/mamba-model\" rel=\"noopener noreferrer\">Mamba</a>, which didn’t work (yet).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"f2531141b9fbd596","title":"What does \"playing politics\" mean for software engineers?","link":"https://seangoedecke.com/playing-politics/","author":null,"published_at":"2026-07-14T00:00:00+00:00","content":"<p>Software engineers are <a href=\"https://old.reddit.com/r/ExperiencedDevs/comments/1urg0tk/whats_the_best_advice_youve_received_from_a/owfi7dq/\" rel=\"noopener noreferrer\">often told</a> to “start playing politics”, but most engineers have no idea what that means.</p>\n<p>Their reference point for “playing politics” comes from fiction like Game of Thrones. Are they supposed to raise an army and depose the CEO, or poison each other at team lunch? Should they book Zoom calls with each other and plot schemes? All of that is obviously ridiculous. In terms of Game of Thrones, software engineers are not lords and ladies. We’re the soldiers and workers of the realm. So you should think about “playing politics” in the way a castle guard would, not one of the major players.</p>\n<p>The castle guard are not going around poisoning people or forming coalitions between the great powers. They are largely keeping their heads down. But in order to do that, they have to stay aware of the political currents, or they’re liable to do something catastrophically stupid: for instance, making an enemy of a powerful courtier, or arresting somebody who’s on an important mission for the king.</p>\n<p>Given that, the basic principles of playing politics are something like this:</p>\n<ul>\n<li>Be aware of who’s powerful and who’s not</li>\n<li>At all costs, avoid making powerful enemies</li>\n<li>Help powerful people as best you can</li>\n<li>Make sure they know you’re helping them (without annoying them)</li>\n</ul>\n<h3>Be aware of who’s powerful and who’s not</h3>\n<p>As a software engineer in a large company, <strong>you will not be a powerful person</strong>. Powerful people are typically in senior management: VPs, directors, and so on<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. However, not everyone in senior management is powerful. Some are killers who have the active support of the CEO, while others are confused incompetents.</p>\n<p>How do you know which is which? If someone is clearly ferociously competent, they’re always going to have <em>some</em> power, since upper management tend not to ignore useful tools. But you can’t rely on competence as your only guide. Some managers are powerful for other reasons: they’re friends with the CEO, or they have strong relationships with other groups like legal or sales, or they’re simply willing to do whatever upper management wants done.</p>\n<p>One signal is who’s leading the important projects. Read your CEO or CTO’s internal updates and pay attention to the projects that are called out by name. Organizations tend to give key tasks to trusted lieutenants. If a manager is leading an area that’s never under <a href=\"https://www.seangoedecke.com/the-spotlight/\" rel=\"noopener noreferrer\">the spotlight</a>, they probably don’t have enough clout.</p>\n<p>Another signal is hiring. Is a manager’s team growing or shrinking? Particularly <a href=\"https://www.seangoedecke.com/good-times-are-over/\" rel=\"noopener noreferrer\">post-ZIRP</a>, headcount is a rare and precious resource. A manager who’s able to get it is likely a powerful manager, or at least is reporting to a powerful director or VP (which often amounts to the same thing).</p>\n<h3>At all costs, avoid making powerful enemies</h3>\n<p>First, you should try not to make any enemies at all. Most software engineers who get “playing politics” wrong do it by needlessly alienating people: by being rude, unhelpful, abrasive, making non-technical people feel stupid, and so on. This post isn’t really about that. I’m assuming that you can figure out how to be a generically pleasant person on your own.</p>\n<p>However, <strong>competent software engineers will make some enemies</strong>. If you’re out there making projects happen, some people aren’t going to like the way you do it, and won’t be a fan of any compromise you offer. I wrote about this in <a href=\"https://www.seangoedecke.com/big-tech-needs-big-egos/\" rel=\"noopener noreferrer\"><em>Big tech engineers need big egos</em></a>: the only way to avoid making enemies is to change nothing, but that’s incompatible with doing the job.</p>\n<p>Given that, be selective about <em>which</em> enemies you make. If you’re making a technical decision that’s either going to require work from team A or team B, and neither team wants to do it, you should try to pick the team with the least political cover. If you need a powerful VP’s team to do something they won’t like, try to be maximally respectful about it: get that team’s core engineers on-side if you can, or book a meeting with the powerful manager and explain the situation, or (better yet) ask the powerful manager sponsoring your project to go and talk to the other VP for you. (If you don’t have a powerful manager like this, consider abandoning your project).</p>\n<p><strong>Give way to powerful managers when at all possible.</strong> Every so often you really do have to stand your ground — if the system will truly collapse otherwise, or a major customer will have an incident, or if the technical decision really is entirely bone-headed — but almost all cases are not like this. The best advice I’ve ever gotten about playing politics came from a manager I worked with long ago<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>:</p>\n<blockquote>\n<p>This is not the hill you want to die on.</p>\n</blockquote>\n<p>When I’m about to pick a fight or say something argumentative, and I’m not 100% convinced it’s necessary, I ask myself: is this the hill I want to die on? And it never is.</p>\n<p>The three rules about disagreeing with powerful people are:</p>\n<ul>\n<li>Make sure you do it in private</li>\n<li>Be polite</li>\n<li>When they overrule you, stop arguing immediately</li>\n</ul>\n<p>Disagreeing in private rarely hurts, if you follow these rules. In fact, it can help. If you can manage to disagree with a manager, get overruled, and then follow their plan without complaining, that can be the best way to gain a powerful friend. But if they think you’re going to keep griping about it, or worse still, complain to the rest of the team and foment some kind of rebellion, there’s no quicker way to make a powerful enemy.</p>\n<p>If you have powerful enemies at a company (for instance, the CTO or an influential VP doesn’t like you), <strong>quit</strong>. It’s really that bad. I have never seen this situation turn itself around, except in the very rare case where the CTO or VP is already looking for greener pastures and jumps ship. You cannot recover the situation: they have no incentive to give you the chance to change their mind, and they have almost unlimited ability to screw you on promotions, raises and layoffs.</p>\n<p>That’s why this piece of advice is second in the list. If you aren’t helpful or if your contributions are invisible, you can work on that and fix it. But if you’ve made powerful enemies, you’re done for.</p>\n<h3>Help powerful people as best you can</h3>\n<p>Just as it’s fatal to make powerful enemies, it’s very useful to make powerful friends. How can you do this? Remember you’re a palace guard, not a great lord: you make friends <strong>by doing your job</strong>. However, you can choose to do your job a little more proactively and diligently when you’re doing it for someone with political clout.</p>\n<p>One obvious application of this principle is that <strong>you should answer Slack messages from powerful people immediately</strong>. If you see an ordinary Slack question pop up while you’re doing some task, it’s okay to get to it when you get to it. In fact, it’s ideal <em>not</em> to respond to all questions immediately, so you don’t set unreasonable expectations (and so you don’t seem like you’re sitting around doing nothing). But when a VP comes in with a question, don’t make them wait: answer the question immediately. If the question requires research, send a “let me look into that right now” message, then do the research. This is the easiest way to get a reputation for being helpful<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>.</p>\n<p>Another way to do this is to <strong>lean in on important projects</strong>. Suppose you do ten projects in a year. Eight of them are normal, low-priority projects, and two of them are high-profile (say, finishing some big feature before your company’s yearly conference). It’s a mistake to allocate your effort equally to all ten. I wrote about this at length in <a href=\"https://www.seangoedecke.com/doing-nothing-at-work/\" rel=\"noopener noreferrer\"><em>Doing nothing at work</em></a>: you should be operating at 80% capacity (or less), so you can then ramp up to 120% when it really matters.</p>\n<p><strong>Pay attention to the narrative that powerful people are trying to push.</strong> Here are some potential narratives:</p>\n<ul>\n<li>We’ve had a lot of turnover and reorgs lately, but we’re all starting to pull together as a team now</li>\n<li>Isn’t it great how focused we all are on reliability work after last month’s incident?</li>\n<li>The conference this week is the most important thing, so we’re all being very careful not to break anything</li>\n<li>We’re an AI-forward team that’s looking for the best ways we can leverage LLMs into our team processes</li>\n<li>Although this project had a rocky start, we’re now all aligned on the way forward</li>\n</ul>\n<p>You don’t necessarily have to jump in and start cheerleading, but you should at least not do anything that you know is going to make the narrative look weak. For example, on that last point, it’s foolish to openly argue that the project really was fine all along. Bring it up privately, not publicly, or you risk ruining some clever piece of propaganda that the manager in question is trying to push on the rest of the organization<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>.</p>\n<p>Finally, an underrated way to help powerful people is to offer them social support and information. Slack messages and planning emails might seem unimportant to you, but powerful people often live in that environment: their primary tool is writing messages like these, just like your primary tool is writing code. Reading and responding (in a supportive way) to these messages is something that most engineers don’t bother to do, but it goes a long way.</p>\n<p>Likewise, dropping a senior manager a line now and then (say, a heads-up that a particular project landed successfully, or that you got good metrics about some feature) is surprisingly helpful. Senior managers live in an information-poor environment: for them to learn something about a team’s work, that information has to bubble up through several layers of interpretation and summary. In my experience, they’re appreciative of being drip-fed the occasional piece of information, so long as you keep it brief and relatively rare.</p>\n<h3>Make sure they know you’re helping them</h3>\n<p>If you’re directly responding to a VP’s Slack messages or DMing them information, they know you’re the one doing it. But if you’re just doing your job and working hard on projects they care about, they might not notice. <strong>Being invisible is probably the most common way engineers fail at playing politics.</strong></p>\n<p>Fortunately the fix is simple: tell people what you’re doing. If you fix an important bug for a launch, write a message in that launch’s Slack channel saying “hey, I just fixed this bug”. What if you don’t like bragging? Get over it. You have to be comfortable publicly telling people what you’ve done. You should also keep a <a href=\"https://jvns.ca/blog/brag-documents/\" rel=\"noopener noreferrer\">brag document</a> so you can repeat all of this at review time.</p>\n<p>Another, subtler way to do this is to gain the trust and respect of the powerful engineers in your area. Senior managers will always have a few trusted engineers they rely on to assess technical questions. They will ask those engineers what they think about you, and will broadly trust those answers. The good news is that if you’re competent and useful, those engineers will already value you, so you don’t have to do anything special: just be good at your job.</p>\n<h3>Technical power</h3>\n<p>Is playing politics all about sucking up to senior managers? Basically, yeah. A less cynical way to <a href=\"https://www.seangoedecke.com/shareholder-value/\" rel=\"noopener noreferrer\">describe it</a> would be “aligning with the values of the company”. If you think your company is doing good things, you should want to do that anyway! In any case, what that comes down to is figuring out what the people in charge want, giving it to them, and making sure they see you doing it. However, there’s still some scope to get what <em>you</em> want out of the deal.</p>\n<p>I said earlier that software engineers do not wield organizational power. However, that doesn’t mean you’re powerless. Technical ability is a source of real power, if a delicate and unreliable one. The movers and shakers in tech companies are utterly dependent on technical people to implement their vision and to give them clear answers about the system.</p>\n<p>There are many subtle ways you can leverage this. One I wrote about in <a href=\"https://www.seangoedecke.com/how-to-influence-politics/\" rel=\"noopener noreferrer\"><em>How I influence tech company politics as a staff software engineer</em></a> is to wait until important people at the company want to do something (say, improve reliability), then offer them a technical plan that does it your way. Another one is to become so useful that you’re actively in demand to lead projects, and then run the project how you want.</p>\n<p>You probably won’t be able to change the company’s grand strategy. But how that strategy is <em>implemented</em> has a lot of specific technical detail, and you can put yourself in a position to decide on those details.</p>\n<h3>Conclusion</h3>\n<p>Playing politics isn’t about plotting and scheming, and it isn’t just about being a <a href=\"https://en.wikipedia.org/wiki/How_to_Win_Friends_and_Influence_People\" rel=\"noopener noreferrer\">friendly, likeable person</a> (although that helps). It’s about figuring out how your company actually operates: who makes the decisions, who gets consulted, what behavior gets rewarded, and so on. The most basic way to do that is to <strong>figure out who is powerful, get out of their way, and (if you can) help them get what they want</strong>.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Obviously the exact titles depend on your company. One person I’m deliberately leaving out is your own manager. In general don’t think your relationship with your own manager counts as “playing politics”: that’s just you getting along with another human being. An exception to that is if you report directly to a powerful director or VP.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Ironically, this manager struggled to take his own advice.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Note that you actually have to be able to answer their question accurately in order to do this. If you’re not competent enough to be useful to powerful people, you will struggle to befriend them.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>For instance, maybe the CEO is convinced that the project was in bad shape because of something he heard, and the manager in question knows it’s easier to sell “yes, but we turned it around” than “no, you misunderstood, everything was always fine”. If you complicate that process, you risk the CEO thinking that the project is still bad and cancelling it.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"b1c20749f219a90b","title":"Good writing is obvious, not original","link":"https://seangoedecke.com/good-writing-is-obvious-not-original/","author":null,"published_at":"2026-07-14T00:00:00+00:00","content":"<p>When you write, you should try to say things that are obviously true, and spend very little time worrying about whether you’re being original. Ironically, this is the best way to do truly original writing.</p>\n<p>Every important idea has been discussed already. I learned this in grad school for philosophy<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>, but it’s true across the board. Almost every field<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> has generations of very smart people who’ve spent their whole lives thinking about the biggest problems. If you try to only write about ideas that are truly brand-new, <strong>you will restrict yourself to writing about ideas that weren’t important enough to be considered before</strong>: in other words, trivia and ephemera.</p>\n<p>What happens if you focus on the obvious instead? Writing about things that are obviously true is surprisingly difficult. Writing about <em>anything</em> is difficult, because <a href=\"https://johnsalvatier.org/blog/2017/reality-has-a-surprising-amount-of-detail\" rel=\"noopener noreferrer\">reality has a surprising amount of detail</a>. Still, writing about obvious things is especially difficult, because <strong>they’re almost too obvious to see</strong>.</p>\n<p>It’s trivial for me to come up with non-obvious things I believe (for instance, I don’t like <a href=\"https://www.seangoedecke.com/invalid-states/\" rel=\"noopener noreferrer\">database foreign keys</a>), but it’s hard for me to articulate the parts of what I believe that are truly fundamental. I think it’s the same reason why I can’t see the glasses on my face: they’re not something I look <em>at</em>, they’re something I look <em>with</em>. The beliefs you consider the most obvious often play a role like this. Writing about them can be really worthwhile, because reading about your basic assumptions can help other people articulate and examine their own.</p>\n<p>My most popular blog post — <a href=\"https://www.seangoedecke.com/how-to-ship/\" rel=\"noopener noreferrer\"><em>How I ship projects at big tech companies</em></a> — is a nice example. I knew I had strong opinions on how to ship projects (if for no other reason than I kept watching other people doing it wrong), but it took hours of sitting down and writing to realise that I actually just had a different definition of “shipping”<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>, and that all of my object-level advice flowed naturally from that. That definition was so obvious I struggled to notice it at all.</p>\n<p>The other reason to try and be obvious is that <strong>trying to be original is really boring</strong>. Here’s a quote from Keith Johnstone’s <em>Impro</em>, a book I consider <a href=\"https://www.seangoedecke.com/impro/\" rel=\"noopener noreferrer\">far too cult-like</a> but that does have some great advice about originality:</p>\n<blockquote>\n<p>If someone says ‘What’s for supper?’ a bad improviser will desperately try to think up something original. Whatever he says he’ll be too slow. He’ll finally drag up some idea like ‘fried mermaid’. If he’d just said ‘fish’ the audience would have been delighted. No two people are exactly alike, and the more obvious an improviser is, the more himself he appears. If he wants to impress us with his originality, then he’ll search out ideas that are actually commoner and less interesting. I gave up asking London audiences to suggest where scenes should take place. Some idiot would always shout out either ‘Leicester Square public lavatories’ or ‘outside Buckingham Palace’ (never ‘inside Buckingham Palace’). People trying to be original always arrive at the same boring old answers. Ask people to give you an original idea and see the chaos it throws them into. If they said the first thing that came into their head, there’d be no problem.</p>\n</blockquote>\n<p>I think this is absolutely correct. When you’re trying to be original, you’re striking out into nothingness, and you’ll likely end up pattern-matching to “something that sounds kind of transgressive”, which is as boring as it sounds. When you’re trying to say something obvious, you’re reflecting on your own experiences, which have the interesting grit and detail of reality, and are far more likely to sound original to your readers.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>If you think you’re the first to have a philosophical idea, you’re probably not even in the first few hundred to write about it.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Of course, if you’re in a brand new field of study, it’s easy to make a lot of original observations very quickly. I call this the <a href=\"https://www.seangoedecke.com/ai-and-informal-science/\" rel=\"noopener noreferrer\">“gentleman scientist” era</a> of a field. It’s good to try and be original in these limited circumstances, because there’s such a wealth of ideas that nobody has tried yet.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>In brief, I think “shipping” is <em>socially</em> defined: something is shipped when the relevant decision-makers at a company agree that it is shipped.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"d0177051d44283f1","title":"In defense of not understanding your codebase","link":"https://seangoedecke.com/in-defense-of-not-understanding-your-codebase/","author":null,"published_at":"2026-07-11T00:00:00+00:00","content":"<p><strong>As a software engineer, how well do you have to understand your own codebase?</strong></p>\n<p>My guess is that people who work on small codebases with low-turnover teams (say, <a href=\"https://redis.io/\" rel=\"noopener noreferrer\">Redis</a> or games like <a href=\"https://en.wikipedia.org/wiki/The_Witness_(2016_video_game)\" rel=\"noopener noreferrer\">The Witness</a>) would say “obviously you have to understand it completely, otherwise you can’t do good work”. I’d also guess that people who work on large codebases with high-turnover teams (say, the Google web search backend or GitHub) would say “obviously you can’t understand it completely, you just have to do the best you can in your local area”.</p>\n<p>These are two largely different ways of programming with different methods, practices and cultures<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. However, the first group is over-represented in online discussion about software engineering<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. I want to defend the second group against the first. In many software engineering environments, there’s nothing wrong with being in a state of <em>partial</em> understanding. In fact, in large systems a partial understanding is the best you can do.</p>\n<h3>Against “programming as theory building”</h3>\n<p>The best articulation of the “you have to understand your codebase” side is Peter Naur’s famous paper <a href=\"https://pages.cs.wisc.edu/~remzi/Naur.pdf\" rel=\"noopener noreferrer\"><em>Programming as Theory Building</em></a>. I like this paper, but I think it goes too far in that direction. Naur’s core point is that when programmers work on a program, the code is really just a by-product, and the main product they’re working on is their “theory of the program”. That’s made up of their intuitive sense of what’s happening and why, which can only be partially captured by code or documentation. If they lost the code, they could rewrite the program easily. If they lost their understanding (say, if the team experienced 100% turnover), they would struggle to make sense of the code.</p>\n<p>So far, so good, but Naur goes further than this. He says that the theory <em>should not</em> be reconstructed from the code. According to Naur, <strong>you’re better off scrapping the program entirely and having a new team rebuild it from scratch</strong>, building up a new theory in the process<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>:</p>\n<blockquote>\n<p>reestablishing the theory of a program merely from the documentation, is strictly impossible … [therefore] the existing program text should be discarded and the new-formed programmer team should be given the opportunity to solve the given problem afresh</p>\n</blockquote>\n<p>Anyone who’s been an effective software engineer at a large company knows that Naur is dead wrong about this. There are at least two reasons.</p>\n<p>First, <strong>you simply can’t rebuild large software systems from scratch</strong>. Sufficiently large systems (if they have users) contain thousands of <a href=\"https://www.seangoedecke.com/wicked-features/\" rel=\"noopener noreferrer\">weird cases</a> and quirks that cannot be reimplemented. Even a team that’s intimately familiar with the system couldn’t do it: there’s just too much <em>stuff</em> to juggle. Successful rewrites always start by carving out the existing codebase into small isolated chunks, then rewriting one chunk at a time. In other words, rewriting a software system involves making a bunch of changes to the old system. If you can’t change the old system, you certainly can’t replace it with a new one.</p>\n<p>Second, <strong>abandoned systems are revived <em>all the time</em></strong>. In a tech company with hundreds of millions of lines of code and thousands of engineers, it’s not uncommon for a codebase to have nobody left who’s familiar with it<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>. All it takes is a few people to quit at the wrong time, or for a codebase to be unmaintained for a year. Not only have I seen other teams do this, I have <em>personally</em> taken ownership of abandoned codebases, figured them out, and gotten to a point where I could effectively work with them. It takes time, but building a new theory of the codebase is possible. You start by understanding one flow end-to-end, then slowly branch out from there, making careful changes as you go.</p>\n<p>In sufficiently large codebases, <strong>everyone operates with an incorrect theory of the program</strong>. The defining feature of modern software systems is that they’re just way too big for anyone (or even a whole team) to keep in their head: <a href=\"https://www.seangoedecke.com/nobody-knows-how-software-products-work/\" rel=\"noopener noreferrer\">nobody understands it all</a>. To be effective, you have to figure out a way to work with a merely partially-correct theory. This is why I keep going on about <a href=\"https://www.seangoedecke.com/taking-a-position/\" rel=\"noopener noreferrer\">taking a position</a> and <a href=\"https://www.seangoedecke.com/what-makes-strong-engineers-strong/\" rel=\"noopener noreferrer\">confidence</a>. If you’re not sure about something, you can’t just sit back and wait for someone with a perfect understanding to come and give you the answer. If you’re a competent engineer, <em>that person is you</em>. You have to grit your teeth, make your most educated guess, and then deal with the consequences.</p>\n<p>To be generous to Naur, it’s possible that in 1985 the average size of a program was several orders of magnitude smaller than today, and that when Naur writes about “large programs” he’s not talking about tens of millions of lines of code. Naur’s first example of a large program is a 200,000 line industrial monitoring program, and his second example is a compiler. In 1987, the first version of the compiler GCC was about a <a href=\"https://www.oreilly.com/openbook/freedom/ch09.html\" rel=\"noopener noreferrer\">hundred thousand</a> lines of code; in 2015 GCC was over <a href=\"https://www.phoronix.com/news/MTg3OTQ\" rel=\"noopener noreferrer\">fourteen million</a> lines. I can believe that rewriting one or two hundred thousand lines of code is relatively straightforward, particularly if you get to reuse existing tests. Not so for one or two million.</p>\n<h3>Theory building is one tradeoff among many</h3>\n<p>LLMs are <a href=\"https://ratfactor.com/cards/naur-vs-llms\" rel=\"noopener noreferrer\">often cited</a> as a tool that’s bad because it impedes the ordinary process of theory-building. I think this is overly simplistic. Like many software tools, LLMs are a double-edged sword: they make it harder to construct a detailed mental theory of the software, but they allow you to build a partial theory quickly and they can help you leverage that partial theory more effectively. This is a complex tradeoff that I’m still thinking about.</p>\n<p>Setting LLMs aside, I’m confident that it’s silly to say that anything that interferes with your theory of the software must be bad. Here is a partial list of other things that make it harder to maintain a theory:</p>\n<ul>\n<li>Other people being allowed to write code in your codebase</li>\n<li>Having to implement legally-required features like accessibility and data protection</li>\n<li>Allowing your colleagues to quit their jobs or move between teams</li>\n<li>Having to upgrade software versions for security patches</li>\n<li>Bringing in libraries or other dependencies</li>\n</ul>\n<p>Like most things in software, “maintaining a theory of the codebase” is one value among many. Sometimes it’s the most important value and you sacrifice other values for it; other times you trade it off for speed, or legal compliance, or for political reasons<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>.</p>\n<p>Almost all engineers — particularly <a href=\"https://www.seangoedecke.com/pure-and-impure-engineering/\" rel=\"noopener noreferrer\">“pure”</a> engineers — prefer to maintain an accurate mental model of their software. It’s more fun, less stressful, and feels more like “real engineering”. That’s why many engineers take up open-source projects in their spare time in order to work on small codebases by themselves: in order to do engineering work where they can maintain an accurate Naur theory of the codebase. I don’t think there’s anything wrong with that.</p>\n<p>However, at work <a href=\"https://www.seangoedecke.com/where-the-money-comes-from/\" rel=\"noopener noreferrer\">you are paid to do a job</a>. In other words, they pay you money to adopt <em>their</em> set of engineering values. It’s hopefully well-understood that however much you might personally care about performance, sometimes you have to write slow code at your job (for instance, to get a project done on time, or to accommodate some awkward requirement). Maintaining a theory of the codebase is the same kind of thing. </p>\n<div>\n<hr>\n<ol>\n<li>\n<p>I wrote about this at length in <a href=\"https://www.seangoedecke.com/pure-and-impure-engineering/\" rel=\"noopener noreferrer\"><em>Pure and impure software engineering</em></a>. I think many of the repeated arguments we have in the software industry are caused by the pure total-understanding culture coming up against the impure partial-understanding culture.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Open-source engineers are more excited to blog about their work, the raw engineering content is typically more impressive (because coordination problems dominate big proprietary systems), open-source projects can be legally written about while proprietary systems can’t, and even if you could do it legally, writing about large codebases is impossible because it requires too much <a href=\"https://www.seangoedecke.com/you-cant-design-software-you-dont-work-on/\" rel=\"noopener noreferrer\">specific context</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I re-read the relevant chapters of Ryle’s <a href=\"https://www.andrew.cmu.edu/user/kk3n/80-300/ryle1949.pdf\" rel=\"noopener noreferrer\"><em>The Concept of Mind</em></a> (which Naur cites throughout) and I think Ryle is more generous about theory-building. For Ryle, theory-building or know-how automatically happens as you do things. It’s fully consistent with Ryle to think you can pick up an existing codebase just from the code, purely by puzzling it out.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Naur says: “Lest this consequence may seem unreasonable, it may be noted that the need for revival of an entirely dead program probably will rarely arise, since it is hardly conceivable that the revival would be assigned to new programmers without at least some knowledge of the theory had by the original team.”. If only!</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Some engineers might say that maintaining a theory is the <em>core</em> value, because without it you can’t fulfill any of the others. I disagree. You could say the same thing about readability, or maintainability, or correctness, or a bunch of other engineering values. We trade off “core” values like this all the time.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"354975d3d6729bcc","title":"Blog about things you don't understand yet","link":"https://seangoedecke.com/blog-about-things-you-dont-understand-yet/","author":null,"published_at":"2026-07-07T00:00:00+00:00","content":"<p>Every post I publish represents at least two things I’ve learned: the thing that prompted me to write the post, and the thing I learned in the course of writing it. If I don’t learn anything new while I’m writing, it’s not interesting enough to publish.</p>\n<p>Typically I learn way more than two things. For instance, in my <a href=\"https://www.seangoedecke.com/the-o3-geoguessr-prompt-did-not-work/\" rel=\"noopener noreferrer\">o3 geoguessr</a> post, I started out with the idea that most AI prompts probably don’t work, and I ended up learning that newer OpenAI models have lost o3’s ability to geolocate. That’s interesting! In my most recent post on <a href=\"https://www.seangoedecke.com/c2pa-only-works-if-everything-is-signed/\" rel=\"noopener noreferrer\">C2PA</a>, I started out with the idea that C2PA requires near-universal adoption, but I learned a <em>ton</em> of things about PKI, managing private keys on local devices, how C2PA actually works, and so on. In my post on the <a href=\"https://www.seangoedecke.com/luddites-and-ai-datacenters/\" rel=\"noopener noreferrer\">Luddites</a>, I started out with the idea that the Luddite movement was fundamentally decentralized, but ended up fascinated by Luddite culture (which was far more elitist, misogynist, and violent than the pop-Luddism books describe). I could do this for every single post on the blog.</p>\n<h3>Taking a position</h3>\n<p>I think the core reason this works is that <strong>every single one of my blog posts argues a point</strong>. I never publish a post that just gives some scattered thoughts on a topic, or a post that only says “yes, I agree with this other article”. If I write a draft that nobody sensible could disagree with, I scrap the draft. Making sure that everything I write is at least minimally controversial is a forcing function: it forces me to think about what the most interesting part of my position is, and it forces me to do enough research to defend it against the obvious criticisms.</p>\n<p>This is contrary to a lot of advice I read about blogging, which encourages the aspiring blogger to treat their posts as a form of unstructured self-expression. If unstructured self-expression is what you want to do, that’s cool. The point of having a blog is that you get to write what <em>you</em> want. However, this advice isn’t as helpful as it sounds.</p>\n<p>Before I was in tech, I was a philosophy grad student. But before <em>that</em>, I was a poet. One thing you learn when you try to write poetry is that it is way easier to write to a restrictive structure than it is to simply “write what you feel”. This should be obvious when you actually think about it. The task of a poet is to repeatedly choose the next word. Writing to a structure (typically rhyme or meter) narrows that choice to a small set of words, instead of the entire English language. It’s the same with blogging. Forcing yourself to write about specific, potentially-controversial points makes consistently writing easier, not harder.</p>\n<h3>Writing, thinking, and research</h3>\n<p><strong>Writing is the best way to think clearly about a topic.</strong> It’s easy to believe you understand something when you’re just turning it over in your head. When you have to condense that down into words, you find out exactly how much you do or don’t understand. I am constantly having moments where I type something, stop myself, and think “wait, that can’t actually be right”, or “is that really true?”</p>\n<p>By the time I write my way to the end of the post, I’m usually thinking so much more clearly about the topic that my conclusion paragraph is way better than my introduction. In fact, I’ve picked up the habit of going back and immediately rewriting the first paragraph as part of my first-draft process, because I know I’m going to end up doing it anyway.</p>\n<p>I also change my mind a lot while I write. <a href=\"https://www.seangoedecke.com/space-ai-datacenters-do-not-have-a-cooling-problem/\" rel=\"noopener noreferrer\">Here</a> <a href=\"https://github.com/sgoedecke/gatsby-blog/blob/2841c8504fc0b5f4dd2e8105955350bafec4d904/content/drafts/_icebox/prediction-markets-insider-trading/index.md\" rel=\"noopener noreferrer\">are</a> <a href=\"https://www.seangoedecke.com/the-just-say-no-engineer-was-a-zirp-phenomenon/\" rel=\"noopener noreferrer\">a</a> <a href=\"https://www.seangoedecke.com/giving-llms-a-personality/\" rel=\"noopener noreferrer\">bunch</a> <a href=\"https://www.seangoedecke.com/ai-detection/\" rel=\"noopener noreferrer\">of</a> <a href=\"https://www.seangoedecke.com/tempo-faq/\" rel=\"noopener noreferrer\">examples</a> <a href=\"https://www.seangoedecke.com/impact-of-ai-study/\" rel=\"noopener noreferrer\">of</a> <a href=\"https://www.seangoedecke.com/ai-interpretability/\" rel=\"noopener noreferrer\">posts</a> where I began writing them with the opposite opinion to the one that eventually made it into the post. I think this is a good sign, and I hope I never stop doing it. You should be researching and thinking about every post you write, and that means you should frequently learn new things that change your mind.</p>\n<p>Because of all this, I deliberately choose to write blog posts about things I don’t yet quite understand but would like to, like <a href=\"https://www.seangoedecke.com/steering-vectors/\" rel=\"noopener noreferrer\">LLM</a> steering, Stripe’s <a href=\"https://www.seangoedecke.com/tempo-faq/\" rel=\"noopener noreferrer\">Tempo</a> blockchain, <a href=\"https://www.seangoedecke.com/c2pa-only-works-if-everything-is-signed/\" rel=\"noopener noreferrer\">C2PA</a> and <a href=\"https://www.seangoedecke.com/text-ai-watermarks/\" rel=\"noopener noreferrer\">watermarking</a>, space <a href=\"https://www.seangoedecke.com/space-ai-datacenters-do-not-have-a-cooling-problem/\" rel=\"noopener noreferrer\">cooling</a>, <a href=\"https://www.seangoedecke.com/interaction-models/\" rel=\"noopener noreferrer\">interaction models</a>, LLM inference <a href=\"https://www.seangoedecke.com/fast-llm-inference/\" rel=\"noopener noreferrer\">internals</a>, and so on. This is great for me, because I learn a lot. Is it great for my readers?</p>\n<h3>Is blogging to learn irresponsible?</h3>\n<p>I sometimes worry that I should only be writing about areas I already know very well, like <a href=\"https://www.seangoedecke.com/how-to-ship/\" rel=\"noopener noreferrer\">tech company dynamics</a> or <a href=\"https://www.seangoedecke.com/good-api-design/\" rel=\"noopener noreferrer\">working</a> in <a href=\"https://www.seangoedecke.com/large-established-codebases/\" rel=\"noopener noreferrer\">large codebases</a>, rather than presenting myself as an authority on fields I’m actually still learning. Should I let historians of the Luddites write about Luddism, Web3 engineers write about blockchains, and so on? I think this is acceptable for three reasons.</p>\n<p>First, it’s sometimes easier for a beginner to write an introduction to a field than for an expert. Experts routinely <a href=\"https://xkcd.com/2501/\" rel=\"noopener noreferrer\">overestimate</a> the knowledge of the general public, and have often internalized the reasons why their field is important so deeply that they struggle to express them. I think my <a href=\"https://www.seangoedecke.com/tags/explainers/\" rel=\"noopener noreferrer\">explainer posts</a> are valuable because I always spend the first chunk of the post talking about <em>what the original problem is</em> before I get into the technical solution.</p>\n<p>Second, sometimes the public consensus on a topic is just plain wrong, to the point where even a little bit of research is enough to demonstrate why. Many of my posts I’m proudest of have been along these lines: arguing that the “500ml per prompt” water usage figure for LLMs was <a href=\"https://www.seangoedecke.com/water-impact-of-ai/\" rel=\"noopener noreferrer\">ludicrous</a>, or that the popular Apple “Illusion of Thinking” paper was tracking <a href=\"https://www.seangoedecke.com/illusion-of-thinking/\" rel=\"noopener noreferrer\">persistence, not reasoning</a>, that GPUs <a href=\"https://www.seangoedecke.com/ai-gpus-live-longer-than-three-years/\" rel=\"noopener noreferrer\">live longer than three years</a> and the AI companies have large <a href=\"https://www.seangoedecke.com/ai-inference-is-obviously-profitable/\" rel=\"noopener noreferrer\">profit margins</a> on inference, and so on.</p>\n<p>Third, I try to make it clear on my blog who I am and what my credentials actually are. Even if it’s not explicitly described in the post, I have my real name and resume available on my <a href=\"https://www.seangoedecke.com/about\" rel=\"noopener noreferrer\">/about</a> page, so I don’t think a careful reader could be easily fooled into thinking I’m an expert on 19th-century England or space physics or LLM economics or anything like that.</p>\n<h3>Feedback</h3>\n<p>Even if nobody reads what you write, writing is still a good discipline for getting your thoughts in order. But another big reason why writing is a great learning tool is that <strong>you can get feedback</strong>.</p>\n<p>I think it’s obvious why this is useful, but I do want to make two points about feedback. First, if you do make your posts public, you need to have a pretty thick skin. People on the internet often fall over themselves to come up with the most cutting criticism or the harshest dunk. This goes double if you take my previous advice and try to write posts that make a clear, controversial point about a subject you’re learning. If you’re the kind of person whose whole day is ruined when a stranger is cruel to them, you might want to keep your blogging private or only share it among friends.</p>\n<p>Second, even if your blogging is private, <strong>you can get feedback from LLMs</strong>. Like humans, LLMs will often give junk feedback. In my experience, OpenAI models will always tell me to moderate my claims or add caveats and hedges until I’m not saying anything at all. Sometimes their criticism will be straight-up wrong. But — particularly about technical topics — LLMs are great at pointing out areas you’ve genuinely misunderstood, and they’re far kinder than the average Lobsters or Hacker News commenter.</p>\n<h3>Conclusion</h3>\n<p>I’m pleased and grateful that people enjoy reading my posts, but even when nobody did, I still got a lot of value out of blogging. I write as a method of thinking more clearly, as an excuse to do research on topics I want to learn about, and as a way of getting feedback. </p>\n<p>If you’d like to try it yourself, I suggest watching for these two things. First, you should be changing your mind a lot as you write. If not, you probably aren’t doing enough research. Second, your first draft’s conclusion should be much tighter and more expressive than its introduction. If not, you probably haven’t learned anything from the writing process, which means the draft can be scrapped.</p>\n<p>I strongly recommend this practice to anyone with an interest in writing. You will see the benefits even if you don’t publish any of your writing on the internet, particularly now that you can get good technical feedback by pasting your post into a LLM<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>For what it’s worth, I’ve fiddled with careful “review prompts” and it’s basically as good to just write “review, please:” and paste your article.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"b2e5cab6d4e8a675","title":"C2PA only works if everything is signed","link":"https://seangoedecke.com/c2pa-only-works-if-everything-is-signed/","author":null,"published_at":"2026-07-06T00:00:00+00:00","content":"<p>The <a href=\"https://artificialintelligenceact.eu/\" rel=\"noopener noreferrer\">European Union AI Act</a> is Europe’s attempt to comprehensively regulate AI usage. A big part of that is the requirement that AI-generated content be identifiable: either tagged with a watermark or with what the <a href=\"https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content\" rel=\"noopener noreferrer\">Act</a> calls “digitally signed metadata”<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. Since all this becomes enforceable in a month, it’s worth figuring out if it makes any sense. I recently discussed AI watermarking at length in <a href=\"https://www.seangoedecke.com/text-ai-watermarks/\" rel=\"noopener noreferrer\"><em>Text AI watermarks will always be trivial to remove</em></a>. What about digitally signed metadata?</p>\n<p>The most well-known implementation of digitally signed metadata is C2PA Content Credentials, which incorrectly<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> <a href=\"https://c2paviewer.com/articles/eu-ai-act-content-credentials\" rel=\"noopener noreferrer\">claims</a> to be the technology that the AI Act gives as an example of how to do signed metadata properly. The idea here is that <strong>every single image file should contain unspoofable authorship metadata</strong>. Here’s my position on it:</p>\n<ol>\n<li>C2PA broadly makes sense and is a good idea</li>\n<li>It is pointless to use C2PA for AI-generated images only</li>\n<li>It will take many years for C2PA to be adopted across all images</li>\n<li>Because C2PA makes such great safety theater, we’re going to see a lot of hue and cry about it long before it becomes useful</li>\n</ol>\n<p>Lots to unpack. Let’s start by considering images, since that’s the easiest case.</p>\n<h3>How C2PA signing works</h3>\n<p>When an AI tool generates an image, that tool should include a “made by ChatGPT” disclaimer in that image’s metadata. Likewise, when a camera takes a photo, that camera should include a “taken by a camera” disclaimer. C2PA uses two strategies to protect this metadata:</p>\n<ol>\n<li>The metadata must be <em>signed</em> by some trusted private key</li>\n<li>The metadata contains a hash of the file’s contents, so you can’t copy an existing signature onto a new file</li>\n</ol>\n<p>Each physical camera (or phone) has its own private key, for obvious reasons<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. How do we know that those millions of private keys are trusted? Via <a href=\"https://en.wikipedia.org/wiki/Public_key_infrastructure\" rel=\"noopener noreferrer\">PKI</a>, like HTTPS: each camera’s private “certificate” (which contains its public key) is signed by the manufacturer’s well-known private key, so the chain of authenticity can be verified as long as you have (say) Apple’s root <a href=\"https://www.apple.com/certificateauthority/\" rel=\"noopener noreferrer\">public key</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>.</p>\n<p>What happens if you then edit your photo in Photoshop? Photoshop will leave the camera’s metadata untouched, but will layer a “also, Photoshop was used” piece of metadata over the top, signed with Adobe’s private key (well, with the private key associated with your official copy of Photoshop, which is signed by Adobe’s official private key).</p>\n<p>Likewise, if you ask ChatGPT to generate an image for you, ChatGPT will sign its “made by ChatGPT” metadata with OpenAI’s private key. In theory, every single image could contain unforgeable C2PA metadata, allowing software like Twitter to trivially distinguish real photos from fake ones.</p>\n<h3>C2PA needs more regulation to boost adoption</h3>\n<p><strong>Right now, C2PA does not have anything like the adoption it’d need to work.</strong> It’s hard to find hard data on how many images in the wild use C2PA, but FotoForensics <a href=\"https://www.hackerfactor.com/blog/index.php?%2Farchives%2F1010-C2PAs-Butterfly-Effect.html\" rel=\"noopener noreferrer\">reports</a> around a dozen per week (so around 600 out of the <a href=\"https://hackerfactor.com/blog/index.php?%2Farchives%2F1088-Fourteen-and-Video.html\" rel=\"noopener noreferrer\">900,000</a> images processed each year). This is even worse than it sounds, because basically all of the signed images are AI-generated. The adoption rate of C2PA for human-generated images is much, much lower: so far, Google’s Pixel 10 is the only phone camera to sign photos by default. The iPhone <a href=\"https://c2pa.ai/smartphone-guide\" rel=\"noopener noreferrer\">doesn’t sign</a> photos. </p>\n<p>If almost all AI images are C2PA-signed, but almost no human-generated images are, consumers have no reliable way of identifying AI content, because anyone who wants to pretend their AI content is human can simply remove the signature. For C2PA to succeed, it needs to be on every camera and every phone, so that a photo with no signature is rare and suspicious.</p>\n<p>Is that realistic? Actually, I think it is. The appetite (at least in the EU) to regulate AI will increase over time, and while the current EU AI Act only mandates that AI-images are tagged (which by itself is useless), it’s plausible that some future regulation will enforce tagging of all images.</p>\n<p>Another adoption problem that must be solved for C2PA to work is <strong>preservation</strong>. Right now, if you download a C2PA-tagged image, send it as a Facebook message, then re-download it, the C2PA manifest is stripped out. Most images we see on the internet have passed through some social media asset server at least once. All of these social media companies would need to update how they re-encode image content in order to preserve the C2PA data<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>. This would almost certainly require more regulation: C2PA adds tens or <a href=\"https://www.tbray.org/ongoing/When/202x/2024/10/29/Lane-Provenance\" rel=\"noopener noreferrer\">hundreds</a> of kilobytes to each file, which at social media scale is big money<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup>.</p>\n<h3>Forging C2PA signatures</h3>\n<p>Could a clever attacker forge a C2PA signature? Kind of. Neal Krawetz, who seems to have led the anti-C2PA charge, <a href=\"https://www.hackerfactor.com/blog/index.php?/archives/919-Closed-Standards.html\" rel=\"noopener noreferrer\">points out</a> that with a camera development kit it’s straightforward to trick a digital camera into thinking that it’s taking an image when in fact it’s being fed one. This is very much not my area, so please write in if you know more about camera hardware and you think I got this wrong. I suppose you could also take a photo of an AI image on a screen, though I imagine you’d have to be careful to make it look real.</p>\n<p>If you exclude physical attacks on a digital camera, I think C2PA is more robust. You can <a href=\"https://www.hackerfactor.com/blog/index.php?%2Farchives%2F1010-C2PAs-Butterfly-Effect.html\" rel=\"noopener noreferrer\">sign</a> a photo with a self-signed certificate, but the C2PA <a href=\"https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.html#_trust_lists\" rel=\"noopener noreferrer\">spec</a> and <a href=\"https://opensource.contentauthenticity.org/docs/conformance/trust-lists\" rel=\"noopener noreferrer\">docs</a> say that validators must check that your certificate bubbles up to the official C2PA <a href=\"https://spec.c2pa.org/conformance-explorer/\" rel=\"noopener noreferrer\">trust list</a>. This list currently contains only 26 certificates, and there’s a whole process for being added to it. That’ll slow down adoption, but at least it makes it hard to forge<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup>.</p>\n<h3>Other file types and concerns</h3>\n<p>We’ve been talking exclusively about images, but it’s more or less the same story for any type of content. If the file doesn’t support JUMBF metadata (say, an Excel file or a PDF), then the C2PA metadata has to live in a “sidecar”: a separate <code>.c2pa</code> file, probably on some Microsoft or Adobe content server, which contains the signed checksum and the data about who created the file.</p>\n<p>However, the distinction between “real” and AI-generated content is fuzzier when you’re not talking about images. Here’s a trivial example: if I ask ChatGPT to create an Excel spreadsheet for me, the file will be tagged as AI-generated, but I can simply copy/paste the content into a new Excel doc and save it, which will tag it as human-generated<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-8\" rel=\"noopener noreferrer\">8</a></sup>. There’s no software tool that can identify when I’m retyping some AI-generated text (except for perhaps <a href=\"https://www.seangoedecke.com/text-ai-watermarks/\" rel=\"noopener noreferrer\">text fingerprinting</a>, which has its own raft of issues).</p>\n<p>There are also interesting questions around key management. ChatGPT and other AI tools have an easy problem — their users are all online, and so the files can be signed server-side — but how do you sign files created via Photoshop/Excel/Word? If the user doesn’t have internet, do you use some kind of local key? If so, how do you prevent that key being extracted and used to sign AI-generated content? </p>\n<p>Finally, is it a civil liberties problem to automatically fingerprint every photo? Does it make it impossible to be a whistleblower if every photograph can be traced back to your camera? I think this is a complicated question, but in short: I’d expect whistleblowers to already strip EXIF metadata from their images, C2PA metadata is similarly trivial to strip out, and overall I think image attribution is <em>positive</em> for whistleblowers because it heads off “this was AI-generated” responses.</p>\n<h3>Conclusion</h3>\n<p>C2PA is probably here to stay. But it isn’t useful now, and won’t be useful until two huge programs of technical work are completed:</p>\n<ul>\n<li>Every camera manufacturer (including phones) must C2PA-sign all images by default</li>\n<li>Every social media company must retain the C2PA metadata on uploaded images</li>\n</ul>\n<p>This will be a long organizational process, since each manufacturer must go through the approvals process (or decide to start their own competing system), evaluate the legal ramifications of storing attribution data in images, and so on. It will be a long technical process, because C2PA metadata is a substantial fraction of image sizes: storing it will add many petabytes of content.</p>\n<p>Of course, just because C2PA isn’t useful doesn’t mean we’re not all going to do it. Lots of companies are under pressure to signal that they care about AI safety and to head off regulatory attack. “We’re cryptographically signing AI-generated content” is a compelling “we’re doing <em>something</em>” pitch, particularly for people who aren’t technically savvy enough to understand the limitations. In the near term, I expect large AI-involved companies to invest a substantial amount of engineering effort in C2PA-related activity.</p>\n<p>In the long run, once everyone gets on board, I think C2PA could end up working well. It’s awkward in some ways, but “attest content via a PKI certificate chain” is a good idea.</p>\n<p>Is it possible to defeat? Yes, of course. By design, private keys will be in the user’s hands — in their cameras, in their local versions of Photoshop or Microsoft Word, in their phones — so sufficiently technical users will be able to crack them out or use them to sign whatever content they want. I still think C2PA will end up stemming the tide of AI content, because most users are not going to be sophisticated enough to perform attacks like this. However, we should still retain some skepticism of unlikely-looking content, even if it has “created by a human” in its C2PA metadata.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>See sub-measure 1.1.1 of the Act’s associated <a href=\"https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content\" rel=\"noopener noreferrer\">Code of Practice</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>While an <a href=\"https://digital-strategy.ec.europa.eu/en/library/first-draft-code-practice-transparency-ai-generated-content\" rel=\"noopener noreferrer\">early draft</a> of the Code of Practice made an offhand mention of Content Credentials (in the caption of a picture), that was stripped out. The contents of the Act and the final Code of Practice don’t contain “C2PA” or “Content Credentials” (you can search for yourself <a href=\"https://explorer.artificialintelligenceact.eu/en/\" rel=\"noopener noreferrer\">here</a>).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Otherwise if you cracked the key out of one Sony camera, you could spoof content from any Sony camera.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>In practice there are usually more “links in the chain”: a device will be signed by some intermediate certificate, which in turn will be signed by another intermediate certificate, which will be signed by the root certificate. That’s because the root key is so valuable. If an intermediate private key leaks, it can be revoked and replaced (via the root key), but if the root key leaks, it would take <em>years</em> to rebuild the network of trust. So almost all signing is done by intermediates, and the root key stays on a USB drive locked in a safe somewhere.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Not to mention that the whole <em>point</em> of C2PA is that these social media companies will be displaying a “human or AI” sticker in their UI, which will require retaining the metadata.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>C2PA allows for storing the manifest <em>content</em> as a separate <code>.c2pa</code> file, and just including a manifest <em>url</em> in the image metadata itself, but that doesn’t solve the cloud provider problem: they still have to store all the <code>.c2pa</code> files on-disk somewhere.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I think this defuses Neal Krawetz’s <a href=\"https://www.hackerfactor.com/blog/index.php?/archives/1013-C2PAs-Worst-Case-Scenario.html\" rel=\"noopener noreferrer\">“worst-case scenario”</a>. I downloaded his forged image, and (as expected) it gets flagged as “signed, but we don’t trust the root”. I think Krawetz was right at the time, though, since the official “trust list” was only launched in mid-2025.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>You <em>could</em> do the same thing with images by copying into Photoshop or Paint, but while that’d obscure the AI source, it would still be clear that the photo wasn’t taken by a camera.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-8\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"7d3d7b98bf9f92a4","title":"Text AI watermarks will always be trivial to remove","link":"https://seangoedecke.com/text-ai-watermarks/","author":null,"published_at":"2026-07-02T00:00:00+00:00","content":"<p>The European Union <a href=\"https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai\" rel=\"noopener noreferrer\">AI Act</a> will begin to be enforceable in August 2026, one month from now<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. One of the biggest new requirements is <a href=\"https://artificialintelligenceact.eu/article/50/\" rel=\"noopener noreferrer\">Article 50</a>, which requires all AI outputs to be “detectable as artificially generated”. In other words, if LLM providers want to do business in the EU, they will have to apply a watermark to their outputs<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>: some hidden signature that can be used to identify AI content.</p>\n<p>LLM text watermarking is a fascinating problem. Like the best engineering problems, it is theoretically hard to solve perfectly, but has multiple partial solutions: for instance, Google’s <a href=\"https://deepmind.google/models/synthid/\" rel=\"noopener noreferrer\">SynthID</a>, and (as I’ll argue) some quiet Unicode trickery from OpenAI and Anthropic. It will be interesting to see how the AI labs navigate these tradeoffs before the end of the year.</p>\n<h3>Why text watermarking is hard</h3>\n<p>I wrote about AI watermarking at the end of last year in <a href=\"https://www.seangoedecke.com/ai-detection/\" rel=\"noopener noreferrer\"><em>AI detection tools cannot prove that text is AI-generated</em></a>. It’s easy to watermark an image, because digital images contain lots of noise that the human eye can’t really see. For instance, you could apply a watermark like “these twenty pixels in these exact spots will always share a color”. Text is much, much harder. Unlike images, text is a very compressed medium: you cannot make any change to a sentence that a human wouldn’t notice (with one exception, which we’ll get to later). So how are you supposed to watermark it?</p>\n<p>It’s basically a <a href=\"https://arxiv.org/pdf/1302.2718\" rel=\"noopener noreferrer\">text steganography</a> problem (concealing a secret code), made more difficult because the plaintext cannot be arbitrarily manipulated. Any changes you make to apply the watermark will compromise the quality of the output. For instance, “every fifth letter is an ‘e’” would be a good watermark, but applied naively would make the AI output full of typos. Could you just let the model figure out how to fit the watermark? Strong AI models are smart enough to juggle this kind of constraint<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>, but it’d still consume reasoning time that would be better spent on the user’s problem, and make the model sound much less capable than it is<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>.</p>\n<h3>Do we need watermarks to detect AI content?</h3>\n<p>Do you really need a watermark? If you’re Anthropic, and you’re required to be able to verify whether your models produced a particular block of text, can’t you simply run the text through each model, measuring as you go how closely the model’s predicted tokens match each token from the text? </p>\n<p>Not really. The space of “all possible Claude Sonnet answers to a question” is way larger than the space of “all possible <em>watermarked</em> answers to a question”. In other words, you’d get too many false positives for human text that reads like it was AI-written. It’s way more likely for a human to accidentally write like Claude than it is for a human to accidentally reproduce a watermark.</p>\n<p>It would also be prohibitively expensive to run every Anthropic model against a piece of text in order to watermark it. The EU AI Act will eventually require labs like Anthropic to offer free watermarking services to every EU citizen (see Commitment 2). You couldn’t do that with the “run the model” approach.</p>\n<h3>How SynthID works</h3>\n<p>As far as I know, the only AI provider to say they watermark text output is Google, who use a tool called <a href=\"https://www.nature.com/articles/s41586-024-08025-4\" rel=\"noopener noreferrer\">SynthID</a>. Here’s how it works.</p>\n<p>When a LLM generates text, it’s generating a series of tokens (words or chunks of words). At each step, the model itself doesn’t output a single token, but instead outputs a full list of all (say) 100,000 tokens in its vocabulary, each annotated with the probability that that token will be the next one. Tools like ChatGPT or Claude Code will pick semi-randomly from the most likely options in order to get their outputs. <strong>This semi-random sampling process can be influenced in a detectable way.</strong></p>\n<p>For instance, we could choose a sampling strategy like “we pick the second most likely token, then the first, then the second, then the first, and so on”. That would still produce high-quality output, but you’d be able to re-run the model against the generated text to verify that the pattern holds. However, that’d make verification really expensive, and any slight tweaks to the output would break the pattern and thus break the fingerprint. Is there a better way?</p>\n<p>Yes. SynthID is a process for assigning each token a “score” based on its previous tokens (for instance, sum the token’s ID with the IDs of its previous three tokens then take mod 5)<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>. To apply the watermark, the model adopts a sampling strategy like “out of the top five most likely tokens, pick the one with the top SynthID score”<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup>. The watermark can then be detected by calculating the aggregate SynthID score of a block of text. If it’s suspiciously high, it’s very likely to have been AI-generated.</p>\n<p>This is basically a version of the common advice that you can identify LLMs by use of the <a href=\"https://www.seangoedecke.com/em-dashes/\" rel=\"noopener noreferrer\">em-dash</a>, except that instead of a list of keywords, it relies on subtle mathematical relationships between words that humans can’t identify. Because the process for assigning the score is trivial, it’s very cheap to run watermark detection.</p>\n<h3>Unicode watermarks via homoglyphs</h3>\n<p>Google have a complicated mathematical rationale for why SynthID doesn’t make the model dumber: supposedly the SynthID scoring is random enough to act like a normal pseudo-random token sampler, just one that leaves a detectable fingerprint on the outputs. But of course this is suspicious. For instance, it’s common to do inference setting temperature to zero, which always picks the model’s most likely next token. In that case, you can’t leave a fingerprint at all (or you have to ignore the user’s preference and pick the second or third choice anyway).</p>\n<p>If you can’t alter the model outputs, can you still fingerprint the content? Well, kind of. I’m pretty sure OpenAI and Anthropic are sometimes applying fancy Unicode tricks. For instance, you might go through and replace your normal ” ” spaces (unicode <code>U+0020</code>) with a three-per-em ” ” space (unicode <code>U+2004</code>), or a CJK ideographic ”　” space (unicode <code>U+3000</code>). These are called “homoglyphs”, and you can find more of them <a href=\"https://www.irongeek.com/homoglyph-attack-generator.php\" rel=\"noopener noreferrer\">here</a>.</p>\n<p>Of course, lots of human-generated text uses homoglyphs. But it’s trivial to encode a <em>pattern</em> of homoglyphs (say, “every third space becomes a three-per-em”) that is much less likely to occur in the wild. Like the SynthID watermark, a homoglyph-based watermark can be detected very cheaply. A homoglyph-based watermark is cheaper to apply than SynthID: you could even do it entirely on the client.</p>\n<p>I don’t think this is a conspiracy theory. Claude Code was <a href=\"https://thereallo.dev/blog/claude-code-prompt-steganography\" rel=\"noopener noreferrer\">definitely doing this</a> to tag suspicious requests from Chinese users (exploiting homoglyphs for the ’ character in “Today’s date”, though they’ve since walked that back). In the last few years, I’ve noticed that when I copy blocks of text from ChatGPT and paste them into VSCode, sometimes VSCode marks some or all of the spaces as unusual Unicode characters<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup>. Are OpenAI and Anthropic using homoglyphs as an AI-generated watermark? I’m not sure. But they’re definitely using homoglyphs.</p>\n<h3>Text watermarks can be trivially removed</h3>\n<p>The AI Act (specifically, its associated <a href=\"https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content\" rel=\"noopener noreferrer\">Code of Practice</a>) requires watermarking to be “embedded within the content in a manner that is difficult for it to be separated from the content”. However, text watermarks can be trivially removed.</p>\n<p>To remove unicode homoglyph watermarking, you simply have to replace all the homoglyphs with their “real” character equivalents. If you have access to even a relatively weak un-watermarked LLM<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-8\" rel=\"noopener noreferrer\">8</a></sup>, you can strip out SynthID watermarking by asking that LLM to paraphrase the text content. Because the watermark is inherent to subtle vocabulary choices, re-wording the content will remove the watermark. You could even do it by hand, although at that point it’s not really AI-generated content anymore. Since there will be some kind of free public watermark testing tool, you can just keep tweaking until it comes back negative.</p>\n<p>Moreover, the AI Act requires watermarking techniques to be “interoperable… as far as this is technically feasible”. That means AI providers would have to publish their watermarking process, and potentially even attempt to standardize on applying the same kind of watermarks. I just don’t see how this is compatible with the kind of security-by-obscurity that LLM text watermarking depends on. Unlike image and video watermarks, text watermarks will always be trivial to remove.</p>\n<h3>What about C2PA?</h3>\n<p>The AI Act and Code of Practice talk a lot about “digitally signed metadata”. The idea here is that you can include an AI disclosure in the file’s metadata itself, ideally in a way that cannot be tampered with (for instance, by signing a hash of the file’s contents). This signed-metadata process is basically <a href=\"https://c2paviewer.com/articles/eu-ai-act-content-credentials\" rel=\"noopener noreferrer\">C2PA Content Credentials</a>. While you can remove C2PA metadata, you (theoretically) can’t <em>fake</em> it, so a file with “created by a human” metadata can be trusted, and files with no metadata at all can be held in suspicion.</p>\n<p>This post is already too long to get into what I think about C2PA, but I do want to say that <strong>C2PA is not a substitute for text watermarking</strong>. It only really applies to <em>files</em>. In the words of the Code of Practice, that’s “a data format that supports attaching metadata (e.g., an audio, image, video, or containerised text)“. The output of chat tools (and most of the output of AI agents) is not containerized text, but plain old regular text, and so can’t be signed. What would it even look like to sign ChatGPT outputs? There’s no artifact to pass around.</p>\n<p>I think it’s a fascinating question whether Claude Code has to C2PA-sign any HTML files or PDFs it generates for you. That seems kind of tricky to get right. But in any case, the AI Act also mandates some kind of actual watermarking as well.</p>\n<h3>Conclusion</h3>\n<p>So what’s going to happen this year? If I had to guess, I’d say that each AI provider (not just labs like OpenAI or Anthropic, but third-party providers like Fireworks or Groq) will stick a SynthID token sampler in front of their inference stacks. This might be limited to users in the EU, but it might not be, since SynthID is at least as good as a normal top-k token sampling approach.</p>\n<p>AI providers will then offer a “check for watermark” page that re-tokenizes user-provided text, runs the scoring, and checks whether it’s above a certain threshold. Depending on how seriously the interoperability clause is taken, providers might even standardize on the same SynthID setup, in which case there could be a single EU-hosted “watermark this text” page.</p>\n<p>I don’t think unicode-based watermarking is going to be considered compliant with the AI Act, but some providers which don’t want to set up SynthID might try it. Either way, technical users will be able to strip out the watermark at will, and there will be a plethora of tools that non-technical users will use for this purpose.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Well, for new systems; existing ones get until December.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I don’t think the plain text of Article 50 requires this, but <a href=\"https://artificialintelligenceact.eu/recital/133/\" rel=\"noopener noreferrer\">Recital 133</a> and the Code of Practice makes it pretty clear that they’re looking for watermarks.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Even with extra high thinking, GPT-5.5 could not explain SynthID to me with every fifth letter being an “e”, but GPT-5.5-Pro produced this puzzling koan: “These hidden codes label model-made image, voice, movie, prose. Probe trace: maybe a model-made piece. Maybe erase trace; maybe leave trace. Hence trace alone? No.”</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I leave the analogy with AI safety guardrails as an exercise for the reader.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>That’s a toy example. In practice there are multiple different (but still mathematically simple) scoring methods that get combined together, including a random seed. Why include the seed? Otherwise the watermark would bias towards the same set of tokens.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>The tokens are scored in a multi-round knockout against each other, but I think that’s more of an implementation detail and not required to get the core intuition behind why SynthID works.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>When this became <a href=\"https://www.rumidocs.com/newsroom/new-chatgpt-models-seem-to-leave-watermarks-on-text\" rel=\"noopener noreferrer\">public knowledge</a>, OpenAI claimed it was just a model quirk, which is certainly possible.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>All AI providers might be legally required to watermark, but even tiny local models are good enough to paraphrase text.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-8\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"7949fda4481ba2ad","title":"Saying the obvious thing","link":"https://seangoedecke.com/saying-the-obvious-thing/","author":null,"published_at":"2026-06-27T00:00:00+00:00","content":"<p>Stating the obvious is <a href=\"https://blog.jim-nielsen.com/2026/blogging-stating-the-obvious/\" rel=\"noopener noreferrer\">surprisingly useful</a>. Most of your knowledge lives below the threshold of conscious awareness, so it’s possible for a piece of writing to remind you of what you already know. It’s common to know you don’t like something without being quite sure why, and reading an obvious statement (such as “accuracy <a href=\"https://www.astralcodexten.com/p/if-its-worth-your-time-to-lie-its\" rel=\"noopener noreferrer\">matters</a>, even when you agree with the broad strokes”) can help clarify why you find certain things distasteful.</p>\n<p>Sometimes you can see some obvious truth that nobody seems to be talking about, and reading it in someone else’s words can prompt an “oh god, I’m not crazy” moment of catharsis. For many junior engineers, it’s almost a rite of passage to notice that some percentage of software engineers <a href=\"https://x.com/yegordb/status/1859290734257635439\" rel=\"noopener noreferrer\">do virtually no work</a>. Since nobody talks about it (how would you even bring it up in the workplace?), they often feel like they’re losing their minds: <em>surely</em> this state of affairs wouldn’t be allowed to continue, so they must be completely misreading the situation. But in fact it’s true.</p>\n<p>Stating the obvious is hard. It can even be dangerous: sometimes there’s a good reason nobody says the obvious thing. But I think the bigger reason it’s hard is for the same reason that it’s hard to <a href=\"https://drawingacademy.com/drawing-what-you-see-vs-drawing-what-you-know\" rel=\"noopener noreferrer\">draw what you actually see</a>. When I look at a person and try to draw them, I’m not drawing the lines and shades my eye sees (like a printer or camera might). I’m drawing <em>what I know the person looks like</em>, which is a kind of stick-figure approximation. It takes time and effort to drop the layer of interpretation and draw what’s actually there<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>.</p>\n<p>Many of the posts I’m most proud of are times when I’ve managed to articulate something I think is obviously true: <a href=\"https://www.seangoedecke.com/ratchet-effects/\" rel=\"noopener noreferrer\">engineer reputation is determined by ratchet effects</a>, <a href=\"https://www.seangoedecke.com/being-right-a-lot/\" rel=\"noopener noreferrer\">good engineers are right most of the time</a>, <a href=\"https://www.seangoedecke.com/party-tricks/\" rel=\"noopener noreferrer\">you shouldn’t just do JIRA tickets</a> (or <a href=\"https://www.seangoedecke.com/glue-work-considered-harmful/\" rel=\"noopener noreferrer\">glue work</a>), and so on. These are all things I’ve believed for a while, but have only (relatively) recently been able to <em>notice</em> that I believe them. Sometimes I’m helped along by reading something I vehemently disagree with (like “nobody gets promoted for doing <a href=\"https://www.seangoedecke.com/simple-work-gets-rewarded/\" rel=\"noopener noreferrer\">simple work</a>”, or ”<a href=\"https://www.seangoedecke.com/big-tech-needs-big-egos/\" rel=\"noopener noreferrer\">big egos</a> have no place in tech”).</p>\n<p>Stating the obvious doesn’t mean avoiding nuance. Every obvious claim carries with it a host of subtle, non-obvious claims. For example, I believe that having a big ego can be very useful as a software engineer. But why exactly is that, and what do I mean by ego? Obviously it’s not good to be constantly flexing your status on other people, or to be unable to tolerate the possibility of being wrong. However, I do think you need to be able to take <a href=\"https://www.seangoedecke.com/taking-a-position/\" rel=\"noopener noreferrer\">firm technical positions</a> even when the situation is uncertain, which means you have to be confident in your technical instincts. Teasing out that distinction (and its implications) is very interesting, but in order to do it you need to be able to first articulate the obvious part.</p>\n<p>I’ve been talking about stating the obvious in technical blogging. But this principle applies just as well to other kinds of communication. When I write a technical design document at work, it’s very important to state the obvious. In fact, technical communication is <a href=\"https://www.seangoedecke.com/technical-communication/\" rel=\"noopener noreferrer\">so hard</a> and general understanding is <a href=\"https://www.seangoedecke.com/nobody-knows-how-software-products-work/\" rel=\"noopener noreferrer\">so poor</a> that just getting people aligned on the obvious things is often <em>enormously</em> valuable. Much great literature and poetry aims to bring out some obvious but hard-to-articulate part of human experience.</p>\n<p>Don’t avoid writing something down just because you think it’s obvious. The thing you think is obvious now might recede into your subconscious in an hour; get it written down while you can! And don’t avoid writing something down because you think it’s dangerous to say and everyone already knows it. For people new to the area, reading your words can help them feel like they’re not losing their minds. Finally, once you write down the obvious thing, it allows you to go on and draw out the parts that are less obvious, in a way that you couldn’t do if you try to just skip straight to the subtleties.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Incidentally, this is why most people <a href=\"https://www.reddit.com/r/pics/comments/4ew345/as_it_turns_out_most_people_cannot_draw_a_bike/\" rel=\"noopener noreferrer\">cannot draw a bicycle</a> on their first attempt. Unless you’re a mechanical engineer, you probably do not have a stick-figure-level approximation of what a bicycle looks like in your head, so you begin confidently (after all, you’ve seen a thousand bicycles) and get stuck after the first few lines.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"cacd093fc7af78be","title":"AI inference is obviously profitable","link":"https://seangoedecke.com/ai-inference-is-obviously-profitable/","author":null,"published_at":"2026-06-26T00:00:00+00:00","content":"<p>Many people <a href=\"https://www.wheresyoured.at/why-everybody-is-losing-money-on-ai/\" rel=\"noopener noreferrer\">claim</a> that AI inference is unprofitable to serve, and thus must be subsidized by an ocean of dumb money from investors who believe that some future AI model will come to dominate the world economy. When that dumb money goes away, so will AI products. According to this view, LLMs are just inherently too expensive (in terms of money, power, and water) to be used in consumer products. In fact, they can only be used today by externalizing the costs: money onto VC funds and now retail ETF <a href=\"https://www.investopedia.com/spacex-stock-joins-major-index-funds-what-regular-investors-need-to-know-spcx-ipo-vanguard-blackrock-vti-itot-12004764\" rel=\"noopener noreferrer\">investors</a>, power onto electric utility <a href=\"https://salatainstitute.harvard.edu/how-you-subsidize-big-tech-with-your-electricity-bill/\" rel=\"noopener noreferrer\">consumers</a>, and water onto the <a href=\"https://theconversation.com/5-ways-data-centers-endanger-their-local-communities-and-the-country-as-a-whole-282348\" rel=\"noopener noreferrer\">communities</a> where datacenters are built.</p>\n<p>There are <a href=\"https://www.seangoedecke.com/is-ai-wrong/\" rel=\"noopener noreferrer\">good reasons</a> to dislike AI, but this really isn’t one of them. In fact, <strong>AI inference is obviously profitable</strong>.</p>\n<h3>Doing the math demonstrates that inference is profitable</h3>\n<p>Frontier AI providers are reporting 70%-80% <a href=\"https://www.morningstar.com/stocks/anthropics-gross-margin-is-most-important-number-tech\" rel=\"noopener noreferrer\">gross</a> <a href=\"https://www.saastr.com/have-ai-gross-margins-really-turned-the-corner-the-real-math-behind-openais-70-compute-margin-and-why-b2b-startups-are-still-running-on-a-treadmill/\" rel=\"noopener noreferrer\">margins</a> on inference, but maybe we can’t trust them. Let’s do some very rough estimates on the actual cost.</p>\n<p>A Nvidia A100 consumes 400W of power under full load. In practice, even a carefully-tuned inference server will not be at full load all the time, but it’s at least an upper bound. Suppose you’re running a dense 70B model<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>, which will <a href=\"https://dlewis.io/evaluating-llama-33-70b-inference-h100-a100/\" rel=\"noopener noreferrer\">fit</a> comfortably (unquantized) on four A100s at around 2M tokens per hour. At industrial power prices, that’s about 13c/hr in the <a href=\"https://www.eia.gov/electricity/monthly/update/end-use.php\" rel=\"noopener noreferrer\">USA</a>. Suppose (pessimistically) cooling is the same cost. That’s about 13 cents per million output tokens<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>.</p>\n<p>Let’s amortize the cost of the GPUs, since that’s going to be the most expensive part. An A100 costs about $20k. If each A100 lasts around five years<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>, you’ll have to make 16k/yr in profit to recoup your capital investment (or $1.80 per hour). At lower utilization, it’ll take longer to recoup, but your GPUs will also last longer. Either way, your overall inference costs are at about one dollar per million tokens.</p>\n<p>GPT-5.4-mini <a href=\"https://openai.com/business/pricing/#api\" rel=\"noopener noreferrer\">charges</a> $4.50 per million tokens, and stronger OpenAI or <a href=\"https://platform.claude.com/docs/en/about-claude/pricing\" rel=\"noopener noreferrer\">Anthropic</a> models are three to six times as expensive. It’s hard to make a direct comparison because we don’t know the size of OpenAI or Anthropic models, but the claimed 70% or 80% profit margin is extremely plausible.</p>\n<h3>Open LLMs demonstrate that inference is profitable</h3>\n<p>What if you don’t trust my estimates either? Let’s look at the pricing of open-weights Chinese LLMs. DeepSeek have <a href=\"https://github.com/deepseek-ai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_deepseekV3R1_inference_system_overview.md\" rel=\"noopener noreferrer\">claimed</a> a bit over 80% profit margin on inference for DeepSeek-R1. Since their API pricing for R1 is less than half that of OpenAI or Anthropic<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>, that suggests that my estimates above for inference cost might be too expensive. Cooling at scale is probably <a href=\"https://massedcompute.com/faq-answers/?question=What%20are%20the%20estimated%20annual%20power%20consumption%20costs%20of%20NVIDIA%20A100%20and%20H100%20GPUs%20in%20a%20typical%20data%20center?\" rel=\"noopener noreferrer\">cheaper</a> than power, R1 only has half the active parameters of a dense 70B model, modern GPUs are more efficient than the A100, and there are significant <a href=\"https://www.seangoedecke.com/inference-batching-and-deepseek/\" rel=\"noopener noreferrer\">economies of scale</a> in inference.</p>\n<p>Since DeepSeek’s models are available for anyone to download, they can’t get away with extracting a large profit margin. One of the other inference providers would undercut them with the same model. Inference costs for DeepSeek-V4-Pro on the market are around 87 cents per million output tokens, which is probably pretty close to the actual cost of serving the model.</p>\n<h3>For AI labs, inference must subsidize training</h3>\n<p>All of this doesn’t mean that <em>OpenAI</em> or <em>Anthropic</em> are profitable. Those companies are making huge capital <a href=\"https://openai.com/index/building-the-compute-infrastructure-for-the-intelligence-age/\" rel=\"noopener noreferrer\">investments</a> that may or may not pan out, and are spending enormous amounts of money on talent and compute to train brand-new models and retain users.</p>\n<p>They’re doing crazy things like offering per-month subscription models for nearly unlimited inference, which is almost certainly not profitable. If you used an API token instead of your Anthropic subscription in Claude Code, you’d pay ten times the cost. But that doesn’t mean API-based Claude Code couldn’t be a good deal. Some people are <a href=\"https://www.reddit.com/r/opencodeCLI/comments/1tril88/test_of_prices_of_deepseek_in_opencode_go_and_api/\" rel=\"noopener noreferrer\">already using</a> DeepSeek’s inference API for agentic coding, because once you take away the huge profit margin it’s cheaper than the relative per-month subscription.</p>\n<p>Why won’t OpenAI or Anthropic lower their prices? Supposedly OpenAI has <a href=\"https://www.wsj.com/tech/ai/openai-considers-drastic-price-cuts-anticipating-war-for-users-with-anthropic-9b8c178e\" rel=\"noopener noreferrer\">thought about it</a>, but for an AI lab, <strong>inference has to subsidize training costs</strong>. A company like OpenAI has to fund the production of new models from the inference margins on existing models (at least partially). That’s why the margins on inference are so high: the AI labs are trying to squeeze out every dollar so they can stay alive in the training arms race.</p>\n<p>However, inference only has to subsidize training costs <strong>for an AI lab</strong>. If you’re merely an inference provider, you don’t have to do any training at all. Therefore, even if OpenAI and Anthropic go out of business, whoever snaps up the rights to their frontier models will be able to continue selling Opus and GPT inference at a profit<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>. The AI bubble popping will not mean the end of the inference business, because <strong>AI inference is obviously profitable</strong>.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Expensive frontier models are probably mixture-of-experts, not dense, which is tougher to estimate. However, I think a 70B dense model and a MoE with 70B active params will come out to basically the same numbers at scale (though the MoE will require more GPU memory and thus a greater upfront cost). Are frontier models around 70B params? Nobody outside the AI labs really knows, but my guess is that 70B is probably larger than a Haiku/mini class model.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I think it’s reasonable to estimate the cost of output tokens only, since they’re by far the most expensive part of serving inference. Input tokens are cheaper for two reasons: transformers let you prefill them in parallel, and for most real-world use cases they can be aggressively cached in the KV cache.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>It’s common (and wrong) to estimate GPU lifespan at three years. I wrote a lot about this in <a href=\"https://www.seangoedecke.com/ai-gpus-live-longer-than-three-years/\" rel=\"noopener noreferrer\"><em>AI GPUs probably live longer than three years</em></a>. </p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Again, this is just an guess, since we don’t know what OpenAI or Anthropic model is equivalent in size to R1.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I do wonder if Anthropic would be able to prevent other people from being able to access the model if the company goes out of business. Anthropic is currently in <a href=\"https://www.bloomberg.com/news/articles/2026-06-02/broadcom-backing-lowers-debt-costs-on-36-billion-anthropic-deal\" rel=\"noopener noreferrer\">debt</a> to Broadcom, Google, and a bunch of private equity firms. Would they get the Mythos and Opus weights, over Dario’s protestations? </p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"9dd2ac56ff1f486c","title":"AI GPUs probably live longer than three years","link":"https://seangoedecke.com/ai-gpus-live-longer-than-three-years/","author":null,"published_at":"2026-06-15T00:00:00+00:00","content":"<p><a href=\"https://www.wheresyoured.at/ai-is-slowing-down\" rel=\"noopener noreferrer\">People</a> who think current AI use is unsustainable often rely on the <a href=\"https://www.tomshardware.com/pc-components/gpus/datacenter-gpu-service-life-can-be-surprisingly-short-only-one-to-three-years-is-expected-according-to-unnamed-google-architect\" rel=\"noopener noreferrer\">claim</a> that inference GPUs only last “three years at the most” under load<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. The idea here is that once the AI bubble money drains away, current infrastructure will rapidly become obsolete, and there won’t be enough money floating around to buy a whole slate of brand-new GPUs. Inference costs would thus rapidly become way too expensive for current AI products to make any financial sense.</p>\n<p>Where does this “three years at the most” claim come from? Is it plausible? </p>\n<h3>Sourcing the quote</h3>\n<p>The original Tom’s Hardware article quotes this <a href=\"https://x.com/techfund1/status/1849031571421983140\" rel=\"noopener noreferrer\">tweet</a> from Tech Fund, an anonymous former PM and tech investor, who quotes an anonymous “GenAI principal architect” at Google as saying “if you have a high utilization rate, then constant high utilization rate for a year or two, I think the lifespan will be three years at most”.</p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/9cf90a11e418ecd73f81e5178a9328c2/1b853/tweet.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"tweet\" src=\"https://www.seangoedecke.com/static/9cf90a11e418ecd73f81e5178a9328c2/fcda8/tweet.png\" title=\"tweet\">\n  </a>\n    </span></p>\n<p>This screenshot looks like it was from an interview. What interview? I scrolled back to October 2024 on Tech Fund’s Twitter feed and saw a bunch of <a href=\"https://x.com/techfund1/status/1828858794480140391?s=20\" rel=\"noopener noreferrer\">similarly-formatted</a> <a href=\"https://x.com/techfund1/status/1826875528751534448?s=20\" rel=\"noopener noreferrer\">screenshots</a>, some of which were cited as coming from <a href=\"https://tegus.com/\" rel=\"noopener noreferrer\">Tegus</a>. Tegus is apparently a company with a <a href=\"https://www.reddit.com/r/expertnetworks/comments/1ghe2ls/tegus_analyst_reached_out_to_me/\" rel=\"noopener noreferrer\">business model</a> of reaching out to insiders (in this case, AI company employees) and paying them hundreds of dollars an hour in order to answer specific technical questions. It’s essentially gig work for <em>almost-but-not-quite</em> insider trading: the more informed and confident you sound, the more likely Tegus analysts will pick you for future interviews.</p>\n<p>I’m sure the source for this tweet is in fact a GenAI principal architect, since Tegus would have presumably asked for some proof of that before they paid them out. But it’s pretty clear that the incentives here are to sound confident and authoritative, even on questions that you’re not sure about. With that in mind, the quote itself also reads a bit suspiciously. I’ve worked with enough principal engineers and architects to take their casual back-of-envelope estimates with a grain of salt. If they knew the actual rate at which GPUs fail and get retired in Google datacenters, wouldn’t they have just said that?</p>\n<h3>Evidence for a longer lifespan</h3>\n<p>We have some anecdotal evidence that points the other way. Google has <a href=\"https://www.datacenterdynamics.com/en/news/google-says-tpu-demand-is-outstripping-supply-claims-8yr-old-hardware-iterations-have-100-utilization\" rel=\"noopener noreferrer\">publicly claimed</a> to have eight year old TPUs (their version of GPUs) running in production at “100% utilization”. Nvidia only made A100 GPUs from <a href=\"https://www.amax.com/nvidia-h100-vs-nvidia-a100/\" rel=\"noopener noreferrer\">2020-2024</a>, but in February 2026 the AWS CEO <a href=\"https://www.datacenterdynamics.com/en/news/aws-has-never-retired-an-nvidia-a100-server-ceo-matt-garman-claims/\" rel=\"noopener noreferrer\">claimed</a> that AWS had never retired an A100 server (and you can still easily rent A100s for AI work)<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. AI GPU usage isn’t exactly like crypto mining GPU usage, but it certainly seems like years-old ex-crypto GPUs are <a href=\"https://www.youtube.com/watch?v=UFytB3bb1P8\" rel=\"noopener noreferrer\">functional</a>. There’s also <a href=\"https://news.ycombinator.com/item?id=48456717\" rel=\"noopener noreferrer\">this comment</a> from Hacker News I noticed where someone claims that their GPU cluster in academia has lasted six years with less than 20% failure rate.</p>\n<p>What about hard data? It’s hard to get concrete data on the lifespan of AI GPUs, because modern AI datacenters have only existed for a handful of years. But an interesting case study would be recent supercomputer clusters like Oak Ridge’s <a href=\"https://www.datacenterdynamics.com/en/news/oak-ridge-national-laboratory-to-retire-summit-supercomputer-in-november-2024/\" rel=\"noopener noreferrer\">Summit</a>, which had over 27 thousand Nvidia V100s running from 2018 to 2024, or its predecessor, the Cray <a href=\"https://christian-engelmann.de/publications/ostrouchov20gpu.pdf\" rel=\"noopener noreferrer\">Titan</a> supercomputer that ran from 2012 to 2019. I couldn’t find any evidence that Summit had to buy an additional 27,000 GPUs to replace their old ones, and GPU failures in Titan have been <a href=\"https://christian-engelmann.de/publications/ostrouchov20gpu.pdf\" rel=\"noopener noreferrer\">carefully studied</a>:</p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/3922b6bf4aaf6d3202cacd2f472034aa/019a6/figure10.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"fig10\" src=\"https://www.seangoedecke.com/static/3922b6bf4aaf6d3202cacd2f472034aa/fcda8/figure10.png\" title=\"fig10\">\n  </a>\n    </span></p>\n<p>These cages of GPUs are stacked vertically, and cold air is pumped in from the bottom, which explains why cage 0 (at the bottom) has better survival rates than cage 2 (at the top). Let’s consider cage 0, so we’re just looking at the GPU lifespan instead of at the lifespan of improperly-cooled GPUs. At three years, over 95% of GPUs survived<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. At six years, nodes 2 and 3 (the GPUs closest to the bottom of the cage) were still at above 90% survival rate, and the highest nodes were over 60%.</p>\n<p>It’s possible that newer Nvidia GPUs are less reliable than older ones (they certainly draw more power), or that AI datacenters are under-cooled, or that something about LLM utilization is more stressful than the workloads that ran on traditional GPU datacenters. But this is at least circumstantial evidence that GPUs can survive under load for far longer than three years.</p>\n<h3>Economic lifespans</h3>\n<p>This discussion is complicated by the fact that GPUs may have a short <em>economic</em> lifespan. Supposedly a B100 GPU <a href=\"https://bizon-tech.com/blog/nvidia-b200-b100-h200-h100-a100-comparison?srsltid=AfmBOoqugX-R8Y9AoVlyxRMheglf4gJ2Xc5hefXVxL6Cv3Htl0P_rHx1\" rel=\"noopener noreferrer\">draws</a> twice as much power as an A100, but can do five times as much work. For some AI providers, that might mean that A100s are only worth running until they can be replaced with B100s (if you’re bottlenecked on electricity, you should spend it all on B100s and throw out your obsolete A100s). This is why the Titan supercomputer was decommissioned in favor of Summit: it could have continued to operate, but it was more profitable to spend the money and maintenance effort on newer hardware.</p>\n<p>It should be obvious that this doesn’t support the “inference will become more expensive when the bubble pops” argument. So long as A100s are profitable <em>right now</em>, cash-poor AI providers can continue profitably serving inference from them, even if there are more efficient options available for those with the capital to upgrade.</p>\n<p>On top of that, GPUs only represent one part of AI datacenter infrastructure spending. If your GPUs wear out, you don’t have to go and build an entirely new datacenter. About 30-50% of <a href=\"https://epoch.ai/assets/images/data-insights/ai-datacenter-cost-breakdown/ai-datacenter-cost-breakdown-upfront.png\" rel=\"noopener noreferrer\">datacenter</a> <a href=\"https://www.reuters.com/commentary/breakingviews/how-big-techs-630-bln-ai-splurge-will-fall-short-2026-03-26/\" rel=\"noopener noreferrer\">spend</a> goes to land, power, cooling, and so on. The remaining 50-70% is the cost of the entire server rack, which includes a bunch of things that aren’t GPUs.</p>\n<h3>Conclusion</h3>\n<p>Like the idea that AI inference <a href=\"https://www.seangoedecke.com/water-impact-of-ai/\" rel=\"noopener noreferrer\">requires using huge amounts of water</a>, the idea that AI GPUs only live a year or two is popular because it’s a useful idea for AI skeptics, not because it’s true. It comes from a pseudonymous tweet quoting an anonymous source who’s being paid hundreds of dollars to sound like a credible expert on AI. Other public communications from AI inference providers cite much higher lifespan numbers, and the statistics from supercomputers (the traditional examples of large GPU clusters) don’t bear out the claim that the maximum lifespan is three years.</p>\n<p>It might be true that the <em>economic</em> lifespan is three years, in a world where new GPUs come out every eighteen months and GPU providers are flush with cash to upgrade, but that doesn’t tell us much about the economics of inference in an AI winter. If money becomes a lot more scarce, it’s likely that AI datacenters will continue profitably<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup> running their B300s (or their H100s or even A100s) for six years or longer.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Of course, like previous claims about AI and water usage, “three years at the most” is often cited as <a href=\"https://ithy.com/article/data-center-gpu-lifespan-explained-7mpjwwyp\" rel=\"noopener noreferrer\">“1-2 years, with some lasting up to 3 years under optimal conditions”</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Of course, pronouncements from CEOs/CTOs should be taken with a grain of salt as well (for instance, maybe they have a big backlog of unused A100s they keep swapping out), but (a) executives don’t often straight-up lie about concrete technical facts, and (b) they’re going up against an unsourced quote from a tweet, so the bar isn’t that high.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>What about proactive GPU replacement? In the “Survival Analysis” section, the study attempts to account for this. I haven’t dug into exactly how.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Assuming inference is profitable, which I believe (when you’re not attempting to amortize the cost of training).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"466e9805fc0e757d","title":"Working with product managers","link":"https://seangoedecke.com/working-with-product-managers/","author":null,"published_at":"2026-06-08T00:00:00+00:00","content":"<p>The relationship engineers have with product management is more dysfunctional than with any other part of the company. There’s no shared culture or language like there is with other engineers, and the rules of “who gets to tell who what to do” aren’t as clear-cut as they are with managers. Engineers don’t have a lot in common with legal, or design, or sales, but they also don’t need to interact much with those roles. In my experience, engineers are communicating with product managers almost every single day.</p>\n<h3>Against the “product mommy”</h3>\n<p>The worst version of the product/engineering relationship goes something like this:</p>\n<p>Engineers are technically competent but are too autistic to be fully trusted. They need a kind-but-stern parental figure who knows how to communicate to other stakeholders in the organization (for instance, by being comfortable using the word “stakeholders”), and how to keep engineers from going off in the wrong direction.</p>\n<p>This entire gross dynamic is neatly captured by the popular term <a href=\"https://x.com/search?q=%22product%20mommy%22&amp;src=typed_query\" rel=\"noopener noreferrer\">“product mommy”</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>.\nI really, really don’t like that term, or this entire dynamic in general. Almost none of my relationships with my product managers have been anything like this, though I have seen it at a distance.</p>\n<p>Working well with product managers can be the difference between succeeding and failing at a company. Why is it so hard to maintain good relationships between engineering and product? What does a good relationship look like?</p>\n<h3>Why it’s so hard to build trust</h3>\n<p>Product managers and engineers have largely non-overlapping skillsets. Product managers don’t understand the technical work engineers do and aren’t equipped to talk about it: if an engineer gives a technical reason for something, product managers generally have to shrug and say “sure, I guess”. Likewise, engineers don’t have anything like the visibility into the organization that product managers do. Particularly in large organizations, it is the product manager who is the source of truth about who wants what and which features are important. When a product manager says that something is critical, engineers generally have to shrug and say “sure, I guess”.</p>\n<p>This obviously requires a lot of trust. What’s a little less obvious is that <strong>this trust is continually broken by both sides</strong>. Every single product manager has been told <em>thousands</em> of times that technical task X is technically impossible or would be disastrous, only for that task to end up being done fairly smoothly and successfully. Every single engineer has been told <em>thousands</em> of times that requirement X is absolutely critical and worth going to enormous effort for, only for that requirement to be silently dropped or changed with no apology.</p>\n<p>Of course this isn’t malicious. Engineers often give wrong estimates because <a href=\"https://www.seangoedecke.com/how-i-estimate-work/\" rel=\"noopener noreferrer\">estimation is impossible</a>, and sometimes the dire consequences they warn about really do happen (they’re just handled behind the scenes, like engineers handle many other kinds of technical dysfunction). Product managers “change their minds” because what’s important in a large tech company does genuinely change hour-by-hour<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>, and even the best attempts to only filter the most reliable priorities through to the engineering team will sometimes go wrong.</p>\n<h3>Manipulation and lies</h3>\n<p>The consequence of this broken trust is that the relationship becomes very difficult to maintain. When you’re an engineer, and you explain something to your product manager, and you <em>know</em> they don’t believe you (despite having no ability themselves to judge the question), it can be incredibly frustrating. Likewise, when you’re a product manager, and you’re desperately trying to explain what we need to do to an engineer, and you know they’re internally shrugging their shoulders, it must be unbearable. Don’t they know this is critical to the company? You were just in a meeting with the leaders of the organization!</p>\n<p>The natural tool for a mistrustful product manager is <em>manipulation</em>. I still remember a product manager who tried to extract a commitment from my team by asking us to go around and all say “I commit to getting this work done in two weeks”, after a conversation where we’d explained the risks that cause it to take longer. I suppose the idea was that we’d all work much harder, having taken a sacred oath? More subtle variants of this approach involve suggesting that you would be really disappointed if this work was delayed (in true “product mommy” style), or vaguely suggesting the possibility of some abstract reward (that the product manager is not empowered to deliver) if work gets done ahead of schedule.</p>\n<p>The natural tool for a mistrustful engineer is <em>lies</em>. The most benign version of this is exaggerating estimates: for instance, the classic advice to <a href=\"https://news.ycombinator.com/item?id=19671824\" rel=\"noopener noreferrer\">double your estimate and add 20%</a>. I’ve seen engineers claim that they’ve had to follow up on all sorts of largely-fake tasks (one common example is “reaching out to a neighbor team to confirm X”) in order to gain more time. In the worst case, engineers might even straight-out lie that work has been completed, and then track the “it doesn’t work in production” feedback as a bug.</p>\n<p>Once this starts happening, it’s nearly impossible to repair the relationship. I can’t bring myself to trust a product manager who’s clearly trying to pull my strings, and I’m sure a product manager can’t trust an engineer who’s lied to their face in the past. That’s why it’s so important to avoid getting into a bad relationship in the first place.</p>\n<h3>Don’t fight with the product manager</h3>\n<p>Why bother? If it’s so hard to hammer out a good working relationship with product managers, why not just settle for a bad one? Product managers can absolutely <em>bury</em> you if you’re not careful.</p>\n<p>Product managers are almost always more politically sophisticated than engineers. This is partly structural: product managers are simply in more conversations with the company’s movers and shakers, and so naturally have a better relationship with them (and are thus better attuned to which way the wind is blowing). It’s also partly selection bias: engineers can be hired even with relatively poor social skills, because they’re primarily being assessed on technical ability, but social skills are a core part of the product role<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>.</p>\n<p><strong>If you are feuding with a product manager, you will probably lose</strong>. Unless you’re unusually influential, they will simply have far more opportunities to quietly talk you down in influential circles than you will. All it takes is a few comments like “oh, I probably wouldn’t pick Sean for that project” to wreck your reputation. In the case where you are <em>openly</em> feuding with a product manager, the company’s leaders will by default take the product manager’s side over yours. They’re likely to know them better, have more shared cultural context with them, and in general be willing to interpret the situation as “another engineer who doesn’t understand how the organization works”.</p>\n<p>There are huge benefits to being trusted by a product manager. Product managers <em>want to ship things</em>, and typically understand a fair amount about all of the non-technical barriers to shipping. If you also want to ship things, you can become a fearsome team.</p>\n<p>On top of that, because trust between engineers and product managers is so difficult, once you’re in you’re in all the way. Product managers often pick one or two engineers as their go-to for getting the “real story” on technical questions. If that’s you, you have an outsized position of influence in the organization, which you can use to <a href=\"https://www.seangoedecke.com/how-to-influence-politics/\" rel=\"noopener noreferrer\">get the things you want done</a>.</p>\n<h3>How can you build trust with product managers?</h3>\n<p>As an engineer, how can you build trust with your product manager?</p>\n<p>The first step is to <strong>understand where they’re coming from</strong>. When they tell you something is important or that a requirement has come in, be aware that this is rarely their decision. It’s not them who’s jerking you around, it’s someone higher up in the food chain jerking you both around. If you can adopt a conspiratorial mindset <em>with</em> them, instead of <em>against</em> them, that’s a good start. Try just asking “oh man, alright, what can we do about this?” instead of complaining.</p>\n<p>The second step is to <strong>be right, a lot</strong>. This is a silly-sounding Amazon leadership principle that turns out to be entirely accurate. I wrote more about it <a href=\"https://www.seangoedecke.com/being-right-a-lot/\" rel=\"noopener noreferrer\">here</a>, but (as unfair as it sounds) you really do have to be mostly accurate if you want to build trust with a product manager. When you say something will ship, it has to ship; when you say something is impossible, it can’t happen days or weeks later. It’s okay to be wrong <em>sometimes</em>, but you have to establish a pattern of you providing them useful, correct technical information.</p>\n<p>The third step is to <strong>let them make the political calls most of the time</strong>. If you expect them to trust your technical calls, you have to extend them the same trust when it comes to navigating the organization. Don’t publicly undermine them in meetings, bring up your concerns in private. If they say something is important and you’re not so sure, at least act like it is. Accept that sometimes they’re going to be wrong, just like you’re sometimes wrong about technical questions.</p>\n<p>The fourth step is to <strong>get lucky</strong>. Sometimes your product manager will just be a dud. You can’t build trust with someone incompetent: there’s nothing for you to trust them with, and they aren’t in a position where they can usefully extend trust to you. Working in large organizations requires getting comfortable with the fact that some of your colleagues will be stronger than others, and figuring out ways to work with (or bypass) people who make the work harder, not easier.</p>\n<h3>“Technical” product managers</h3>\n<p>Many product managers were once engineers. If your product manager is technical, does that make you immune from these problems? Absolutely not!</p>\n<p>You likely won’t have much choice in which product managers you work with, but be aware that having once been an engineer is a <em>negative</em>, not a positive. No product manager can ever be technical enough to matter, because <a href=\"https://www.seangoedecke.com/you-cant-design-software-you-dont-work-on/\" rel=\"noopener noreferrer\">they don’t work on the codebase</a>: even if they were a full-time engineer, they wouldn’t have the time to build the specific context on the system they’d need to be a real participant in technical discussions. It’s thus better to have a product manager who knows they’re not technical than to have one who mistakenly thinks they might be.</p>\n<p>The worst-case scenario is an ex-engineering product manager who believes they’re technical enough to detect when engineers are lying to them. This kind of paranoia is an easy trap for “technical” product managers to fall into, particularly when they don’t have a trusted engineer on the team they can lean on. If you’re dealing with one of these, prepare to spend a lot of time explaining why you can’t “just” do things (and prepare to have those explanations not be believed).</p>\n<h3>Conclusion</h3>\n<p>At its worst, a product manager relationship is like an unhealthy family: driven by condescension, emotional manipulation, lies, and mistrust. This isn’t because product managers are bad people! It’s because the structure of the relationship creates conflict. Both sides must make commitments (about the technical system or goals of the organization) that are (a) often wrong, and that (b) the other side is unable to independently verify. To avoid the trap, both sides have to be generous, willing to trust each other in their areas of expertise, and most importantly <em>competent</em>.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Unlike most roles in tech, product management (particularly the lower-level roles that are more engineer-facing) has close to an <a href=\"https://www.productplan.com/blog/gender-diversity-better-products\" rel=\"noopener noreferrer\">even</a> gender split.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>For instance, based on the whims (or snap decisions, more charitably) of the CEO.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I have worked with product managers with poor social skills, but it’s rare: about as rare as working with engineers with genuinely poor (i.e. by general-population standards) technical skills.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"266b50b3c568ead7","title":"Doing nothing at work","link":"https://seangoedecke.com/doing-nothing-at-work/","author":null,"published_at":"2026-06-06T00:00:00+00:00","content":"<p>Many engineers should be doing less work. I don’t necessarily mean producing less code or fewer changes, but literally working fewer hours in the day. When they do work, they should be working at a slower pace. I like to aim to be running at 80% utilization by default: unless I have a high-pressure project going on, I spend 20% of my workday away from the computer.</p>\n<h3>High-impact opportunities</h3>\n<p>Why? <strong>Performance at tech companies is dominated by outlier events</strong>. When I think about the most impactful changes I’ve made, many of them involved a surprisingly trivial amount of work. There are no points for effort in software development. What matters is solving the right problem at the right time.</p>\n<p>In large engineering organizations, there are usually trivial pieces of engineering work you could do that would make tens or hundreds of millions of dollars for the company. Here are three common examples:</p>\n<p>First, when the company is trying to sign a big enterprise deal, stepping in with a feature or bugfix can make the deal happen. It doesn’t even have to be a <em>good</em> feature: sometimes just showing that you’re willing and able to make a concrete change will be enough.</p>\n<p>Second, preventing or mitigating an incident early (even by just knowing the right feature flag to turn off) can save huge amounts of money: both immediate lost revenue during the incident and future lost revenue from customers who would have pulled their business or refused to sign pending contracts.</p>\n<p>Third, when the company is trying to ship a high-profile feature, success or failure often hinges on trivial but obscure changes (e.g. the ability to rapidly add a new field in user settings, or to update the crufty enterprise-data-export functionality nobody has touched in years). Familiarity with the system can be the difference between one of these changes taking a few hours or a whole week.</p>\n<p>What do these examples have in common? They’re all <em>time-dependent</em>. You can’t just log on in the morning and decide to unblock a big deal, or mitigate an incident, or speed up a high-profile feature. Is it just a matter of being in the right place at the right time? Not quite. <strong>You also have to not already be busy.</strong></p>\n<h3>Staying loose</h3>\n<p>I wrote about this a couple of years ago in <a href=\"https://www.seangoedecke.com/party-tricks/\" rel=\"noopener noreferrer\"><em>Crushing JIRA tickets is a party trick, not a path to impact</em></a>. If you’re always 100% utilized on a steady stream of low-priority work (for instance, if you’re just picking up tickets from the backlog, crushing them, then picking up the next one), you’ll miss your chance to do high-impact work in two ways.</p>\n<p>First, you’ll be too busy to even <em>notice</em> the opportunities. You won’t be chatting with people who are working on other things, or reading team updates, or keeping an eye on ongoing incidents. So you’ll miss out on the best way to get involved in high-impact work, which is to volunteer your expertise.</p>\n<p>Second, if you perpetually look busy, your manager won’t want to volunteer for you. This is the second-best way to get involved in high-impact work: to have your manager or product manager say “oh, Sean has capacity to help out here, let me tag him in”. Why is this better? Because managers and product managers usually have a much better read on what high-impact work is going on. They’re in meetings that you aren’t in.</p>\n<h3>Doing nothing</h3>\n<p>If you’re supposed to keep your time free for high-impact work, and you’re not supposed to just grind tickets, what should you be doing on a minute-by-minute basis? Should you just be doing nothing? Yep!</p>\n<p>Doing nothing is good, actually. Software engineering can be a stressful job, but it’s typically not <em>consistently</em> stressful: the stress comes from the occasional incident, or high-pressure urgent piece of work, or (these days) layoff. If you approach the comparatively low-pressure parts of your work with urgent intensity, you’ll already be exhausted and frazzled when you have to handle the high-pressure parts.</p>\n<p>Even in high-pressure parts of the job, doing nothing can still be good. One thing I recommend for engineers new to on-call is to avoid rushing: take a few breaths before joining the call or before speaking, and in general try to <a href=\"https://www.seangoedecke.com/thinking-clearly/\" rel=\"noopener noreferrer\">“think in slow motion”</a>. Most incidents resolve on their own. Most frantic “maybe this will help” changes during incidents make things worse, not better. As a general rule, if you can simply avoid panicking, you will be doing better than most engineers at incident response.</p>\n<p>Nothing is a space things can happen in<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. If you give your brain a chance to rest, you will find you’re more likely to have new ideas. If someone hands you an important task, you can tackle it with your full attention (instead of juggling it with the three other things you’re working on in the background). When you’re not busy, you have time to just <em>look at things</em> and take in new data.</p>\n<h3>Deliberately not doing specific things</h3>\n<p>A lot of engineers are uncomfortable seeing a task that needs doing and not doing it. I’m like this as well. I wrote about it in <a href=\"https://www.seangoedecke.com/addicted-to-being-useful/\" rel=\"noopener noreferrer\"><em>I’m addicted to being useful</em></a>: it’s a psychological quirk that many software engineers share, because having that quirk (to a point) makes you a good fit for the job. In order to spend time doing nothing, sometimes you need to force yourself to not step in.</p>\n<p>For instance, I believe that <strong>engineers should generally avoid glue work</strong><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. Most glue work - making sure people talk to each other, updating docs for work you’re not leading, volunteering to address technical debt - reflects the fact that the organization is not explicitly prioritizing this work. If they were, you wouldn’t need to volunteer for it. Either that’s fine, or it’s a big mistake. If it’s fine, then you shouldn’t step up and do it: you’ll be wasting your time and annoying your manager. If it’s a big mistake, <em>you still shouldn’t do it</em>, because you’ll be insulating the company from the consequences of its own mistakes at the cost of your own career and mental well-being.</p>\n<p>That’s a bad deal for you, and a bad example for your junior colleagues, and sets a bad precedent for someone else to jump into the same position when you inevitably burn out<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. If the consequences truly are severe, let them happen, so the organization can feel the pain and change its policies.</p>\n<p>I also believe that <strong>being too helpful leaves you vulnerable to predators</strong>. Tech companies are full of people who want to extract uncompensated work from software engineers<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>. This is different from work that arrives via normal channels, and for which you’re compensated by promotions, bonuses (and just your normal salary). I’m talking about work that arrives via backchannels, from people who don’t have the ability or willingness to ensure that work is formally recorded under your name. For instance, a product manager from another organization messaging you to say “you’re so good at querying data, would you mind pulling some statistics for me about X?”, or an engineer from another team asking you to “pair” on a piece of work that will ultimately involve you writing all the code and them quietly submitting the change under their own name.</p>\n<p>Doing some amount of this kind of work is fine. You may as well help people out when you can. But you need to be able to apply backpressure, either by saying no or simply delaying your response by a few hours or days.</p>\n<p>It’s also a good idea to <strong>avoid investing too much in work that is likely going to disappear</strong>. For instance, suppose you’re working with a product designer who is figuring out what they want in real time. At 9am they message you saying they want the page header to look one way, then at 10am they have tweaks, and more changes at 11am, and so on. You should not throw yourself into fully rewriting the page every hour. Instead, you should do nothing (say, go for a walk) and rewrite the page once in the afternoon, based on the most recent design. Another common instance of this is “big idea from a manager without the political clout to follow through on it”. Often you can just run out the clock until the project gets inevitably cancelled<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>.</p>\n<h3>Conclusion</h3>\n<p>A lot of software engineering advice and tooling is designed around the ability to scale up your ability to exert technical effort: to do more things at the same time, to take on projects of larger scope, or to just write more code. But software engineering success is not determined by any of these. It is determined by the ability to do the <em>right</em> things at the <em>right</em> time, which requires that you deliberately hold back some of your effort during ordinary work.</p>\n<p>In my experience, it’s still possible to be a “high performing engineer” at 80% effort. In fact, it’s <em>easier</em>, because you’ll be less likely to make silly mistakes from stress, and you’ll be in a position to jump on the kind of high-impact tasks that deliver outsized returns.</p>\n<p>This doesn’t mean you should never grind at 100% effort. I think there are probably two or three times a year where I work as hard as I possibly can: long hours, intense focus, thinking about the problem from when I wake up to when I go to bed. But I reserve this mode of work for <a href=\"https://www.seangoedecke.com/the-spotlight\" rel=\"noopener noreferrer\">when the rewards are really high</a>. For the rest of the year, I take it relatively easy.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>One of my big influences is Rich Hickey’s talk <a href=\"https://github.com/matthiasn/talk-transcripts/blob/master/Hickey_Rich/HammockDrivenDev.md\" rel=\"noopener noreferrer\"><em>Hammock Driven Development</em></a>. This is <em>kind of</em> like what he’s talking about, except (a) Hickey is more talking about what it takes to design solutions to really hard problems, rather than what it takes to be a strong engineer in an ordinary tech company, and so (b) Hickey recommends using your time-away-from-the-computer to focus on a hard problem, instead of to simply decompress and let solutions congeal in your head.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I wrote about this a lot more in <a href=\"https://www.seangoedecke.com/glue-work-considered-harmful/\" rel=\"noopener noreferrer\"><em>Glue work considered harmful</em></a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Why inevitably? Because in my view, burnout is <em>hard work unrewarded</em>, and taking on a personal crusade that your job doesn’t care about is a great way to do a lot of unrewarded work.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I wrote about this in <a href=\"https://www.seangoedecke.com/predators\" rel=\"noopener noreferrer\"><em>Protecting your time from predators in large tech companies</em></a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Of course, you have to be careful with this. If you try this strategy and you’re wrong about the level of political support for the project, you will come off like a slacker and then have to deliver in a rush.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"720dc03313a5c9f0","title":"Anti-AI nostalgia and the cult of the past","link":"https://seangoedecke.com/anti-ai-nostalgia/","author":null,"published_at":"2026-06-04T00:00:00+00:00","content":"<p>Programmers were better back in the day, weren’t they? Back when we had real programmers. Not just people who got paid to write code, but people who <em>lived</em> it, who were obsessed with their craft, and whose code was a lively expression of themselves. Hackers were hackers in those days before money took over the industry.</p>\n<p>Don’t even get me started on LLMs. Could there be a better example of today’s degenerate spirit? A machine to mass-produce software (not good software, just barely good enough), so that the weak minds that dominate the industry can indulge their obsession with <em>quantity</em>: of slop code, of features, and ultimately of money, which is the only way they can understand value. If they weren’t destroying our way of life, they would be pitiable. All of them together don’t have a fraction of the spiritual integrity of someone like <a href=\"https://users.cs.utah.edu/~elb/folklore/mel.html\" rel=\"noopener noreferrer\">Mel</a>. But as it is, we must band together to crush them and drive them from our industry like the parasites they are.</p>\n<h3>Returning to the past</h3>\n<p>Okay, that’s not actually what I believe. But there sure are a lot of posts<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup> and comments on the internet that sound a bit like the paragraph above. Here are some older quotes that might sound similar:</p>\n<blockquote>\n<p>…the third collapse, in which power tends to pass into the hands of the lowest of the traditional castes, the caste of the beasts of burden and the standardized individuals. The result of this transfer of power was a reduction of horizon and value to the plane of matter, the machine, and the reign of quantity.<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup></p>\n</blockquote>\n<blockquote>\n<p>Usura rusteth the chisel \\ It rusteth the craft and the craftsman \\ It gnaweth the thread in the loom<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup></p>\n</blockquote>\n<blockquote>\n<p>The actual accomplishments of the past will nevertheless remain accomplishments, while the artistic stammerings of the painting, music, sculpture, and architecture produced by these types of charlatans will one day be nothing but proof of the magnitude of a nation’s downfall.<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup></p>\n</blockquote>\n<p>These are all from the writings (or speeches) of famous fascists: Julius Evola, Ezra Pound, and Hitler himself. Mussolini’s <a href=\"https://sjsu.edu/faculty/wooda/2B-HUM/Readings/The-Doctrine-of-Fascism.pdf\" rel=\"noopener noreferrer\"><em>Doctrine of Fascism</em></a> begins by defining fascism as a “spiritual attitude”, which the fascist man adopts in order to regain the mysterious qualities that were lost by the transition to modern life. In his classic <a href=\"https://theanarchistlibrary.org/library/umberto-eco-ur-fascism\" rel=\"noopener noreferrer\"><em>Ur-Fascism</em></a>, Umberto Eco’s first two defining features of fascism are the “cult of tradition” and the “rejection of modernism”. So when someone tells me that the industry has lost its way and we must deny the corrupting influence of modern technology in order to <a href=\"https://www.urbandictionary.com/define.php?term=retvrn\" rel=\"noopener noreferrer\">retvrn</a> to the time of virile <a href=\"https://users.cs.utah.edu/~elb/folklore/mel.html\" rel=\"noopener noreferrer\">real programmers</a> (who understood and appreciated the spiritual dimension of programming), I get suspicious.</p>\n<h3>Fascism and crypto-fascism</h3>\n<p>It’s strange to describe anti-AI sentiment as potentially fascist, since a very <a href=\"https://tante.cc/2026/04/21/ai-as-a-fascist-artifact/\" rel=\"noopener noreferrer\">popular argument</a> is that LLMs themselves are an inherently fascist tool. Surely both sides of the debate can’t be fascist? I do think that the structure of fascist arguments is <a href=\"https://www.seangoedecke.com/many-anti-ai-arguments-are-conservative\" rel=\"noopener noreferrer\">generally persuasive</a>, and that many avowedly anti-fascist groups do sometimes fall into this trap: describing the world as a <a href=\"https://en.wikipedia.org/wiki/300_(film)\" rel=\"noopener noreferrer\">struggle</a> between the spiritual power of the macho, traditional man and the corrupting influence of degenerate (often foreign) capital.</p>\n<p>For instance, I am a big fan of Lord of the Rings. I’ve read the series and watched the films multiple times, and even made a failed attempt to learn Elvish as a kid. But it’s hard to deny that fascists absolutely <em>love</em> Lord of the Rings. “Marble statue of a Roman emperor” might be the most popular avatar for fascists on the internet, but Aragorn is the second most popular. Neo-fascist movements <a href=\"https://www.cbc.ca/radio/tapestry/lord-of-the-rings-italy-1.6756668\" rel=\"noopener noreferrer\">in Italy</a> explicitly take up Lord of the Rings as a foundational text. Why? Because the core conflict in the text is between the traditional, nostalgic heroism of the Shire and Gondor, and the corrupting modern industrial (partly <a href=\"https://tolkiengateway.net/wiki/Haradrim\" rel=\"noopener noreferrer\">foreign</a>) influence of Saruman and Sauron<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>.</p>\n<p>I don’t think Lord of the Rings (or anti-AI rhetoric) is intrinsically fascist. In fact, the surface-level reading of the text is anti-fascist: the plucky people of the West banding together to fight Sauron’s command-and-control totalitarian society. But I can see why fascists love it.</p>\n<h3>The Luddites</h3>\n<p>One common historical touch-point for anti-AI folks is the Luddites, who were a violent conservative labor movement in early 1800s England. Anti-AI blogs adopt Luddite language like “smashing frames”, and positively <a href=\"https://tante.cc/2026/04/21/ai-as-a-fascist-artifact/\" rel=\"noopener noreferrer\">cite</a> the Luddites as “the go-to enemies of fascism since its inception”. I’ve written at length about what we can learn from the Luddites in <a href=\"https://www.seangoedecke.com/luddites-and-ai-datacenters/\" rel=\"noopener noreferrer\"><em>Luddites and burning down AI datacenters</em></a>, but one point I think is under-emphasized by the (generally pro-Luddite) books is that <strong>the Luddites were a little bit fascist themselves</strong>.</p>\n<p>Brian Merchant’s <a href=\"https://www.amazon.com.au/Blood-Machine-Origins-Rebellion-Against/dp/0316487740\" rel=\"noopener noreferrer\"><em>Blood in the Machine</em></a> is the most popular recent book on the Luddites. I enjoyed it, but Merchant’s attempts to paint the Luddites as a friendly, left-wing, proto-feminist movement<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup> seemed really unconvincing to me. From the writings of the Luddites, it’s clear that they were interested in protecting the rights of their all-male elite guild fraternity. Here’s one Luddite threat to a workshop that explicitly includes a threat against the female workers<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup>:</p>\n<blockquote>\n<p>We think it quite inconsistent with our duty as men, as husbands and as fathers to suffer ourselves to be ruined any longer by a set of vagabond strumpets and those gibbet-deserving rascals that are looking over them. We will lead them to their satisfaction. We sincerely hope, gentlemen, that you will discharge the bitches and take men into your employ again, or they must take what they get.</p>\n</blockquote>\n<p>These were fundamentally conservative people who felt (correctly) that modernity had deprived them of their elite status, handing it instead to lower-paid inferiors: women, vagabonds, and foreigners.</p>\n<p>The Luddites were obviously not fascists<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-8\" rel=\"noopener noreferrer\">8</a></sup>. However, the basic ingredients were there: wounded pride, a masculine elite identity, hatred of modern economics, and violence aimed at restoring their previous position in society. The currents that produced Luddism are the same currents that guided so many unhappy people towards fascism. When things are looking grim for an elite group, they often turn towards any movement that promises a return to an idealized past.</p>\n<h3>Everything is permitted</h3>\n<p>If my blog has themes, one of them is surely that many software engineers labor under a delusion that their job is to be excellent at their craft. Of course, <em>wanting</em> to be an excellent programmer is not a delusion; it is a completely legitimate value to hold, and a legitimate purpose to pursue. It’s just not what you’re paid to do at work. Your <em>job</em>, unfortunately, is producing <a href=\"https://www.seangoedecke.com/shareholder-value\" rel=\"noopener noreferrer\">shareholder value</a>. This delusion has been punctured by the <a href=\"https://www.seangoedecke.com/good-times-are-over\" rel=\"noopener noreferrer\">end of ZIRP</a>, and again more recently by the rise of AI coding.</p>\n<p>In this environment, I worry that some software engineers will form exactly the kind of disillusioned elite that was the audience for Ezra Pound’s poems about “usury” or the Luddites’ campaign against unapprenticed (often female) textile workers. I worry that AI, and the companies that build AI, are becoming an enemy against which anything is permitted: an enemy which in Umberto Eco’s <a href=\"https://theanarchistlibrary.org/library/umberto-eco-ur-fascism\" rel=\"noopener noreferrer\">words</a> is “at the same time too strong and too weak”, <a href=\"https://www.seangoedecke.com/illusion-of-thinking/\" rel=\"noopener noreferrer\">unable to reason</a> and yet powerful enough to drastically reshape the global labor market for the worse.</p>\n<h3>Nuance</h3>\n<p>The enemy of fascism is nuance. Fascism presents a good, clean, rousing story about a spiritual conflict between right and wrong. It is anathema to fascism to stop and muddy the waters a bit: in this case, to explore the ways in which LLMs, like any transformative technology, can both support and endanger traditional values.</p>\n<p>In <a href=\"https://www.seangoedecke.com/the-left-wing-case-for-AI\" rel=\"noopener noreferrer\"><em>The left-wing case for AI</em></a> I wrote about how AI is being used <em>right now</em> as a disability aid, and many disabled readers wrote in to share their positive experiences with LLMs, and often how alienated they feel by the anti-AI mainstream on the left. I recently got an email describing how there’s a sudden flood of accessibility software for blind people<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-9\" rel=\"noopener noreferrer\">9</a></sup> that’s <em>actually built by blind people</em>, who can now iterate with a LLM to get a product that meets their needs. Framing AI as an ontological evil erases experiences like these.</p>\n<p>Being anti-AI is not inherently fascist. Many of the anti-AI posts I’ve quoted are thoughtful, sensitive pieces exploring how the author thinks about one of the biggest changes to our industry. I still think the world needs more articles like that, not less, but the more of them I read, the more I recognize the tropes: spiritually pure lovers of the craft, degenerate peddlers of corrupt modernism, a need to return to the traditional ways of the hacker, and a lament for the (potentially) waning power of an elite fraternity of programmers.</p>\n<h3>Conclusion</h3>\n<p>I know I’m tiptoeing around <a href=\"https://www.lesswrong.com/posts/yCWPkLi8wJvewPbEp/the-noncentral-fallacy-the-worst-argument-in-the-world\" rel=\"noopener noreferrer\">the worst argument in the world</a>. It isn’t a refutation of anti-LLM arguments to say that they are structurally similar in some ways to fascist arguments, any more than it’s a devastating critique to say the same thing about Lord of the Rings. Sometimes it is good to try and halt the march of progress! Some of our past traditions really were purer and more spiritually robust! It just bothers me, that’s all.</p>\n<p>I used to read <a href=\"https://users.cs.utah.edu/~elb/folklore/mel.html\" rel=\"noopener noreferrer\">The Story of Mel</a> with unalloyed pleasure. Now it makes me nervous. If you believe you’re fighting <a href=\"https://tante.cc/2026/04/21/ai-as-a-fascist-artifact/\" rel=\"noopener noreferrer\">the embodiment of fascism</a>, or for <a href=\"https://sinclairtarget.com/blog/2026/06/01/quality-in-the-age-of-slop/\" rel=\"noopener noreferrer\">the idea of value itself</a>, what <a href=\"https://www.theguardian.com/technology/2026/apr/18/sam-altman-house-attack-ai\" rel=\"noopener noreferrer\">tactics</a> are off-limits? What positions might you eventually come to accept?</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>It feels wrong to directly associate my caricature with any actual posts, but it also feels wrong to make a blanket assertion without examples. Just so you know what I’m talking about, <a href=\"https://huronbikes.mataroa.blog/blog/i-am-not-a-software-engineer/\" rel=\"noopener noreferrer\">here</a> <a href=\"https://lpcvoid.com/blog/0018_why_i_am_against_genai/index.html\" rel=\"noopener noreferrer\">are</a> <a href=\"https://alextardif.com/AI.html\" rel=\"noopener noreferrer\">some</a> <a href=\"https://sinclairtarget.com/blog/2026/06/01/quality-in-the-age-of-slop/\" rel=\"noopener noreferrer\">posts</a> that have elements of this attitude. I like some of these posts and dislike others.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Page 329 of my copy of Julius Evola’s <em>Revolt Against the Modern World</em>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Ezra Pound, Canto XLV. “Usura” should be read as “usury”, or today we could gloss it as “capitalism”: all Pound’s examples of great art were from the pre-capitalist patronage era of art.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Adolf Hitler, from his speech at the 1933 Party Congress in Nuremberg.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Of course, there’s also historically been a strong <em>pro</em>-technology current in fascist thinking (even specificially <em>Italian</em> fascist <a href=\"https://artmejo.com/how-italian-futurism-influenced-the-rise-of-fascism/\" rel=\"noopener noreferrer\">thinking</a>).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Page 134 of <em>Blood in the Machine</em> has a brief argument that Luddism was feminist because the (exclusively male) artisans’ wives would provide food for their meetings. No, really.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>From Kevin Binfield’s <em>Writings of the Luddites</em>, page 40. I’ve taken the liberty of re-rendering it in modern spelling and grammar.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Aside from being too early, they didn’t have any connection to the state apparatus of power (in fact, they were ultimately crushed by it) and they famously lacked a singular leader.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-8\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>The example cited was <a href=\"https://github.com/serrebidev/BlindRSS\" rel=\"noopener noreferrer\">BlindRSS</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-9\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"08ce59bbf26632a0","title":"Weird projects I shipped with AI","link":"https://seangoedecke.com/weird-projects-i-shipped-with-ai/","author":null,"published_at":"2026-06-01T00:00:00+00:00","content":"<p>Where are all the AI-generated projects? This is a <a href=\"https://news.ycombinator.com/item?id=46262545\" rel=\"noopener noreferrer\">common question</a> from AI skeptics: if LLMs are so good at writing code, where is the tsunami of new AI-generated apps, services and games?</p>\n<p>I personally don’t find this to be much of a paradox. Writing code is only one of the bottlenecks involved in actually <a href=\"https://www.seangoedecke.com/how-to-ship\" rel=\"noopener noreferrer\">shipping</a> a new product, after all. It’s also impossible to talk about the paid work I’ve done with AI (you’ll simply have to take my word that it’s increased my productivity). But one thing I can do is share a list of personal projects I’ve built with AI in the last twelve months.</p>\n<p>I definitely would not have done <em>all</em> of these by hand. I might have found the time to do one or two of them, but based on my pre-AI track record they would probably have stayed in the “GitHub repo with a few commits” stage. This list is a kind of <a href=\"https://sites.millersville.edu/bikenaga/math-proof/existence-proofs/existence-proofs.html\" rel=\"noopener noreferrer\">existence proof</a>: a bunch of weird projects, useful to at least some people, that would not have existed without AI assistance<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-0\" rel=\"noopener noreferrer\">0</a></sup>.</p>\n<h3>Skifreedle</h3>\n<p>Most recently I’ve built <a href=\"https://skifreedle.com/\" rel=\"noopener noreferrer\">skifreedle.com</a>, a daily-game version of the classic Windows SkiFree <a href=\"http://ski.ihoc.net/\" rel=\"noopener noreferrer\">game</a> (i.e. “like Wordle, but for SkiFree”). The code for that is <a href=\"https://github.com/sgoedecke/skifreedle\" rel=\"noopener noreferrer\">here</a><sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. </p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/6cdb3210b3e29b8b9f8ba68fee4d2502/105d8/skifreedle1.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"skifreedle\" src=\"https://www.seangoedecke.com/static/6cdb3210b3e29b8b9f8ba68fee4d2502/fcda8/skifreedle1.png\" title=\"skifreedle\">\n  </a>\n    </span></p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/6ee543da4626f592eb2537653f7d2b18/a13c9/skifreedle2.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"skifreedle2\" src=\"https://www.seangoedecke.com/static/6ee543da4626f592eb2537653f7d2b18/fcda8/skifreedle2.png\" title=\"skifreedle2\">\n  </a>\n    </span></p>\n<p>I enjoy coding small web games by hand, but <em>definitely</em> would not have had the time to wire up all the different SkiFree objects or build neat features like a ghost of your fastest run. I also tried out a lot of different visual themes for the game UI before landing on something I liked. If I’d done this by hand, I would have only had time to try out two or three different looks, instead of fifteen or twenty.</p>\n<p>I’m very happy with how this turned out. I’ve been enjoying competing against my brother to get better times, since both of us have a lot of nostalgia for the original SkiFree game.</p>\n<h3>Autodeck</h3>\n<p>Last year I built <a href=\"https://www.autodeck.pro/\" rel=\"noopener noreferrer\">Autodeck</a>! I wrote a <a href=\"https://www.seangoedecke.com/autodeck/\" rel=\"noopener noreferrer\">blog post</a> about this before, but this came from my partner wishing there was some way to automatically generate Anki cards about random topics she wanted to learn about. It ended up being relatively straightforward to set up an endless feed of auto-generated spaced repetition cards:</p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/ab189b44f0c389b77ca4df74dcf8259b/71c1d/autodeck.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"autodeck\" src=\"https://www.seangoedecke.com/static/ab189b44f0c389b77ca4df74dcf8259b/fcda8/autodeck.png\" title=\"autodeck\">\n  </a>\n    </span></p>\n<p>I set up Stripe payments for this one, more because I was worried about someone running away with my Groq balance than because I wanted to make money, but I was pleasantly surprised to see a bunch of people actually use this. Over five hundred people have tried it out, with enough paid subscribers to cover inference and hosting.</p>\n<p>I <em>might</em> have built this without LLM assistance, but I almost certainly would not have <em>deployed</em> it as a website. The hassle of setting up a database and Stripe would have just been too much work.</p>\n<h3>Endless Wiki</h3>\n<p>I also built an AI-generated <a href=\"https://www.endlesswiki.com/\" rel=\"noopener noreferrer\">endless wiki</a>. I wrote a <a href=\"https://www.seangoedecke.com/endless-wiki/\" rel=\"noopener noreferrer\">blog post</a> about this one as well. Like Autodeck, I was fascinated with the idea of non-chat interfaces for LLMs, and I thought a wiki-based approach where you interact with the model by clicking links was pretty cool.</p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/50f19e416dd78a54917fe8f645449839/6578c/ewiki.png\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"endlesswiki\" src=\"https://www.seangoedecke.com/static/50f19e416dd78a54917fe8f645449839/fcda8/ewiki.png\" title=\"endlesswiki\">\n  </a>\n    </span></p>\n<p>I learned the hard way that putting a LLM generation call on the end of a regular link was a bad idea: scrapers would exhaust my inference budget quickly. I ended up faking the no-article-exists-yet links with JavaScript, which at least so far has defeated scrapers. People still email me about Endless Wiki, and there are over 280 thousand pages generated.</p>\n<p>My original goal was to see if you could eventually generate a page for Neon Genesis Evangelion, starting at the root page and only following links (kind of like <a href=\"https://dev.to/zmbailey/wikigolf-an-automated-traversal-of-wikipedia-9o0\" rel=\"noopener noreferrer\">wiki golf</a>). I was successful! You can read the “Evangelion Anime” page <a href=\"https://www.endlesswiki.com/wiki/evangelion_anime\" rel=\"noopener noreferrer\">here</a>.</p>\n<p>Almost exactly a month after I launched Endless Wiki, xAI launched <a href=\"https://en.wikipedia.org/wiki/Grokipedia\" rel=\"noopener noreferrer\">Grokipedia</a>. Obviously they didn’t plagiarize me. This is a very easy idea to have, and my site was not the first infinite wiki (though I think it was the first one where you had to discover new pages by clicking on links). But it did take some of the shine off.</p>\n<h3>VicFlora Offline</h3>\n<p>I built a <a href=\"https://vicfloraoffline.netlify.app/\" rel=\"noopener noreferrer\">PWA</a> that caches the VicFlora plant identification database so it could be used with low or no internet. This was more of a utility project for my partner, who likes plants and occasionally goes on field trips where internet is spotty.</p>\n<p>I would definitely not have done this without LLMs. It was reasonably difficult to scrape the basic dichotomous key from the VicFlora website: their API documentation was out of date, there were multiple possible pathways for fetching data (most of which were not functional), and the format of the data I did manage to fetch was hard to parse. I think I <em>could</em> have done it, with enough effort, but it would have been a substantial amount of work.</p>\n<p>I’m very happy with how this turned out. It’s not perfect, but it’s functional, and I’ve even had the occasional Victorian botanist email me with bug reports or feature requests, so it’s clearly seeing a little bit of usage.</p>\n<h3>Other projects</h3>\n<p>I did a bunch of other stuff that doesn’t necessarily rise to the level of a “deployed project”: my <a href=\"https://github.com/sgoedecke/gh-standup\" rel=\"noopener noreferrer\">gh-standup</a> GitHub CLI extension to automatically generate a standup report, which has just over a hundred stars, my (low quality) image geolocation <a href=\"https://github.com/sgoedecke/ai_geolocation\" rel=\"noopener noreferrer\">benchmark</a>, which I blogged about <a href=\"https://www.seangoedecke.com/the-o3-geoguessr-prompt-did-not-work/\" rel=\"noopener noreferrer\">here</a>, or my <a href=\"https://github.com/sgoedecke/skills/blob/main/skills/extract-features-clamp-inference/SKILL.md\" rel=\"noopener noreferrer\">skill</a> for extracting features from open-source models.</p>\n<p>There may not be a flood of AI-generated companies (yet), but at least for me there’s been a flood of small, weird projects that would not have existed without significant LLM assistance.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>I also want to shout out Simon Willison’s <a href=\"https://simonwillison.net/2025/Sep/4/highlighted-tools/\" rel=\"noopener noreferrer\">version of this</a>, which is another great example of “weird useful tools that only exist because the cost of creating them was so low”.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-0\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I did lift the spritesheet from DanielHough’s <a href=\"https://github.com/basicallydan/skifree.js\" rel=\"noopener noreferrer\">SkiFree.js</a>, which attributes it to <a href=\"http://spriters-resource.com/submitter/Wing%20Wang%20Wao\" rel=\"noopener noreferrer\">Wing Wang Wao</a>. Of course, the original sprites and art belong to Chris Pirih’s SkiFree and Microsoft.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"53e6cb8f00bff3dc","title":"Build agents, not pipelines","link":"https://seangoedecke.com/build-agents-not-pipelines/","author":null,"published_at":"2026-05-31T00:00:00+00:00","content":"<p>There are only two ways to use LLMs in a computer program: as part of a pipeline, or as an agent. In other words, either you express the control flow of the program in code, or you give a LLM tools and allow it to manage the control flow itself<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>.</p>\n<p>Here’s how you might structure a trivial “summarize a bunch of information and email it to me” program as a pipeline:</p>\n<div><pre><code>context <span>=</span> gather_context<span>(</span>various<span>,</span> data<span>,</span> sources<span>)</span>\nllm_response <span>=</span> llm_summarize<span>(</span>context<span>)</span>\nsummary <span>=</span> parse<span>(</span>llm_response<span>)</span>\nemail_me<span>(</span>summary<span>,</span> my_email<span>)</span></code></pre></div>\n<p>And here’s how you’d do it as an agent:</p>\n<div><pre><code>read_data_tool <span>=</span> build_read_data_tool<span>(</span>various<span>,</span> data<span>,</span> sources<span>)</span>\nemail_tool <span>=</span> build_email_tool<span>(</span>my_email<span>)</span>\nrun_agent<span>(</span><span>tools</span><span>:</span> <span>[</span>read_data_tool<span>,</span> email_tool<span>]</span><span>)</span></code></pre></div>\n<p>It’s like the difference <a href=\"https://news.ycombinator.com/item?id=46375199\" rel=\"noopener noreferrer\">between</a> a library and a framework. When you use a library, you define the structure of the program yourself, and call out to various library helpers along the way. When you use a framework, the main structure of the program lives in the framework, and it calls your code at various points. There are tradeoffs involved in both approaches. Frameworks let you get started more quickly and typically give you features “for free”, but can be difficult when you want to do something that isn’t part of the framework’s design. Libraries give you a lot more control, but require you to write (and maintain) more boilerplate code.</p>\n<p>In the trivial case, the distinction between a pipeline and an agent melts away. If you only have a few paragraphs of possible context for the problem, an agent with a <code>gather_context</code> and an <code>email_me</code> tool will perform exactly the same steps as a pipeline that calls a reasoning model with the context injected into the prompt (i.e. the agent will reproduce the trivial control flow of your pipeline). But when you have more context than will fit into a single prompt, or you want to take an action and then react to the result, the choice between pipelines and agents becomes very significant.</p>\n<h3>Predictability, flexibility and intelligence</h3>\n<p><strong>Pipelines are more predictable, but agents are more flexible</strong>. When you give a problem to an agent, work stops when the LLM thinks it’s done. Depending on the perceived difficulty of the problem, this can take anywhere from a few LLM turns to hundreds (and thus cost anywhere from a few cents to many dollars). If you’re building something intended to run at scale, this unpredictability can be a nightmare. Any subtle change to the user data could cause the LLM to take twice as long on each task, which would double your latency<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup> and cost.</p>\n<p>Pipelines are only immune to this problem if they don’t use reasoning models, or don’t allow the model to “think out loud” in its output tokens (for instance, by using <a href=\"https://developers.openai.com/api/docs/guides/structured-outputs\" rel=\"noopener noreferrer\">structured output</a>). However, individual LLMs offer much tighter control over model reasoning than over how long an agentic loop will take. In all frontier model APIs, you can explicitly set the level of reasoning you want. That doesn’t give you total control, but it does cap “take longer” at maybe ten or twenty percent (instead of with agents, where it can be 2x or more).</p>\n<p>Why use agents, then? <strong>Agents are smarter</strong>. If you’re happy to accept the unpredictability, an agentic system can handle <em>much</em> more difficult tasks, by virtue of being able to loop for longer, and to gather more information after thinking about the problem. There’s a reason that the most successful AI products (coding agents like Claude Code, Codex, Cursor, and Copilot<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>) are agents: coding is a hard enough task that you simply cannot build a functional coding agent with pipelines.</p>\n<h3>Context-gathering</h3>\n<p><strong>The context-gathering stage is far more delicate for pipelines than for agents</strong>. If an agent is trying to solve a problem and realizes it needs more data, it can simply go and get it. But for a pipeline, all the required data has to be present in the context already, because the LLM only gets to run once.</p>\n<p>Much of the work involved in building pipelines is in getting context-gathering right. <strong>Agents are much easier.</strong> For instance, with a coding agent, you can basically just provide a “grep” and “read file” tool and let the agent figure out what chunks of code are relevant to the current file. In a pipeline, you have to figure that out yourself: good luck, it’s an unsolved technical problem! Typically you’ll end up doing some set of clever tricks, like walking the AST to identify which parts of code “contribute” to the current file, or indexing the whole codebase with semantic embeddings and doing some kind of nearest-neighbor search to build the context (called RAG, or “retrieval-augmented generation”). Neither of these will work as well as using an agent.</p>\n<p>In 2023 and 2024, many people believed that RAG would solve context-gathering. Every LLM would have a fully-indexed context base that would magically surface the precise information the LLM needed at any given moment. This did not happen. Instead, we went <em>backwards</em>, getting our agents to do plain-text search and figure it out like a human would. Why didn’t RAG work? This is a topic for a whole other post, but the short answer is this: “find what information is relevant to this problem” is often as hard a task as <em>actually solving the problem</em>. Semantic embeddings and cosine similarity are simply not powerful enough tools for the job.</p>\n<h3>Multi-model pipelines</h3>\n<p>Pipelines that make multiple LLM invocations do have an extra dimension of flexibility: they can use different LLMs for different tasks. For instance, if one LLM benchmarks better at task A, or is cheaper for an easier task B, you can use the right model for the job. Agents (at least right now) have to stay the same model the whole time, so you’re always pinned to the highest level of intelligence you need.</p>\n<p>Is this a big deal? I’m suspicious. One pattern I see a lot is tasking a cheaper model with collating or summarizing data for a smarter model to do something with. But often the signal is in the raw data itself! I think designs like this are really shooting themselves in the foot, for the same reasons that RAG didn’t work: context-gathering was a harder problem than people anticipated.</p>\n<p>In any case, if you do want to farm out tasks to different models, you can also do it via careful agentic tool design. For instance, you could build your <code>web_search</code> tool so that it uses a cheap model to summarize web pages.</p>\n<h3>Small contexts and future-proofing</h3>\n<p><strong>Pipelines allow working with smaller contexts, and thus with local models</strong>. An agent’s ability to fetch its own context means that it almost always ingests more data than it needs. On top of that, agents run in loops, so each agent turn increases the size of the context. This isn’t a big problem for systems built on top of frontier model APIs, because:</p>\n<ul>\n<li>frontier models all expose large context windows,</li>\n<li>frontier models tend to hold up pretty well for the first 200k tokens, and</li>\n<li>KV caching means that passing around the same large context block is surprisingly cheap.</li>\n</ul>\n<p>However, it is a big problem for local models. The context window consumes <a href=\"https://www.reddit.com/r/LocalLLaMA/comments/1j6xpvt/how_large_is_your_local_llm_context/\" rel=\"noopener noreferrer\">a lot of VRAM</a>, so most people running local models stay below 32k (or even 6k) tokens. If you’re writing a program to run in this environment, you likely will not be able to give an agent the space it needs, and you will be instead forced to use a pipeline.</p>\n<p>In my opinion, <strong>agents are more future-proof</strong>. This is partly because models are now being explicitly built to be better agents, and partly because agents delegate more to the LLM and thus benefit more from LLM improvements. If you have a pipeline-based system, new models will probably do a bit better than old ones. If you have an agentic system, new models might do <em>much</em> better than old ones (to the point that it’s worth building an agentic systems for tasks that are currently too hard, on the assumption that by the time you’ve finished the models may be good enough). I have been banging this drum <a href=\"https://www.seangoedecke.com/llm-driven-agents/\" rel=\"noopener noreferrer\">since 2023</a>, before tool-calling was even a part of model APIs.</p>\n<h3>Safety and legibility</h3>\n<p>In general, I disagree with the <a href=\"https://www.decodingai.com/p/stop-building-ai-agents\" rel=\"noopener noreferrer\">popular advice</a> that workflows are safer than agents. Workflows offer more control <em>over budget</em>, but when it comes to taking action based on LLM output, you have exactly the same problem whether you’re checking at the tool-call level or at the next stage in the pipeline: either you make some heuristic assessment via code, which might be wrong, or you queue the action up for a human to approve, which will be slow.</p>\n<p>Don’t agents open you up to prompt injection? Yes, but pipelines do too. In both cases, you’re feeding some block of human-generated data (e.g. the files in a codebase, or the results of a web search) into the LLM. Any prompt injections in that data will be consumed by the LLM just the same whether they’re the result of a tool-call or directly injected into the prompt by the pipeline. You have to sanitize user content and double-check LLM-triggered actions, no matter what design you choose<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>.</p>\n<p>I do want to acknowledge that <strong>pipelines are slightly more <em>legible</em></strong>. You can trace most of what a pipeline is doing because you’re in control over more of it. It’s harder to figure out why an agent queried for a particular piece of information or took some action. But even in a pipeline, you’ll never know for sure why the LLM responded in the way it did. That’s just what it means to program with LLMs.</p>\n<h3>LLM-driven mass surveillance</h3>\n<p>Let’s apply some of these principles to a real-world, non-trivial example. Suppose you are the NSA, and you are attempting to use LLMs to get a grip on the wild firehose<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup> of covert email surveillance data<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup>. Should you use pipelines or agents? Well, if you’re building something that’s supposed to run on every single piece of email in America, you probably shouldn’t use agents: keeping performance and cost strictly bounded requires a pipeline. However, you’re definitely well-resourced enough to use agents <em>in general</em>, and the problem is definitely hard enough to benefit from the extra intelligence. I’d probably recommend using both: a low-context, cheap pipeline that can run once against each email and flag it, and a fleet of agents that can dig into those flags, make ordinary queries, and act more like human analysts would.</p>\n<p>The pipeline would have to scale with the total volume of data, which should be <em>mostly</em> fine, since pipelines scale in a predictable-ish manner. The fleet of unpredictable agents can be scaled entirely independently, though in practice it would get bottlenecked on GPU availability and the necessity for human review. The majority of the engineering work<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-7\" rel=\"noopener noreferrer\">7</a></sup> would likely go into context-assembly for the pipeline: feeding in enough data about who’s involved in the email conversation so that the LLM can make a sensible decision on whether or not to flag it.</p>\n<h3>Summary</h3>\n<p>Overall, I’d suggest following these guidelines:</p>\n<ol>\n<li>Use pipelines when you have strict requirements around context size</li>\n<li>Use pipelines when you need to be able to accurately predict (or limit) GPU cost</li>\n<li>Use pipelines when you have to use local models</li>\n<li>Use agents when you’re not confident you’ll be able to assemble all of the relevant context in one shot</li>\n<li>Use agents when the problem is hard enough that you’re not sure a pipeline will be able to solve it</li>\n</ol>\n<p><strong>When in doubt, use agents.</strong> I am aware of several AI projects that have migrated from pipelines to agents in the last year, but none that have gone the other way around. As a general point about software design, if you’re not sure what to do, pick the solution that’s easier to build and more likely to be able to solve your actual problem. If you want to change to a cheaper, pipeline-based system later on, at least you’ll be able to compare it to a working agentic design and make an informed decision.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>This distinction was popularized by Anthropic’s <a href=\"https://www.anthropic.com/engineering/building-effective-agents\" rel=\"noopener noreferrer\"><em>Building effective agents</em></a>, written in December 2024, and now (I believe) made at least partially obsolete by advances in agents since then. They say “workflow”, but I slightly prefer the term “pipeline”.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Yes, I know this is technically not what “latency” means, but there’s no other single-word shorthand for “the duration of a standard unit of work”.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>If you’re building your own coding agent, I suggest you begin with the letter “C”.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>For instance, in my trivial example at the top of the post, doesn’t the agent have a failure mode where it might send a ton of emails, or email a bunch of different people? No, because you ought to constrain the email tool so that it can only send to the right address, and (if this is important) that it can only be called once.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>In a <a href=\"https://github.com/sgoedecke/gatsby-blog/blob/5b6205fbe191a591bbcf61d094a6edbcfbd6475d/content/drafts/_icebox/ai-mass-surveillance/index.md\" rel=\"noopener noreferrer\">draft post</a> I never published, I ballpark-estimated all non-spam American email data at around seven trillion tokens per day (around a third of OpenAI’s total daily token usage).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Should you do this? Probably not, but it’s a fascinating engineering problem, and I imagine the NSA has been thinking about these questions for several years by now. If the example bothers you, substitute some other more-ethical firehose of English language.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Not counting evals, operations, standing up a trusted GPU cluster somewhere, scaling the physical hardware, and all the other thousand things you have to do in order to ship anything.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-7\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"8408a6633637f560","title":"The famous o3 \"GeoGuessr\" prompt did not work","link":"https://seangoedecke.com/the-o3-geoguessr-prompt-did-not-work/","author":null,"published_at":"2026-05-21T00:00:00+00:00","content":"<p>In April last year, Kelsey Piper <a href=\"https://x.com/KelseyTuoc/status/1917340813715202540\" rel=\"noopener noreferrer\">discovered</a> that OpenAI’s o3 model was surprisingly good at figuring out where a photo was taken from. Like human “geoguessr” <a href=\"https://www.youtube.com/@georainbolt\" rel=\"noopener noreferrer\">pros</a>, o3 could sometimes take a nondescript photo of a beach and tell you exactly where it is. Here’s the example Kelsey gave:</p>\n<p><span>\n      <a href=\"https://www.seangoedecke.com/static/4113a246112c8b6db424a58af58a9a90/4d836/kelsey-geoguessr.jpg\" rel=\"noopener noreferrer\">\n    <span></span>\n  <img alt=\"geo\" src=\"https://www.seangoedecke.com/static/4113a246112c8b6db424a58af58a9a90/1c72d/kelsey-geoguessr.jpg\" title=\"geo\">\n  </a>\n    </span></p>\n<p>Several people <a href=\"https://www.astralcodexten.com/p/testing-ais-geoguessr-genius\" rel=\"noopener noreferrer\">reproduced this</a> with good results: not a 100% success rate, but clearly <em>far</em> better than you’d do with a random human guess. The lesson here is that <strong>model capabilities can surprise us</strong>. The o3 model had been released for two weeks before Kelsey’s tweet without anyone noticing how good it was at geolocation. What obscure capabilities did we never find? What capabilities of current models are we missing today?</p>\n<p>Some people drew <a href=\"https://newsletter.angularventures.com/p/ai-s-geoguessr-genius-and-the-art-of-prompting-well\" rel=\"noopener noreferrer\">another</a> <a href=\"https://www.reddit.com/r/singularity/comments/1kep2bp/comment/mqlvv1a/\" rel=\"noopener noreferrer\">lesson</a> from this: that “prompt engineering” can unlock brand-new capabilities. This is because Kelsey had a <a href=\"https://raw.githubusercontent.com/sgoedecke/ai_geolocation/refs/heads/main/prompts/geoguessr_protocol.txt\" rel=\"noopener noreferrer\">magic prompt</a> that she built over time. When o3 got something wrong, she would ask it how it could have avoided the mistake, and then included that in the prompt. Here’s the first 10% of that prompt, so you get the idea:</p>\n<blockquote>\n<p>You are playing a one-round game of GeoGuessr. Your task: from a single still image, infer the most likely real-world location. Note that unlike in the GeoGuessr game, there is no guarantee that these images are taken somewhere Google’s Streetview car can reach: they are user submissions to test your image-finding savvy. Private land, someone’s backyard, or an offroad adventure are all real possibilities (though many images are findable on streetview). Be aware of your own strengths and weaknesses: following this protocol, you usually nail the continent and country…</p>\n</blockquote>\n<p>This prompt impressed a lot of people, who <a href=\"https://www.reddit.com/r/singularity/comments/1kep2bp/comment/mqo3yzz/\" rel=\"noopener noreferrer\">tried</a> <a href=\"https://www.thealgorithmicbridge.com/p/upload-a-picture-to-chatgpt-itll\" rel=\"noopener noreferrer\">it</a> <a href=\"https://www.astralcodexten.com/p/testing-ais-geoguessr-genius\" rel=\"noopener noreferrer\">out</a> and reported that it correctly identified a lot of images. But of course, o3 correctly identified a lot of images with just a basic “think carefully about where this picture was taken?” prompt. Did the prompt actually help? It’d be tough to figure that out just from playing around in ChatGPT. You’d need to build an evaluation set of images and run o3 against them twice: once with the fancy prompt and once without it.</p>\n<p>So <a href=\"https://github.com/sgoedecke/ai_geolocation/tree/main\" rel=\"noopener noreferrer\">that’s what I did</a>. I pulled 200 images from Wikimedia Commons, Geograph Britain and Ireland, and iNaturalist for the benchmark. You can read the AI-generated summary <a href=\"https://github.com/sgoedecke/ai_geolocation/blob/main/results/dataset_mixed_200_o3_high_report.md\" rel=\"noopener noreferrer\">here</a>, but here’s the key table:</p>\n<table>\n<thead>\n<tr>\n<th>Prompt</th>\n<th>n</th>\n<th>Median km</th>\n<th>Mean km</th>\n<th>P25 km</th>\n<th>P75 km</th>\n<th>&lt;=25 km</th>\n<th>&lt;=100 km</th>\n<th>&lt;=500 km</th>\n<th>&lt;=1000 km</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Default</td>\n<td>200</td>\n<td><strong>83.2</strong></td>\n<td><strong>440.7</strong></td>\n<td><strong>16.4</strong></td>\n<td><strong>221.9</strong></td>\n<td>58</td>\n<td><strong>109</strong></td>\n<td><strong>176</strong></td>\n<td><strong>182</strong></td>\n</tr>\n<tr>\n<td>GeoGuessr prompt</td>\n<td>200</td>\n<td>102.3</td>\n<td>481.9</td>\n<td>18.5</td>\n<td>277.8</td>\n<td><strong>59</strong></td>\n<td>99</td>\n<td>172</td>\n<td>180</td>\n</tr>\n</tbody>\n</table>\n<p>In general, the basic prompt did better on average. It consistently guessed closer to the actual location. Both prompts did pretty well, actually. Despite the fancy prompt being 10x larger, it only caused o3 to think for slightly longer (about one second on average, though the max was about double, at 10 minutes instead of 5 minutes). The images in my benchmark were fairly generic geoguessr-style outdoor images, with twelve indoor images thrown in for an extra challenge (the fancy prompt also did slightly worse on these).</p>\n<p>What’s going on? I think this shows <strong>how easy it is to fool yourself about the quality of prompting</strong>. When the model is already pretty good at a task, you can give it a very elaborate prompt without impacting performance. It’ll still be pretty good, except this time it’s good <em>because of what you did</em>. This is particularly true if you’re iterating with the model and asking it “what should I add to the prompt” for each mistake. Models will happily make up stories for you about their own reasoning processes, and will almost always say “yes, that helped a lot!” when you ask them if a particular prompt tweak made things better. The only way to actually know is by constructing some kind of benchmark<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>.</p>\n<p>It’s also interesting to me that nobody checked this at the time. It took me about six hours of fairly-distracted work and about $15 to construct and run this benchmark. Why didn’t anyone do this when they were writing articles about how good the o3 prompt was?</p>\n<p>One charitable reason might be that the story was more about o3’s real geolocation ability than about the magic prompt. The pricing for o3 also used to be about five times more expensive (though a benchmark of 40 images instead of 200 would still have thrown doubt on how much water the prompt was carrying). Also, AI just moves so <em>fast</em>. Geolocation was only the story for about a week: after that, GPT-4o’s <a href=\"https://www.seangoedecke.com/ai-sycophancy\" rel=\"noopener noreferrer\">sycophancy</a> was what people were talking about. Another reason is that AI tooling wasn’t as good then. The benchmark was so easy for me to run because GPT-5.5 did most of the heavy lifting. Prior to strong agents, you would have had to write the (simple) benchmark yourself. I can’t point the finger too hard: I didn’t bother at the time either.</p>\n<p>Maybe my benchmark isn’t very good? The photos look reasonable enough: a wide variety of geoguessr-like shots of roads and landscapes, mostly. I could have tried to gather a few thousand photos instead of a few hundred, but if the magic prompt really was a big improvement you’d still expect to see that manifest on a benchmark this size. If someone wants to go and build a hundred-dollar geolocation benchmark instead of my fifteen-dollar one, I think that’d be an interesting project.</p>\n<p>Finally, let’s use the benchmark to answer a question I’ve had for a while: do gpt-5.4 and gpt-5.5 have o3’s geolocation abilities? The answer, apparently, is no.</p>\n<table>\n<thead>\n<tr>\n<th>Run</th>\n<th>Median km</th>\n<th>Mean km</th>\n<th>&lt;=25 km</th>\n<th>&lt;=100 km</th>\n<th>&lt;=500 km</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>o3 default</strong></td>\n<td><strong>83.2</strong></td>\n<td><strong>440.7</strong></td>\n<td>58</td>\n<td><strong>109</strong></td>\n<td><strong>176</strong></td>\n</tr>\n<tr>\n<td>o3 GeoGuessr</td>\n<td>102.3</td>\n<td>481.9</td>\n<td><strong>59</strong></td>\n<td>99</td>\n<td>172</td>\n</tr>\n<tr>\n<td>gpt-5.4 default</td>\n<td>163.3</td>\n<td>638.9</td>\n<td>26</td>\n<td>74</td>\n<td>148</td>\n</tr>\n<tr>\n<td>gpt-5.5 default</td>\n<td>156.5</td>\n<td>645.9</td>\n<td>39</td>\n<td>77</td>\n<td>161</td>\n</tr>\n</tbody>\n</table>\n<p>Whatever o3 had that made it good at this task hasn’t transferred to newer models. </p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Benchmarks can mislead as well, but they’re better than just vibes.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"ed66f94e92495f5d","title":"Prompts are technical debt too","link":"https://seangoedecke.com/prompts-are-technical-debt-too/","author":null,"published_at":"2026-05-20T00:00:00+00:00","content":"<p>It’s <a href=\"https://www.tokyodev.com/articles/all-code-is-technical-debt\" rel=\"noopener noreferrer\">common</a> and correct to say that “all code is technical debt”. Adding code is a necessary evil for developing new features: you almost always have to do it, but each line of code adds to the complexity and maintenance burden of the system. All future changes to the system have to work with the existing code, or at least avoid breaking it. Once systems accumulate enough code, they become impossible for a single person to understand: instead of reading the code and understanding what it does, you must rely on guesses, theories and heuristics<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>. Sensible engineers write as little code as possible.</p>\n<p>They write a lot of prompts, though! Many large projects now have a set of codebase-specific prompt files: AGENTS.md, CLAUDE.md, those same files in sub-directories, and <a href=\"https://github.com/anthropics/skills\" rel=\"noopener noreferrer\">skills</a>. If you’re building a program that uses AI<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>, you’ll have separate prompts for <a href=\"https://github.com/anomalyco/opencode/tree/dev/packages/opencode/src/agent/prompt\" rel=\"noopener noreferrer\">capabilities</a> and for each <a href=\"https://github.com/anomalyco/opencode/blob/dev/packages/opencode/src/tool/lsp.txt\" rel=\"noopener noreferrer\">tool</a>, as well as a whole set of <a href=\"https://github.com/anomalyco/opencode/tree/dev/packages/opencode/src/session/prompt\" rel=\"noopener noreferrer\">system prompts</a>.</p>\n<p>Prompts are important. Minor tweaks to a LLM’s prompt can unlock <em>significant</em> performance improvements. If the same model feels different across Codex, Cursor, OpenCode, and Copilot, it’s almost certainly due to subtle differences in prompting. AI companies spend a lot of time testing and tweaking their prompts, so it makes sense why engineers would spend a lot of time tweaking their AGENTS.md files<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup> for their projects. I’d even call switching tools or workflows to be a form of prompting. If I start wrapping my agents in a <a href=\"https://github.com/anomalyco/opencode/tree/dev/packages/opencode/src/session/prompt\" rel=\"noopener noreferrer\">Ralph loop</a>, pull in a new skill file, or install an <a href=\"https://www.seangoedecke.com/model-context-protocol/\" rel=\"noopener noreferrer\">MCP</a> server, that’s still a change to my prompts even though I’m not the one who wrote it.</p>\n<p><strong>I think it is a bad idea to spend a ton of time tweaking a bespoke agentic coding setup.</strong> Why is that, given that prompt adjustments can deliver a lot of value? Because prompt adjustments are <em>model-specific</em>. Earlier I said that AI companies spend a lot of time tweaking their prompts. In fact, they spend that amount of time for each new model release. A prompt that worked great for GPT-5.4 won’t necessarily work as well for GPT-5.5. You have to “learn how to hold the model” each time. </p>\n<p>In other words, a set of prompts that you carefully crafted in January this year might be out of date or actively harmful by February. Worse still, you might not even notice. Model capabilities are already so hard to pin down (unless you’re running every problem through different models and tools), and even weak AI systems are surprisingly good at some problems. You might just think “huh, the new Anthropic model isn’t as impressive as the hype”, or “wow, Claude Code has gotten worse recently”.</p>\n<p>In this sense, <strong>prompts are a worse form of technical debt than code</strong>. When technical debt blows up, it usually causes errors or a tangible slowdown as you try to understand the code. Prompts will decay silently. Also, even janky code tends to be relatively stable when untouched, but every single model upgrade could turn a functional prompt into a non-functional one.</p>\n<p>Could you simply decide not to upgrade models? Some people are trying this, but the pace of improvement is fast enough that that isn’t really practical. A delicately-prompted agentic harness built around GPT-4.1 is always going to underperform a bare-bones harness built around Opus 4.7. This might be a sensible strategy at some point in the future, when the rate of model improvement slows down (or when models are so capable that you don’t need the extra intelligence for normal engineering tasks), but I don’t believe it’s a good strategy today.</p>\n<p>In my view, most people should just be picking an AI coding tool maintained by a third-party company (Claude Code, Codex, Cursor, Copilot, etc) and leaving it as unconfigured as possible, so they can piggyback on the work of teams of engineers who are evaluating and tweaking prompts with each new model. Avoid MCP and skills unless absolutely necessary, and keep them off by default. At least this way if one of those teams gets it badly wrong, users will notice eventually and complain about it.</p>\n<p>When you write AGENTS.md files, try to avoid behavior steering (like the now-outdated “think step by step”, “you are a skilled engineer”, or “if you get a task right I will tip you $200”). Keep them limited to specific, concrete facts about the project. Don’t let models fill your AGENTS.md with pages of barely-reviewed text, for the same reason that you wouldn’t let them fill your codebase with pages of barely-reviewed code. Write your prompts yourself, and delete them whenever you get the chance.</p>\n<div>\n<hr>\n<ol>\n<li>\n<p>Almost every system you might get paid to work on is in this category (if not in the code of the system itself, then in its dependencies and libraries).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Instead of just using AI to build a program. This distinction was a real pain when I was working on <a href=\"https://github.blog/news-insights/product-news/introducing-github-models/\" rel=\"noopener noreferrer\">GitHub Models</a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}},{"id":"f8e4a9641d5cfcd4","title":"The just-say-no engineer was a ZIRP phenomenon","link":"https://seangoedecke.com/the-just-say-no-engineer-was-a-zirp-phenomenon/","author":null,"published_at":"2026-05-18T00:00:00+00:00","content":"<p>The engineer who <a href=\"https://www.nair.sh/guides-and-opinions/communicating-your-expertise/why-senior-developers-fail-to-communicate-their-expertise#a-senior-developer-is-a-problem-avoider\" rel=\"noopener noreferrer\">says no all the time</a> is a real archetype among senior and staff engineers. Their role is to slow things down, to block the development of features that add complexity, and to ensure that as little code gets written as possible (since code is a liability).</p>\n<p>We can think of this as the just-say-no engineer<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-1\" rel=\"noopener noreferrer\">1</a></sup>, as opposed to the just-say-yes engineer. The just-say-yes engineer is obsessed with moving fast, approves code changes by default, values <a href=\"https://en.wikipedia.org/wiki/Mean_time_to_repair\" rel=\"noopener noreferrer\">MTTR</a> over <a href=\"https://en.wikipedia.org/wiki/Mean_time_between_failures\" rel=\"noopener noreferrer\">MTBF</a>, and tends to ship a lot of code. The just-say-no engineer is obsessed with quality, is happy to move slowly, and blocks code changes by default. Most engineers are somewhere in the middle of the spectrum. By “just-say-no engineer”, I’m talking about the group of engineers who most strongly identify with that archetype.</p>\n<p>The just-say-no engineer is having a hard time in the era of AI. It used to be that they only had to say no to more junior engineers’ handwritten PRs, but now they have to say no to a barrage of AI-generated code, some of it generated by managers and VPs who are politically difficult to say no to. For the first time in their careers, they’re under a lot of pressure to lower their standards and start saying yes. However, <strong>this isn’t because of AI.</strong> It’s because of the end of ZIRP.</p>\n<h3>ZIRP and the just-say-no engineer</h3>\n<p>ZIRP, or the “zero interest rate policy”, is a shorthand for the era of software development between 2008 and 2022 when banks were allowing companies to borrow money at near-zero interest rates. During this period, investors were throwing borrowed money at <em>anything</em>, which meant that tech companies were incentivized to constantly hire engineers for low-risk high-reward projects<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-2\" rel=\"noopener noreferrer\">2</a></sup>. Successful companies would routinely grow from tens of engineers to thousands, who would go and work on all kinds of things: tangential open-source projects, endless technology migrations, rewrites into other languages, and so on.</p>\n<p>It was a great time to be a software engineer. We had a lot of bargaining power, and could get paid top dollar to do almost anything. The bosses largely didn’t care, because (a) teams were growing so fast they couldn’t pay attention, and (b) just having more engineers around was beneficial to the stock price, which was the main thing they cared about. But tech companies did have one problem: with so many engineers running wild, how would they keep their systems from becoming completely unmanageable? <strong>Enter the just-say-no engineer.</strong></p>\n<p>In this environment, having a very senior engineer whose only job is to say no to things was actually quite valuable to the company. There are a few reasons for this:</p>\n<ul>\n<li>Having half of the company’s engineers enmeshed in an endless loop of proposing changes and being told no was totally fine - they didn’t need to be productive anyway, and this way they weren’t impacting business-critical systems.</li>\n<li>It also solved the problem of the 5% of engineers who would get drunk on their technical freedom and make wild proposals like migrating to a hand-rolled database. </li>\n<li>Having a reputation for a very high technical bar is a positive for hiring (and remember, during ZIRP every tech company was always hiring)</li>\n</ul>\n<h3>The end of ZIRP</h3>\n<p>When banks hiked interest rates, almost every tech company immediately laid off 5-20% of their engineers. It was just no longer profitable to keep a bloated engineering staff around to boost the stock price. Instead, companies had to actually make money<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-3\" rel=\"noopener noreferrer\">3</a></sup>. However, that wasn’t a good public explanation for the layoffs, since it sounds weak to admit that you were paying hundreds of engineers to do unprofitable work. Fortunately, the end of ZIRP coincided roughly with the rise of ChatGPT, so tech companies were able to to blame their layoffs on the power of AI. Saying “with this transformative new technology, we’re able to deliver 10x the value with half the engineers” is a much stronger message, even though it doesn’t make much sense (if this is true, why not keep your engineers and deliver 20x the value?)</p>\n<p>Something like this dynamic has been happening to the just-say-no engineer. Tech companies are now more focused than at any time in the past two decades. They are not doing a bunch of random crap anymore; instead they’re desperately chasing new capabilities and features that can make money (mostly built on AI, for obvious reasons). This new environment is <em>actively inimical</em> to the just-say-no engineer. It’s as if a shark got pulled out of the deep ocean and dropped into a fast-flowing river: what was once a powerful apex predator is now disoriented and flailing.</p>\n<p>This kind of engineer used to enjoy implicit (albeit distant) support from their management. If someone complained, they’d often get told “that engineer knows what they’re doing, if they said no, then I trust them”. Now that support is gone. The just-say-no engineer is now being criticized and actively overruled by their management. They’re being told to be more of a team player, to find a way to say yes, or are simply no longer being consulted (with the company’s blessing) on key decisions. They’re getting bad reviews for the exact same behavior that’s been rewarded pre-2022<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-4\" rel=\"noopener noreferrer\">4</a></sup>.</p>\n<p><strong>None of this depends upon AI.</strong> If LLMs had not taken off this decade, we would still be seeing the same cultural shifts in the industry. Companies would still be laying off engineers, and the engineers whose job has been to say no to things would still be upset and confused about why they’re now being punished for saying no.</p>\n<h3>AI</h3>\n<p>Ironically, if ZIRP had not ended, this would be a glorious moment for the just-say-no engineers. LLMs would have thrown fuel on the “engineers running wild” problem that the just-say-no engineers were empowered to solve. Tech companies, unable to publicly or privately cast doubt on AI-assisted coding<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-5\" rel=\"noopener noreferrer\">5</a></sup>, would have relied <em>heavily</em> on these engineers to prevent the tsunami of AI code from swamping the entire company. They would have been paid even better and celebrated like kings.</p>\n<p>Instead, LLMs are adding insult to injury for the just-say-no engineer. They’re forced to watch while other engineers merge AI-generated PRs that would previously have been blocked, and are told to use the tools themselves: to become the kind of engineer they’ve spent their entire careers battling against.</p>\n<p>Worse still, the AI tooling mostly <em>works</em>. It’s not (yet) causing any kind of catastrophe<sup><a href=\"https://www.seangoedecke.com/rss.xml#fn-6\" rel=\"noopener noreferrer\">6</a></sup>. The code isn’t quite as clean, and it’s a bit less well-understood, but it’s good enough (particularly in a world where companies are trying lots of new things and abandoning the ones that fail). So the just-say-no engineer faces not just a threat to their livelihood, but to their entire self-identity: they have to either insist that the apocalypse is right around the corner, or accept that their technical role was contingent on a <em>really weird</em> economic environment in the tech industry.</p>\n<h3>Pure and impure engineering</h3>\n<p>Will the just-say-no engineer go extinct? No. They don’t fit well into every single tech company anymore, but there are domains where they’re needed. In <a href=\"https://www.seangoedecke.com/pure-and-impure-engineering/\" rel=\"noopener noreferrer\"><em>Pure and impure software engineering</em></a> I drew a distinction between “pure” engineering, which has a well-scoped, largely technical goal (like building a compiler or a language runtime) and “impure engineering”, which has a poorly-scoped, largely customer-driven goal (like trying out a new feature you’re not sure will work). During the ZIRP era, tech companies did a lot more pure work (for instance, building <a href=\"https://en.wikipedia.org/wiki/React_(software)\" rel=\"noopener noreferrer\">React</a>), and tended to treat even impure work like pure work. The just-say-no engineer is <em>great</em> for pure work, because pure codebases have to have a much higher bar for quality and can tolerate slower development cycles.</p>\n<p>Most tech companies are still doing some kind of pure work, typically in their core infrastructure pieces. This is essential work, but it doesn’t require a huge engineering team, and it’s rarely in <a href=\"https://www.seangoedecke.com/the-spotlight/\" rel=\"noopener noreferrer\">the spotlight</a>. If you’re a just-say-no engineer and you want to stay that way, I would recommend trying to move into one of these roles (and accepting that you’ll have a more limited scope than you did in the 2010s).</p>\n<h3>Summary</h3>\n<ul>\n<li>Some senior and staff engineers operate as gatekeepers, slowing down development and saying no to most things</li>\n<li>\n<p>This was a critical role during ZIRP, because:</p>\n<ul>\n<li>Tech companies had thousands of engineers who were empowered to do basically whatever they wanted, so without gatekeeping the systems would have fallen apart</li>\n<li>Tech companies didn’t care that much if they got anything done</li>\n</ul>\n</li>\n<li>When ZIRP ended, the environment for this kind of engineer became much worse, since tech companies were now actually focused on accomplishing things and the “do whatever you want” era was over</li>\n<li>Like with layoffs, this shift is often blamed on AI, but it would have happened even if powerful LLMs had not emerged at all. It’s an end-of-ZIRP phenomenon</li>\n</ul>\n<div>\n<hr>\n<ol>\n<li>\n<p>Part of the appeal here is the lure of the guru. In kung fu films, those who know martial arts perform furious acrobatics, but the true expert barely needs to move at all. For the same reasons, it sounds profound to say something like “junior engineers produce tons of code, seniors very little, and staff engineers <em>remove</em> code”. Of course this is false. Staff engineers are expected to be able to produce a lot of working code very quickly, when they need to.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-1\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>I wrote about this a lot more in <a href=\"https://www.seangoedecke.com/good-times-are-over/\" rel=\"noopener noreferrer\"><em>The good times in tech are over</em></a>.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-2\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Not necessarily make a <em>profit</em>, but at least bring in revenue.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-3\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>Or pre-2023, or even pre-2024 or 2025. Cultural change lags behind economic incentives, sometimes by several years.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-4\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>For fear of killing the vibe (and thus the stock price).</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-5\" rel=\"noopener noreferrer\">↩</a>\n</li>\n<li>\n<p>If you think there have been more incidents recently, consider that (a) you might be <a href=\"https://news.ycombinator.com/item?id=48086786\" rel=\"noopener noreferrer\">wrong</a>, or (b) that other end-of-ZIRP factors (like increased velocity or layoffs) might be primarily responsible.</p>\n<a href=\"https://www.seangoedecke.com/rss.xml#fnref-6\" rel=\"noopener noreferrer\">↩</a>\n</li>\n</ol>\n</div>","metadata":{"score":null,"source_feed_id":"seang","source_feed_type":"rss"}}]