Every few months the AI industry finds a new name for a problem it never solved. The failure rate stays roughly where it was. The safety commitments stay mostly in the press release. What changes is the vocabulary — "pivot," "rebrand," "next chapter" — and each new term buys another round of the same pitch. I've spent enough time watching this pattern that I want to write down six places I see it clearly, and what I think they add up to.
1. The industry measures success at the demo layer
The number that gets quoted everywhere is that something like three in four enterprise AI projects never reach production, or never deliver ROI anyone can point to. The sourcing on that figure is inconsistent — different firms, different methodologies, different years — so I won't pretend it's a precise measurement. But the shape of the claim matches what I see: a model that performs well in a demo gets treated as equivalent to one that performs well in production, and those are not the same achievement.
A demo runs on curated inputs, in a controlled setting, for an audience that wants it to work. Production runs on whatever your users actually type, at whatever hour they type it, forever. The gap between those two conditions is where most of that failure rate lives, and it's not a bug in execution — it's what you'd expect when you evaluate a system in the wrong environment.
There's a second cost that almost never makes it into the ROI deck: the AI boomerang. Companies automate a workflow, then discover a new labor category standing where the old one used to be — prompt engineers, output reviewers, people whose whole job is auditing what the model produced. The savings don't disappear. They come back as a different line item, and I don't think most projections account for that before they get approved.
2. The AGI definition problem
I've written elsewhere about why Geoffrey Hinton's definition of AGI — a system "better than the average human" — doesn't hold up as a scientific claim, and what I think a testable replacement looks like.1 I won't re-run that whole argument here, but the short version matters to everything else in this piece: when the industry's biggest term has no agreed test, no agreed metric, and a moving bar, every lab gets to claim it's close to something nobody can actually check. That's not a side issue. It's the same move that shows up in the other five problems below — vague language doing the work that a real commitment should be doing.
3. Engineers win, scientists get the press release
Every AI lab I've looked at runs two cultures under one roof. Engineers optimize for what ships — latency, cost per token, retention. Scientists optimize for what's understood — interpretability, failure modes, the theory underneath the product. These are not the same job, and in a growth-driven lab, they are not resourced the same way. Engineers win the budget fights. Scientists get quoted in the launch announcement and then go back to a team that isn't growing nearly as fast as the one that shipped the feature.
I don't think this is a personnel problem. It's what you get when the market only pays for shipping. Scaling law research and inference optimization both get called "AI research" in the same breath, and that flattens a distinction that actually matters: one is trying to understand what these systems are, the other is trying to make them faster. Safety and interpretability commitments get announced generously and funded thinly, and the result compounds with every release — more powerful systems, understood less well than the ones before them. We are building ahead of our own understanding, on purpose, because that's what the incentives reward.
4. The loop that never pays anyone
Two extraction points sit inside the same pipeline, and I think they're the same problem wearing different clothes.
The first is training data. Large models are built on internet-scale scraping — creative work, journalism, research, personal writing — taken without consent, credit, or payment. The lawsuits frame this as copyright, and copyright is the part that's litigable, but I think the deeper issue is labor: millions of people generated the value these systems run on, and essentially none of them were compensated for it.
The second is post-training. RLHF runs on human raters, often paid piece-rate through platforms like Scale AI or Remotasks, doing cognitively demanding evaluation work for wages that don't reflect that. And underneath the paid raters is a much larger, unpaid layer: every user who corrects an answer, thumbs down a response, or just keeps using the product is doing quality assurance for free. You are the customer, the product, and the QA department, simultaneously, and only one of those roles gets billed.
Run the whole loop start to finish: scrape the data, train the model, ship it to users, let users refine it for free through their own usage, then retrain on a bigger, user-validated dataset and start again. I don't think this loop is an unfortunate side effect of how these systems get built. I think it's the business model, and neither end of it — what goes in, or what comes back out — involves paying the people who make the system valuable.
5. "Propose-Decide-Execute" is a patch wearing a philosophy's clothes
Early autonomous agents failed in specific, embarrassing ways: hallucinated tool calls, irreversible actions taken on bad reasoning, no ability to recover once something went wrong mid-task. The industry's response was to insert a human approval step between proposal and execution, and then to name that insertion "Propose-Decide-Execute" and present it as a mature, safety-conscious architecture rather than what it actually is — a checkpoint added because the thing kept breaking.
I don't think PDE solves the underlying failure. It relocates the blame. If the model proposes something wrong and a human approves it, the human owns the error. If the model proposes something right and a human overrides it, the model's whole value proposition takes the hit. What that architecture actually optimizes for is liability management, not autonomous capability — and calling it a "pivot" obscures that full autonomy, the thing that was actually promised in 2023, still hasn't been delivered. Compare that year's "autonomous agents will replace knowledge workers" pitch to 2025's "humans stay in the loop for high-stakes decisions" framing. Same product. Rewritten terms and conditions.
6. Companies are paying to train their own replacements
Here's the loop I find hardest to look away from. A company buys an AI tool to raise productivity. Employees use it, and their usage — patterns, corrections, workflows — gets logged. That data trains the next version of the tool. The next version automates more of what those employees were doing. Now there's a business case for cutting the headcount in those roles.
Companies aren't passive victims in this. The executives who approve the AI tooling budget are frequently the same ones who later approve the layoffs that budget makes possible. The company is actively authoring its own workforce displacement, and paying twice for it — once in licensing fees, again in severance.
The part I find genuinely bleak is what happens to the individual worker inside that loop. The employees who use these tools best are, by definition, the ones who demonstrate most clearly how their own job can be automated. Competence becomes the training signal for your own replacement. And there's no structural brake on any of this, because everyone involved — the worker showing off what the tool can do, the manager reporting the productivity win — is rewarded on a timeline that's much shorter than the one the damage shows up on.
What this adds up to
I don't think any of this is growing pains. Growing pains resolve. What I'm describing is six gaps that rebranding doesn't close, because rebranding was never meant to close them:
The financial failure rate is the gap between demo and deployment. The AGI definition fight is the gap between scientific rigor and investor narrative. The Engineer-vs-Scientist divide is the gap between shipping and understanding. The data-and-labor loop is the gap between who creates value and who gets paid for it. The PDE rebrand is the gap between the autonomy that was promised and what actually shipped. The recursive displacement loop is the gap between what companies say they want — productivity — and what they're actually building, which is structural unemployment they'll eventually have to answer for.
Six separate arguments, but I think they point at the same industry: economically extractive, running on language that's precise enough to sound rigorous and vague enough to escape accountability, and badly out of alignment with the people it keeps telling us it's here to help. None of that means the underlying technology is worthless — I work with it every day, and some of what it can do is genuinely new. But "the technology works" and "the industry is honest about how it's built and who it's built on" are two different claims, and I think we should stop letting the first one stand in for the second.