"Today’s AI Is Great at Inconsequential Use Cases"
This continues to be true after I made this statement 2 years ago, which makes is just like any other piece of tech we have seen through out history.
Two years ago, in a forum thread I have long since lost, I made a claim that struck several people as needlessly pedantic. Generative AI, I argued, cannot be trusted to produce deterministic output without heavy guardrails built around it, and most businesses do not actually want a clever answer. They want a set of rules that returns the same result every time the same question is asked, so that the discount a customer is quoted on Tuesday is the discount they are quoted on Thursday, and a credit decision does not turn on which analyst happened to run it.
That point read as hair-splitting at the time. I keep watching it turn out to matter. A business leader asks ChatGPT, or Claude, or Gemini, a real question, one with money or direction riding on it, and takes the answer that comes back as a settled position. From that point on the leader has quietly left the conversation. What I am contending with is the output of a probabilistic system, one that will almost certainly hand back a different answer, sometimes a contradictory one, if the same question reaches it an hour later. The confidence of the prose hides the fact that nothing underneath it is fixed.
There is a gentler version of the same encounter, and it is the version that has made the technology popular. A partner at a mid-market firm asks a frontier model to draft a client memo. Ninety seconds later she has something clean, well argued, and close enough to final that she edits for twenty minutes and sends it. Here the non-determinism costs her nothing, because she was never after a repeatable answer. She wanted a first draft, and any of the ten slightly different drafts the model might have produced would have done the job. The memo is good. Her firm makes decisions the next morning in exactly the way it made them the morning before, and nothing about how work moves through the building has changed, because the model was asked to do a task the building already knew how to route.
Multiply that scene across every function and you have the current state of enterprise AI. The demonstrations are real. It drafts a passable brief in the time it takes to describe the assignment, and turns a sentence of plain instruction into working code, at a level that would have read as science fiction in 2021. It has also, in most organizations that have deployed it, changed almost nothing that reaches an operating statement. The two observations are usually held apart, the impressiveness filed as progress and the flatness filed as an implementation problem that scale will eventually fix. They belong together.
Ethan Mollick has a name for the reason. He calls the boundary of what these models do well the “jagged frontier”. Capability does not fall away smoothly as tasks get harder, the way a person’s would. The same model that passes a bar exam will miscount the letters in a short word, and one that produces a competent legal brief can, without warning, cite a case that never existed. The line separating reliable output from failure has the geometry of a coastline, jagged and unpredictable, and it cannot be seen from the outside. You cannot tell which side of the frontier a given task sits on until you have checked the output, at which point you have already done the work you were trying to hand off.
When the Boston Consulting Group put more than seven hundred of its consultants to work with GPT-4 in 2023, the ones who used it on tasks that sat inside the frontier produced work that was faster and graded higher. The ones who used it on a task designed to fall just outside the frontier did worse than colleagues who had no model at all, because the output was fluent enough to be trusted and wrong enough to lead them off a cliff. The lesson most people drew was that AI helps with some things and not others, which is true and close to useless. The uncomfortable lesson is about where the reliable region falls.
Plot the strengths honestly and they cluster in one place. The models handle knowledge, prose, mathematics, and the kind of bounded reasoning you can state in a paragraph and settle in a page. They struggle to hold state across a long task, to remember reliably what happened several steps earlier, to retrieve the right thing at the right moment without being handed it. Its strengths sit almost entirely in tasks that a single request can contain. The weakness surfaces in the connective tissue of work that has to run across people and across days, which is where most of what an organization does actually happens. The distribution is not random. The frontier stands highest where a task is a self-contained act, and it falls away the moment the work becomes a process that has to be carried from one hand to the next.
Which is where a paper from 1990 becomes useful. Paul David, an economic historian at Stanford, was trying to explain why the computers by then sitting on every desk had left no visible mark on the productivity figures, a puzzle Robert Solow had captured in the observation that you could see the computer age everywhere except in the statistics. David’s answer was to go back and look at electricity.
The dynamo was a general-purpose technology, the sort whose value arrives only after everything around it has been rebuilt. Thomas Edison patented a workable incandescent bulb in 1880. Twenty years on, electric motors still turned under five percent of the mechanical power in American factories, and the productivity gains engineers had confidently predicted were nowhere in the figures. The technology worked. Factory owners had simply installed it as a replacement part.
A nineteenth-century factory was organized around its power source. A single steam engine or water wheel turned a central shaft that ran the length of the building and drove every machine through belts and pulleys hung from the ceiling. Machines were placed according to how much power they drew and how close they had to sit to the shaft. Buildings were tall and narrow because a shaft could only run so far. When electricity arrived, the obvious move was to pull out the steam engine and drop in one large electric motor to turn the same shaft. Everything else stayed. The belts, the pulleys, the layout dictated by a driveline that no longer needed to exist.
This did almost nothing for productivity, and in the accounts it often looked worse, because the firm had spent capital on a new motor while keeping every dollar of the old transmission equipment in place, raising what it had sunk into the plant with no matching rise in what came out of it. The gains arrived decades later, when a generation of managers stopped treating the motor as a cleaner steam engine and started asking what a factory would look like if power were free to go anywhere. The answer was the unit drive: a small motor on each machine, drawing power through wire instead of belt. Once every machine carried its own motor the central shaft became unnecessary, and with it the constraints it had imposed. Factories could spread across a single floor. Machines could be arranged in the order the work actually flowed. Buildings could carry windows and proper ventilation, because the layout no longer bent around a driveline. Productivity did not move until the 1920s, some forty years after the enabling technology existed, and it moved because the work had been redesigned around the motor. Until that redesign, the motor was a tidier way to do what the steam engine had always done.
Anyone who lived through the arrival of the personal computer watched a compressed version of the same story. The first thing offices did with a PC was use it as a better typewriter. You typed a document on a screen instead of a page, fixed mistakes without correction fluid, and printed the result. It was faster, and it changed nothing about how information moved through a company, because the printed page still had to be carried down the hall to whoever needed it and filed in a cabinet once they were done with it. The machine became consequential when it stopped producing documents and became a connected node, wired first to the other machines in the building and the databases that held the real records, and eventually to the network that turned into the internet. The value was never in the box on the desk. It was in the system that grew up around the box, and that system took years to build and required people to change how they worked.
Today’s AI sits in the replacement-part phase. The partner drafting her memo in ninety seconds is using a frontier model the way a 1900 factory used its first electric motor and a 1985 office used its first PC, as a faster way to perform a task the surrounding system already knows how to handle. This is why the impressive use cases are so reliably inconsequential. Consequence lives in the parts of an organization the frontier is worst at reaching, the work that runs across functions, that has to remember its own history, that carries state from one week to the next. Those are the same capabilities the models are weakest at, and they belong to the same processes that would have to be torn up and rebuilt for the models to matter, which is slow, expensive, and cuts across whose job is whose in a way that drafting a memo never does. Those processes also need the same answer every time they run, a fixed rule rather than a fresh composition, and a probabilistic model cannot be trusted to hold a rule until it sits inside enough constraint and checking that the surrounding work has, in effect, been rebuilt to contain it.
It would be easy to read this as a failure of nerve, a charge that the people running these companies are too timid or too dim to see what is in front of them. That reading is wrong, and unfair. The factory owners of 1900 who kept their central shaft were not being stupid. Scrapping a serviceable plant built for steam, one with years of use left in it, to rebuild around a technology whose payoff was still speculative, would have been a poor decision at the time, and most of the firms that did reorganize waited until their old plants wore out anyway. Using AI trivially is, for most organizations right now, the rational move. The sunk cost this time is an entire architecture of who reports to whom, which system owns which record, and how work is handed off, one that runs well enough and that no single executive has the authority or the appetite to tear up on the strength of a return the pilots have not yet shown.
The analogy should not be pushed into a promise. Electricity took forty years in part because its reorganization was physical, a matter of pouring foundations and stringing wire, and capital that heavy moves slowly. The reorganization AI demands is mostly organizational, which could make it faster, since nobody has to pour concrete. It could as easily make it slower. Concrete, once poured, stays where you put it, whereas the habits, incentives, and quiet status hierarchies that AI would rearrange push back, and they do not wear out on a schedule that forces the question.
There is a further complication the electricity story does not carry. The dynamo, once invented, sat still and waited for factories to catch up. The frontier does not sit still. The distance a model travels in two years, from one that could barely sustain an argument to one that reasons through problems that stop graduate students, means the reliable region keeps widening while organizations are still finishing their first clumsy retrofit. That cuts both ways. It could mean the models grow capable enough to matter before anyone reorganizes at all, or it could mean firms spend a decade retrofitting against a target that has moved on by the time they reach it.
And there is a possibility the comfortable version of this argument leaves out. David was writing the history of a technology that paid off. Electricity did, in the end, remake the industrial world. Not every general-purpose technology does, and the ones that fail leave behind the same field of impressive, inconsequential demonstrations that a merely early one does, which makes the two hard to tell apart from the inside. The flatness of today’s returns fits a technology that will reshape everything once the work is redesigned. It also fits one that never will. Nothing in the evidence available now separates the two.
Which returns us to the memo. To say that today’s AI is superb at things that do not matter is to describe a phase rather than to lodge a complaint, and it carries a fairly precise diagnostic. The trivial use cases are trivial because they ask nothing of the organization, and the organization is where consequence gets manufactured. A company covered in clever, pointless applications of AI has not adopted AI. It has located every place a powerful new capability can be dropped into the existing machine without disturbing it, which is worth knowing, because those are the places where the machine was never the constraint. The work that decides whether the company wins is still organized around a driveline nobody has yet thought to pull out.


