Wisdom that outlasts the algorithm.

THE CURVE

You are the primary source.

The board asks the question every board is asking right now:
Are we behind on AI?

The executive answering it has no number about his own company. Not a bad number, none. There is no line anywhere in his systems that says what moved. So he does the responsible thing, the thing I would have done: he goes and finds the best outside evidence available. He brings a study.

The study is a survey.

It surveyed twelve hundred people in exactly his position, and asked them exactly the question his board just asked him. Not one of them had a number either.

He walks into that room holding his own uncertainty, formatted. With a sample size on it.

And it works. The board nods. The slide is credible. Sound methodology, real sample, disclaimed confidence interval. Nothing false added to the conversation.

Nothing in the room is firsthand, either.

(That meeting is a composite and it describes the shape of the scenario. I have watched several versions of this and this does not describe one room, one company, or one study.)

Now derive it bottom-up, because the whole issue is the shape.

There are two types of evidence about whether a technology actually works:
- Firsthand: somebody ran the thing, wrote down what happened before, wrote down what happened after, and dated it.
- Secondhand: somebody asked the people who ran it.

Both are legitimate. The second is built by asking around. It collects memories of the first and tells you where you sit in the crowd.

So what happens when the first kind was never written down?

You get a survey of impressions. That measures something real: what a population believes. It does not measure what happened, because nobody in the population recorded the ground truth. The aggregate tries to contain what none of the instances held. It inherits the absence within each node.

The debate over whether the AI spend is working is not wrong. It is hollow. And a hollow number is much harder to catch than a wrong one, because a wrong number leaves evidence of being wrong. A hollow one leaves no receipt. No error to find. No methodology to attack. Everyone did their job correctly all the way down, and at the bottom of the stack is a hole where somebody's operating record was supposed to be.

The debate over whether the AI spend is working is not wrong. It is hollow.



Here is what that looks like after it has run its full course.

In 1995 intelligent agents were much hyped. The promise? Software that would understand what you meant and go do it for you.

Michael Mullany went back through twenty years of the industry's technology charts and found that the hype died with Clippy (remember the persistent and omnipresent paperclip in MS Word). The hype returned two decades later as chatbots, with the same unsolved problem underneath: getting a machine to understand context.

Sit with the paperclip a second, because everyone remembers it as a joke. Almost nobody remembers it as a measurement problem.

By the visible signals of its moment, that deployment was a triumph. It was funded. The idea behind it topped every list of what mattered. And it had distribution almost no product in history has had — it appeared, unbidden, in front of anyone who typed the word "Dear" into a blank draft. Adoption was significant. Reach, exposure, sessions: every metric pointing to arrival.

History reveals: the capability was not there.

Here’s my read: Clippy did not have bad numbers. Clippy had perfect numbers. It won on adoption. But did it actually deliver? Here’s my read: the dashboard said triumph while the users said punchline. Only one of those made it into the metrics. Clippy did not have bad numbers. The numbers were answering the wrong question, and numbers answering the wrong question look exactly like numbers answering the right one.

By now you might be thinking of a specific failure with a name, so let’s go there. Circular reporting - where an article cites a source that cites the article, and a claim becomes true by repetition - propagates a false fact. And a false fact is eventually debunkable: someone finds the origin, the origin is empty, the claim collapses. The industry surveys circulating now propagate a hollow shell. An absence cannot be debunked. No origin to go find. No retraction to issue.

Then there is the part about scale.

Ordinary noise averages out. That is the whole reason we collect more data. Errors are random, they scatter in every direction, they cancel, and the signal emerges from the pile.

Now look at the errors in the evidence the AI spend produces about itself. A pilot that works becomes a case study; a pilot that dies becomes…silence. An adoption metric counts seats, never outcomes. A survey samples the executives who approved the budget. No dishonesty in that list. You only ever see what survived.

Run down the list and notice what every entry has in common. The errors all lean the same way: toward the money. And same-way errors do not cancel when you collect more of them. They stack, and the larger sample just makes the lean look like precision. I have not measured the size, and neither has anyone I can cite. Check the direction against your own org: which of your AI signals exist because somebody paid for them?

We’re deploying unprecedented investment into something nobody can currently measure. The reason nobody can measure it is not that we are early. The money went into capability and skipped the measurement. The spending did not create the blindness; it inherited it and made it expensive. The bill is now large enough that the blindness finally has a price.

That is the diagnosis. The turn is where it gets useful.

If firsthand evidence is genuinely this scarce, then whoever holds it has something rare.

Not virtuous. Rare.

One clean before-and-after, out of your own system. Dated, with a name attached. That carries more weight in a room than an external study of twelve hundred Technology executives — not because your sample is better. Because your sample is real, while theirs is remembered.

All the evidence is secondhand. Yours can be the only firsthand account.

──────────────────────────────


THE SIGNALS

Half-Baked.
I have been assuming this is a gap that closes. I cannot find a time when it was ever closed. My argument above rests on an unstated premise: that the missing record is a new problem, a byproduct of moving fast, something we are early for and will grow out of. So I went looking for the period when companies could say what their last big technology investment returned. I cannot find it. Not enterprise resource planning. Not cloud. Not mobile. I cannot locate the decade when a normal company could pull the file and tell you what the money delivered. I think AI is different because the spend is large enough to force the question to the top of the org. The outputs are legible enough to actually count: a draft produced, a claim adjudicated, a ticket closed. I cannot defend those yet. If you have ever watched a company reconstruct after the fact, from its own records, what a technology investment returned, reply and tell me. That one would move this from an assumption to a finding.

Hot Take.
Your AI strategy is a literature review. Read the evidence underpinning your strategy. The market sizing came from an analyst. The capability claims came from vendors. The competitive urgency came from an article…or from a peer at a networking event. Within you’ll find a list of what you have started. That tracks intent, not outcome. What you have is a well-researched document about what other people think, presented as a plan for what you will do. Observations of the company are virtually absent in every strategy document I’ve read this year. That is not for lack of rigor; every citation within is solid. It is a failure of provenance, and nobody checks provenance because a footnote and a measurement look identical once they land in a slide.

Confession.
I built an instrument this year that takes dated readings of the AI market, and it has three hard rules:
1) Never edit a placement in place. A new reading becomes a new dated file, so the old one survives.
2) Expired readings degrade to UNRATED rather than a low score, because a stale reading that still answers confidently is worse than no reading.
3) Absence is not zero. When the source cannot see something, the answer is "no data," never "bad score."
All three exist for one reason: to stop the instrument from placing a number where it finds a hole. I wrote these rules after catching my own instinct to always derive a number. Then I looked at my own project registry. No version control. One status field per project, overwritten in place. I couldn’t reconstruct what I believed about my own work six months ago, because I edited over it every time I learned something. I built a firsthand instrument to move beyond secondhand impressions.

──────────────────────────────


THE NEXUS

There is one number in the public record I keep coming back to, and it is not the one people quote: At least half of generative AI projects are projected to overrun their budgeted costs. The reason given is not model quality or talent. It is architectural choices and thin operational know-how. The same reporting notes that most organizations attempting to build custom models are expected to abandon the effort for three reasons: cost, complexity, and technical debt. (Simon Sharwood, The Register, 28 May 2026, reporting on Gartner's research.)

Notice why that particular number can exist at all. Spend is the one thing companies record firsthand: every invoice, every cloud bill, dated and named. The return side has no ledger. That is why the overrun can be projected and the payoff cannot: one half of the equation was written down and the other half never was.

You cannot overrun a budget you never committed. The spend arrived, then the pilot arrived, and the evidence never landed.

We are not early. Early would mean the money is still deciding. The money has decided. It is committed, being spent now, and it is already projected to exceed its own estimates. We are mid-flight with no instruments. If you are early you can wait for better information. If you are mid-flight you have to go get it.

An outside chart measures a category and cannot measure your instance. If companies cannot measure their own instances, the aggregate was compiled from companies that cannot measure themselves. This runs at every scale — the analyst's chart and your own board agenda are the same instrument; both log what is still being argued rather than what got settled.

I argued a few issues back that under real uncertainty you commit in the dark and take the bets you will not regret. I still believe that, and this is its other half. Committing in the dark is how you act when there is no reading. Producing a firsthand record creates a lantern to illuminate the next decision. Do only the first and you build a career of brave guesses. Do both and the guesses get shorter.

Where this is heading: the standard we are about to hold machine agents to — earn trust by producing evidence on demand, not by reputation — will come back around to land on the humans and institutions deploying them.

Pick the single strongest thing you currently believe about AI in your organization. Now trace where you learned it. Can you get to a measurement, or does the trail end at somebody's summary?

Reply with where your trail ended.

──────────────────────────────


THE MONDAY MOVE

10 minutes, no tools. Write down the three things you currently believe about AI in your organization: the three you would actually say out loud in a board meeting. Next to each one, write a single word for where you learned it: vendor, analyst, article, peer, or internal. One word, no explaining, no hedging the uncomfortable ones. You are not grading the beliefs. The finding is their relative weighting. If the last column comes back empty, the contrast reveals how much you have riding on outward looking beliefs.

THE ASYMMETRIC MOVE

This quarter. Produce exactly one firsthand reading. Not a dashboard, not a program. One number out of your own systems with five things attached: a before, an after, a date, and a name. And the number that would make you stop, written down in advance. Then hand it to someone whose budget does not depend on the answer and ask them to pressure test. If it survives, you own something almost nobody in your market owns. If it does not, you’ve only invested an afternoon instead of staking your entire strategy. Both outcomes pay. The only way to lose is to not have the number when someone asks.

THE DECADE MOVE

This decade. Become an institution that can answer questions about itself from its own records. Not a dashboard project. A habit: every significant investment gets an owner, a measure, and a date at the moment it is made, so that a year later the answer to "did it work" is a lookup instead of an argument. And because almost nobody does this, the institution that does ends the decade holding the scarcest thing in its market. You advance beyond your competitors still reading each other's summaries.

──────────────────────────────


THE COMPOUNDING ASSET

The Provenance Page. One page, two columns, five rows maximum. Left column: a belief you hold about AI in your organization. Right column: one word for where it came from. That is the whole artifact. Deliberately too small to be a project — five rows, because the discipline is choosing which five beliefs are actually driving decisions.

Run it with your leadership team and do not let anyone use the right column to defend the left. You’re looking for the conversation that starts when two people point to different provenance for the same belief.

──────────────────────────────


THE GROUNDING

This newsletter is called Signals from the Curve because there are two kinds of forecasting: the curve and the cliff. The cliff says everything changes at once. The curve says it is already changing, you just have to know where to look.

Where to look assumes somebody is looking — that somewhere there is a record of what actually happened, made by someone who was there. Most of what is circulating right now is not that. It is careful, competent, well-sampled work resting on a layer that was never written. We don’t need more analysis. What we’re lacking is people who actually checked.

Michael Mullany did the thing almost nobody bothers to do: he went back through twenty years of the industry's technology charts and graded more than two hundred technologies. He found the marker for having arrived was used only a handful of times in two decades — cloud computing, 3D printing, natural language search, electronic ink. He found major technologies that went mainstream while nobody marked it, and others that died while nobody was looking. Then he published the whole thing on LinkedIn in December 2016 and gave it away.

Ten years later the pattern still holds.

The opportunity is sitting in your own systems.

If something here changed how you are thinking, hit reply. I will respond.

— Chris


Chris Huber Reitz

Chief of AI & Strategy at Essential Innovations · Founder, Attainable AI · Adjunct Faculty, Columbia University