Building Easier AI
Teach or Build: When to Train Your Team vs. Buy It Done
The question owners actually ask about AI isn't whether it works. It's where the money goes: do I teach my own people to use this, or do I pay somebody who already knows how?
Before either, size the prize. The St. Louis Fed's FRED Blog put a number on it on August 27: the share of work hours saved by generative AI rose from 1.6% to 2.2% between the third quarter of 2024 and the second quarter of 2026, while the share of employed adults using it for work rose from 28.2% to 39.2%. Call it roughly an hour a week on the Fed's measure, and carry the Fed's caveat with it: people report their own saved time, so the figure is "inherently approximate."
An hour a week per person is real. It is also the entire pot. Every training day, retainer and subscription has to pay itself back out of an hour a week, and a pitch whose arithmetic needs more is a pitch to ask for the working.
Which door you go through has actually been tested, properly, with a control group. A randomized trial of 758 consultants included an arm that got AI plus training materials — the teach-your-team condition, run as an experiment. Training won on quality across the tasks the tool could do, and lost on speed. On the one task it couldn't do, the trained group was hurt worse than the untrained one.
Size the prize before you take the meeting
Put the Fed's hour against your own payroll first. Five office seats, one hour a week each, forty-eight working weeks: 240 hours. At a loaded $40 an hour that's about $9,600 a year — the ceiling, before a dollar is spent and before anything comes off for the time a human still spends checking output.
One more thing sits inside the Fed's figures: over those seven quarters usage climbed about eleven points while hours saved climbed six tenths of one. Adoption is running well ahead of measured return, and that is the market your quote was priced in.
Training helped where the tool worked, and hurt where it didn't
The best evidence on the teach-or-build fork is a pre-registered randomized trial run with BCG by researchers from Harvard, Wharton, MIT Sloan and Warwick, circulated in September 2023 and later published in Organization Science. 758 consultants, three arms: no AI, AI only, and AI plus instructional videos and documents on using it well. That third arm is the teach-your-team condition — an overview, not a training week, which bounds what it proves.
Inside what the authors call the frontier — eighteen tasks GPT-4 could do — the trained arm lifted quality 42.5% over the control, against 38% for AI alone. Speed went the other way: the untrained arm finished 27.6% faster, the trained arm 22.5%.
Then one task, built deliberately to sit outside the frontier. The control group, working with no AI at all, got it right 84.5% of the time. The AI groups came in at 70% and 60% — and the trained arm was the 60%, down 24 percentage points on the paper's regression, against 13 for the untrained. It also produced better-sounding wrong answers: recommendation quality up 25.1% against 17.9%, delivered 30% faster against 18%.
Training bought speed, polish and confidence. It did not teach anyone where the tool stops working. The trial ran on GPT-4 in spring 2023, and the frontier has moved since — but it moved, it didn't disappear. Your outside-the-frontier task is the one where the model has no way to know: the soil report that contradicts the plan, the inspector who fails a detail everybody else passes.
Buying beats building about two to one, in a report that half takes it back
MIT's NANDA initiative published the most-quoted figure in enterprise AI in July 2025 — 95% of organizations getting zero return on GenAI spend. Further into the same report sits the number relevant to this fork: external partnerships with learning-capable, customized tools "reached deployment ~67% of the time, compared to ~33% for internally built tools." Two to one for buying. MIT's own URL for that PDF now returns the lab's overview page; the copy quoted here is a mirror.
Now the report's own limitation, filed under the heading Important Limitation: "These success rate differences may reflect organizational capabilities rather than implementation approach alone." The sample is 52 interviews and 153 survey responses, and success is self-reported. So the most-cited statistic in enterprise AI carries a caveat that mostly withdraws it, and almost nobody quoting it has read that far. Read it against us, because we are one of the firms you would be buying from: the honest version is that the companies who buy well were better run before they bought anything.
Four of RAND's five failure causes aren't technical
RAND's 2024 report on why AI projects fail opens: "By some estimates, more than 80 percent of AI projects fail—twice the rate of failure for information technology projects that do not involve AI." Note the preface. RAND did not measure that figure, it cites someone else's — and in an article partly about how AI statistics get laundered, saying so is the job.
What RAND did do is interview 65 data scientists and engineers with at least five years building these systems, then sort the wreckage into five root causes: a misunderstood problem, inadequate data, chasing the newest technology instead of a user's problem, missing infrastructure, and applying AI to problems that are, in RAND's words, "too difficult for AI to solve."
Only the last is a limit of the technology. RAND ranks the first one first — misunderstandings about a project's intent and purpose "are the most common reasons for AI project failure" — and stating exactly which problem is being solved is the one thing you cannot buy from anybody.
Nobody can tell you what it costs
I went looking for a market rate for AI consulting the way you'd check a going rate for framing labor. Ten searches returned ten AI consultancies publishing their own rate cards: no government wage series, no independent survey, no trade-association benchmark. None of those numbers appear here — they are marketing collateral from the people quoting them, and that test includes us. We don't publish a rate card either.
So price it the way you price a sub instead: make the quote name the hours, name the seats it touches, and say what happens if it doesn't work.
You cannot trust your own read of it
Every figure above is self-reported, including the Fed's hour and MIT's 67%. The one number here that came off a stopwatch ran here last week: METR's 2025 trial of 16 experienced open-source developers, who forecast AI would make them 24% faster, believed afterwards it had made them 20% faster, and were measured 19% slower. In February 2026 METR said developers are likely faster with current tools, while calling its own follow-up data "only very weak evidence" for how much. The 39-point gap between what those developers felt and what the stopwatch said is the part that doesn't age.
What I'd actually do
Buy the tool. You are not building a model, and the trial above puts the quality gain mostly in the tool: 38% over the control for AI alone, 42.5% with the materials on top.
Write the problem definition yourself. RAND says that's where projects die, and no vendor can supply it: one page naming the task, who does it now, how often, and what "working" looks like in numbers.
Train the newest person, not the best one. Brynjolfsson, Li and Raymond's study of 5,179 support agents — which ran in last week's AI tip — found a 14% average lift, 34% for novices, close to nothing for the experienced.
Spend part of the training on where it stops. Teach failure modes, not just prompts, and name the three jobs in your business where the model cannot know the answer.
Measure with a stopwatch. Time ten of the task this week and ten in four weeks. Nobody's feelings about it are admissible.
Our own tools — PO Builder, Permit Tracker, Robyn, Jobs to Pay — are bought models running problem definitions we wrote ourselves, and the writing took longer than the building every time. No savings percentage gets printed for any of them until the audited before-and-after exists; the minute one does, I become the source I just spent two sections telling you to discount.
Teach or build stops being a fork the minute that one-page problem definition is written. Written properly, it tells you which door the work goes through — and how you'd know if it went through the wrong one.