Building Easier AI

The Back-Office Jobs AI Is Good At, and the Ones It Isn't

Construction sits at about 8.9% AI adoption. Information services run 39.3%, professional services 30.3%. What makes those three figures worth more than the adoption surveys you have been shown all year is where they came from: the JPMorganChase Institute pulled them off de-identified bank transaction data covering 4.6 million small businesses (April 2026), and identified adopters by actual payments to more than 500 named AI tools. Nobody was asked to estimate anything about themselves.

The usual reading of that gap is that construction is behind, and the usual next sentence is that field work is harder to automate than office work. Both are unhelpful. Every business that builds something runs a back office — invoices coded, POs cut, schedules maintained, permits chased, client emails triaged — and the back office is where almost every trade, field-service and manufacturing firm starts. The question was never office versus field.

One randomized trial gets closer than anything else published — Microsoft's own, on its own product: over 6,000 workers at 56 firms, six months, measured by telemetry rather than by asking people how it felt. It looked at three back-office tasks and got three different answers. One clearly worked. One worked about half as well as the headline says. One didn't move at all on average — while moving enormously in both directions underneath.

The line it draws isn't office against field. It's whether the output can be checked.

Three tasks, three answers

Dillon, Jaffe, Peng and Cambon randomized Microsoft 365 Copilot licenses within 56 firms and watched for six months (arXiv 2504.11443, April 2025; the firms entered in late 2023 and early 2024). Say the obvious thing first: all four authors work at Microsoft and the product tested is Microsoft's. It is still the largest randomized trial in the area and the outcomes are telemetry, not self-report — but the vendor ran it.

Email worked. Against a control group spending 2.8 hours a week reading email, workers who actually used the tool read 23.58 fewer messages and 31.16 fewer minutes a week, an 18% cut, and replies went out about 10% faster. They did not simply reply less — licensees replied to 13.3 conversations a week against the control's 13.5. The authors read that as triage, not inattention.

Documents worked, at half the size of the number in circulation. Time to finish a Word document fell 12.4% against a control mean of 7.5 days — but that is the effect on people who used it. The effect of buying someone a license is −6.4%. Both sit in the same table, and nearly 60% of licenses handed out went largely unused, which is why they separate at all.

Meetings didn't move, and the average hides it. Total meeting time shifted +2.38 minutes a week, not statistically significant. Underneath the null: six firms where it rose a significant 26.52 minutes a week, against others where it fell a significant 25.19. Same tool, same task, opposite signs — and the direction wasn't predicted by how much a firm used Teams beforehand. The authors conclude that "other organizational factors determine" whether workers react.

The meetings result is the finding. Email and documents produce something a person can check: the reply went, the draft is done. A meeting's outcome depends on how your organization behaves, and no license changes that.

What it takes off you is the rough draft

Noy and Zhang ran a pre-registered experiment with 444 college-educated professionals on mid-level professional writing (MIT working paper, March 2023; published in Science that July). Time taken fell by 0.8 standard deviations and quality rose by 0.4 — standard deviations, not the tidy percentages that saturate the secondary coverage. Low-ability workers gained most.

Their sentence on mechanism answers this article's title. The tool "mostly substitutes for worker effort rather than complementing worker skills, and restructures tasks towards idea-generation and editing and away from rough-drafting."

Rough-drafting is what it takes off you. Work that is mostly deciding, it doesn't.

The jobs it isn't good at

Scheduling and dispatch are what vendors market hardest to construction and field-service firms, and they carry the clearest published failure curve of anything here. Valmeekam, Stechly and Kambhampati ran o1-preview through PlanBench (arXiv 2409.13373, submitted September 2024). Short Blocksworld planning problems: 97.8%. The 110 problems needing at least 20 dependent steps: 23.63%. Scramble the vocabulary so it can't lean on familiar patterns and the set drops to 37.3%. A classical planner solved all 600 correctly, averaging 0.265 seconds each.

A construction schedule is both hard cases at once — hundreds of dependent steps, in your own idiosyncratic vocabulary. Date that result September 2024 and grant that models have moved. The direction is the durable part: accuracy falls as dependent steps accumulate, and falls again when the words are unfamiliar.

Unchecked document production is the other one, and it carries a price tag. Deloitte Australia delivered a 237-page report to the Australian government on a $290,000 contract containing fabricated academic citations and a fabricated quote attributed to a federal court judge; it disclosed Azure OpenAI use in a corrected version dated 26 September 2025 and agreed on 7 October to refund the final instalment. The model didn't fail there. A 237-page deliverable went out the door without anyone reading the citations.

Damien Charlotin's AI Hallucination Cases database counts how often courts catch that — 2,016 cases as of its own stamp, last updated 4 September 2026, and moving daily. Every count quoted in a blog post is stale by the time you read it, this one included.

Nobody has measured your industry

Go looking for measured AI results in field service, the trades or accounts payable and every road ends at a vendor. On contractors the best available read is ServiceTitan's 2026 Commercial Specialty Contractor Industry Report, a survey of 1,000-plus commercial construction leaders — the same source we cited two weeks ago, for the same reason: it sells software to the people it surveyed. It's what exists.

Accounts payable is worse. The touchless-processing and cost-per-invoice benchmarks every AP vendor quotes trace back to one compilation drawn from Ardent Partners' State of ePayables 2024, built on 212 self-reporting AP professionals and handed out as gated collateral by the automation vendors selling against it. For these functions there is no government series and no benchmark with a published sample frame.

Even the good source launders one: the JPMorganChase report carries a borrowed line that "a 2025 survey showed that more than 80 percent of small businesses using AI report productivity gains" — exactly the kind of statistic the rest of it makes unnecessary.

The part that argues against us

We sell this work, so read the Copilot trial's adoption figure against us first. Nearly 60% of purchased licenses went largely unused. That is the most likely outcome of any software purchase, ours included, and it is why the honest version of the documents result is 6.4% and not 12.4%.

No savings percentage for PO Builder, Robyn, Permit Tracker or Jobs to Pay appears anywhere above. They run in our own companies every week and we keep paying for them, which is evidence of a kind. The audited before-and-after that would let me print a number doesn't exist yet, and printing one anyway would make me the source this article spends a section telling you to discount.

Sorting your own back office

One question sorts the list, and it is not whether the work happens in an office. Can somebody check the output in under a minute against something that already exists — the contract, the budget line, the plan set, the schedule, the sent folder?

Draft the client update: checkable. Triage a week of messages into change orders and decisions: checkable. Reconcile a selection sheet against the allowances: checkable. Build the schedule: no — your scheduling software's dependency engine already does that, correctly, in a fraction of a second. Send anything unread: no, at any price.

Entry costs almost nothing, which is the last useful thing in the JPMorganChase data: in 2025 the 25th percentile of adopting firms spent about $20 a month. It counts paid services only, so treat 8.9% as a floor. The money was never the constraint. The hour somebody spends checking the output is — and either that hour is in somebody's week, or the job doesn't go on the list.

Building Easier Weekly

One of these lands in your inbox every week.

Every week: what moved in the market, design trends in new construction and remodeling, hot markets, and one AI tip you can use the same afternoon.

Read past issues  ·  Get it by email