Claims about AI and productivity swing between revolution and hype. The most useful evidence sits in between: controlled studies that gave some workers AI and others not, then measured the work. Their results are encouraging, specific, and in a few places sobering.

Customer support: the biggest gains went to newer workers

One of the largest field studies so far followed 5,172 customer support agents at a Fortune 500 software company as an AI assistant was rolled out. Published in The Quarterly Journal of Economics in 2025, it found that AI assistance "increases worker productivity, as measured by issues resolved per hour, by 15% on average."

The average hides the story. Less skilled and less experienced agents improved by about 30 percent, and agents in the lowest skill group by 36 percent. New hires climbed the learning curve faster. The most experienced, highest skilled agents saw "small gains in speed and small declines in quality."

Treated agents with two months of tenure perform just as well as untreated agents with more than six months of tenure.

Brynjolfsson, Li and Raymond, The Quarterly Journal of Economics, 2025

Professional writing: less time, better drafts

In an experiment published in Science in 2023, 453 college educated professionals completed writing tasks drawn from their own occupations. Those given ChatGPT finished faster and better: "The average time taken decreased by 40% and output quality rose by 18%." The tasks were short and done online in a single session, so they show what happens on a focused assignment, not across months of real work.

Consulting: the jagged frontier

A study of 758 Boston Consulting Group consultants using GPT-4, now published in Organization Science, coined a phrase that stuck: the jagged technological frontier. AI is excellent at some tasks and poor at others that look just as hard, and the line between them is not obvious.

On 18 tasks inside the frontier, consultants with AI completed 12.2 percent more tasks and finished 25.1 percent faster on average, with quality scores up roughly 30 to 34 percent against the control group. On a business problem chosen to fall outside it, the picture reversed. 84.5 percent of consultants without AI reached the correct answer, compared with 70.6 percent of those using GPT-4 and 60 percent of those given GPT-4 plus a prompting overview. Overall, the authors report that "subjects using AI were 19% less likely to produce correct solutions compared with those without AI."

Software: fast on a fresh task, slower in a familiar codebase

In GitHub's 2023 experiment, 95 developers built a small web server in JavaScript. Those with GitHub Copilot finished 55.8 percent faster, in 71.17 minutes on average against 160.89. The uncertainty was wide, and the researchers came from GitHub, Microsoft Research, and MIT Sloan, so the company had an interest in the result.

A 2025 trial by the research group METR tested the opposite setting: 16 experienced open source developers working 246 real tasks in large projects they already knew well. Allowing early 2025 AI tools "increases completion time by 19%," even though the developers believed the tools had made them about 20 percent faster. The study is small, but the gap between perceived and measured speed is a lesson on its own.

So why has the whole economy not sped up?

Adoption is still partial. As of May 3, 2026, 19.8 percent of U.S. businesses reported using AI in any business function in the Census Bureau's Business Trends and Outlook Survey, with the Information sector at 39.7 percent and Finance and Insurance at 33.9 percent. Census changed the question's wording in November 2025, so this figure cannot be compared directly with older readings. A Census study of firms found that 66 percent of AI users rely on it only to augment tasks, and that AI related employment decreases happened at just 2 percent of firms.

Estimates of the total effect remain modest. Economists at the Federal Reserve Bank of St. Louis estimate that workers' time savings from generative AI equal 1.6 percent of all U.S. work hours, and that generative AI "may have increased labor productivity by up to 1.3% since the introduction of ChatGPT," an estimate built on self reported data. In Denmark, researchers linking surveys to official records found no measurable effect on earnings or recorded hours two years after ChatGPT's launch, ruling out effects larger than 2 percent.

Reading the evidence for your own work

  • Start where tasks are well defined and the output is easy to check: drafts, support replies, routine code.
  • Expect the largest lift for people newer to a task, and use AI as a way to share what your best people know.
  • Check hardest where the task falls outside AI's strengths. Confident wrong answers cost accuracy.
  • Measure your own before and after. In the METR trial, the developers' own sense of speed pointed the wrong way.