Skip to content
← Writing · Research

Do Businesses That Use AI Report Better Results?

Public Census data cannot show whether businesses that use AI report better results: the within-group comparisons are uninterpretable and the overall verdict is inconclusive.

Quick answer

These public Census data cannot say. In 22 two-week periods, the share of businesses reporting AI use moves too little beyond sampling noise for a comparison over time to show a relationship, so the verdict: inconclusive. Across states there is no sign that more AI use goes with better self-reported results, if anything the opposite; that compares states, not businesses.

Key takeaways

  • Verdict: inconclusive. The data could not tell whether more reported AI use goes with better reported revenue or performance, which is different from finding that it does not.
  • Across 52 units (50 states, DC and Puerto Rico), those with more reported AI use report lower results (r = -0.49 for the revenues index, r = -0.54 for the performance index); see the caveat in the states section.
  • Sector and metro comparisons are inconclusive too: with 18 sectors and 25 metros, only a large correlation could have shown up.
  • A claim that companies using AI do better is a comparison of groups; check how it was measured before relying on it.

Do businesses that use AI report better results?

Public data cannot tell you whether businesses that use AI do better. The Census Business Trends and Outlook Survey publishes, for each sector, state and big metro, the share of businesses using AI and the share reporting rising revenue or good performance, but the AI-use share moves so little beyond sampling noise from one two-week period to the next that no comparison over time could show a relationship. Noise here means the random wobble you get from surveying only some businesses each fortnight.

Here is what was measured. The AI figure is the share of businesses answering "Yes" to a question about using AI in any of their business functions; it does not ask about AI agents as such. The results figures are two Census indexes: the revenues index, which summarizes whether businesses said their revenues rose, fell or stayed the same compared with the two weeks before, and the performance index, which summarizes how businesses rated their own current performance, from poor to excellent. No public dataset measures "AI agents" or "value", so this article cannot speak to either. The verdict is inconclusive: the data could not tell, which is not the same as finding no relationship.

What did we measure, and how?

We compared groups of businesses, never single businesses. We used 18 sectors, 52 state units (50 states, DC and Puerto Rico) and the 25 largest metros, over 22 periods from 202524 to 202619. The files were retrieved on 2026-10-06.

We started from the assumption that outside the tech sectors, reported AI use would not go with better reported results. The data could neither support nor contradict that assumption. Before we fixed our method we had already seen which sectors report more AI use, so that part was not a blind test.

You can compare groups with each other, or ask whether, within one group, periods with more AI use are periods with higher results. The second comes closer to your question, and the data cannot answer it. The reason is reliability, which is the share of a figure's movement that is real change rather than sampling noise. For the AI-use share within a group over time, the registered reliability is 0.01 for states, 0.04 for metros and 0.12 for sectors. In plain words, almost all of the period-to-period movement in a state's AI-use share is sampling noise.

Noise of that size pulls any correlation toward zero. (A correlation, written r, is a score from -1 to 1 for how closely two figures move together.) A near-zero correlation is therefore expected from noise alone, and a small one cannot be read as a relationship in either direction. For example, the state revenues result (r = 0.09) and the metro performance result (r = 0.16) must not be read as relationships in either direction. The study's pre-set rule called all eight within-group tests too noisy to read (its formal label is "uninterpretable").

You can check every figure in the Census files: the BTOS data page links the state, sector and top 25 metro workbooks (the AI-use share is the "Yes" answer to question 7; the indexes are in the "Index Estimates" sheet), and the questionnaire and methodology explain the wording and how the indexes are built.

Technical detail (skippable): the Holm-adjusted p-values for the within-group slopes, in a family of 16 tests, were p = 0.184 (states, revenues), p = 0.398 (metros, revenues), p = 0.417 (metros, performance) and p = 1.000 for both sector tests and for states on performance. The smallest within-group correlations the design could have detected, using an effective sample size adjusted for the clustering of periods within each group, were minimum detectable r = 0.60 for sectors, minimum detectable r = 0.25 for states and minimum detectable r = 0.39 for metros. These assume the survey estimates have no sampling noise, but they do have it, so these figures understate what could not be seen. For the between-group comparisons below, with 18 sectors and 25 metros, the smallest correlations that could have been detected were minimum detectable r = 0.62 and minimum detectable r = 0.54; in plain words, only a large correlation could have shown up. The 16 non-tech sectors gave the same uninterpretable result as all 18.

What do states, sectors and metros show when compared with each other?

Across states, higher AI use goes with lower self-reported results, but that compares states, not businesses. Across states, there is no sign that more AI use goes with better self-reported results; if anything the opposite (r = -0.49 for the revenues index, 95% CI -0.66 to -0.29; r = -0.54 for the current-performance index, 95% CI -0.72 to -0.31; 52 states, DC and Puerto Rico), and this comparison cannot separate AI from everything else that differs between states (industry mix, business size, costs, how businesses rate themselves). The interval is the range of values that the data leave plausible. This comparison was not designed to detect harm, and the within-state data cannot test it.

Sectors and metros are inconclusive. The table gives each correlation between a group's average AI-use share and its average index.

Groups compared Units Revenues index Performance index Reading
States (50, DC, Puerto Rico) 52 r = -0.49 r = -0.54 Opposite direction; see the caveat above
Sectors 18 r = 0.33 r = 0.24 Inconclusive
Top 25 metros 25 r = -0.09 r = -0.06 Inconclusive

For sectors, the intervals around both correlations are wide and include zero. For metros, both correlations are near zero. With only 18 sectors and 25 metros, these comparisons cannot exclude a moderate correlation of either sign. Do not read the sector pattern as saying that tech sectors do better.

Exploratory: secondary checks, which were not corrected for multiple tests, gave small within-metro correlations that were mostly positive; with this level of noise they cannot be told apart from sampling noise shared by the two answers, so we do not interpret them.

What can this data not show?

It cannot show which way any link runs. The reverse direction is also plausible: businesses that are doing well may be the ones that adopt AI. Both figures are self-reported survey estimates from the same respondents, and the two-week periods overlap, so they are not independent observations.

Who is not in the sample matters too:

  • Businesses with no employees. The survey covers employer businesses only.
  • Multi-state firms, which are left out of the state figures.
  • Management of companies (sector 55), which was dropped because AI use was published in only 5 of the 22 periods.
  • Metros outside the top 25, and anyone who did not respond.

The AI-use figures themselves have gaps: the Census suppresses cells it considers too imprecise, and the study never filled them in. Only the current question wording is used, which is why the window starts at 202524. California appears here only as one state among 52 units and four of the 25 metros; nothing in this article is a California claim.

For the AI-use figures themselves, see Is AI only used by tech companies? (tech versus other sectors), What do businesses use AI for? (uses outside tech), AI use by city (metro shares) and AI use by business size and industry (employee counts).

How should you read a claim that AI users do better?

Treat it as a comparison of groups, and ask three questions.

  1. Is it a comparison of groups or of the same businesses before and after?
  2. Does it allow for the reverse direction described above, or ignore it?
  3. How was AI use measured, and how precisely? If the measure moves mostly through sampling noise, the comparison cannot tell you much.

Two things make a claim more informative. It follows the same businesses before and after a change, and it compares businesses in the same industry and size band. The public tables used here give group figures by sector, state and metro, not figures for the same businesses over time or within an industry and size band.

Analysis by Deivy Hernandez, from public data files.

FAQ

Does an inconclusive result mean AI does not work?

No. Inconclusive means the data could not tell, in either direction. The study was designed to detect a relationship and could not, because the measure of AI use is too noisy over time.

Why not compare each business that uses AI with each that does not?

The public Census files give figures for groups (sectors, states and large metros), not for individual businesses. Any comparison from public data is therefore between groups.

Is the revenues index the same as profit or sales growth?

No. It summarizes whether businesses said their revenues rose, fell or stayed the same over the last two weeks. The performance index summarizes a self-rating from poor to excellent.

What would give a clearer answer?

A design that measures AI use precisely for each business, rather than as a share of respondents in a cell. The public files do not offer that.

Sources
  1. U.S. Census Bureau — Business Trends and Outlook Survey, data page
  2. U.S. Census Bureau — BTOS, State data file
  3. U.S. Census Bureau — BTOS, Sector data file
  4. U.S. Census Bureau — BTOS, Top 25 metropolitan areas data file
  5. U.S. Census Bureau — BTOS core questionnaire
  6. U.S. Census Bureau — BTOS methodology

Get the next guide by email

Practical AI and automation guides for people who ship. No theory.

NO SPAM · UNSUBSCRIBE ANYTIME