Issue 1 · July 14 · Issue 2 · July 21 · Issue 3 · July 28 · Issue 4 · August 5 · Issue 5 · August 11 · Issue 6 · August 19 · Issue 7 · August 26 · Issue 8 · September 1 · Current issue
A weekly read on what the AI noise is hiding.
Two agent systems finished three tenths of a point apart on autonomous accuracy, and one of them needed about a third more of its cases sent for human review to hit the same reliability target. The first number is the one that gets compared. The same bill turns up when a model is swapped out from under an agent's memory, and when one teammate is traded for another — charged, both times, to whoever stayed. One item ran the other way, and it was the task with no other people in it.
Two systems three tenths of a point apart on autonomous accuracy, and one needed about a third more of its cases sent for human review to reach the same reliability target. Every stat sourced. The hype named as hype. One number worth acting on. Read the full brief below, free, no email required.
Read the first signal and the fifth together. Where the work needed a human in the loop, two systems tied on autonomous accuracy and one of them needed about a third more of its cases sent for human review before it could be trusted. Where the work needed nobody at all — a Lean proof a checker either accepts or does not — swapping the model alone took the result from zero to something. That is the same finding twice: the frontier moved, and the coordination did not move with it.
Two agent systems separated by 0.3 points of autonomous accuracy, 72.8% against 72.5%, needed 39.2% against 29.6% of cases sent to a human to reach the same 76% reliability target. That is one clinical-audit case study, sixteen agent systems, 750 cases, and no firm’s P&L was measured.
Population and screen. Sixteen agent systems — not firms and not people. Each had to emit an outcome on a clinical-audit case and be scored against a 76% reliability target under a selected oversight policy. One workflow class, not a cross-firm sample. No human reviewer was employed and no salary attached: the cost the study minimises is a modeled operating cost, not a staffing line. The paper is a Scale AI paper. Its title block carries the Scale AI logo and numbered affiliations: Scale AI, with co-authors at the University of California, Santa Cruz and at Vanderbilt University Medical Center, and correspondence to two scale.com addresses. Read it as vendor-adjacent research: the lead affiliation sells enterprise AI evaluation.
Correction, 11 September 2026. As published on 8 September this section read: “The paper does not state its authors’ affiliations, so nothing here attributes it to a lab, and no reader should.” That was wrong. The arXiv abstract page states no affiliation, and this issue checked only the abstract page; the PDF’s first page names Scale AI in the logo, in footnote 1 and in both correspondence addresses. The error was ours, it was a claim about provenance, and the affiliation it missed is the one a reader most needed.
Ranked on autonomous accuracy, the two tie. Ranked on how much of the work still has to reach a person, one of them costs about a third more. The first ranking is the one that gets printed. The second is the one you live with. Source: READY or Not: Reliable Enterprise Agent Deployment, a September arXiv preprint, arXiv:2609.02095, submitted 2 September 2026. Not peer reviewed.
Four memory-store designs, 48 synthetic histories, and two open-weight models under 10B parameters. Synthetic histories, not any firm’s records; small open-weight models, not the frontier models the upgrade conversation is actually about.
With that stated: the fixed-schema knowledge graph barely noticed the swap, moving by no more than 0.0020 in either direction. The compressed free-text notes store moved asymmetrically and by a wide margin — it gained 9.91 percentage points migrating one way and lost 13.28 migrating the other. The difference is the shape the memory was written in, and whether that shape survives being read by something else. Repairing the free-text store in place did not reach the study’s recovery target in any of the 48 cases; going back to the raw history and rebuilding did, in most.
This is the item most likely to be over-read this week. It shows a mechanism at small scale. Anyone who turns ±0.0020 into a migration business case is inventing a denominator the study never had. Source: an arXiv preprint on memory portability, arXiv:2609.05339, Goyal and Ray, submitted 4 September 2026. Not peer reviewed.
Eight teams per setting, each formed independently from one base model, each agent keeping a private notebook across ten formation episodes. Role-matched agents were then traded between teams. Game environments, LLM agents only — no humans, no firm, no payroll.
The study compares each swap against a control that reproduces the disruption of a roster change without changing who occupies the seat. Without it the finding is “change is disruptive,” which nobody needs a paper for. With it, the cost is attributable to who is in the seat rather than to the fact that something was disturbed.
Against that placebo, a swap costs little in task score and raises communication spent per unit of progress by 16 to 63 percent. In one environment, when the agenda-setting agent is replaced, most of the additional communication comes from the agent that stayed. The bill went to whoever did not move. Source: an arXiv preprint on interchangeability in LLM agent teams, arXiv:2609.05279, Gao, Yu, Deng, Li and Wang, submitted 4 September 2026. Not peer reviewed.
CentaurBench is framing rather than news; its issue date is August 2026. Its population and screen matter more than its result. It is model-on-model: an assistant model writes assistance text for a standardized lower-capacity worker model across seven economically grounded tasks, scored by blind pairwise comparison by a panel of LLM judges over ten runs. No human being was assisted in this study.
In summary rather than in the paper’s words: the automation winner loses on augmentation in five of the seven tasks, the unaided worker model outranks every assisted condition on three of them, and only one model’s guidance beats no guidance on average.
It is the same swap problem from the other side. The capability transfers; the fit with whatever is already in the seat has to be found again. Source: CentaurBench, NBER Working Paper 35663, issue date August 2026, posted 31 August 2026. Working paper, not peer reviewed.
Twenty-one interviews inside one large software-services company, roughly a year after its AI-adoption launch. Non-probability, one firm, one industry. No percentage, rate or magnitude from this study appears anywhere in this issue, because none of them would survive that screen.
It is here for the shape. The researchers name five classes of cost carried by the people in the seats: accountability anxiety, craft-identity disruption, meaning and satisfaction erosion, cognitive and workload intensification, and uncertainty distress.
This study is about adopting AI, not about swapping teammates, and it is not evidence for the interchangeability preprint’s finding. What it does is name the kind of cost a task score does not capture, in the words of the people carrying it. Source: an arXiv preprint reporting a qualitative case study, arXiv:2609.03456, Alami, Paja and Tiwari, submitted 3 September 2026. Not peer reviewed.
Sixty-eight Erdős problems, roughly a tenth of the 652 catalogued as unsolved as of August 2026, selected as significant and hard, and formalized in Lean. One official attempt per problem, capped at $300 and 72 hours, and a solve counts only if the model emits a Lean proof or disproof that passes verification. No partial credit, no human in the loop, no judge panel. One model solved two; the rest solved none.
Two limits hold here and neither works without the other. The measurement limit: two solves out of sixty-eight is a floor rather than a capability level, because of how it was produced, and a larger budget is a different measurement, not a better model. The organizational limit: formally verifiable mathematics is the least organizationally entangled work that exists. No handoff, no review board, no colleague who has to be re-briefed. The only variable that moved was the model.
That is why it runs here: where work is fully specified, machine-verifiable and needs no other people, capability alone was the binding constraint in this benchmark, and this week it visibly loosened. It locates the boundary rather than refuting the rest of this issue. Source: Epoch AI, Announcing FrontierMath Erdős, 1 September 2026.
IDC/Microsoft, “The Agent Production Gap,” covered 6 September 2026: a 171% global ROI claim for agents that reach production, and 192% in the US. Vendor-run research, and the coverage does not state the sample behind the 2026 figures. There is no reachable denominator, so no way to say who was asked or what qualified them. Named here so you recognize it when it arrives. This brief did not open the underlying study.
McKinsey, The state of AI in 2026, recirculated on 6 September. The survey was fielded 4 May to 8 June 2026 and published 25 August 2026. A recirculation date is not a publication date. Standing disclosure: mckinsey.com is not reachable from the machine this brief is written on, and no McKinsey page was opened for this issue.
The claim that the Minneapolis District attributed productivity gains in roughly equal measure to AI and non-AI causes. What is refused is the attribution, not the document. The Beige Book edition released 2 September 2026 is the August 2026 edition, not the September one. It was opened for this issue — cover and front matter, the National Summary in full, and the entire Minneapolis District section. The claim is not in the document, and the word “AI” does not appear in the Minneapolis section at all. The likely origin: Minneapolis is the Reserve Bank that prepared this edition, stated in the same page-1 footnote as the collection date, and a preparing bank is not a finding. The other eleven District sections were not read, so the finding is that the claim as attributed is absent, not that no District mentions AI.
“Eurostat’s standing enterprise AI figure is 19.95 percent.” This came back from a machine sweep. Eurostat was not opened, the figure is unverified here, and an unverified two-decimal number is not improved by its decimals.
No nationally representative firm-adoption number was published between 1 and 7 September 2026. That was established two ways rather than assumed. First, the Census Bureau’s Business Trends and Outlook Survey national data file was downloaded and parsed for this issue: the newest cycle in it is the one this brief already reported, reference period 27 July to 9 August 2026, release CB26-TPS.48. The next cycle publishes 10 September, two days after this issue ships. Second, a window sweep of the other official statistical agencies returned empty. Searched and not found, which is not proof of absence.
So the strongest firm-level number available to you this week is still last week’s, and it should be cited as BTOS, reference period 27 July to 9 August 2026, release CB26-TPS.48 — never as the bare URL, because that file is replaced on 10 September and the address will keep resolving to different numbers.
The one official-sector item that did land in this window is not a statistic. The Federal Reserve Beige Book, August 2026 edition, released 2 September 2026, says one thing about AI in its National Summary, in one sentence: “Districts also reported both positive and negative effects of artificial intelligence on labor demand.” Population and screen: anecdotal contacts across twelve Districts. No sample size is stated. Contacts are not selected at random and this is not a probability sample. The information in it was collected on or before 24 August 2026.
It is one qualitative sentence and it carries the weight of one. Every vendor number in this issue points one way; the one official-sector source in the window reports both directions at once, with no magnitude, no direction and no count of Districts beyond twelve.
arXiv:2609.02095, READY or Not. Abstract page opened 7 September 2026; PDF full text opened 11 September 2026. All four figures and the sixteen systems over 750 cases were read off the abstract page. Affiliations: Scale AI (lead), University of California, Santa Cruz, and Vanderbilt University Medical Center, read off the PDF title block. Preprint, not peer reviewed. Correction, 11 September 2026: this row previously read “Author affiliations are not stated on the page, so this issue attributes it to no lab.” True of the abstract page, false of the paper, and the row did not say which of the two had been opened. The abstract page alone is no longer sufficient for any affiliation claim.
arXiv:2609.05339, memory portability. Opened, 7 September 2026. All figures, the 48 synthetic histories and the sub-10B model class read off the page. Preprint, not peer reviewed.
arXiv:2609.05279, interchangeability. Opened, 7 September 2026. The abstract was transcribed verbatim into the research notes and the placebo control is quoted from it. Preprint, not peer reviewed.
arXiv:2609.03456, psychological costs. Opened, 7 September 2026. Qualitative case study; the five cost classes are the paper’s own labels and no figure was lifted from it. Preprint, not peer reviewed.
NBER Working Paper 35663, CentaurBench. Opened, 7 September 2026. Abstract read verbatim; the ten-run replication and the blind pairwise LLM-judge scoring were added from the page. Working paper, not peer reviewed, August 2026 issue date.
Epoch AI, FrontierMath Erdős. Opened, 7 September 2026. Protocol, per-model solve counts and per-solve cost and time all confirmed on the page.
U.S. Census Bureau, BTOS National data file. Downloaded and parsed, 7 September 2026. It establishes the negative finding: the newest cycle in the file is the one reported last week.
Federal Reserve Beige Book, released 2 September 2026. OPENED, 7 September 2026. Cover and front matter, the National Summary in full, and the entire Minneapolis District section were read. The other eleven District sections were not, and that is the stated limit of the check. The August 2026 edition; collection date and preparing bank taken from the page-1 footnote.
NBER Working Paper 35275, Writing Code vs. Shipping Code. OPENED, 7 September 2026. Title, authors, the May 2026 issue date and the population of more than 100,000 GitHub developers were read off the page. No figure from it appears in this issue, because a four-month-old paper does not become this week’s news by being linked again. The MIT Sloan write-up of it was NOT OPENED, and this brief makes no claim about what that write-up put back into circulation.
mckinsey.com. NOT OPENED. Unreachable from this machine. No McKinsey page was loaded for this issue and no McKinsey figure is printed in it.
IDC/Microsoft, “The Agent Production Gap.” NOT OPENED. Only the 6 September coverage was seen. Named as a foil in the refusals and used as evidence nowhere.
Eurostat. NOT OPENED. The sweep’s 19.95 percent is unverified and unused.
NBER Working Paper 35684, on organization capital. NOT OPENED. Association rather than experiment, carrying no magnitude this brief could print. Named so you know it was seen and passed over.
Method. One machine sweep of the 1 to 7 September window was run for coverage and its returns were treated as leads rather than findings. Every item above that carries a figure was opened at its own URL; two of the items in this issue were not in the sweep’s results at all and came from querying the arXiv API by submission date. The sweep’s one independently checkable claim — that no official statistical product landed in the window — was confirmed against the Census file directly rather than accepted.
Standing caveat on the preprints. Four of the six sources used above are arXiv preprints. They have not been peer reviewed, their figures may change between versions, and none of them measured a firm’s costs. They are printed here because their populations and tests are stated plainly enough to argue with, which is more than most of what will reach you this week.
One weekly read on what the AI noise is hiding. Every stat sourced, the hype named as hype, one number worth acting on. No spam, unsubscribe anytime.
Start by measuring the coordination tax you’re paying right now. Free, a few minutes.
Next Tuesday: another scan, same rule. Cut the hype. Show you what the noise is hiding.