ENTRY 023 · ANSWER-ENGINES · By Answer Engineered Research
2,601vs183
Two Experts, One 8.2 Million Conversation Sample, and the Numbers 59,545 and 51
Microsoft's summary judgment filing prints both counts on one page. They measure different publishers at different thresholds, and the page never says so.
What the document actually is
Everything below was read out of the PDF itself, downloaded from CourtListener’s RECAP archive and extracted with pdftotext during this run. Fact numbers are the filing’s own.
| Field | Value |
|---|---|
| Document | Defendant Microsoft Corporation’s Rule 56.1 Statement of Undisputed Material Facts |
| Case | In Re: OpenAI, Inc. Copyright Infringement Litigation |
| Number | 1:25-md-03143 (SHS)(OTW), S.D.N.Y. |
| Filed | 4 September 2026 |
| Length | 65 pages, public redacted version |
| Docket entry | Document 1827 in the MDL; entry 1557 in NYT v. Microsoft, 1:23-cv-11195 |
One caveat governs this whole post. A Rule 56.1 statement is a party’s own list of facts it says are not in dispute. Microsoft wrote this. The plaintiffs have not conceded any of it, the court has not ruled on the motion, and every sentence quoted here is Microsoft’s characterisation of the plaintiffs’ expert work, not a finding. The docket is public and the filing itself is downloadable, so none of this has to be taken on anyone’s word.
Where the 8.2 million came from
Fact 29 describes the sample:
News Plaintiffs’ expert Dr. Goldstein analyzed logs of 8.2 million real-world conversations with Microsoft’s consumer version of Copilot produced by Microsoft. The 8.2 million conversations were selected from a population of approximately 1,134,799,717 total conversations between users and Microsoft’s consumer Copilot occurring between July 2023 to September 2024 based on matches with certain keyword terms in any text-containing field.
So the 8.2 million is not a random sample of Copilot usage. It is a keyword-filtered slice of roughly 1.13 billion conversations, filtered on terms the plaintiffs supplied. That matters for anyone tempted to convert either count into a rate: the denominator was chosen to be enriched for hits.
Why 59,545 and 51 are not the same measurement
Three differences, all of them in the filing, none of them in fact 79.
Scope. Goldstein’s 59,545 covers conversations matching keywords associated with the News Plaintiffs as a group. Wenger’s 51, per fact 60, is specific to Mother Jones. Different publisher sets.
Threshold. Fact 60 states Wenger’s rule:
Dr. Wenger treated as a “putative regurgitation” any Copilot that contained at least five 16-gram matches to an individual Mother Jones Asserted Work.
Five separate 16-word runs matching one article. Goldstein’s analysis, per fact 32, uses two different scores in parallel, a 4x4-gram continuity score and a 16-gram score, and reports both.
Exclusions. Wenger’s count is taken “after exclusion of public-domain text (e.g., the Constitution or the Bible).” Fact 79 does not mention that either.
A stricter threshold on a narrower corpus with an extra exclusion step returns a smaller number. That is not a contradiction, it is arithmetic. What is genuinely notable is that a document arguing these facts are undisputed sets the two figures side by side without a sentence explaining that they answer different questions.
The comparison that does hold
There is a like-for-like pair in the same filing, and it is more informative than the headline gap. For each publisher, Goldstein reports matches twice: once in Copilot’s grounding data, the material retrieved to inform an answer, and once in the outputs actually displayed to a user. Same expert, same score, same corpus. Fact 32(b) and (c), for The New York Times:
Goldstein found that of The New York Times’s 4,519,337 Asserted Works, 2,601 articles matched in the Copilot fields (the grounding data) using the 4x4 gram continuity score and 2,462 using the 16-gram score.
Goldstein also found that of The New York Times’s 4,519,337 Asserted Works, 183 New York Times articles matched in the Copilot fields (the outputs displayed to the user) using the continuity score and 104 using the 16-gram score.
The same pattern repeats across four publishers, all figures verbatim from facts 32, 38, 47 and 50:
| Publisher | Asserted works | Grounding matches (4x4) | Output matches (4x4) | Grounding (16-gram) | Output (16-gram) |
|---|---|---|---|---|---|
| The New York Times | 4,519,337 | 2,601 | 183 | 2,462 | 104 |
| Chicago Tribune | 1,205,350 | 1,200 | 82 | 1,159 | 45 |
| The Mercury News | 130,331 | 931 | 94 | 877 | 57 |
| The Denver Post | 89,269 | 1,005 | 123 | 969 | 78 |
Retrieval is not display. Between an order of magnitude and fifteen times as much matched text sits in the grounding layer as reaches a user, on every publisher in the table. Anyone building a measurement practice around “how much of my text does this thing reproduce” is measuring one of two very different quantities, and most tools do not say which.
Fact 32(d) adds one line worth quoting on its own: “Goldstein did not report a match to any entire article.” The same sentence appears for each of the four publishers.
The number nobody produced
Fact 146 is the one that matters most to this field, and it is a single sentence:
News Plaintiffs’ experts produced no analysis quantifying any loss of traffic, subscriptions, or revenue as a result of readers choosing to read a Copilot output with “regurgitated” content (as measured by Goldstein) as a substitute for reading any of their Asserted Works.
Facts 88 and 95 make the same point publisher by publisher. Of The New York Times, the filing says its experts “performed no analysis of that traffic data,” despite the paper collecting it. Of the Daily News, that its experts “did not analyze whether Copilot affected actual traffic volumes to its website.”
This is a defendant’s framing and should be read as one. Microsoft has an obvious interest in saying nobody measured the harm, and the absence of an expert analysis in one case is not evidence that no such effect exists. It is also possible the plaintiffs made a deliberate legal choice about what they needed to prove.
But strip the advocacy out and a fact remains. In the most heavily resourced piece of AI-and-publishers litigation currently running, with billions of conversation logs produced in discovery and named experts on both sides, the filing asserts that the traffic-substitution number this entire field argues about was not calculated by the people best placed to calculate it. Every “AI is costing publishers X percent” figure in circulation still comes from panel surveys and third-party traffic estimates, not from this.
The mechanism, in the filing’s own words
For anyone tracking how these systems are described in legal documents, fact 70 and fact 71 are the plainest definition of an answer engine yet entered into this docket:
Web grounding enhances a user’s experience; for example, by accessing information from Microsoft’s Bing search infrastructure, Copilot with web grounding improves knowledge discovery, synthesis, and location.
Copilot’s web grounding works by connecting the LLM with Microsoft Bing via Retrieval Augmented Generation, or RAG.
Facts 81 to 83 cover the control surface publishers were given. The filing dates it to a Microsoft Bing Blogs post of 22 September 2023, describing the repurposing of the NOARCHIVE and NOCACHE meta tags, with NOARCHIVE-tagged content that “will not be included in Bing Chat answers, not be linked to in the answers,” and NOCACHE signalling Copilot to exclude from grounding data everything on a page other than its title, URL and a snippet.
Fact 145 is Microsoft’s own experts: “consumers rarely use Copilot to obtain information about current events.” No rate is stated. This post does not repeat it as a measurement, because the filing does not give one.
What would change this reading
The plaintiffs’ opposing Rule 56.1 statement and their own experts’ reports would. Goldstein’s supplemental report and Wenger’s opening report are cited throughout this document as exhibits; the underlying reports are not in the public version, so the thresholds and exclusions here are known only through Microsoft’s summary of them. If the plaintiffs’ filings show the two counts were framed differently in the original reports, or that a traffic analysis exists that Microsoft has characterised away, that changes the picture.
A ruling on the motion would change more. Nothing here has been decided.
What would not change is the methodology point. Two experts, one dataset, 59,545 and 51, and a document that prints both without defining either. Until someone publishes a threshold, a corpus and an exclusion rule alongside a regurgitation figure, comparing any two such figures is a category error.
All figures in this post were read from the filing PDF directly on 8 September 2026. The docket and the document are linked above and both returned HTTP 200 that day.