The FBI's Review of the Epstein Files: A Critical Review
In March 2025 the FBI reviewed roughly 100,000 Epstein-related records the way reviews were run in 1985 — by hand, around the clock, on end-of-life software. Forty years of eDiscovery research predicted the result. This is a sourced analysis of how that review was run, and why it failed.
This review draws only on the public record — Congressional correspondence, internal FBI emails released to reporters, court filings, and the government's own releases. Four findings recur across those sources:
Finding 1 · A 1985 methodology
Linear, page-by-page human review at scale — the exact approach empirical research discredited four decades ago.
Finding 2 · End-of-life tooling
Review software bought in 2011 (~$11.9M, per procurement records) — then bought again in 2021–22 ($4M more) instead of being replaced.
Finding 3 · No measurement
No stated recall target, sampling protocol, or quality-control metric — no way to know what the review missed.
Finding 4 · Capabilities left unused
Machine translation, handwriting OCR, and entity resolution exist in unclassified form — and were absent from the process.
The science the review ignored
The idea that armies of human reviewers are the "careful" way to work through a document population was tested — and disproven — four decades ago, in a study built on litigation over a Bay Area Rapid Transit (BART) accident. Every rigorous test since has widened the gap. Any critical review of the FBI's process has to start here, because none of this was obscure: it is the settled, judicially recognized baseline of the field.
1985 — The BART study: lawyers found 20%, believed they'd found 75%
David Blair and M.E. Maron studied the discovery database from the BART litigation — roughly 40,000 documents, about 350,000 pages. The attorneys running the review were confident they had retrieved more than 75% of the relevant documents. Measured against the actual population, they had found about 20%. Human judgment about review completeness was off by a factor of nearly four — and the reviewers never knew.1
2011 — Technology-assisted review formally outperforms manual review
Grossman & Cormack, analyzing TREC 2009 Legal Track data in the Richmond Journal of Law & Technology, showed that technology-assisted review (TAR) can be more effective and more efficient than exhaustive human review, on both recall and precision.2
2012 — Courts agree
In Da Silva Moore v. Publicis Groupe (S.D.N.Y., Feb. 24, 2012), a federal court judicially accepted TAR for the first time, citing the superiority of computer-assisted review over the alternatives.3
2020s — The gap goes exponential
Modern review layers large-language-model analysis on top of TAR: entity resolution across aliases, timeline construction, machine translation, and handwriting OCR — capabilities that once existed only inside classified intelligence systems and are now available in unclassified form to any agency or litigant that chooses to use them. None of it requires a clearance — only the decision to use it.
The findings in detail
Finding 1 — The methodology was the one discredited in 1985.
In March 2025, roughly 1,000 FBI Information Management Division personnel were put on 24-hour shifts to comb through approximately 100,000 Epstein-related records by hand, supplemented by hundreds of New York Field Office personnel — many without training in identifying statutorily protected victim information or handling FOIA material.4 That is linear manual review: the approach whose recall Blair & Maron measured at roughly 20%, and which both the research literature and the courts have treated as inferior to technology-assisted review since 2011–2012.
Finding 2 — The tooling predates the modern field.
The software behind FBI eDiscovery review is the Clearwell platform, a product of the late 2000s that the Bureau adopted in 20115 — and whose product line, after passing from Clearwell Systems to Symantec (2011) to Veritas, has since reached end-of-life. Federal procurement records show the purchase: in September 2011 the FBI obligated roughly $11.9 million for "CLEARWELL SOFTWARE" across three reseller awards — $1,986,794 via ThunderCat Technology (PIID DJFA1D104061, Sept. 9), and $4,714,875 plus $5,177,960 via Software Information Resource Corp. (PIIDs DJFA1D1131101 and DJFA1D1131102, Sept. 29–30, under IDC DJFJFBI11311; the first order may duplicate the parent vehicle's reported value).8 What followed was maintenance, not modernization: a $30,956 upkeep order in March 2016 (PIID DJF161200D0004095), and then — a decade after the original buy — a $4.04 million order for a "60 TB license, Clearwell/Veritas software" for the BIDMAS program ($2,037,000 on Feb. 24, 2021, plus a $2,000,000 option exercised Feb. 16, 2022; PIID 15F06721F0000517).8 In other words, when the platform aged out, the Bureau's response was to buy more of it. The procurement record shows no FBI award for any modern review platform — Relativity, Everlaw, Casepoint, or Exterro — only small forensic-tool purchases (Nuix, Intella) in the $875–$15,000 range. A review run in 2025 on a platform bought in 2011 cannot apply fifteen years of accumulated method.
This is the procurement data behind the FBI’s eDiscovery tooling:
Date | What the FBI bought | Vendor | PIID | Amount
---|---|---|---|---
Sept 9, 2011 | "CLEARWELL SOFTWARE" | ThunderCat Technology | DJFA1D104061 | $1,986,794
Sept 29, 2011 | "CLEARWELL SOFTWARE" (order under IDC DJFJFBI11311) | Software Information Resource Corp. | DJFA1D1131101 | $4,714,875
Sept 30, 2011 | "CLEARWELL SOFTWARE" | Software Information Resource Corp. | DJFA1D1131102 | $5,177,960
June 25, 2012 | "CLEARWELL TRAINING" | Symantec Corp. | DJFA2DOE0030 | $3,325
Mar 7, 2016 | "CLEARWELL SYMC SOFTWARE MAINTENANCE" | Software Information Resource Corp. | DJF161200D0004095 | $30,956
Feb 24, 2021 | "EDISCOVERY 60 TB LICENSE, CLEARWELL/VERITAS SOFTWARE FOR BIDMAS PROGRAM" | ThunderCat Technology | 15F06721F0000517 | $2,037,000
Feb 16, 2022 | Option exercise on the same 60 TB Clearwell/Veritas license | ThunderCat Technology | 15F06721F0000517 P00003 | $2,000,000
2015–2022 | Small forensic tools only (Nuix licenses, Intella Pro) — no review platform | Nuix, Vound, resellers | various | $875–$15,000 each
Finding 3 — Nothing was measured.
The central lesson of the BART study is that reviewer confidence is not a metric: the attorneys who had found 20% believed they had found 75%. Nothing in the public record describes a recall target, a sampling protocol, or any quality-control measurement for the FBI's review. The observable outcome matches the prediction: the government has since acknowledged fixing thousands of documents in the released files that may have contained victim information.7
Finding 4 — Available capabilities went unused.
Machine translation, handwriting OCR, entity resolution, and image analysis exist in unclassified form and are in routine use across government and private litigation. Internal FBI emails from the 2025 review instead show personnel asking for capabilities their tooling simply lacked — including help identifying public officials and celebrities in images.6 The gap was not one of clearance or availability, but of process.
Sources
Blair & Maron, "An Evaluation of Retrieval Effectiveness for a Full-Text Document-Retrieval System," Communications of the ACM (1985)
Grossman & Cormack, "Technology-Assisted Review in E-Discovery Can Be More Effective and More Efficient Than Exhaustive Manual Review," 17 Rich. J.L. & Tech. 11 (2011)
Da Silva Moore v. Publicis Groupe, 287 F.R.D. 182 (S.D.N.Y. 2012)
Sen. Durbin press release on the March 2025 FBI review (July 2025)
"The FBI buys Clearwell eDiscovery Platform," eDisclosure Information Project (Dec. 2011)
Bloomberg, "Epstein Files: New FBI Emails Detail Review, 'Special Redaction Project'" (Nov. 2025)
PBS NewsHour, "Government says it's fixing thousands of documents in Epstein-related files that may have had victim information"
Federal Procurement Data System (FPDS), FBI "Clearwell" contract actions: PIIDs DJFA1D104061, DJFJFBI11311, DJFA1D1131101, DJFA1D1131102, DJFA2DOE0030, DJF161200D0004095, 15F06721F0000517
Scope of this review. This analysis draws only on publicly released, already-redacted records and public reporting. Redactions in those records protect victims and third parties — including minors — and this review neither reproduces nor speculates about protected content. The critique here is of the review process, not of the redactions' purpose.
Evidentiary.ai — enterprise eDiscovery & litigation support. This analysis is based on public records and public reporting, for reference and research; it is not legal advice.
Evidentiary.ai · Analysis

Expert Analysis Resources
This page host an expert analysis of the Epstein files, providing critical insights and defensible methodologies that legal teams can apply to complex data reviews.
Note: This file is a web-converted, read-only resource.
Consultation Inquiry
Contact Evidentiary.ai for similar defensible AI reviews of complex matters. We invite you to reach out to our team to discuss your specific requirements.
Request a Forensic Audit
Invite Evidentiary.ai to review your complex matters with defensible AI. We provide secure review for legal teams tackling intricate data sets.