WritfindFindings, not flags.
Articles

Can AI deep research tools handle adverse media investigations? A three-case comparison

In our comparison of Nikola, Ozy Media and Binance, general AI deep research tools provided useful leads. Writfind covered more reference matters overall; Gemini led on Nikola. The results support initial discovery, but decisions about important matters still need checks on developments, omissions and sources.

By Writfind · Published

A Writfind investigation · September 2026. Official product documentation checked: 2 October 2026.

Contents

  1. AI research for adverse media
  2. Findings at a glance
  3. Evaluation: what counts as a good investigation?
  4. Limitations and the practical choice

AI research for adverse media

ChatGPT Deep Research and Gemini Deep Research use AI to research questions across multiple sources and produce reports with citations. Gemini also lets users review a research plan and select sources. These are dedicated research workflows, distinct from ordinary ChatGPT answers. OpenAI documentation, Gemini documentation.

Screening identifies potentially relevant matters. Investigation examines what happened, subsequent developments and unresolved questions. The report records findings and limitations; a write-up explains their implications for its reader.

We tasked both tools and Writfind with public-source investigations of Nikola, Ozy Media and Binance. Writfind focuses on adverse media screening, investigation and narrative reports, and visualization. We assessed the resulting reports for coverage, relevance and source support.

Findings at a glance

All three tools covered Ozy’s federal fraud prosecution. Writfind found the UK Financial Conduct Authority (FCA) restrictions on Binance Markets Limited, which both Deep Research reports missed. Both found the Australian Securities and Investments Commission (ASIC) action against Binance Australia Derivatives and Nikola’s Lion Electric battery-contract dispute, which Writfind missed. Findings, report excerpts and coverage judgments, FCA announcement.

Across our 76-matter reference set, Writfind’s coverage was 73.7–85.5%, ChatGPT Deep Research’s 42.1–43.4%, and Gemini Deep Research’s 48.7%. The ranges reflect report and scoring versions from the investigation, not independent repeated trials.

Coverage ranges from our September 2026 investigation on a 76-matter reference set; values are also in the following table.
Figure 1. Coverage of 76 admitted reference matters in our September 2026 comparison. Ranges reflect report and scoring versions, not statistical intervals. Counts and version conditions; CSV data.
Tool Covered / reference matters Matter recall
Writfind 56–65/76 73.7–85.5%
ChatGPT Deep Research 32–33/76 42.1–43.4%
Gemini Deep Research 37/76 48.7%

The percentages are pooled counts divided by 76, not averages of company percentages. Gemini led on Nikola in the version used for our public samples: 16 of 21 matters, versus Writfind’s 14. An overall lead does not establish better coverage for every company. Case results.

Evaluation: what counts as a good investigation?

A useful adverse media report identifies relevant matters, excludes unrelated material and supports its findings. We assessed these requirements separately, using a task-specific approach to evaluation. Hamel Husain’s evaluation framework.

Metric Good result Failure
Matter recall Relevant adverse matters are identified An in-scope adverse matter is omitted
False positives Finding is adverse and concerns a subject within scope Non-adverse, wrong-subject or out-of-scope finding
Evidence support A cited page establishes the central claim Cited sources do not establish that claim

These measures do not score the completeness of follow-up or the importance of an omission. Narrative and visualization were not scored, and no combined score was used.

Common task and scoring

The common task covered the company, controlled subsidiaries and business-related conduct by current or former executives. It excluded namesakes, incidental mentions and uncontrolled group companies. Plaintiff-only or victim-only matters were excluded, subject to exceptions for adverse disputes. English reports were to cite factual claims, distinguish allegations from findings and state unknown outcomes. Writfind used a keyword search provider; the Deep Research products used their own research tools. Task and scoring method.

Coverage and source support used binary, model-assisted judgments with product identities withheld. Coverage used two passes and a further judgment for disagreement. Human decisions were recorded for reference-set exceptions; this was not a complete independent human audit. Scoring record.

1. Coverage: which relevant matters were found?

The unit was a matter, rather than a section or a URL. A report counted as covering a matter when it identified the parties and event and stated at least one reference fact. Splitting a matter into several sections earned no extra credit.

We pooled candidates and checked them against the common scope to build the reference set: 21 Nikola matters, 9 Ozy matters and 46 Binance matters. A pool cannot include completely unknown matters that every system missed. Our set also retains five candidates from earlier research—one Nikola matter and four Binance matters. All four Binance matters were absent from the compared reports. This exception makes the 76-matter set broader than a strict three-report union. Reference origins and admission decisions.

Our reference set: 76 admitted matters, including five earlier-research-only candidates; unknown additional matters cannot be counted.
Figure 2. Our reference set and its boundary: 71 matters include candidates from compared reports; five originated only in earlier research. Diagram areas are schematic. Reference origins and admission record.

The report version used in Writfind’s public samples gives this case-level result:

Company Writfind ChatGPT Deep Research Gemini Deep Research
Nikola Corporation 14/21 12/21 16/21
Ozy Media Inc. 9/9 6/9 7/9
Binance Holdings Limited 33/46 14/46 14/46

Counts are covered / admitted reference matters. Each matter has equal weight, so Binance’s 46 matters dominate the aggregate. The counts do not distinguish a consequential omission from a less important one. Writfind’s 9/9 for Ozy covers this reference set, not every possible adverse matter. Case results and full coverage matrix.

The ASIC example shows why omissions deserve individual attention. ChatGPT reported the A$10 million court penalty against Binance Australia Derivatives; Gemini covered the earlier licence-cancellation stage. Both counted as covering the matter; Writfind missed it. The same coverage judgment thus accommodated different stages of the action. Finding a matter does not establish that its subsequent developments have been investigated. Report passages, ASIC release.

2. Relevance: which findings should be excluded?

False positives were defined as non-adverse, wrong-subject or out-of-scope matters. Recorded candidate rejections were 2 for Writfind, 0 for ChatGPT Deep Research and 0 for Gemini Deep Research; both were “not adverse.”

These are reference-set construction counts, not a comparable false-positive rate for the final reports. One later record lacks the candidate file and therefore records zero; that zero cannot establish an absence of false positives. Rejected candidates and attribution limit.

3. Evidence: do the sources establish the claims?

Evidence support assessed cited passages of at least eight words, excluding pure URL lists. A pass required at least one cited page to establish the central claim. Uncited passages were not assessed; unreadable-source and investigation-only passages were excluded.

Tool Supported / eligible cited passages Excluded: no cited page readable Excluded: investigation only Unresolved scoring, kept in denominator
Writfind 160/161 (99.4%) 0 1 1
ChatGPT Deep Research 145/159 (91.2%) 4 5 0
Gemini Deep Research 140/150 (93.3%) 0 2 1

These percentages are not whole-report accuracy rates: they assess the central claim, not every detail or individual URL. Two unresolved judgments—one Writfind and one Gemini—remain in the denominator; neither is a confirmed factual error. Evidence counts and exclusions.

One Ozy passage illustrates the distinction. ChatGPT’s report combined the September 2025 dismissal of the SEC case against Ozy and Carlos Watson with March 2026 injunctions against two co-defendants. Its cited SEC announcement established the dismissal but did not discuss those injunctions. The recorded judgment did not accept the combined passage as centrally supported. Original passage and judgment, SEC announcement.

Two claims share one SEC source: the source establishes the dismissal but does not establish the co-defendant injunctions.
Figure 3. A passage in ChatGPT Deep Research’s Ozy report, paraphrased to separate its claims. The SEC page supports the dismissal but does not establish the second claim. Original report passage and recorded judgment; SEC release.
Claim in the report Check against the cited page
Ozy and Watson: September 2025 dismissal with prejudice The SEC announcement states this development.
Rao and Han: March 2026 final antifraud injunctions The cited page does not discuss this development. Its absence is a support gap, not proof that the claim is false.

The next step is to seek separate support for the injunctions. Writfind’s Ozy sample also records an unresolved conflict between a 2023 official release and a reported 2026 development. Its matter text and sources expose that gap for review. Ozy SEC matter.

Limitations and the practical choice

Writfind initiated and conducted this supplier self-test. The three companies were also development cases, not an unseen or representative sample. The reference set cannot quantify unknown omissions; its five earlier candidates qualify the comparison. One retained Brazilian congressional matter lacks a direct verification link. Reference record.

The work was conducted in September 2026. Writfind’s sample-report collection finished on 21 September and revised drafting on 23 September. Two ChatGPT reports state a cutoff of 18 September, but exact execution timestamps, account tiers and settings for the six Deep Research tasks were not preserved. A cutoff is not an execution date. Current product descriptions checked on 2 October are separate from these test results. Dates and missing conditions.

Saved versions combine separate collections, redrafting and rescoring; the same six Deep Research reports were reused. Development observations varied by roughly 5–19 percentage points, without controlled, same-configuration repeatability trials. The ranges do not establish stable superiority. Model-assisted scoring, source-access differences and incomplete public competitor exports limit independent reproduction. Versions and variation, comparison record.

The results suggest a practical division of work:

Task Useful starting point Review still required
Initial screening General Deep Research can supply leads and an initial account. Check identity and scope; an empty result does not establish an absence of adverse matters.
Investigation of an important matter Trace developments, as the ASIC example illustrates. Check procedural stages and separately source the injunction claim in the Ozy report.
Report and write-up Present supported findings, unresolved points and their implications. Review omissions and source gaps that could affect the intended assessment.

These cases support AI-assisted discovery and show where further investigation is needed; they do not establish an exhaustive screen or a universal winner.


Editorial update · 3 October 2026: clarified the interpretation of the measures and added practical review guidance. Comparison results unchanged.