Back to Blog
Compliance Collects Evidence. It Doesn't Grade It.

Compliance Collects Evidence. It Doesn't Grade It.

Andrew Roe

In March 2026, an anonymous investigator who publishes as DeepDelver posted an analysis of a leaked spreadsheet from Delve, a compliance automation startup that had raised $32 million at a $300 million valuation. Of 494 SOC 2 report files in the leak, 493 were nearly identical: the same paragraphs, the same grammatical errors, one string of text appearing in 99.8% of the files (DeepDelver, 2026-03-19). Every one of the 259 Type II reports in the set reached the same conclusion for a different company each time: zero security incidents during the observation period.

Delve disputes the allegations and says it is investigating the leak (Corporate Compliance Insights, 2026-05-21). What DeepDelver's spreadsheet showed, if accurate, was more specific than "the reports looked similar." Conclusions and board meeting minutes were reportedly pre-written by Delve itself, before the client had submitted any evidence at all, and the named "US-based" auditors traced back to certification shells overseas (DeepDelver, 2026-03-19).

A SOC 2 report is built out of a single word doing three different jobs: evidence. A screenshot of a settings toggle is evidence. A log an auditor actually reviewed across a year is evidence. A conclusion typed before the client sent anything is, technically, also evidence, in the sense that it sits in the same folder as the rest. Nothing in the process registers the difference between those three things. That gap is not what let Delve happen. It is what let it go unnoticed for as long as it did.

Evidence isn't one thing

Compliance frameworks ask for evidence per control, and the tooling built around SOC 2, ISO 27001, and their peers has gotten very good at collecting it. What it has never done is grade it.

Some of what gets filed as evidence is a state check: an API pull showing a config flag is set, a screenshot of a settings panel, an attestation that a policy exists. Call it configuration evidence. It proves a setting sat in a particular position at a particular moment, and nothing else. "MFA enabled: true" pulled from an identity provider does not say whether MFA can be bypassed with a fatigue attack or a call to the help desk.

Some of it is observation evidence: a log reviewed across the audit window, a quarterly access review someone signed. It proves a process ran, or that someone looked. It does not prove the process caught anything, or that looking changed a result.

And some of it is test evidence: a penetration test that actually tried to get past the control, an incident and the response to it. This is the only tier that answers the question a security control exists to answer, which is whether it holds when someone tries to get through it.

A SOC 2 matrix does not distinguish between these. A control can be satisfied with configuration evidence when the risk it is meant to cover is exactly the kind only test evidence could speak to, and the report reads the same either way.

The audit samples evidence. It doesn't weigh it.

The reason this gap survives contact with a real audit is procedural, not just human error. A SOC 2 Type II engagement tests a sample of an observation window, typically three to twelve months, rather than every instance of a control operating. A control that ran correctly eleven months out of twelve can still read as a clean pass if the sample never lands on the twelfth month.

Sampling decides how much configuration or observation evidence is enough. It has never been the layer that decides whether a control needed test evidence in the first place. That call sits with the auditor's judgment, engagement by engagement, with no structural requirement forcing it.

The AICPA's own response to the Delve allegations shows what that gap looks like from the regulator's side, independent of whether any single allegation holds up. On May 14, 2026, the AICPA's Peer Review Board issued guidance stating that "reviewing a single SOC 2 engagement file is often insufficient" and directing reviewers to compare a firm's engagements against each other, since identical risk assessments, sample sizes, and testing procedures across different clients make an engagement "nonconforming" (Journal of Accountancy, 2026-05). A structured monitoring program for firms doing SOC 2 work started June 1, 2026. The guidance names no specific firm; it cites "recent feedback" as the trigger. But the fix it lands on is instructive: the tell for ungraded evidence was never going to be inside any one file. It was going to be that every file looked the same.

What ungraded evidence looks like at the extreme

Delve is the version of this where nobody was checking at all. 259 different companies, in different industries, at different maturities, do not produce 259 identical "zero incidents" conclusions if anyone read the evidence behind each one (DeepDelver, 2026-03-19). That outcome only happens if the conclusion was never conditioned on the evidence to begin with, which is what DeepDelver's spreadsheet alleges.

The uncomfortable part is how little Delve had to invent. It did not need to fake a working product. Configuration and observation evidence collection is exactly what a compliance automation platform is built to do well, and by most accounts Delve's underlying platform did that part fine. What it apparently skipped was the layer where a human decides whether the evidence collected actually answers the question the control exists to answer, and writes a conclusion that depends on the answer. Remove that layer and a real evidence pipeline still produces a real-looking report, because nothing downstream ever required a check on whether the report's claims track its evidence.

Most of this is not fraud, and the gap still shows

Most SOC 2 reports are not fabricated. Most auditors read what they are given. The gap between evidence tiers still costs real companies something, because Vanta, the largest vendor in this exact market, published its own data on it. In an October 2025 survey of 3,500 IT and business leaders, 61% said they spend more time proving security than improving it, and 64% said today's security frameworks feel like "security theater" (Vanta, 2025-10-29).

That is not a fraud story. Those are honest audits, honest evidence, an honest platform, and a majority of the people doing the work still describe it as theater. The likeliest reason is the one this argument has been building toward: a great deal of what gets filed as evidence was only ever going to prove that something was configured, never that it works, and the people producing it know the difference even when the report in front of them does not.

Why the flattening persists

The incentive is straightforward. Configuration and observation evidence are cheap and fast to collect. Test evidence is slow and expensive, because it requires someone to actually try to break the thing. Delve marketed speed and scale as its differentiator on the way to a $300 million valuation (Corporate Compliance Insights, 2026-05-21); it is far from the only compliance vendor whose pitch leans on how fast a report can be produced.

A report format that never discloses which tier of evidence backs each control gives a buyer, an enterprise procurement team reading a vendor's SOC 2 report, no way to reward the vendor who paid for a real penetration test over the one who pulled a screenshot. Both reports say evidence was provided. Nothing on the page says which kind.

The question worth asking is not whether your auditor or your compliance platform has evidence for a given control. Almost everyone does, Delve included. Ask what tier it is. A screenshot proves a setting exists. A log proves someone looked. Only a test proves the control holds when someone tries to get past it, and for the controls where that is the actual risk, a report that never crossed into that tier was never going to tell you what you needed to know, whichever auditor signed it.


Sources: DeepDelver, "Delve - Fake Compliance as a Service - Part I", 2026-03-19; Corporate Compliance Insights on the Delve scandal, 2026-05-21; Journal of Accountancy on AICPA's SOC 2 peer review guidance, 2026-05; Vanta State of Trust 2025, 2025-10-29.