Blog Diligence

Speed vs. Accuracy in Diligence: Is There Actually a Tradeoff?

Clausebeam Team 8 min read
Abstract concept of speed and accuracy in legal review

The standard framing of AI-assisted contract review is that it trades speed for accuracy: you get results faster, but you accept more risk of missing something. We ran a structured test to find out whether that tradeoff is real, and what it actually looks like in practice.

The premise of our test was straightforward: take a set of contracts, run AI-assisted clause review on each one, then have experienced attorneys complete the same review independently. Compare the results. Where does the model catch what humans miss? Where does it miss what humans catch? And where does AI review genuinely add speed without any accuracy cost?

This is not a peer-reviewed study. It is internal test data from a set of 80 agreements: 40 NDAs and 40 MSAs, reviewed both with Clausebeam and by practicing attorneys with experience in commercial contracts. We are sharing it because the findings do not fit neatly into either the "AI is perfect" or "AI cannot be trusted" narratives, and practitioners deserve a more nuanced picture.

What we tested and how

The 80 agreements were selected to span a range of complexity. Some were short-form NDAs of 3 to 4 pages with fairly standard language. Others were longer-form bilateral NDAs with carve-outs, residuals provisions, and non-solicitation language that interacted with the confidentiality obligations. The MSAs ranged from 8-page standard SaaS agreements to 25-page professional services contracts with detailed statements of work and complex indemnity structures.

Each attorney reviewer was asked to identify and assess: the top three clause-level risk items in each agreement, any provisions they considered off-market for the contract type, and any provisions they would mark up if representing the receiving party. Attorney reviewers worked independently without access to Clausebeam's output first. We compared outputs after both were complete.

Timing was measured informally. We are not claiming precise time-per-document figures, because attorney review speed varies based on experience and attention, and some attorneys work faster under observation. What we can say is the order of magnitude: NDA review by experienced attorneys ran in the range of 15 to 45 minutes per document. MSA review ran in the range of 45 minutes to 2 hours per document, depending on length and complexity. Clausebeam's clause analysis for all 80 documents completed in under 25 minutes total.

Where AI added speed without accuracy cost

On the NDAs, the AI and attorney reviews were in strong agreement on the core structural issues: which confidentiality obligations were mutual vs. one-way, whether residuals clauses were present and how broad they were, and whether the non-solicitation scope exceeded what would typically be acceptable. For straightforward NDAs with clean language, the agreement rate on "what are the top risk items" was very high. In those cases, the AI review genuinely did add speed with no accuracy cost. Getting to the same risk identification in a fraction of the time is real value, not a marketing claim.

The same pattern held for certain categories of MSA clause: liability cap structure, payment term deviations from standard, and governing law / jurisdiction selection all showed high agreement rates. These are categories where the risk lives in identifiable structural patterns rather than in contextual judgment, and the AI model performed well.

Where human judgment still dominates

The gaps appeared in three categories.

The first was contextual risk: clauses that look standard in isolation but are problematic given the specific nature of the transaction. One MSA in our test set contained a standard limitation of liability provision that most reviewers rated as acceptable. Two attorney reviewers flagged it as problematic because, given the nature of the services described in the statement of work, nearly all foreseeable damages from a breach would be consequential, making the consequential damages exclusion functionally eliminate the entire liability cap benefit for the customer. The AI analysis did not pick up that contextual inference. It correctly identified the limitation of liability structure and flagged the consequential damages exclusion, but it did not reason through why those provisions were particularly problematic for this specific service type.

The second gap was cross-document pattern recognition at a semantic level. Within a single agreement, the AI model handles cross-references well for structured clause relationships. But when the risk emerges from comparing the substance of obligations across two separate provisions that are not formally cross-referenced, such as an indemnity obligation in one section and a force majeure provision in another that could undermine the indemnity trigger, human reviewers caught this connection more reliably than the model did in ambiguous cases.

The third gap was negotiating strategy. Attorney reviewers did not just flag what was risky; they assessed what was worth fighting over given the client's position and the likely counterparty response. That strategic overlay is not a clause analysis function. It requires understanding the deal context, the counterparty's likely posture, and the client's actual risk tolerance. No AI clause analysis tool should be expected to provide that judgment, and Clausebeam does not try to.

The false precision problem

One finding we want to highlight because it has implications for how practitioners use AI clause review: false precision on confidence.

Some AI clause analysis tools present flags with a confidence score or risk rating that can create a misleading impression of certainty. In our test, the most dangerous pattern was when the model assigned a moderate-risk flag to a clause that experienced attorneys rated as high-risk, creating a risk of under-prioritization. The model was not wrong to flag the clause; it was wrong in its risk calibration relative to context.

Our approach to this problem is to flag what we find and explain the structural basis for the flag, rather than assigning a single numeric risk score. We want practitioners to read the flag explanation and apply their own contextual judgment to calibrate severity, rather than trusting a score that may not account for deal-specific factors. This adds a small amount of work for the reviewer, but we think it produces better outcomes than false precision.

What the tradeoff actually looks like

Based on our testing, the real picture of the speed-accuracy tradeoff in AI-assisted clause review is not a flat line. It is category-specific.

For structural clause detection, standard deviation identification, and risk flagging based on language patterns, AI review performs comparably to experienced attorney review while being dramatically faster. For high-volume, routine contract review, such as an in-house team's NDA and vendor MSA queue, the tradeoff is not really a tradeoff at all. You get thorough structural review at speed without meaningful accuracy loss.

For contextual risk assessment, deal-strategy judgment, and the kind of reasoning that connects provisions across documents at a semantic level, attorney judgment remains essential and should not be replaced. AI review is a first-pass tool, not a substitute for the judgment that experienced practitioners bring to complex analysis.

We are not saying that AI will never improve on contextual reasoning in contracts. We are saying that right now, today, the appropriate use case is enhancing the efficiency of attorney review, not replacing it. Getting to a thorough structural risk analysis in minutes instead of hours means attorneys can spend their judgment where it actually matters, rather than on the structural detection work that a model can handle reliably and quickly.

A note on what consistency adds

One finding that does not fit neatly into a speed-accuracy framing but is genuinely important: AI review is more consistent than attorney review at scale. Experienced attorneys get tired. They work under time pressure. They develop preferences about which clauses to focus on based on experience that may not match the priorities of a specific deal. Clause review consistency across a large contract stack, where every indemnity clause gets the same scrutiny regardless of which document in the stack it appears in, is something AI does better than humans by design. That consistency is worth something independent of the speed question.