Most content about AI legal document review tools is a ranking. Nine vendors, five scoring columns, a comparison table at the bottom. Those lists are useful for learning the names, and useless six months later, because the category moves faster than anyone refreshes them — features ship, pricing changes, companies get acquired, and the tool that topped the list is now a module inside something else.
A durable framework beats a perishable ranking. Whatever legal AI software you are evaluating — for contract analysis, discovery review, or investigations — ask these ten questions first. They are ordered roughly the way a buying conversation actually goes: what the thing is for, whether it works, whether it is defensible, what it costs, and what it takes to run it.
How to use this: ask all ten of every vendor on your shortlist, including us. The answers that matter are the specific ones. “Yes, we have an audit trail” is not an answer; “here is an exported audit log from a real matter” is.
Short answer: most buyers conflate three different product categories that share the phrase “AI legal document review.” Contract-lifecycle tools work forward on documents you are drafting or negotiating. Discovery and investigation tools work backward on documents that already exist and must be classified for a matter. Research tools do neither — they answer legal questions against case law. Map the vendor’s real workflow to your matter types before anything else, because the rest of the evaluation depends on it.
This is the single most common source of wasted evaluation time. A litigation boutique reads a roundup of “the best AI legal tools,” books four demos, and discovers on call three that the product is a Word add-in for redlining an NDA — excellent software, entirely unrelated to the 90,000 emails sitting in the matter they actually need help with.
Scroll table horizontally →
| Category | Core job | Where it lives | Typical buyer |
|---|---|---|---|
| Contract lifecycle / contract analysis | Drafting, redlining, clause extraction, obligation tracking | Word, a CLM, or a browser | Transactional and in-house commercial teams |
| Discovery & investigation review | Classification, privilege, redaction, production | A review platform holding the collected population | Litigation firms, in-house litigation and compliance |
| Legal research | Answering questions of law against case law and secondary sources | A research database | Anyone briefing a motion |
Within discovery there is a further split worth asking about: does the platform handle the data types your matters actually produce? Email and loose files are table stakes. Modern matters increasingly turn on Slack and Teams, Discord and other chat platforms, and mobile device data — formats that behave nothing like documents and that many review tools still flatten into unreadable transcripts. Ask for the supported format list in writing.
What to ask: “Walk me through a matter that looks like mine, from collection to production.” If the demo starts with a single contract, you are in the wrong category.
Short answer: ask for precision and recall on a named task — responsiveness, privilege, entity extraction — and ask who measured it. Self-reported accuracy is a marketing number until it has been reproduced on a document set the vendor did not choose. The only figure that ultimately matters is the one you generate yourself on a validation sample from your own population.
There is good empirical reason to be skeptical of vendor claims. In 2024, researchers at Stanford’s RegLab and Institute for Human-Centered AI tested the leading AI legal research tools and found hallucination rates above 17% for one major product and above 34% for another — on tools that had been marketed as eliminating hallucination. Those were research tools rather than review classifiers, so the numbers do not transfer directly. The transferable lesson does: independent testing produced materially worse results than the vendors’ own descriptions, on a task the vendors had every incentive to get right.
“Accurate” is also not one number. A privilege classifier tuned for recall will over-call privilege and hand you a bloated log to cut down; one tuned for precision will miss documents and create waiver risk. Both can be described as 95% accurate. Ask which error the tool is built to prefer, and whether you can move that threshold per matter.
For a sense of what a serious benchmark looks like from the inside, we have published a nine-model comparison on legal responsiveness review and the evaluation framework underneath it, including where our own numbers fall short.
Short answer: yes is the only acceptable answer, and the override has to be logged. Every classification the model makes — responsive, privileged, confidential — must be visible to counsel, reversible by counsel, and recorded as having been reversed. A tool that produces conclusions you cannot inspect or change is not a review tool; it is a liability.
This is where the defensibility conversation actually starts. Courts assessing AI-assisted review are applying existing standards of reasonableness rather than writing new doctrine, and the question they ask is whether a competent attorney supervised the process. Attorney direction also affects the protection your work product receives: output generated from attorney-crafted instructions and reviewed by counsel is treated as reflecting counsel’s mental impressions, while unsupervised machine classification generally is not.
The practical test is whether the interface makes the override cheap. If correcting a wrong call takes six clicks and a support ticket, reviewers will stop correcting them, and the audit trail you were counting on becomes a record of everyone agreeing with the model. Our note on what “defensible” actually means covers the standard in more detail.
What to ask: “Show me a document where a reviewer disagreed with the model. What does the record of that look like six months later?”
Short answer: two things. First, whether privilege determinations produce a defensible, exportable privilege log as a by-product of the review rather than as a separate manual project. Second, what stops privileged material from leaving your control — both into a production and into a vendor’s training data.
Privilege is where review costs concentrate and where mistakes are least recoverable. It is also the part of the workflow most vendors gloss over in a demo, because the log is tedious and the log is the deliverable. A tool that classifies privilege beautifully but leaves you drafting several thousand entries by hand at $5–$15 apiece has moved the cost rather than removed it. We have written on automating log creation and on what changes at scale.
The second half is a vendor-diligence question rather than a product question. Recent judicial attention to AI in discovery has focused less on whether the model is right and more on what happens to the documents after upload — specifically the risk that a tool retains submitted material or uses it to train shared models, which is impossible to unwind once privileged material has been ingested. Get the answer in the contract, not the sales deck.
Short answer: measure it in the units your deadlines use. Ask for a stated turnaround from upload to first usable review set, then ask for a reference matter with a document count comparable to yours. “Fast” is not a metric; “100,000 documents processed and classified in two days” is.
Time-to-value is where legacy discovery platforms quietly cost the most. Processing queues measured in days, an implementation call before the first upload, and a project manager between you and your own data all convert into weeks that a compressed discovery schedule does not have. For a firm responding to a subpoena or working an early case assessment against a court-set date, throughput is not a convenience feature.
What to ask: “If I upload 250 GB on Friday afternoon, when can a reviewer start, and what is the charge for that processing?” The second half of that question is where processing surcharges tend to surface.
Short answer: find out which unit you are billed on — users, documents, or gigabytes — and model what happens when that unit doubles. Seat-based pricing penalizes you for adding a reviewer or bringing in outside counsel; per-document pricing penalizes you for collecting broadly; per-GB pricing is the most predictable, provided processing, exports and support are inside the number rather than beside it.
Scroll table horizontally →
| Pricing model | Common in | What it punishes | What to check |
|---|---|---|---|
| Per seat / per user | Contract AI, legal research, some review platforms | Adding reviewers, co-counsel and experts mid-matter | Annual commitment, minimum seats, cost of a temporary reviewer |
| Per document | Managed review services | Broad collection and second passes | Whether privilege review and log drafting are add-ons on top of the base rate |
| Per gigabyte hosted | Traditional eDiscovery platforms | Long-running matters that sit in hosting for years | Separate processing, user, project-management and egress charges |
| Flat all-in per gigabyte | DecoverAI | Nothing structurally — the unit is the data, not the team | That “all-in” is literal: review, logs, redaction, production, support |
The trap is rarely the headline rate. It is the line items beside it: per-seat fees that make bringing in a second firm expensive, project-management hours billed at $250, egress fees charged when you leave, and credits that expire before your matter does. The most useful single question a buyer can ask is the one vendors most resist: what is the all-in number for this matter?
What to ask: “Price me a 20 GB matter with 4 reviewers running 14 months, including processing, exports and support — one number.”
Short answer: SOC 2 Type II is the minimum bar, and the distinction from Type I matters — Type I attests to controls at a point in time, Type II to their operation over a period. Regulated matters add requirements: HIPAA for anything touching health data, and a documented retention and deletion policy in every case. Ask where the data is hosted and get written confirmation that client data is never used to train shared models.
Scroll table horizontally →
| Requirement | Why it matters | What satisfies it |
|---|---|---|
| SOC 2 Type II | Independent attestation that controls operated over time, not just existed on one day | A current report under NDA — not a “SOC 2 compliant” claim or a Type I |
| HIPAA | Required whenever the population contains protected health information | A signed BAA, plus documented safeguards |
| No training on client data | The waiver exposure courts have actually focused on | A contractual term, not a policy page |
| Data residency | Cross-border matters and GDPR obligations | A named region and a written commitment to it |
| Retention & deletion | Matters end; the data should too | A stated retention period and verifiable deletion on request |
For matters with a European dimension, add the cross-border and GDPR questions to this list. DecoverAI’s own posture is documented on our security page and in our Trust Center.
Short answer: ask whether the tool connects to what your team already opens every day — Word, email, your document management or practice management system — or whether it becomes a parallel system with its own login, its own export step, and its own re-import step. Every handoff between systems is a place where versions diverge and work gets redone.
The honest counterpoint: for discovery review specifically, a dedicated platform is the right answer, because the work does not belong in Word. What you are testing there is not whether the platform is separate but whether getting data in and out of it is friction-free. The questions that matter are ingestion breadth on the way in and, on the way out, whether you can leave — production formats, load files, and whether your data comes back in a usable shape when the matter ends or you switch vendors. Our migration guide covers what that looks like in practice, and egress fees are where the answer usually gets expensive.
What to ask: “If we terminate in month three, what exactly do we get back, in what format, and at what cost?”
Short answer: for anything touching litigation or regulatory production, require chain-of-custody logging, Bates numbering, and an exportable audit trail that distinguishes what the AI proposed from what a human confirmed. “AI-assisted” output with no record of the assist is worse than no AI at all, because you cannot defend a process you cannot describe.
The test is whether you could reconstruct, eighteen months later and under cross-examination, how a specific document ended up classified the way it did: which model version scored it, what it scored, which reviewer confirmed or overrode it, when, and against what protocol. Most platforms log something. Fewer log the model’s proposal and the human decision as separate, exportable facts.
Pair this with a production quality-control process before anything goes out the door — our production QC checklist covers the failures that actually draw sanctions, and how to use AI for document review without getting sanctioned covers the documentation to keep alongside it. If your matter has an ESI protocol, check that the tool can produce to it rather than to its own default.
Short answer: ask directly whether your team can start a matter today, by itself, or whether the first matter requires a vendor-led implementation, a training period, and a minimum contract. For a small litigation team this is often the deciding difference, and it is rarely on the pricing page.
Legacy discovery platforms were designed for buyers with a litigation support department and an IT function to absorb the implementation. A five-attorney firm has neither, which is why the practical barrier to adopting enterprise review software is usually not the license fee but the six weeks and the dedicated person it assumes you have. If your firm runs discovery without a litigation support team, weight this question heavily.
What to ask: “Can we upload data and start reviewing this afternoon without talking to anyone? If not, what is the shortest path, and what does it cost?”
Take this into the demo. One row per question, what a good answer sounds like, and the answer that should give you pause.
Scroll table horizontally →
| Question | What good looks like | Red flag |
|---|---|---|
| 1. Built for what? | A worked example of a matter shaped like yours, end to end | The demo is a single contract when your problem is 90,000 emails |
| 2. Accuracy | Precision and recall per task, plus a pilot on your own data | One blended accuracy figure, self-reported, no test set named |
| 3. Attorney override | One-click override, logged as an override, visible later | Model output presented as final, or overrides that leave no record |
| 4. Privilege | Log generated from the review, exportable, in your format | Privilege classification with the log left as manual work |
| 5. Time-to-value | A stated turnaround and a reference matter at your volume | “It depends on the queue” and no comparable engagement |
| 6. Pricing | One all-in number for a modeled matter, in writing | Seat minimums, annual lock-in, expiring credits, egress fees |
| 7. Security | Current SOC 2 Type II, HIPAA where relevant, no-training term | “SOC 2 compliant,” a Type I, or training rights buried in the MSA |
| 8. Workflow fit | Broad ingestion in, clean load files and full export out | Data that is easy to put in and expensive to get back |
| 9. Audit trail | Exportable log separating model proposal from human decision | “AI-assisted” with no record of what the AI decided |
| 10. Onboarding | Self-serve start, same day, no minimum commitment | Mandatory implementation project before the first upload |
These questions are not neutral, and it would be dishonest to pretend otherwise — we built DecoverAI around a particular set of answers to them, for a particular buyer: the 5-to-50-attorney litigation firm and the small in-house team that has real discovery volume, no enterprise budget, and no litigation support department. Here is where we land, stated specifically enough that you can hold us to it.
The throughline is a gap these ten questions tend to expose. Legacy discovery platforms — Relativity, Everlaw, CS Disco, Logikcull — were built for the AmLaw 100’s budget and IT department, and their pricing and onboarding still assume both. Most of the newer AI legal software getting written about was built for contract lifecycle management, which is a different job. That leaves discovery-grade, defensible review packaged for a firm that does not have an enterprise contract in it, which is the space we set out to occupy. If you want to see the alternatives side by side first, our comparison of eDiscovery software tools and our buyer’s checklist both name names.
Run the ten questions at every vendor on your list, us included. If our answers do not hold up against someone else’s, that is worth knowing before you sign, not after.
What is AI legal document review software?
AI legal document review software applies machine learning and large language models to a collected document population to classify it — responsive or not, privileged or not, confidential or not — so attorneys read a fraction of the set rather than all of it. It sits inside the discovery workflow, between processing and production, and the attorney still makes the final call.
Is AI-assisted document review legally defensible in court?
Yes, when the process is documented. Courts assess the reasonableness of the process rather than the tool, so what matters is a written protocol, statistical validation of the results, attorney supervision of every consequential call, and an audit trail showing what the model proposed and what a human confirmed or overrode.
How much do AI legal document review tools cost?
Pricing follows three patterns: per-user seat licenses (common in contract AI, roughly $100–$200 per user per month), per-document managed review (commonly $0.75–$0.90 a document plus add-ons), and per-gigabyte hosting (roughly $25–$100/GB, often with processing and export charged separately). DecoverAI is flat at $60/GB/month, all-in.
Can AI legal tools replace document review attorneys?
No. They change what the attorney does. AI removes the requirement that a lawyer open every document to reach a defensible call, but the privilege determination, the responsiveness judgment on close calls, and the sign-off on production remain the attorney’s. A tool that markets itself as removing the lawyer is describing a defensibility problem.
What is the difference between AI contract analysis and AI document review?
Contract analysis works forward on documents you are creating or negotiating — drafting, redlining, clause extraction, obligation tracking — and usually lives in Word. Document review works backward on documents that already exist and must be classified for a matter: responsiveness, privilege, confidentiality, redaction and production. The workflows barely overlap.
Is it safe to upload privileged documents to an AI tool?
It depends entirely on the vendor’s data handling, not the model. Before uploading privileged material, get contractual confirmation that client data is never used to train shared models, confirm where data is hosted, confirm deletion is real and verifiable, and confirm the platform carries SOC 2 Type II with matter-level isolation. Public consumer chatbots fail all four.
How accurate is AI at detecting privilege in legal documents?
Accuracy varies widely by tool and by how it is measured, and self-reported numbers are not a substitute for testing on your own data. The practical answer is that no vendor’s privilege classifier should be trusted unsupervised: run a validation sample from your own population, measure precision and recall on it, and keep an attorney on the final call.
The rate ranges here are market benchmarks rather than quotes, and they move. The Stanford figures come from a 2024 RegLab and HAI study of AI legal research tools, cited to make a point about the gap between vendor claims and independent testing rather than as a measure of document review classifiers, which were not what was tested. DecoverAI’s throughput and volume figures are platform-reported, the forty-minute comparison is a customer’s account of one matter rather than an audited benchmark, and our pricing reflects published rates subject to change. Verify current terms with every vendor directly, including us.
This article is general information about legal technology and discovery practice, not legal advice for any particular matter. Vendor descriptions and rate ranges reflect publicly available information at the date of writing and may change.