Every legal AI vendor demos well. The interface is clean, the sample answer is confident, and the room nods. That is what a demo is for. A demo is built to look finished. A platform has to hold up when it is pointed at ten years of real matters, real clients, and real privilege, and asked a question no one rehearsed. The distance between those two things is the entire evaluation.
This guide lays out the criteria that actually decide the purchase and, under each one, the questions that separate a serious vendor from a good slide deck. Work through them in roughly the order below. The early criteria, grounding and accuracy and where the system runs, eliminate candidates fastest, so there is little reason to spend a pilot on a vendor that cannot clear them.
Does every answer trace back to a source?
The first thing to test is whether the product can show its work. A legal answer is only usable if you can follow it to the document it came from, read the surrounding context, and confirm the tool did not invent the connection. Ask the vendor to produce an answer and then click straight through to the exact passage that supports each claim. Not a list of documents the system consulted. The specific line.
Push on the edges. Ask what the tool does when the answer is not in the corpus. A grounded system says it does not know or returns nothing; an ungrounded one fills the gap with a plausible sentence, which is the failure mode that ends careers. If citations are an afterthought bolted on for the demo, this is where it shows. We wrote about why this matters, and how to test for it, in our piece on grounded citations and legal AI hallucinations.
Accuracy and hallucination controls
Every vendor will tell you their system is accurate. The useful question is how they know, and what happens when it is not. Ask how the vendor measures accuracy, whether they can share the methodology, and how they handle the cases the model gets wrong. A team that has thought seriously about this can describe their evaluation harness, their known failure modes, and the guardrails that catch a bad answer before an attorney does.
Ask specifically about refusal behavior. A trustworthy legal AI product is comfortable saying "I cannot answer that from your documents." A product tuned to always produce something confident is optimized for demos, not for diligence. Ask to see it decline. If it never does, that is the answer.
Where does the platform actually run?
This is the criterion that reweights all the others, so ask it early: where does the software run, and whose infrastructure holds your client data while it does? There is a meaningful difference between a platform that pulls your documents into the vendor's cloud and one that deploys inside your firm's own cloud, under your own keys, in a dedicated and isolated enclave that your security team already governs.
The deployment model decides who you are trusting. When the platform runs in your own environment, the data, the encryption keys, the network boundary, and the audit logs sit inside controls your firm already operates and your clients already examine. Ask the vendor plainly: can this deploy in our own cloud? Who holds the keys? Can we revoke them? Our security page lays out how Reframe answers those questions, and our comparison of private legal AI versus public chatbots covers why the boundary matters more than any single feature.
SOC 2 and data handling
A SOC 2 logo on a slide is the start of the security conversation, not the end of it. Ask for the full Type II report under NDA and read it. Confirm the product you are buying is inside the audited scope, check the exceptions the auditor found, and see which subprocessors are carved out. Then get the data handling commitments in writing: what is retained, for how long, and a contractual warranty that your documents are never used to train anyone's model.
These are the questions that reveal how a vendor really operates once the deal is signed. We put the full list, and how to read what comes back, in our guide to the SOC 2 questions to ask legal AI vendors.
Does it connect to your document system?
A legal AI tool is only as good as the corpus it can see, and for most firms that corpus lives in a document management system. Ask whether the platform integrates natively with iManage or NetDocuments, how it handles the permission model those systems enforce, and whether an attorney only ever sees results from matters they are already entitled to open. A tool that flattens your access controls to make search easier has created a confidentiality problem, not solved a knowledge one.
Ask how ingestion stays current. A one-time export goes stale the moment it finishes; a live integration keeps the AI working from what the firm knows today. We go deeper on the mechanics in our note on DMS integration with iManage and NetDocuments.
Governance and audit trails
Assume a client, a regulator, or your own general counsel will one day ask what the AI did and why. The vendor needs an answer built in. Ask whether every query, retrieval, and answer is logged, whether those logs are yours to keep and inspect, and whether permissions can be scoped by matter, practice group, and ethical wall. Ask what an administrator can see and what they cannot.
Governance is also about change control. Who can alter how the system behaves, and is that change recorded? For a firm that answers to clients and courts, an unauditable AI is not a convenience with a small risk. It is a liability you cannot explain after the fact.
Is it built on a legal ontology, or bolted onto search?
This is the criterion most evaluations miss, and it separates a durable platform from a wrapper. Many legal AI products are keyword or vector search with a language model stapled to the front. They can find documents that look similar to your question. What they cannot do is understand that a party, a matter, an obligation, and a deadline are different kinds of things that relate to each other in specific ways.
A platform built on a legal ontology models those relationships explicitly. It knows a counterparty across a hundred matters is one entity, that a clause creates an obligation, that an obligation has a trigger and a deadline. That structure is what lets the AI answer questions about your firm's actual situation rather than retrieve documents that share words with your prompt. Ask the vendor to explain how their system represents entities and relationships. If the answer is "we search really well," you are buying search. If it is a coherent account of a knowledge model, you are buying a platform that will still be answering harder questions in three years.
References and pilots
Finally, make the vendor prove it outside the demo. Ask for references at firms of comparable size and complexity, and ask those references the questions the vendor cannot answer for themselves: what broke, what support looked like, and whether attorneys actually use it. Adoption is the honest metric. A tool no one opens after the launch email failed regardless of how it scored.
Then structure a pilot with teeth. Define success before it starts, in your terms, on your documents, with your hardest real questions rather than the vendor's curated set. Set a fixed window, name the attorneys who will run it, and agree in advance what result would make you buy and what result would make you walk. A confident vendor will welcome that; the pilot is where a real platform pulls away from a good deck.
None of this is adversarial. The vendors worth buying from have heard these questions before and answer them quickly, because they built the product expecting to be asked. The ones hearing them for the first time are telling you, precisely and at no cost, exactly where they sit. The demo starts the conversation. The diligence is the decision.