5 Questions to Ask AI Legal Software Vendors Before You Buy

Jessica Nguyen
September 14, 2026 · 5 min read
Jessica Nguyen is President, Chief Strategy and Legal Officer at Sandstone. She most recently served as Deputy General Counsel for AI Innovation and Trust at DocuSign.
The legal market is crowded with general-purpose AI wrappers marketed as purpose-built in-house tools.
Skipping structured vendor evaluation is how in-house legal teams end up with tools that solve the wrong problems, fail compliance review, or produce AI outputs no lawyer can stand behind. This guide covers the five questions that matter, the follow-up questions vendors won't volunteer, and the red flags that should end an evaluation early.
What Problem Does This AI Legal Software Actually Solve?
Before evaluating any vendor, legal teams need to name the specific operational problems they are trying to fix. Vendors most eager to close deals quickly usually sell the most generic solution.
Purpose-built AI legal software solves defined workflow problems. The most common operational failures worth scoping before any vendor conversation:
- Scattered requests: Intake fragmented across Slack, email, and ticketing tools, with no central view of what is pending or who owns it
- Inconsistent positions: No systematic way to enforce standard clause language or fallback terms, leaving position consistency up to individual lawyers
- Lost institutional knowledge: Precedent, negotiation history, and policy decisions locked in individual counsel's inboxes or memory, inaccessible at the moment they're needed
- Limited visibility: No unified view of workload, capacity, or request status, making it impossible to report on legal's operational performance
A vendor worth evaluating should map their product to at least one of these problems with specificity. If the answer to "what does this solve" is a feature list rather than a workflow description, that is a signal.
The best AI for legal questions addresses a defined operational failure, not a general chatbot interface pointed at legal documents. Ask vendors to demonstrate the tool on a problem your team actually has, not a scenario they constructed, and request a sandbox environment that mirrors your real workflow without exposing actual matter data.
How Does This Vendor Protect Our Data and Legal Privilege?
Legal data is categorically different from most enterprise data. Attorney-client privilege is a structural legal protection that can be waived inadvertently. Any vendor that cannot answer data handling questions in writing is not a vendor an in-house team should trust with their matters.
Written documentation of data handling practices is non-negotiable. Verbal assurances during a sales call are not a substitute, and legal teams should treat unsigned data processing agreements as an unresolved evaluation item, not a post-signature detail.
Data Residency and Encryption Standards
Data residency refers to the physical and jurisdictional location where data is stored and processed. For legal teams operating across borders, this matters both for regulatory compliance and for controlling which legal frameworks govern your data. Ask where data is stored (geography and cloud provider), whether it is encrypted at rest and in transit, and whether the vendor uses customer data to train its models. Many vendors bury this in their terms of service. The answer should be a clear no, documented in writing.
Also ask about zero-data-retention agreements with the underlying AI model providers. Many legal AI vendors sit on top of foundation models from OpenAI, Anthropic, Google, or others, and some of those providers retain query data by default unless enterprise agreements specify otherwise. The data handling terms with those providers matter as much as the vendor's own policies.
Client Confidentiality Safeguards
Ask whether queries submitted to the AI are logged, stored, or accessible to vendor staff, and how the tool handles documents containing privileged communications or personally identifiable information. Ask specifically about data portability: whether your institutional knowledge, playbooks, and document history can be exported when you leave, and in what format. Vendor lock-in on your own institutional knowledge is a structural risk that most teams don't price in at signing.
Compliance Certifications to Request
Three certifications signal baseline maturity in data security and privacy practice:
- SOC 2 Type II: Validates that security controls have been audited over time, not just at a single point in time.
- ISO 27001: The international standard for information security management systems
- GDPR and CCPA compliance: Required for any legal team handling personal data across jurisdictions
Beyond certifications, ask whether the vendor's product has been reviewed against ABA Formal Opinion 512 on generative AI use by lawyers, and whether they have guidance on compliance with state bar rules in jurisdictions where your lawyers are licensed. Competence, confidentiality, and supervision requirements apply to AI-assisted legal work and vary by jurisdiction.
How Do You Prevent Hallucinations and Ensure Accuracy?
Hallucinations are confident, coherent AI outputs that are factually wrong — and in legal work, a liability risk that is difficult to detect and costly to remediate.
A 2026 survey found only 22.1% of legal professionals report high trust in AI outputs initially. That trust comes from verifiability — the ability to trace an AI claim back to its source. Vendors that can't explain how they verify outputs are asking legal teams to work on faith.
Character-Level Citations and Audit Trails
Ask whether the AI can cite its sources at the character level — exact, clickable references back to specific language in source documents, not paraphrased summaries. Page-level citations are insufficient for legal work; a lawyer who cannot verify exactly what language the AI relied on cannot confidently stand behind the output. Audit trails serve a second function: they create a defensible record of how legal conclusions were reached, which matters when a contract position is later questioned internally or in a dispute.
Human Oversight in AI Workflows
The right architecture for AI in legal work keeps lawyers in the decision seat. Supervised agents should handle drafting and first-pass redlines based on institutional knowledge and playbooks, while lawyers apply judgment, make exceptions, and approve outputs before anything moves forward. Ask how the vendor's product enforces this structure and where it builds human checkpoints into the workflow.
A tool that routes final legal outputs directly to business teams without a lawyer review step will eventually create a problem.
Training Data Sources, Validation, and Model Drift
Ask where the model's legal knowledge comes from: proprietary legal corpora, public case law, or customer data. Ask whether attorney expertise was involved in developing the system's legal reasoning, or whether it is a general large language model pointed at legal documents. Ask about model drift — how accuracy is monitored when the vendor updates its underlying model, and whether customers are notified before changes that could affect legal output behavior.
Can This AI Integrate With Our Tools and Learn From Our Team?
This question separates point solutions from platforms. Point solutions do one thing in isolation and require your team to change how they work to use them. Rip-and-replace tools fail because they add friction rather than removing it — the legal team has to adopt a new destination to do work, adoption stalls, and the tool goes unused. The better architecture layers AI on top of workflows where legal work already happens.
Tech Stack Compatibility Questions
The integrations that matter most for in-house legal:
- Communication tools: Slack, Microsoft Teams, Outlook, and Gmail — wherever requests actually originate, not where legal wishes they originated
- CLM and contract tools: Ironclad, DocuSign, and existing contract repositories where executed agreements live
- Business systems: Salesforce, Jira, ServiceNow, and HRIS — the systems that carry business context attached to every legal request
- Document workflow: Microsoft Word integration with native redlining capability matters for any team doing substantive contract work inside Word
Ask about single sign-on (SSO) compatibility with your existing enterprise identity infrastructure. A legal AI tool that requires a separate login creates an adoption barrier that compounds over time.
Institutional Knowledge and Playbook Support
Playbooks are the encoded negotiation preferences, fallback positions, and clause libraries that represent how your legal team actually operates. They are the difference between AI that produces generic first drafts and AI that produces drafts grounded in your specific positions.
Ask whether the tool can ingest past redlines, templates, and negotiation notes to learn your team's positions. Ask whether playbooks update over time as your positions evolve, or whether they require ongoing manual maintenance. Ask whether attorney expertise was used to seed pre-built playbooks, or whether your team is building from a blank slate. Platforms like Sandstone allow teams to drag and drop contracts to build self-improving playbooks that get sharper with every use, capturing institutional knowledge that would otherwise leave when a lawyer does.
Unified Intake Across Channels
Ask whether the tool can surface requests from multiple channels into a single view, with business context attached automatically. Routing and triage that understands the intent behind a request — not just its keywords — reduces the time lawyers spend figuring out what a request actually requires before they can start working on it. A tool that centralizes intake without enriching it with business context has solved half the problem.
What Does ROI Look Like and How Do You Measure Success?
Vendor ROI decks are marketing materials. The more useful question is: what does success look like in the first 90 days, and how will we measure it?
A straightforward ROI framework for legal AI: multiply weekly hours saved per lawyer by the number of lawyers on the tool, multiply by the team's fully loaded hourly rate, then multiply by 47 working weeks. Set that against total annual cost of the software. That calculation — done with your numbers, not the vendor's — is a more credible starting point than any vendor-provided ROI projection.
Pricing models vary — per-seat, per-matter, and flat-fee structures carry different total cost of ownership depending on team size and volume. Ask vendors to walk through the full cost model, including implementation, integration, and any fees that surface post-signature. Implementation timelines belong in the ROI model too: tools that layer on existing workflows deploy in weeks, while solutions requiring process redesign or change management can take months, and that gap has a real dollar cost. Ask for references from customers at comparable scale and in comparable industries.
The metrics worth tracking from day one:
- Time to first response: How quickly business teams receive answers after submitting a request — a proxy for legal's accessibility to the business
- Request backlog: Volume of pending matters over time, as a measure of throughput
- Consistency: Reduction in off-playbook positions or escalations that required ad hoc legal judgment
- Capacity visibility: Ability to benchmark workload by team or business unit, enabling more accurate resource planning and reporting to leadership
How to Run a Meaningful Pilot
A 14-day pilot on your own documents is more revealing than any vendor demo. Before the pilot begins, agree on specific success metrics: which workflows will be tested, what accuracy threshold is acceptable, and what adoption rate justifies a full rollout. Run on actual matters from your pipeline, not simplified examples, and test specifically for the failure modes most likely to affect your team — hallucination rates on your document types, integration friction with your specific tools, and whether business teams actually use the intake channel.
Vendors who resist independent pilot access or insist on using their own sample documents are protecting a gap between demo performance and real-world performance.
Red Flags That Should End the Evaluation
A few vendor behaviors signal a product that is not ready for enterprise legal use:
- Pricing hidden behind NDAs or available only after a full sales cycle
- Data retention answers that are vague or require interpretation — vendors built for legal's risk tolerance can explain their data practices in plain language
- No zero data retention agreement with underlying model providers, which is standard in enterprise AI agreements
- Accuracy claims without methodology — benchmark data and validation processes should be available on request, not described in marketing language
- A law firm tool retrofitted for in-house use — law firm AI optimizes for billable hour efficiency; in-house AI needs to solve intake, knowledge capture, and cross-functional collaboration
Choosing the Right AI Partner for Your Legal Team
These five questions separate vendors who have thought seriously about in-house legal from those selling horizontal AI tools with a legal label applied. Answering them well requires a vendor to understand how legal work actually operates, which is the baseline requirement for any tool serious about improving it.
The goal is a platform that unifies context, surfaces institutional knowledge, keeps lawyers in control, and improves with use. When legal has those things, it stops operating as a bottleneck and starts functioning as a strategic partner to the business.
Learn how Sandstone enables in-house legal departments with AI.
FAQs About AI Legal Software Vendor Evaluation
What questions should you always ask before using AI in legal work?
The five core questions: what operational problem does this solve, how does the vendor protect privileged data and comply with bar ethics rules, what safeguards exist against hallucinated outputs, how does the tool integrate with existing systems and learn from institutional knowledge, and what does ROI look like in measurable terms. These apply whether evaluating a standalone tool or a full platform.
How do you compare AI legal software vendors side by side?
Build a scorecard based on the five questions above and weight each criterion by your team's priorities. Run a pilot on your own documents — a 14-day pilot on actual matters is more informative than any vendor demo. Speak with customer references directly rather than through vendor-arranged calls.
What is the difference between legal AI for law firms and in-house teams?
Law firm tools optimize for billable hour efficiency and client-facing research. In-house tools need to solve intake management, internal knowledge capture, cross-functional visibility, and reporting to leadership on operational performance. A tool built for law firms applied to an in-house context will optimize for the wrong outcomes.
How long does it take to implement AI legal software?
It depends on integration complexity and how much process change the tool requires. Tools that layer on existing systems deploy in weeks. Solutions requiring new portals, process redesign, or extensive change management take months — and that time belongs in total cost of ownership calculations.
What compliance and ethics rules apply to AI legal software?
ABA Formal Opinion 512 addresses competence and confidentiality obligations when using generative AI. State bar rules vary, with many jurisdictions issuing guidance on supervision, confidentiality, and disclosure for AI-assisted legal work. Evaluate whether a vendor's tool has been assessed against these standards and whether they provide jurisdiction-specific compliance guidance.