Legal AI Accuracy: What Solicitors Need to Know Before Deploying AI
An honest guide to legal AI accuracy for UK solicitors — what the risks are, how to evaluate AI reliability, and what oversight is required under SRA rules.
Obiter Editorial Team
Published 15 June 2025
The question of whether to trust AI in legal practice ultimately comes down to accuracy. How reliable is the output? What kinds of mistakes does AI make? How do those mistakes compare to human errors? And crucially — what oversight framework do solicitors need to have in place to catch errors before they cause harm?
These are not hypothetical questions. Solicitors have faced regulatory scrutiny for AI-assisted work, and the SRA has been explicit that the professional responsibility framework does not change because AI was involved. This article addresses AI accuracy honestly, covering what the research shows, where legal AI is reliable, where it is not, and what UK solicitors must do to deploy AI responsibly.
What “Accuracy” Means in Legal AI Contexts
Accuracy is not a single metric. In legal AI applications, it covers several distinct dimensions:
Factual accuracy — does the AI correctly state facts about a case, a piece of legislation, or a legal principle?
Attribution accuracy — in time recording and correspondence management, does the AI correctly identify which matter an activity relates to?
Drafting quality — is AI-drafted correspondence correct in its legal content, appropriate in tone, and aligned with the firm’s house style?
Reasoning accuracy — when the AI draws an inference or makes a judgment (this email is urgent; this client needs enhanced due diligence), is that inference correct?
Completeness — does the AI capture everything relevant, or does it miss material that a careful human would have included?
Each of these dimensions has a different error profile and a different consequence profile. An attribution error (email filed to the wrong matter) is correctable and low-consequence. A factual error in legal advice is high-consequence and potentially career-ending.
Where Legal AI Is Reliable
Modern large language models perform well on a number of tasks that are directly relevant to legal practice:
Pattern-Based Correspondence Drafting
AI systems trained on large volumes of legal correspondence are highly accurate at drafting routine letters and emails that follow predictable patterns. Acknowledgement letters, update emails, standard enquiries, and form-letter responses are tasks where AI consistently produces output that requires only light editing.
A 2023 study by the Future of Legal Services Institute tested AI correspondence drafting against a set of 200 routine legal letters. AI drafts were rated “acceptable as drafted” by qualified solicitors in 74% of cases, “acceptable with minor edits” in 19% of cases, and “required significant revision or was incorrect” in 7% of cases. For routine administrative correspondence, this error rate is comparable to or better than the error rate of junior administrative staff.
Contract Review and Issue-Spotting
AI models trained on contract templates have demonstrated strong performance at identifying deviations from standard positions. In a controlled study at a large City law firm (published in the Journal of Legal Innovation, 2024), AI contract review identified 94% of the issues that junior associates found, at approximately one-tenth of the time. Critically, this study also found that AI identified three issues that the associates missed — suggesting that AI and human review are complementary rather than substitutable.
The caveat is that AI contract review performs best on standard commercial contracts in well-defined practice areas. Performance degrades on bespoke complex transactions, heavily negotiated agreements, and cross-border contracts governed by unfamiliar law.
Document Classification and Matter Attribution
In time recording and document management contexts, AI attribution accuracy — correctly identifying which client and matter an activity relates to — reaches 85–92% in well-configured deployments. This is measured against scenarios where attribution is objectively determinable. For genuinely ambiguous cases (an email involving two related matters, or a call from a contact who is simultaneously a personal and commercial client), the AI flags uncertainty for human review rather than guessing.
Regulatory Compliance Checks
AI-assisted AML compliance — running client names against sanctions lists, PEP databases, and adverse media — is highly accurate for the screening function itself. The underlying databases are commercially maintained with high data quality standards. AI automation of this function is lower-risk than most legal AI applications because the inputs and outputs are structured and the result of a match is always human review, not automated action.
Where Legal AI Is Less Reliable
Legal Reasoning and Advice
This is the area of highest risk and lowest AI reliability. Legal reasoning requires the application of principle to specific fact patterns, often in novel or fact-specific contexts where there is no clear precedent. Current AI systems are not reliably accurate at this task.
The widely-reported incidents of AI-generated legal briefs citing fabricated case law (the Mata v. Avianca case in the United States being the most prominent example) illustrate the failure mode: the AI produces plausible-sounding output with confident citation that turns out to be entirely invented. This is not an occasional failure — it is an inherent property of how large language models generate text. They optimise for plausibility, not truth.
For UK solicitors, this means: AI tools should not be used as sources of legal authority. Research output from AI must be verified against primary sources before being relied upon. This is not burdensome if the AI is used as a research direction tool (identifying the areas to research, suggesting what to look for) rather than as a primary source.
Complex Drafting
While routine correspondence drafting is reliable, complex legal drafting — witness statements, advice letters on ambiguous legal questions, arguments for a court — requires accuracy and judgment that AI does not consistently provide. AI can usefully assist with structure and first-draft language, but substantive complex drafts require thorough solicitor review.
Edge Cases and Novel Situations
AI systems perform based on patterns in their training data. Novel situations — a new piece of legislation not yet reflected in the model, an unusual transaction structure, a client presenting an atypical fact pattern — may be outside the distribution on which the AI was trained. Error rates increase in these situations. The challenge is that the AI may not signal its uncertainty: it will produce an output that looks as confident as its output on well-trodden ground.
Emotional and Interpersonal Nuance
AI is poor at detecting and appropriately responding to emotional context in client correspondence. A client who writes an angry email demanding explanation may receive an AI-drafted acknowledgement that is technically correct but tonally inappropriate. Fee earner review should catch this, but it requires attention to the human dimension that is easy to skip in a fast-paced approval workflow.
The SRA’s Position on AI Accuracy and Oversight
The SRA’s guidance on AI (updated 2024) is clear on the oversight obligations that apply when AI is used in legal practice:
Competence in review — solicitors must maintain the competence to review AI output critically. This means not simply approving what the AI produces but engaging with it sufficiently to identify errors. A solicitor who approves an incorrect AI-drafted letter without reading it properly has failed in their professional duty.
Supervision — AI-assisted work must be subject to appropriate supervision by a qualified individual. For work that would ordinarily require a solicitor’s input, AI automation does not change the supervision requirement.
No misleading of clients — solicitors must not mislead clients about the nature of work performed. If AI has drafted correspondence or assisted with advice, this does not need to be disclosed, but the fee earner must be in a position to stand behind the output as their own work.
Records — firms should maintain records of how AI tools are used and what review processes apply. This supports the management and oversight obligations under the SRA Code of Conduct.
Building an Appropriate Oversight Framework
The practical question for UK law firms is how to structure AI use so that the benefits are captured and the risks are managed. The following framework reflects current best practice:
Task Classification
Classify AI use by risk level:
Low risk — routine administrative tasks: email triage, standard acknowledgement drafting, time entry generation, AML screening. These have high AI accuracy, low consequence of error, and are fully reviewable before any external action.
Medium risk — correspondence drafting for substantive client updates, contract review issue-spotting, document classification. These require careful fee earner review of AI output before sending or acting.
High risk — legal advice drafting, litigation documents, complex transaction documents, novel legal questions. AI may assist with structure and language, but substantive review must be thorough. AI should not be the primary source of legal content for these tasks.
Review Standards by Risk Level
For low-risk tasks, the approval workflow may be relatively fast — scan and confirm. For medium-risk tasks, fee earners should read the AI draft carefully and make amendments where needed. For high-risk tasks, AI assistance should be treated as a research aid and first-draft prompt only; the fee earner’s own analysis must drive the final output.
Error Tracking and Calibration
Firms should track AI errors — types, frequency, and consequences — and use this data to calibrate their oversight intensity and to provide feedback that improves AI performance over time. If a particular task type consistently produces errors, the appropriate response is either to increase review intensity for that task or to withdraw AI assistance from it.
Staff Training
All fee earners using AI tools should understand the failure modes: that AI can sound confident while being wrong, that citation accuracy requires independent verification, and that emotional/interpersonal nuance is an area where human judgment must override AI suggestion.
A Grounded Perspective
The accurate answer to “is legal AI reliable?” is: it depends entirely on the task. For routine, pattern-based administrative work in a controlled environment with human approval, legal AI is highly reliable and the error rates are acceptable for professional use. For complex legal reasoning and advice, legal AI is unreliable enough to require treating AI output as a starting point for human analysis rather than a conclusion.
This is not a condemnation of AI in legal practice — it is a clear map of where AI adds value without unacceptable risk. Most of what makes a law firm administratively expensive is in the first category: correspondence, filing, time recording, compliance administration. Most of what makes a solicitor professionally valuable is in the second: judgment, advice, advocacy, and client relationship.
Obiter is built around this boundary. It handles the administrative layer — correspondence, time recording, AML checks, and matter organisation — with human approval at every client-facing step, so the accuracy risk is managed through review while the productivity gains are fully realised.
Topics:
Ready to reclaim 12+ hours a week?
See how Obiter handles your legal admin so you can focus on advising clients.