How to pilot AI contract review with source checks
A practical law-firm pilot for testing citation fidelity, false positives, approval steps, and reviewer confidence before wider AI contract review rollout.
An AI contract-review pilot should answer a narrow operational question: can the team reach a verified issue list more reliably, with less reconstruction, while the lawyer still controls every conclusion and proposed change?
That is different from asking whether a model can produce an impressive summary. A fluent summary can still omit the clause that matters, detach a finding from its source, or make a reviewer spend longer proving it wrong than they would have spent reading the contract.
This is a practical way for an England & Wales firm to test the workflow. It is not a legal-performance benchmark, and it does not replace the firm’s procurement, information-governance, or professional-supervision requirements.
Start with a decision the pilot must support
Choose one recurring review job with a known reviewer and a stable definition of done. Examples include extracting the commercial spine of supplier agreements, checking a defined set of clauses against a playbook, or preparing a first issue list from a small transaction pack.
Write down the boundary before uploading anything:
- which document types are in scope
- which fields or playbook positions matter
- who may access the pack
- who verifies findings and citations
- whether the pilot may propose wording or only identify issues
- what must happen when the system is uncertain or fails
Avoid a mixed bundle of unrelated contracts on the first run. If the task is vague, the result will be impossible to score and easy to oversell.
Build a human-reviewed reference
Before testing AI output, have an appropriately experienced reviewer create the reference extract or issue list for the chosen documents. Record the source clause for every material entry and note where reasonable reviewers may disagree.
The reference does not need to pretend that contract review is perfectly objective. It needs to separate three things:
- facts that should be extracted consistently, such as parties, dates, term, governing law, and notice details;
- playbook questions with a defined firm position;
- judgments that depend on wider matter context or negotiation strategy.
That separation prevents the pilot from giving the model credit for confident prose where the real task was source retrieval, and it keeps professional judgment outside the automated score.
Test the source trail before the prose
For every extracted field or finding, open the cited passage and ask:
- Does the citation contain the claimed fact or clause?
- Is the surrounding context needed to understand it?
- Did the system select the operative provision rather than a definition, recital, or superseded schedule?
- Can the reviewer move from the answer to the source without searching the document again?
A citation is useful only when it shortens verification. A page number attached to an unsupported conclusion is not evidence.
Pillars is Ardela’s product for this part of the job: the team defines fields, extracts them into a reviewable Table, and keeps answers connected to uploaded source documents. The commercial product detail belongs on the Pillars page; the pilot should judge the source trail against the firm’s own documents and reference review.
Count misses and unsupported findings separately
One headline “accuracy” percentage hides the errors a legal team actually needs to manage. Track at least:
- missed material items — the reference contains an issue the system did not surface;
- unsupported findings — the system raised an issue that the cited text does not support;
- wrong-source findings — the answer may be plausible, but the citation points to the wrong provision;
- ambiguous cases — the document or playbook leaves room for reasonable disagreement;
- reviewer corrections — the finding was useful but needed a material change before it could be relied on.
Review the failures, not only the aggregate. A tool that performs well on standard dates but misses a particular carve-out may still be unsuitable for the chosen workflow.
Separate finding an issue from drafting the change
Detection and redlining are different decisions. First verify that the issue exists and understand the commercial position. Only then ask for proposed wording.
When a proposed change is generated, compare it with the operative clause and the approved playbook position. The reviewer should see the exact replacement, accept or reject it deliberately, and preserve the original when it is rejected.
Clara presents proposed rewrites and redlines for approval rather than applying them silently. In a pilot, record which proposals were accepted unchanged, edited, or rejected. That is more informative than counting how many suggestions appeared.
Exercise permissions and failure states
A clean sample contract proves very little about a production workflow. Test the controls around the review:
- a user without access to the matter or pack;
- two similar matters that must not be mixed;
- a scanned or poorly structured document;
- a missing schedule or incorporated document;
- a retry after provider or network failure;
- an answer where the source is insufficient.
The correct outcome may be a refusal, an unresolved field, or a request for human input. A system that always returns a complete-looking answer is harder to supervise than one that exposes uncertainty.
Measure the reviewer’s finished job
At the end of the pilot, compare the complete workflow with the reference process. Useful measures include:
- time from opening the pack to a verified extract or issue list;
- number and severity of missed material issues;
- number of unsupported findings;
- reviewer corrections per document;
- time spent locating source passages;
- redline acceptance, edit, and rejection counts;
- whether the audit trail is sufficient to reconstruct what was checked and approved.
Do not turn a tiny pilot into a universal productivity claim. The outcome applies to the tested document set, playbook, reviewers, and product version. Expand only when the next document group has its own reference and accountable owner.
A defensible next step
Run the first pilot on a representative, appropriately protected pack with an agreed field list and review standard. Use the same reviewers for the reference and assisted pass where practical, record disagreements, and retain the failed cases for the next evaluation.
If the job is primarily structured extraction and cited questions, start with Pillars. If the source has been verified and the next job is a proposed rewrite, use Clara with approval on every change. For data-location, access, and identity controls, review Security before the pilot begins.