Legal Contract Intelligence — Corporate Law Firm, New York, USA
The Situation
A mid-size US litigation and corporate law firm was running a generic LLM across contract review. On paper the setup made sense. In practice, the model flagged irrelevant clauses, missed jurisdiction-specific risk language, and required attorney correction on most outputs. Senior associates were spending more time fixing the model's work than they would have spent doing the review themselves.
The firm had 14 years of internal contract history: annotated reviews, negotiated redlines, approved playbooks by deal type, and documented risk decisions. None of it was in the model.
What Amorisoft Did
Amorisoft ran an 8-week data preparation process to structure the firm's historical review decisions into a training dataset. This covered contracts across 6 practice areas, with clause-level annotations drawn from 4,200 reviewed documents.
Fine-tuning ran on top of a base LLM selected for its performance on long-context legal text. The firm's own playbooks were used to set deviation thresholds by clause type and counterparty profile.
RLHF ran across 5 feedback cycles. Raters were 4 senior associates from the corporate and M&A teams. Each cycle took one week. By cycle 3 the model's deviation flags were matching senior associate judgment at a rate above 80%. By cycle 5 it was at 89%.
The production model went live 6 weeks after data handover.
Implementation Note
The firm's document management system had no consistent naming convention across practice areas. Contracts from the litigation team were filed differently from those in corporate. Amorisoft's data team spent the first 10 days of the data preparation phase on classification before any annotation work could start.
Results
First-pass review time dropped from 4.5 hours per contract to 1.3 hours, a 71% reduction. Clause extraction accuracy went from 61% to 94% after the RLHF cycles, and attorneys accepted 89% of the model's deviation flags without revision. On the same headcount, senior associate throughput ran at 3.2x baseline.
