ResearchWorkSystemsAboutTalk shop
Research/Fine-Tuning Made a Small Model More Precise, Not More Complete: A Contract-Clause Extraction Experiment
Small Language ModelsLoRAContract AnalysisEvaluationPre-registrationOriginal Research

Fine-Tuning Made a Small Model More Precise, Not More Complete: A Contract-Clause Extraction Experiment

A pre-registered experiment on public CUAD contracts: can a LoRA-tuned 4B model close the gap to a larger model at finding 41 kinds of clause and quoting them verbatim? Four tuning rounds later, the tuned model quoted far more faithfully and made fewer false claims, but still missed more clauses than the untuned base. The pre-registered gate was not cleared, and the main comparison was never run.

Porsync Research · Published 2026-10-03

Finding

On 41-category contract-clause extraction, LoRA-tuning a 4B model raised precision (0.82 to 0.93) and verbatim-quote fidelity (83% to 97%) but not macro-F1: the best of four tuned arms scored 0.697 against 0.727 for the untuned base on validation (a difference the sample could not separate from zero), and no tuned arm cleared the pre-registered gate to be tested against the larger model.

CUAD clause extraction — small-model tuning experiment (Oct 2026)

Clause detection (41 categories) plus verbatim span extraction on public CUAD contracts: 347 train, 61 validation, 102 test. A larger model (one sample) and an untuned 4B model were scored on the test set; four LoRA-tuned variants of the 4B model were scored on validation. One consumer GPU. Scored by a fixed script against a pre-registered protocol.

MeasurementValueNote
Larger model, test (macro-F1)0.783one sample, 102 contracts
Untuned 4B model, test (macro-F1)0.7156.8 points behind; paired 95% CI -10.0 to -3.0
Untuned 4B model, validation (macro-F1)0.727precision 0.824, recall 0.867, quotes 83.2% verbatim
Best tuned arm (S3), validation (macro-F1)0.697precision 0.925, recall 0.816, quotes 97.1% verbatim; paired difference -2.8, CI -7.3 to +3.9
Tuned arms S1 / S2 / S3 / S4 (macro-F1)0.585 / 0.671 / 0.697 / 0.668the last round, with a more positive-heavy mix and more epochs, got worse
Pre-registered gate (tuned beats untuned on validation)Not clearedso no tuned model was run on the test set

The question

Contract review is a good test for small models: the answers are checkable, the text is long, and a wrong or invented quote is costly. The question I wrote down before running anything was narrow. Can a small model (gemma-4-e4b, about 4B parameters), LoRA-tuned on the public CUAD contracts, land within five macro-F1 points of a larger model at two jobs: deciding which of 41 clause types a contract contains, and quoting the supporting text verbatim?

The protocol, the splits (347 train, 61 validation, 102 test) and the rules for what could be changed after seeing results were fixed in a pre-registration file first. The test set was held back for one final comparison.

The baselines

On the 102 test contracts, the larger model scored a macro-F1 of **0.783**. The untuned small model scored **0.715**. That 6.8 point gap is real: a contract-level paired bootstrap puts the difference at -6.8 points with a 95% interval of -10.0 to -3.0. The untuned model had a recognisable failure profile: decent recall, but only about three quarters of its quotes appeared verbatim in the contract.

Four tuning rounds

Contracts run to tens of thousands of tokens, so training used 1,000-token chunks, and at inference every chunk is read and the results are merged (a clause counts as present if any chunk says so). That merge rule was frozen before any test run. The tuned arms differ in how many negative chunks were kept and how long they trained, and all four were scored on the validation set, never the test set.

ArmMacro-F1PrecisionRecallVerbatim quotes
Untuned base0.7270.8240.86783.2%
S1 (110 steps)0.5850.9250.62196.4%
S2 (300 steps, fewer negatives)0.6710.9440.76197.3%
S3 (600 steps, fewer still)0.6970.9250.81697.1%
S4 (900 steps, fewest negatives)0.6680.9030.78497.3%

What tuning bought, and what it didn't

Tuning bought two things, immediately and consistently. **Precision** rose from 0.82 to about 0.93: the tuned models made far fewer false claims that a clause was present. And **quote fidelity** rose from 83% to 97%: the tuned models almost never invented text.

It did not buy **recall**. Every tuned arm missed more clauses than the untuned base, and the arm that came closest (S3, recall 0.816) still trailed it (0.867). The tuned models lean toward answering "absent". The rare categories, such as source-code escrow and unlimited-license clauses, were the weakest.

The best arm, S3, ended 2.8 points behind the untuned base on validation. The paired interval, from -7.3 to +3.9, spans zero. With 60 contracts the data cannot say whether S3 is slightly worse, equal or slightly better. What it can say is that it is not meaningfully better, which is what the gate required.

The last round made it worse

The first three rounds improved steadily, which suggested the answer was more of the same. The fourth round pushed further: an even more positive-heavy training mix and about 2.2 passes over the data. It got worse (0.668), losing both recall and precision against S3. That is the signal to stop. I ran out of cheap levers at this scale rather than out of patience.

I also got something wrong along the way. I had assumed the third round saw only part of the training data and that training longer was the obvious next step. It had already covered it about 1.3 times. That claim is corrected in the project write-up.

Where the gate landed

The pre-registered gate was that a tuned arm must beat the untuned base on validation before it earns a single run on the test set. None did, so none was run, and the main planned comparison, tuned small model against the larger model, did not happen. That question is **unanswered**, not answered no. I also did not run S3 on the test set after the fact, because that would turn a clean held-out split into an exploratory number.

Limits worth stating

- Four tuned arms were chosen on the same 61 validation contracts, so their validation scores are, if anything, optimistic.
- The larger model is a single sample from a mid-tier model, not a ceiling.
- One small base model, one seed, one prompt.
- A repair rule (appending a single closing brace when a chunk's JSON came back one brace short) was derived from validation failures and then frozen for every arm. It moved the first arm from 0.537 to 0.585.
- 102 test contracts cannot reliably separate a 2-3 point deficit from equivalence, so a smaller gap would have been reported as inconclusive rather than a win.

What I would take from it

1. **Tuning a small model trades completeness for fidelity.** If your cost is a wrong or made-up quote, the tuned model is the better extractor even at similar macro-F1. If your cost is a missed clause, it is not.
2. **Pre-register the gate, then respect it.** The most useful thing the protocol did was stop me from quietly running the test set on an arm that had not earned it.
3. **Watch for the point where a curve turns.** Three rounds of improvement made a fourth feel obvious, and it was not.
4. **Report the negative result with its intervals.** "Not cleared, and here is how uncertain we are" is a finding.

Nothing from this experiment is deployed. The data is the public CUAD set; the protocol, scores and scripts are in the project repository.