The Ivo Blog
→
Introducing Sage: An Open-Source Model Trained for Contract Work
Product news

Introducing Sage: An Open-Source Model Trained for Contract Work

Kelsey Eisen
Kelsey Eisen

Contract work is a difficult task for most LLMs. General purpose AI models excel at short, one-off tasks, particularly those that require broad knowledge and quantitative, non-context-dependent analysis. Contract work is the opposite of this. Contracts are long and complex, requiring nuanced judgment, layering of tasks, and heavy dependence on context. Often there is no clean, single “right answer” to a contract question, only a better judgment call requiring nuanced analysis. General purpose models have improved enormously, but none of them can yet perform this kind of work the same way an experienced human attorney could. 

We believe that we will see significant gains in AI performance from specialized research efforts to teach models the specific work of contracting. Ivo has started doing this research, and today we're sharing the first major result.

Sage is our free, open-source model, post-trained for long-horizon contract work. It's available to download today, along with the method we used to build it and the full dataset we used to train it. Read on for details on why we built it, what we found, and why we chose to release it openly.

What is post-training, and why does it matter?

Post-training is how a model learns a specific job. One way to frame it is that a general-purpose model is like a recent law school graduate: trained in many different areas of law, but not yet an expert in any. A post-trained model is like an attorney who has practiced for some time and specializes in their practice area. While they may be equally intelligent and capable, the second knows how this specific work is done, and can do it faster and more effectively. 

When applied to contract work, post-training a model changes how it behaves across a whole assignment. A post-trained model can learn to use its tools more efficiently, pull in the right context at the right moment, reach conclusions in fewer steps, and produce a finished deliverable that is much closer to the desired result than what a non-post-trained model would produce. It also makes performance measurable against the metrics that matter; in this case, whether the work product meets the standard a legal team would expect.

We think this is one of the most important and least understood levers in legal AI. General purpose models are a strong starting point, but the biggest gains come when we teach them the specialized work of contracts, to the standard a legal team expects.

How we built Sage

Sage is built on DeepSeek V4 Flash, a publicly available open model. We post-trained it using reinforcement learning, a method that improves a model by rewarding outputs that score well against defined criteria.

Our training ground was the Legal Agent Benchmark (LAB), a public set of long, multi-step legal assignments covering general legal practice and contract work. We trained Sage on 1,040 LAB tasks. For each one, the model worked through the assignment in the benchmark's own environment, using a small set of tools to read, write, edit, and search files. Its output was then graded against expert-developed criteria, and those scores guided the model toward better work over repeated rounds of testing.

To make sure Sage learned the work rather than memorizing answers, we set aside 180 tasks from the 1,040 and kept them out of training entirely. Only after training did we use them to test the model (picture giving a student an exam made of problems they've never practiced). We also kept tasks built on the same agreement together—either all in training or all in the test—so Sage was never tested on a document it had already seen.

No customer data was used at any point, and Sage was trained only on public tasks. We built Sage in partnership with River AI, who provided valuable training infrastructure support. Training a model of this size takes serious computing capability, and River's support helped make this work possible.

What this means for cost

Contract work happens at volume, across thousands of agreements, so what a model costs per task matters as much as how well it performs. A model that does excellent work at an unsustainable price may be an interesting research result, but it’s not something a real legal team can practically build on.

Thus, in making Sage, we made cost a central metric. We ran Sage alongside 13 other models, including the leading frontier models, on the same 180 held-out tasks. We then tracked cost per task, a measurement that reflects each model's own token usage at its provider's list price.

The Results

We scored each model two ways. The first is the criterion pass rate, or the number of individual requirements a model satisfies across assignments. The second is the all-criteria pass rate, a much stricter measure that tracks the percentage of tasks that the model completed while satisfying every requirement. 

General legal tasks

On general tasks, Sage's all-criteria pass rate of 15.2% is comparable to Fable 5.1's at 15.7%. However, a remarkable difference is seen when we compare the price. Sage costs $0.40 per task, while Fable 5.1 costs $92.46. This puts Sage at an almost identical effectiveness, just at about 0.4% of the cost. 

Sage also exceeded Kimi K3, GPT-6 Astra, and GPT-5.6 Sol on the same measure. Among the models that reached a criterion pass rate above 90%, Sage was the least expensive by roughly a factor of nine.

Chart 1: LAB General, criterion pass rate vs. cost per task
Chart 2: LAB General, all-criteria pass rate vs. cost per task

Contract tasks

Contract work is where post-training made the biggest difference. Sage met 91% of grading criteria, compared to 70% for the DeepSeek V4 Flash base model, representing a gain of more than 20 points from training on contract work alone. It did this at a cost of $1.24 per task, while Fable 5.1 cost $142.39 on the same suite of tasks. Among models above the 90% mark, the next least expensive costs roughly eight times as much.

Chart 3: LAB Contracts, criterion pass rate vs. cost per task
Chart 4: LAB Contracts, all-criteria pass rate vs. cost per task

Overall, Sage offers the best balance of accuracy and cost of any model we tested. On both LAB General and LAB Contracts tasks, and on both ways we score the work, Sage sits on the frontier of performance against cost (as represented by the light green box in the graphs above).

A more efficient way of working

Post-training changed how Sage works, not just how well it scores. Across all 180 held-out tasks, the average assignment took 25 turns, down from 31 for the base model, and Sage used 21% fewer tokens. In practice, Sage learned to reach conclusions in fewer steps and to use its tools more effectively. That efficiency is also why it costs less per task than the model it was built on.

Holding up beyond the training set

We also tested Sage on benchmarks outside its training. On RedlineBench, an independent benchmark for contract redlining, its score rose from 30.8 to 45.7. And on LegalBench, a broad test of short-form legal reasoning, it scored 83.0 against the base model's 83.1. Specializing in contract work didn't come at the expense of general legal reasoning.

Why we're openly releasing Sage 

Most AI used for legal work is closed. Some legal AI companies have announced post-trained models, but few have made theirs open-sourced. Ivo is the first to release an open-source model post-trained for long-horizon contract work. 

We chose to make Sage open-source for three reasons.

Progress in a field like this depends on shared work. There's remarkably little open research on AI for contracts. We'd rather contribute to a foundation for this work than to guard one, because the whole field—and the field of law at large—improves faster when people can build on each other's results.

Transparency lets people check our work. Alongside the model, we're publishing our training method and evaluation record: task lists, evaluation transcripts, deliverables, and the judge's decisions. Anyone can inspect what the model did and how it was graded, rather than just taking our word for it.

Every legal team's work is different. Contract complexity varies enormously, from a Fortune 500 legal department to small businesses to agreements between individuals. An open model can be fine-tuned on a team's own data, adjusted with its own rules and safeguards, and adapted to its own use cases.

Sage is released on Hugging Face under an MIT license, which generally permits commercial use, modification, and redistribution. It remains subject to the terms of its underlying base model. To our knowledge, Sage is the first open model from a legal AI company post-trained for contract work. Open legal models have existed before, and we're building on that foundation.

What this means for what's next

Advancing contract AI takes building the technology, measuring it honestly, and pushing the entire industry forward. We see Sage as an early step in this important work, and we're excited about where it leads. 

In the meantime, we’ll keep researching, keep sharing, and keep being your trusted source for contract AI. 

Download Sage on Hugging Face here.