Judgment is hired, not scored.
Legal AI systems take positions on open questions that no benchmark measures. Newcomb is built so the only position delivered is the one your counsel actually took.
Why we built it this way
Ambiguity is why the practice of law exists.
Written rules cannot anticipate every fact pattern. Someone has to decide what a rule means in this instance — whether this vendor, this data, and this purpose are what the policy meant by “necessary.” That decision is legal judgment, and companies pay for it precisely because it cannot be reduced to a lookup.
Consider how the same companies select outside counsel. No general counsel has ever chosen a law firm because it was less likely to misstate the law. Every firm on a corporate panel knows the law; knowing it is the price of admission. Panel decisions turn on whose read you want on a close call. The profession has evaluated judgment for a century without a benchmark — through prior matters, referral, and the knowledge that a named partner stands behind a position. Crude as those proxies are, they measure the right thing.
Legal AI now takes positions on the same open questions, and nothing scores them. Ask a system a close question and it will do one of three things: commit to a reading, hedge, or decline and route the question to a lawyer. Declining may well be the right answer. But the output rarely tells you which of the three happened. It arrives as finished prose, carrying a risk posture nobody set. And when an answer looks polished, it is easy to stop checking the reasoning behind it — and checking the reasoning is the job.
Deliver the position counsel took. Do not propose one nobody set.
A system that delivers a position the lawyer has already taken produces a different document, and a different record, than one that proposes a position for the lawyer to edit. The first kind carries the lawyer's judgment. The second kind invites the lawyer to stamp a stranger's. Newcomb is deliberately the first kind.
It begins with the policies, memos, and trainings Legal already paid to create. A model may help extract their structure; it cannot authorize a response. When existing guidance does not resolve a real question, Newcomb brings it to counsel once, with two straightforward questions: What do you think? What would make you think differently? The second question is the important one. The answer to it — the facts that would change counsel's mind — is what makes a one-off answer safe to reuse.
Counsel can answer just that employee and identify a conditional answer worth reusing. Putting it into use still requires separate review and testing. If counsel later activates automatic delivery for a separately tested class of approved questions, the next applicable question can receive the exact approved response; uncertainty still returns to Legal. A question that cannot be matched safely returns to Legal even when guidance may exist.
Our design principles.
- Automate structure and repetition, not legal judgment. Models may propose structure, and, where matching is separately activated, candidate matches. Counsel alone approves the response and the conditions under which it may be reused.
- Never substitute software's risk tolerance for counsel's. No document holds every answer, and no software should quietly decide how much risk your company takes. That decision belongs to a person with a name.
- Make non-answers visible. The supervised pilot shows counsel every submitted question — answered, unanswered, and declined. A success-only dashboard is not evidence of safety; it is evidence of a filter.