In my last piece I argued that legal work is configuration, and I ended on the two stages ahead: delegated execution, where the lawyer configures and the agent works unsupervised inside those parameters, and the stage after that, where the configuration itself is the representation. A few people asked the right follow-up. What actually has to exist for delegated execution to work? Not in theory. In a product, with a login.
Look at what the market has built so far and the gap shows up fast. We have contract lifecycle tools that manage documents. We have e-sign. We have automation that assembles templates, and now agents that can run a negotiation playbook end to end. Every one of those products treats the document as the unit of work. But watch how the judgment behind the document is stored. The indemnity cap lives in the template. The reason it's that cap lives in the lawyer's head, or in an email from 2023, or in a playbook doc that three people edit and nobody versions. The industry built elaborate custody chains for paper and left the actual legal judgment in an unversioned free-for-all.
That arrangement survives the supervised era because the lawyer reviews everything anyway. It collapses at delegated execution. If an agent is going to work unsupervised inside my parameters, I need to know exactly what those parameters are, which version the agent ran, and why they say what they say. Ask a lawyer today to produce the reasoning behind a fallback position their template has carried for two years and most can't. Now imagine a client, a court, or a malpractice carrier asking the same question about a contract the lawyer never touched.
The product that has to exist next is a system of record for legal judgment. Call it version control for the configuration layer. It does a handful of things, and all of them are about making supervision demonstrable rather than asserted.
It captures parameters as data, not prose. Every position in the playbook is a typed value with an owner, a timestamp, and a written reason. Down to who set that cap, when, and what fact pattern justified it. It versions those parameters the way engineering versions code, with diffs a lawyer can actually read, so when a position changes there's a record of what changed and who changed it. It keeps provenance: every document the agent produces traces back to the exact parameter version that produced it, no exceptions. It builds audit in as a first-class feature, because the lawyer's job shifts from reviewing every output to sampling outputs against the configuration, and sampling only works if the tool schedules it, records it, and surfaces the misses. And it watches for drift. When the case law moves, or market terms move, the configuration goes stale, and someone has to own that diff. Today nobody does, because today's tools don't know the configuration exists.
There's a regulatory reason this product is the unlock and not just good hygiene. The bar's supervision rules already reach nonlawyer assistance, and the ABA read them onto generative AI in Formal Opinion 512. Supervision is a duty you have to be able to show, not just claim. "I configured it carefully" is a claim. A versioned parameter set, a provenance trail, and an audit log is a showing. The first malpractice dispute over an agent-negotiated contract will turn on exactly that difference, and the lawyers who can produce the record will be the only ones whose agents get to keep working unsupervised.
Who builds it? I'd bet against the incumbents. The lifecycle tools are document-centric to their core, and retrofitting a judgment layer onto a document database is a rewrite, not a feature. My guess is it comes from the people already practicing this way, the same way the best contract automation came from lawyers who got tired of their own templates. The firms that treat their judgment as versioned infrastructure will compound it. Everyone else will keep renegotiating from memory and calling it experience.
The agents are ready for delegated execution. The record keeping isn't. That's the product.