Bo[a]B is an AI system that runs on a computer in your office. The corpus stays. Every action writes to an audit log you hold and can hand to a client, a regulator, or opposing counsel — without asking us for anything.
The objection is rarely “your security is worse than ours.” Often it isn't. The objection is a sentence the firm has to be able to say out loud — to a client, a regulator, or opposing counsel:
Here is everywhere that matter file went. Here is who touched it, when, and under what authority. Here is the record proving nothing left the perimeter that wasn't meant to.
That sentence needs an audit trail the firm holds, for a retention period the firm sets. Not a vendor attestation about subprocessors and regions that can be revised in a policy update you learn about by email.
Confidentiality obligations don't transfer to somebody else's compliance page.
Documents are processed on your hardware, under your keys. There is no upload step, because there is nowhere to upload to.
Anything that does leave is filtered at the egress point, and the filter is inspectable. It fails closed — see below, because how it fails is the part that matters.
Every action is recorded in a cryptographically chained log. You can hand it to an auditor directly, without a support ticket.
A daily ceiling that stops work rather than a number on a dashboard. It raises an error before spending, not after.
No subscription that can be repriced. No vendor that can read your files. No migration project when the contract ends.
None of these mechanisms are novel, and I'd rather say so than have a partner discover it. Redaction at an egress boundary has been shipping commercially for years. Any firm evaluating this should find who else does it and make them compete.
What is harder to find at this size is the combination: you own the hardware, you hold the keys, you keep the log, and the vendor can't read the corpus.
The redaction engine is deterministic and fails closed. On a corpus it has not been tuned for, it over-redacts substantially — it will mask things that did not need masking, and the output is noisier for it.
That is the deliberate default, and for privileged material it is the correct one. A gate that errs toward leaking is not a gate. Tuning against a firm's own corpus and vocabulary is part of an engagement, and it is where the false-positive rate comes down.
I'd rather tell you the shape of the failure than quote you a headline accuracy number. If you want the measured recall and precision against a third-party-annotated legal corpus, ask me and I will send the run, the method, and the caveats — including where it performs worse than you would like.
Only the second one is a breach. A missed name identifies someone on its own. A missed date does not — unless the city and the profession were missed alongside it. Standard redaction scoring counts both as the same miss, which tells you how the engine did against an answer key, not whether anyone could be re-identified.
So we measured re-identification directly, on the keyed corpus the pipeline would actually emit, against the hardest adversary — the model provider itself, holding every document the firm has ever sent.
No pair of documents from different matters shared a unique identifier. No case number, no matter number, no email, no account tail. That is structural rather than lucky: those identifiers are scoped to a matter by construction.
The residual was 5.9% — a shared personal-name fragment, where the same person genuinely appears in two matters. Rotating keys per matter takes that to zero while preserving the within-matter linkage the work requires.
We also decomposed our own first number, which is the part worth knowing about us. The naive reading said 72% of document pairs were linked. Most of that was the fail-closed catch-all joining on ordinary English — one pseudonym resolved to the word “as.” The real figure is 5.9%. An outside technical advisor will do that decomposition eventually. We would rather have done it first, and told you.
What we do not claim: that this makes re-identification impossible. Dates, amounts, jurisdictions and document-type vocabulary are preserved because the downstream work needs them, and in combination they constitute a docket selector. We treat that as accepted residual risk, not a solved problem.
A California practice — five lawyers, all remote — could not answer a basic question about their own files: what is in here, and what state is it in?
Bo[a]B read the titles and metadata across the corpus, stripped the identifying material, and produced a shared map all five could work from.
Not a pilot deck. A thing that ran, on their material, and changed what the practice could see about itself.
The largest firms are running several bespoke AI builds at once — one for M&A, another somewhere else — with internal teams to staff them. They don't need an appliance. They need headcount.
Bo[a]B is for the firm carrying the same confidentiality obligation and the same client pressure to adopt AI, without a department to build it. Roughly AmLaw 100–200, or any practice where the corpus is privileged and you know the IT team by name.
Year one runs $17K–$79K depending on scope. Currently in beta with a small number of firms.
I'm Patrick Inverso. I ran P&L for two decades — a marketplace I took to profitable in six months and held for nearly nine through an exit, and a seven-company division I merged into a single book at Cast & Crew. Before that, revenue work at Google, The New York Times, and MoMA.
For the last two years I've been building Bo[a]B: the fleet, the orchestration, the audit chain, the spend caps, the redaction boundary. I operate it myself.
That matters for one reason. There are three questions worth asking about any AI system a firm depends on: what does it cost, what does it leak, and what does it refuse to do. All three require a running system and a bill in your own name. I can answer all three about mine, and show you the log.
I'd rather show you the thing running than send a deck. Thirty minutes, and you'll know whether it fits.