Quick answer: the single biggest blocker to AI adoption in law firms isn't capability — it's confidentiality. Sending privileged client material to a third-party API raises questions about duty of confidentiality, client consent, and where data ultimately resides that many firms are unwilling to answer with "we think it's probably fine." A self-hosted LLM removes the question entirely: the model runs on infrastructure the firm controls, and privileged material never leaves it. That's why this is the fastest-growing question we get from legal clients, and it's a genuinely good answer. Here's the practical picture.
Why confidentiality is the blocker
Lawyers operate under a duty of confidentiality that's broader than privilege and doesn't have a "but it was convenient" exception. When an associate pastes a client's contract into a public AI tool to get a summary, that material has been transmitted to a third party — and the questions that follow are uncomfortable: What are that provider's terms? Is the data retained? Could it be used for training? Who at the provider could access it? Does this require client consent? Does our engagement letter cover it? Would we be comfortable explaining this to the client, or to a bar committee?
Most firms, when they actually sit with those questions, conclude that routine use of public AI tools on client material is not something they can defend. So they do one of two things: they ban AI outright (losing real productivity and watching people use it anyway on personal accounts), or they proceed nervously without a clear policy. Neither is good. The technology promises meaningful gains on exactly the work that consumes the most expensive hours in the building — document review, summarization, research synthesis, drafting — and firms are stuck on a data question rather than a capability one.
How self-hosting resolves it
A self-hosted LLM changes the analysis at its root. The model runs on the firm's own server or private cloud instance. When a lawyer asks it to summarize a deposition or review a contract, the document goes to a machine the firm controls and the response comes back from that same machine. Nothing is transmitted to a vendor. There is no third-party terms of service governing the data, no retention question, no training question, and no external party with access. The answer to "where did the client's material go?" is "nowhere — it stayed on our system."
That's a categorically different posture than "we use a vendor with good privacy terms," and it's one that's straightforward to explain to a client, a partner, or a regulator. It also tends to satisfy the increasingly common client demand — particularly from corporate clients with their own compliance obligations — that outside counsel disclose whether and how AI is used on their matters. A firm that can say "our AI runs entirely in-house and your material never leaves our infrastructure" is in a strong position with sophisticated clients rather than an awkward one.
What a self-hosted LLM can realistically do for a firm
Set expectations correctly and the value is substantial. Self-hosted open-weight models handle the language-heavy, high-volume work that fills legal practice well: summarizing long documents, depositions, and case files; extracting structured information — dates, parties, obligations, clauses — from unstructured documents; first-pass review that flags relevant passages, anomalies, or missing provisions for an attorney to examine; drafting routine correspondence and first versions of standard documents; and searching the firm's own accumulated work product in natural language so past knowledge is actually reachable.
What they are not is a substitute for legal judgment, and this needs to be stated plainly because the profession has already produced cautionary tales. AI output — self-hosted or otherwise — must be verified by an attorney before it's relied on, filed, or sent. Courts have sanctioned lawyers for submitting AI-fabricated citations, and no hosting arrangement changes that duty. The correct framing is that a self-hosted LLM is a very fast, very tireless assistant working on privileged material safely, whose output a lawyer reviews. That framing delivers the productivity without the professional risk.
The practical deployment picture
Deployment for a firm is more approachable than most partners assume. The model runs on a server the firm controls — either hardware in the office or a dedicated instance in a cloud account the firm owns — with a runner that exposes it to internal tools. Sizing is a real decision: smaller models handle summarization and extraction capably on modest hardware, while heavier reasoning benefits from stronger, GPU-equipped machines. Most firms are well served starting with a model sized to their highest-volume tasks rather than buying for the most demanding hypothetical.
Around the model you want the things that make it usable and defensible: access controls so only authorized people can use it, logging so you can demonstrate what was done, integration with your document systems so lawyers aren't copying and pasting, and a written AI policy that defines what it may be used for and what always requires attorney review. That governance layer is not optional — it's what turns a capable tool into something the firm can stand behind. We build these deployments with those controls from the start, because for a law firm the compliance architecture is the product.
Where it fits with the rest of a firm's AI
Self-hosting doesn't have to be all-or-nothing, and for most firms it shouldn't be. The sensible architecture is to route work by sensitivity: privileged and client-confidential material goes to the self-hosted model, while genuinely non-sensitive tasks — drafting a marketing email, summarizing a public court opinion, general research on public sources — can use whatever tool is most capable, including cloud APIs. That hybrid gives you frontier capability where it's safe and airtight confidentiality where it matters, rather than forcing a single compromise across everything.
The broader point for firms weighing this: the reason to look at self-hosted AI isn't ideology about open source, it's that it makes AI usable on the work that actually matters to your practice. Intake, billing, and scheduling automation are valuable and don't touch privilege — those we cover in our law firm automation playbook. But the deep productivity gains sit in the document-heavy work, and that work is privileged. Self-hosting is how you unlock it without asking your clients to trust a vendor they never chose.
FAQ
Is it ethical for a law firm to use AI on client matters?
Yes, with proper supervision and confidentiality safeguards — the profession's technological-competence expectations assume lawyers will use modern tools competently. The ethical problems arise from unsupervised reliance on AI output (filing unverified citations) and from transmitting confidential material to third parties without adequate consideration. Self-hosting addresses the confidentiality half directly; attorney review of every output addresses the supervision half.
Does self-hosting really mean client data never leaves the firm?
That's precisely the point. With a self-hosted model, inference happens on infrastructure the firm controls, so prompts and documents are processed locally and no third party receives them. The important caveat is that this is only true if the whole system is built that way — a self-hosted model connected to some other cloud service could still leak data, so the architecture around it matters as much as the model.
Are open-weight models good enough for legal work?
For the high-volume language work that dominates practice — summarizing, extracting, first-pass review, drafting routine documents, searching your own work product — capable open-weight models perform well and are improving rapidly. For the most demanding reasoning, leading proprietary models still tend to lead. Since the highest-value legal use cases are largely the document-heavy ones, self-hosted models cover a great deal of what firms actually need.
What about using a cloud provider's enterprise terms instead?
That's a legitimate alternative many firms choose — enterprise agreements with strong confidentiality and no-training terms meaningfully reduce risk. The difference is that it's a contractual protection rather than an architectural one: you're trusting a vendor's terms and controls. Self-hosting removes the third party from the equation entirely. Firms with the most sensitive matters, or clients who demand it, tend to prefer the architectural answer.
How much does a self-hosted LLM cost a small firm?
It's a fixed cost rather than a per-use one: hardware or a dedicated cloud instance, plus the setup and ongoing maintenance. Sized to a small firm's realistic workload it's a manageable investment, and because it doesn't meter per document, heavy use doesn't increase the bill. Weighed against the billable hours recovered on document review and drafting, firms typically find the economics work quickly.
Can a small firm afford a self-hosted LLM?
Yes — the economics are more approachable than most partners expect. It's a fixed cost (a suitably specified server or a dedicated cloud instance, plus setup and maintenance) rather than a per-document charge, so heavy use doesn't increase it. Sized to a small firm's realistic workload it's a manageable investment, and measured against the billable hours currently consumed by document review, summarization, and drafting, firms generally find it pays back quickly.
What should we tell clients about our AI use?
Increasingly, sophisticated clients ask — and being able to answer clearly is an advantage. A firm running self-hosted AI can state plainly that AI assists with document review and drafting, that all processing happens on infrastructure the firm controls, that no client material is transmitted to third parties, and that an attorney reviews everything before it's relied upon. That's a far stronger position than either an evasive answer or a blanket "we don't use AI" that may not survive contact with reality.
Want AI your firm can actually use on privileged work? We build self-hosted deployments with the access controls, logging, and governance a firm needs — owned by you. Get a free audit, or read the law firm automation playbook. Related: self-hosted AI explained and the AI policy template.
