For the first phase of generative AI adoption, most law firms treated legal AI cost much like any other software expense. Buy licenses, then allocate seats, run a pilot and then negotiate an enterprise agreement with another line added to the technology budget.
Agentic AI is beginning to challenge that model.
A conventional AI assistant might summarize a document or answer a question. An AI agent can potentially undertake a sequence of activities: search multiple sources, analyze documents, call tools, compare results, reconsider earlier conclusions and assemble a final work product. For a law firm, that could mean reviewing a transaction data room, comparing hundreds of contracts, researching authorities, analyzing previous matters and producing a structured diligence report.
Each additional step requires computation. More importantly, much of that computation happens before the lawyer sees the answer.
This is why legal AI cost is becoming an economic issue.
One of the most important pieces of recent research comes from Bai and colleagues in How Do AI Agents Spend Your Money?, a study involving researchers affiliated with the University of Michigan, Stanford, MIT, Google DeepMind and All Hands AI.
The researchers analyzed eight frontier models performing agentic software engineering tasks. They found that agentic tasks consumed vastly more tokens than conventional reasoning or chat, with input tokens rather than final output tokens driving much of the cost. Token consumption was also highly variable. Repeated runs of the same task could differ by as much as 30 times, while spending more tokens did not consistently produce greater accuracy. Performance frequently peaked at an intermediate level of expenditure rather than continuing to improve as token consumption increased.
This research concerns coding agents and its numerical findings should not be directly generalized to law firms. The underlying mechanism, however, deserves the attention of legal technology leaders.
An agent does not simply generate an answer. It accumulates context. As Stanford's Digital Economy Lab explains in its discussion of the research, an agent can repeatedly reread the original task, previous responses and subsequent observations as it moves through a workflow. The context therefore grows as the task progresses.
That distinction matters enormously in legal work because legal tasks are document and knowledge intensive.
Ask an AI system to summarize one agreement and the computational boundary is relatively clear. Ask it to review an acquisition data room, identify change of control provisions, compare them with the firm's preferred position, assess material deviations and produce a diligence report, and the system may need to retrieve and process hundreds or thousands of pieces of information.
The economics of legal AI therefore increasingly depend not simply on how much an AI writes, but on how much information it must process before it can produce useful legal work.
A 2026 paper, Token Economics for LLM Agents: A Dual View Study from Computing and Economics, helps explain this shift. Chen and colleagues argue that tokens are no longer just units of information processing; in agentic AI, they also function as factors of production, a medium of exchange and a unit of account.
They frame agentic AI as a resource allocation problem: minimizing total cost while still achieving the required output quality. Their model also includes human cognitive labor, recognizing that prompting, alignment and review carry economic value alongside machine computation.
That distinction is crucial in legal services.
The cheapest AI interaction is not necessarily the most economically efficient legal interaction. A low cost model that produces an unreliable answer requiring extensive associate review, repeated prompts and partner correction could ultimately cost the firm more than a more expensive system that produces a reliable result with substantially less human intervention.
Legal AI is already demonstrating a different productivity model
In Lawyering in the Age of Artificial Intelligence, Jonathan Choi, Amy Monahan and Daniel Schwarcz conducted a randomized controlled trial in which law students completed realistic legal tasks with and without GPT 4. They found that AI assistance produced large and consistent improvements in speed, while improvements in the quality of legal analysis were smaller and less consistent.
More recent research by Schwarcz and colleagues examined newer AI approaches, including retrieval augmented generation and reasoning models. Their randomized controlled trial found statistically significant productivity gains of between 50% and 130% in five of six tested legal tasks, alongside improvements in work quality.
These studies do not establish the economics of commercial law firm deployments, but they illustrate why the commercial model matters.
If AI reduces the lawyer time required to produce some categories of legal work, then the economic input does not simply disappear. Part of it moves from human labor toward technology, inference, data, knowledge infrastructure and human validation.
For law firms, that has profound implications.
The profession has spent decades measuring the production of legal work primarily in units of lawyer time. Agentic AI introduces another production input whose consumption may vary according to the complexity of the task, the amount of information processed, the model selected and the number of actions the system performs.
In other words, AI can reduce the cost of human production while simultaneously creating a new category of machine production cost.
In June 2026, Legora announced that it was moving Agent Pro, its most capable agent product, to consumption based pricing. The significance is not simply that one vendor changed the pricing of one product. It demonstrates that at least part of the legal AI market is beginning to separate the economics of agentic work from conventional per user software licensing.
The broader market context is also instructive. The Financial Times reported in August 2026 on Legora's move toward consumption based pricing while examining the economics and competitive pressures affecting specialist legal AI providers.
What about Harvey? It would be wrong to claim that Harvey will start charging law firms by the token. What can reasonably be said is that the underlying economics make some form of usage sensitive pricing increasingly plausible across the legal AI market as agentic workloads become larger and more computationally intensive.
That does not mean law firms will necessarily receive invoices denominated in millions of tokens. Vendors have many ways to translate underlying computational consumption into a commercial model. They could use agent credits, workflow allowances, pooled consumption, usage tiers, premium model charges, task based pricing or subscriptions that include a consumption threshold.
The unit visible to the customer may therefore remain different from the unit driving the vendor's underlying cost.
For law firm leaders, the strategic point is more important than the billing terminology: as AI moves from answering questions to performing substantial workflows, firms should not assume that unlimited agentic work will always behave economically like a conventional fixed price software seat.
This brings us to an aspect of legal AI economics that has received considerably less attention - what is the AI consuming?
It is easy to think about tokens as an abstract technical measure. In enterprise legal AI, however, a large proportion of those inputs can represent something much more familiar to knowledge professionals: context.
That context may include instructions, retrieved documents, previous conversation, tool results, research sources, matter information and knowledge from the firm's internal systems.
The token economics research describes the context window as a scarce resource and examines memory architecture, information retrieval and context management as important dimensions of agent efficiency. It also reviews approaches that impose explicit token budgets on retrieval and filter information according to relevance rather than repeatedly replaying large quantities of context.
This provides an important connection between AI economics and knowledge management.
For a law firm, the question is no longer simply whether the AI can access the firm's knowledge. The question is whether it can identify the right knowledge efficiently.
Consider the typical information environment of a large firm. Valuable knowledge may be distributed across the document management system, SharePoint, Microsoft Teams, email, precedent collections, matter files, CRM, experience systems, research services and specialist practice applications.
The relevant knowledge may exist somewhere in that environment, but existence is not the same as machine readiness.
A precedent may apply only in a particular jurisdiction. A research note may have been superseded by a change in law. A model agreement may be authoritative for one practice group but inappropriate for another. Two similar documents may reflect completely different negotiating positions. A previous matter may contain useful reasoning but also be subject to client confidentiality restrictions.
Legal knowledge therefore requires more than semantic similarity. It requires authority, provenance, permissions and context.
Imagine a lawyer asks an AI agent to find the firm's preferred limitation of liability wording for a particular type of transaction and explain what the firm normally recommends.
In a poorly structured knowledge environment, the system might discover multiple precedents, several historic versions, previous transaction documents containing similar clauses, practice group variations and meeting notes. Somewhere among them is the approved precedent, but the repositories provide weak signals about which source is authoritative.
An AI system can potentially compensate for that ambiguity by retrieving more information, processing more documents, comparing more alternatives and conducting additional reasoning. But that compensation consumes computational resources and may also increase the amount of human verification required.
The traditional KM problem therefore acquires a new economic dimension. When lawyers cannot find trusted knowledge, they lose time. When AI cannot efficiently identify trusted knowledge, it can consume additional compute. Both are costs of poor knowledge infrastructure.
This connection is consistent with established KM research. Aviv, Hadar and Levy's research into knowledge management infrastructure argues that effective KM infrastructure needs to integrate knowledge procedures into the operational flow of knowledge intensive business processes rather than treating knowledge management as a separate activity. Their framework combines technological, cultural and knowledge process dimensions and emphasizes structuring organizational knowledge assets around operational needs.
For law firms, where the service itself is knowledge intensive, that principle becomes particularly important when AI is another consumer of the firm's knowledge.
KMWorld's Building a KM Foundation for Enterprise AI identifies data integration and interoperability, knowledge graphs, metadata management and tagging, and collaboration and knowledge sharing as important building blocks for enterprise AI. It also highlights persistent problems such as fragmented repositories, information silos and the quality and reliability of organizational information.
For legal AI, the objective should be to retrieve the smallest useful set of authoritative, relevant and permission appropriate knowledge required for the task.
LegalBench, the interdisciplinary benchmark published at NeurIPS, was created precisely because legal reasoning is not one homogeneous capability. Its 162 tasks cover six different forms of legal reasoning and were developed with substantial input from legal professionals.
The implication for knowledge architecture is important. Giving an AI system a large volume of legal documents is not the same as giving it the contextual signals required to determine which knowledge is relevant to a particular legal question.
A law firm's product is professional knowledge and judgment. If AI increasingly contributes to producing that service, its economics become intertwined with matter economics.
Under the traditional model, the cost and value of legal production were heavily associated with lawyer time. If an associate spent 30 hours reviewing contracts, those hours were visible in the matter economics. AI may allow the same category of work to be performed with considerably less human time, but the underlying cost of production has not vanished. It has changed composition.
The firm may now be paying for AI subscriptions, agent consumption, model inference, integrations, legal information services, knowledge infrastructure and human validation.
As AI performs a greater proportion of substantive workflow activity, firms may therefore need to understand at least some AI expenditure as a cost of legal service delivery rather than simply undifferentiated IT overhead. This is particularly significant for fixed fee and alternative fee matters.
If a firm agrees a fixed fee on the assumption that AI will reduce lawyer hours, but the agent then consumes unexpectedly high computational resources while analyzing a large document set, a new form of margin variability enters the matter.
Conversely, measuring only AI consumption would also be misleading. A more expensive AI workflow could still produce better matter economics if it substantially reduces lawyer time while meeting the firm's quality and risk requirements. The metric that ultimately matters is therefore not cost per token. It is closer to cost per trusted legal outcome. That combines machine consumption with the human effort needed to produce and validate the result.
This leads to a broader proposition for law firm KM leaders.
Knowledge management has traditionally been justified through improvements in search, reuse, lawyer productivity, precedent quality, expertise location, knowledge retention and risk management. Those benefits remain important.
Agentic AI adds another potential source of value: making organizational knowledge more efficient for machines to consume.
If an AI agent can identify the authoritative precedent quickly because the content has clear ownership, metadata, jurisdiction, lifecycle status and relationships to relevant matters and expertise, it may need less irrelevant context than an agent forced to infer those distinctions from a large collection of poorly differentiated documents.
That does not mean knowledge management can guarantee lower AI costs. It means the architecture of knowledge becomes one of the variables influencing how efficiently AI can retrieve and use enterprise context.
The token economics research is particularly interesting here because it treats memory as something closer to reusable capital than disposable context. Structured memory and efficient retrieval can reduce the need to repeatedly reconstruct information during agent workflows.
There is a direct conceptual parallel with enterprise KM. A law firm should not have to make every lawyer rediscover what the firm already knows. Increasingly, it should not make every AI agent rediscover it either.
Where AtlasFuse fits
AtlasFuse provides knowledge infrastructure for firms operating in the Microsoft ecosystem. Rather than requiring all enterprise knowledge to be moved into another repository, AtlasFuse connects and structures knowledge across Microsoft 365 and connected systems, creating a governed and permission aware knowledge layer that can be reused by people and AI.
For law firms, the opportunity is to make knowledge more understandable before it reaches Microsoft Copilot, specialist legal AI platforms or intelligent agents.
That means capturing and enriching knowledge with the context machines need, connecting content with expertise and related concepts, identifying authoritative sources, preserving permissions and governance, and making the resulting knowledge reusable across multiple AI experiences.
The economic argument is that better prepared knowledge could reduce unnecessary retrieval and repeated reconstruction of context while making AI interactions easier to validate. This should be treated as a hypothesis to measure rather than a guaranteed cost saving.
The objective is not simply to minimize tokens but to increase legal value per unit of AI consumption.
Your legal AI bill is growing, look beneath the model
There are many technical ways to manage legal AI cost. Model routing, caching, prompt design, context compression, knowledge graphs, budget controls and vendor negotiation all have a role.
But firms should also investigate why the AI needs to consume particular context in the first place.
If an agent must repeatedly search through duplicate, outdated and poorly contextualized information before finding the knowledge that matters, part of the AI bill is effectively paying for information friction.
That leads to a question every law firm investing seriously in AI should ask:
How hard are we making AI work to understand what our firm already knows?
The answer brings legal AI economics directly into the domain of knowledge management.
A firm's knowledge has always been one of its most valuable assets. In the agentic era, the structure, quality and accessibility of that knowledge may also influence how efficiently the firm can turn machine intelligence into legal value.
That makes knowledge infrastructure more than a KM investment and part of the operating model for legal AI.
Legal AI cost is the total cost associated with using AI to produce legal work. It can include platform subscriptions, consumption or agent usage, model inference, integrations, knowledge and data infrastructure, external information services and the human time required to review and validate AI outputs.
Why can AI agents cost more than AI chat?AI agents can perform multiple steps, retrieve information, call tools and repeatedly process growing context during a workflow. Academic research on agentic coding found substantially higher token consumption than conventional LLM interactions, with input context responsible for much of the expenditure. The precise ratios from coding research should not be assumed to apply to legal work.
Are legal AI platforms moving to consumption based pricing?At least some are. Legora announced consumption based pricing for Agent Pro in June 2026. This does not establish that the entire legal AI market will adopt the same model, but it provides evidence that agentic legal AI is beginning to be commercialized differently from conventional per user software.
Can knowledge management reduce legal AI cost?There is not yet sufficient empirical evidence to claim a guaranteed reduction. However, better metadata, authoritative content, lifecycle management, permissions and more precise retrieval can make it easier for AI systems to identify relevant context. Whether that translates into lower costs will depend on the architecture, vendor pricing model, workflow and model being used.
What should law firms measure?Rather than measuring tokens alone, firms should consider the total cost required to achieve an acceptable legal outcome. That includes AI consumption, lawyer review, retries, corrections and other human involvement. The emerging token economics literature similarly frames agent optimization as minimizing total cost subject to a required quality threshold.