Key Takeaways
  1. Decreasing marginal differentiation in LLM models.

  2. Narrowing in benchmark leadership metrics and clustering in capabilities.

  3. A likely collision course with hyperscaler partners up the software stack.

  4. Inference is inflecting upward from enterprise adoption of AI and reasoning, a very positive mix shift to monetization.

Introduction

Three years following the launch of ChatGPT, AI is being implemented across broad enterprise software markets. At the center of this shift is the LLM: the foundation model that provides inference to deliver business insights and automate tasks.

To support the financial commitments LLM vendors are making for data center capacity, and to realize their full market opportunity, they need to evolve their business models. The LLM market is on a path toward commoditization in the long tail of models, and turning the 5 leading oligopoly vendors into a durable and profitable business requires moving up the tech stack to deliver higher value.

This is the path Microsoft took with Windows and Oracle took with its database, leveraging a foundational layer to build comprehensive platforms, rich ecosystems, and durable market leadership. The question is whether today’s leading LLM vendors can do the same, and what happens when that ambition collides with the software platforms their hyperscaler partners have already built.

The stakes are significant: a substantial amount of market capitalization and a considerable TAM as the prize. In the following paper we will explore:

  • The Commoditization Problem. Why benchmark leadership is fleeting and LLMs alone may not be a durable business.
  • The Limits of LLMs. What foundation models are good at, and where they fall short.
  • The Stack Imperative. How value moves up the tech stack, and the historical precedent for this path.
  • The Inference Engine. What’s driving inference demand and why it matters for LLM business models.
  • Data Center Realities. The infrastructure investment required to meet inference demand.
  • The Collision. What happens when LLM vendors, third-party AI-native software companies and hyperscalers compete for the same platform position.
  • Differentiators Beyond the Model. What creates durable competitive advantage.
  • Implications. Who is positioned to win, and what it takes to get there.

The Commoditization Problem

The LLM market is on a path to decreasing marginal returns in innovation, imposed by the scaling laws in AI. The result is increasingly rapid benchmark leadership changeover, something like the frequency of leadership changes in a close basketball game.

We are seeing a narrowing in benchmark leadership metrics and a clustering of capabilities. Sentiment as to who is leading with the best model has shifted multiple times in just the last year, as Claude, Grok, ChatGPT, and Gemini have traded positions. Market lead times are compressing to several months, and no durable advantage has emerged.

Adding pressure: open source. Open-weight LLMs from China (DeepSeek V3, Qwen 3 from Alibaba), Mistral in France, and Meta’s Llama are pushing the frontier on multiple dimensions. They are free to adopt.

There are now an estimated 1,900 different language models globally, including small language models, domain-specific models, and variants of free open-source models. The LLM market could be viewed as a commoditizing utility layer: little durable differentiation, low switching costs, and pricing erosion from the sheer volume of alternatives.

To be fair, an estimated 88% of global LLM revenues still come from the top five vendors. An oligopoly is forming. But if the technology progression is constrained to incremental improvements and vendors can only progress linearly, no one gains a lasting market lead. Pricing and margins on LLMs alone must necessarily derive from other forms of differentiation.

The question becomes: differentiation how? The answer is not likely to simply build a fundamentally better model. This runs into a ceiling that even the best LLMs cannot break through.

We conclude that market differentiation will be derived from offering a comprehensive software stack or platform that supports a robust partner ecosystem, and to some degree by distribution scale in the consumer/mobile space and broader financial scale advantage to do acquisitions and broaden the product line.

The Limits of LLMs

LLMs are good at language: text summarization, translation, report generation, document creation, coding and debugging, image recognition, and image/audio/video generation from natural language inputs. In the enterprise context, they can connect to data, orchestrate workflows, and act as a natural language interface over enterprise systems. They can reimagine personal productivity software and automate many knowledge tasks that previously required human effort.

But LLMs are bad at true reasoning. They struggle with math and long chains of logic, where errors compound along intermediate steps. LLMs make probabilistic assessments, not deterministic, repeatable, and consistent deductions. They do high-level pattern matching, not structured symbolic inference, and break down when the patterns become less familiar. In large context windows, factors and constraints can simply be dropped from consideration. They are prone to hallucinations and, without guardrails for verification, cannot be trusted for critical decisions in fields such as medicine, law, finance, and safety-related operations. They infer patterns rather than “understanding” in a human sense, and they are weak at tasks requiring common sense grounded in physical reality and nuanced human intent.

Some industry visionaries such as Yann LeCun, “the grandfather of AI,” have very contrarian views on LLMs and are skeptical of the ability to linearly extrapolate from current LLM technology to get to the end goals of AGI or eventually super-intelligence.

On the path toward AGI, new technologies will need to complement LLMs to address these shortcomings: producing repeatable, deterministic output that can be relied upon in mission-critical applications and meet rigorous regulatory requirements. This is the key difference in passing the Turing test vs. achieving true intelligence: true intelligence is not probabilistic, though probabilistic systems can mimic intelligence. Enterprises will require true intelligence.

If the model itself has a ceiling, durable value must come from somewhere else. Historically, we’ve seen this path before.

The Stack Imperative

The LLM can be viewed as a foundational component of an AI system, similar in strategic significance to an operating system or database as building blocks for traditional enterprise systems. But as with Windows and Oracle’s database, the foundation is just the cornerstone of a more complete software stack.

Value derives up the stack, closer to the applications layer, leveraging domain expertise and managing business workflows around customer data. The strategic significance of building a comprehensive stack is to support a robust ecosystem of partners and customers, creating sticky relationships and a protective moat that supports premium pricing and higher margins.

A comprehensive software stack also insulates the LLM vendor from short-term changes in model benchmark leadership and commoditizing pricing pressures. If the value lives up the stack, the quarterly benchmark horse race matters less.

This means a more durable business can be built on top of the LLMs, providing the leading LLM vendors with market differentiation up the stack, scale advantage, and viable business models. We might also expect that rather than being disrupted by important new model innovations on the path to AGI, these scale vendors might have the opportunity to acquire, license, or build as an innovator or fast follower the new frontier and foundational technology that is likely to emerge through ongoing industry innovation.

We have seen this path before. Microsoft moved up the tech stack beyond Windows to database and infrastructure, security, monitoring and observability, DevOps, personal productivity apps, Dynamics small business applications, and business intelligence, ultimately becoming a hyperscaler with Azure. Oracle migrated beyond the database foundation layer to infrastructure software and into enterprise horizontal and vertical applications, and more recently into the OCI hyperscale business.

The key for LLM vendors who have the vision to build a broader business is to build a comprehensive software stack on top of their models: attracting an ecosystem of partners and customers, driving scale volume, and building richer business models as a full-stack provider with higher margins and sticky relationships.

What does the stack look like? The software stack extends upward from the foundation models through small language models with domain-specific knowledge, data management and data quality, RAG (retrieval-augmented generation to incorporate proprietary knowledge), middleware and tooling, security, coding assistants, multi-step reasoning, agentic workflow orchestration, knowledge graphs, governance, observability and monitoring, guardrails, vertical market AI agents with domain-specific data, and ultimately applications at the top of the stack in a highly fragmented space.

As we have suggested, there is business model risk of being a “one trick pony” provider of LLMs in a potentially less differentiated market. This could potentially influence the larger, leading, and better-financed LLM vendors to consider moving up the software stack and becoming full-stack providers. This will, however, also put them on a collision course with their hyperscaler partners, not in hosting but in the highly strategic full software stack.

Building out the layers of the stack will require significant technical expertise, scale advantage, financial resources, and likely partnerships and acquisitions. The competitive moat will likely be determined by the differentiation of the overall software stack, and much of this differentiation may take place higher up in the stack.

By moving up the tech stack, we would expect LLM vendors to develop robust agentic frameworks that compete with those now being introduced by leading enterprise software vendors. The goal is to capture enterprise design wins and become standards that businesses adopt as a core platform for the development and deployment of next-generation applications and infrastructure.

The ultimate realization of an LLM vendor’s full market potential is to build a platform that attracts a rich ecosystem of partners and entrenched customers building around the full stack for the next generation of AI-enabled enterprise systems. Our point is that the latest LLM benchmarks, which may shift public sentiment around apparent LLM market leadership, may prove to be a poor measure of longer-term business model success in AI. Rather, the ability to execute on a vision to become a full-stack vendor is likely more important, and scale to finance ongoing heavy investments is important to do both.1

The Inference Engine

With the enterprise adoption of AI, we are seeing a shift in data centers from model training to inference. This mix shift is critical to LLM vendors’ business models.

Compute capacity has been largely dedicated to training models episodically, with the balance going to inference. GPUs are good for high-speed, parallel processing of floating-point calculations and are flexible and optimal for training models. But a new class of Inference Processing Units (IPUs), LPUs, ASICs, and various accelerator chips are emerging that are more optimized for the specialized task of inference. Going forward, we would expect enterprise inference and reasoning to account for over 80% of future data center capacity.

Critically, enterprise inference is monetizable in a way consumer usage is not. We estimate 50% to 70% of enterprises pay for inference, given the high velocity and volume required for real-time business applications. Consumer usage, by contrast, is roughly 95% free, with under 5% paying roughly $20 per month. The shift from consumer to enterprise is as important as the shift from training to inference.2

User Monetization Rates
Enterprise50–70% pay
Consumer<5% pay

Roughly 50–70% of enterprises pay for inference given the velocity and volume of real-time business applications; consumer usage is roughly 95% free, with under 5% paying about $20 per month.

What drives enterprise inference volume? Agents leverage foundation models to automate business processes and orchestrate workflows. A simple chatbot question may involve several calls to the LLM. A complex multistep workflow may require an order of magnitude more.

Even more intensive: reasoning. LLMs are not good at reasoning, so they compensate with brute force computation. The approach has been to expand context windows and work through problems step by step in “deep thinking” or chain-of-thought mode, producing long sequences of intermediary deductions, checks, and competing conclusions before selecting an answer. This can involve hundreds or thousands of tokens per query. These drivers of inference are a shot in the arm for LLM vendor monetization and the ongoing expansion of demand for data center capacity.

Beyond agentics and reasoning, other inference drivers are emerging. Large quantitative models that use physics, biology, math, and science for modeling new protein molecules for new medicines and new material science will further drive inference demand. A broad spectrum of autonomous devices, from autos to humanoid robots, will leverage domain-specific and small language models (SLMs), distributing processing to the device level for low latency but placing further demands on components along the supply chain.

Inference will be the heartbeat of business automation, reasoning, agentics, and personal assistants. Inference costs are declining rapidly with improvements in model efficiency (cost per token) and faster, more optimized processors. How margins evolve for inference is unclear and underscores the importance of adding value up the stack.

It is easy to conclude that we are at a very early stage in data center demand. The amount of inference involved in reasoning is substantial, as is the volume of calls from agents orchestrating complex business workflows. And as agents become more autonomous and begin to be trusted to interoperate across business operations and interact with other enterprises, demand will only grow.

Data Center Realities

The inference opportunity requires massive infrastructure investment. If we look beyond the LLM vendors to the hyperscalers who run the models and are making many of the data center expansion commitments, we are looking at some of the largest and most profitable enterprise companies in the world, with robust cash flows and access to capital markets outside their own financial statements.

The top five hyperscalers spent close to $400 billion on capex in 2025, about 23% of their aggregate revenues. About 90% to 95% of cash flow from operations was spent on capex. These same hyperscalers have also increased corporate debt and construction loans by an additional $120 billion to $130 billion in 2025. The magnitude of the aggregate numbers for data center expansion is large ($600 billion in 20253), but so is the market opportunity if inference demand is just ramping up as we suggest.

A broader perspective: an estimated $2.9 trillion will be spent on data center expansion in the four years between 2025 and 2028, with half funded by the largest vendor balance sheets and the other half from debt capital markets and various partners. If we are correct that inference is likely to inflect significantly upward from enterprise demand, and there are ongoing capacity constraints, we would not expect to see an overall bubble of supply on the horizon.

Supply chain constraints. The supply chain has many constraints that imply we will not be in an oversupply position anytime soon.4 These include:

  • GPUs, and the emerging class of Inference Processing Units (IPUs), LPUs, and ASICs optimized for inference
  • High bandwidth memory (HBM) chips, increasingly in short supply and high priced
  • Power, power transformers, and delivery grids that in some cases date back over a hundred years
  • Large gas turbines, now backordered five to seven years; small gas turbines sold out through 2028
  • Natural gas pipeline distribution capacity
  • Fiber optics to reduce latency among data centers and their users
  • AI expertise and capital

To put the scale in perspective: an estimated 9 gigawatts of AI data center capacity was added in 2025 in the US to an estimated 50 GW base, with forecasts for over 100 GW to be added from 2024 to 2030. One gigawatt is the equivalent of a large nuclear reactor, adequate to service a city of about 800,000 to one million homes. A large data center now refers to a cluster of about 100,000 GPUs per building supported by 0.5 GW of power. Plans for hundreds of thousands of GPUs may have a combined cost of about $30 billion with 1 GW of power.

This leaves open the question of whether the business models for all vendors will support the cash flows to service debt and finance ongoing expansion. Some vendors have been more aggressive in making financial commitments than others. For leading LLM vendors, we would think it prudent to allocate adequate resources to build the software stack, not just data center capacity.

The Collision

The stack vision is compelling. But awkwardly, this vision is similar to that of the hyperscaler hosting partners who have continually built rich tech stacks to drive market leadership in their own businesses.

We see LLM providers and hyperscalers necessarily cooperating as partners yet also competing to deliver comprehensive and differentiated software stacks. Some of the leading LLM vendors will host their own models as hyperscalers but also have arrangements to be supported by other strategic hyperscalers and others who are making investments in data center capacity. As LLM vendors move up the tech stack to grow their TAM and derive strategic advantage, they will begin to collide with not only their hyperscaler partners but also some of their customers: the AI-native companies and traditional SaaS vendors who integrate via APIs to build on top of foundation models. An early example of this is in coding assistants, where Anthropic and OpenAI have enhanced their own coding assistants in competition with market leaders Cursor and Cognition who run on these models.

Enterprise software vendors are also building out their own agentic frameworks and related AI platforms, striving for customer adoption by promoting their own tech stacks. Only a very few will possess the resources to go beyond their installed base of users and compete broadly as full-stack vendors. They will likely want to be open and partner, integrating and leveraging component layers with various other stack component suppliers as much as possible.

We might see the hyperscalers and enterprise software vendors coexisting in a familiar coopetition market even as they overlap in agentic frameworks and other parts of the stack. The enterprise vendors are more likely to focus on enabling their installed base to build and deploy agents and interoperate with others rather than trying to be a full-stack provider, which would be a significant diversion of resources and a very different market than the applications business.

The hyperscalers already deliver robust full-stack platforms that support multiple LLM models, which their customers demand. The industry has learned not to lock in architecturally to one foundational dependency, so we observe vendors and users integrating lightly through open APIs. The efficiency benefits of an integrated full stack are seductive, however, which is the mission of full-stack vendors striving to attract an ecosystem and leverage network effects.

LLM vendors will need to open the stack with open interfaces (APIs) to support application development from third-party apps players and nurture and support an ecosystem of application partners who drive traffic through the stack to the underlying LLM and related products and services. These partners may well be reluctant to commit to proprietary components of the stack and risk getting locked in architecturally, an always tricky issue in the software market. Full software stack vendors will likely compete with some of these partners as well as they broaden their aperture in the apps space.

As the market evolves, it becomes increasingly important that agents from different systems can interoperate in complex workflows. There are agent protocols being adopted for this, including Model Context Protocol (MCP), Agent-to-Agent Protocol (A2A), and Agent Communications Protocol (ACP), to name a few. As reliability and trust in the models build, we would expect increasing levels of autonomous complex workflow orchestration and automation, within the enterprise but also across supply chains and customer relationships.

The strategic battle to control the stack and become a standard for the next generation of computing is critical. One might conclude that it is strategically imprudent for an LLM provider to pour all its cash flow into data center buildout when those resources might be better spent on M&A and partnering to add robustness to the broader tech stack.

There is historical precedent for this dynamic. Oracle dominated the database market, and as they expanded (by acquisition mostly) into the applications space, Oracle began to compete with the ecosystem of applications companies that ran on the Oracle database platform. The app vendors were compelled to run on top of Oracle given the popularity and market presence of this platform. Same with Microsoft Windows. These platforms became de facto market standards, and competitors up the tech stack needed to support the standard platform.

There is historical precedent for the risk of not controlling the platform. In an earlier generation, Netscape was the leading browser vendor with ambitions to disintermediate Windows by becoming the de facto platform for a new generation of web-based applications. Microsoft responded by bundling its own browser with Windows. While there were later antitrust issues, Netscape was frustrated in its ambitions and disintermediated in the market. LLM vendors need to make sure they are successful moving up the tech stack and not limited to the LLM market alone, which may have uncertain standalone value in a competitive market where scaling laws impose diminishing returns. This disintermediation risk underscores the imperative.

Differentiators Beyond the Model

If benchmark leadership is fleeting and the stack is the path to durable value, what else differentiates?

Distribution. Gemini is embedded into the Samsung Bixby virtual assistant and Apple’s anticipated new Siri intelligent assistant, a consumer market with the potential of over a billion users. Google can bundle Gemini into search, Gmail, Docs, and YouTube, a market of potentially 2 billion weekly users in 2026 versus ChatGPT’s current 900 million.

Enterprise focus. Anthropic has focused on enterprise scale versus the consumer market, which relates directly to the value of moving up the software stack where enterprises pay for inference and demand robust, integrated solutions.

Financial scale. Access to large sources of capital enables vendors to fund data center capacity, pursue acquisitions, and devote resources to expanding the stack. A broad software stack can be assembled through internal development, partnerships, and acquisitions, but all require financial scale and access to capital from private or public markets.

Acquisition currency. An IPO may be advantageous to access public capital markets and provide a currency for acquisitions. The full stack requires significant new skill sets, broader security and infrastructure capabilities, and domain expertise the further one moves up the stack into applications. This cost of acquisition is often now measured in billions of dollars.

The greater the stickiness in the market created by the software stack and other differentiators such as distribution and revenue scale, the greater the chance of success in adapting and evolving with ongoing innovations in model technologies. This has normally been the case in the software sector: incremental innovators want to disrupt but lack the resources to build a scale business or the time to build market presence and distribution, so they are consolidated by larger players with financial resources, scale advantage, and established distribution.

Implications

The market will likely support many LLM vendors and small language models, and it may fragment into hundreds of alternatives. But it is not clear that more than a few have the scale opportunity to be full-stack vendors and fight for a leading position as an emerging market standard. The ultimate realization of an LLM vendor’s full market potential will be to leverage network effects and attract a rich ecosystem of partners and entrenched customers that build around the full stack.

Platform shifts provide fertile ground for consolidation. M&A is used by larger scale vendors to acquire the technology needed, domain expertise, AI skills and IP, and shorten critical time-to-market. As LLM vendors migrate up into the applications layer, they will require domain expertise they do not currently possess. This is the path Oracle took as it acquired its way into the enterprise application space in the prior two generations (initially into the client/server market with PeopleSoft and Siebel, and then into the cloud/SaaS market with Taleo, RightNow, Eloqua, and NetSuite, among others). Initially, we would expect LLM vendors to focus on target markets such as financial analysis, planning, budgeting, personal productivity, and vertical applications such as legal, healthcare, and office of the CFO.

Success for incumbent vendors requires that they demonstrate both architectural agility and the mindset to impart organizational urgency to change. The latter is surprisingly difficult for companies who may resist the near-term impacts of reinventing themselves. They may lack technical skills, and both organizational inertia and resistance to change play primary roles. To adapt, most are compelled to acquire companies, IP, or “acquihire” the people who can bring in new skills and culture of innovation and disruption to enable the required time-to-market changes.

It would be overly ambitious to try and own more than a portion of the fragmented apps layer at the top of the stack. The applications market is just too broad and requires extensive domain and vertical market expertise, and in some markets extensive dedicated sales and support services.

Low-hanging fruit will be AI-native applications that play to the strengths of LLMs in text, language, images, video, coding, and domain knowledge, where agents can assist in automating new and previously manual human tasks. Opportunities likely include next-generation personal productivity software with agents to assist in automation, customer support and call center automation with virtual agents, legal document drafting and e-discovery, healthcare administration processes, and financial planning and modeling. Over time, we would expect this to evolve into a new generation of broader horizontal apps currently addressed by CRM and HR apps, starting at the low end of the market.

We would expect LLM vendors to acquire AI-native apps vendors who need leverage in go-to-market, and out-of-favor SaaS vendors who have domain expertise that can be repurposed in the AI-native stack. Acquihires may prove particularly attractive. They eliminate time-consuming product and organizational integrations that derail innovations and distract management. They provide structural flexibility to design non-exclusive licenses and sidestep regulatory objections to traditional M&A. Take the people and skills you need, leave the rest, and focus on the urgency to execute.

The regulatory environment appears favorable. The current administration is supportive of US efforts to stay ahead of China in AI and is open to industry participation. We have seen less resistance to consolidation broadly by the DOJ and FTC, and some may feel there is a window of opportunity to pursue larger deals.

The IPO window is likely to open as well, in an environment of high industry growth, growing scale among AI-native players, an accommodative Fed, and currently reduced market volatility. We would expect IPOs from emerging AI-native software vendors who are ramping to scale, as well as select traditional SaaS companies who missed the IPO windows of 2021 and have been able to layer on agentics but, more importantly, have begun to rearchitect with the LLM as the underlying inference engine to unlock insights about the business, enhance revenues, and improve productivity by orchestrating workflows without adding to headcount.

It may be instructive to take a quick look at the monolithic Mag 7 AI trade of earlier in 2025, which broadened over the year, cascading downstream to second, third, and fourth derivative beneficiaries among an enabling supply chain and across many different industry groups. Investors have expressed apprehensions that the AI trade has been overdone after the large run-up in 2025 centered around the Mag 7, which in aggregate represents about a third of the weighting of the S&P 500 and is responsible for nearly 45% of the performance of the index over the year. We would expect a further broadening of the AI trade as AI finds its way to enhancing productivity, in varying degrees, in every industry, globally.

We are optimistic about industry growth dynamics in 2026 as AI goes mainstream in the enterprise market.

We welcome your thoughts and a follow-up conversation.

Notes
  1. An important element will be the extension of vibe coding and coding assistants, and the ability to extend coding to enable Python customizations. Having a robust platform and open APIs makes it attractive for enterprises to support the platform, winning over more users, which drives more appeal of the platform in a network effect.
  2. Advertising could help monetize consumer usage and is likely upcoming. And as smartphones and a variety of wearable devices begin to scale and use personal agents, we will see further growth in inference demand on the consumer side as well.
  3. The $600B direct spend on data center expansion in 2025 helped fuel economic activity broadly for the US economy. It is estimated that investments in AI infrastructure could add 5% to GDP over the next 3 years.
  4. China has much greater energy expansion efforts underway and will be the chief rival to the US in power to fuel AI and is working to build its own chip fabrication capabilities. Taiwan’s TSMC currently produces over 90% of the world’s most advanced logic chips and provides about half of the world’s foundry capacity and 20% of total global chip production overall, an obvious geopolitical risk to AI and every industry that relies on advanced chips.

This material is for informational purposes only and does not constitute an offer, recommendation, or solicitation to buy or sell any security. Forward-looking statements are not guarantees of future performance and involve risks and uncertainties. Sherlund Partners may have existing positions in, or may seek to advise or invest in, companies mentioned herein.

Fundamental data and forecast supplied by management of the Company. Such projections may not be realized. Past performance may not recur, and there is no guarantee of future results.

This material has been prepared for information and educational purposes only, and it is not intended to provide, nor should it be relied on for tax, legal, or investment advice. You should consult with your own tax, legal, and financial professionals for your specific situation. The views and opinions expressed in this article are those of the author and do not necessarily reflect the views or opinions of Finalis Securities, LLC. Securities offered through Finalis Securities LLC Member FINRA/SIPC. Sherlund Partners LLC and Finalis Securities LLC are separate, unaffiliated entities.

This content is for general educational purposes only and not intended as legal, tax, accounting, securities or investment advice nor an opinion regarding the appropriateness of any investment, nor a solicitation of any type.

← Back to all insights