On August 11, 2026, IBM and Together AI signed a multi-year $240 million agreement to build a large-scale open-source AI inference cluster on IBM Cloud using NVIDIA HGX B300 systems. The deal is the clearest signal yet that enterprise AI’s next competitive front is not which company trains the best model. It is who can run the best open-source models the cheapest, at regulated-enterprise scale.
What IBM and Together AI Actually Built
The announcement, published on IBM’s newsroom, describes the first dedicated large-scale inference cluster on IBM Cloud built specifically around NVIDIA HGX B300 systems, paired with NVIDIA Spectrum-X Ethernet networking. IBM says NVIDIA’s own assessment is that the configuration delivers 30 times more AI factory output compared to prior-generation hardware.
Together AI will use that cluster to run inference on open-source models: DeepSeek, Kimi, MiniMax, Nemotron, and GLM are among the models the company currently supports. Enterprises building on Together AI’s platform get access to this capacity without provisioning their own infrastructure. The cluster is expected to be available in Q1 2027.
For context on scale: Together AI currently reports serving around 400 trillion tokens per month across its existing infrastructure. The IBM Cloud cluster adds dedicated, enterprise-grade capacity on top of that baseline.
Why IBM Chose This Strategy
IBM’s calculation is straightforward, even if the execution is complex. The company was never positioned to out-train Anthropic, OpenAI, or Google on frontier models. Competing head-to-head with Amazon, Microsoft, or Google on hyperscaler cloud scale is similarly difficult. What IBM does have is deep regulatory trust with banks, hospitals, government agencies, and industrial enterprises.
Vipul Ved Prakash, CEO at Together AI, put the enterprise pitch plainly in the press release: “Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale.”
That framing points to something the last twelve months of enterprise AI deployment have confirmed: cost per token is now a board-level concern. The enterprise AI pricing revolt documented in Q2 2026 showed that organizations running agentic workflows at scale were hitting monthly AI bills that dwarfed their earlier estimates. The answer for many has been to shift workloads to open-source models, where pricing is set by infrastructure economics rather than frontier-model subscription tiers.
IBM, by providing that infrastructure with enterprise-grade SLAs and regulatory pedigree, gets a seat at a table it was otherwise excluded from.
The Inference Economy: What the Numbers Show
The timing of this deal is not coincidental. Gartner, cited by Network World, published a forecast this week that makes the infrastructure economics explicit:
| Metric | 2026 | 2027 |
|---|---|---|
| AI-optimized IaaS market size | $42 billion | $66 billion |
| YoY growth through 2026 | 96% | |
| Inference share of AI IaaS | 55% | 59% |
| Global inference spend | $23.3 billion | Rising |
| Global training spend | $19.0 billion | Rising but slower |
Source: Gartner forecast, cited in Network World, August 11, 2026
The shift is structural. Inference spending surpassing training in 2026 reflects a market that has moved from building models to deploying them at scale. Gartner’s analysts note that agentic AI amplifies this pattern: multi-step autonomous execution means each user query drives multiple inference calls, not one. An enterprise running AI agents across sales, operations, and customer support generates orders of magnitude more inference load than the same organization running a single chatbot.
Together AI’s platform already reflects this reality. The company’s August 2026 Series C at an $8.3 billion valuation was explicitly tied to scaling inference capacity. IBM’s $240 million cluster commitment accelerates that capacity build substantially.
The Open-Source Momentum in Enterprise AI
The IBM partnership also reflects a quieter shift in enterprise AI procurement. Security concerns and cost pressure are combining to push regulated-industry buyers toward models they can inspect, run on their own infrastructure, or access through a provider with clear data handling guarantees.
The Next Web’s analysis noted that “cybersecurity worries about closed models from Anthropic, OpenAI and Meta are cited as one reason some firms would rather run something they can inspect and host themselves.” For a hospital or a bank, keeping sensitive data on infrastructure it controls is frequently a regulatory requirement, not just a preference.
IBM Cloud’s position in regulated industries (the company cites thousands of government and financial services deployments on its hybrid cloud platform) gives Together AI something it lacked: a route into enterprise accounts that procurement, legal, and compliance teams have already approved. Together AI provides the inference technology and open-source model catalog. IBM provides the contractual and regulatory wrapper.
This mirrors a pattern visible in NVIDIA’s Nemotron routing infrastructure and Fireworks AI’s enterprise inference platform: the enterprise AI stack is disaggregating, and the winners are those building the most defensible layer of the stack, whether that is model quality, inference speed, pricing, or regulatory compliance.
What Enterprise Leaders Should Take from This Deal
Token cost is now a procurement criterion. The IBM-Together AI deal frames the enterprise AI market as an economics competition, not just a capability competition. Organizations evaluating AI vendors in H2 2026 should be modeling token cost at production scale for their anticipated agentic workloads, not just capability benchmarks.
Open-source inference now has enterprise plumbing. Historically, the friction for regulated enterprises was not open-source model quality but rather the operational overhead of standing up compliant, reliable inference infrastructure. IBM Cloud’s entry into this space removes that friction for a large segment of the enterprise buyer market.
The inference supply build will continue. The Gartner data shows inference demand growing at 96% YoY through 2026, and the investments reflect that. Between IBM ($240M), Together AI’s own $800M Series C, NVIDIA’s Nemotron router, and a wave of specialist inference providers, the market is building toward a surplus of open-source inference capacity. For enterprise buyers, that surplus eventually means lower prices.
The open-vs-closed question has a third answer. Much enterprise discussion frames AI vendor choice as “frontier closed model” versus “self-hosted open-source.” IBM and Together AI are building a third option: enterprise-grade open-source inference-as-a-service, with the compliance and SLA posture of a regulated cloud vendor. That option will appeal to organizations that need the cost profile of open-source but cannot absorb the operational complexity of building their own serving infrastructure.
If you are structuring your organization’s AI infrastructure strategy for 2026 and beyond, the Enera team can help you model the right inference architecture for your workloads and risk profile.