Introduction
IBM has stopped trying to outgun AWS, Azure, or Google Cloud. By signing a multi-year, $240M agreement with Together AI on August 11, 2026, the company is openly claiming a different identity: that of a specialized AI inference "neocloud." This deliberate repositioning says a great deal about the structural shift underway in the cloud market — and about the opportunities emerging for players who choose not to fight a war of attrition against the big three.
An Inference Cluster Built for Industrial-Scale Production
The deal calls for the deployment, in Q1 2027, of the largest dedicated inference cluster ever hosted on IBM Cloud. It will be built on approximately 2,000 NVIDIA Blackwell HGX B300 chips, interconnected via NVIDIA's Spectrum-X networking fabric. These systems are reported to deliver thirty times the processing capacity of previous generations. This is not general-purpose infrastructure: it is engineered end-to-end for a single type of workload — industrial-scale AI inference.
Together AI: A Case Study in Open Source Going Enterprise
Founded in 2022, Together AI embodies the broader market transformation. In July 2026, the company closed an $800M funding round at a valuation of $8.3 billion — 2.5 times the figure it commanded at its Series B just eighteen months earlier. It now processes 400,000 trillion tokens per month for more than one million developers, running exclusively on open-source models — DeepSeek, MiniMax, Kimi. Its customers reportedly cut their inference costs by up to sixty times compared to equivalent proprietary alternatives. Usage of open-source models in production has reportedly tripled within a year, according to the company's own data.
The Quiet Restructuring of the Cloud Market
This is where the signal turns strategic. IBM is not positioning itself as the fourth hyperscaler. It is targeting a distinct segment: providing raw, reliable, economically predictable compute to operators whose volumes outgrow what general-purpose cloud can competitively handle. The specialized neocloud model is not a new concept — CoreWeave and Lambda Labs blazed this trail — but the entry of an established enterprise operator of IBM's scale gives the phenomenon a different dimension entirely.
The move also reflects a fundamental shift in procurement criteria. Large-scale inference has different requirements from a standard cloud application project: what matters is cost per token, latency, compute density, and guaranteed availability — not the breadth of a managed-services catalog.
What This Means for IT and Infrastructure Leaders
For CIOs and infrastructure leaders at large organizations, this deal is an invitation to revisit the standard cloud vendor selection playbook. Multicloud no longer simply means splitting workloads across AWS, Azure, and GCP. It can now encompass specialized operators, purpose-built for specific AI workloads, operating under structurally different commercial models.
The rise of open-source models in production environments is accelerating this shift. As AI inference becomes a significant budget line item in its own right, the choice of underlying infrastructure — and of the provider behind it — will carry increasing weight in architectural decisions. IBM is betting that this window stays open for years to come, and that specialization beats generalization in a market where the three leading hyperscalers have already built a lead that is very hard to close.

