Short Summary
IBM and Together AI have announced a multi-year, $240 million partnership aimed at expanding access to enterprise-grade open-source artificial intelligence. Under the agreement, IBM Cloud will host a massive computing cluster equipped with NVIDIA’s advanced HGX B300 hardware and high-speed Spectrum-X Ethernet networking. Scheduled to go live in early 2027, the infrastructure will allow Together AI to serve high-volume open-source model workloads at significantly lower costs, offering businesses a high-performance alternative to expensive proprietary AI ecosystem lock-ins.

Introduction
As artificial intelligence transitions from experimental pilots to core business operations, enterprises are hitting a financial and technical wall. Running frontier AI models in production requires massive computing power, and reliance on closed, proprietary models often brings unpredictable costs and restricted customization.
To address this challenge, IBM and AI startup Together AI signed a multi-year agreement valued at $240 million. The initiative brings together IBM’s cloud architecture, Together AI’s open-source inference engine, and NVIDIA’s next-generation hardware. By building a dedicated high-performance computing environment on IBM Cloud, the companies aim to redefine the economics of open-source artificial intelligence, proving that open models can match closed alternatives in both speed and reliability at a fraction of the price.
What Happened?
IBM will build a dedicated, large-scale inference cluster within IBM Cloud specifically designed for Together AI. Powered by NVIDIA HGX B300 systems and connected via NVIDIA Spectrum-X Ethernet networking, the deployment marks the first dedicated cluster of its kind built on IBM’s cloud platform.
Expected to become operational in the first quarter of 2027, the hardware will serve as an engine for Together AI, a company focused on hosting, fine-tuning, and deploying open-source AI models. Together AI—which recently raised $800 million in a Series C funding round at an $8.3 billion valuation—already processes 400 trillion AI tokens per month for over a million developers. The new deal gives Together AI guaranteed GPU capacity and raw computing power to scale its platform as enterprise demand explodes.
Why It Matters
The AI landscape is witnessing a structural shift. Historically, tech companies relied heavily on hyper-scalers like AWS, Microsoft Azure, and Google Cloud for raw GPU capacity. However, the rise of specialized “neoclouds”—infrastructure providers built exclusively for AI workloads—has disrupted traditional cloud economics.
For IBM, this deal anchors its position as a primary backend provider for top-tier AI platforms. Rather than competing head-to-head for developer mindshare against specialized AI platforms, IBM secures a reliable, multi-million-dollar tenant while showcasing its cloud security and reliability.
For enterprises, the partnership signals an opportunity to lower operational expenditures. Open-source models (such as Meta’s Llama series, DeepSeek, and MiniMax) have reached performance parity with many proprietary systems. However, serving these models to millions of end-users in real time requires specialized hardware optimized for rapid token output. This cluster provides the necessary throughput to make open-source deployments economically viable for global organizations.
Technical Explanation
To understand why this architecture represents a step forward, it helps to distinguish between two main phases of machine learning: training and inference.
- Training is the initial process where an AI model learns patterns from massive datasets. It requires raw compute power over weeks or months.
- Inference is the live execution phase where a trained model processes user prompts and generates responses in real time.
Inference is an always-on, high-volume process where latency (speed) and token economics (cost per word/character generated) directly dictate profitability.
+-------------------------------------------------------------------+
| TOGETHER AI PLATFORM |
| (Open-Source Models, Agentic Workflows, API Keys) |
+-------------------------------------------------------------------+
|
v
+-------------------------------------------------------------------+
| IBM CLOUD ARCHITECTURE |
| (Enterprise Governance, Hybrid Cloud, Hybrid Security) |
+-------------------------------------------------------------------+
|
v
+-------------------------------------------------------------------+
| NVIDIA HGX B300 & SPECTRUM-X |
| (30x Output Boost, High-Bandwidth Interconnect, FP4/FP8 Precision) |
+-------------------------------------------------------------------+
The underlying hardware driving this cluster consists of NVIDIA HGX B300 systems, built on NVIDIA’s Blackwell GPU architecture. According to NVIDIA, these systems deliver up to 30 times more output performance compared to previous-generation hardware.
Connecting thousands of these GPUs requires specialized network design. The cluster uses NVIDIA Spectrum-X Ethernet networking, a network technology engineered specifically for multi-tenant AI clouds. Standard Ethernet often suffers from packet collisions and latency spikes when handling distributed AI tasks; Spectrum-X prevents network congestion and maintains maximum data flow between computing nodes, ensuring real-time response speeds even under heavy traffic load.
Key Highlights
- $240 Million Investment: A multi-year infrastructure agreement binding IBM, Together AI, and NVIDIA tech stacks.
- Next-Gen Hardware: Deployment of NVIDIA HGX B300 accelerators paired with Spectrum-X Ethernet.
- Launch Timeline: Target availability set for Q1 2027 on IBM Cloud.
- High-Volume Scale: Together AI currently handles 400 trillion tokens monthly and plans to expand enterprise agentic workflows.
- Significant Funding: Supported by Together AI’s recent $800M Series C round at an $8.3B valuation.

Benefits
For Enterprise Customers
- Reduced Token Costs: Running open-source models on optimized inference hardware lowers the cost-per-response compared to closed proprietary APIs.
- Enterprise Security and Compliance: Leveraging IBM Cloud’s underlying infrastructure ensures strict data governance, crucial for heavily regulated sectors like finance, healthcare, and government.
- Elimination of Vendor Lock-In: Open stacks allow businesses to migrate models, tweak parameters, and host weights wherever they choose without being tied to a single vendor’s API.
For the Open-Source Ecosystem
- Production-Grade Reliability: Historically, open-source models suffered from fragmented hosting options. Dedicated enterprise clusters bridge the gap between open code and high availability.
- Faster Agentic AI Execution: Complex AI agents that call multiple tool APIs require low latency. Improved hardware reduces output lag, making multi-step reasoning systems feel instantaneous.
Challenges
While the infrastructure promises high performance, several technical and market hurdles remain:
- Hardware Deployment Timelines: With availability targeted for early 2027, maintaining a competitive edge requires seamless hardware delivery and integration over the next 18 months.
- Thermal and Energy Demand: Next-generation GPU clusters require immense power and liquid cooling infrastructure. IBM Cloud must ensure sustainable energy provisioning at scale.
- Evolving Model Architectures: AI model design changes rapidly. The underlying hardware and software stack must adapt if new architecture paradigms bypass standard transformer-based inference methods.
Future Outlook
The partnership highlights where enterprise software is heading. As businesses integrate autonomous agents to handle workflows, daily query volume will shift from millions to billions of transactions. At that scale, paying premium prices for closed APIs becomes unsustainable.
By mid-2027, dedicated open-source inference clusters will likely become standard across major cloud ecosystems. This initiative positions IBM and Together AI at the center of that transition, setting a template for how traditional cloud providers and specialized AI platforms can co-exist.
Our Analysis
This agreement represents a strategic win for all three players involved:
- IBM proves that its cloud platform can support bleeding-edge AI workloads, shifting market perception beyond legacy hybrid cloud services.
- Together AI secures guaranteed enterprise-grade computing power without having to expend capital constructing its own physical data centers.
- NVIDIA reinforces its hardware dominance across both legacy hyper-scalers and specialized AI providers.
More importantly, it serves as a reality check for the broader market: the debate between open-source and proprietary AI will not be decided solely by model performance, but by the efficiency and cost-effectiveness of the infrastructure hosting them.
FAQ
1. What is AI inference, and how does it differ from training?
Training is the initial stage where an AI model learns from large amounts of data. Inference is the operational stage where the trained model processes user queries and generates outputs in real time.
2. Why are companies shifting toward open-source AI models?
Open-source models give organizations complete control over their data, eliminate vendor lock-in, allow custom fine-tuning, and significantly lower token costs compared to closed proprietary APIs.
3. What makes the NVIDIA HGX B300 hardware unique?
Built on NVIDIA’s Blackwell architecture, the HGX B300 is optimized specifically for heavy AI factory throughput and fast inference, offering up to 30 times the output performance of prior generations.
4. When will this new cluster be operational?
The cluster is expected to be fully deployed and available on IBM Cloud in the first quarter of 2027.
5. How much data is Together AI currently processing?
Together AI currently serves approximately 400 trillion AI tokens per month across its developer and enterprise user base.
Conclusion
The $240 million partnership between IBM and Together AI signals a turning point in enterprise AI infrastructure. By uniting NVIDIA’s HGX B300 processors with IBM’s cloud foundation and Together AI’s open-source orchestration, the initiative addresses the growing enterprise demand for fast, cost-effective, and transparent artificial intelligence. As the cluster comes online in early 2027, it stands to accelerate the mainstream adoption of open-source models across global industries.
