Together AI's inference cloud platform: pricing, key features, funding history, and how it compares for AI developers in 2026.
Together AI is an AI-native cloud computing company founded in June 2022 by CEO Vipul Ved Prakash together with Ce Zhang, Chris Re, Percy Liang, and Tri Dao, and headquartered in San Francisco, California. Prakash brought a background as a repeat founder, having previously built the email security company Cloudmark and the social analytics company Topsy, which he sold to Apple in 2013, while his co-founders contributed deep academic AI research experience from institutions including Stanford, ETH Zurich, and the University of Chicago.
The company was built around the thesis that open-source AI models, run on efficient, purpose-built infrastructure, could match or beat closed proprietary models on cost and, in many cases, performance. Rather than building its own foundation models exclusively, Together AI operates a cloud platform that hosts and serves a wide catalog of open-source and partner models, alongside the GPU infrastructure needed to train and fine-tune custom models.
Together AI has raised substantial venture funding to build out its GPU infrastructure, including a 305 million dollar Series B round led by General Catalyst, and an 800 million dollar raise in July 2026 that valued the company at 8.3 billion dollars, with backing from major investors including Nvidia, Vista Equity Partners, and Aramco Ventures.
Together AI's core product is serverless inference, a pay-per-token API for running popular open-source language models such as Llama and DeepSeek variants, plus image generation, video generation, audio transcription, and embedding and reranking models, all accessible through a unified API without managing any GPU infrastructure directly.
For teams that need more control or predictable performance, Together AI offers dedicated inference on single-tenant GPU instances, provisioned throughput for reserved capacity at a fixed rate, and on-demand or reserved GPU clusters that can be rented by the hour for custom training or high-throughput workloads, with discounts of 10 to 30 percent for longer reservation terms.
The platform also includes fine-tuning infrastructure for training custom versions of open-source models on proprietary data, a code execution sandbox for running AI-generated code safely, and managed storage, positioning Together AI as a full-stack AI development cloud rather than just an inference API.
Together AI uses entirely usage-based pricing with no flat subscription fee. Serverless chat and language model inference is billed per million input and output tokens, with rates spanning from roughly 5 cents per million tokens for the smallest open models up to several dollars per million tokens for large flagship models like DeepSeek's reasoning variants; a batch processing API offers a 50 percent discount for workloads that don't need real-time responses.
GPU compute is billed by the hour, with on-demand clusters running roughly 4 to 8 dollars per GPU-hour depending on the chip type, and dedicated single-tenant inference instances starting around 5.49 dollars per hour for H100 GPUs, with reserved capacity for 7 to 180-plus days offering meaningful discounts. Fine-tuning is billed per token processed during training, starting around 48 cents per million tokens for standard supervised fine-tuning and rising for specialized model types, with a minimum charge per job.
Together AI provides an AI-native cloud platform offering pay-per-token inference APIs, GPU cluster rental, and fine-tuning infrastructure for open-source AI models.
Together AI was founded in June 2022 by CEO Vipul Ved Prakash along with Ce Zhang, Chris Re, Percy Liang, and Tri Dao.
Together AI has raised roughly 900 million dollars across multiple rounds, including an 800 million dollar raise in July 2026 at an 8.3 billion dollar valuation.
Together AI uses usage-based pricing: inference is billed per token, GPU clusters are billed per GPU-hour, and fine-tuning is billed per token processed during training.
Together AI does not have a persistent free plan for production use, but it offers trial credits for new accounts to test the platform.
Together AI hosts a wide catalog of open-source models including Llama and DeepSeek variants, along with image, video, audio, and embedding models.
Together AI focuses on hosting and serving open-source models at competitive per-token prices, rather than developing and selling access to a single proprietary closed model family.
Yes, Together AI offers fine-tuning infrastructure billed per token processed, letting teams train custom versions of open-source models on their own data.