Usage-based (pay-per-token) with on-demand GPU and custom enterprise pricing, from Pay-as-you-go from a fraction of a cent per 1M tokens, with $1 in free credits for new accounts
Verified
Not yet
Last updated
July 18, 2026
Founded
2022
Headquarters
Redwood City, California, United States
Web AppFree TrialAPIAIFreemium
Overview
Fireworks AI is a generative AI inference platform that gives developers fast, production-grade access to open-weight and custom large language models through a single OpenAI-compatible API. Instead of building its own foundation models, Fireworks focuses on serving models efficiently, hosting more than 100 open-source and proprietary models including Llama, DeepSeek, Qwen, Mixtral, and Mistral.
Founded in late 2022 by Lin Qiao, former head of PyTorch at Meta, and six co-founders also from Meta's AI infrastructure teams, Fireworks has grown from a niche inference API into a broad platform covering serverless inference, dedicated on-demand GPU deployments, batch inference, and managed fine-tuning. The company is headquartered in Redwood City, California, and reported an annualized revenue run rate exceeding $1 billion as of its July 2026 funding round.
Key Features
The platform's core technology is FireAttention, a custom CUDA-kernel inference engine that the company says serves models several times faster than standard open-source serving frameworks, while preserving output quality through targeted quantization.
Beyond raw speed, Fireworks offers managed fine-tuning (supervised fine-tuning, direct preference optimization, and reinforcement fine-tuning), Multi-LoRA hosting for running up to 100 fine-tuned model adapters at base-model pricing, function calling and structured outputs for agentic workflows, and dedicated on-demand GPU deployments for predictable, high-volume production traffic.
Pricing
Fireworks uses metered, pay-per-token billing for serverless inference, with new accounts receiving a small amount of free credit to start. Cached input tokens and batch inference are discounted fifty percent relative to standard rates, and fine-tuned models are served at the same price as their base models.
For predictable, high-volume workloads, Fireworks offers dedicated on-demand GPU deployments billed per GPU-second, at roughly $7 per GPU-hour for H100 and H200 chips, $10 per GPU-hour for B200 chips, and $12 per GPU-hour for B300 chips, with no separate charge for start-up time. Fine-tuning is billed per million training tokens, ranging from about $0.50 for LoRA supervised fine-tuning on smaller models up to $40 for full-parameter DPO on models above 300 billion parameters. Enterprise customers can negotiate custom contracts with volume discounts and private deployment options.
Key Features
FireAttention inference engine — A custom CUDA-kernel-based inference engine built in-house that Fireworks says delivers substantially faster token generation than standard open-source serving frameworks, including on large mixture-of-experts models like DeepSeek.
Broad open-model catalog — Hosts more than 100 open-source and proprietary models, including Llama, DeepSeek, Qwen, Mixtral, Kimi, GLM, and Mistral, accessible through a single OpenAI-compatible API.
Serverless pay-per-token inference — Metered, postpaid API access with no infrastructure to manage, including discounted pricing for cached input tokens and batch inference jobs.
On-demand dedicated GPU deployments — Reserved H100, H200, B200, and B300 GPU capacity billed per GPU-second for predictable, high-volume production workloads, with no extra charge for start-up time.
Managed fine-tuning — Fully managed supervised fine-tuning, direct preference optimization, and reinforcement fine-tuning pipelines priced per million training tokens by model size.
Multi-LoRA model hosting — Deploy and instantly switch between up to 100 fine-tuned model adapters on shared infrastructure at the same price as the base model, useful for per-customer or per-task personalization.
Function calling and structured outputs — Native support for function calling, JSON mode, and structured outputs designed for agentic and compound AI applications that chain multiple model and tool calls together.
Enterprise security and compliance — SOC 2 Type II and HIPAA compliance plus ISO 27001, ISO 27701, and ISO 42001 certifications, with private and virtual-private-cloud deployment options for regulated industries.
Very broad catalog of open-weight models plus support for custom and fine-tuned model uploads
Fine-tuned models and Multi-LoRA adapters are served at the same price as base models, lowering the cost of customization
Strong compliance posture (SOC 2 Type II, HIPAA, ISO 27001/27701/42001) suitable for regulated enterprise buyers
Cons
Usage-based, per-token and per-GPU-hour pricing across many model sizes and deployment modes can be harder to predict and budget than a flat subscription
No proprietary flagship foundation model of its own, so output quality ceiling depends on the open and third-party models it hosts
On-demand dedicated GPU deployments require infrastructure planning and capacity commitment that smaller teams may find unnecessary for low-volume use
Rapid pricing and product changes driven by fast company growth mean published rates for niche models can shift and should be reverified before large commitments
Pricing
Serverless Inference Pay-per-token, from a fraction of a cent per 1M tokens for smaller models Usage-based, postpaid
On-Demand Deployments $7 per GPU-hour on H100/H200, $10 per GPU-hour on B200, $12 per GPU-hour on B300 Billed per GPU-second
Fine-Tuning From $0.50 to $40 per 1M training tokens depending on model size and method Pay per training token
Enterprise Custom pricing Annual contract
Frequently Asked Questions
What is Fireworks AI?
Fireworks AI is a generative AI inference and fine-tuning platform that lets developers run open-weight and custom large language models through a fast, OpenAI-compatible API, without managing their own GPU infrastructure.
Who founded Fireworks AI and when?
Fireworks AI was founded in late 2022 by Lin Qiao, former head of PyTorch at Meta, together with six co-founders who also came from Meta's AI infrastructure and PyTorch teams.
How much has Fireworks AI raised, and what is it worth?
Fireworks has raised capital across a Series A, Series B, Series C, and Series D, most recently a $1.5 billion round in July 2026 at a $17.5 billion valuation led by Atreides Management, Index Ventures, and TCV, with NVIDIA and Lightspeed among the participating investors.
How is Fireworks AI priced?
Fireworks uses pay-per-token metered billing for serverless inference, per-GPU-hour billing for dedicated on-demand deployments, and per-million-training-token pricing for fine-tuning, with new accounts receiving a small amount of free credit to get started.
Which models can I run on Fireworks AI?
Fireworks hosts more than 100 open-source and proprietary models, including Llama, DeepSeek, Qwen, Mixtral, Kimi, GLM, and Mistral, and also supports deploying your own custom or fine-tuned models.
How does Fireworks AI compare to Together AI, Groq, and AWS Bedrock?
Fireworks is generally considered a price-and-performance all-rounder with a broad model catalog and strong fine-tuning support, similar in positioning to Together AI. Groq focuses on raw speed using custom chip hardware but offers fewer models and no fine-tuning, while AWS Bedrock is a hyperscaler-hosted alternative favored by teams already standardized on AWS infrastructure.