Map → L6 AI Runtime, Serving & Model Access → Fireworks AI
Fireworks AI
L6 · AI Runtime, Serving & Model Access
inference & servingstartupprivateseries C+
Fireworks AI operates an enterprise inference platform engineered to optimize, fine-tune, and serve generative AI models. Using custom compilation stacks and FireAttention kernels, it accelerates token generation speed while substantially reducing compute costs. Its Multi-LoRA multiplexing technology allows thousands of custom model adapters to execute efficiently on shared GPU clusters.
Ecosystem functionActs as a high-throughput token serving stack that leverages custom kernel optimizations and Multi-LoRA multiplexing to maximize serving efficiency.
Business modelUsage-based per-token API and dedicated GPU hosting
Revenue$1.0B+reported
Valuation$17.5Breported
Last round$1.505Bled by Atreides Management / Index Ventures / TCV · 2026
Sources · 3 ↓
Suggest an editLast verified 2026-07-26