Startup profile for DeepInfra: latest funding, latest known valuation, total disclosed funding, investors, status, sources, and why the company is interesting. Part of the Venture Capital Tracker startup directory.

Startup profile · funding coverage

DeepInfra funding, valuation and investors

Managed AI inference cloud — OpenAI-compatible API for 190+ open models at production scale.

Keep track of DeepInfra

Save this profile to your VCT watchlist for a quick return.

View watchlist

Funding, valuation & investors

Answer-first snapshot

Latest funding

Series B

$107M · May 2026

Latest known valuation

Not publicly disclosed

Total disclosed equity funding

$125M

Excludes debt, grants, acquisitions, secondaries, and IPO proceeds.

Current status

private

series b · Palo Alto, CA

Investors in latest funding

Lead: 500 Global , Georges Harik

Other: Felicis , NVIDIA , Samsung Next , Supermicro

Sources: latest funding Last verified: 2026-07-25

Overview

DeepInfra is a Palo Alto-based cloud inference platform that lets developers and enterprises run open-source and proprietary LLMs via OpenAI-compatible APIs without owning GPU fleets. Founded in 2022 by Nikola Borisov and team (ex-imo messenger, 200M+ users), the company owns and operates its GPU infrastructure, supports 190+ models, and holds SOC 2 and ISO 27001 certifications with zero-data-retention options. In May 2026 it raised $107M in Series B co-led by 500 Global and Georges Harik, with NVIDIA, Felicis, Samsung Next, and Supermicro participating. The company also partners with Nvidia on Nemotron model serving.

Why DeepInfra is interesting

DeepInfra's $107M Series B in May 2026 came as the platform processed nearly five trillion tokens per week — proof that inference economics, not training clusters, is where recurring AI infrastructure revenue concentrates.

Product & use cases

DeepInfra provides a fully managed inference cloud with pay-per-token pricing, OpenAI-compatible endpoints, and enterprise security (zero data retention, compliance certifications) for running open and proprietary models in production.

  • Startups and scaleups deploying open-weight LLMs without capital expense on GPU clusters
  • Enterprises routing inference workloads to cost-optimized providers with compliance requirements
  • Agentic AI applications needing high-throughput, low-latency model serving at scale

Key facts

  • Series B (May 2026): $107M co-led by 500 Global and Georges Harik (GlobeNewswire)
  • Processes nearly 5 trillion tokens per week on owned GPU infrastructure
  • 190+ open-source models; OpenAI-compatible API; SOC 2 and ISO 27001 certified
  • Series A (Apr 2025): $18M; team built imo messenger scaling to 200M+ users
  • Nvidia partnership for Nemotron inference; Supermicro hardware investor

Funding history (newest first)

Series B

2026-05 $107M
  • 500 Global (lead)
  • Georges Harik (lead)
  • Felicis (participant)
  • NVIDIA (participant)
  • Samsung Next (participant)
  • Supermicro (participant)

Source: https://deepinfra.com/series-b

Competitive landscape

Edge: DeepInfra combines owned infrastructure with imo-era scaling expertise — processing trillions of tokens weekly while undercutting some GPU-cloud incumbents on per-token economics for supported open models.

Inference is the recurring cost center for most AI products as open models proliferate. DeepInfra competes in a crowded layer where margins depend on GPU supply, utilization, and routing — differentiation comes from model catalog breadth, compliance, and cents-per-token pricing rather than proprietary silicon.

  • Together AI direct

    Inference and fine-tuning platform for open models; similar developer audience.

  • Fireworks AI direct

    Fast inference for open and custom models; competes on latency and price.

  • Baseten direct

    Model serving and deployment platform; stronger on custom model ops.

  • AWS Bedrock / Azure OpenAI incumbent

    Hyperscaler bundled model APIs; DeepInfra wins on open-model breadth and cost for some workloads.

Notable stories

  • DeepInfra's founding team built imo, a messenger app with 200M+ users — bringing hyperscale consumer infra experience to inference economics (GlobeNewswire, May 2026).
  • Georges Harik, one of Google's earliest engineers, co-led the Series B — a strategic signal for production-grade inference at scale.

Industries

AI & Machine Learning Infrastructure & Cloud Developer Tools

Related funding articles

Venture Capital Tracker pieces that cover DeepInfra's financing or category context.

FAQs about DeepInfra

Practical answers founders, operators, and investors typically search for.

DeepInfra runs a managed inference cloud with OpenAI-compatible APIs for 190+ open and proprietary models, on infrastructure the company owns and operates.
$107 million in Series B funding co-led by 500 Global and Georges Harik.
Both serve open-model inference; DeepInfra emphasizes owned infra, token volume (5T/week), and compliance certs; Together also offers fine-tuning and training.
Nikola Borisov (CEO) and team from imo messenger; founded 2022 in Palo Alto.
The platform offers zero-data-retention options and holds SOC 2 and ISO 27001 certifications, per company disclosures.
190+ open-source models including Llama, Mistral, Qwen, and Nemotron family, via OpenAI-compatible endpoints.
500 Global and Georges Harik co-led Series B; NVIDIA, Felicis, Samsung Next, and Supermicro participated. None are in the VCT fund directory.
April 2025 — $18 million, per Data Center Dynamics reporting.

By Venture Capital Tracker

Last updated:

Editorial note: AI tools assisted with research, structure, or drafting. Venture Capital Tracker retains human editorial responsibility for factual accuracy, relevance, and source quality before publication.