# Nebius buys Inferize to tackle AI’s cold-start problem

> The cloud provider is folding a nine-month-old startup into its Token Factory platform to cut model load times and idle GPU capacity.

- Publisher: NextDiff (https://nextdiff.com/)
- Section: Companies
- Author: NextDiff Editorial
- Published: 2026-10-06T18:44:57.607Z
- Updated: 2026-10-07T04:51:55.968Z
- Canonical URL: https://nextdiff.com/posts/nebius-acquires-inferize-ai-cold-starts/
- Topics: nebius, acquisitions, inference, gpus

## In short

Nebius announced on October 1, 2026 that it has acquired Inferize, a startup founded in January 2026 whose technology shortens AI model cold starts and reduces idle GPU capacity. The team and technology join Nebius Token Factory, its managed inference platform. Financial terms were not disclosed.

## Key takeaways

- Nebius announced on October 1 that it has acquired Inferize, an inference-optimisation startup founded in January 2026.
- Inferize’s technology cuts “cold starts” — the time a model takes to load before it can serve requests — and reduces idle GPU capacity.
- Deal terms were not disclosed; the team joins Nebius Token Factory, its managed inference platform.

Nebius said on October 1 that it has acquired **Inferize**, a young startup working on one of the least glamorous but most expensive problems in production AI: the **cold start**.

## What does Inferize do?

When traffic to a model spikes, new GPU capacity has to load the model’s weights before it can answer a single request. That delay — the cold start — forces providers to choose between slow responses and keeping expensive GPUs warm and idle “just in case.” Inferize builds technology that shortens model load times and reduces idle capacity, improving what Nebius calls token economics: how much useful output each GPU-hour produces.

According to [Nebius’ announcement](https://nebius.com/newsroom/nebius-acquires-inferize-to-strengthen-nebius-token-factorys-production-inference-stack), Inferize was founded in January 2026 and had a working prototype within three months. Its technology and team will be integrated into **Nebius Token Factory**, the company’s managed platform for running AI models in production. Financial terms were not disclosed.

Nebius CTO Danila Shtan framed it as a whole-system problem: the platform has to react when demand changes, including how quickly extra capacity is ready.

## Why it matters

Training gets the headlines, but inference is where AI spends money every day. As usage becomes spikier — agents that suddenly fan out into dozens of calls, launches that drive traffic overnight — the cost of idle GPUs and slow scale-up grows. Buying a team that attacks that cost directly is a sign of where cloud providers now compete: not only on how many GPUs they have, but on how efficiently they keep them busy.

## What builders should take from it

If you run your own inference, measure your cold-start time and idle GPU percentage — they are often bigger costs than the per-token price on your invoice. If you buy inference, ask providers how quickly they scale under bursty load, not just what they charge per million tokens.

Related: [Amazon wants investors to own  billion of its Nvidia chips — and rent them back](/posts/amazon-8-billion-nvidia-chips-sale-leaseback/).

## Questions readers ask

### What is an AI cold start?

A cold start is the delay while a model's weights load onto new GPU capacity before it can serve requests; it forces providers to choose between slow responses and keeping expensive GPUs idle.

### What does Inferize do?

Inferize builds inference-optimisation technology that shortens model load times and reduces idle GPU capacity, improving how much useful output each GPU-hour produces.

### How much did Nebius pay for Inferize?

Nebius did not disclose the financial terms of the acquisition.

## Sources

- [Nebius newsroom — Nebius acquires Inferize to strengthen Nebius Token Factory's production inference stack](https://nebius.com/newsroom/nebius-acquires-inferize-to-strengthen-nebius-token-factorys-production-inference-stack)
