Bitget App
Trade smarter
Buy cryptoMarketsTradeFuturesStocksEarnInstitutionAI & More
Nebius (NBIS.US) acquires Israeli startup Inferize to strengthen AI inference business, transaction amount could reach up to $150 million

Nebius (NBIS.US) acquires Israeli startup Inferize to strengthen AI inference business, transaction amount could reach up to $150 million

智通财经智通财经2026/10/01 15:11
Show original
By:智通财经

AI cloud infrastructure service provider Nebius announced the acquisition of Israeli AI startup Inferize, further enhancing its AI inference infrastructure capabilities.

According to Zhitong Finance APP, AI cloud infrastructure service provider Nebius (NBIS.US) has announced the acquisition of Israeli AI startup Inferize to further enhance its AI inference infrastructure capabilities. Inferize specializes in reducing idle time of Graphics Processing Units (GPUs) and accelerating the deployment speed of large AI models. According to Israeli tech media Calcalist, the transaction is valued between $100 million and $150 million, though Nebius did not disclose the specific financial terms.

The core of this acquisition lies in improving the efficiency of GPU resource utilization, particularly in the context of rapidly shifting AI inference demands, enabling Nebius to allocate computing resources more quickly and reduce idle time of costly computing equipment.

Aiming at AI Inference Efficiency Bottlenecks, Reducing GPU Idle Time

Inferize’s core technology mainly addresses the “cold start” issue in AI model deployment. Cold start refers to the process where an AI model must load and initialize related resources before it can begin processing user requests. When a new inference instance is launched, the GPU may have to wait for the model to finish loading before it can start computations, resulting in temporary resource idleness.

This issue becomes particularly prominent when AI inference demands surge suddenly. As user requests soar, cloud service providers need to quickly launch more inference instances. However, if model loading takes too long, newly added GPU resources cannot be utilized immediately, which affects service response speed and overall computational efficiency.

Inferize’s technology is designed to shorten this waiting period, so that new inference resources can be put into operation faster, allowing capacity expansion to better align with real demand fluctuations.

For Nebius, this means the company can more flexibly meet customer demand and improve the utilization of its existing infrastructure without having to maintain a large number of idle GPUs as standby resources over the long term.

Nebius CTO Danila Shtan stated that efficiently running AI inference services requires not only faster GPUs and optimized models, but also a system that can promptly respond to shifting demands, including rapidly providing additional computing power when customer needs increase. He noted that Inferize brings not only technology for accelerating this process but also an engineering team with rich experience in GPU systems.

Nebius plans to incorporate Inferize’s engineers into its inference business team and to integrate the relevant technology into its AI inference platform Token Factory, enhancing the platform’s ability to respond to customer demands and enabling the existing infrastructure to handle more practical computing tasks.

Shtan also emphasized that the Inferize team’s future contributions will not be limited to this initial technology integration.

Reducing Standby Computing Costs, Enhancing AI Infrastructure Operational Efficiency

Inferize co-founder and CEO Guy Bortnikov stated that, to respond promptly to customer requirements, cloud service providers usually must maintain a certain scale of standby GPU resources—these idle compute devices themselves represent additional costs. He explained that Inferize was established precisely to reduce these costs, and by joining Nebius, the related technology can be directly applied to operational AI cloud platforms.

In the AI inference business, the efficiency of compute resource scheduling is directly tied to the operational costs of the infrastructure. Since GPU equipment is expensive, enterprises that need to keep a large amount of idle computing power on hand to meet potential demand peaks could see their overall return on investment affected.

By shortening model startup times and accelerating resource deployment, Nebius is expected to improve service flexibility while reducing its reliance on standby GPU capacity.

Notably, Bortnikov had previously co-founded Granulate, a company focused on optimizing compute infrastructure software. Granulate mainly developed technology to improve the efficiency of computational resources, and was acquired by Intel (INTC.US) in 2022. This background also illustrates that the Inferize team has relevant experience in the field of computational infrastructure optimization.

Continuous Acquisitions of AI Tech Firms—Nebius Further Strengthens Inference Platform

The acquisition of Inferize is one of a recent series of M&A moves by Nebius in the field of AI inference and model optimization technology. Previously, Nebius acquired Eigen AI in a deal valued at $643 million. The company also recently acquired AI tech company Clarifai, although the transaction amount was not disclosed.

All of these acquisitions are aimed at enhancing Nebius’s AI inference and model optimization capabilities, further improving its AI infrastructure services.

Unlike infrastructure investments mainly centered around AI model training compute power, AI inference focuses more on the operational efficiency after models are deployed, including request processing speed, computing resource scheduling, model deployment, and cost per unit of computation.

As Nebius continues to integrate related technologies, its business is evolving from providing GPU computing resources toward further encompassing technologies and services that improve the operational efficiency of AI models.

This acquisition of Inferize will further supplement Nebius’s technology in dynamic compute scheduling and rapid model deployment. However, to what extent this integration can deliver cost savings and efficiency gains remains to be validated by future operational performance.

0
0

Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.

Understand the market, then trade.
Bitget offers one-stop trading for cryptocurrencies, stocks, and gold.
Trade now!

You may also like

Anthropic moves closer to IPO: Reportedly aiming for a $2 trillion valuation before Thanksgiving, Investor Day scheduled for October 14th

According to reports, Anthropic may kick off its roadshow as early as November, aiming to complete its public listing before the U.S. Thanksgiving holiday on November 26. The company will hold a Pre-IPO Investor Day with potential investors on October 14 and has already sent invitations to several institutional investors. Some potential investors believe the company's reasonable valuation may reach $1.8 trillion to $2 trillion.

华尔街见闻•2026/10/01 20:36

High U.S. Treasury yields exert pressure, bank stock index deeply in correction zone, Citigroup drops nearly 5% intraday

During Thursday's trading, the KBW Bank Index hit a four-month low, falling 14% from its mid-August peak. Analysts noted that in September, the financial sector's performance relative to other industries was the worst for the same period since 1990. Some analysts also pointed out that concerns over the threat of AI agents, potential impacts from the US midterm elections, and rising interest rates have been overblown. The major banks' earnings season begins on October 13, with profitability becoming a key point of validation.

华尔街见闻•2026/10/01 18:41