Nebius (NBIS.US) Acquires Israeli Startup Inferize to Double Down on AI Inference Business, Deal Value Could Reach Up to $150 Million
AI cloud infrastructure service provider Nebius has announced the acquisition of Israeli AI startup Inferize, further strengthening its AI inference infrastructure capabilities.
AI cloud infrastructure service provider Nebius (NBIS.US) has announced the acquisition of Israeli AI startup Inferize to further strengthen its AI inference infrastructure capabilities. Inferize focuses on reducing idle time for graphics processing units (GPUs) and accelerating the deployment of large AI models. According to Israeli tech media Calcalist, the deal value is estimated to be between $100 million and $150 million, though Nebius did not disclose specific financial terms.
The core of this acquisition lies in improving GPU resource utilization efficiency, particularly amid rapidly changing AI inference demand, helping Nebius allocate computing resources more quickly and reduce idle time for expensive computing equipment.
Targeting the AI Inference Efficiency Bottleneck to Reduce GPU Idle Time
Inferize's core technology primarily addresses the "cold start" problem in AI model deployment. Cold start refers to the process during which an AI model must complete model loading and related resource initialization before it begins processing user requests. When a new inference instance starts up, the GPU may need to wait for model loading to complete before it can formally begin computation, resulting in brief periods of resource idling.
This problem is especially pronounced when AI inference demand suddenly increases. As user requests surge, cloud service providers need to rapidly spin up more inference instances, but if model loading takes a long time, the newly added GPU resources cannot be immediately put into use, affecting service response speed and overall computing utilization efficiency.
Inferize's technology aims to shorten this waiting process, enabling newly added inference resources to be brought online more quickly and making compute scaling more closely aligned with actual demand changes.
For Nebius, this means the company is expected to more flexibly meet customer needs and improve the utilization efficiency of its existing infrastructure without having to maintain large amounts of idle GPUs as backup resources over the long term.
Nebius Chief Technology Officer Danila Shtan stated that efficiently running AI inference services requires not only faster GPUs and optimized models, but also a system that can respond to demand changes in a timely manner, including rapidly providing additional computing capacity when customer demand increases. He noted that Inferize not only brings technology that accelerates this process but also has an engineering team with extensive experience in GPU systems.
Nebius plans to incorporate Inferize's engineers into its inference business team and integrate the related technology into its AI inference platform Token Factory to enhance the platform's ability to respond to customer needs and enable existing infrastructure to handle more actual computing tasks.
Shtan also stated that the Inferize team's future contributions will not be limited to this initial technology integration.
Reducing Backup Compute Costs and Improving AI Infrastructure Operational Efficiency
Inferize co-founder and CEO Guy Bortnikov stated that in order to respond to customer demand at any time, cloud service providers typically need to maintain a certain scale of backup GPU resources, and these standby computing devices themselves represent additional costs. He noted that Inferize was founded precisely to reduce this portion of costs, and joining Nebius will enable the related technology to be directly applied to an operational AI cloud platform.
In the AI inference business, the scheduling efficiency of computing resources is directly related to infrastructure operating costs. Because GPU equipment is expensive, if enterprises need to retain large amounts of idle computing power over the long term to cope with potential demand peaks, it may affect overall return on investment.
By shortening model startup time and accelerating the deployment of computing resources, Nebius is expected to reduce its reliance on backup GPU capacity while improving service elasticity.
It is worth noting that Bortnikov previously co-founded Granulate, a computing infrastructure optimization software company. The company mainly developed technology to improve the operating efficiency of computing resources and was acquired by Intel Corporation (INTC.US) in 2022. This background also reflects the Inferize team's relevant experience in the field of computing infrastructure optimization.
Successive Acquisitions of AI Technology Companies as Nebius Continues to Strengthen Its Inference Platform
This acquisition of Inferize is one of a series of recent acquisitions by Nebius centered on AI inference and model optimization. Previously, Nebius completed the acquisition of Eigen AI for a deal value of $643 million. In addition, the company recently acquired AI technology company Clarifai, but did not disclose the specific amount of that transaction.
These acquisitions are all aimed at enhancing Nebius's AI inference and model optimization capabilities and further improving its AI infrastructure services.
Unlike infrastructure investments that mainly focus on the computing power required for AI model training, AI inference places greater emphasis on operational efficiency after models are put into actual use, including request processing speed, computing resource scheduling, model deployment, and cost per unit of computing power.
As Nebius continues to integrate related technologies, its business layout is also extending from providing GPU computing resources to technologies and services that improve the actual operating efficiency of AI models.
This acquisition of Inferize will further supplement Nebius's technical capabilities in dynamic compute scheduling and rapid model deployment. However, the extent to which the related technology integration can bring cost savings and efficiency improvements still remains to be verified through subsequent operational performance.
Related Articles

NAND supply tightness may persist until 2028! Citi is bullish on AI demand driving memory chip price increases, reiterates "Buy" rating on SanDisk (SNDK.US).

Down 24% year-to-date! McDonald's Corporation (MCD.US) stock could be headed for its longest weekly losing streak in over 20 years; $8.5 billion new strategy sparks market concerns.

UBS Group AG: iPhone 18 Pro series performance during the waiting period is lackluster, assigns a "Neutral" rating to Apple Inc. (AAPL.US)
NAND supply tightness may persist until 2028! Citi is bullish on AI demand driving memory chip price increases, reiterates "Buy" rating on SanDisk (SNDK.US).

Down 24% year-to-date! McDonald's Corporation (MCD.US) stock could be headed for its longest weekly losing streak in over 20 years; $8.5 billion new strategy sparks market concerns.

UBS Group AG: iPhone 18 Pro series performance during the waiting period is lackluster, assigns a "Neutral" rating to Apple Inc. (AAPL.US)






