ML Model Swapping With Similarity-Based Subnet Scaling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing ML model optimization techniques are time-consuming, resource-intensive, and fail to adapt efficiently to varying hardware platforms and environmental conditions, leading to suboptimal performance for users with different hardware and performance metric preferences.

Innovation Solution

A machine learning model scaling (MLMS) system that selects and deploys ML models in an energy and communication-efficient manner, using a similarity-based subnet selection process to minimize memory write operations and adapt to real-time system changes, while being agnostic to specific ML model search approaches.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional NAS techniques are used to automatically discover ideal ML models, then model performance is improved, but training time and computational resources increase significantly

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent pre-trains a supernet model that encompasses multiple possible subnet architectures before deployment. This preliminary action allows the system to quickly swap between pre-configured subnets based on runtime conditions without performing time-consuming NAS training at deployment time, thus resolving the contradiction between model performance and training time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically selects and swaps between different subnet architectures from the pre-trained supernet based on real-time hardware conditions, power modes, and performance requirements. This dynamic adaptation eliminates the need for repeated NAS training while maintaining optimal model performance across varying operational conditions.

Inventive Principle:
Principle #15Dynamics

2Productivity

If ML models are optimized for specific hardware platforms, then inference efficiency is improved, but adaptability to different hardware platforms decreases

Engineering Contradiction:
Improveinference efficiencyVSAvoidhardware platform adaptability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal supernet model that can adapt to multiple hardware platforms through subnet swapping. The supernet is designed to encompass various subnet architectures that can be selectively activated based on the target hardware platform, thus achieving both high inference efficiency on specific platforms and broad adaptability across different platforms.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes architectural parameters by swapping between different subnet configurations within the supernet to match specific hardware platform characteristics. This allows the model to optimize inference efficiency for each platform without requiring separate training processes, maintaining both efficiency and adaptability.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If larger ML architectures are used, then model accuracy is improved, but resource consumption and training time increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the large supernet model into multiple smaller subnet architectures that can be selectively deployed. Each subnet represents a different trade-off between accuracy and resource consumption, allowing the system to choose the appropriate subnet size based on available resources while maintaining the option to access higher-accuracy configurations when resources permit.

Inventive Principle:
Principle #1Segmentation

4Reliability

If ML models are manually designed and tuned, then model performance is improved, but design time and expertise requirements increase

Engineering Contradiction:
Improvemodel performanceVSAvoiddesign time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs the time-consuming model design and tuning process in advance by pre-training the supernet and configuring multiple subnet architectures. This preliminary action eliminates the need for manual design and iterative tuning at deployment time, as the system can directly utilize the pre-optimized subnets that are ready for immediate deployment.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260050655A1Machine learning model scaling system with energy efficient network data transfer for power aware hardware
Publication Date: 2026.02.19 INTEL CORP
  • US20260050655A1 patent drawing
  • US20260050655A1 patent drawing
  • US20260050655A1 patent drawing

AI summary

The present disclosure is related to machine learning model swap (MLMS) framework for that selects and interchanges machine learning (ML) models in an energy and communication efficient way while adapting the ML models to real time changes in system constraints. The MLMS framework includes an ML model search strategy that can flexibly adapt ML models for a wide variety of compute system and/or environmental changes. Energy and communication efficiency is achieved by using a similarity-based ML model selection process, which selects a replacement ML model that has the most overlap in pre-trained parameters from a currently deployed ML model to minimize memory write operation overhead. Other embodiments may be described and/or claimed.