Inference Service Version Selection for Heterogeneous Deployment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The deployment of AI inference services is labor-intensive and inefficient due to the need for manual intervention in various links, leading to high human costs and low overall efficiency, particularly in managing heterogeneous models and optimizing performance across different computing environments.

Innovation Solution

An automated inference service deployment method that selects and deploys a target version of the service based on performance information of the runtime environment, using modules for obtaining, selecting, and deploying the service, thereby reducing manual intervention and improving efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual intervention is used in deployment links, then deployment can be carefully controlled, but human cost increases and overall efficiency decreases

Engineering Contradiction:
Improvedeployment controlVSAvoiddeployment efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system automatically selects and deploys the appropriate inference service version by self-evaluating runtime environment performance information, eliminating the need for manual intervention while maintaining reliable deployment control through automated decision-making mechanisms

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical deployment operations are replaced with an automated selection mechanism that uses performance information to intelligently choose and deploy the appropriate inference service version, thereby improving efficiency while maintaining control

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If multiple candidate versions of inference service are maintained, then adaptability to different runtime environments is improved, but system complexity increases

Engineering Contradiction:
Improveenvironment adaptabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system manages multiple candidate versions by changing parameters such as quantization precision (e.g., FP16, INT8, INT4) and model architecture configurations, allowing adaptability to different runtime environments while maintaining a structured approach to version management

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

Multiple candidate versions are pre-prepared with different performance characteristics before deployment. The system evaluates runtime environment information and selects the most appropriate pre-prepared version, avoiding the complexity of dynamic adaptation while maintaining high adaptability

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12591462B2Inference service deployment method, device, and storage medium
Publication Date: 2026.03.31 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12591462B2 patent drawing
  • US12591462B2 patent drawing
  • US12591462B2 patent drawing

AI summary

Provided are an inference service deployment method, a device and a storage medium, relating to the field of artificial intelligence technology, and in particular to the field of machine learning and inference service technology. The inference service deployment method includes: obtaining performance information of a runtime environment of a deployment end; selecting a target version of an inference service from a plurality of candidate versions of the inference service of a model according to the performance information of the runtime environment of the deployment end; and deploying the target version of the inference service to the deployment end.