Inference Model Manager Resource Reassignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In a distributed environment, data processing systems hosting inference models may face insufficient computing resources, leading to inadequate performance and downtime, especially when prioritizing higher priority models over lower priority ones during resource constraints.
Innovation Solution
An inference model manager dynamically re-assigns data processing systems to prioritize higher priority inference models by distributing inference model portions across multiple systems based on available computing resources and priority rankings, ensuring timely execution and maintaining functionality during disruptions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data processing systems host multiple inference models with different priority levels, then the system can meet diverse consumer needs, but computing resource capacity becomes insufficient when prioritizing higher priority models
Solution Approach 1:
The patent segments inference models into priority levels (first priority, second priority, third priority) and divides computing resources into allocatable portions. This segmentation allows the system to allocate resources differently based on model importance, ensuring high-priority models receive sufficient computing capacity while lower-priority models share remaining resources, thus resolving the contradiction between versatility and resource capacity.
Solution Approach 2:
The system applies partial action by allocating computing resources proportionally based on priority levels rather than uniformly. Higher-priority models receive excessive action (more resources) while lower-priority models receive partial action (fewer resources). This partial allocation strategy enables the system to maintain diverse model capabilities while managing limited computing resource capacity through prioritization.
2Productivity
If the system dynamically re-assigns data processing systems to prioritize higher priority models, then timely execution of high priority models is ensured, but downtime increases for lower priority models
Solution Approach 1:
The patent implements dynamic resource reassignment where the system continuously monitors computing resource capacity and dynamically adjusts model deployments based on current conditions. When capacity constraints are detected, the system dynamically migrates lower-priority models to different data processing systems or reduces their resource allocation, ensuring high-priority models maintain timely execution while accepting temporary downtime for lower-priority models. This dynamic adaptation resolves the contradiction between productivity and time loss.
Solution Approach 2:
The system maintains continuity of useful action for high-priority models by ensuring they continue to execute timely regardless of resource fluctuations. The dynamic reassignment mechanism ensures that high-priority model execution remains uninterrupted, while lower-priority models experience temporary disruptions. This continuity focus on high-priority models resolves the contradiction between maintaining productivity and accepting downtime for lower-priority models.
3Reliability
If inference models are distributed across multiple data processing systems, then system reliability improves, but computing resource capacity is分散 (dispersed) and insufficient for all models
Solution Approach 1:
The patent applies local quality by assigning different resource allocation characteristics to different models based on their priority levels. High-priority models receive concentrated resource allocation at specific data processing systems, while lower-priority models receive dispersed allocation across remaining systems. This local differentiation ensures that reliability is maintained through distributed deployment while concentrating resources where they are most needed, resolving the contradiction between reliability and concentrated resource capacity.
Data Source
AI summary
Methods and systems for managing execution of inference models across multiple data processing systems are disclosed. To manage execution of inference models across multiple data processing systems, a system may include an inference model manager and any number of data processing systems. The inference model manager may obtain operational capability data for the inference models from the data processing systems. The inference model manager may use the operational capability data to determine whether the data processing systems have access to sufficient computing resources to complete timely execution of the inference models. If the data processing systems do not have access to sufficient computing resources to complete timely execution of the inference models, the inference model manager may re-assign one or more data processing systems to re-balance the computing resource load and support continued operation of at least a portion of the inference models.


