Reinforcement Learning Allocation Model for Computing Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing allocation models for managing computing systems, particularly those using reinforcement learning, face challenges such as high computational requirements, limited applicability in small systems with limited power, and the need for extensive manual labor and expert knowledge integration, leading to inefficiencies and unsatisfactory performance.
Innovation Solution
An initial allocation model based on reinforcement learning is constructed, incorporating a man-machine interaction process to generate more effective training data, combining machine learning with human expertise to improve the accuracy and efficiency of workload allocation across multiple computing units.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If reinforcement learning-based allocation models are used to manage computing systems, then allocation accuracy can be improved, but computational requirements and training complexity increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-training the allocation model using historical operation data and computing unit state information before actual deployment. This pre-training phase prepares the model in advance with baseline knowledge, reducing the complexity of subsequent real-time allocation decisions while maintaining high accuracy.
Solution Approach 2:
The patent implements partial action by selectively training the allocation model with a subset of critical features and parameters rather than processing all possible data dimensions. This approach achieves sufficient allocation accuracy without the excessive computational burden of comprehensive training, balancing precision and complexity.
2Reliability
If extensive manual labor and expert knowledge integration are applied in model training, then model performance can be improved, but training time and human resource requirements increase
Solution Approach 1:
The patent applies self-service by enabling the allocation model to automatically learn from operational data and improve its performance through continuous feedback loops. The system self-adjusts parameters and optimizes allocation strategies without requiring extensive manual intervention or expert knowledge input, reducing training time while maintaining model reliability.
Solution Approach 2:
The patent implements feedback mechanisms where the allocation model receives performance metrics and allocation outcomes, automatically adjusting its parameters based on this feedback. This closed-loop approach allows the model to improve performance over time with minimal human intervention, reducing both training time and dependency on expert knowledge while maintaining high reliability.
3Productivity
If allocation models are trained with limited data in small computing systems, then training resources are reduced, but model accuracy and applicability deteriorate
Solution Approach 1:
The patent applies universality by designing an allocation model trained on diverse operation types and computing unit configurations that can generalize across different system sizes and workloads. The model learns transferable patterns from varied training data, enabling it to maintain high accuracy in small computing systems with limited training data by applying knowledge from broader operational contexts.
Data Source
AI summary
A method includes: acquiring a set of operations to be performed on multiple computing units in the computing system; determining, based on the set of operations, the state of the multiple computing units, and an allocation model, an allocation action for allocating the set of operations to the multiple computing units and a reward for the allocation action, wherein the allocation model describes an association relationship among a set of operations, the state of multiple computing units, the allocation action for allocating the set of operations to the multiple computing units, and the reward for the allocation action; receiving an adjustment for the reward in response to determining that a match degree between the reward for the allocation action and a performance index of the computing system after the allocation action is performed satisfies a predetermined condition; and generating, based on the adjustment, training data for updating the allocation model.


