Cloud Resource Allocation via Prediction Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cloud virtualization systems, the unpredictable nature of AI application service computing resources leads to issues such as insufficient resource allocation causing service interruptions or excessive resource allocation resulting in resource waste, posing challenges for smart factory automation and energy conservation.
Innovation Solution
A resource allocation system and method for cloud environments that employs prediction models to dynamically adjust resource allocation. This involves obtaining initial resource requirements, generating deployment parameter sets using prediction models, and reconfiguring the cloud environment to ensure optimal resource utilization and prevent service congestion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual resource allocation based on past experience is used, then resource allocation can be performed without complex prediction systems, but resource allocation accuracy deteriorates leading to service interruptions or resource waste
Solution Approach 1:
The system performs preliminary actions by using prediction models to forecast future resource requirements before actual demand occurs. The first prediction model generates deployment parameter sets in advance based on initial resource requirements, and the second prediction model continuously updates predictions based on real resource requirements, enabling proactive resource allocation that prevents service interruptions and optimizes resource utilization.
Solution Approach 2:
The system implements feedback mechanisms where the second prediction model uses real resource requirements (obtained from actual AI workload execution) to generate updated predicted resource requirements. This feedback loop allows the system to continuously learn and improve prediction accuracy, adjusting deployment parameter sets based on actual performance data and changing production line conditions.
2Reliability
If excessive resource allocation is performed to prevent service interruptions, then service reliability is improved, but resource waste increases
Solution Approach 1:
The system applies dynamics by making resource allocation flexible and adaptive rather than static. The deployment parameter sets are dynamically adjusted based on predicted resource requirements generated by the prediction models. As production line conditions change and real resource requirements are observed, the system continuously updates predictions and reconfigures resource allocation, ensuring reliability while avoiding excessive resource allocation.
3Adaptability or versatility
If manual resource allocation is used, then system complexity is reduced, but adaptability to changing production line conditions deteriorates
Solution Approach 1:
The system implements self-service by enabling automated resource allocation through prediction models that independently forecast resource requirements and generate deployment parameter sets. The resource allocation system serves itself by using the second prediction model to monitor real resource requirements and automatically adjust predictions without manual intervention, allowing the system to adapt to changing production line conditions autonomously.
Data Source
AI summary
A resource allocation method for a cloud environment includes steps performed by one or more servers. These steps include: obtaining an initial resource requirement, generating a first deployment parameter set according to the initial resource requirement by a first prediction model, configuring the cloud environment according to the first deployment parameter set by a resource allocator, obtaining and inputting the requests to a machine learning model in the cloud environment and generating a real resource requirement, generating a predicted resource requirement at least according to the real resource requirement by a second prediction model, when detecting that the cloud environment is in a busy state according to the predicted resource requirement by the resource allocator, generating a second deployment parameter set according to the real resource requirement by the first prediction model; and reconfiguring the cloud environment according to the second deployment parameter set by the resource allocator.

