Method and device for adjusting number of algorithm containers, electronic equipment and storage medium

By dynamically adjusting the number of algorithm containers and combining a weighted summation mechanism of measured and predicted loads, the inefficiency caused by static resource configuration is solved, achieving precise matching of resources and loads and improving the system's resource utilization and performance.

CN120849023AActive Publication Date: 2025-10-28PEOPLES DAILY NEWS AGENCY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511349028.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-10-28
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

In existing technologies, static resource configuration methods result in low GPU resource utilization, which cannot adapt to dynamic business needs, leading to resource waste and limitations on the execution of other tasks.

Method used

By acquiring measured and predicted loads, dynamically updating the prediction confidence, using weighted summation to form a self-correction mechanism, and combining adjustment coefficients to optimize the number of containers in the algorithm, resource and load matching is achieved.

Benefits of technology

It improved resource utilization, optimized resource allocation, ensured that the system could respond to actual needs in a timely manner, and enhanced system performance and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849023A_ABST
    Figure CN120849023A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of container arrangement, and discloses an algorithm container number adjusting method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining an actual measurement load of a current monitoring period and a predicted load predicted in a previous monitoring period for the current monitoring period; updating the prediction confidence coefficient of the previous monitoring period according to the actually measured load of the current monitoring period to obtain the prediction confidence coefficient of the current monitoring period; based on the prediction confidence coefficient of the current monitoring period, performing weighted summation on the actually measured load of the current monitoring period and the predicted load of the current monitoring period to obtain a comprehensive load of the current monitoring period; determining an adjustment coefficient according to a difference value between the comprehensive load and a preset standard load; and obtaining the number of the adjusted algorithms Pod according to the total number of the current algorithms Pod and the adjustment coefficient. The problem that the resource utilization rate is low due to a static resource configuration mode is solved, accurate matching of load requirements is achieved, and the resource utilization rate is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of container orchestration technology, and in particular to a method, apparatus, electronic device, and storage medium for adjusting the number of algorithmic containers. Background Art

[0002] Currently, for applications that require GPUs to process complex tasks, the common practice is to pre-deploy GPU algorithm modules at the maximum scale (i.e., launch as many Pods as possible) when the task pool is expected to reach its peak. This approach is based on a conservative estimate, aiming to ensure sufficient computing power to handle all requests even under the busiest conditions. This model is a static resource configuration model, using a Deployment configuration with a fixed number of replicas, and the number of Pods is set based on historical peak load, which cannot adapt to dynamically changing business needs.

[0003] However, while this static resource configuration method can guarantee service performance under high load, it also leads to low resource utilization: long-term occupation of excessive GPU resources without full utilization not only increases hardware costs but also limits the execution of other tasks that may need these resources more urgently. Summary of the Invention

[0004] The purpose of this invention is to provide at least one method, apparatus, electronic device, and storage medium for adjusting the number of algorithm containers, which can at least solve the problem of low resource utilization caused by static resource configuration, and at least achieve the effect of matching resource allocation with load and improving resource utilization.

[0005] To address the aforementioned technical problems, at least one embodiment of this application provides a method for adjusting the number of algorithm containers, including: Obtain the measured load for the current monitoring period and the predicted load for the current monitoring period from the previous monitoring period. The prediction confidence level of the current monitoring period is obtained by updating the prediction confidence level of the previous monitoring period based on the measured load of the current monitoring period. Based on the prediction confidence of the current monitoring period, the measured load and the predicted load of the current monitoring period are weighted and summed to obtain the comprehensive load of the current monitoring period. The adjustment coefficient is determined based on the difference between the comprehensive load and the preset standard load; The adjusted number of algorithm Pods is obtained based on the current total number of algorithm Pods and the adjustment coefficient.

[0006] At least one embodiment of this application also provides an apparatus for adjusting the number of algorithm containers, comprising: The load acquisition module is used to acquire the measured load of the current monitoring period and the predicted load for the current monitoring period predicted in the previous monitoring period. The confidence calculation module is used to update the predicted confidence of the previous monitoring period based on the measured load of the current monitoring period, so as to obtain the predicted confidence of the current monitoring period. The comprehensive load calculation module is used to perform a weighted summation of the measured load and the predicted load of the current monitoring period based on the prediction confidence level of the current monitoring period, so as to obtain the comprehensive load of the current monitoring period. The adjustment coefficient determination module is used to determine the adjustment coefficient based on the difference between the comprehensive load and the preset standard load; The algorithm Pod number adjustment module is used to obtain the adjusted number of algorithm Pods based on the current total number of algorithm Pods and the adjustment coefficient.

[0007] At least one embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described method for adjusting the number of algorithm containers.

[0008] At least one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for adjusting the number of algorithm containers.

[0009] The algorithmic container number adjustment method, apparatus, electronic device, and storage medium provided in the embodiments of this application dynamically update the prediction confidence level by acquiring the measured load and the predicted load, based on the error between the current measured load and the predicted load for the current period. Higher confidence levels indicate more reliable predictions, and the prediction error is directly fed back into the confidence level calculation, forming a self-correcting mechanism. Using the prediction confidence level of the previous period as a weight, the current measured load and the predicted load for the current period are weighted and summed to obtain a comprehensive load, balancing real-time performance and predictability. An adjustment coefficient is determined based on the difference between the comprehensive load and a preset standard load, transforming load differences into quantifiable adjustment signals and avoiding subjective judgment. Combining the current total number of Pods and the adjustment coefficient, the target number of Pods is calculated, achieving precise matching of load demands and optimizing resource utilization.

[0010] In some optional embodiments, updating the prediction confidence of the previous monitoring period based on the measured load of the current monitoring period includes: After the actual load of the current monitoring period arrives, the prediction error of the current monitoring period is calculated, and the error record in the error sliding window is updated. The error sliding window includes error data from multiple historical monitoring periods. Calculate the statistical characteristics of the error within the updated error sliding window, wherein the statistical characteristics include at least one of the mean error, standard deviation, and range; Take the quantile of the historical values ​​of the statistical features under each historical monitoring period within the updated error sliding window according to a preset proportion to generate the feature threshold of the corresponding statistical features; The correction coefficient is calculated based on the statistical characteristics and the characteristic threshold. The prediction confidence of the previous monitoring period is adjusted using the correction coefficient to obtain the prediction confidence of the current monitoring period.

[0011] In this embodiment, a fixed-size error sliding window is maintained to store prediction error data from multiple historical monitoring periods. The window is updated over time, automatically discarding outdated data to ensure that the confidence assessment reflects the latest load characteristics. Statistical characteristics such as the mean error, standard deviation, and range of the error within the sliding window are calculated to more comprehensively quantify the error distribution and improve the accuracy of the prediction confidence for the current monitoring period.

[0012] In some optional embodiments, when there are multiple statistical features, the step of calculating the correction coefficient based on the statistical features and the feature threshold includes: For each statistical feature, the sub-correction coefficient of the statistical feature is calculated based on the ratio of the statistical feature to the feature threshold corresponding to the statistical feature. The correction coefficients are obtained by weighted summation of the multiple sub-correction coefficients.

[0013] In this embodiment, the ratio of each statistical feature (such as mean error, standard deviation, and range) to its corresponding threshold is calculated separately to generate a sub-correction coefficient, instead of directly using the original value. This standardizes the impact of error and unifies the evaluation scale. The structure of independent calculation of sub-correction coefficients plus weighted fusion retains the sensitivity of individual features while achieving global stability through weighting.

[0014] In some optional embodiments, the prediction confidence level for the current monitoring period is calculated according to the following formula:

[0015] =

[0016] in, The prediction confidence level for the current monitoring period. The prediction confidence level for the previous monitoring period. For statistical characteristics, , and These are the mean error, standard deviation, and range, respectively. The weights of the sub-correction coefficients for statistical feature i. This is the feature threshold.

[0017] In some optional embodiments, the overall load for the current monitoring period is calculated according to the following formula:

[0018] in, The overall load for the current monitoring period, The predicted load for the previous monitoring period, The measured load for the current monitoring period. This represents the prediction confidence level for the current monitoring period.

[0019] In this embodiment, when the prediction confidence is high, the number of containers is adjusted in advance according to the predicted trend to avoid resource waste or shortage; when the prediction confidence is low, it is quickly adjusted according to the measured load to ensure that the system can respond to actual needs in a timely manner and improve the performance and stability of the system.

[0020] In some optional embodiments, the adjustment factor is calculated according to the following formula:

[0021] in, To adjust the coefficient, The overall load for the current monitoring period, For the preset standard load, This is the preset adjustment range.

[0022] In this embodiment, the adjustment coefficient K is linearly proportional to the difference between the comprehensive load and the standard load. This ensures that the adjustment range is directly proportional to the load deviation; that is, the larger the deviation, the larger the adjustment range, thereby enabling rapid correction of resource shortages or excesses. By setting an appropriate N value, the step size of each adjustment can be controlled, making the adjustment process smoother.

[0023] In some optional embodiments, it also includes: The duration of the next monitoring cycle is adjusted based on the prediction confidence level of the current monitoring cycle.

[0024] In this embodiment, the inverse linkage between confidence level and monitoring frequency helps to accurately allocate limited resources to periods of high uncertainty, thereby maximizing resource efficiency while ensuring the system's perception capabilities.

[0025] In some optional embodiments, adjusting the duration of the next monitoring period based on the prediction confidence of the current monitoring period includes: Based on a preset mapping relationship between confidence level and cycle duration, the duration of the next monitoring cycle is determined according to the predicted confidence level of the current monitoring cycle, where there is a positive correlation between confidence level and cycle duration.

[0026] In this embodiment, the monitoring cycle is shorter when the confidence level is low, so as to shorten the interval for intensive monitoring and thus improve the prediction accuracy; when the confidence level is high, the monitoring cycle is longer, so as to extend the interval to save resources, thus achieving an intelligent balance between monitoring accuracy and resource consumption.

[0027] In some optional embodiments, the duration of the next monitoring cycle is calculated according to the following formula: ,

[0028] in, The duration of the next monitoring cycle, This is the minimum value of the preset monitoring cycle. This is the maximum value of the preset monitoring cycle. The prediction confidence level for the current monitoring period. This is the preset minimum prediction confidence level. This is the preset maximum prediction confidence level. It is an exponential coefficient used to adjust the shape of the mapping curve between confidence level and period duration.

[0029] In this embodiment, through exponential mapping, and This enables a rapid decrease in monitoring cycle length at low confidence levels, thereby improving the system's response speed. Attached Figure Description

[0030] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.

[0031] Figure 1 This is a flowchart of an algorithm container number adjustment method provided in one embodiment of this application. Figure 1 ; Figure 2 This is a flowchart of an algorithm container number adjustment method provided in another embodiment of this application. Figure 2 ; Figure 3 This is a schematic diagram of an algorithm container number adjustment device provided in another embodiment of this application. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.

[0033] This invention proposes a method for adjusting the number of algorithm containers. The implementation details of the method for adjusting the number of algorithm containers in this embodiment are described below. The following content is only for the convenience of understanding the implementation details and is not necessary for implementing this solution.

[0034] Example 1: The algorithm container number adjustment method of this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities. Its specific process can be as follows: Figure 1 As shown, it includes: Step 110: Obtain the measured load for the current monitoring period and the predicted load for the current monitoring period predicted in the previous monitoring period. In this embodiment, the monitoring system triggers a monitoring and adjustment operation at preset fixed time intervals, such as 5 minutes or 10 minutes. The measured load for the current monitoring period is the actual load value measured at the end of the current period, such as CPU utilization, request QPS, and memory usage. The predicted load is the predicted load for the current period obtained from the previous monitoring period using a preset prediction model, such as a time series prediction model. The load for the next period is predicted at the end of each monitoring period.

[0035] Step 120: Update the prediction confidence of the previous monitoring period based on the measured load of the current monitoring period to obtain the prediction confidence of the current monitoring period. Specifically, prediction confidence reflects the accuracy of the prediction model. If the error between the current measured load and the predicted load is small, the confidence increases; conversely, it decreases. The confidence index is dynamically updated based on the current measured load to verify the prediction accuracy of the previous period, thus reflecting the reliability of the prediction model.

[0036] In practical applications, prediction models may experience short-term inaccuracies due to sudden traffic spikes, data noise, or other factors. If the actual load for the current period differs significantly from the predicted load, it indicates that the prediction model may be inaccurate under the current conditions, requiring a reduction in the confidence level. This reduces the weight of the predicted load in subsequent weighted summations. By dynamically updating the confidence level, the system can quickly respond to these changes, preventing resource allocation errors caused by prediction mistakes.

[0037] Step 130: Based on the prediction confidence of the current monitoring period, the measured load of the current monitoring period and the predicted load of the current monitoring period are weighted and summed to obtain the comprehensive load of the current monitoring period. In this embodiment, a comprehensive load balancing approach is used to smoothly adjust the number of algorithm Pods. Adjusting solely based on measured load can easily lead to overreaction to instantaneous peaks, but scaling down after the peak disappears results in frequent fluctuations in the number of containers, increasing system overhead and risk. Adjusting solely based on predicted load struggles to handle unexpected situations not captured by the model, such as traffic surges caused by hotspot events, where the predicted load is lower than the surged load, making the adjusted number of algorithm Pods insufficient to cope with the traffic spike. The comprehensive load balancing approach combines the stability of prediction with the real-time nature of measured loads, facilitating smooth scaling up and down and avoiding resource fluctuations.

[0038] Step 140: Determine the adjustment coefficient based on the difference between the comprehensive load and the preset standard load; In this embodiment, the standard load is the reasonable load threshold expected to be borne by the algorithm Pod, which can be expressed in the form of resource utilization (such as CPU, memory) or business metrics (such as QPS, latency). For example: CPU standard load: 70% (with a 30% buffer reserved to handle sudden traffic surges); Standard QPS load: 5000 requests / second; Or define a custom metric: the load level corresponding to a prediction model inference latency of ≤200ms.

[0039] The target line is defined by the standard load, and the direction of the difference directly indicates the direction of adjustment. When the total load is greater than the standard load, the capacity needs to be expanded to reduce the pressure on a single Pod; if the total load is less than the standard load, the capacity can be reduced to save resources.

[0040] Step 150: Based on the current total number of algorithm Pods and the adjustment coefficient, obtain the adjusted number of algorithm Pods.

[0041] In this embodiment, if the adjustment coefficient is greater than 1, the capacity is expanded; if the adjustment coefficient is less than 1, the capacity is reduced. When calculating the adjustment coefficient, the difference can be mapped to the adjustment coefficient to achieve proportional control. Alternatively, a piecewise function can be used: if the difference is greater than or equal to the first threshold but less than the second threshold, the adjustment coefficient is 1.2; if the difference is greater than the second threshold, the adjustment coefficient is 1.5.

[0042] In summary, the algorithm for adjusting the number of containers provided in this embodiment dynamically updates the prediction confidence level based on the error between the current measured load and the predicted load for the current period, by acquiring the measured load and the predicted load. Higher confidence levels indicate more reliable predictions, and the prediction error is directly fed back into the confidence level calculation, forming a self-correcting mechanism. Using the prediction confidence level of the previous period as a weight, the current measured load and the predicted load for the current period are weighted and summed to obtain the comprehensive load, balancing real-time performance and predictability. An adjustment coefficient is determined based on the difference between the comprehensive load and a preset standard load, transforming load differences into quantifiable adjustment signals and avoiding subjective judgment. Combining the current total number of Pods and the adjustment coefficient, the target number of Pods is calculated, achieving precise matching of load requirements and optimizing resource utilization.

[0043] In some embodiments, updating the prediction confidence of the previous monitoring period based on the measured load of the current monitoring period includes: calculating the prediction error of the current monitoring period after the actual load of the current monitoring period arrives, and updating the error records in the error sliding window, wherein the error sliding window includes error data from multiple historical monitoring periods; calculating the statistical characteristics of the error in the updated error sliding window, wherein the statistical characteristics include at least one of mean error, standard deviation, and range; taking a preset proportion of quantiles for the historical values ​​of the statistical characteristics in each historical monitoring period in the updated error sliding window to generate a feature threshold for the corresponding statistical characteristics; calculating a correction coefficient based on the statistical characteristics and the feature thresholds, and adjusting the prediction confidence of the previous monitoring period using the correction coefficient to obtain the prediction confidence of the current monitoring period.

[0044] In this embodiment, prediction confidence is an indicator representing the reliability of the prediction result, typically ranging from 0 to 1. A value closer to 1 indicates a more reliable prediction, while a value closer to 0 indicates a less reliable prediction. By adjusting the correction coefficient, the prediction confidence can more accurately reflect the actual situation of the current prediction. For example, if the prediction confidence of the previous monitoring period was 0.8, after adjusting the correction coefficient, the prediction confidence of the current monitoring period becomes 0.7. This indicates that the reliability of the prediction has decreased due to a certain deviation between the measured load and the predicted value in the current monitoring period.

[0045] After the actual load of the current monitoring period arrives, the system compares the actual load value with the predicted load value for the same period, which was previously derived based on the prediction model. The prediction error for the current monitoring period is obtained through specific error calculation methods, such as absolute error (absolute value of actual load - absolute value of predicted load) or relative error ((absolute value of actual load - absolute value of predicted load) / absolute value of predicted load).

[0046] An error sliding window is a data structure with a fixed capacity used to store error data from multiple historical monitoring periods. When new error data for the current monitoring period is generated, the oldest error data is removed from the window according to a first-in, first-out (FIFO) principle, while the new error data is added. This ensures that the error sliding window always contains the latest and most continuous error records over a period of time, providing a data foundation for subsequent analysis and prediction of error trends. For example, after the actual load arrives, the prediction error at the current time point is calculated. =|Measured Load - Predicted Load|; Error sliding window W records historical error sequences. , The size of the error sliding window, It can be 30.

[0047] The mean error is the arithmetic mean of all error data within the error sliding window. It reflects the average level of prediction error over a period of time and provides an overall indication of prediction accuracy. The standard deviation measures the dispersion of error data, reflecting the degree of deviation between individual error values ​​and the mean error. A larger standard deviation indicates greater fluctuation in error data and poorer prediction stability; a smaller standard deviation indicates relatively concentrated error data and more stable prediction results. The range is the difference between the maximum and minimum error values ​​within the error sliding window, showing the range of error data variation. A larger range means that the prediction error may fluctuate significantly in extreme cases; a smaller range indicates a relatively narrow range of error variation. In practical applications, the feature thresholds for corresponding statistical features can be set by taking the quantiles of historical values ​​set using any one of the mean error, standard deviation, or range, or multiple feature thresholds can be set by combining various statistical features, and the correction coefficient can be determined by combining these feature thresholds.

[0048] The system calculates quantiles for historical values ​​of statistical features across different monitoring periods within the updated error sliding window, using a preset percentage. A quantile is the value at a specific position when a set of data is arranged in ascending order. For example, the median is the 50th quantile, representing a critical value that divides the data into upper and lower parts. By selecting different preset percentages, such as 25% and 75%, corresponding quantiles can be calculated, reflecting the distribution of statistical features at different levels. Based on the calculated quantiles, feature thresholds for the corresponding statistical features are generated. These thresholds serve as a criterion to determine whether the current statistical feature is within a normal range. If it is not within the normal range, a correction coefficient is calculated based on the statistical feature and the feature threshold. The calculation method for the correction coefficient can be designed according to specific application scenarios and requirements. The calculated correction coefficient is used to adjust the prediction confidence level of the previous monitoring period to obtain the prediction confidence level for the current monitoring period.

[0049] In some embodiments, when there are multiple statistical features, the step of calculating the correction coefficient based on the statistical features and the feature threshold includes: for each statistical feature, calculating a sub-correction coefficient of the statistical feature based on the ratio of the statistical feature to the feature threshold corresponding to the statistical feature; and obtaining the correction coefficient based on the weighted sum of the multiple sub-correction coefficients.

[0050] In some embodiments, the prediction confidence level for the current monitoring period is calculated according to the following formula:

[0051] =

[0052] in, The prediction confidence level for the current monitoring period. The prediction confidence level for the previous monitoring period. For statistical characteristics, , and These are the mean error, standard deviation, and range, respectively. The weights of the sub-correction coefficients for statistical feature i. This is the feature threshold.

[0053] Specifically, this embodiment combines mean error, standard deviation, and range to construct a multidimensional confidence function, wherein, For sub-correction coefficients, These are the adjustable weighting coefficients corresponding to the sub-correction coefficients. For example, the weight of the sub-correction coefficient corresponding to the mean error is 0.4, the weight of the sub-correction coefficient corresponding to the standard deviation is 0.3, and the weight of the sub-correction coefficient corresponding to the range is 0.3. The overall value is a correction factor.

[0054] This embodiment also adds confidence level boundary constraints to avoid extreme values ​​affecting system stability. The confidence level ranges from 0.1 to 0.9.

[0055] In some embodiments, the overall load for the current monitoring period is calculated according to the following formula:

[0056] in, The overall load for the current monitoring period, The predicted load for the previous monitoring period, The measured load for the current monitoring period. This represents the prediction confidence level for the current monitoring period.

[0057] When the prediction confidence is high, the number of containers is adjusted in advance according to the predicted trend to avoid resource waste or shortage; when the prediction confidence is low, it is quickly adjusted according to the measured load to ensure that the system can respond to actual needs in a timely manner and improve the system's performance and stability.

[0058] In some embodiments, the adjustment factor is calculated according to the following formula:

[0059] in, To adjust the coefficient, The overall load for the current monitoring period, For the preset standard load, This is the preset adjustment range.

[0060] Based on the current total number of algorithm Pods and the adjustment coefficient, the adjusted number of algorithm Pods can be obtained as follows:

[0061] in, The adjusted number of algorithm Pods, This represents the total number of Pods in the current algorithm.

[0062] In some embodiments, the method further includes: adjusting the duration of the next monitoring period based on the prediction confidence level of the current monitoring period.

[0063] In this embodiment, when the prediction confidence is high, extending the monitoring cycle can reduce redundant data collection and processing, thereby lowering computational costs. When the confidence is low, shortening the cycle allows limited monitoring resources to be concentrated on high-risk periods of prediction failure, preventing the escalation of faults due to excessively long monitoring intervals.

[0064] In some embodiments, adjusting the duration of the next monitoring period based on the predicted confidence level of the current monitoring period includes: determining the duration of the next monitoring period based on the predicted confidence level of the current monitoring period, according to a preset confidence level-period duration mapping relationship, wherein the confidence level and the period duration are positively correlated.

[0065] In this embodiment, the mapping relationship can be a proportional mapping, an exponential mapping, or a piecewise function mapping, as long as it satisfies the requirement that the higher the confidence level, the longer the period duration, and the lower the confidence level, the shorter the period duration.

[0066] In some embodiments, the duration of the next monitoring cycle is calculated according to the following formula: ,

[0067] in, The duration of the next monitoring cycle, This is the minimum value of the preset monitoring cycle. This is the maximum value of the preset monitoring cycle. The prediction confidence level for the current monitoring period. This is the preset minimum prediction confidence level. This is the preset maximum prediction confidence level. It is an exponential coefficient used to adjust the shape of the mapping curve between confidence level and period duration.

[0068] In this embodiment, A value greater than 1 indicates a steep curve in the low-confidence region and a flat curve in the high-confidence region. This allows the system to quickly shorten the duration of the next monitoring cycle at low confidence levels, thereby improving the prediction accuracy for the next monitoring cycle.

[0069] Low confidence levels mean that the predictive model is failing in the current environment, such as a sudden traffic spike or a change in business model. In this case, rapidly shortening the monitoring cycle allows for faster capture of actual load changes. Previously, monitoring was done every 10 minutes; now it might be done every 5 minutes. The time to detect traffic surges is reduced from an average of 5 minutes to 2.5 minutes, better handling scenarios with rapid business growth.

[0070] Example 2: Based on the above embodiments, this embodiment provides a specific application example. The method for adjusting the number of algorithm containers provided in this embodiment is applied to a system for batch proofreading articles. This system has the following technical components: Kubernetes (K8s) API: Used to interact with K8s clusters and manage the creation, deletion, and replication of algorithm module Pods.

[0071] Unified image repository: Stores container images of algorithm modules, ensuring that all nodes can pull the latest image version.

[0072] Deployment YAML file: Predefined deployment configuration for each algorithm module, including parameters such as resource requests and limits.

[0073] Database - Redis: Used as a cache database for the task pool, enabling fast reading and writing of the number of unprocessed tasks.

[0074] Database - MySQL: Persistently stores the results of completed tasks, providing reliable data storage and query services.

[0075] Front-end management page: Responsible for building the user interface, allowing administrators to maintain algorithm tasks and display information such as the current number of tasks, the number of tasks processed, and the usage of K8s cluster resources.

[0076] Core scheduling service: Implements the main business logic, including communicating with the K8s API, monitoring the status of the task pool, evaluating the performance of the GPU algorithm module, and making scaling decisions.

[0077] GPU resource pool management: All GPUs are divided into three parts: basic resource pool, shared resource pool, and dedicated resource pool, each corresponding to different types of resource allocation strategies.

[0078] like Figure 2 The diagram illustrates the entire process of adjusting the number of algorithm Pods when using the aforementioned system for article review. For example, if 1000 articles are uploaded for review simultaneously, after task splitting, 15,000 review tasks (text + images) are generated. The system dynamically scales up according to the algorithm container number adjustment method described in the embodiment, automatically expanding from 5 Pods to 35 Pods. Tasks are distributed according to the Pod's processing capacity. The 35 GPU Pods process the tasks simultaneously, and finally, the results are aggregated using Spark streaming. After processing is complete, the system shrinks back to 5 Pods.

[0079] The specific implementation details and core processes between the technical components are as follows: I. Task Reception and Distribution The data processing flow is as follows: 1.1 Input data: Details of the task submitted by the user (document content, review rules, QoS level tags).

[0080] Task parameter validation: The core scheduling service validates the validity of input parameters (such as file format and rule ID) using regular expressions or JSON Schema.

[0081] 1.2 Processing steps: Task ID generation: Use UUID to generate unique task IDs to ensure global uniqueness.

[0082] Task serialization: Serializes task metadata (ID, timestamp, status, rules, QoS level) into a JSON object.

[0083] Task storage: Redis write: Push the task JSON into a Redis List structure (such as task_queue:algorithm_service) using the RPUSH command.

[0084] MySQL persistence: Execute `INSERT INTO tasks (task_id, status, created_at, qos_level) VALUES (?, 'PENDING', NOW(), ?)` via JDBC.

[0085] 1.3 Output: The task ID is returned to the user, and the task metadata is stored in Redis and MySQL.

[0086] II. Task Monitoring and Analysis The data processing flow is as follows: 2.1 Input data: Redis task queue length, Kubernetes cluster resource metrics (CPU / GPU utilization), and historical load data.

[0087] 2.2 Processing steps: Task queue monitoring: Predicted load calculation: The load amount for the next 10 minutes (i.e., the next monitoring period) is predicted using a pre-trained prediction model to obtain the predicted load; Actual load calculation: Before the end of the current period, the number of unprocessed tasks (i.e., the current number of tasks) is obtained through LLEN task_queue:algorithm_service. The total Pod processing capacity is obtained through Kubernetes API calls. The actual load for the current monitoring period is calculated using the following formula: Predicted load = (current number of tasks / total Pod processing capacity) * 100, where Pod processing capacity is based on the historical average task processing rate.

[0088] III. Dynamic expansion and contraction 3.1. Update the prediction confidence of the previous monitoring period based on the measured load of the current monitoring period to obtain the prediction confidence of the current monitoring period: Initial confidence =0.5, calculate the prediction error at the current time point after the actual load arrives. =|Measured Load - Predicted Load|, maintaining an error sliding window W. The error sliding window W records N error data points. Latest... Upon arrival, update the historical error sequence. Based on the updated error sequence, the following statistical characteristics are calculated: average error =

[0089] Standard deviation

[0090] Extremely poor

[0091] Constructing a multidimensional confidence function:

[0092] in, , , The 95th percentile is the value within the error sliding window.

[0093] 3.2. Based on the prediction confidence level of the current monitoring period, the measured load and the predicted load of the current monitoring period are weighted and summed to obtain the comprehensive load of the current monitoring period:

[0094] in, The overall load for the current monitoring period, The predicted load for the previous monitoring period, The measured load for the current monitoring period. This represents the prediction confidence level for the current monitoring period.

[0095] 3.3. Determine the adjustment coefficient based on the difference between the comprehensive load and the preset standard load:

[0096] in, To adjust the coefficient, The overall load for the current monitoring period, For the preset standard load, This is the preset adjustment range.

[0097] 3.4. Based on the current total number of algorithm Pods and the aforementioned adjustment coefficient, obtain the adjusted number of algorithm Pods:

[0098] in, The adjusted number of algorithm Pods, This represents the total number of Pods in the current algorithm.

[0099] 3.5. Adjust the duration of the next monitoring period based on the prediction confidence level of the current monitoring period: ,

[0100] in, The duration of the next monitoring cycle, This is the minimum value of the preset monitoring cycle. This is the maximum value of the preset monitoring cycle. The prediction confidence level for the current monitoring period. This is the preset minimum prediction confidence level. This is the preset maximum prediction confidence level. It is an exponential coefficient used to adjust the shape of the mapping curve between confidence level and period duration.

[0101] IV. Resource Pool Allocation Basic resource pool: Fixed allocation to core services (core algorithm modules).

[0102] Shared resource pool: allocated on demand by adjusting the number of replicas of the corresponding algorithm service Deployment.

[0103] Exclusive resource pool: Exclusively owns a portion of the GPU card, and binds to a specific point through nodeSelector.

[0104] Kubernetes API calls: After calculating the number of Pods that need to be scaled up or down, the Deployment configuration in the Kubernetes cluster needs to be modified to trigger the change. Simultaneously, the Kubernetes scheduler will automatically schedule the Pods based on cluster resource availability.

[0105] Update Deployment: Use the Kubernetes Java client to perform a patch operation to update the .spec.replicas field.

[0106] Pod scheduling: The Kubernetes scheduler selects the optimal node based on the node's GPU resources (reported via nvidia.com / gpu by the Device Plugin).

[0107] Output: The new Pod starts and is added to the task processing queue.

[0108] V. GPU Resource Pool Management Data processing flow: Input data: Task QoS level, GPU node resource information.

[0109] Processing steps: Resource pool division: Basic resource pool: Nodes are marked with nodeLabel (e.g., gpu-pool: base) and only allocated to high-priority tasks.

[0110] Shared resource pool: Supports dynamic GPU allocation (such as CUDA_VISIBLE_DEVICES subdivisions of NVIDIA DCGM or mGPU) via Device Plugin.

[0111] Dedicated resource pool: By binding a specific node through nodeSelector, the corresponding algorithm Pod has exclusive access to the entire GPU card.

[0112] Resource scheduling strategy: Shared resource allocation: Use the `--gpus` parameter of `nvidia-docker` to specify the ratio of video memory to computing power (and how to utilize this data later) (e.g., `--gpus '"device=0, compute_cap=3.5, memory=4096"').

[0113] Output: The task is allocated to the corresponding GPU resources in the resource pool according to the QoS level.

[0114] VI. Task Lifecycle Management Data processing flow: Input data: Task ID, processing status, timeout.

[0115] Processing steps: Task assignment and locking mechanism: Redis lock: The worker obtains the task through the BLPOP atomic operation and sets SET task_lock:ID 1 EX120 NX.

[0116] Status Update: Change the task status from PENDING to PROCESSING and record the start time.

[0117] Task timeout handling: Scheduled scanning: Use Redis's SCAN command to check for expired locks (EXPIRED keys) and reset the task status to PENDING.

[0118] Result persistence: MySQL insert: Execute INSERT INTO results (task_id, result, duration, algorithm_version) VALUES (?, ?, ?, ?).

[0119] Security and fault tolerance mechanisms: Task retry: Automatically retry timed-out tasks 3 times, and mark them as ERROR if they fail.

[0120] Database transactions: Task status updates and result storage use BEGIN...COMMIT transactions to ensure consistency.

[0121] Output: The task status is updated to COMPLETED, and the result is stored in MySQL.

[0122] New tasks can be added to the system through the front-end interface, and results can be obtained after processing. The management platform allows users to view the current task progress and the status of processed tasks.

[0123] Resource configuration: Administrators can set up different types of GPU resource pools, specifying which algorithm services should exclusively use specific GPU resources, and which can be allocated on demand from the shared resource pool.

[0124] Real-time monitoring: The front end provides an intuitive dashboard that displays key metrics such as the current number of tasks, the number of completed tasks, and the utilization rate of K8s cluster resources.

[0125] Logs and Alarms: The system supports log recording, and can trigger alarms to notify administrators to handle abnormal situations in a timely manner.

[0126] This embodiment can significantly improve GPU resource utilization, enhance the flexibility and efficiency of applications deployed on Kubernetes, reduce operational costs while improving user experience, and greatly improve the efficiency of article review. Specific advantages and features are as follows: 1. High resource utilization efficiency Dynamic adjustment: The system can automatically adjust the number of algorithm module Pods based on real-time task volume and resource usage to ensure that expensive computing resources such as GPUs are fully utilized.

[0127] Pooled management: By dividing GPU resources into basic resource pools, shared resource pools, and dedicated resource pools, resource allocation can be controlled more precisely to meet the needs of different types of tasks.

[0128] 2. High flexibility and adaptability Rapid response: The system can quickly react to changes in workload, instantly increasing or decreasing the number of Pods to maintain optimal service performance and adapt to request fluctuations over different time periods.

[0129] Customizable strategies: Allows users to customize resource allocation strategies according to their own business needs, flexibly responding to changing work environments.

[0130] 3. High degree of automation Intelligent scheduling: It has a built-in intelligent decision-making mechanism based on task pool status and algorithm module performance, which reduces the need for manual intervention and improves operation and maintenance efficiency.

[0131] Continuous optimization: A process of continuous monitoring, analysis, and adjustment to ensure that the system is always in optimal configuration.

[0132] 4. Scalability Microservice architecture: Utilizing containerized deployment and a microservice architecture, it is easy to scale horizontally and supports larger-scale workload processing capabilities.

[0133] Modular design: Each functional component is relatively independent, which facilitates maintenance and upgrades without affecting the normal operation of other parts.

[0134] 5. User-friendly interface Intuitive Display: The user interface developed with Vue.js provides a clear and intuitive display of information such as task statistics and resource usage, helping administrators to better understand and manage the entire system.

[0135] Easy to use: The simple operation process and user-friendly interactive design make it easy for even non-technical people to get started.

[0136] 6. Robust monitoring and feedback Real-time monitoring: The system provides comprehensive real-time monitoring functions, including key indicators such as task progress and resource utilization, to ensure that problems can be detected and resolved in a timely manner.

[0137] Alarm mechanism: When an abnormal situation is detected, the system will trigger an alarm to notify the administrator to prevent potential risks from escalating.

[0138] 7. Data persistence and security Stable storage: Using a MySQL database for persistent data storage ensures the security and reliability of task results.

[0139] Access control: Combined with the security mechanisms of the Kubernetes cluster, strict access control is implemented to protect sensitive data from unauthorized access.

[0140] 8. Technological advancement Modern technology stack: It adopts the currently popular front-end and back-end separation architecture (Vue.js + Java) and the container orchestration platform Kubernetes, reflecting its forward-looking technology.

[0141] Integrated Innovation: It integrates a variety of advanced technologies and services, such as Redis caching and unified image repository, to build an efficient and stable ecosystem.

[0142] In summary, this solution not only solves the problems of resource waste, slow response speed, and complex technical operation and maintenance in existing technologies, but also provides users with an intelligent, flexible, and easy-to-use solution, which greatly improves the efficiency of resource management and task processing.

[0143] Example 3: Another embodiment of this application relates to a device for adjusting the number of algorithm containers. The implementation details of this device are described below. The following details are for ease of understanding and are not essential for implementing this solution. A schematic diagram of the device for adjusting the number of algorithm containers in this embodiment can be seen as follows: Figure 3 As shown, it includes a load acquisition module 310, a confidence calculation module 320, a comprehensive load calculation module 330, an adjustment coefficient determination module 340, and an algorithm Pod number adjustment module 350.

[0144] The load acquisition module 310 is used to acquire the measured load of the current monitoring period and the predicted load for the current monitoring period predicted in the previous monitoring period. The confidence calculation module 320 is used to update the predicted confidence of the previous monitoring period based on the measured load of the current monitoring period, so as to obtain the predicted confidence of the current monitoring period. The comprehensive load calculation module 330 is used to perform a weighted summation of the measured load and the predicted load of the current monitoring period based on the prediction confidence of the current monitoring period, so as to obtain the comprehensive load of the current monitoring period. The adjustment coefficient determination module 340 is used to determine the adjustment coefficient based on the difference between the comprehensive load and the preset standard load; The algorithm Pod number adjustment module 350 is used to obtain the adjusted number of algorithm Pods based on the current total number of algorithm Pods and the adjustment coefficient.

[0145] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.

[0146] In some embodiments, the confidence calculation module includes: The prediction error calculation unit is used to calculate the prediction error of the current monitoring period after the actual load arrives in the current monitoring period, and update the error record in the error sliding window, wherein the error sliding window includes error data from multiple historical monitoring periods. An error statistical feature calculation unit is used to calculate the statistical features of the error within the updated error sliding window, wherein the statistical features include at least one of mean error, standard deviation, and range; The feature threshold calculation unit is used to take the quantile of the historical values ​​of the statistical features under each historical monitoring period within the updated error sliding window at a preset ratio to generate the feature threshold of the corresponding statistical features. The confidence correction unit is used to calculate a correction coefficient based on the statistical characteristics and the characteristic threshold, and to adjust the prediction confidence of the previous monitoring period using the correction coefficient to obtain the prediction confidence of the current monitoring period.

[0147] In some embodiments, when there are multiple statistical features, the confidence correction unit includes: The sub-correction coefficient calculation unit is used to calculate the sub-correction coefficient of each statistical feature based on the ratio of the statistical feature to the feature threshold corresponding to the statistical feature. The correction coefficient calculation unit is used to obtain the correction coefficient based on the weighted summation of multiple sub-correction coefficients.

[0148] In some embodiments, it also includes: The monitoring cycle adjustment module is used to adjust the duration of the next monitoring cycle based on the prediction confidence level of the current monitoring cycle.

[0149] In some embodiments, the monitoring cycle adjustment module is used to determine the duration of the next monitoring cycle based on the predicted confidence of the current monitoring cycle, according to a preset confidence-cycle duration mapping relationship, wherein the confidence and cycle duration are positively correlated.

[0150] Example 4: Another embodiment of this application relates to an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the algorithm container number adjustment method in the above embodiments.

[0151] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.

[0152] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.

[0153] Example 5: Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.

[0154] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0155] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.

Claims

1. A method for adjusting the number of containers in an algorithm, characterized in that, include: Obtain the measured load for the current monitoring period and the predicted load for the current monitoring period from the previous monitoring period. The prediction confidence level of the current monitoring period is obtained by updating the prediction confidence level of the previous monitoring period based on the measured load of the current monitoring period. Based on the prediction confidence of the current monitoring period, the measured load and the predicted load of the current monitoring period are weighted and summed to obtain the comprehensive load of the current monitoring period. The adjustment coefficient is determined based on the difference between the comprehensive load and the preset standard load; The adjusted number of algorithm Pods is obtained based on the current total number of algorithm Pods and the adjustment coefficient.

2. The method for adjusting the number of algorithm containers according to claim 1, characterized in that, The step of updating the prediction confidence of the previous monitoring period based on the measured load of the current monitoring period includes: After the actual load of the current monitoring period arrives, the prediction error of the current monitoring period is calculated, and the error record in the error sliding window is updated. The error sliding window includes error data from multiple historical monitoring periods. Calculate the statistical characteristics of the error within the updated error sliding window, wherein the statistical characteristics include at least one of the mean error, standard deviation, and range; Take the quantile of the historical values ​​of the statistical features under each historical monitoring period within the updated error sliding window according to a preset proportion to generate the feature threshold of the corresponding statistical features; The correction coefficient is calculated based on the statistical characteristics and the characteristic threshold. The prediction confidence of the previous monitoring period is adjusted using the correction coefficient to obtain the prediction confidence of the current monitoring period.

3. The method for adjusting the number of algorithm containers according to claim 2, characterized in that, When there are multiple statistical features, the step of calculating the correction coefficient based on the statistical features and the feature threshold includes: For each statistical feature, the sub-correction coefficient of the statistical feature is calculated based on the ratio of the statistical feature to the feature threshold corresponding to the statistical feature. The correction coefficients are obtained by weighted summation of the multiple sub-correction coefficients.

4. The method for adjusting the number of algorithm containers according to claim 3, characterized in that, The prediction confidence level for the current monitoring period is calculated using the following formula: = in, The prediction confidence level for the current monitoring period. The prediction confidence level for the previous monitoring period. For statistical characteristics, , and These are the mean error, standard deviation, and range, respectively. The weights of the sub-correction coefficients for statistical feature i. This is the feature threshold.

5. The method for adjusting the number of algorithm containers according to claim 1, characterized in that, The overall load for the current monitoring period is calculated using the following formula: in, The overall load for the current monitoring period, The predicted load for the previous monitoring period, The measured load for the current monitoring period. This represents the prediction confidence level for the current monitoring period.

6. The method for adjusting the number of algorithm containers according to claim 1, characterized in that, Calculate the adjustment factor using the following formula: in, To adjust the coefficient, The overall load for the current monitoring period, For the preset standard load, This is the preset adjustment range.

7. The method for adjusting the number of algorithm containers according to any one of claims 1-6, characterized in that, Also includes: The duration of the next monitoring cycle is adjusted based on the prediction confidence level of the current monitoring cycle.

8. The method for adjusting the number of algorithm containers according to claim 7, characterized in that, The adjustment of the duration of the next monitoring cycle based on the prediction confidence of the current monitoring cycle includes: Based on a preset mapping relationship between confidence level and cycle duration, the duration of the next monitoring cycle is determined according to the predicted confidence level of the current monitoring cycle, where there is a positive correlation between confidence level and cycle duration.

9. The method for adjusting the number of algorithm containers according to claim 8, characterized in that, The duration of the next monitoring cycle is calculated using the following formula: , in, The duration of the next monitoring cycle, This is the minimum value of the preset monitoring cycle. This is the maximum value of the preset monitoring cycle. The prediction confidence level for the current monitoring period. This is the preset minimum prediction confidence level. This is the preset maximum prediction confidence level. It is an exponential coefficient used to adjust the shape of the mapping curve between confidence level and period duration.

10. A device for adjusting the number of algorithm containers, characterized in that, include: The load acquisition module is used to acquire the measured load of the current monitoring period and the predicted load for the current monitoring period predicted in the previous monitoring period. The confidence calculation module is used to update the predicted confidence of the previous monitoring period based on the measured load of the current monitoring period, so as to obtain the predicted confidence of the current monitoring period. The comprehensive load calculation module is used to perform a weighted summation of the measured load and the predicted load of the current monitoring period based on the prediction confidence level of the current monitoring period, so as to obtain the comprehensive load of the current monitoring period. The adjustment coefficient determination module is used to determine the adjustment coefficient based on the difference between the comprehensive load and the preset standard load; The algorithm Pod number adjustment module is used to obtain the adjusted number of algorithm Pods based on the current total number of algorithm Pods and the adjustment coefficient.

11. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the algorithm container number adjustment method as described in any one of claims 1 to 9.

12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for adjusting the number of algorithm containers as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Container resource control method and device and computer storage medium

    CN113886010A

  • Predictive elastic scaling method and system considering multi-dimensional load characteristics

    CN118260089A

  • Cluster dynamic scaling method and system based on periodic data prediction

    CN118467098A

  • Allocating background workflows in a data storage system using autocorrelation

    US8140775B1