Big data resource capacity management method and device
Through the LSTM model and business priority strategy, the problems of resource usage fluctuation prediction deviation and untimely allocation were solved, the refinement and stability of resource management were achieved, and costs were reduced.
Patent Information
- Application Number
- CN202510828844.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies are unable to cope with the nonlinear, periodic and sudden fluctuations in resource utilization, and the prediction deviation is large. The scheduling system lacks multi-dimensional indicator evaluation and autonomous learning capabilities, resulting in untimely or excessive resource allocation, affecting the stable operation of core businesses.
The LSTM model is combined with multi-dimensional resource data for trend prediction, and the changing patterns are captured through sliding windows. A business priority model is built, the resource allocation strategy is dynamically adjusted, and a prediction error feedback mechanism is set up to achieve closed-loop optimization.
It improves the real-time and accuracy of resource management, realizes refined regulation based on demand and optimization, improves the operational stability of core businesses, and reduces resource waste and operation and maintenance costs.
Smart Images

Figure CN120743508A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and more specifically, to a method and device for managing big data resource capacity. Background Art
[0002] With the rapid development of technologies such as cloud computing, big data, and artificial intelligence, data centers and enterprise IT systems face the challenge of managing massive, multi-dimensional, and time-varying resources. In this context, accurately predicting and efficiently scheduling key resources such as computing, storage, and network bandwidth has become a critical issue for ensuring stable system performance, reducing resource waste, and improving service quality. Existing big data resource capacity management methods primarily rely on fixed threshold rules or static configuration policies based on historical averages. These approaches suffer from the following major shortcomings: Most systems use simple sliding averages or linear extrapolation to predict resource usage trends, which are difficult to address with nonlinear, cyclical, and sudden fluctuations in resource usage, resulting in large prediction errors and frequent policy failures. Current scheduling systems generally lack comprehensive assessment of multi-dimensional indicators such as service level, access frequency, and SLA requirements. Instead, they statically allocate resources based solely on total resources, failing to achieve refined, on-demand, and optimized resource control. Most resource management systems still rely on manually set thresholds and trigger conditions and lack autonomous learning and dynamic adjustment capabilities, leading to frequent issues such as delayed response, untimely scheduling, and overscheduling. Existing systems generally lack deep integration with factors such as the real-time operating status, priority, and service-level agreements (SLAs) of the business, and are unable to dynamically adjust resource supply strategies based on business sensitivity or criticality, affecting the stable operation of core businesses.
[0003] Although some studies have attempted to introduce machine learning-based prediction models for resource trend judgment, there are often problems such as shallow models, low integration, and isolated training mechanisms, making it difficult to form a closed-loop linkage with actual resource management processes. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and device for big data resource capacity management to solve the problems raised in the above-mentioned background technology: Most systems use simple sliding average or linear extrapolation to predict resource usage trends, which is difficult to cope with nonlinear, periodic and sudden fluctuations in resource usage, resulting in large prediction deviations and frequent policy failures. Current scheduling systems generally lack comprehensive evaluation of multi-dimensional indicators such as business level, access frequency, SLA requirements, etc., and only perform static allocation based on the total amount of resources, which cannot achieve on-demand and optimized fine-grained regulation of resources. Most resource management systems still rely on manual setting of thresholds and trigger conditions, and lack the ability to learn and adjust dynamically, resulting in frequent problems such as delayed response, untimely scheduling or over-scheduling. Existing systems generally lack deep integration with factors such as the real-time operating status, priority, service level agreement (SLA) of the business, and cannot dynamically adjust resource supply strategies based on business sensitivity or criticality, affecting the stable operation of core businesses.
[0005] Technical solution: A method for managing big data resource capacity, comprising the following steps:
[0006] S1. Periodically collect system resource data, data behavior characteristics, and business operation indicators according to a set collection period T. The resource data includes CPU usage, memory usage, disk I / O, and network bandwidth usage. The data behavior characteristics include data access frequency, data growth rate, and activity. The business indicators include request frequency, number of concurrent connections, and SLA level.
[0007] S2. After uniform normalization of the collected data, input it into the cache pool as model training data. The normalization formula is:
[0008]
[0009] Among them, x is the original index value, x′ is the normalized result, and x min ,x max is the historical extreme value of the indicator;
[0010] S3. Build and train an LSTM model, using historical multi-dimensional resource data within a fixed window size W as input features to predict resource usage trends for various types of resources over the next M time periods. The prediction results serve as the basis for scheduling strategies.
[0011] S4. Based on the access frequency, SLA level, and response delay of the business, a multi-factor weighted function is used to calculate the business priority and classify the business into three categories: core business, routine business, and low-sensitivity business;
[0012] S5. Determine whether to trigger resource expansion based on the difference between the predicted resource usage and the current remaining resource amount. When resource usage falls below the set threshold for multiple consecutive cycles, execute a resource recovery strategy. The scheduling process prioritizes resources for core businesses based on business priority.
[0013] S6. By comparing the error between the predicted value and the actual value, the error average and sliding cumulative value are calculated; when the error exceeds the threshold, the model parameters and scheduling strategy sensitivity are adjusted through the prediction error feedback mechanism to achieve closed-loop optimization of resource capacity management.
[0014] Preferably, the acquisition period T in S1 is set to an adjustable value between 1 second and 300 seconds, preferably 60 seconds; the sampling window size W in S3 is given by the formula Calculated.
[0015] Preferably, in the resource prediction step of the LSTM model in S3: the input matrix is constructed using historical data of nearly W time periods Where N is the resource indicator dimension; the LSTM model structure uses a two-layer stack, each layer contains 64 hidden units, and the activation function is tanh; the model output is the predicted value of each resource in M future time periods, and the prediction loss function is the mean square error (MSE): If the average error If the value exceeds the set threshold ε (such as 5%), the model retraining mechanism is triggered.
[0016] Preferably, in S4: a weighted scoring function is used to score the priority of the service, and the scoring formula is:
[0017]
[0018] Among them, F r is the access frequency of business units; SLA is the service level indicator, P1 is 3, P2 is 2, and P3 is 1; D l is the maximum response delay requirement (ms); w1 = 0.4, w2 = 0.4, w3 = 0.2.
[0019] Preferably, the services are divided into the following categories according to the score S: core services: S≥2.5; conventional services: 1.5≤S<2.5; low-sensitivity services: S<1.5;
[0020] The classification results are used to determine resource scheduling priorities in resource scheduling. The scheduling rules are as follows: core businesses enjoy exclusive resource quotas, priority expansion, and prohibited from shrinking; conventional businesses obtain flexible quotas when resources are sufficient; low-sensitivity businesses are scheduled on demand and can be given priority in resource recovery.
[0021] Preferably, the S5 resource expansion operation is triggered when the following conditions are met: Calculate the predicted difference Δ t =P t -R t , where P t To predict resource requirements, R t is the currently available resources; when Δ t >Δ th (The recommended threshold is 10% of the system capacity), capacity expansion starts; the expansion resource calculation formula is:
[0022] Q add =γ·Δ t ;
[0023] Among them, γ is the safety redundancy coefficient, ranging from 1.2 to 2.0, which is used to reserve capacity in advance.
[0024] Preferably, the resource recycling strategy in S5 is triggered when the following conditions are met: within D consecutive periods (such as 30 minutes), if the average resource usage rate meets the following conditions:
[0025]
[0026] Among them, U i is the resource utilization rate of the i-th cycle, θ is the idle threshold (30%);
[0027] When the conditions are met, resources are released starting with low-sensitivity services according to priority until resource utilization returns to a safe range.
[0028] Preferably, the prediction error feedback mechanism in S6 includes: accumulating the prediction error using an exponentially weighted moving average (EWMA):
[0029] E sum (t) = α·E(t) + (1-α)·E sum (t-1);
[0030] Among them, α is the smoothing coefficient (taken as 0.3);
[0031] The scheduling sensitivity is dynamically adjusted based on the error accumulation value, and the following threshold adjustment formula is executed:
[0032]
[0033] θ new =θ old ·(1-μ·E sum (t));
[0034] Among them, λ and μ are strategy adjustment factors.
[0035] A big data resource capacity management device, which specifically includes a cloud native platform with distributed task scheduling and load migration capabilities, and can realize automated resource allocation, container copy adjustment, and node task migration operations based on a big data resource capacity management method.
[0036] Compared with the prior art, the advantages of the present invention are:
[0037] (1) This paper introduces a deep recurrent neural network (LSTM) to learn and model the trends of historical multi-dimensional resource indicators, and captures the periodic and sudden change patterns through a sliding window mechanism. Compared with traditional methods such as linear regression and sliding average, it has better prediction accuracy and robustness in multiple scenarios.
[0038] (1) The present invention combines the prediction results with the set resource threshold difference to dynamically trigger the expansion and recovery strategy, avoiding the lag and misjudgment caused by manually configuring scheduling rules, and significantly improving the real-time and accuracy of system resource management.
[0039] (3) The present invention introduces indicators such as service access frequency, SLA level, and response delay, adopts a multi-factor weighted model to calculate service priority, and divides services into three levels: core, conventional, and low-sensitivity. This implements a refined resource allocation strategy that prioritizes key services, thereby improving the operational stability of core services.
[0040] (4) The present invention sets up a prediction error monitoring and model adaptive retraining mechanism, which can automatically trigger retraining when the model accuracy decreases, effectively overcoming the impact of environmental changes on model performance, and maintaining the long-term effectiveness of the model and the adaptability of the scheduling strategy.
[0041] (5) The model training and reasoning process of the present invention is deployed in the form of microservices, which can be seamlessly integrated with existing scheduling frameworks (such as Kubernetes, Hadoop YARN, etc.), support horizontal expansion of various resource types, and have good platform versatility and industrial implementation potential.
[0042] (6) The present invention effectively avoids excessive resource reservation or temporary preemption by accurately predicting resource capacity in advance and scheduling it on demand, thereby optimizing resource utilization efficiency and reducing infrastructure operation and maintenance and energy consumption costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a schematic diagram of the overall process of the big data equipment digital inspection and evaluation method of the present invention. DETAILED DESCRIPTION
[0044] For examples, see Figure 1 , a big data resource capacity management method, comprising the following steps:
[0045] S1. Periodically collect system resource data, data behavior characteristics, and business operation indicators according to the set collection period T. Resource data includes CPU usage, memory usage, disk I / O, and network bandwidth usage. Data behavior characteristics include data access frequency, data growth rate, and activity. Business indicators include request frequency, number of concurrent connections, and SLA level. The collection period T is set to an adjustable value between 1 and 300 seconds, preferably 60 seconds.
[0046] S2. After uniform normalization of the collected data, input it into the cache pool as model training data. The normalization formula is:
[0047]
[0048] Among them, x is the original index value, x′ is the normalized result, and x min ,x max is the historical extreme value of the indicator;
[0049] S3. Build and train an LSTM model, using historical multi-dimensional resource data within a fixed window size W as input features to predict the usage trends of various resources in the next M time periods. The prediction results are used as the basis for the scheduling strategy. The sampling window size W is given by the formula Calculated; In the resource prediction step of the LSTM model: use historical data of nearly W time periods to build the input matrix Where N is the resource indicator dimension; the LSTM model structure uses a two-layer stack, each layer contains 64 hidden units, and the activation function is tanh; the model output is the predicted value of each resource in M future time periods, and the prediction loss function is the mean square error (MSE): If the average error If the threshold ε is exceeded (e.g., 5%), the model retraining mechanism is triggered. The LSTM model training steps are as follows:
[0050] Step 1: Input data construction and processing
[0051] 1. Define resource indicator dimension N:
[0052] This includes N indicators, including CPU usage, memory usage, disk read / write I / O, and network bandwidth usage. In this example, N=4.
[0053] 2. Sampling period T and window size W:
[0054] The default sampling period T is 60 seconds;
[0055] Use the following formula to determine the window size W:
[0056]
[0057] Indicates the number of sampling points in the past 24 hours. If T = 60, then W = 1440.
[0058] 3. Input matrix construction:
[0059] For each prediction, construct an input matrix Each row represents the N indicator values at a time point;
[0060] Use MinMax normalization or Zscore normalization to process all historical data to ensure that the model input range is stable.
[0061] Step 2: LSTM model structure design
[0062] 1. Model hierarchy:
[0063] Input layer: shape is (W, N);
[0064] LSTM layer 1: 64 units, tanh activation, returns the entire sequence;
[0065] LSTM layer 2: 64 units, tanh activation, returns the last state;
[0066] Dense fully connected output layer: The output dimension is M×N, that is, the predicted value of each resource in the next M time periods.
[0067] 2. Loss function definition:
[0068] Use mean squared error (MSE) as the loss function:
[0069]
[0070] 3. Optimizer and training parameters:
[0071] Optimizer: Adam;
[0072] Learning rate: default 1×10 -3 , support automatic adjustment;
[0073] Batch size: 64;
[0074] Number of training rounds: 50-200 rounds;
[0075] The EarlyStopping mechanism automatically terminates training based on the error fluctuations in the validation set.
[0076] Step 3: Training and prediction process
[0077] 1. Training data set construction:
[0078] The sliding window method collects multiple sequence segments of length W;
[0079] Each segment corresponds to a target value interval of length M;
[0080] The ratio of training set, validation set, and test set is 7:2:1.
[0081] 2. Model training process:
[0082] It is started every time the system is in the early stage of operation or when the error exceeds the threshold to trigger retraining;
[0083] Daily offline training + real-time fine-tuning mechanism:
[0084] Retrain the full model in batches every day;
[0085] Online fine-tuning is performed every H hours using the most recent K samples.
[0086] Step 4: Deep integration of the model and the scheduling system
[0087] 1. Prediction-driven scheduling mechanism:
[0088] The model outputs a forecast curve for each type of resource in the next M time periods every period (e.g., every 5 minutes). The scheduling control module calculates whether the expansion or recycling threshold is reached based on the forecast results.
[0089] If the predicted value continues to be higher than the resource threshold, a capacity expansion request will be sent to the resource management platform;
[0090] If the predicted value continues to be lower than the set ratio and the current load is low, resource release is triggered.
[0091] 2. Error feedback control mechanism:
[0092] Compare the deviation between actual usage and predicted values:
[0093]
[0094] Calculate the sliding error average:
[0095]
[0096] like (If set to 5%), then: start the model retraining process; simultaneously adjust the resource scheduling sensitivity (such as expansion threshold, recycling delay, etc.);
[0097] Step 5: Algorithm Deployment and Elastic Integration Recommendations: Deploy the model as a microservice (such as Flask / TorchServe), which can be called in real time by the scheduling module through the REST API. All prediction results are returned in a JSON structure, including time series prediction values for each resource type. Link with the Kubernetes HPA (Horizontal Pod Autoscaler) to dynamically adjust the number of pods or node resource specifications based on the predicted values.
[0098] S4: Based on the access frequency, SLA level, and response delay of the business, a multi-factor weighted function is used to calculate the business priority, and the business is divided into three categories: core business, routine business, and low-sensitivity business. In S4: a weighted scoring function is used to score the priority of the business. The scoring formula is:
[0099]
[0100] Among them, F r is the access frequency of business units; SLA is the service level indicator, P1 is 3, P2 is 2, and P3 is 1; D l is the maximum response delay requirement (ms); w1 = 0.4, w2 = 0.4, w3 = 0.2.
[0101] According to the score S, the business is divided into: core business: S≥2.5; conventional business: 1.5≤S<2.5; low-sensitivity business: S<1.5;
[0102] The classification results are used to determine resource scheduling priorities in resource scheduling. The scheduling rules are as follows: core businesses enjoy exclusive resource quotas, priority expansion, and prohibited from shrinking; conventional businesses obtain flexible quotas when resources are sufficient; low-sensitivity businesses are scheduled on demand and can be given priority in resource recovery.
[0103] S5. Determine whether to trigger resource expansion based on the difference between the predicted resource usage and the current remaining resource amount. When the resource usage rate is lower than the set threshold for multiple consecutive cycles, execute the resource recovery strategy. The scheduling process ensures that core businesses are allocated resources first based on the business priority results. The resource expansion operation is triggered when the following conditions are met: Calculate the predicted difference Δ t =P t -R t , where P t To predict resource requirements, R t is the currently available resources; when Δ t >Δ th (The recommended threshold is 10% of the system capacity), capacity expansion starts; the expansion resource calculation formula is:
[0104] Q add =γ·Δ t ;
[0105] Among them, γ is the safety redundancy coefficient, ranging from 1.2 to 2.0, which is used to reserve capacity in advance.
[0106] The resource recycling policy is triggered when the average resource usage meets the following conditions:
[0107]
[0108] Among them, U i is the resource utilization rate of the i-th cycle, θ is the idle threshold (30%);
[0109] When the conditions are met, resources are released starting with low-sensitivity services according to priority until resource utilization returns to a safe range.
[0110] S6. By comparing the error between the predicted value and the actual value, the error average and sliding cumulative value are calculated; when the error exceeds the threshold, the model parameters and scheduling strategy sensitivity are adjusted through the prediction error feedback mechanism to achieve closed-loop optimization of resource capacity management.
[0111] The forecast error feedback mechanism includes: accumulating forecast errors using the exponentially weighted moving average (EWMA):
[0112] E sum (t) = α·E(t) + (1-α)·E sum (t-1);
[0113] Among them, α is the smoothing coefficient (taken as 0.3);
[0114] The scheduling sensitivity is dynamically adjusted based on the error accumulation value, and the following threshold adjustment formula is executed:
[0115]
[0116] θ new =θ old ·(1-μ·E sum (t));
[0117] Among them, λ and μ are strategy adjustment factors.
[0118] A big data resource capacity management device, specifically including a cloud native platform with distributed task scheduling and load migration capabilities, can realize automated resource allocation, container copy adjustment, and node task migration operations based on a big data resource capacity management method.
[0119] The above shows and describes the basic principles, main features and advantages of the present invention; those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements all fall within the scope of the present invention to be protected; the scope of protection claimed in the present invention is defined by the attached claims and their equivalents.
Claims
1. A method for managing the capacity of big data resources, characterized in that: The steps include: S1. Periodically collect system resource data, data behavior characteristics, and business operation indicators according to a set collection period T. The resource data includes CPU usage, memory usage, disk I / O, and network bandwidth usage. The data behavior characteristics include data access frequency, data growth rate, and activity. The business indicators include request frequency, number of concurrent connections, and SLA level. S2. After uniform normalization of the collected data, input it into the cache pool as model training data. The normalization formula is: Among them, x is the original index value, x′ is the normalized result, and x min ,x max is the historical extreme value of the indicator; S3. Build and train an LSTM model, using historical multi-dimensional resource data within a fixed sampling window size W as input features, to predict resource usage trends for various types of resources over the next M time periods. The prediction results serve as the basis for scheduling strategies. S4. Based on the access frequency, SLA level, and response delay of the business, a multi-factor weighted function is used to calculate the business priority and classify the business into three categories: core business, routine business, and low-sensitivity business; S5. Determine whether to trigger resource expansion based on the difference between the predicted resource usage and the current remaining resource amount. When resource usage falls below the set threshold for multiple consecutive cycles, execute a resource recovery strategy. The scheduling process prioritizes resources for core businesses based on business priority. S6. By comparing the error between the predicted value and the actual value, the error average and sliding cumulative value are calculated; when the error exceeds the threshold, the model parameters and scheduling strategy sensitivity are adjusted through the prediction error feedback mechanism to achieve closed-loop optimization of resource capacity management.
2. A method for managing the capacity of large data resources according to claim 1, characterized in that: The acquisition period T in S1 is set to an adjustable value between 1 second and 300 seconds; the sampling window size W in S3 is determined by the formula Calculated.
3. A method for managing big data resource capacity according to claim 1, characterized in that: In the resource prediction step of the LSTM model in S3: the input matrix is constructed using historical data of nearly W time periods Where N is the resource indicator dimension; the LSTM model structure uses a two-layer stack, each layer contains 64 hidden units, and the activation function is tanh; the model output is the predicted value of each resource in M future time periods, and the prediction loss function is the mean square error (MSE): If the average error If the value exceeds the set threshold ε5%, the model retraining mechanism is triggered.
4. A method for managing the capacity of large data resources according to claim 1, characterized in that: In S4, a weighted scoring function is used to score the priority of the services. The scoring formula is: Among them, F r is the access frequency of business units; SLA is the service level indicator, P1 is 3, P2 is 2, and P3 is 1; D l is the maximum response delay requirement; w1 = 0.4, w2 = 0.4, w3 = 0.
2.
5. A method for managing the capacity of large data resources according to claim 4, characterized in that: According to the score S, the business is divided into: core business: S≥2.5; conventional business: 1.5≤S<2.5; low-sensitivity business: S<1.5; The classification results are used to determine resource scheduling priorities in resource scheduling. The scheduling rules are as follows: core businesses enjoy exclusive resource quotas, priority expansion, and prohibited from shrinking; conventional businesses obtain flexible quotas when resources are sufficient; low-sensitivity businesses are scheduled on demand and can be given priority in resource recovery.
6. A method for managing big data resource capacity according to claim 1, characterized in that: The S5 resource expansion operation is triggered when the following conditions are met: Calculate the predicted difference Δ t =P t -R t , where P t To predict resource requirements, R t is the currently available resources; when Δ t >Δ th When Δ t h represents the threshold parameter for triggering resource expansion. The expansion resource calculation formula is: Q add =c·D t ; Among them, γ is the safety redundancy coefficient, ranging from 1.2 to 2.0, which is used to reserve capacity in advance.
7. A method for managing big data resource capacity according to claim 1, characterized in that: The resource recycling strategy in S5 is triggered when the following conditions are met: within D consecutive cycles, if the average resource usage rate meets the following conditions: Among them, U i is the resource utilization rate of the i-th cycle, θ is the idle threshold; When the conditions are met, resources are released starting with low-sensitivity services according to priority until resource utilization returns to a safe range.
8. A method for managing big data resource capacity according to claim 1, characterized in that: The prediction error feedback mechanism in S6 includes: accumulating the prediction error using an exponentially weighted moving average (EWMA): E sum (t)=α·E(t)+(1-α)·E sum (t-1); Among them, α is the smoothing coefficient and is set to 0.3; The scheduling sensitivity is dynamically adjusted based on the error accumulation value, and the following threshold adjustment formula is executed: i new =θ old ·(1-μ·E sum (t)); Among them, λ and μ are strategy adjustment factors.
9. A big data resource capacity management device, characterized in that: The big data resource capacity management device specifically includes a cloud native platform with distributed task scheduling and load migration capabilities, and can realize automated resource allocation, container copy adjustment, and node task migration operations based on a big data resource capacity management method described in any one of claims 1-8.
Citation Information
Patent Citations
Resource management method and device, equipment and storage medium
CN117544635A
Dynamic setting method and device for CPU (Central Processing Unit) usage amount threshold value
CN118210627A
Method for automatically configuring resources of data center
CN119094335A