An intelligent operation and maintenance method and system for a data center based on an internet of things

By using median filtering and long short-term memory network models in the IoT control system, combined with time series decomposition and logistic regression algorithms, a resource allocation scheme is generated, which solves the problem of unreasonable resource allocation in traditional methods and achieves efficient resource utilization and system stability.

CN120929273BActive Publication Date: 2026-02-03BEIJING LONGKUN INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511445203.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-02-03
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Traditional IoT control system intelligent operation and maintenance methods struggle to accurately capture the nonlinear relationships between data and lack adaptive adjustment capabilities, leading to unreasonable resource allocation and low system operation and maintenance efficiency.

Method used

The median filtering algorithm is used to smooth the time series. Combined with the long short-term memory network model and time series decomposition tool, nonlinear influencing factors are captured. Resource allocation schemes are generated through logistic regression and genetic algorithm, and the resource ratio is dynamically adjusted to achieve the preset performance target.

Benefits of technology

It enables precise allocation of resources, improves system operation and maintenance efficiency, reduces resource waste, and enhances overall system performance and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929273B_ABST
    Figure CN120929273B_ABST
Patent Text Reader

Abstract

The application provides an Internet of Things-based data center intelligent operation and maintenance method and system. The method adopts a long short-term memory network model to capture nonlinear influence factors of the second structured resource data set to obtain a nonlinear influence factor set; determines a current prediction demand deviation value according to a prediction trend curve and a prediction error range, calculates a first resource proportion or a second resource proportion according to a current resource allocation state and the prediction demand deviation value; calculates an allocation priority score for the current first resource or the second resource, determines a first priority sequence of the current resource allocation according to the priority score, and generates a first resource allocation scheme according to the first priority sequence. Through implementation of the application, accurate allocation of system resources and efficient operation and maintenance of system resources can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent operation and maintenance of data centers, and in particular to an intelligent operation and maintenance method and system for data centers based on the Internet of Things. Background Technology

[0002] As a core pillar of the information technology industry, the operation and maintenance of data centers in industrial control systems directly determines the success or failure of enterprise digital transformation, and their stability and efficiency are crucial to enterprise operations. With the booming development and widespread penetration of IoT control systems, data centers can collect and accumulate massive amounts of operation and maintenance data in real time, laying a solid data foundation for intelligent operation and maintenance management. However, in current operation and maintenance practices, traditional operation and maintenance models driven by manual experience and simple rules still dominate. Faced with complex and ever-changing resource demands, these models reveal obvious limitations, frequently leading to resource waste or performance bottlenecks. In particular, in predicting the changing trends of industrial control system resource performance and optimizing resource allocation, traditional methods, limited by their own technical characteristics, struggle to accurately capture the nonlinear relationships between data, lack adaptive adjustment capabilities, and cannot meet the stringent requirements of modern data centers for high reliability and efficient resource utilization.

[0003] In the field of intelligent operation and maintenance of IoT control systems, numerous technical challenges arise: extracting patterns and accurately predicting resource performance trends from massive amounts of high-dimensional, dynamic historical operation and maintenance data presents significant difficulties; even if trend prediction is achieved, dynamic optimization of resource allocation remains challenging; and in the process of resource allocation optimization, balancing performance improvement with resource costs under complex constraints presents a technical obstacle that traditional algorithms struggle to overcome. These intertwined problems severely restrict the development and application of intelligent operation and maintenance technologies. Therefore, the field of intelligent operation and maintenance of IoT control systems suffers from problems of unreasonable resource allocation and low system operation and maintenance efficiency. Summary of the Invention

[0004] This invention provides an intelligent operation and maintenance method and system based on the Internet of Things, which realizes precise resource allocation and improves the operation and maintenance efficiency of the system.

[0005] An embodiment of the present invention provides an intelligent operation and maintenance method based on the Internet of Things, comprising:

[0006] Extract the indicator resource dataset from historical resource usage records, and smooth the time series of the indicator resource dataset using the median filtering algorithm to obtain the time series resource dataset;

[0007] Multi-dimensional indicators are extracted and outliers are filtered from the time-series resource dataset to obtain the first structured resource dataset;

[0008] The first structured resource dataset is extracted using a time series decomposition tool to extract trend and periodic terms, resulting in a second structured resource dataset. The second structured resource dataset contains periodic pattern feature values ​​of resource performance changes.

[0009] A long short-term memory network model is used to capture the nonlinear impact factors of the second structured resource dataset to obtain a set of nonlinear impact factors. The first nonlinear impact factor that exceeds the corresponding preset threshold in the set of nonlinear impact factors is selected to obtain a first set of nonlinear impact factors.

[0010] A future trend simulation is performed on the first set of nonlinear influencing factors to obtain the predicted trend curve of resource performance changes and the prediction error range.

[0011] The steps for calculating resource ratios are as follows: First, determine the current predicted demand deviation value based on the predicted trend curve and the predicted error range. Then, calculate either a first resource ratio or a second resource ratio based on the current resource allocation status and the predicted demand deviation value. The first resource is a scarce resource, and the second resource is a surplus resource.

[0012] Steps for generating a resource allocation scheme: Calculate an allocation priority score for the current first resource or second resource, determine a first priority sequence for the current resource allocation based on the priority score, and generate a first resource allocation scheme based on the first priority sequence;

[0013] When the system resource utilization efficiency fails to reach the preset performance target, the parameter update rate of the long short-term memory network model is adjusted, the prediction trend curve and prediction error range are redefined, and the steps of calculating the resource ratio and generating the resource allocation scheme are repeated until the system resource utilization efficiency reaches the preset performance target.

[0014] Furthermore, the step of calculating the resource ratio includes:

[0015] The bottleneck response delay data and bottleneck distribution area data of the first nonlinear influence factor are obtained from the historical data of the first nonlinear influence factor. The periodic feature values ​​of the bottleneck response delay data and the distribution density of the bottleneck distribution area data are extracted using a time series decomposition tool to obtain the periodic pattern of the bottleneck response delay and the spatial distribution characteristics of the bottleneck distribution area.

[0016] Based on the periodic pattern and spatial distribution characteristics, feature mapping is performed on the current resource allocation status dataset to obtain a first resource allocation status dataset. The predicted demand deviation of the first resource allocation status dataset is calculated to obtain the first resource ratio or the second resource ratio.

[0017] Furthermore, the step of generating a resource allocation scheme includes:

[0018] When the proportion of the first resource or the proportion of the second resource exceeds the corresponding preset threshold, a logistic regression algorithm is used to calculate the priority score of the current resource, and the first priority sequence of the current resource allocation is determined based on the priority score.

[0019] For the first priority sequence, a deviation analysis is performed based on the system load balancing index and response time threshold, and a first resource allocation scheme with high resource utilization is generated based on the results of the deviation analysis.

[0020] Furthermore, the step of using a long short-term memory network model to capture the nonlinear influence factors of the second structured resource dataset to obtain a set of nonlinear influence factors specifically involves:

[0021] A first training set is established based on the periodic pattern feature values ​​of the second structured resource dataset. The long short-term memory network model is trained using the first training set to capture the nonlinear influencing factors of resource performance changes, thereby obtaining a set of nonlinear influencing factors.

[0022] Furthermore, future trend simulations are performed on the first set of nonlinear influencing factors to obtain the predicted trend curve and prediction error range of resource performance changes, specifically:

[0023] For the first set of nonlinear influencing factors, a time series forecasting tool is used to simulate the future trends of the nonlinear influencing factors to obtain the predicted trend curve of resource performance changes, and an error estimation tool is used to perform deviation analysis on the predicted trend curve to obtain the predicted error range of resource performance changes.

[0024] Furthermore, for system resources in the first priority sequence whose allocation priority score is higher than the corresponding preset threshold, a second resource allocation scheme is generated based on bottleneck impact range data, resource load data, resource load peak fluctuation data, and performance stability conditions.

[0025] Furthermore, generating the second resource allocation scheme includes:

[0026] The resource load data, resource load peak fluctuation data, and bottleneck impact range data are obtained from the historical resource usage records. Data comparison tools are used to process and analyze the bottleneck impact range data and resource load peak fluctuation data to obtain the load distribution characteristic value of the system.

[0027] Based on the load distribution characteristic value, the system throughput data is mapped to obtain the first load distribution characteristic data. The first load distribution characteristic data that meets the preset conditions is selected as the second load distribution characteristic data. The preset conditions are that the first load distribution characteristic data is generated during the high-frequency task scheduling period and the system meets the performance stability conditions. The system throughput data is obtained from the historical resource usage records.

[0028] The first fluctuation pattern of the second load distribution feature data is extracted using a time series decomposition tool, and the second resource allocation scheme is generated based on the first fluctuation pattern.

[0029] Furthermore, bottleneck distribution area data and predicted resource demand deviation data of the second resource allocation scheme are obtained, and a genetic algorithm is used to search the current resource allocation status of the system. The current resource allocation status is the resource allocation status adjusted according to the second resource allocation scheme.

[0030] The second fluctuation pattern of the current resource allocation status is extracted using a time series decomposition tool, and a third resource allocation scheme is generated based on the second fluctuation pattern.

[0031] Furthermore, the distribution characteristics of the resource bottleneck impact range and the load balancing data corresponding to the resource bottleneck impact range are obtained from the historical resource usage records. The distribution characteristics and response time are analyzed by a data comparison tool to obtain resource deviation data.

[0032] Based on the resource deviation data and the current resource allocation status, for periods when the current resource scheduling frequency is higher than the average level, a second fluctuation pattern of the current resource allocation status is extracted using a time series decomposition tool. Based on the second fluctuation pattern, the second resource allocation scheme is adjusted to obtain a third resource allocation scheme.

[0033] Based on the above method embodiments, the present invention provides corresponding system embodiments;

[0034] An embodiment of the present invention provides an intelligent operation and maintenance system for data centers based on the Internet of Things, including: a structured resource data module, a trend prediction module, a resource ratio calculation module, a resource allocation scheme generation module, and a module for achieving preset performance targets;

[0035] The structured resource data module is used to extract indicator resource datasets from historical resource usage records, smooth the time series of the indicator resource datasets using a median filtering algorithm to obtain a time series resource dataset; extract multi-dimensional indicators and filter outliers from the time series resource datasets to obtain a first structured resource dataset; use a time series decomposition tool to extract trend and periodic items from the first structured resource dataset to obtain a second structured resource dataset; the second structured resource dataset contains periodic pattern feature values ​​of resource performance changes;

[0036] The prediction trend module is used to capture the nonlinear influencing factors of the second structured resource dataset using a long short-term memory network model to obtain a set of nonlinear influencing factors, and to filter the first nonlinear influencing factors in the set of nonlinear influencing factors that exceed the corresponding preset threshold to obtain a first set of nonlinear influencing factors; and to perform future trend simulation on the first set of nonlinear influencing factors to obtain the prediction trend curve and prediction error range of resource performance changes.

[0037] The resource ratio calculation module is used to determine the current predicted demand deviation value based on the predicted trend curve and the predicted error range, and to calculate a first resource ratio or a second resource ratio based on the current resource allocation status and the predicted demand deviation value; the first resource is a scarce resource, and the second resource is a surplus resource;

[0038] The resource allocation scheme generation module is used to calculate an allocation priority score for the current first resource or second resource, determine a first priority sequence for the current resource allocation based on the priority score, and generate a first resource allocation scheme based on the first priority sequence.

[0039] The module for achieving the preset performance target is used to adjust the parameter update rate of the long short-term memory network model and redetermine the prediction trend curve and prediction error range when the system resource utilization efficiency fails to reach the preset performance target. The module for calculating the resource ratio and the module for generating the resource allocation scheme are then executed repeatedly until the system resource utilization efficiency reaches the preset performance target.

[0040] The embodiments of the present invention have the following beneficial effects:

[0041] This invention discloses a prediction-based dynamic resource adjustment method and system. By constructing a time-series feature and a long short-term memory (LSTM) network model, it predicts resource performance change trends and error ranges, analyzes bottleneck response delays and distribution areas, and calculates resource allocation priority scores. When the score exceeds a threshold, a resource adjustment process is initiated, combining bottleneck impact and load peaks to generate a preliminary resource allocation scheme. A genetic algorithm is used to optimize the scheme, generating dynamic resource allocation instructions and changing the system resource status in real time. When the preset performance target is not met, nonlinear influencing factors are re-analyzed, LSM network model parameters are adjusted, the predicted trend is updated, the score is recalculated, and a secondary resource allocation scheme is generated. The final resource adjustment instruction updates the system resource status, achieving dynamic optimization of resource utilization efficiency and ensuring performance stability. This invention effectively improves the accuracy and flexibility of resource allocation, reduces resource waste, enhances overall system performance, achieves precise resource allocation, and improves system operation and maintenance efficiency. Attached Figure Description

[0042] Figure 1 This is a flowchart illustrating an IoT-based intelligent operation and maintenance method according to an embodiment of the present invention.

[0043] Figure 2 This is a schematic diagram of the structure of an IoT-based intelligent operation and maintenance system provided in an embodiment of the present invention. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0045] See Figure 1 This is a flowchart illustrating an intelligent operation and maintenance method based on the Internet of Things (IoT) according to an embodiment of the present invention, including:

[0046] Step S01: Extract the indicator resource dataset from the historical resource usage records, and smooth the time series of the indicator resource dataset using the median filtering algorithm to obtain the time series resource dataset.

[0047] In a preferred embodiment, step S01 includes the following steps:

[0048] Step S011: Obtain historical resource usage records from the operation and maintenance logs, and use regular expressions to parse the timestamp and resource usage fields in the historical resource usage records to generate the first time series dataset.

[0049] Step S012: Calculate various indicator data of the resource based on the first time series dataset, and save the indicator data to the indicator resource dataset after filtering.

[0050] Specifically, the resource usage metrics in the first time-series dataset are calculated, and the usage metrics exceeding the first preset threshold are selected as the resource load peaks. The time-series volatility of the load peaks is then calculated using a sliding window. Specifically, when CPU utilization exceeds 80%, this data is marked as the CPU load peak. The daily CPU utilization volatility for each server is calculated using the standard deviation formula. When the daily CPU utilization volatility exceeds 10%, the server is marked as a high-volatility device. The load peaks and volatility corresponding to 100 high-risk servers are selected, and the load peaks and volatility are associated and saved to the first metric dataset.

[0051] Step S013: Apply median filtering algorithm to the index resource dataset to smooth the time series, reduce the proportion of data noise, and obtain the time series resource dataset.

[0052] In a preferred embodiment, the operation and maintenance logs are extracted in a structured manner using a log parsing tool. Specifically, the operation and maintenance logs contain resource usage data such as CPU, memory, and network bandwidth. Regular expressions are used to match the timestamps and resource indicator fields in the operation and maintenance logs to convert unstructured logs into structured data, extracting hourly resource usage records. For example, a server's CPU utilization rate was 75.3% at 10:00 AM on October 1, 2023.

[0053] In a preferred embodiment, the time series of the indicator resource dataset is smoothed using a median filter algorithm to obtain a time series resource dataset. The window size of the median filter algorithm is 5 hours. Specifically, if the CPU utilization of the server is abnormally high at 95.1% during a certain period, it is adjusted to 82.4% after a moving average, effectively reducing the impact of sudden noise and lowering the noise ratio from the initial 8.5% to 3.2%.

[0054] In a preferred embodiment, the data in the first structured resource dataset is stored in CSV format in a preset order, containing approximately 5 million records. The preset order includes fields such as server ID, timestamp, and resource type. A verification algorithm such as MD5 is used to verify that the structured resource dataset has not been tampered with, ensuring data integrity. The verification algorithm is preferably the MD5 verification algorithm. Specifically, the data is verified to have not been tampered with based on a file checksum such as a1b2c3d4e5.

[0055] Step S02: Extract multi-dimensional indicators and filter outliers from the time series resource dataset to obtain the first structured resource dataset.

[0056] In a preferred embodiment, the second structured resource dataset is correlated with operation and maintenance alarm records to analyze whether high load peaks have triggered alarms, thereby verifying the accuracy of the screening results. Specifically, if the server triggers three alarms during the peak period, the accuracy of the screening results can be verified, thus providing a reliable basis for subsequent resource optimization.

[0057] Step S03: Use a time series decomposition tool to extract trend and periodic terms from the first structured resource dataset to obtain a second structured resource dataset; the second structured resource dataset contains periodic pattern feature values ​​of resource performance changes.

[0058] In a preferred embodiment, in a server resource management scenario, the structured resource dataset records the CPU utilization of a server cluster over a one-year period, with data collected hourly. A time series decomposition tool, such as STL, is used to split the data into trend terms, periodic terms, and residual terms. Preferably, the time series decomposition tool is STL. The trend term reflects the long-term trend of CPU utilization, such as a slow increase with business growth; the periodic term captures recurring patterns, such as a peak utilization rate of 70% from 8 AM to 8 PM daily, while the low utilization rate is only 20% in the early morning. Through decomposition, the periodic pattern characteristic value is obtained as a daily 24-hour cycle or a weekly 7-day cycle.

[0059] Step S4: Use a long short-term memory network model to capture the nonlinear impact factors of the second structured resource dataset to obtain a set of nonlinear impact factors. Filter the first nonlinear impact factors in the set of nonlinear impact factors that exceed the corresponding preset threshold to obtain a first set of nonlinear impact factors.

[0060] Step S05: Perform future trend simulation on the first set of nonlinear influencing factors to obtain the predicted trend curve of resource performance changes and the prediction error range.

[0061] In a preferred embodiment, for the first set of nonlinear influencing factors, a time series forecasting tool is used to simulate the future trends of the nonlinear influencing factors to obtain the predicted trend curve of resource performance changes, and an error estimation tool is used to perform deviation analysis on the predicted trend curve to obtain the predicted error range of resource performance changes.

[0062] In a preferred embodiment, a first training set is established based on the periodic pattern feature values ​​of the second structured resource dataset. The first training set is used to train a Long Short-Term Memory (LSTM) network model to capture nonlinear influencing factors of resource performance changes, resulting in a set of nonlinear influencing factors. Specifically, the initial training set is mapped based on the periodic pattern feature values ​​to establish the first training set. For example, the first training set includes CPU utilization and its nonlinear influencing factors, such as historical data like user access volume and task scheduling frequency. The first training set is used to train the LSM network model to capture nonlinear influencing factors of resource performance changes. Specifically, through capture analysis, it is found that peak user access volume is highly correlated with peak CPU utilization, and there is a nonlinear relationship; for example, when access volume exceeds 10,000 times / hour, CPU utilization increases exponentially. Therefore, the set of nonlinear influencing factors includes user access volume, data transmission volume, and the number of background tasks.

[0063] In a preferred embodiment, a first set of nonlinear impact factors is obtained by screening the nonlinear impact factor set and selecting the first nonlinear impact factor that exceeds a corresponding preset threshold. A time series forecasting tool is used to simulate the future trend of the first nonlinear impact factor to obtain a predicted trend curve and fluctuation direction reflecting the changes in resource performance. An error estimation tool is used to perform deviation analysis on the predicted trend curve to obtain the prediction error range of resource performance changes. Specifically, for the nonlinear impact factor, such as user visits exceeding 12,000 times / hour, a future trend simulation is performed. Using a time series forecasting tool such as the ARIMA model, the trend of user visits over the next 24 hours is predicted based on historical visit data, resulting in a predicted trend curve for user visits. The predicted trend curve shows that the visit volume will reach 15,000 times / hour on the next working day, and the predicted CPU utilization may rise to 85%, with an upward fluctuation direction. The prediction results of this embodiment help to allocate computing resources in advance and avoid system performance bottlenecks.

[0064] The method employs error estimation tools to perform deviation analysis on the predicted trend curve, obtaining the prediction error range for resource performance changes. Specifically, the difference range between actual and predicted values ​​is obtained from historical data to determine the prediction error range for resource performance changes. Error estimation tools such as root mean square error (RMSE) are used to analyze the prediction result deviation, finding that the difference between the actual CPU utilization and the predicted value is within ±5%. For example, a predicted CPU utilization of 80% and an actual value of 78% result in an error of 2%. Analysis of multi-day data yields a prediction error range of ±4% to ±6%. This prediction error range, as described in this embodiment, can guide resource scheduling and set safety thresholds to address prediction deviations.

[0065] The embodiments of this invention demonstrate significant effectiveness. Extracting the periodic pattern feature values ​​helps identify regular patterns in resource usage. Mapping the first training set based on these periodic pattern feature values ​​helps determine key influencing factors on system performance. The predicted trend curve provides a basis for resource planning, and error analysis enhances prediction reliability. Adjusting server load balancing in advance based on the predicted trend curve can reduce downtime risk and improve system stability by approximately 20%. Furthermore, combining analysis with nonlinear influencing factors and optimizing scheduling strategies can improve resource utilization by approximately 15%, saving costs for enterprises. These closely integrated and logically sound techniques collectively support scientific decision-making in resource performance management.

[0066] Step S06: Calculate the resource ratio: Determine the current predicted demand deviation value based on the predicted trend curve and the predicted error range, and calculate the first resource ratio or the second resource ratio based on the current resource allocation status and the predicted demand deviation value; the first resource is a scarce resource, and the second resource is a surplus resource.

[0067] In a preferred embodiment, step S06 includes the following steps:

[0068] Step S061: Obtain bottleneck response delay data and bottleneck distribution area data of the first nonlinear influence factor from historical data of the first nonlinear influence factor. Use a time series decomposition tool to extract the periodic feature values ​​of the bottleneck response delay data and the distribution density of the bottleneck distribution area data to obtain the periodic pattern of the bottleneck response delay and the spatial distribution characteristics of the bottleneck distribution area. Specifically, using a time series decomposition tool to extract the periodic feature values ​​of the bottleneck response delay data involves splitting the bottleneck response delay data into trend data, periodic data, and residual data. Specifically, the response delay data of the current server cluster is collected minute by minute over a one-month period. Decomposition reveals that response delay bottlenecks mostly occur between 10:00 AM and 11:00 AM on weekdays, with the delay value increasing from an average of 0.5 seconds to 2 seconds, while it remains stable at around 0.2 seconds during non-working hours. This periodic pattern indicates significant diurnal fluctuations in resource demand. This embodiment of the invention can identify specific nodes currently experiencing resource shortages based on the spatial distribution characteristics, such as the distribution density, of the bottleneck distribution area. Within the same cluster, the distribution density shows that three server nodes handled 80% of the request load during peak periods, while other nodes accounted for only 20%, indicating a significant uneven distribution of resources. This spatial distribution characteristic provides a crucial basis for subsequent resource adjustments.

[0069] Step S062: Based on the periodic pattern and spatial distribution characteristics, perform feature mapping on the current resource allocation status dataset to obtain a first resource allocation status dataset. Calculate the predicted demand deviation of the first resource allocation status dataset to obtain the first resource ratio or the second resource ratio. Specifically, calculate the predicted demand deviation of the first resource allocation status dataset based on the predicted trend curve and the predicted error range. Preferably, after calculating the predicted demand deviation of the first resource allocation status dataset, perform correlation factor analysis on the first resource allocation status dataset and the predicted demand deviation. In this embodiment of the invention, for a situation where the current system resource shortage ratio reaches 30% during peak periods and the resource surplus ratio is 40% during trough periods, analysis based on the correlation factors determines that the main factors affecting the predicted error range include task scheduling frequency and request volume fluctuations, laying the foundation for subsequent system resource optimization. In this embodiment, in the field of server resource management, the analysis of bottleneck response latency and bottleneck distribution areas in historical data can be approached from the perspective of periodicity and spatial distribution, combined with resource allocation status for optimization and adjustment.

[0070] Step S07: Generate resource allocation scheme: Calculate allocation priority score for the current first resource or second resource, determine the first priority sequence of the current resource allocation based on the priority score, and generate a first resource allocation scheme based on the first priority sequence.

[0071] When the proportion of the first resource or the proportion of the second resource exceeds a corresponding preset threshold, a logistic regression algorithm is used to calculate the priority score of the current resource, and a first priority sequence for the current resource allocation is determined based on the priority score. The corresponding preset threshold refers to a first preset threshold for the proportion of the first resource or a second preset threshold for the proportion of the second resource. Specifically, the allocation priority score is calculated using a logistic regression algorithm based on historical data of the task scheduling frequency and data processing throughput of the first or second resource. By prioritizing resource allocation to nodes with high loads, this embodiment can effectively identify key areas requiring resource optimization.

[0072] For the first priority sequence, deviation analysis is performed based on the system load balancing index and response time threshold. A first resource allocation scheme for resource utilization is generated based on the results of the deviation analysis. Specifically, the current system response time threshold is 1 second, while the response time of some nodes has reached 1.5 seconds. An optimized allocation scheme is derived through deviation analysis, which increases computing power by 20% to high-load nodes. This embodiment can significantly alleviate the system resource bottleneck problem through this adjustment. The first resource allocation scheme in this embodiment includes task splitting and node expansion. Task splitting involves breaking large tasks into smaller tasks and distributing these smaller tasks to low-load nodes. Node expansion involves temporarily increasing virtual machine resources to cope with sudden resource shortages. These measures in this embodiment support each other and jointly improve resource utilization efficiency. Through the analysis of periodic patterns and spatial distribution characteristics, problem areas can be located more accurately, while priority scoring and deviation analysis ensure the rationality of resource allocation, ultimately forming a complete optimization loop.

[0073] In a preferred embodiment, a second resource allocation scheme is generated for system resources in the first priority sequence whose allocation priority score is higher than a corresponding preset threshold. Specifically, the second resource allocation scheme is generated based on bottleneck impact range data, resource load data, resource load peak fluctuation data, and performance stability conditions.

[0074] In a preferred embodiment, generating the second resource allocation scheme includes the following steps:

[0075] Step S071: Obtain the resource load data, resource load peak fluctuation data, and bottleneck impact range data from the historical resource usage records. Use a data comparison tool to process and analyze the bottleneck impact range data and resource load peak fluctuation data to obtain the system load distribution characteristic value.

[0076] Step S072: Based on the load distribution characteristic value, map the system throughput data to obtain first load distribution characteristic data. Select the first load distribution characteristic data that meets preset conditions as second load distribution characteristic data. The preset conditions are that the first load distribution characteristic data is generated during high-frequency task scheduling periods and the system meets performance stability conditions. Specifically, meeting performance stability conditions means meeting response speed requirements. The system throughput data is obtained from the historical resource usage records.

[0077] Step S073: Use a time series decomposition tool to extract the first fluctuation pattern of the second load distribution feature data, and generate the second resource allocation scheme based on the first fluctuation pattern.

[0078] In a preferred embodiment, when the first resource allocation scheme does not meet the performance stability condition, the resource occupancy ratio within the bottleneck's influence range is obtained through load balancing adjustment rules. The adjusted resource allocation ratio is calculated, and the final allocation feasibility is determined. Combining the allocation feasibility result and real-time dynamic adjustment data, a second resource allocation scheme is generated using a task scheduling frequency control tool, resulting in an adjustment execution sequence for the second resource allocation scheme that meets the constraints. Specifically, when the performance stability condition requires a response time of no more than 1.2 seconds, but the response time of some nodes during peak periods has reached 1.5 seconds, this indicates a deviation. At this time, the resource occupancy ratio within the resource bottleneck range is analyzed through load balancing adjustment rules. It is found that high-load nodes occupy 70% of the cluster's computing resources, while other nodes have low resource utilization. The adjusted resource allocation ratio of the first resource allocation scheme is set to transfer 20% of the computing resources from low-load nodes to high-load nodes, and a feasibility analysis of the first resource allocation scheme is performed to ensure that no new bottlenecks are introduced.

[0079] In a preferred embodiment, during the feasibility analysis of the first resource allocation scheme, combined with dynamically adjusted real-time data, if the resource load still shows that some nodes are overloaded, the first resource allocation scheme is adjusted using a task scheduling frequency control tool to obtain a second resource allocation scheme. Specifically, in the first resource allocation scheme, the current task scheduling frequency is 120 times per minute. After adjustment, the second resource allocation scheme reduces the scheduling frequency of some non-urgent tasks to 80 times per minute, while distributing tasks to low-load nodes. The resource allocation scheme of this embodiment can effectively balance the overall load.

[0080] In a preferred embodiment, the resource allocation tasks in the first resource allocation scheme are hierarchically classified, specifically divided into high-priority tasks, medium-priority tasks, and low-priority tasks. For high-priority tasks, they are ensured to be allocated to nodes with optimal performance, while low-priority tasks are assigned to nodes with lower loads. This hierarchical strategy and resource ratio adjustment in this embodiment support each other, ensuring efficient resource utilization and stable system performance.

[0081] In a preferred embodiment, analyzing the distribution information of the bottleneck's impact range includes analyzing the resource usage of each node. Specifically, if a first node is identified within the current system cluster, and this first node is consistently under high load during peak periods, real-time monitoring and historical data comparison of the first node confirms it as the primary bottleneck area, then additional resources are prioritized for allocation to the first node. This precise positioning method in this embodiment helps to quickly alleviate pressure on critical areas.

[0082] In a preferred embodiment, a dynamic monitoring mechanism is established to ensure system performance stability. When the response time approaches a preset response threshold, resource adjustments are automatically triggered to ensure that the system always meets performance stability conditions during resource adjustments. This embodiment combines the dynamic monitoring mechanism with fluctuation patterns extracted by time series decomposition tools to ensure that the resource allocation scheme meets both periodic requirements and can cope with sudden load changes, thereby improving overall system stability.

[0083] In a preferred embodiment, the peak load fluctuation data of the server cluster shows that during peak hours on weekdays, the average load rate reaches 85%, while during off-peak hours it is only 30%. This fluctuation difference suggests the need to dynamically adjust the resource allocation strategy. After processing with a data comparison tool, the load distribution characteristic value is obtained. This load distribution characteristic value indicates that some nodes have a load close to 100% during peak hours, while other nodes have only 20%, showing a significant uneven distribution.

[0084] In a preferred embodiment, by mapping the load distribution characteristics and combining them with historical throughput data, it can be found that during periods when task scheduling frequency is higher than average, such as from 9:00 AM to 11:00 AM daily, throughput demand surges to 8,000 messages per second, while the average is only 3,000 messages per second. After extracting the fluctuation patterns using time series decomposition tools, it is found that resource demand exhibits periodic peaks in specific time periods. Therefore, the priority should be to focus adjustments on high-load nodes during these periods, increasing resource supply to alleviate pressure.

[0085] In a preferred embodiment, bottleneck distribution area data and predicted resource demand deviation data of the second resource allocation scheme are obtained, and a genetic algorithm is used to search for the current resource allocation status of the system. The current resource allocation status is the resource allocation status adjusted according to the second resource allocation scheme. A second fluctuation pattern of the current resource allocation status is extracted by a time series decomposition tool, and a third resource allocation scheme is generated based on the second fluctuation pattern.

[0086] In a preferred embodiment, the second resource allocation scheme is adjusted according to the following steps to meet the stability requirements of predicted resource trends:

[0087] Step S074: Obtain the distribution characteristics of the resource bottleneck impact range and the load balancing data corresponding to the resource bottleneck impact range from the historical resource usage records. Perform deviation analysis on the distribution characteristics and response time using a data comparison tool to obtain resource deviation data.

[0088] Step S075: Based on the resource deviation data and the current resource allocation status, for periods when the current resource scheduling frequency is higher than the average level, extract the second fluctuation pattern of the current resource allocation status using a time series decomposition tool. Adjust the second resource allocation scheme according to the second fluctuation pattern to obtain a third resource allocation scheme. Use a genetic algorithm to search for the current resource allocation status.

[0089] Step S076: When the third resource allocation scheme does not meet the predicted resource trend stability requirements, the resource occupancy ratio within the bottleneck's influence range is obtained through load balancing adjustment rules. Based on the resource occupancy ratio within the bottleneck's influence range and the response speed requirements, a feasibility analysis is performed on the third resource allocation scheme. Based on the feasibility analysis results and real-time resource occupancy ratio data from dynamic resource adjustments, the third resource allocation scheme is regenerated using the system task scheduling frequency control tool, resulting in an adjustment execution sequence for the third resource allocation scheme that meets the performance stability requirements.

[0090] In a preferred embodiment, within the field of server resource management, the distribution characteristics analysis of the bottleneck's impact range can extract key information from historical load data and, combined with the distribution of load balancing data, delve deeper into potential problems in resource allocation. Specifically, in a server cluster with 8 nodes, 2 nodes consistently achieve load rates exceeding 90% during peak hours, while the load rates of the other nodes are only around 25%. This uneven distribution indicates a need for targeted adjustments. This embodiment uses data comparison tools to process the relationship between these distribution characteristics and response speed, discovering that the response time of high-load nodes exceeds a preset 1.3 seconds, while the response time of low-load nodes is only 0.5 seconds, thus initially establishing a basis for deviation correction.

[0091] In a preferred embodiment, based on the deviation correction criteria and real-time resource occupancy ratio data, during periods with higher-than-average scheduling frequency, such as 10:00 AM to 12:00 PM on weekdays, the fluctuation pattern of the allocation ratio can be extracted using a time series decomposition tool. Analysis results show that resource demand fluctuates slightly every 30 minutes during this period. Therefore, the priority adjustment should focus on high-load nodes, increasing their resource supply ratio to 1.5 times the original level to alleviate resource pressure. When a deviation between the priority direction and trend stability is found, such as the load rate of some nodes remaining unstable after adjustment, further adjustments using load balancing rules are needed. Resource occupancy ratio data from bottleneck areas should be obtained, combined with response speed constraints (e.g., response time must not exceed 1.2 seconds), to determine the feasibility of the final allocation scheme.

[0092] In a preferred embodiment, based on the feasibility analysis of the allocation ratio and combined with the long-term fluctuations in resource demand forecasts, it is predicted that peak load demand will increase by 15% within the next week. A specific resource combination plan can be generated using a task scheduling frequency control tool. Specifically, the task scheduling frequency is adjusted from 100 times per minute to 80 times per minute, while some non-critical tasks are allocated to low-load nodes to alleviate the pressure on high-load nodes. The final resource adjustment plan must meet trend stability requirements, such as ensuring that load rate fluctuations are controlled within 10%. This embodiment, from adjusting for short-term fluctuations to long-term demand forecasts, forms a complete resource management closed loop, which helps improve the overall operating efficiency and stability of the system.

[0093] In a preferred embodiment, the analysis of the distribution characteristics of the bottleneck's impact range includes node-level resource monitoring. Specifically, additional resources are prioritized for high-load nodes, and the rationality of resource adjustment directions is confirmed through comparison with historical data.

[0094] In a preferred embodiment, a real-time resource monitoring mechanism is set up to ensure the stability of system performance. Specifically, resource adjustments are automatically triggered when the response time approaches a preset response threshold to ensure that the system always meets performance stability conditions during resource adjustments, reducing the delay caused by manual intervention. For periods when the scheduling frequency is higher than average, the resource allocation ratio of system tasks is dynamically adjusted to avoid overloading of a single node. This embodiment, through the mutual support of multiple implementation methods, from node distribution to task scheduling and response speed monitoring, forms a complete resource optimization strategy, effectively improving resource utilization and system stability.

[0095] Step S08: When the system resource utilization efficiency does not reach the preset performance target, adjust the parameter update rate of the long short-term memory network model, redetermine the prediction trend curve and prediction error range, and repeat the steps of calculating the resource ratio and generating the resource allocation scheme until the system resource utilization efficiency reaches the preset performance target.

[0096] In a preferred embodiment, a dynamic resource allocation instruction is generated according to the third resource allocation scheme; the current system state is updated in real time according to the dynamic resource allocation instruction; adjusted resource utilization efficiency data is obtained; and it is determined whether the preset performance target has been achieved based on the resource utilization efficiency data. Specifically, the steps include:

[0097] Step S081: Extract the fluctuation characteristics of resource adjustment frequency using a time series decomposition tool. Based on the fluctuation characteristics and monitored real-time resource occupancy ratio data, change the priority order of the dynamic resource allocation instructions. According to the priority order, use an instruction generation tool to hierarchically classify the dynamic resource allocation instructions, obtain the latest resource utilization efficiency data from system status monitoring, and determine the execution order of resource instructions based on the hierarchically classified dynamic resource allocation instructions. Specifically, in the field of server resource management, the fluctuation characteristics of adjustment frequency can be extracted using a time series decomposition tool to analyze the dynamic allocation demand data. The time series decomposition tool can break down the changing trend of resource demand into periodic fluctuations and random noise, helping to identify possible regular changes in resource allocation. In a server cluster, resource demand shows a significant peak every 2 hours within a certain period. This embodiment can clearly capture this periodic characteristic using the decomposition tool, providing a basis for subsequent adjustments.

[0098] Specifically, this embodiment, by combining real-time resource utilization data, can further determine the priority order of dynamic resource allocation change instructions. The real-time resource utilization data reflects the operational status of server nodes in real time, such as key indicators like CPU utilization and memory usage. Specifically, in a cluster of 10 nodes, 3 nodes consistently have CPU utilization above 80%, while the other nodes are only around 30%. Prioritization will place these 3 high-load nodes as the primary adjustment targets. This resource adjustment mechanism in this embodiment ensures timely resource allocation.

[0099] Specifically, by using an instruction generation tool to hierarchically classify the dynamic resource allocation instructions, the latest resource utilization efficiency data can be obtained from system resource status monitoring, thereby determining the execution order of the dynamic resource allocation instructions. Hierarchical classification categorizes system resource adjustment tasks into three types: urgent, general, and low priority. Urgent tasks involve resource replenishment for high-load nodes, general tasks involve optimization and adjustment for medium-load nodes, while low-priority tasks handle long-term planning. This embodiment ensures a reasonable and efficient execution order of resource instructions through this method.

[0100] Step S082: Using the execution order, compare the resource utilization efficiency data with a preset efficiency threshold using a data comparison tool. If the resource utilization efficiency data is lower than the preset efficiency threshold, adjust the dynamic resource allocation command using a frequency control tool, and determine whether the adjusted resource utilization efficiency data reaches the preset performance target. This embodiment uses a data comparison tool to compare the resource utilization efficiency data with the preset efficiency threshold, providing a clear basis for adjusting the dynamic resource allocation command. Specifically, the response time threshold is 1.2 seconds, while the current high-load node's response time is 1.5 seconds, significantly lower than the standard value. In this case, it is necessary to adjust the change rhythm using a frequency control tool, such as reducing the task scheduling frequency from 90 times per minute to 70 times per minute, to alleviate node pressure, and observe whether the response time recovers to within 1.2 seconds after the adjustment.

[0101] Step S083: The adjusted resource utilization efficiency data is analyzed from multiple dimensions using a target determination tool to obtain a combination of resource instructions that matches the preset performance target, serving as the final dynamic resource allocation instruction. Specifically, after adjustment, the response time is reduced to 1.1 seconds, but the load rate of some nodes remains uneven. Multi-dimensional analysis is used, combining indicators such as load rate, response time, and resource occupancy ratio, to generate a comprehensive dynamic resource allocation scheme and corresponding dynamic resource allocation instructions. For example, some tasks are transferred to low-load nodes, while increasing resource supply to high-load nodes by 10%, ensuring the achievement of the overall performance target. The final determination of the dynamic resource allocation scheme requires comprehensive multi-dimensional analysis results to ensure that resource allocation matches the predicted demand trend. Specifically, the prediction shows that resource demand during peak hours will increase by 20% in the next 24 hours. The dynamic resource allocation scheme prioritizes reserving resources for high-load nodes while dynamically adjusting the task allocation ratio to avoid overloading a single node. The dynamic resource allocation method in this embodiment can effectively improve the rationality of resource utilization and ensure the stable operation of the system during peak periods.

[0102] The logical progression from the core solution to the extended solution in this embodiment of the invention can further enrich the diversity of dynamic resource allocation. The core solution of this embodiment focuses on the emergency adjustment of high-load nodes, while the extended solution considers the pre-allocation of resources for low-load nodes to cope with sudden demands. This multi-faceted support approach ensures the integrity and adaptability of the resource allocation solution, providing a more flexible response strategy for server resource management.

[0103] In a preferred embodiment, when the system resource utilization efficiency fails to meet the preset performance target, adjusting the parameter update rate of the Long Short-Term Memory network model and determining the prediction trend curve and prediction error range includes the following steps:

[0104] Step S084: When the system resource utilization efficiency fails to meet the preset performance target, system resource usage data and system resource demand prediction data are obtained from historical resource usage records using a data backtracking tool. In the field of server resource management, when system resource utilization efficiency fails to meet the preset performance target, a series of tools and methods can be used for in-depth analysis and adjustment to gradually optimize resource allocation strategies. First, regarding the application of the data backtracking tool, system resource usage data and system resource demand prediction data can be extracted from historical resource usage records to analyze the nonlinear factor dependencies. Specifically, in a server cluster, historical resource usage records show a complex nonlinear correlation between resource demand and time periods. For example, resource demand always shows a surge trend from 9:00 AM to 11:00 AM on weekdays, while it remains relatively stable at other times. This embodiment uses a data backtracking tool to extract the distribution characteristics of this dependency relationship, providing a basis for subsequent adjustments.

[0105] Step S085: Analyze the first nonlinear influence factor between the system resource usage data and the system resource demand prediction data, extract the dependency relationship features from the first nonlinear influence factor to obtain the distribution map of the dependency relationship features, and recalculate the feature weights of the Long Short-Term Memory network model using a weight allocation tool based on the distribution map of the dependency relationship features and the system resource demand prediction results to obtain the second network weight parameters. The weight allocation tool can recalculate the feature weights based on the analysis of the dependency relationship distribution features. Specifically, the influence weight of time period on resource demand in historical resource usage records was originally 0.3, but analysis shows that its actual impact is greater, so it can be adjusted to 0.5, while reducing the weights of other secondary factors. This recalculation method in this embodiment combines prediction trends and can generate updated network weight parameters, laying the foundation for model optimization.

[0106] Step S086: Based on the second network weight parameters, the update rate of the Long Short-Term Memory network model parameters is corrected using a parameter adjustment tool to generate an updated prediction trend curve and prediction error range. When the prediction trend curve has not reached a stable state, fluctuation characteristics are extracted from the prediction trend data using a fluctuation analysis tool, and combined with a trend stability index to generate the final prediction trend curve and prediction error range. For cases where the prediction trend curve has not reached a stable state, the fluctuation analysis tool can extract fluctuation characteristics from the prediction trend data. When analysis reveals that fluctuations are mainly concentrated during peak periods and are related to certain abnormal events in historical data, such as sudden traffic surges, the final prediction result can be determined by combining the trend stability index. For example, it can be predicted that resource demand will increase by 20% during peak periods, and resource allocation can be adjusted accordingly.

[0107] The core solution of this invention focuses on optimizing the prediction model through data backtracking and weight adjustment, while the extended solution considers reserving a certain resource buffer based on fluctuation analysis to cope with emergencies. The core solution prioritizes resource allocation for high-load nodes, while the extended solution reserves 5% resource margin for medium-load nodes to improve overall stability. This multi-faceted support ensures the comprehensiveness and adaptability of the resource management strategy. Simultaneously, the data backtracking tool can be implemented in conjunction with a historical log system, the weight allocation tool can adjust parameters through a simple rule engine, and the fluctuation analysis tool relies on a time series analysis module to extract features. These methods in this embodiment work together to form a complete optimization chain, providing strong support for server resource management, especially effectively avoiding resource shortages or uneven allocation during peak periods.

[0108] In a preferred embodiment, the final dynamic resource allocation instruction is determined according to the following steps:

[0109] Step S087: Based on the predicted trend curve and the predicted error range, obtain the second bottleneck distribution area data and the second bottleneck influence range data. Use a data integration tool to classify the second bottleneck distribution area data and the second bottleneck influence range data, obtain key indicators for resource allocation, and calculate the priority score of system resources based on the key indicators. When the priority score is inconsistent with the corresponding preset score threshold, use a weight adjustment tool to perform a secondary calculation on the second bottleneck distribution area data and the second bottleneck influence range data. In the field of server resource management, based on the analysis results of predicted trends and trend stability, the data integration tool can be used to classify bottleneck distribution and influence range, and generate a preliminary priority score list. Specifically, in a server cluster in a data center, assuming that the predicted trend shows that the peak resource demand is from 2 pm to 4 pm, some nodes become bottlenecks due to excessive CPU load. The data integration tool analyzes historical load logs and real-time monitoring resource usage data to classify bottleneck nodes into three categories: high load, medium load, and low load, and determines the influence range, such as high load nodes affecting 10% of the cluster throughput. The key indicators include CPU utilization and memory usage, generating a preliminary priority score list. The priority score list is sorted based on a comprehensive score of node load and impact range. A node with CPU utilization reaching 90% and impact range covering critical business operations receives the highest priority score. If the sorting criteria of the priority score list are inconsistent with preset thresholds, such as a high-load node's score not meeting expectations, the weight adjustment tool can perform a secondary calculation of the priority scores. Specifically, in the initial weight allocation, CPU utilization has a weight of 0.4, and impact range has a weight of 0.3. However, analysis shows that impact range has a greater impact on business continuity, so the impact range weight can be adjusted to 0.5, and the CPU utilization weight reduced to 0.3. The updated priority score data reflects more accurate resource allocation needs; for example, if a node's score increases from 80 to 95, its priority is adjusted accordingly. This resource adjustment method in this embodiment ensures that resource allocation is more closely aligned with actual business needs.

[0110] Step S088: The scheme generation tool is used to refine the third resource allocation scheme corresponding to the priority score. Specifically, constraints are extracted from the data of the second bottleneck distribution area to generate refined allocation scheme content that matches the adjustment instructions. It is then determined whether the refined content meets the preset performance target requirements. Based on the updated priority score data, the scheme generation tool is used to refine the priority score data and extract constraints from the bottleneck distribution. Specifically, bottleneck nodes may be limited by hardware performance or network bandwidth. The tool generates refined allocation scheme content accordingly, such as prioritizing the allocation of additional virtual machine resources to high-load nodes and limiting the resource consumption of low-priority tasks. The refined content must meet the target requirements, such as ensuring that the overall cluster response time is less than 100 milliseconds during peak periods. The scheme generation tool verifies the feasibility of the refined content by simulating the allocation effect.

[0111] Step S089: Use an instruction conversion tool to convert the detailed content into specific resource adjustment instructions, and determine the final dynamic resource allocation instructions based on the latest predicted trend curve. Preferably, the instruction conversion tool converts the detailed allocation scheme into specific resource adjustment instructions, and generates the final dynamic resource allocation instructions by combining the updated trend data. Specifically, the dynamic resource allocation instructions are: "Add 2 virtual CPU cores to node A and limit the bandwidth usage of node B to 50%". The application scope of the dynamic resource allocation instructions is limited to bottleneck areas, such as targeting only high-load nodes, to ensure the targeted nature of the adjustment. Preferably, the instruction conversion tool interfaces with the cluster management platform and automatically issues instructions via API to achieve rapid resource adjustment.

[0112] The core solution in this embodiment focuses on generating an accurate priority list through data integration and weight adjustment, while the extended solution considers reserving a 10% resource buffer for potential bottleneck nodes to cope with sudden loads. This multi-faceted approach, combined with data integration, weight adjustment, solution generation, and instruction conversion tools, forms a complete resource optimization chain, supporting the efficient operation of the server cluster.

[0113] In a preferred embodiment, after allocating resources according to the final dynamic resource allocation instruction, it is determined whether the system meets the preset performance stability requirements, specifically as follows:

[0114] Based on the final dynamic resource allocation instruction, a data acquisition tool extracts real-time data from the resource surplus ratio and bottleneck response latency. Combined with the dynamic adjustment rules, an initial system status update dataset is generated, resulting in an updated result reflecting the system resource utilization rate. Using this initial system status update dataset, a load balancing tool recalculates the resource load peak value. Constraints are extracted from the bottleneck distribution area data to determine the adjusted resource load peak value. If the adjusted resource load peak value exceeds a preset response threshold, a classification processing tool analyzes the load balancing status to obtain a load distribution result that meets the system performance stability requirements. Based on the load distribution result, an instruction conversion tool matches the resource load peak value with preset performance stability requirements, generating a third dynamic resource allocation instruction. After resource allocation according to the third dynamic resource allocation instruction, it is determined whether the system meets the performance stability requirements.

[0115] This embodiment uses data acquisition tools to extract real-time data from resource surplus ratio and bottleneck response latency, providing a foundation for generating system status update datasets.

[0116] In a preferred embodiment, within the field of server resource management, during the operation of a data center cluster, the node resource surplus ratio is 30%, meaning CPU and memory utilization are far below expectations, while another node's response latency reaches 150 milliseconds, exceeding the normal threshold of 100 milliseconds. A data acquisition tool monitors and records these key indicators in real time and, combined with dynamic adjustment rules (such as prioritizing nodes with excessive latency), generates an initial dataset. This initial dataset reflects resource utilization, such as the inefficient use of surplus nodes and the performance bottlenecks of bottleneck nodes, providing a basis for subsequent optimization. Simultaneously, a load balancing tool recalculates the peak resource load based on the initial dataset and extracts constraints from the bottleneck distribution area data.

[0117] In a preferred embodiment, analysis revealed that the bottleneck node's response delay was due to insufficient network bandwidth, with the limiting condition being a bandwidth utilization rate of 90%. The load balancing tool simulated resource reallocation and calculated adjusted peak load data. For example, migrating some tasks from the bottleneck node to surplus nodes reduced bandwidth utilization to 70%, restoring the response time to 90 milliseconds. This resource adjustment method ensures more balanced resource utilization. When the adjusted peak load data still exceeds the preset response time threshold, a classification tool analyzes the load balancing status. Specifically, if the adjusted node's response time is 120 milliseconds, still exceeding the limit, the classification tool categorizes nodes into high-load, medium-load, and low-load types. Analyzing the task characteristics of high-load nodes reveals that some non-critical tasks are consuming excessive resources. By limiting task priorities, a load distribution result conforming to performance stability standards is generated, ensuring that the response time of critical tasks is below 100 milliseconds.

[0118] Preferably, the instruction conversion tool matches peak load data with performance stability standards to generate the dynamic resource allocation instruction. The dynamic resource allocation instruction migrates non-critical tasks from node C to node D and increases the bandwidth allocation on node C to 2Gbps. The instruction conversion tool automatically issues instructions by interfacing with the cluster management platform, ensuring that the adjustments take effect in real time. It should be noted that the generation of the dynamic resource allocation instruction considers the specificity of bottleneck areas, adjusting only high-load nodes to avoid affecting other nodes. This resource adjustment method in this embodiment improves system stability.

[0119] The extended solution in this embodiment can reserve a resource buffer for potential bottleneck nodes. Specifically, it includes reserving 10% of bandwidth resources to cope with sudden surges in tasks. This embodiment, through a multi-faceted support solution, forms a complete optimization chain from data acquisition to instruction conversion, ensuring the efficient operation of the server cluster.

[0120] Based on the above method embodiments, corresponding apparatus embodiments are provided;

[0121] like Figure 2 As shown, another embodiment of the present invention provides an intelligent data center operation and maintenance system based on the Internet of Things, including: a structured resource data module, a trend prediction module, a computing resource ratio module, a resource allocation scheme generation module, and a preset performance target achievement module;

[0122] The structured resource data module is used to extract indicator resource datasets from historical resource usage records, smooth the time series of the indicator resource datasets using a median filtering algorithm to obtain a time series resource dataset; extract multi-dimensional indicators and filter outliers from the time series resource datasets to obtain a first structured resource dataset; use a time series decomposition tool to extract trend and periodic items from the first structured resource dataset to obtain a second structured resource dataset; the second structured resource dataset contains periodic pattern feature values ​​of resource performance changes;

[0123] The prediction trend module is used to capture the nonlinear influencing factors of the second structured resource dataset using a long short-term memory network model to obtain a set of nonlinear influencing factors, and to filter the first nonlinear influencing factors in the set of nonlinear influencing factors that exceed the corresponding preset threshold to obtain a first set of nonlinear influencing factors; and to perform future trend simulation on the first set of nonlinear influencing factors to obtain the prediction trend curve and prediction error range of resource performance changes.

[0124] The resource ratio calculation module is used to determine the current predicted demand deviation value based on the predicted trend curve and the predicted error range, and to calculate a first resource ratio or a second resource ratio based on the current resource allocation status and the predicted demand deviation value; the first resource is a scarce resource, and the second resource is a surplus resource;

[0125] The resource allocation scheme generation module is used to calculate an allocation priority score for the current first resource or second resource, determine a first priority sequence for the current resource allocation based on the priority score, and generate a first resource allocation scheme based on the first priority sequence.

[0126] The module for achieving the preset performance target is used to adjust the parameter update rate of the long short-term memory network model and redetermine the prediction trend curve and prediction error range when the system resource utilization efficiency fails to reach the preset performance target. The module for calculating the resource ratio and the module for generating the resource allocation scheme are then executed repeatedly until the system resource utilization efficiency reaches the preset performance target.

[0127] It is understood that the above system embodiments correspond to the method embodiments of the present invention, and can implement the IoT-based intelligent operation and maintenance method for data centers provided by any of the above method embodiments of the present invention.

[0128] The above are preferred embodiments of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A data center intelligent operation and maintenance method based on the Internet of Things, characterized in that, include: Extract the indicator resource dataset from historical resource usage records, and smooth the time series of the indicator resource dataset using the median filtering algorithm to obtain the time series resource dataset; Multi-dimensional indicators are extracted and outliers are filtered from the time-series resource dataset to obtain the first structured resource dataset; The trend and period items of the first structured resource dataset are extracted using a time series decomposition tool to obtain the second structured resource dataset; The second structured resource dataset contains periodic pattern feature values ​​of resource performance changes; A long short-term memory network model is used to capture the nonlinear impact factors of the second structured resource dataset to obtain a set of nonlinear impact factors. The first nonlinear impact factor that exceeds the corresponding preset threshold in the set of nonlinear impact factors is selected to obtain a first set of nonlinear impact factors. A future trend simulation is performed on the first set of nonlinear influencing factors to obtain the predicted trend curve of resource performance changes and the prediction error range. The steps for calculating resource ratios are as follows: First, determine the current predicted demand deviation value based on the predicted trend curve and the predicted error range. Then, calculate either a first resource ratio or a second resource ratio based on the current resource allocation status and the predicted demand deviation value. The first resource is a scarce resource, and the second resource is a surplus resource. Steps for generating a resource allocation scheme: Calculate an allocation priority score for the current first resource or second resource, determine a first priority sequence for the current resource allocation based on the priority score, and generate a first resource allocation scheme based on the first priority sequence; When the system resource utilization efficiency fails to reach the preset performance target, the parameter update rate of the long short-term memory network model is adjusted, the prediction trend curve and prediction error range are redefined, and the steps of calculating the resource ratio and generating the resource allocation scheme are repeated until the system resource utilization efficiency reaches the preset performance target. The step of calculating the resource ratio includes: The bottleneck response delay data and bottleneck distribution area data of the first nonlinear influence factor are obtained from the historical data of the first nonlinear influence factor. The periodic feature values ​​of the bottleneck response delay data and the distribution density of the bottleneck distribution area data are extracted using a time series decomposition tool to obtain the periodic pattern of the bottleneck response delay and the spatial distribution characteristics of the bottleneck distribution area. Based on the periodic pattern and spatial distribution characteristics, feature mapping is performed on the current resource allocation status dataset to obtain a first resource allocation status dataset. The predicted demand deviation of the first resource allocation status dataset is calculated to obtain the first resource ratio or the second resource ratio.

2. The data center intelligent operation and maintenance method based on the Internet of Things as described in claim 1, characterized in that, The step of generating a resource allocation scheme includes: When the proportion of the first resource or the proportion of the second resource exceeds the corresponding preset threshold, a logistic regression algorithm is used to calculate the priority score of the current resource, and the first priority sequence of the current resource allocation is determined based on the priority score. For the first priority sequence, a deviation analysis is performed based on the system load balancing index and response time threshold, and a first resource allocation scheme with high resource utilization is generated based on the results of the deviation analysis.

3. The IoT-based intelligent operation and maintenance method for data centers as described in claim 2, characterized in that, The method of using a long short-term memory network model to capture the nonlinear influence factors of the second structured resource dataset to obtain a set of nonlinear influence factors is as follows: A first training set is established based on the periodic pattern feature values ​​of the second structured resource dataset. The long short-term memory network model is trained using the first training set to capture the nonlinear influencing factors of resource performance changes, thereby obtaining a set of nonlinear influencing factors.

4. The IoT-based intelligent operation and maintenance method for data centers as described in claim 3, characterized in that, A future trend simulation is performed on the first set of nonlinear influencing factors to obtain the predicted trend curve and prediction error range of resource performance changes, specifically: For the first set of nonlinear influencing factors, a time series forecasting tool is used to simulate the future trends of the nonlinear influencing factors to obtain the predicted trend curve of resource performance changes, and an error estimation tool is used to perform deviation analysis on the predicted trend curve to obtain the predicted error range of resource performance changes.

5. The IoT-based intelligent operation and maintenance method for data centers as described in claim 4, characterized in that, The method further includes: for system resources in the first priority sequence whose allocation priority score is higher than the corresponding preset threshold, generating a second resource allocation scheme based on bottleneck impact range data, resource load data, resource load peak fluctuation data, and performance stability conditions.

6. The IoT-based intelligent operation and maintenance method for data centers as described in claim 5, characterized in that, The generation of the second resource allocation scheme includes: The resource load data, resource load peak fluctuation data, and bottleneck impact range data are obtained from the historical resource usage records. Data comparison tools are used to process and analyze the bottleneck impact range data and resource load peak fluctuation data to obtain the load distribution characteristic value of the system. Based on the load distribution characteristic value, the system throughput data is mapped to obtain the first load distribution characteristic data. The first load distribution characteristic data that meets the preset conditions is selected as the second load distribution characteristic data. The preset conditions are that the first load distribution characteristic data is generated during the high-frequency task scheduling period and the system meets the performance stability conditions. The system throughput data is obtained from the historical resource usage records. The first fluctuation pattern of the second load distribution feature data is extracted using a time series decomposition tool, and the second resource allocation scheme is generated based on the first fluctuation pattern.

7. The IoT-based intelligent operation and maintenance method for data centers as described in claim 6, characterized in that, The method further includes: Obtain bottleneck distribution area data and predicted resource demand deviation data of the second resource allocation scheme, and use a genetic algorithm to search for the current resource allocation status of the system. The current resource allocation status is the resource allocation status adjusted according to the second resource allocation scheme. The second fluctuation pattern of the current resource allocation status is extracted using a time series decomposition tool, and a third resource allocation scheme is generated based on the second fluctuation pattern.

8. The IoT-based intelligent operation and maintenance method for data centers as described in claim 5 or 6, characterized in that, The method further includes: obtaining the distribution characteristics of the resource bottleneck impact range and the load balancing data corresponding to the resource bottleneck impact range from the historical resource usage records, and performing deviation analysis between the distribution characteristics and the response time using a data comparison tool to obtain resource deviation data; Based on the resource deviation data and the current resource allocation status, for periods when the current resource scheduling frequency is higher than the average level, a second fluctuation pattern of the current resource allocation status is extracted using a time series decomposition tool. Based on the second fluctuation pattern, the second resource allocation scheme is adjusted to obtain a third resource allocation scheme.

9. A data center intelligent operation and maintenance system based on the Internet of Things, characterized in that, The method for implementing an IoT-based intelligent operation and maintenance method for data centers as described in any one of claims 1 to 8 includes: a structured resource data module, a trend prediction module, a resource ratio calculation module, a resource allocation scheme generation module, and a preset performance target achievement module. The structured resource data module is used to extract indicator resource datasets from historical resource usage records, smooth the time series of the indicator resource datasets using a median filtering algorithm to obtain a time series resource dataset; extract multi-dimensional indicators and filter outliers from the time series resource datasets to obtain a first structured resource dataset; use a time series decomposition tool to extract trend and periodic items from the first structured resource dataset to obtain a second structured resource dataset; the second structured resource dataset contains periodic pattern feature values ​​of resource performance changes; The prediction trend module is used to capture the nonlinear influencing factors of the second structured resource dataset using a long short-term memory network model to obtain a set of nonlinear influencing factors, and to filter the first nonlinear influencing factors in the set of nonlinear influencing factors that exceed the corresponding preset threshold to obtain a first set of nonlinear influencing factors; and to perform future trend simulation on the first set of nonlinear influencing factors to obtain the prediction trend curve and prediction error range of resource performance changes. The resource ratio calculation module is used to determine the current predicted demand deviation value based on the predicted trend curve and the predicted error range, and to calculate a first resource ratio or a second resource ratio based on the current resource allocation status and the predicted demand deviation value; the first resource is a scarce resource, and the second resource is a surplus resource; The resource allocation scheme generation module is used to calculate an allocation priority score for the current first resource or second resource, determine a first priority sequence for the current resource allocation based on the priority score, and generate a first resource allocation scheme based on the first priority sequence. The module for achieving the preset performance target is used to adjust the parameter update rate of the long short-term memory network model and redetermine the prediction trend curve and prediction error range when the system resource utilization efficiency fails to reach the preset performance target. The module for calculating the resource ratio and the module for generating the resource allocation scheme are then executed repeatedly until the system resource utilization efficiency reaches the preset performance target.

Citation Information

Patent Citations

  • Dynamic load balancing method and system for network routing

    CN119854301A

  • Dynamic scheduling system and method for container calculation

    CN119862037A