Intelligent monitoring method of AI computing power data center

By collecting and preprocessing the load and temperature data of the server in the AI ​​computing power data center, using autoregressive moving average algorithm and multi-scale analysis technology to dynamically adjust the temperature threshold, the problems of monitoring false alarms and hysteresis in the existing technology are solved, and more accurate and timely abnormal identification and processing are achieved.

CN120045420AActive Publication Date: 2025-05-27北京英沣特能源技术有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510518003.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

When monitoring the load and temperature of AI computing power data center servers, the existing technology lacks correlation analysis in complex scenarios, resulting in false positives and lag, making it difficult to optimize scheduling in a timely manner and protect hardware security.

Method used

The intelligent monitoring method of AI computing power data center is adopted to collect CPU average load data and temperature data, and the load is predicted using the autoregressive moving average algorithm, and analyzed it under high-frequency, medium-frequency, and low-frequency sampling scales to comprehensively obtain accurate predicted loads, dynamically adjust the temperature threshold, and realize adaptive abnormal identification.

Benefits of technology

Improve data quality and analysis accuracy, reduce false alarms and lag, and timely identify abnormal situations in the data center, ensuring hardware security and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045420A_ABST
    Figure CN120045420A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of load monitoring analysis, in particular to an intelligent monitoring method of an AI computing power data center. The method comprises the following steps: obtaining an average load data sequence and a temperature sequence; acquiring a predicted load at the next moment by using an autoregressive moving average algorithm; under the high-frequency sampling scale, calculating a high-frequency prediction load at the next moment; under the intermediate frequency sampling scale, calculating an intermediate frequency prediction load at the next moment; under the low-frequency sampling scale, the low-frequency prediction load at the next moment is calculated; obtaining the weights of the high-frequency prediction load, the intermediate-frequency prediction load and the low-frequency prediction load, and carrying out weighted summation to obtain an accurate prediction load at the next moment; calculating a temperature threshold value of the next moment by using the predicted load and the accurate predicted load; predicting temperature data at the next moment according to the temperature sequence, and recording the temperature data as predicted temperature; and if the predicted temperature is greater than or equal to the temperature threshold, sending out a prompt. According to the invention, the temperature abnormity of the computing power data center can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of load monitoring and analysis, and particularly relates to an intelligent monitoring method for an AI computing power data center. Background Art

[0002] With the rapid development of artificial intelligence, big data, and cloud computing, AI computing power data centers are undertaking increasingly high computing loads, and servers are in a running state of high power consumption and high heat output for a long time. Server overload and local overheating problems are one of the main risks affecting the stability, energy efficiency, and hardware life of data centers. In severe cases, it may lead to frequency reduction, downtime of computing nodes, and even hardware damage. Therefore, how to accurately and real-time monitor the running state of servers and give early warnings and optimize scheduling before problems occur is a key challenge for the intelligent management of data centers.

[0003] Currently, the intelligent monitoring of data centers mainly relies on sensors to collect key running data such as CPU load, temperature, fan speed, power consumption, etc., and uses fixed threshold methods for anomaly detection. Among them, the traditional processing method judges whether the CPU is overloaded and whether the temperature is too high based on pre-set fixed thresholds. The autoregressive moving average algorithm can also be used to predict the CPU load and temperature, and give early warnings according to fixed thresholds. However, these are fixed threshold warnings for single indicators and lack relevance in complex scenarios. For example, only monitoring the CPU load may lead to false alarms. In some high-load situations, the server is still within a safe temperature range. At the same time, only monitoring the temperature has a lag. When the system overheats and alarms, it has exceeded the critical point, making it difficult to optimize scheduling in a timely manner, and it may cause irreversible damage to the hardware. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of the present invention is to provide an intelligent monitoring method for an AI computing power data center, and the specific technical solution adopted is as follows: An embodiment of the present invention provides an intelligent monitoring method for an AI computing power data center, and the method includes: Collect the average load data of the CPU and the temperature data in the server and perform preprocessing to obtain an average load data sequence and a temperature sequence; Use the autoregressive moving average algorithm to predict the average load data at the next moment of the current moment based on the average load data sequence, denoted as predicted load; set different sampling scales, including high-frequency sampling scale, medium-frequency sampling scale, and low-frequency sampling scale; At a high-frequency sampling scale, calculate the high-frequency predicted load for the next moment based on the average load data before the next moment; at a medium-frequency sampling scale, calculate the medium-frequency predicted load for the next moment based on the average load data before the next moment; at a low-frequency sampling scale, calculate the low-frequency predicted load for the next moment based on the average load data before the next moment; Obtain the weights of the high-frequency predicted load, medium-frequency predicted load, and low-frequency predicted load based on the high-frequency predicted load and medium-frequency predicted load, and perform weighted summation to obtain the accurate predicted load for the next moment; Calculate the temperature threshold for the next moment using the predicted load and the accurate predicted load; predict the temperature data for the next moment based on the temperature sequence, denoted as the predicted temperature; if the predicted temperature is greater than or equal to the temperature threshold, issue a reminder.

[0005] Preferably, collect the average load data of the CPU and the temperature data inside the server and perform preprocessing to obtain an average load data sequence and a temperature sequence, including: The preprocessing mainly includes data cleaning and normalization processing. Data cleaning includes filling in missing values, identifying outliers, and processing duplicate data; the preprocessed average load data and the temperature data inside the server are respectively formed into an average load data sequence and a temperature sequence.

[0006] Preferably, at a high-frequency sampling scale, calculating the high-frequency predicted load for the next moment based on the average load data before the next moment includes: Set a high-frequency sampling window at the high-frequency sampling scale. The high-frequency sampling window slides with a preset step length. After sliding to the current moment, obtain the distance in time series between an average load data within the high-frequency sampling window and the average load data at the current moment, denoted as the time series distance; perform a negative correlation mapping on the product of the time series distance and the first adjustment coefficient using the exponential function with the natural constant as the base to obtain the weight of this average load data; use the weights of each average load data within the high-frequency sampling window for weighted summation and averaging to obtain the high-frequency predicted load.

[0007] Preferably, at a medium-frequency sampling scale, calculating the medium-frequency predicted load for the next moment based on the average load data before the next moment includes: Set a medium-frequency sampling window at the medium-frequency sampling scale. The medium-frequency sampling window slides with a preset step length. After sliding to the current moment, use the average load data within the medium-frequency sampling window to calculate the medium-frequency predicted load for the next moment; the specific formula for calculating the medium-frequency predicted load is: , where, represents the medium-frequency predicted load for the next moment; represents the mean of a preset number of average load data before the next moment; T represents the number of average load data within the intermediate-frequency sampling window, , μ represents the window center, σ represents the smoothing coefficient; k and j represent the step indices within the intermediate-frequency sampling window; represents the (t - j)-th average load data within the intermediate-frequency sampling window; represents the first data within the intermediate-frequency sampling window; e represents the natural constant.

[0008] Preferably, at the low-frequency sampling scale, calculating the low-frequency predicted load for the next moment based on the average load data before the next moment includes: Setting a low-frequency sampling window at the low-frequency sampling scale, the low-frequency sampling window slides with a preset step. After sliding to the current moment, the average load data within the low-frequency sampling window is used to calculate the low-frequency predicted load for the next moment; the specific formula for calculating the low-frequency predicted load is: , where, represents the low-frequency predicted load for the next moment; A represents the number of average load data within the low-frequency sampling window; and respectively represent the (t - j)-th and (t - j - 1)-th average load data within the low-frequency sampling window; β represents the scaling coefficient; e represents the natural constant; represents the t-th average load data within the low-frequency sampling window.

[0009] Preferably, obtaining the weights of the high-frequency predicted load, intermediate-frequency predicted load, and low-frequency predicted load according to the high-frequency predicted load and intermediate-frequency predicted load includes: Using the exponential function with the natural constant as the base to perform a negative correlation mapping on the difference between the high-frequency predicted load and the low-frequency predicted load to obtain a first mapping result, multiplying the first mapping result by a first scaling coefficient and adding a preset value to obtain a first addition result, and the reciprocal of the first addition result is the weight of the high-frequency predicted load; using the exponential function with the natural constant as the base to perform a negative correlation mapping on the difference between the low-frequency predicted load and the high-frequency predicted load to obtain a second mapping result, multiplying the second mapping result by the first scaling coefficient and adding the preset value to obtain a second addition result, and the reciprocal of the second addition result is the weight of the low-frequency predicted load; subtracting the sum of the weights of the high-frequency predicted load and the low-frequency predicted load from the preset value to obtain the weight of the intermediate-frequency predicted load.

[0010] Preferably, calculating the temperature threshold for the next moment using the predicted load and the accurate predicted load includes: Multiply the exact predicted load by the difference between the predicted load and the first adjustment coefficient, and perform a negative correlation mapping using the exponential function with the natural constant as the base to obtain the third mapping result; calculate the reciprocal of the sum of the third mapping result and the preset value, calculate the difference between the reciprocal and the first adjustment coefficient, and multiply it by the second adjustment coefficient to obtain the multiplication result; the sum of the multiplication result and the temperature threshold at the current moment is the temperature threshold at the next moment.

[0011] Preferably, predicting the temperature data at the next moment according to the temperature sequence, denoted as the predicted temperature, includes: At the high-frequency sampling scale, the high-frequency sampling window slides with a preset step length. After sliding to the current moment, in the temperature sequence, obtain the difference between the temperature data at the first moment within the high-frequency sampling window and the temperature data at the current moment, and divide it by the number of temperature data within the high-frequency sampling window to obtain the change rate of the average difference; add the temperature data at the current moment to the average difference to obtain the temperature data at the next moment, denoted as the predicted temperature.

[0012] The embodiments of the present invention have at least the following beneficial effects: This application preprocesses the collected average load data of the CPU and the temperature data in the server to obtain the average load data sequence and the temperature sequence, which can improve the data quality and ensure the accuracy of subsequent analysis; then, based on the autoregressive moving average algorithm, the predicted load at the next moment of the current moment is predicted based on the average load data sequence. Then, different sampling scales are set, and the collected average load data is analyzed at different sampling scales to obtain the high-frequency predicted load, the medium-frequency predicted load, and the low-frequency predicted load. These are combined to obtain the exact predicted load. The temperature threshold at the next moment is calculated using the predicted load and the exact predicted load, and at the same time, the temperature data at the next moment is predicted to adaptively adjust the temperature threshold, thereby accurately identifying anomalies in the data center. Description of the Drawings

[0013] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0014] Figure 1 It is a flowchart of a method for intelligent monitoring of an AI computing power data center provided by an embodiment of the present invention. Detailed Embodiments

[0015] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, a method for intelligent monitoring of an AI computing power data center according to the present invention, including its specific implementation manner, structure, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0017] The following specifically describes, in conjunction with the accompanying drawings, the specific solution of a method for intelligent monitoring of an AI computing power data center provided by the present invention.

[0018] Embodiment:

[0019] The main application scenario of the present invention is that the AI computing power data center needs to continuously monitor the CPU load to ensure the safe operation and smooth progress of processes.

[0020] Please refer to Figure 1 , which shows the flowchart of a method for intelligent monitoring of an AI computing power data center provided by an embodiment of the present invention. The method includes the following steps: Step S1, collect the average load data of the CPU and the temperature data inside the server and perform preprocessing to obtain the average load data sequence and the temperature sequence.

[0021] Collect the operation status data of the server, including the average load data of the CPU and the temperature data inside the server. The average load data (Load Average) refers to the average number of processes in the running state (Running / Runnable) and the uninterruptible state within a unit time, that is, the average number of active processes. The smaller this value, the less workload the system has and the lower the load; conversely, the higher the load. The average load data can be directly obtained by the system. For example, in the Linux system, use the top command to view the average load data.

[0022] Furthermore, it is necessary to perform preprocessing on the average load data and the temperature data inside the server. The preprocessing mainly includes data cleaning and normalization processing. This is because missing values, outliers, or duplicate data may occur during the data collection process, so data cleaning is required. Data cleaning includes filling in missing values, identifying outliers, and processing duplicate data; at the same time, since the dimensions of the average load data and the temperature data are different, normalization processing is required for convenient subsequent processing.

[0023] For data cleaning, if there are short-term missing values in the average load data or temperature data of the CPU (such as the average load data or temperature data of a certain minute of a certain server is missing), the sliding window average value is used for filling. For outliers, box plot analysis (IQR) is used to detect outliers in the CPU average load data, and the outlier points outside the normal range are removed. The Z-score method is used to detect whether the temperature data is abnormal. If the temperature reading exceeds 3 times the standard deviation of the mean, it may be a sensor failure or an extreme situation, and further confirmation is required. For duplicate data, if there are multiple identical records within a short period of time (such as within 1 second), only the first record is retained to avoid duplicate calculations.

[0024] For normalization, the Min-Max normalization method is adopted to map the data to the interval [0,1] for subsequent model calculations. Finally, the preprocessed average load data and temperature data in the server are respectively composed of an average load data sequence and a temperature sequence, where the last moment in the sequence is the data at the current moment. The sampling frequency of the average load data in the average load data sequence and the temperature data in the temperature sequence is once per second.

[0025] Step S2, use the autoregressive moving average algorithm to predict the average load data at the next moment of the current moment based on the average load data sequence, denoted as the predicted load; set different sampling scales, including high-frequency sampling scale, medium-frequency sampling scale, and low-frequency sampling scale.

[0026] Through the autoregressive moving average algorithm, an adaptive model of the server's load status is established, so that the dynamic distribution of the temperature threshold can be adjusted under different workloads. This can automatically relax the temperature detection threshold under high load to avoid false alarms, and appropriately lower the threshold in the case of low load but possible heat dissipation abnormalities, thereby improving the sensitivity of anomaly detection.

[0027] Since the change of the average load data may be affected by various factors such as sudden computing tasks and long-term trend fluctuations, the traditional autoregressive moving average algorithm can make predictions for a single feature. By using the CPU load condition to adjust the temperature threshold, the CPU load condition is first monitored, which can distinguish normal high load (such as batch computing tasks) from abnormal high load (such as overload caused by system resource exhaustion), thereby reducing the occurrence of false alarms.

[0028] A simple overview of the autoregressive moving average algorithm is as follows: The autoregressive part is based on past observations, and these historical data are linearly combined to predict the data at the next moment. The autoregressive model relies on the values at previous time points for prediction; The moving average part makes predictions based on the error terms at the current and previous times, correcting the errors generated by the autoregressive part. The moving average model uses a linear combination of error terms to improve predictions; By combining autoregression and moving average, it captures the linear trends and random fluctuations in the data, providing more accurate predictions for time series.

[0029] Set the data in the average load data sequence as: , where represents the average load data at the current time, that is, the average load data at time t. The autoregressive moving average algorithm (a well-known technique) combines the autoregressive and moving average parts to represent the internal law of time series data. The specific formula can be expressed as: , where, represents the predicted average load data at the next time, denoted as the predicted load. μ represents the constant term, with the initial value set to 0. p and q are the autoregressive order and moving average order of the model respectively, represents the autoregressive coefficient, indicating the influence of the average CPU load data at the previous p times on the predicted average load data. i represents the distance of the historical data from the current time. For example, the distance from time t - 1 to the current time t is one time; represents the moving average coefficient, indicating the influence of the error terms at the past q times on the current load; is the white noise error term, representing the unexplained part in the observed values, represents the white noise error at time t - j. Through the historical load data, first estimate the parameters and , and use the least squares method for parameter estimation to update the model parameters. Thus, the predicted average load data at the next time of the current time can be obtained, that is, the predicted load.

[0030] Since there are limitations when using the autoregressive moving average algorithm for prediction, there may be a certain difference between the predicted load and the real data. Therefore, on this basis, progressive multi-scale analysis needs to be introduced to deeply model the load-temperature change trend. The system can distinguish short-term fluctuations (such as temperature fluctuations caused by instantaneous computing tasks) from long-term trends (such as the decrease in heat dissipation efficiency due to the continuous high temperature of the server), so as to accurately identify real abnormal situations instead of misreporting instantaneous temperature peaks.

[0031] In this application, multi-scale analysis refers to setting different sampling scales. At different sampling scales, the average load data is collected and then analyzed. The sampling scales include high-frequency sampling scale, medium-frequency sampling scale, and low-frequency sampling scale. Preferably, in the embodiments of this application, the high-frequency sampling scale is once per second, the medium-frequency sampling scale is once per minute, and the low-frequency sampling scale is once per hour. The implementer can also adjust according to the actual situation.

[0032] Step S3, at the high-frequency sampling scale, calculate the high-frequency predicted load at the next moment based on the average load data before the next moment; at the medium-frequency sampling scale, calculate the medium-frequency predicted load at the next moment based on the average load data before the next moment; at the low-frequency sampling scale, calculate the low-frequency predicted load at the next moment based on the average load data before the next moment.

[0033] The multi-scale sampling of this solution is three sampling scales: high-frequency, medium-frequency, and low-frequency. Analyzing in multiple scales is to make the predicted data closer to the real value, and it can be compared with the average load data predicted by the autoregressive moving average algorithm. According to the comparison results, analyze whether there is a risk of CPU overload in the future average. The high-frequency sampling scale represents the load change with a relatively high instantaneous sampling frequency, the medium-frequency sampling scale represents the average load change within a short period of time, and the low-frequency sampling scale represents the cumulative load change over a relatively long period of time. Adjust the temperature threshold in combination with the CPU load conditions at multiple scales, predict the load level in the future for a period of time through the historical load sequence, and adjust the temperature threshold based on the prediction results.

[0034] The high-frequency sampling scale under multiple scales mainly represents the instantaneous fluctuation of the CPU load at each moment. At the high-frequency sampling scale, calculate the high-frequency predicted load at the next moment based on the average load data before the next moment. Set the high-frequency sampling window at the high-frequency sampling scale. The length of the window is N, and the reference value is 30s. It can also be expressed that the amount of data in the window is 30, that is, the time difference between two adjacent moments is 1 second. The high-frequency sampling window is the average load data of the 30 moments before the next moment.

[0035] The high-frequency sampling window slides with a preset step length. After sliding to the current moment, obtain the distance in time series between an average load data in the high-frequency sampling window and the average load data at the current moment, denoted as the time series distance; perform a negative correlation mapping on the product of the time series distance and the first adjustment coefficient using the exponential function with the natural constant as the base to obtain the weight of this average load data; perform weighted summation and averaging on the weights of each average load data in the high-frequency sampling window to obtain the high-frequency predicted load.

[0036] Its specific calculation formula is: , Among them, represents the high-frequency predicted load at the next moment after the current moment; N represents the number of average load data within the high-frequency sampling window; i represents the temporal distance from the moment corresponding to an average load data to the current moment in the time series, that is, the temporal distance, and the weight of the average load data closer to the t moment is greater. represents the weight of the average load data corresponding to the (t - i)th moment. represents the instantaneous CPU average load data, which can capture the instantaneous change of the average load data. α represents the first adjustment coefficient. The reference value is 0.5, and the reference value of the preset step size is 1 second.

[0037] Further, at the medium-frequency sampling scale, the medium-frequency predicted load at the next moment is calculated based on the average load data before the next moment. Specifically, a medium-frequency sampling window at the medium-frequency sampling scale is set, and the medium-frequency sampling window slides at a preset step size. After sliding to the current moment, the medium-frequency predicted load at the next moment is calculated using the average load data within the medium-frequency sampling window. The specific calculation formula is: , where represents the medium-frequency predicted load at the next moment; represents the mean of a preset number of average load data before the next moment; T represents the number of average load data within the medium-frequency sampling window. ,μ represents the window center, σ represents the smoothing degree coefficient; k and j represent the step indices within the medium-frequency sampling window; represents the (t - j)th average load data within the medium-frequency sampling window; represents the first data within the medium-frequency sampling window; e represents the natural constant.

[0038] represents the medium-frequency predicted load obtained by predicting the average load data of the CPU within the medium-frequency sampling window, which balances the noise interference in the short-time window; represents the Gaussian weight, representing the Gaussian weight distribution, which is used to control the smoothing effect of the time series, making the data points closer to the center have higher weights, while the weights of the data points farther away gradually decrease. represents the difference in the CPU average load data at the (t - j) moment weight of represents the weighted average difference within the medium-frequency sampling window, making the data points closer have higher weights. During the operation of the CPU, its load value changes at all times, and a residual term within a short time needs to be added on top of the reference average value to make the predicted value more accurate; It represents the mean of a preset number of average load data before the next moment. The value of the preset number is 5, which can also be called the mean of the CPU average load data of the last 5 sampling points within the window. Since the data obtained under the medium-frequency sampling scale can represent the fluctuations within a short period, the average value of the recent several points is taken as the reference value. The length of the medium-frequency sampling window is 30 minutes, that is, the window can accommodate 30 average load data under the medium-frequency sampling scale.

[0039] Finally, under the low-frequency sampling scale, the low-frequency predicted load for the next moment is calculated based on the average load data before the next moment.

[0040] Specifically, a low-frequency sampling window under the low-frequency sampling scale is set. The low-frequency sampling window slides with a preset step length. After sliding to the current moment, the average load data within the low-frequency sampling window is used to calculate the low-frequency predicted load for the next moment. The specific calculation formula is: , where, represents the low-frequency predicted load for the next moment; A represents the number of average load data within the low-frequency sampling window; and respectively represent the (t - j)-th and (t - j - 1)-th average load data within the low-frequency sampling window; β represents the scaling coefficient; e represents the natural constant; represents the average load data at the t-th moment within the low-frequency sampling window, that is, the average load data at the t-th moment (the current moment).

[0041] In the formula represents the CPU load situation under the predicted low-frequency sampling window, representing the cumulative CPU average load data within a longer window, which can represent the change trend of the long-term CPU average load data. represents the weight of the average load data of adjacent CPUs. The smaller the difference between the two, the more stable the data, so the greater the weight is given. represents the weighted difference of adjacent data at the t - j moment. After summing, and then adding the actual value at the current moment, it can represent that at this moment, the difference of the data points closer to the t-th moment is larger, making the closer data points more representative and more capable of accurately predicting the true data of the CPU average load data at the t + 1 moment.

[0042] Under the multi-scale analysis, by combining the short-term time window, the medium-term time window, and the long-term time window, the predicted values of the average load data for the next moment at the high-frequency sampling scale, the medium-frequency sampling scale, and the low-frequency sampling scale can be obtained, and then subsequent analysis can be carried out to obtain more accurate predicted data compared to the predicted load obtained by using the autoregressive moving average algorithm.

[0043] Step S4: Obtain the weights of the high-frequency predicted load, medium-frequency predicted load, and low-frequency predicted load based on the high-frequency predicted load and medium-frequency predicted load, and perform weighted summation to obtain the accurate predicted load at the next moment.

[0044] In step S3, the high-frequency predicted load, medium-frequency predicted load, and low-frequency predicted load at the high-frequency sampling scale, medium-frequency sampling scale, and low-frequency sampling scale are obtained. Further, it is necessary to fuse these three, and comprehensively combine the average load data predicted at different scales to obtain a more accurate predicted value.

[0045] Use the exponential function with the natural constant as the base to perform a negative correlation mapping on the difference between the high-frequency predicted load and the low-frequency predicted load to obtain a first mapping result. Multiply the first mapping result by a first scaling coefficient and add a preset value to obtain a first addition result. The reciprocal of the first addition result is the weight of the high-frequency predicted load; use the exponential function with the natural constant as the base to perform a negative correlation mapping on the difference between the low-frequency predicted load and the high-frequency predicted load to obtain a second mapping result. Multiply the second mapping result by the first scaling coefficient and add the preset value to obtain a second addition result. The reciprocal of the second addition result is the weight of the low-frequency predicted load; subtract the sum of the weights of the high-frequency predicted load and the low-frequency predicted load from the preset value to obtain the weight of the medium-frequency predicted load.

[0046] Use the weights of the high-frequency predicted load, medium-frequency predicted load, and low-frequency predicted load to perform weighted summation on the high-frequency predicted load, medium-frequency predicted load, and low-frequency predicted load to obtain the accurate predicted load at the next moment.

[0047] Its specific calculation formula is: , where, represents the accurate predicted load at the next moment; , and represent the high-frequency predicted load, low-frequency predicted load, and medium-frequency predicted load respectively.

[0048] represents the weight of the high-frequency predicted load, represents the weight of the low-frequency predicted load, represents the weight of the medium-frequency predicted load, e represents the natural constant, γ represents the first scaling coefficient, and the value of the preset value is 1. Implementers can adjust it according to the actual situation; the reference value of the first scaling coefficient is 3. When and are very close or equal, the weights of the high-frequency predicted load and the low-frequency predicted load can be made around 0.25. Because at the medium-frequency sampling scale, the collected data can better illustrate the short-term changes in the data, and the described short-term trend changes are more representative. Therefore, its weight should be the largest, close to 0.5.

[0049] and respectively represent the first mapping result and the second mapping result. When the high-frequency prediction load is greater than the low-frequency prediction load, the first mapping result is less than the second mapping result. At this time, the weight of the high-frequency prediction load is greater than the weight of the low-frequency prediction load, indicating that there is data mutation. Amplifying its weight makes subsequent predictions more accurate. If the weight of the low-frequency prediction load is less than the weight of the high-frequency prediction load, it means that the trend (short-term trend) of the short-term data collected at the high-frequency sampling scale may be noise or misjudgment. Therefore, the weight of its low-frequency prediction load (long-term window) is increased, so that the finally synthesized predicted value is closer to the true value and the sensitivity of the subsequent temperature threshold is improved.

[0050] The multi-scale weighting here is to make the predicted CPU load more accurate. The core purpose is to balance short-term fluctuations and long-term trends, ensuring that the system can not only quickly respond to load changes but also avoid the impact of short-term anomalies on predictions. The introduction of the exponential decay function can further ensure smoothness, enabling information at different scales to dynamically adjust the contribution degree and improving prediction stability and accuracy.

[0051] Step S5: Calculate the temperature threshold for the next moment using the predicted load and the precise predicted load; predict the temperature data for the next moment based on the temperature sequence, denoted as the predicted temperature; if the predicted temperature is greater than or equal to the temperature threshold, issue a reminder.

[0052] Compare the predicted load obtained by using the autoregressive moving average algorithm with the precise predicted load obtained by using multi-scale sampling frequency analysis to obtain the temperature threshold for the next moment.

[0053] Calculate the temperature threshold for the next moment using the predicted load and the precise predicted load. Specifically, multiply the difference between the precise predicted load and the predicted load by the first adjustment coefficient, and perform a negative correlation mapping using the exponential function with the natural constant as the base to obtain the third mapping result; obtain the reciprocal of the sum of the third mapping result and the preset value, obtain the difference between the reciprocal and the first adjustment coefficient, and multiply it by the second adjustment coefficient to obtain the multiplication result; the sum of the multiplication result and the temperature threshold at the current moment is the temperature threshold for the next moment.

[0054] The specific calculation formula is: , where represents the temperature threshold for the next moment, represents the temperature threshold at the current moment. An initial temperature threshold will be set at the beginning, and it will be continuously updated at each moment as the model processes; λ represents the second adjustment coefficient, with a reference value of 0.1, and α represents the first adjustment coefficient, with a reference value of 0.5. and respectively represent the accurate predicted load and the predicted load at the next moment, where the value range of the temperature threshold is within 0 to 1.

[0055] The third mapping result is , at greater than when, belongs to the average load data at the predicted t+1 moment under multi-scale sampling, which can better reflect the average load data of the real CPU. represents the average load data of the CPU predicted in a short time. being large indicates that the average load data of the real CPU may be large (the CPU volatility at the t+1 moment is relatively strong, resulting in the autoregressive moving average data being unable to predict this large fluctuation trend), so there is a risk of CPU overload at the t+1 moment, and it is necessary to lower the temperature threshold and optimize heat dissipation in advance. If at less than or equal to when, it indicates that the CPU load tends to be stable or decreasing, and the temperature threshold can be appropriately relaxed to reduce the energy waste caused by excessive cooling.

[0056] Thus, the temperature threshold at the next moment is obtained. Further, it is necessary to predict the temperature data at the next moment, and then monitor and warn the CPU ambient temperature of the computing power data center.

[0057] At the high-frequency sampling scale, the high-frequency sampling window slides at a preset step length. After sliding to the current moment, in the temperature sequence, the difference between the temperature data at the first moment in the high-frequency sampling window and the temperature data at the current moment is obtained, and compared with the number of temperature data in the high-frequency sampling window to obtain the change rate of the average difference; adding the temperature data at the current moment to the average difference gives the temperature data at the next moment, denoted as the predicted temperature.

[0058] Its specific calculation formula is: , where, represents the temperature data at the next moment, that is, the predicted temperature; represents the temperature data at the current moment; represents the first temperature data in the high-frequency sampling window, and N represents the number of temperature data in the high-frequency sampling window. is the change rate of the average difference. Thus, the predicted temperature can be obtained.

[0059] If the predicted temperature is greater than or equal to the temperature threshold, a reminder is issued. The staff adds coolant according to the reminder, or increases the power of the cooling system to increase the circulation speed of the coolant to cope with the increase in temperature and ensure the normal operation of the computing power data center.

[0060] It should be noted that in the embodiments of the present application, when the window moves at each scale, it slides with the sampling unit at the high-frequency sampling scale as the preset step length, that is, 1 second. The implementer can adjust according to the actual situation to update the temperature threshold per second to ensure the normal operation of the computing power center. At the same time, for the high-frequency sampling scale, the medium-frequency sampling scale, and the low-frequency sampling scale, as well as their respective corresponding windows, experimental adjustments can be made according to the error between the actual value and the predicted value and updated in a timely manner to ensure the robustness of the model provided by the present application, improve the accuracy of prediction, and make the average load data predicted according to the multi-scale analysis closer to the true value.

[0061] It should be noted that the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes specific embodiments of this specification. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0062] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.

[0063] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent monitoring method for an AI computing power data center, characterized in that: The method includes: Collect the average load data of the CPU and the temperature data in the server and pre-process them to obtain the average load data series and the temperature series; The autoregressive moving average algorithm is used to predict the average load data sequence based on the current time to obtain the average load data of the next moment, which is recorded as the predicted load; different sampling scales are set, including high-frequency sampling scale, medium-frequency sampling scale and low-frequency sampling scale; Under the high-frequency sampling scale, the high-frequency predicted load at the next moment is calculated based on the average load data before the next moment; under the medium-frequency sampling scale, the medium-frequency predicted load at the next moment is calculated based on the average load data before the next moment; under the low-frequency sampling scale, the low-frequency predicted load at the next moment is calculated based on the average load data before the next moment; Obtain weights of the high-frequency predicted load, the medium-frequency predicted load, and the low-frequency predicted load according to the high-frequency predicted load and the medium-frequency predicted load, and perform weighted summation to obtain the accurate predicted load at the next moment; The temperature threshold at the next moment is calculated using the predicted load and the accurate predicted load; the temperature data at the next moment is predicted based on the temperature sequence and recorded as the predicted temperature; if the predicted temperature is greater than or equal to the temperature threshold, a reminder is issued.

2. According to claim 1, the intelligent monitoring method of an AI computing power data center is characterized in that: The collecting and preprocessing of the average load data of the CPU and the temperature data in the server to obtain the average load data sequence and the temperature sequence includes: Preprocessing mainly includes data cleaning and normalization. Data cleaning includes missing value completion, outlier identification and duplicate data processing. The preprocessed average load data and temperature data in the server are respectively composed of an average load data sequence and a temperature sequence.

3. According to claim 1, the intelligent monitoring method of an AI computing power data center is characterized in that: The method of calculating the high-frequency predicted load at the next moment based on the average load data before the next moment under the high-frequency sampling scale includes: A high-frequency sampling window is set under a high-frequency sampling scale, and the high-frequency sampling window slides with a preset step size. After sliding to the current moment, the time series distance between an average load data in the high-frequency sampling window and the average load data at the current moment is obtained, which is recorded as the time series distance; an exponential function with a natural constant as the base is used to negatively correlate the product of the time series distance and the first adjustment coefficient to obtain the weight of the average load data; the weights of the average load data in the high-frequency sampling window are weighted and summed and averaged to obtain the high-frequency predicted load.

4. The intelligent monitoring method of an AI computing power data center according to claim 1 is characterized in that: The method of calculating the intermediate frequency predicted load at the next moment based on the average load data before the next moment under the intermediate frequency sampling scale includes: The intermediate frequency sampling window under the intermediate frequency sampling scale is set, and the intermediate frequency sampling window slides with a preset step size. After sliding to the current moment, the intermediate frequency predicted load at the next moment is calculated using the average load data in the intermediate frequency sampling window; the intermediate frequency predicted load calculation formula is specifically as follows: , in, Indicates the medium frequency predicted load at the next moment; represents the mean of the preset number of average load data before the next moment; T represents the number of average load data in the intermediate frequency sampling window, , μ represents the window center, σ represents the smoothing coefficient; k and j represent the step index within the intermediate frequency sampling window; Represents the tjth average load data in the intermediate frequency sampling window; Represents the first data in the intermediate frequency sampling window; e represents a natural constant.

5. The intelligent monitoring method of an AI computing power data center according to claim 1 is characterized in that: The step of calculating the low-frequency predicted load at the next moment based on the average load data before the next moment under the low-frequency sampling scale includes: The low-frequency sampling window under the low-frequency sampling scale is set, and the low-frequency sampling window slides with a preset step size. After sliding to the current moment, the low-frequency predicted load at the next moment is calculated using the average load data in the low-frequency sampling window; the specific formula for calculating the low-frequency predicted load is: , in, represents the low-frequency predicted load at the next moment; A represents the number of average load data in the low-frequency sampling window; and They represent the tj-th and tj-1-th average load data in the low-frequency sampling window respectively; β represents the scaling factor; e represents the natural constant; Represents the tth average load data in the low-frequency sampling window.

6. The intelligent monitoring method of an AI computing power data center according to claim 1 is characterized in that: The step of obtaining the weights of the high-frequency predicted load, the medium-frequency predicted load, and the low-frequency predicted load according to the high-frequency predicted load and the medium-frequency predicted load includes: An exponential function with a natural constant as the base is used to negatively map the difference between the high-frequency predicted load and the low-frequency predicted load to obtain a first mapping result, the first mapping result is multiplied by a first scaling factor and added to a preset value to obtain a first addition result, and the inverse of the first addition result is the weight of the high-frequency predicted load; an exponential function with a natural constant as the base is used to negatively map the difference between the low-frequency predicted load and the high-frequency predicted load to obtain a second mapping result, the second mapping result is multiplied by the first scaling factor and added to a preset value to obtain a second addition result, and the inverse of the second addition result is the weight of the low-frequency predicted load; the weight of the medium-frequency predicted load is obtained by subtracting the sum of the weight of the high-frequency predicted load and the weight of the low-frequency predicted load from the preset value.

7. The intelligent monitoring method of an AI computing power data center according to claim 1 is characterized in that: The method of calculating the temperature threshold at the next moment by using the predicted load and the accurate predicted load includes: Multiply the difference between the precise predicted load and the predicted load and the first adjustment coefficient, and use an exponential function with a natural constant as the base for negative correlation mapping to obtain a third mapping result; calculate the inverse of the third mapping result plus the preset value, calculate the difference between the inverse and the first adjustment coefficient, and multiply it with the second adjustment coefficient to obtain the multiplication result; the sum of the multiplication result and the temperature threshold at the current moment is the temperature threshold at the next moment.

8. The intelligent monitoring method of an AI computing power data center according to claim 1 is characterized in that: The step of predicting the temperature data at the next moment according to the temperature sequence, recorded as predicted temperature, includes: At the high-frequency sampling scale, the high-frequency sampling window slides with a preset step size. After sliding to the current moment, in the temperature sequence, the difference between the temperature data at the first moment in the high-frequency sampling window and the temperature data at the current moment is obtained, and compared with the number of temperature data in the high-frequency sampling window, the rate of change of the average difference is obtained; the temperature data at the current moment is added to the average difference to obtain the temperature data at the next moment, which is recorded as the predicted temperature.

Citation Information

Patent Citations

  • Data center task temperature prediction and scheduling method based on RBF neural network

    CN109375994A

  • Short-term power load hybrid prediction method

    CN115438833A

  • Virus detection method and device, electronic equipment and storage medium

    CN116938496A

  • Transformer end electricity consumption monitoring method and system

    CN119165240A

  • Method and apparatus for temperature control, and device and storage medium

    WO2024160122A1