An energy efficiency evaluation method for an artificial intelligence data center

By constructing a dynamic feature library and an AI adaptive weight model, and combining LSTM baseline prediction and isolated forest thresholding, the real-time performance and scenario adaptability issues of energy efficiency evaluation in existing technologies are solved, and intelligent energy efficiency management of AI data centers is realized.

CN120973622BActive Publication Date: 2026-02-03BEIJING QINGYUAN CHUANGYAN TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511071950.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2026-02-03
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing technologies cannot update energy efficiency assessments in real time, cannot adapt to changes in the correlation of energy efficiency factors caused by business switching in AI data centers, and lack real-time performance and scenario adaptability.

Method used

A dynamic feature library is constructed, an AI adaptive weight model is introduced, and LSTM baseline prediction and isolated forest thresholding are combined. The entire process is iteratively optimized through online learning to achieve dynamic energy efficiency evaluation and scenario adaptability.

Benefits of technology

It has achieved real-time energy efficiency evaluation and scenario adaptability, improved the accuracy of anomaly monitoring and the pertinence of optimization suggestions, and formed a fully intelligent energy efficiency management system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973622B_ABST
    Figure CN120973622B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data centers, and discloses an energy efficiency evaluation method of an artificial intelligence data center. By constructing a dynamic feature library and an AI adaptive weight model, the problem that a static evaluation method in the prior art is disconnected with the real-time running state of a data center is solved. The dynamic feature library collects and associates business scene labels and space-time coupling data in real time, provides a scene-based basis for the AI adaptive weight model, changes weight adjustment from passive response to index mutation in the prior art to active adaptation to business scene changes, effectively captures the nonlinear influence of dynamic factors such as server load fluctuation and environmental regulation on energy efficiency, makes the energy efficiency evaluation result more suitable for the actual running state of the data center, adapts to the correlation change of energy efficiency factors caused by business switching of the AI data center, realizes scene-based adaptive adjustment of the weight through a deep learning model, and solves the limitation that fixed weight distribution ignores differences between different business scenes in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data center technology, specifically a method for evaluating the energy efficiency of artificial intelligence data centers. Background Technology

[0002] According to the center, energy efficiency assessment is a core link in improving energy utilization efficiency, but existing technologies have significant limitations:

[0003] A search revealed that the invention patent with announcement number CN119250591A discloses an incremental energy efficiency evaluation system and method based on medium and low voltage distribution networks. Although the evaluation is optimized through "indicator mutation weight adjustment", the weight adjustment depends on the triggering of indicator mutations and does not achieve real-time dynamic updates. It is not directly related to real-time data collection and does not associate business scenario labels (such as deep learning training and real-time inference) with the data. The weight adjustment is irrelevant to the scenario and is difficult to adapt to the changes in the correlation of energy efficiency factors caused by business switching in AI data centers.

[0004] A search revealed that the invention patent with announcement number CN113449949A discloses an evaluation method for electrical equipment based on a weighted rating combination algorithm. This method uses static parameters such as equipment purchase cost and expected lifespan as the main evaluation criteria. The weight allocation is fixed and the update cycle is long. It does not dynamically adjust the static parameters in conjunction with real-time operating data (such as instantaneous power consumption and ambient temperature and humidity). The static parameters lack real-time performance, the LSTM prediction does not incorporate business tags, and the static parameters are completely disconnected from the model update. This makes it impossible to form a data feedback link and respond to real-time changes in the operating status of the equipment in a timely manner.

[0005] While the energy consumption prediction method based on the attention mechanism LSTM enhances the ability to capture trends in time series data, it does not integrate real-time environmental variables and business scenario characteristics. It relies solely on historical energy consumption data for prediction, resulting in insufficient adaptability to changes in energy efficiency baselines under different business scenarios. Summary of the Invention

[0006] The purpose of this invention is to provide an energy efficiency evaluation method for artificial intelligence data centers to solve the problems mentioned in the background art.

[0007] The core innovation of this invention lies in the collaborative closed-loop of four modules: a dynamic feature library provides scenario-based data support, an AI adaptive weight model enables scenario-based weight adjustment, LSTM baseline prediction and isolated forest thresholding construct a precise anomaly detection mechanism, and finally, the entire process is iteratively optimized through online learning. This collaborative mechanism solves the technical challenges of traditional methods in simultaneously addressing dynamism, scenario adaptability, and long-term effectiveness.

[0008] To achieve the above objectives, the present invention provides the following technical solution: an energy efficiency evaluation method for artificial intelligence data centers, the specific steps of which are as follows:

[0009] S1. Construct a dynamic feature library: Utilize distributed sensor networks and edge computing nodes to collect data such as power consumption and computing power of IT equipment (servers, GPU clusters, etc.) in real time, as well as information such as the energy efficiency ratio of cooling systems and ambient temperature and humidity. At the same time, associate it with tags such as business type and time period to form a spatiotemporally coupled dynamic feature library, which provides comprehensive and real-time basic data support for the entire process of subsequent model training, weight adjustment, etc.

[0010] S2. Introduce AI adaptive weight model: Based on the dynamic feature library constructed in step S1, use deep reinforcement learning technology to train the adaptive weight model. This model can receive scene labels and data distribution in the feature library in real time, automatically adjust the weight ratio of each energy efficiency factor, and output dynamic weight.

[0011] S3. Deep Learning Energy Efficiency Baseline Prediction: Using historical data from the dynamic feature library in step S1 as training samples, an energy efficiency baseline prediction model is constructed using a Long Short-Term Memory (LSTM) network. The model, combined with real-time data, can predict the energy efficiency baseline value for the next 15 minutes and output the energy efficiency baseline value.

[0012] S4. Dynamic energy efficiency value calculation: Input the real-time data collected in step S1 into the AI ​​adaptive weight model trained in step S2 to obtain the dynamic weight of each energy efficiency factor. Then, combine it with the baseline value predicted in step S3 to calculate the dynamic energy efficiency value and output the dynamic energy efficiency value.

[0013] S5. Multidimensional anomaly threshold setting: Based on the historical anomaly data and business requirements in the dynamic feature library in step S1, the isolated forest algorithm is used to generate dynamic anomaly thresholds for each energy efficiency factor. The thresholds will be automatically adjusted according to the scene changes, and will work together with the dynamic energy efficiency value obtained in step S4 to accurately determine whether the current energy efficiency is in an abnormal state.

[0014] S6. Generate dynamic optimization suggestions: When the dynamic energy efficiency value in step S4 exceeds the multidimensional anomaly threshold set in step S5, the AI ​​model will call the relevant data in the feature library of step S1, analyze the correlation between abnormal factors, generate targeted suggestions in combination with the historical optimization case library, push them to the management system, and at the same time feed back the processing data to the dynamic feature library of step S1.

[0015] S7. Model Iteration Update: Every 24 hours, the dynamic energy efficiency value from step S4, the anomaly handling results from step S6, and the effects of optimization suggestions are fed back to the dynamic feature library in step 1. The AI ​​adaptive weight model in step S2 and the energy efficiency baseline prediction model in step S3 are updated through online learning algorithms to continuously improve the evaluation accuracy.

[0016] Preferably, the specific steps for constructing the dynamic feature library in step S1 are as follows:

[0017] S11. Multi-dimensional data acquisition: Utilizing distributed sensor networks and edge computing nodes, comprehensive and real-time data capture is performed on the data center, including operating parameters such as instantaneous power consumption, computing power output, and memory usage of IT equipment; energy-related data such as the energy efficiency ratio of the cooling system and the conversion efficiency of the UPS; and environmental information such as the temperature and humidity field distribution in the computer room. At the same time, scenario tags such as business type (e.g., deep learning training, real-time inference) and time period characteristics (peak / off-peak load) are associated with these data to ensure the contextual attributes of the data.

[0018] S12. Data Foundation Support: Through the above collection and correlation operations, a dynamic feature library with spatiotemporal coupling characteristics is formed, providing rich samples for subsequent model training, providing real-time basis for weight adjustment, and providing historical and real-time data support for baseline prediction, so that all energy efficiency evaluation steps can be carried out smoothly and the comprehensiveness and real-time nature of the evaluation process can be guaranteed.

[0019] Preferably, the specific steps for introducing the AI ​​adaptive weight model in step S2 are as follows:

[0020] S21. Model training basis: Based on the dynamic feature library constructed in step S1, the adaptive weight model is trained using deep reinforcement learning (DRL) technology. During the training process, the model can receive business scenario labels and dynamically changing data distribution transmitted by the feature library in real time. By continuously learning the inherent laws of data in different scenarios, it gradually acquires the ability to intelligently adjust the weights of each energy efficiency factor.

[0021] S22. Dynamic weight output: The trained model can automatically adjust the weight ratio of each energy efficiency factor for different scenarios. For example, in a deep learning training scenario, the model will automatically increase the weight of the factor in the overall evaluation based on the correlation data between GPU power consumption and computing power output in the feature library. This directly provides key parameters for the accurate calculation of dynamic energy efficiency values ​​in subsequent steps, ensuring the scenario adaptability of energy efficiency evaluation.

[0022] Preferably, the specific steps for deep learning energy efficiency baseline prediction in step S3 are as follows:

[0023] S31. Model Construction and Training: Select historical data from the dynamic feature library in Step 1 as training samples, and use a Long Short-Term Memory (LSTM) network to construct an energy efficiency baseline prediction model. With its ability to process time series data, the LSTM network can effectively mine the pattern of energy efficiency changes over time in historical data. Through training with a large number of samples, the model can learn and predict energy efficiency trends.

[0024] S32. Baseline Prediction Application: During operation, this model combines real-time updated load trends and environmental changes from the feature library to predict the data center energy efficiency baseline value for the next 15 minutes. This predicted baseline value becomes the key reference benchmark for calculating the dynamic energy efficiency value in step S4, providing a scientific basis for judging whether the real-time energy efficiency is within a reasonable range.

[0025] Preferably, the specific steps for calculating the dynamic energy efficiency value in step S4 are as follows:

[0026] S41. Data and Model Integration: The dynamic data collected in real time in step S1 is input into the AI ​​adaptive weight model trained in step S2. The model outputs the dynamic weights of each energy efficiency factor in the current scenario based on the current data distribution and scenario label. Then, these dynamic weights are integrated with the energy efficiency baseline value predicted by the deep learning model in step S3.

[0027] S42. Energy Efficiency Value Calculation and Function: Using the dynamic energy efficiency formula, the integrated parameters are calculated to obtain the real-time dynamic energy efficiency value, which is the core indicator for measuring the current energy efficiency level of the data center. It is directly used in the anomaly monitoring link in step S5 and becomes an important basis for judging whether energy efficiency is abnormal.

[0028] Dynamic energy efficiency value calculation formula:

[0029]

[0030] In the formula:

[0031] DEE stands for Dynamic Energy Efficiency Value, a key indicator for measuring the real-time energy efficiency of a data center under different times and business scenarios.

[0032] F i These are real-time factor values, derived from the dynamic feature library construction step, where dynamic data is collected in real time through a distributed sensor network and edge computing nodes. This data includes information such as the instantaneous power consumption, computing power output, memory usage of IT equipment, energy efficiency ratio of cooling systems, UPS conversion efficiency, and environmental temperature and humidity field distribution. The specific values ​​reflect the actual operating status of various energy efficiency-related factors of the data center at the current moment.

[0033] Fi includes, but is not limited to, the following 12 key energy efficiency factors:

[0034] (1) IT equipment: instantaneous power consumption of servers, GPU computing power output, memory utilization, network transmission energy consumption, and storage device power consumption;

[0035] (2) Energy system: refrigeration system energy efficiency ratio, UPS conversion efficiency, battery backup system loss, power supply and distribution line loss;

[0036] (3) Environmental parameters: computer room temperature, humidity, and airflow speed;

[0037] W i The dynamic weights are generated by introducing an AI adaptive weight model. The adaptive weight model, trained based on deep reinforcement learning (DRL), can receive business scenario labels (such as deep learning training and real-time inference) from a dynamic feature library in real time, as well as dynamic data distribution. It automatically adjusts the relative importance of each energy efficiency factor (such as device power consumption and environmental parameters) to the overall energy efficiency in the current scenario, and its output value is the dynamic weight.

[0038] B represents the energy efficiency baseline, obtained through a deep learning-based energy efficiency baseline prediction step. Using historical data from a dynamic feature library as training samples, a Long Short-Term Memory (LSTM) network is employed to construct an energy efficiency baseline prediction model. This model combines real-time data from the dynamic feature library, such as real-time load trends and environmental changes, to predict the data center's energy efficiency benchmark value for the next 15 minutes, serving as a reasonable reference level for data center energy efficiency under current operating conditions.

[0039] Preferably, the specific steps for setting the multidimensional anomaly threshold in step S5 are as follows:

[0040] S51. Threshold generation basis: Based on the historical abnormal data in the dynamic feature library in step S1, and combined with the specific energy efficiency requirements of different business scenarios, the isolated forest algorithm is used to automatically generate the dynamic abnormal threshold of each energy efficiency factor. The isolated forest algorithm can define a reasonable fluctuation range for different energy efficiency factors in different scenarios by accurately identifying abnormal data.

[0041] S52. Function of the Threshold System: The generated dynamic anomaly thresholds are adjusted in real time according to changes in the scenario. For example, in a high-load server scenario, the normal fluctuation range of power consumption is larger, and the corresponding threshold will be higher than in a low-load scenario. These dynamic thresholds together constitute a multi-dimensional anomaly monitoring network, which works in conjunction with the dynamic energy efficiency value calculated in step S4. Through comparative analysis, it accurately determines whether the energy efficiency of the current data center exceeds the normal range, thus improving the accuracy of anomaly monitoring.

[0042] Preferably, the specific steps for generating dynamic optimization suggestions in step S6 are as follows:

[0043] S61. Anomaly Analysis Trigger: When the dynamic energy efficiency value calculated in step S4 exceeds the multidimensional anomaly threshold set in step S5, the AI ​​model will immediately start the anomaly analysis process. At this time, the model will call the historical and real-time data related to the anomaly in the dynamic feature library of step S1 to deeply analyze the correlation between the anomaly factors, such as exploring whether the abnormal energy consumption of the air conditioner is related to the heat dissipation efficiency of the server, so as to provide a basis for the generation of subsequent optimization suggestions.

[0044] S62. Suggestion Generation and Feedback: After completing the anomaly correlation analysis, the model combines successful handling experience of similar scenarios in the historical optimization case library to generate targeted optimization suggestions, such as "adjusting the room temperature to a specific range under the current GPU load can reduce air conditioning energy consumption." These suggestions will be pushed to the data center management system in real time. At the same time, relevant data in the anomaly handling process will be fed back to the dynamic feature library in step S1 to further enrich the content of the feature library.

[0045] Preferably, the specific steps of model iterative update in step S7 are as follows:

[0046] S71. Update data source: Every 24 hours, the system will summarize the dynamic energy efficiency value calculated in step S4, the results of anomaly handling in step S6, and the effects of actual application of optimization suggestions, and feed these data back to the dynamic feature library in step 1. These feedback data contain the latest energy efficiency status and optimization information.

[0047] S72. Model Optimization Closed Loop: Utilizing online learning algorithms, based on new data fed back to the feature library, the AI ​​adaptive weight model in step S2 and the energy efficiency baseline prediction model in step S3 are continuously updated. By continuously learning new scenario features and data patterns, the evaluation accuracy of the model is gradually improved, enabling it to continuously adapt to the dynamic changes of the data center and ensuring the long-term effectiveness of the evaluation method.

[0048] LSTM model update: The online learning algorithm prioritizes fine-tuning the output layer weights (adjustment range is ±3% of the initial weights), and the hidden layer weights are updated every 48 hours based on the incremental new data, preserving the learning results of historical time series patterns;

[0049] AI Adaptive Weight Model (PPO) Update: The state space retains 6 core features (GPU power consumption, server load, ambient temperature, cooling energy efficiency ratio, business label, and time period label), and the weight adjustment range of the action space dynamically shrinks / expands according to the distribution of scene data (expanded to [-0.15, +0.15] for high-fluctuation scenes).

[0050] The beneficial effects of this invention are as follows:

[0051] 1. This invention solves the problem of disconnect between static evaluation methods (such as fixed parameters or adjustments triggered by sudden changes in indicators) and the real-time operating status of data centers by constructing a dynamic feature library and an AI adaptive weight model. The dynamic feature library collects and associates business scenario labels and spatiotemporally coupled data in real time, providing a scenario-based basis for the AI ​​adaptive weight model. This transforms weight adjustment from "passively responding to sudden changes in indicators" to "actively adapting to changes in business scenarios," effectively capturing the nonlinear impact of dynamic factors such as server load fluctuations and environmental adjustments on energy efficiency. This makes the energy efficiency evaluation results more consistent with the actual operating status of data centers and adapts to changes in the correlation of energy efficiency factors caused by business switching (such as switching between deep learning training and real-time inference) in AI data centers.

[0052] 2. This invention achieves scenario-based adaptive adjustment of weights through a deep learning model, overcoming the limitation of fixed weight allocation in existing technologies that ignores differences in different business scenarios. In different scenarios such as deep learning training and real-time inference, the model can automatically identify the correlation changes of factors such as GPU power consumption and network transmission energy consumption and dynamically allocate weights. This avoids ignoring scenario differences with fixed weights, reduces the deviation between evaluation results and actual energy efficiency, and provides more targeted quantitative support for energy efficiency optimization in different business scenarios.

[0053] 3. This invention constructs a fully intelligent energy efficiency management system through LSTM baseline prediction, isolated forest dynamic threshold, and closed-loop iteration mechanism, breaking through the limitations of existing technologies that passively handle anomalies and lack advanced early warning and data feedback links. LSTM baseline prediction integrates scenario-based historical and real-time data, solving the problem of insufficient scenario adaptability of traditional prediction that only relies on historical data. The linkage between isolated forest dynamic threshold and dynamic energy efficiency value enables advanced anomaly identification. Combined with historical optimization case library and model iteration update, a closed loop of "anomaly early warning - precise optimization - model evolution" is formed, replacing empirical rules and static parameters, and meeting the needs of proactive energy efficiency management and long-term effectiveness in AI data centers. Attached Figure Description

[0054] Figure 1 This is a flowchart of the energy efficiency evaluation method for the artificial intelligence data center of the present invention. Detailed Implementation

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] like Figure 1As shown in the figure, this embodiment of the invention provides an energy efficiency evaluation method for artificial intelligence data centers. The specific steps of the method are as follows:

[0057] S1. Construct a dynamic feature library: Utilize distributed sensor networks and edge computing nodes to collect data such as power consumption and computing power of IT equipment (servers, GPU clusters, etc.) in real time, as well as information such as the energy efficiency ratio of cooling systems and ambient temperature and humidity. At the same time, associate it with tags such as business type and time period to form a spatiotemporally coupled dynamic feature library, which provides comprehensive and real-time basic data support for the entire process of subsequent model training, weight adjustment, etc.

[0058] S2. Introduce AI adaptive weight model: Based on the dynamic feature library constructed in step S1, use deep reinforcement learning technology to train the adaptive weight model. This model can receive scene labels and data distribution in the feature library in real time, automatically adjust the weight ratio of each energy efficiency factor, and output dynamic weight.

[0059] S3. Deep Learning Energy Efficiency Baseline Prediction: Using historical data from the dynamic feature library in step S1 as training samples, an energy efficiency baseline prediction model is constructed using a Long Short-Term Memory (LSTM) network. The model, combined with real-time data, can predict the energy efficiency baseline value for the next 15 minutes and output the energy efficiency baseline value.

[0060] S4. Dynamic energy efficiency value calculation: Input the real-time data collected in step S1 into the AI ​​adaptive weight model trained in step S2 to obtain the dynamic weight of each energy efficiency factor. Then, combine it with the baseline value predicted in step S3 to calculate the dynamic energy efficiency value and output the dynamic energy efficiency value.

[0061] S5. Multidimensional anomaly threshold setting: Based on the historical anomaly data and business requirements in the dynamic feature library in step S1, the isolated forest algorithm is used to generate dynamic anomaly thresholds for each energy efficiency factor. The thresholds will be automatically adjusted according to the scene changes, and will work together with the dynamic energy efficiency value obtained in step S4 to accurately determine whether the current energy efficiency is in an abnormal state.

[0062] S6. Generate dynamic optimization suggestions: When the dynamic energy efficiency value in step S4 exceeds the multidimensional anomaly threshold set in step S5, the AI ​​model will call the relevant data in the feature library of step S1, analyze the correlation between abnormal factors, generate targeted suggestions in combination with the historical optimization case library, push them to the management system, and at the same time feed back the processing data to the dynamic feature library of step S1.

[0063] S7. Model Iteration Update: Every 24 hours, the dynamic energy efficiency value from step S4, the anomaly handling results from step S6, and the effects of optimization suggestions are fed back to the dynamic feature library in step 1. The AI ​​adaptive weight model in step S2 and the energy efficiency baseline prediction model in step S3 are updated through online learning algorithms to continuously improve the evaluation accuracy.

[0064] The specific steps for constructing the dynamic feature library in step S1 are as follows:

[0065] S11. Multi-dimensional data acquisition: Utilizing distributed sensor networks and edge computing nodes, comprehensive and real-time data capture is performed on the data center, including operating parameters such as instantaneous power consumption, computing power output, and memory usage of IT equipment; energy-related data such as the energy efficiency ratio of the cooling system and the conversion efficiency of the UPS; and environmental information such as the temperature and humidity field distribution in the computer room. At the same time, scenario tags such as business type (e.g., deep learning training, real-time inference) and time period characteristics (peak / off-peak load) are associated with these data to ensure the contextual attributes of the data.

[0066] Spatiotemporal coupling is achieved through 'timestamp + spatial coordinates': the time dimension uses millisecond-level timestamps to record the data acquisition time, and the spatial dimension assigns a unique cabinet number to each sensor node (such as 'Area A-Column 3-Cabinet 5'); when data is associated, data from multiple nodes under the same timestamp are aggregated into 'regional energy efficiency snapshots' according to spatial coordinates, and snapshots with different timestamps form a time series, thus achieving spatiotemporal coupling;

[0067] The distributed sensor network uses three temperature and humidity sensors (model: SHT30) and one power consumption sensor (model: DS2438) deployed per rack, and transmits data via the LoRa protocol (433MHz band); the edge computing nodes use NVIDIA Jetson Nano (4-core ARM CPU, 4GB memory) to be responsible for real-time data preprocessing.

[0068] Example of scenario-based association: Under the scenario label of 'Deep Learning Training', the system automatically associates the power consumption data of Area A-3-5 cabinets (GPU cluster) with the energy efficiency ratio data of the cooling system in the same area through the same timestamp, forming a scenario-based data pair of 'GPU load-cooling efficiency', providing scenario-based samples for the weight adjustment of S2;

[0069] S12. Data Foundation Support: Through the above collection and correlation operations, a dynamic feature library with spatiotemporal coupling characteristics is formed, providing rich samples for subsequent model training, providing real-time basis for weight adjustment, and providing historical and real-time data support for baseline prediction, so that all energy efficiency evaluation steps can be carried out smoothly and the comprehensiveness and real-time nature of the evaluation process can be guaranteed.

[0070] The data preprocessing of the dynamic feature library includes: smoothing instantaneous power consumption data using the moving average method (window size 5 minutes); filling missing values ​​using the Lagrange interpolation method; and performing Min-Max normalization (mapping to the 0-1 interval) on all data to ensure data consistency in the input model.

[0071] The specific steps for introducing the AI ​​adaptive weight model in step S2 are as follows:

[0072] S21. Model training basis: Based on the dynamic feature library constructed in step S1, the adaptive weight model is trained using deep reinforcement learning technology (specifically the PPO algorithm). The model receives business scenario labels and dynamic data distribution transmitted from the feature library in real time, and forms the ability to adjust weights by learning the inherent laws of data in different scenarios.

[0073] The state space of deep reinforcement learning contains 6 features (GPU power consumption, server load, ambient temperature, cooling energy efficiency ratio, business type label, and time period label); the action space is the weight adjustment amount of each energy efficiency factor (range [-0.1, +0.1]).

[0074] The reward function of the PPO algorithm is defined as: R = 1 - (DEE - B) / B, where DEE is the dynamic energy efficiency value and B is the baseline value. When DEE is closer to B, the reward R is closer to 1, which guides the model to learn a weight allocation strategy that keeps energy efficiency stable within a reasonable range.

[0075] The training parameters for the PPO algorithm are: 1000 iterations, 256 data points sampled per iteration, clipping parameter ε = 0.2, discount factor γ = 0.99, Adam optimizer, and learning rate 0.0003.

[0076] The PPO algorithm's editing parameter ε = 0.2 was determined through comparative experiments. This value can ensure the stability of the model while avoiding evaluation oscillations caused by excessive weight adjustment. When ε = 0.1, the adjustment is too slow, and when ε = 0.3, the stability decreases.

[0077] S22. Dynamic weight output: The trained model automatically adjusts the weight ratio of each energy efficiency factor for different scenarios. For example, in a deep learning training scenario, the model automatically increases the weight ratio of the factor based on the correlation data between GPU power consumption and computing power output, providing scenario-based parameters for dynamic energy efficiency value calculation.

[0078] The dynamic weights range from 0 to 1, and the sum of the weights of all energy efficiency factors is 1. Normalization is achieved through the Softmax function to ensure the rationality of DEE calculation.

[0079] The specific steps for deep learning energy efficiency baseline prediction in step S3 are as follows:

[0080] S31. Model Construction and Training: Historical data from the dynamic feature library in step S1 is selected as training samples. A long short-term memory network (LSTM) is used to construct an energy efficiency baseline prediction model. The network contains 3 hidden layers with 64 neurons in each layer. The activation function is ReLU. The pattern of energy efficiency changing over time is explored through training.

[0081] The LSTM model takes 12-dimensional feature data from the past 30 minutes as input (sampled every 5 minutes) and outputs a baseline value for the next 15 minutes. Dropout (ratio 0.2) is used between the input layer and the hidden layer to prevent overfitting.

[0082] The loss function of the LSTM model is mean squared error (MSE), the optimizer is Adam, the initial learning rate is 0.001, the training iteration is 50 rounds, and the batch size of each round is 64.

[0083] S32. Baseline Prediction Application: The model combines real-time updated load trend and environmental change data from the feature library to predict the energy efficiency baseline value for the next 15 minutes, which serves as a reference benchmark for dynamic energy efficiency value calculation.

[0084] The specific steps for calculating the dynamic energy efficiency value in step S4 are as follows:

[0085] S41. Data and Model Integration: The dynamic data collected in real time in step S1 is input into the AI ​​adaptive weight model trained in step S2. The model outputs the dynamic weights of each energy efficiency factor in the current scenario based on the current data distribution and scenario label. Then, these dynamic weights are integrated with the energy efficiency baseline value predicted by the deep learning model in step S3.

[0086] S42. Energy Efficiency Value Calculation and Function: Using the dynamic energy efficiency formula, the integrated parameters are calculated to obtain the real-time dynamic energy efficiency value, which is the core indicator for measuring the current energy efficiency level of the data center. It is directly used in the anomaly monitoring link in step S5 and becomes an important basis for judging whether energy efficiency is abnormal.

[0087] Dynamic energy efficiency value calculation formula:

[0088]

[0089] In the formula:

[0090] DEE stands for Dynamic Energy Efficiency Value, a key indicator for measuring the real-time energy efficiency of a data center under different times and business scenarios.

[0091] F i These are real-time factor values, derived from dynamic data collected in real-time through distributed sensor networks and edge computing nodes during the construction of the dynamic feature library. The real-time factor values ​​Fi include, but are not limited to: instantaneous power consumption, computing power output, and transmission rate of IT equipment (servers, GPU clusters, storage devices, network switches); energy efficiency ratio of cooling systems and UPS conversion efficiency; ambient temperature and humidity, and airflow speed—covering energy efficiency-related factors across the entire chain of data center IT equipment, energy systems, and environmental control. The specific values ​​reflect the actual operating status of various energy efficiency-related factors in the data center at the current moment.

[0092] W i The dynamic weights are generated by introducing an AI adaptive weight model. The adaptive weight model, trained based on deep reinforcement learning (DRL), can receive business scenario labels (such as deep learning training and real-time inference) from a dynamic feature library in real time, as well as dynamic data distribution. It automatically adjusts the relative importance of each energy efficiency factor (such as device power consumption and environmental parameters) to the overall energy efficiency in the current scenario, and its output value is the dynamic weight.

[0093] B represents the energy efficiency baseline, obtained through a deep learning-based energy efficiency baseline prediction step. Using historical data from a dynamic feature library as training samples, a Long Short-Term Memory (LSTM) network is employed to construct an energy efficiency baseline prediction model. This model combines real-time data from the dynamic feature library, such as real-time load trends and environmental changes, to predict the data center's energy efficiency benchmark value for the next 15 minutes, serving as a reasonable reference level for data center energy efficiency under current operating conditions.

[0094] The specific steps for setting the multidimensional anomaly threshold in step S5 are as follows:

[0095] S51. Threshold generation basis: Based on the historical abnormal data and business requirements in the dynamic feature library in step S1, the isolated forest algorithm (with the number of trees set to 100) is used to generate the dynamic abnormal threshold of each energy efficiency factor, and to define the reasonable fluctuation range of different energy efficiency factors in a scenario-based manner.

[0096] The anomaly score threshold for isolated forests is dynamically set based on the business scenario: the threshold is reduced by 10% in deep learning training scenarios (high load fluctuation) to adapt to a larger range of energy consumption fluctuations; the threshold is kept at the baseline value in real-time inference scenarios (low load stability) to balance the accuracy and efficiency of anomaly identification.

[0097] The way to integrate isolated forests with business scenarios is as follows: For the 'deep learning training' scenario (high GPU load), the number of trees is increased by 20% to improve the accuracy of anomaly identification for highly volatile data, and the anomaly score threshold is reduced by 10% (because the energy consumption fluctuation range is larger in this scenario); For the 'real-time inference' scenario (stable operation under low load), the number of trees is reduced by 10% to improve efficiency, and the anomaly score threshold remains unchanged.

[0098] Initial labeling rules for historical anomaly data: When the real-time value of a certain energy efficiency factor exceeds the 95% confidence interval of its 30-day historical data, it is marked as an anomaly; subsequent manual review and correction are used to form a labeled anomaly dataset for training the Isolation Forest.

[0099] The maximum tree depth of the isolated forest is set to 20. Two features are randomly selected for sample division. The anomaly score threshold is set to 0.7 (values ​​higher than this are considered anomalies).

[0100] S52. Function of the threshold system: The dynamic abnormal threshold is adjusted in real time according to the scene changes (e.g., the power consumption threshold is higher in the high load scene of the server than in the low load scene), and compared with the dynamic energy efficiency value in step S4 to determine whether the current energy efficiency is in an abnormal state.

[0101] The isolated forest anomaly score baseline threshold of 0.1 is determined based on the annotation of 1000 normal / abnormal data in typical scenarios. This value corresponds to a balance between anomaly identification accuracy and recall. A value below 0.1 will lead to an increase in false positive rate, while a value above 0.1 will lead to an increase in false negative rate.

[0102] The specific steps for generating dynamic optimization suggestions in step S6 are as follows:

[0103] S61. Anomaly Analysis Trigger: When the dynamic energy efficiency value in step S4 exceeds the multidimensional anomaly threshold in step S5, the AI ​​model calls the relevant data in the feature library in step S1, calculates the correlation coefficient between anomaly factors (|r|>0.8 is determined to be a strong correlation), and analyzes the correlation (such as the correlation between abnormal air conditioning energy consumption and server heat dissipation efficiency).

[0104] Multi-factor anomaly priority determination: By calculating the 'energy efficiency impact weight' of the anomaly factors (determined by the dynamic weight of S2), the correlation between the top two factors with the highest weights is analyzed first (e.g., if the GPU power consumption weight is 0.4 and the air conditioner energy efficiency ratio weight is 0.3, then the correlation between the two is analyzed first).

[0105] S62. Suggestion generation and feedback: Based on the historical optimization case library (categorized by business scenario and anomaly type), generate targeted optimization suggestions (such as "adjust the computer room temperature to a specific range to reduce air conditioning energy consumption"), push them to the management system, and simultaneously feed the processed data back to the dynamic feature library of step S1.

[0106] The historical optimization case library is constructed as follows: Each case contains 7 fields—business scenario tag (e.g., 'deep learning training - high GPU load'), abnormal factor combination (e.g., 'GPU power consumption > threshold + air conditioning energy efficiency ratio < threshold'), optimization measures (e.g., 'adjust the computer room temperature to 22℃ + limit GPU computing power peak'), energy efficiency value before implementation, energy efficiency value after implementation, effect improvement rate (%), and timestamp; the case library is updated every 24 hours with model iteration, and new cases must meet the requirement of effect improvement rate ≥ 5%, while old cases without matching scenarios within 3 months are deleted; during retrieval, a two-dimensional matching of 'scenario tag + abnormal factor cosine similarity' is used, and cases are called when the similarity is ≥ 0.9;

[0107] The scene label + anomaly factor feature vector contains 7 dimensions: business type (one-hot encoding), time period type (peak / valley), proportion of anomaly factor 1 exceeding the threshold, proportion of anomaly factor 2 below the threshold, ambient temperature deviation, load rate, and historical optimization effect level (level 1-5). Cosine similarity is calculated based on this vector.

[0108] Historical optimization effectiveness levels are quantified by 'effectiveness improvement rate': ≥15% is level 5, 10%-15% is level 4, 5%-10% is level 3, 2%-5% is level 2, and <2% is level 1, ensuring the operability of similarity calculation.

[0109] The specific steps of model iterative update in step S7 are as follows:

[0110] S71. Update data source: Every 24 hours, summarize the dynamic energy efficiency value of step S4, the anomaly handling results of step S6, and the optimization suggestion effect data, and feed them back to the dynamic feature library of step S1.

[0111] The optimization suggestion effect data is quantified by the 'deviation rate between dynamic energy efficiency value and baseline value': when the deviation rate decreases (energy efficiency improves), the data is used as a positive sample, and the model updates the relevant weight parameters more frequently; when the deviation rate increases (energy efficiency deteriorates), the data is used as a negative sample, and the model updates the relevant weight parameters less frequently; when the deviation rate does not change significantly, the data is only used for model stability training and does not change the parameter adjustment range.

[0112] S72, Model Optimization Closed Loop: Using online learning algorithms, based on new feedback data, update the AI ​​adaptive weight model in step S2 and the energy efficiency baseline prediction model in step S3 to continuously adapt to the dynamic changes in the data center.

[0113] The online learning algorithm uses incremental gradient descent with a learning rate of 0.001. The training sample size for each batch is 50% of the new data added within 24 hours. For the AI ​​adaptive weight model, the focus is on updating the weight parameters associated with the new scene labels. For the LSTM baseline prediction model, only the output layer weights are fine-tuned to retain the learning results of historical time series patterns.

[0114] Online learning parameter update logic: When the optimization suggestion improves energy efficiency by ≥5%, the model fixes the weight parameters (such as GPU power consumption weight) in that scenario as the baseline value; if the effect does not meet expectations, the parameter adjustment range is reduced by 10% through incremental gradient descent until convergence.

[0115] When the difference in distribution between new data and historical data (calculated by KL divergence) exceeds 0.1, the learning rate is automatically increased to 0.002, and the weight of new data samples is increased to 70% to quickly adapt to sudden changes in the scenario.

[0116] The formula for calculating KL divergence is: Where P represents the new data distribution, Q represents the historical data distribution, and the sample window is the data collected in the last 2 hours (sampling frequency 1 minute / time);

[0117] The KL divergence threshold of 0.1 was verified through 50 sets of scene mutation experiments (including 20 sets of 'deep learning training → real-time inference' switching and 30 sets of 'peak → trough load' switching). When the KL divergence ≤ 0.1, the model's adaptation error to the new scene is < 10%; when it exceeds 0.1, the error suddenly increases to > 15%. Therefore, this is used as the threshold for judging scene mutation.

[0118] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0119] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for evaluating the energy efficiency of an artificial intelligence data center, characterized in that: The specific steps of this method are as follows: S1. Construct a dynamic feature library: Collect power consumption and computing power data of IT equipment in real time, as well as energy efficiency ratio of cooling system and environmental temperature and humidity information, and associate them with business type and time period tags to form a spatiotemporally coupled dynamic feature library. S2. Introduce an AI adaptive weight model: Based on the dynamic feature library constructed in step S1, use deep reinforcement learning technology to train an adaptive weight model to automatically adjust the weight ratio of each energy efficiency factor. S3. Deep learning energy efficiency baseline prediction: Using the historical data from step S1 as training samples, a long short-term memory network is used to build an energy efficiency baseline prediction model to predict the energy efficiency baseline value. S4. Calculation of dynamic energy efficiency value: Input the real-time data collected in step S1 into the AI ​​adaptive weight model trained in step S2 to obtain the dynamic weight of each energy efficiency factor, and then combine it with the baseline value predicted in step S3 to calculate the dynamic energy efficiency value. S5. Multidimensional anomaly threshold setting: Based on the historical anomaly data and business requirements in step S1, the isolated forest algorithm is used to generate dynamic anomaly thresholds for each energy efficiency factor to determine whether the current energy efficiency is in an abnormal state. S6. Generate dynamic optimization suggestions: When the dynamic energy efficiency value in step S4 exceeds the multidimensional anomaly threshold set in step S5, the AI ​​model calls the relevant data in the feature library in step S1 and generates targeted suggestions in combination with the historical optimization case library. S7. Model Iteration Update: Feed back the dynamic energy efficiency value, anomaly handling results and optimization suggestions to the dynamic feature library, and update the AI ​​adaptive weight model of step S2 and the energy efficiency baseline prediction model of step S3 through online learning algorithms.

2. The energy efficiency evaluation method for an artificial intelligence data center according to claim 1, characterized in that: The specific steps for constructing the dynamic feature library in step S1 are as follows: S11, Multi-dimensional data acquisition: Utilize distributed sensor networks and edge computing nodes to capture data from the data center in a comprehensive and real-time manner, and associate business types and time period characteristic scene tags with these data to ensure the contextual attributes of the data; S12. Data Foundation Support: Through data collection and correlation operations, a dynamic feature library with spatiotemporal coupling characteristics is formed, providing samples for subsequent model training, providing real-time basis for weight adjustment, and providing historical and real-time data support for baseline prediction.

3. The energy efficiency evaluation method for an artificial intelligence data center according to claim 1, characterized in that: The specific steps for introducing the AI ​​adaptive weight model in step S2 are as follows: S21. Model training basis: Based on the dynamic feature library constructed in step S1, the adaptive weight model is trained using deep reinforcement learning technology. The model can receive business scenario labels and dynamically changing data distribution transmitted by the feature library in real time. S22. Dynamic weight output: The trained model automatically adjusts the weight ratio of each energy efficiency factor for different scenarios, providing key parameters for the accurate calculation of dynamic energy efficiency values ​​in step S4.

4. The energy efficiency evaluation method for an artificial intelligence data center according to claim 1, characterized in that: The specific steps for deep learning energy efficiency baseline prediction in step S3 are as follows: S31. Model Construction and Training: Historical data from the dynamic feature library in step S1 is selected as training samples. A long short-term memory network is used to construct an energy efficiency baseline prediction model, which can effectively explore the pattern of energy efficiency changes over time in historical data. S32. Baseline Prediction Application: When the model is running, it combines the load trend and environmental change data updated in real time in the feature library to predict the data center energy efficiency baseline value for the next 15 minutes and predict the energy efficiency baseline value.

5. The energy efficiency evaluation method for an artificial intelligence data center according to claim 1, characterized in that: The specific steps for calculating the dynamic energy efficiency value in step S4 are as follows: S41. Data and Model Integration: The dynamic data collected in real time in step S1 is input into the AI ​​adaptive weight model trained in step S2. The model outputs the dynamic weights of each energy efficiency factor in the current scenario based on the current data distribution and scenario label. Then, these dynamic weights are integrated with the energy efficiency baseline value predicted by the deep learning model in step S3. S42. Energy Efficiency Value Calculation and Function: Using the dynamic energy efficiency formula, the integrated parameters are calculated to obtain the real-time dynamic energy efficiency value, which is directly used in the anomaly monitoring stage in step 5.

6. The energy efficiency evaluation method for an artificial intelligence data center according to claim 1, characterized in that: The specific steps for setting the multidimensional anomaly threshold in step S5 are as follows: S51. Threshold generation basis: Based on the historical abnormal data in the dynamic feature library in step S1, and combined with the specific energy efficiency requirements of different business scenarios, the isolated forest algorithm is used to automatically generate the dynamic abnormal threshold of each energy efficiency factor, and define the reasonable fluctuation range of different energy efficiency factors in different scenarios. S52. Function of the threshold system: The generated dynamic anomaly thresholds are adjusted in real time as the scene changes. These dynamic thresholds together constitute a multi-dimensional anomaly monitoring network. They are compared and analyzed with the dynamic energy efficiency value calculated in step S4 to determine whether the energy efficiency of the current data center exceeds the normal range.

7. The energy efficiency evaluation method for an artificial intelligence data center according to claim 1, characterized in that: The specific steps for generating dynamic optimization suggestions in step S6 are as follows: S61, Anomaly Analysis Trigger: When the dynamic energy efficiency value calculated in step S4 exceeds the multidimensional anomaly threshold set in step S5, the AI ​​model immediately starts the anomaly analysis process to analyze the correlation between anomaly factors. S62. Recommendation Generation and Feedback: After completing the anomaly correlation analysis, the model combines successful handling experience of similar scenarios in the historical optimization case library to generate targeted optimization recommendations, which are pushed to the data center management system in real time.

8. The energy efficiency evaluation method for an artificial intelligence data center according to claim 1, characterized in that: The specific steps for model iterative update in step S7 are as follows: S71. Update data source: Every 24 hours, the system will summarize the dynamic energy efficiency value calculated in step S4, the results of anomaly handling in step S6, and the effect data after the actual application of optimization suggestions, and feed these data back to the dynamic feature library in step S1. S72. Model optimization closed loop: Using online learning algorithms, based on the new data fed back to the feature library, the AI ​​adaptive weight model in step S2 and the energy efficiency baseline prediction model in step S3 are continuously updated.

Citation Information

Patent Citations

  • Electric equipment evaluation method based on empowerment rating combination algorithm

    CN113449949A

  • Energy-saving loss-reduction increment energy efficiency evaluation system and method based on medium and low voltage power distribution network

    CN119250591A

  • Intelligent task alarm rule self-learning method and system based on support priority

    CN119441832A

  • Multi-scene visual algorithm evaluation method based on electric power artificial intelligence platform

    CN119445185A