AI-based IDC data center energy efficiency optimization and fault prediction system and method

By building an AI-based IDC data center energy efficiency optimization and fault prediction system, the system analyzes the health status of equipment in real time and dynamically adjusts energy efficiency optimization strategies, solving the problem of the disconnect between energy efficiency optimization and equipment health status in existing technologies, and achieving a balance between energy efficiency improvement and equipment reliability.

CN122152655APending Publication Date: 2026-06-05JILIN XINMANFENG E-COMMERCE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JILIN XINMANFENG E-COMMERCE CO LTD
Filing Date
2026-03-05
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

In existing technologies, data centers do not consider the synergistic relationship between equipment health status and fault warning when pursuing energy efficiency optimization, which may lead to the implementation of aggressive energy-saving strategies and affect the reliability of equipment operation.

Method used

An AI-based IDC data center energy efficiency optimization and fault prediction system is constructed, including a data acquisition module, a fault prediction module, an energy efficiency optimization module, and a collaborative control module. By collecting real-time data on the energy consumption of IT equipment and cooling system parameters in the data center, the system uses AI models to analyze the health status of the equipment, generates graded early warning signals, and dynamically adjusts energy efficiency optimization strategies to avoid over-adjustment of equipment.

Benefits of technology

It achieves the goal of improving energy efficiency while ensuring the reliability of data center operation. Through the collaborative control module, it avoids the conflict between energy efficiency optimization and fault early warning, and ensures the proactive avoidance of equipment health status and the differentiated execution of optimization strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122152655A_ABST
    Figure CN122152655A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, in particular to an IDC data center energy efficiency optimization and fault prediction system and method based on AI, which comprises the following modules: a data acquisition module, which is used for collecting IT equipment energy consumption data, refrigeration system operation parameters and server state log data of a data center in real time; a fault prediction module, which is used for outputting fault early warning information and equipment health state indexes, and performing hierarchical processing on the fault early warning information; an energy efficiency optimization module, which is used for determining a real-time power utilization efficiency value of the data center, generating an energy efficiency optimization strategy, and dynamically adjusting the optimization amplitude of the energy efficiency optimization strategy according to the grade of the fault early warning signal; and a collaborative control module, which is used for performing intersection comparison on the early warning equipment list and the to-be-adjusted equipment list, determining the issuing strategy of a final control instruction according to a comparison result, and evaluating and feeding back the effect of the collaborative control. The application improves the reliability and energy efficiency level of data center operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an AI-based IDC data center energy efficiency optimization and fault prediction system and method. Background Technology

[0002] With the rapid development of cloud computing, big data, and artificial intelligence technologies, the scale and computing power demand of global data centers are growing exponentially, highlighting their strategic importance as core infrastructure of the digital economy. However, while supporting massive data processing and complex computing tasks, data centers also face unprecedented energy consumption pressures and reliability challenges. Cooling systems, as key auxiliary facilities ensuring the stable operation of IT equipment, typically account for a large proportion of the total energy consumption of data centers, leading to generally high energy efficiency, high operating costs, and significant environmental pressure. Simultaneously, to cope with the continuously growing computing demands, the scale and complexity of data center infrastructure are constantly increasing, raising the risk of server, storage, and network equipment failures. Traditional operation and maintenance models relying on manual inspections and reactive responses are not only inefficient but also struggle to identify potential systemic risks in massive amounts of monitoring data in a timely manner. Sudden failures causing business interruptions can result in huge economic losses.

[0003] Chinese Patent Application Publication No. CN114970358A discloses a data center energy efficiency optimization method and system based on reinforcement learning. The data center energy efficiency optimization system of this invention consists of a data integration and management system, an IDC environmental monitoring system, and a DRL (Data Center Reliability Management) center model system. To address the problems of low energy efficiency and high cost in existing data center energy consumption optimization methods, an offline reinforcement learning approach is adopted. An initial control strategy is designed, and data on the system under the initial strategy is recorded, including states, actions, and rewards. A value function is learned using a standard reinforcement learning / function simulation combination to estimate the cumulative expected reward corresponding to a specific action in a given state, thereby obtaining the predicted PUE (Power Usage Effectiveness) value.

[0004] However, the existing technology has the following problems: the energy efficiency optimization process of the existing technology is based only on energy consumption data and does not consider the synergistic relationship between equipment health status and fault warning. This may lead to the implementation of aggressive energy-saving strategies when there are potential risks in the equipment, which will aggravate equipment wear and tear and affect operational reliability. Summary of the Invention

[0005] To address this, the present invention provides an AI-based IDC data center energy efficiency optimization and fault prediction system and method, which overcomes the problems in the prior art where the dynamic synergy between equipment health status and energy efficiency optimization is not considered, leading to the neglect of equipment failure risks when pursuing energy efficiency optimization, or the blind execution of optimization strategies during equipment warning periods, which exacerbates equipment wear and affects the reliability of data center operation.

[0006] To achieve the above objectives, this invention provides an AI-based IDC data center energy efficiency optimization and fault prediction system, comprising: The data acquisition module is used to collect real-time data on the energy consumption of IT equipment in the data center, the operating parameters of the cooling system, and the status log data of the server. The fault prediction module is connected to the data acquisition module and is used to analyze the health status of the equipment based on the status log data through the AI ​​fault prediction model, output fault warning information and equipment health status index, and perform graded processing on the fault warning information based on the comparison result of the equipment health status index and the preset health threshold to generate fault warning signals of different levels. An energy efficiency optimization module, connected to the data acquisition module, is used to determine the real-time power utilization efficiency value of the data center based on the energy consumption data of the IT equipment and the operating parameters of the cooling system, compare the real-time power utilization efficiency value with the preset energy efficiency target value, generate an energy efficiency optimization strategy through an AI energy efficiency optimization model based on the comparison result, and dynamically adjust the optimization range of the energy efficiency optimization strategy according to the level of the fault warning signal output by the fault prediction module. The collaborative control module is connected to the fault prediction module and the energy efficiency optimization module respectively. It is used to compare the intersection of the list of warning devices in the fault warning signal and the list of devices to be adjusted in the energy efficiency optimization strategy, determine the final control command issuance strategy based on the comparison result, and evaluate and provide feedback on the effect of collaborative control based on the comparison result of the actual temperature change value within a preset time period after the final control command is issued and the preset temperature change threshold.

[0007] Furthermore, the fault prediction module determines the level of the fault warning signal based on the comparison between the equipment health status index and a preset health threshold, wherein, When the device health status index is greater than or equal to the first preset health threshold, a low-level fault warning signal is generated. When the device health status index is less than the first preset health threshold and greater than or equal to the second preset health threshold, an intermediate fault warning signal is generated. When the device health status index is less than the second preset health threshold, an advanced fault warning signal is generated.

[0008] Furthermore, the equipment health status index is determined by load deviation and temperature deviation.

[0009] Furthermore, when generating low-to-medium level or medium-level fault warning signals, the fault prediction module corrects the level of the fault warning signal based on the comparison between the temperature change rate of the corresponding device within a preset time period and a preset temperature change rate threshold. Specifically, when the temperature change rate exceeds the preset temperature change rate threshold, the level of the fault warning signal is increased by one level.

[0010] Furthermore, the energy efficiency optimization module dynamically adjusts the optimization magnitude of the energy efficiency optimization strategy based on the level of the fault warning signal output by the fault prediction module, wherein: When a low-level fault warning signal is received, maintain the initial optimization level; When a medium-level fault warning signal is received, the optimization range of the relevant warning equipment is compressed according to the first ratio. When a high-level fault warning signal is received, the optimization range of the warning equipment is compressed according to the second ratio.

[0011] Furthermore, when the difference between the real-time energy utilization efficiency value and the preset energy efficiency target value is greater than a preset deviation threshold, the energy efficiency optimization module reverses the first ratio or the second ratio according to the magnitude of the utilization efficiency difference. The utilization efficiency difference is the difference between the real-time energy utilization efficiency value and the preset energy efficiency target value, and the correction magnitude of the first ratio or the second ratio is positively correlated with the utilization efficiency difference.

[0012] Furthermore, the collaborative control module determines the final control command issuance strategy based on the comparison between the list of warning devices in the fault warning signal and the list of devices to be adjusted in the energy efficiency optimization strategy, including: If the number of intersecting devices is zero, then the final control command is directly issued according to the energy efficiency optimization strategy described above; If the number of intersecting devices is greater than zero and less than a preset threshold, the adjustment instructions involving intersecting devices in the energy efficiency optimization strategy will be transferred to non-intersecting devices before being issued. If the number of overlapping devices is greater than or equal to a preset threshold, the energy efficiency optimization strategy will be suspended and the energy efficiency optimization module will be notified to regenerate the energy efficiency optimization strategy.

[0013] Furthermore, when transferring adjustment commands, the collaborative control module determines the feasibility of the transfer operation based on a comparison between the current temperature value of the target non-intersecting device and a preset temperature threshold. If the current temperature value of the target non-intersecting device is less than the preset temperature threshold, the transfer operation is deemed feasible, and the adjustment command is executed to transfer the device. If the current temperature value of the target non-intersecting device is greater than or equal to the preset temperature threshold, the transfer operation is determined to be infeasible, triggering the overall shelving of the energy efficiency optimization strategy.

[0014] Furthermore, the collaborative control module evaluates the effectiveness of the collaborative control based on the comparison between the actual temperature change value within a preset time period after the final control command is issued and the preset temperature change threshold, and adjusts the preset health threshold when the actual temperature change value is less than the preset temperature change threshold.

[0015] This invention also provides an AI-based method for energy efficiency optimization and fault prediction in IDC data centers, comprising: Step S1: Collect real-time data on the energy consumption of IT equipment in the data center, the operating parameters of the cooling system, and the status log data of the server. Step S2: Based on the status log data, analyze the health status of the equipment using an AI fault prediction model, output fault warning information and equipment health status index, and classify the fault warning information according to the comparison results between the equipment health status index and the preset health threshold to generate fault warning signals of different levels. Step S3: Determine the real-time power utilization efficiency value of the data center based on the energy consumption data of the IT equipment and the operating parameters of the cooling system, compare the real-time power utilization efficiency value with the preset energy efficiency target value, and generate an energy efficiency optimization strategy through the AI ​​energy efficiency optimization model based on the comparison result; Step S4: Dynamically adjust the optimization magnitude of the energy efficiency optimization strategy according to the level of the fault warning signal output by the fault prediction module; Step S5: Compare the list of warning devices in the fault warning signal with the list of devices to be adjusted in the energy efficiency optimization strategy, and determine the final control command issuance strategy based on the comparison result. Step S6: Based on the comparison between the actual temperature change value within the preset time period after the final control command is issued and the preset temperature change threshold, evaluate and provide feedback on the effect of the coordinated control.

[0016] Compared with existing technologies, the advantages of this invention lie in its ability to effectively solve the problem of the disconnect between energy efficiency optimization and equipment health status in traditional data center management solutions by constructing a collaborative working mechanism between a fault prediction module and an energy efficiency optimization module. The fault prediction module analyzes the equipment health status in real time based on an AI model and generates tiered early warning signals. The energy efficiency optimization module dynamically adjusts the magnitude of its optimization strategy according to the level of the early warning signal, thus proactively avoiding equipment health risks. When equipment is in a sub-healthy state, the system automatically reduces the optimization magnitude related to that equipment, avoiding accelerated equipment aging due to excessive adjustments to load or cooling parameters, thereby improving energy efficiency while ensuring the reliability of data center operation.

[0017] Furthermore, this invention uses a collaborative control module to compare the intersection of the early warning device list and the device list to be adjusted, and formulates differentiated control command issuance strategies based on the comparison results, avoiding direct conflicts between energy efficiency optimization commands and fault early warning devices. When there are a few overlapping devices, the system intelligently transfers optimization commands to other healthy devices, ensuring that the energy efficiency optimization target is partially achieved; when the number of overlapping devices is too large, the system automatically pauses the issuance of optimization strategies and triggers strategy replanning, preventing blindly executing optimization operations during large-scale device early warning periods, which could lead to a sharp increase in operational risks.

[0018] Furthermore, this invention utilizes a feedback mechanism for evaluating the effectiveness of the collaborative control module to adaptively adjust the preset health threshold based on actual temperature changes after the control command is issued. This enables the system to continuously optimize the early warning criteria based on historical control performance. This closed-loop feedback design gives the system's collaborative control strategy the ability to learn and continuously evolve, adapting to the impact of long-term operating characteristics such as aging data center equipment and environmental changes, ensuring that the system maintains good control performance throughout its lifecycle. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the module connections of the IDC data center energy efficiency optimization and fault prediction system according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating the process of determining the fault warning signal level in an embodiment of the present invention; Figure 3 A flowchart illustrating the strategy for determining the final control command issuance in an embodiment of the present invention; Figure 4 This is a flowchart of the IDC data center energy efficiency optimization and fault prediction method according to an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0021] It should be noted that the data in this embodiment are all derived from a comprehensive analysis and evaluation of historical test data and corresponding historical test results from the three months prior to this test. Those skilled in the art will understand that the determination of the above-mentioned parameters for any single item in this invention can be achieved by selecting the value with the highest percentage based on the data distribution as the preset standard parameter, using weighted summation to obtain the value as the preset standard parameter, substituting each historical data point into a specific formula and using the value obtained from that formula as the preset standard parameter, or other selection methods, as long as the invention can clearly define different specific situations in the single-item judgment process through the obtained values.

[0022] Please see Figure 1 As shown, it is a schematic diagram of the module connection of the IDC data center energy efficiency optimization and fault prediction system according to an embodiment of the present invention; The data acquisition module is used to collect real-time data on the energy consumption of IT equipment in the data center, the operating parameters of the cooling system, and the status log data of the server. The fault prediction module is connected to the data acquisition module and is used to analyze the health status of the equipment based on the status log data through the AI ​​fault prediction model, output fault warning information and equipment health status index, and perform graded processing on the fault warning information based on the comparison result of the equipment health status index and the preset health threshold to generate fault warning signals of different levels. An energy efficiency optimization module, connected to the data acquisition module, is used to determine the real-time power utilization efficiency value of the data center based on the energy consumption data of the IT equipment and the operating parameters of the cooling system, compare the real-time power utilization efficiency value with the preset energy efficiency target value, generate an energy efficiency optimization strategy through an AI energy efficiency optimization model based on the comparison result, and dynamically adjust the optimization range of the energy efficiency optimization strategy according to the level of the fault warning signal output by the fault prediction module. The collaborative control module is connected to the fault prediction module and the energy efficiency optimization module respectively. It is used to compare the intersection of the list of warning devices in the fault warning signal and the list of devices to be adjusted in the energy efficiency optimization strategy, determine the final control command issuance strategy based on the comparison result, and evaluate and provide feedback on the effect of collaborative control based on the comparison result of the actual temperature change value within a preset time period after the final control command is issued and the preset temperature change threshold.

[0023] Specifically, the data acquisition module includes current and voltage sensors deployed inside the intelligent power distribution unit to collect real-time power and cumulative energy consumption data of IT equipment; return air temperature sensors, chilled water inlet and outlet water temperature sensors, and compressor frequency sensors deployed inside the server room air conditioner to collect operating parameters of the refrigeration system; temperature sensors deployed near the server motherboard and key chips to collect core temperature data of the CPU, GPU, and memory chips; and a baseboard management controller connected to the server via an intelligent platform management interface protocol to collect hardware health status data of the server, including but not limited to CPU temperature, motherboard temperature, fan speed, power supply voltage, and hardware-level error logs.

[0024] Specifically, the AI ​​fault prediction model can be implemented using a time-series prediction model, such as a Long Short-Term Memory (LSTM) network, a gated recurrent unit (GRU), or a hybrid model combining a convolutional neural network and a long short-term memory network (CNN-LSTM). As long as it can learn the evolution law of the device's operating status based on historical state log data and output a quantitative index reflecting the health of the device, the embodiments of this invention do not specifically limit the specific network structure, training algorithm, and implementation method of the AI ​​fault prediction model.

[0025] Specifically, the AI ​​energy efficiency optimization model can be implemented using a reinforcement learning model, such as a Deep Q-Network (DQN), Deep Deterministic Policy Gradient (DDPG), or Proximal Policy Optimization (PPO). It uses the deviation between the real-time energy utilization efficiency value and the preset energy efficiency target value as the input to the reward function. Through continuous interaction with the environment, it learns the optimal control strategy and outputs energy efficiency optimization strategies, including adjustments to the cooling system setpoint, power consumption limits for IT equipment, and airflow organization optimization. This embodiment of the invention does not specifically limit the specific algorithm type, network structure, or training method of the AI ​​energy efficiency optimization model, as long as it can generate an executable energy efficiency optimization strategy based on the current operating state.

[0026] Please see Figure 2 As shown, it is a flowchart for determining the fault warning signal level in an embodiment of the present invention; Specifically, the fault prediction module determines the level of the fault warning signal based on a comparison between the equipment health status index and a preset health threshold. When the device health status index is greater than or equal to the first preset health threshold, a low-level fault warning signal is generated. When the device health status index is less than the first preset health threshold and greater than or equal to the second preset health threshold, an intermediate fault warning signal is generated. When the device health status index is less than the second preset health threshold, an advanced fault warning signal is generated.

[0027] Specifically, the equipment health status index is determined by load deviation and temperature deviation. Load deviation refers to the degree of deviation between the current load value of the equipment and its historical baseline load value, while temperature deviation refers to the degree of deviation between the current core temperature value of the equipment and its historical baseline temperature value. The equipment health status index is calculated using the following formula: , In the formula, This refers to the equipment health status index. As the first weighting coefficient, This is the current load value. This is the historical baseline load value. This is the second weighting coefficient. This is the current core temperature value. Using historical reference temperature values, in this embodiment of the invention, The value is 0.6. The value is set to 0.4. The historical baseline load value and historical baseline temperature value can be taken as the average load and average temperature of the equipment at the same time point in the past week. For newly launched equipment or equipment without historical data, the factory nominal value of the same model equipment or the average value of similar equipment in the same period in the computer room can be taken as the baseline value. There is no specific limitation.

[0028] In this embodiment of the invention, the first preset health threshold is 0.85 and the second preset health threshold is 0.6. However, the above values ​​are not limited to these values, and those skilled in the art can adjust the above values ​​according to actual needs.

[0029] Specifically, when generating a low-to-medium level fault warning signal or a medium level fault warning signal, the fault prediction module corrects the level of the fault warning signal based on the comparison between the temperature change rate of the corresponding device within a preset time period and a preset temperature change rate threshold. Specifically, when the temperature change rate exceeds the preset temperature change rate threshold, the level of the fault warning signal is increased by one level.

[0030] In this embodiment of the invention, the temperature change rate refers to the average rate of change of the core temperature of the equipment per unit time. The preset duration is 5 minutes, and the preset temperature change rate threshold is 5°C / min. It is understood that when the temperature change rate exceeds the preset temperature change rate threshold, it indicates a risk of sudden overheating of the equipment. Even if the current health index is still in the low-to-medium or medium warning range, the warning level needs to be raised by one level to intervene in advance and prevent the failure from occurring.

[0031] Specifically, the energy efficiency optimization module dynamically adjusts the optimization magnitude of the energy efficiency optimization strategy based on the level of the fault warning signal output by the fault prediction module, wherein, When a low-level fault warning signal is received, maintain the initial optimization level; When a medium-level fault warning signal is received, the optimization range of the relevant warning equipment is compressed according to the first ratio. When a high-level fault warning signal is received, the optimization range of the warning equipment is compressed according to the second ratio.

[0032] It is understandable that the optimization range of the warning equipment involved in the medium-level fault warning signal is the product of the initial optimization range and the first proportion, while the optimization range of the warning equipment involved in the high-level fault warning signal is the product of the initial optimization range and the second proportion.

[0033] In this embodiment of the invention, the initial optimization range includes adjustments to the setpoint of the cooling system and adjustments to the power consumption of the IT equipment. For the cooling system, the initial optimization range may be, for example, increasing the chilled water outlet temperature setpoint by 2°C or decreasing the air conditioner fan speed by 10%; for the IT equipment, the initial optimization range may be, for example, decreasing the CPU frequency by 15% or increasing the idle server hibernation ratio by 20%. The first ratio is 50%, and the second ratio is 20%, but the values ​​are not limited to these, and those skilled in the art can adjust the values ​​according to actual needs.

[0034] Specifically, when the difference between the real-time energy utilization efficiency value and the preset energy efficiency target value is greater than a preset deviation threshold, the energy efficiency optimization module reverses the first ratio or the second ratio according to the magnitude of the utilization efficiency difference. The utilization efficiency difference is the difference between the real-time energy utilization efficiency value and the preset energy efficiency target value, and the correction magnitude of the first ratio or the second ratio is positively correlated with the utilization efficiency difference.

[0035] In this embodiment of the invention, the real-time power utilization efficiency value is determined by the energy efficiency optimization module, the preset energy efficiency target value is 1.3, and the preset deviation threshold value is 0.1. However, the above values ​​are not limited to these, and those skilled in the art can adjust the above values ​​according to actual needs.

[0036] Specifically, the corrected ratio is calculated using the following formula: , In the formula, This is the corrected proportion. As the initial ratio, To utilize the efficiency difference, To set a preset deviation threshold, The value is a correction coefficient. In this embodiment of the invention, the preset deviation threshold is 0.1, and the correction coefficient is 2.0.

[0037] Please see Figure 3 As shown, it is a flowchart of the final control command issuance strategy in an embodiment of the present invention; Specifically, the collaborative control module determines the final control command issuance strategy based on the comparison between the list of warning devices in the fault warning signal and the list of devices to be adjusted in the energy efficiency optimization strategy, including: If the number of intersecting devices is zero, then the final control command is directly issued according to the energy efficiency optimization strategy described above; If the number of intersecting devices is greater than zero and less than a preset threshold, the adjustment instructions involving intersecting devices in the energy efficiency optimization strategy will be transferred to non-intersecting devices before being issued. If the number of overlapping devices is greater than or equal to a preset threshold, the energy efficiency optimization strategy will be suspended and the energy efficiency optimization module will be notified to regenerate the energy efficiency optimization strategy.

[0038] In this embodiment of the invention, the preset quantity threshold is 20% of the sum of the warning device list and the device list to be adjusted, but this value is not limited to this, and those skilled in the art can adjust this value according to actual needs.

[0039] Understandably, in traditional control systems, fault diagnosis modules and energy efficiency optimization modules are often parallel and independent. The fault diagnosis module focuses on equipment safety, while the energy efficiency optimization module focuses on operational economy, and the two lack information exchange. By comparing intersection data, isolated fault data and energy efficiency data are fused at the decision-level, resolving the command conflict problem caused by the lack of information connectivity between safety control and energy-saving control in existing technologies.

[0040] Specifically, when transferring adjustment commands, the collaborative control module determines the feasibility of the transfer operation based on a comparison between the current temperature value of the target non-intersecting device and a preset temperature threshold. If the current temperature value of the target non-intersecting device is less than the preset temperature threshold, the transfer operation is deemed feasible, and the adjustment command is executed to transfer the device. If the current temperature value of the target non-intersecting device is greater than or equal to the preset temperature threshold, the transfer operation is determined to be infeasible, triggering the overall suspension of the energy efficiency optimization strategy. The reason for the suspension and the target device information are fed back to the fault prediction module and the energy efficiency optimization module. After waiting for the preset retry cycle, the transfer is attempted again, or manual intervention is required.

[0041] In this embodiment of the invention, the preset temperature threshold is 75°C, but this value is not limited to this. The specific value can be set according to the specifications of different devices or the security operation specifications of data centers.

[0042] Understandably, this step aims to prevent heat buildup caused by command transfer. If the target device intended to receive the transferred load is already nearing its temperature threshold, forcibly transferring the load to it could lead to overheating and system failure. Therefore, by verifying the temperature feasibility of the target device, ensuring that the load is transferred to a truly healthy device with sufficient capacity, the safety and reliability of the control strategy are further enhanced.

[0043] Specifically, the collaborative control module evaluates the effectiveness of the collaborative control based on the comparison between the actual temperature change value within a preset time period after the final control command is issued and the preset temperature change threshold, and adjusts the preset health threshold when the actual temperature change value is less than the preset temperature change threshold.

[0044] In this embodiment of the invention, the preset time period is 10 minutes and the preset temperature change threshold is 2°C. However, the above values ​​are not limited to these, and those skilled in the art can adjust the above values ​​according to actual needs.

[0045] Specifically, the adjustment range of the preset health threshold is positively correlated with the temperature change difference, where the temperature change difference is the absolute value of the difference between the actual temperature change value and the preset temperature change threshold. The adjusted preset health threshold is calculated using the following formula: , In the formula, The adjusted preset health threshold, Set the current preset health threshold (either the first preset health threshold or the second preset health threshold). To preset the temperature change threshold, This represents the actual temperature change value within a preset time period. Adjust the coefficients for learning.

[0046] In this embodiment of the invention, the learning adjustment coefficient is set to 0.1. When the actual temperature rise effect does not meet expectations, the system appropriately lowers the health threshold, making the fault prediction module's evaluation criteria for the equipment's health status more stringent, thereby issuing early warnings and achieving closed-loop self-optimization of the control system.

[0047] Please see Figure 4 As shown, it is a flowchart of the IDC data center energy efficiency optimization and fault prediction method according to an embodiment of the present invention; This invention provides an AI-based method for optimizing energy efficiency and predicting faults in IDC (Internet Data Center) systems, comprising: Step S1: Collect real-time data on the energy consumption of IT equipment in the data center, the operating parameters of the cooling system, and the status log data of the server. Step S2: Based on the status log data, analyze the health status of the equipment using an AI fault prediction model, output fault warning information and equipment health status index, and classify the fault warning information according to the comparison results between the equipment health status index and the preset health threshold to generate fault warning signals of different levels. Step S3: Determine the real-time power utilization efficiency value of the data center based on the energy consumption data of the IT equipment and the operating parameters of the cooling system, compare the real-time power utilization efficiency value with the preset energy efficiency target value, and generate an energy efficiency optimization strategy through the AI ​​energy efficiency optimization model based on the comparison result; Step S4: Dynamically adjust the optimization magnitude of the energy efficiency optimization strategy according to the level of the fault warning signal output by the fault prediction module; Step S5: Compare the list of warning devices in the fault warning signal with the list of devices to be adjusted in the energy efficiency optimization strategy, and determine the final control command issuance strategy based on the comparison result. Step S6: Based on the comparison between the actual temperature change value within the preset time period after the final control command is issued and the preset temperature change threshold, evaluate and provide feedback on the effect of the coordinated control.

[0048] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the invention (including the claims) is limited to these examples; within the framework of the invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of the different aspects of the invention as described above, which are not provided in the details for the sake of brevity.

[0049] This invention is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An AI-based energy efficiency optimization and fault prediction system for IDC data centers, characterized in that, include: The data acquisition module is used to collect real-time data on the energy consumption of IT equipment in the data center, the operating parameters of the cooling system, and the status log data of the server. The fault prediction module is connected to the data acquisition module and is used to analyze the health status of the equipment based on the status log data through the AI ​​fault prediction model, output fault warning information and equipment health status index, and perform graded processing on the fault warning information based on the comparison result of the equipment health status index and the preset health threshold to generate fault warning signals of different levels. An energy efficiency optimization module, connected to the data acquisition module, is used to determine the real-time power utilization efficiency value of the data center based on the energy consumption data of the IT equipment and the operating parameters of the cooling system, compare the real-time power utilization efficiency value with the preset energy efficiency target value, generate an energy efficiency optimization strategy through an AI energy efficiency optimization model based on the comparison result, and dynamically adjust the optimization range of the energy efficiency optimization strategy according to the level of the fault warning signal output by the fault prediction module. The collaborative control module is connected to the fault prediction module and the energy efficiency optimization module respectively. It is used to compare the intersection of the list of warning devices in the fault warning signal and the list of devices to be adjusted in the energy efficiency optimization strategy, determine the final control command issuance strategy based on the comparison result, and evaluate and provide feedback on the effect of collaborative control based on the comparison result of the actual temperature change value within a preset time period after the final control command is issued and the preset temperature change threshold.

2. The AI-based IDC data center energy efficiency optimization and fault prediction system according to claim 1, characterized in that, The fault prediction module determines the level of the fault warning signal based on the comparison between the equipment health status index and a preset health threshold. When the device health status index is greater than or equal to the first preset health threshold, a low-level fault warning signal is generated. When the device health status index is less than the first preset health threshold and greater than or equal to the second preset health threshold, an intermediate fault warning signal is generated. When the device health status index is less than the second preset health threshold, an advanced fault warning signal is generated.

3. The AI-based IDC data center energy efficiency optimization and fault prediction system according to claim 2, characterized in that, The equipment health status index is determined by load deviation and temperature deviation.

4. The AI-based IDC data center energy efficiency optimization and fault prediction system according to claim 3, characterized in that, When generating a low-to-medium level fault warning signal or a medium level fault warning signal, the fault prediction module corrects the level of the fault warning signal based on the comparison between the temperature change rate of the corresponding device within a preset time period and a preset temperature change rate threshold. Specifically, when the temperature change rate exceeds the preset temperature change rate threshold, the level of the fault warning signal is increased by one level.

5. The AI-based IDC data center energy efficiency optimization and fault prediction system according to claim 4, characterized in that, The energy efficiency optimization module dynamically adjusts the optimization magnitude of the energy efficiency optimization strategy based on the level of the fault warning signal output by the fault prediction module, wherein, When a low-level fault warning signal is received, maintain the initial optimization level; When a medium-level fault warning signal is received, the optimization range of the relevant warning equipment is compressed according to the first ratio. When a high-level fault warning signal is received, the optimization range of the warning equipment is compressed according to the second ratio.

6. The AI-based IDC data center energy efficiency optimization and fault prediction system according to claim 5, characterized in that, When the difference between the real-time energy utilization efficiency value and the preset energy efficiency target value is greater than a preset deviation threshold, the energy efficiency optimization module reversely corrects the first ratio or the second ratio according to the magnitude of the utilization efficiency difference. The utilization efficiency difference is the difference between the real-time energy utilization efficiency value and the preset energy efficiency target value, and the correction magnitude of the first ratio or the second ratio is positively correlated with the utilization efficiency difference.

7. The AI-based IDC data center energy efficiency optimization and fault prediction system according to claim 6, characterized in that, The collaborative control module determines the final control command issuance strategy based on the comparison between the list of warning devices in the fault warning signal and the list of devices to be adjusted in the energy efficiency optimization strategy, including: If the number of intersecting devices is zero, then the final control command is directly issued according to the energy efficiency optimization strategy described above; If the number of intersecting devices is greater than zero and less than a preset threshold, the adjustment instructions involving intersecting devices in the energy efficiency optimization strategy will be transferred to non-intersecting devices before being issued. If the number of overlapping devices is greater than or equal to a preset threshold, the energy efficiency optimization strategy will be suspended and the energy efficiency optimization module will be notified to regenerate the energy efficiency optimization strategy.

8. The AI-based IDC data center energy efficiency optimization and fault prediction system according to claim 7, characterized in that, When transferring adjustment commands, the collaborative control module determines the feasibility of the transfer operation based on a comparison between the current temperature value of the target non-intersecting device and a preset temperature threshold. If the current temperature value of the target non-intersecting device is less than the preset temperature threshold, the transfer operation is deemed feasible, and the adjustment command is executed to transfer the device. If the current temperature value of the target non-intersecting device is greater than or equal to the preset temperature threshold, the transfer operation is determined to be infeasible, triggering the overall shelving of the energy efficiency optimization strategy.

9. The AI-based IDC data center energy efficiency optimization and fault prediction system according to claim 8, characterized in that, The collaborative control module evaluates the effectiveness of the collaborative control based on the comparison between the actual temperature change value within a preset time period after the final control command is issued and the preset temperature change threshold. When the actual temperature change value is less than the preset temperature change threshold, the preset health threshold is adjusted.

10. An AI-based method for energy efficiency optimization and fault prediction in IDC data centers, applied to the IDC data center energy efficiency optimization and fault prediction system described in any one of claims 1-9, characterized in that, include: Step S1: Collect real-time data on the energy consumption of IT equipment in the data center, the operating parameters of the cooling system, and the status log data of the server. Step S2: Based on the status log data, analyze the health status of the equipment using an AI fault prediction model, output fault warning information and equipment health status index, and classify the fault warning information according to the comparison results between the equipment health status index and the preset health threshold to generate fault warning signals of different levels. Step S3: Determine the real-time power utilization efficiency value of the data center based on the energy consumption data of the IT equipment and the operating parameters of the cooling system, compare the real-time power utilization efficiency value with the preset energy efficiency target value, and generate an energy efficiency optimization strategy through the AI ​​energy efficiency optimization model based on the comparison result; Step S4: Dynamically adjust the optimization magnitude of the energy efficiency optimization strategy according to the level of the fault warning signal output by the fault prediction module; Step S5: Compare the list of warning devices in the fault warning signal with the list of devices to be adjusted in the energy efficiency optimization strategy, and determine the final control command issuance strategy based on the comparison result. Step S6: Based on the comparison between the actual temperature change value within the preset time period after the final control command is issued and the preset temperature change threshold, evaluate and provide feedback on the effect of the coordinated control.

Citation Information

Patent Citations

  • CN114970358A