Data center active preventive operation and maintenance method based on intelligent group control
Through the proactive preventive operation and maintenance method of the data center with intelligent group control, the data center environment and equipment parameters are collected and analyzed in real time, and the cooling tower fan frequency and water pump speed are dynamically adjusted, which solves the problem of insufficient adaptability in the operation and maintenance management of the data center and achieves efficient and stable fault response and energy consumption optimization.
Patent Information
- Application Number
- CN202510348168.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-09-09
AI Technical Summary
The existing data center operation and maintenance management lacks adaptability, making it difficult to achieve overall performance improvement among devices. Fault prediction relies on statistical models and simplified machine learning algorithms, resulting in poor responsiveness.
A proactive preventive operation and maintenance method for data centers based on intelligent group control is adopted. By collecting environmental parameters and equipment operating parameters in real time, a state association matrix and state space model are constructed. Combined with the Markov decision process and SAC algorithm, control parameters such as the cooling tower fan frequency and water pump speed are dynamically adjusted to form a closed-loop control mechanism.
It significantly improves fault response speed and operation and maintenance efficiency, reduces the energy consumption ratio of cooling equipment to IT equipment, improves resource utilization, enhances the system's adaptability and stability under complex working conditions, and reduces operating costs.
Smart Images

Figure CN120608846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic communication technology, and in particular to the transmission of digital information. Background Art
[0002] Data center operations and maintenance have traditionally relied on a reactive approach, relying on manual monitoring and post-fault handling. With the advancement of digital and intelligent technologies, more and more companies are exploring proactive and preventative operations and maintenance approaches. However, existing technologies often overlook the interconnectivity and integrated management of data center equipment, making it difficult to improve overall efficiency. Existing technologies primarily rely on statistical models and simplified machine learning algorithms for fault prediction.
[0003] For example, Chinese patent publication number CN118590370A discloses a device management method, system, and device for intelligent operation and maintenance, providing the following technical solution: This application provides a device management method, system, and device for intelligent operation and maintenance, belonging to the field of data center operation and maintenance technology. The method establishes a management network for each operation and maintenance network device based on configuration parameters pre-assigned to each operation and maintenance network device. Based on user input operations on a preset front-end page, the installation environment information of the SDN controller is determined to configure the SDN management system based on the management network and installation environment information. Using the SDN front-end page corresponding to the SDN management system and preset panel generation rules, a device panel image corresponding to the operation and maintenance network device is constructed and displayed on the user terminal. Based on the user's command issuance operation to the operation and maintenance network device based on the device panel image, a command line template corresponding to the operation and maintenance network device is matched, and the corresponding command line is issued to the operation and maintenance network device according to the command line template, so as to perform operation and maintenance management of the operation and maintenance network device. However, the above-mentioned device management method, system, and device for intelligent operation and maintenance lack adaptability and intelligence, and cannot meet rapidly changing operational needs. Summary of the Invention
[0004] The present invention solves the problems of the existing technology in terms of insufficient real-time dynamic adjustment, lack of effective integration of complex interactions between data streams and devices at different levels, and poor operation and maintenance effects and responsiveness. It proposes a proactive preventive operation and maintenance method for data centers based on intelligent group control, which achieves the goals of dynamic adjustment according to the status of data center equipment and environmental conditions, adaptation to environmental changes and system parameter fluctuations, high fault recognition rate and strong responsiveness.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A proactive preventive operation and maintenance method for a data center based on intelligent group control includes the following steps: S1: Real-time collection of data center environmental parameters and equipment operating parameters and data preprocessing; S2: Preliminary definition of state space, construction of state correlation matrix and state space model between devices; S3: Define the control action space and establish an optimization function with the goal of minimizing energy efficiency and avoiding equipment overheating; S4: Improve the state space by fusing historical means with real-time states and quantifying system dynamic trends, train intelligent agents based on Markov decision making, and embed dynamic adjustment strategies; S5: The intelligent agent generates control instructions, and the experimental platform executes the instructions and provides feedback, forming a dynamic closed loop.
[0006] It can dynamically adjust control strategies according to the status of data center equipment and environmental conditions, adapt to environmental changes and system parameter fluctuations, and thus improve fault identification and response capabilities.
[0007] Preferably, step S2 includes defining key indicators of the state space, which include environmental parameters; constructing a state association matrix reflecting the device dependency by analyzing the operating data between devices and calculating the correlation coefficients of parameters between devices; establishing a state space model based on the state association matrix, and predicting the system state at the next moment through the current state parameters and the device interaction relationship.
[0008] By constructing a state association matrix through correlation coefficients, we can quantify the dynamic dependencies between devices, enhance the model's ability to model complex interactions, and thus improve the accuracy of system state prediction.
[0009] Preferably, in step S3, the control action space is defined by constructing a reward function based on the cooling tower fan frequency and the water pump speed, and establishing an optimization function with the goal of minimizing energy efficiency and avoiding equipment overheating. The optimization function integrates the goals into a unified reward mechanism through weight factors and dynamically adjusts the control strategy.
[0010] A balance between energy consumption and equipment temperature is achieved through multi-objective optimization functions and weight factors, avoiding system imbalance caused by single-objective optimization. At the same time, dynamic adjustment strategies are used to adapt to real-time operating condition changes.
[0011] Preferably, step S5 specifically includes receiving the current status information of the experimental platform environment, making decisions based on the strategy obtained through training, dynamically adjusting control parameters including the cooling tower fan frequency and the water pump speed, and then conveying the decisions to the control strategy execution device; the experimental platform executes the decisions conveyed by the intelligent agent, and feeds back the state and reward value at the next moment to the intelligent agent based on the execution results, thus forming a closed-loop interactive continuous optimization control strategy.
[0012] Through the closed-loop feedback mechanism, the strategy can be continuously iterated and optimized, which can enhance the system's response capability to emergencies and improve the real-time and robustness of the control strategy.
[0013] Preferably, the step S4 specifically includes the following steps: S4.1: The integrated average value of the parameters within the historical control cycle is combined with the real-time value at the current moment to comprehensively represent the system status. S4.2: The appropriate control mode is selected based on whether the working conditions are stable. If the working conditions remain stable for a period of time, the process proceeds to step S4.3. If the working conditions cannot remain stable for a period of time, the process proceeds to step S4.4. S4.3: Select a relatively long fixed control interval and proceed to step S5; S4.4: Use the SAC algorithm to dynamically adjust the control interval and add an additional term to the reward function to optimize the adjustment.
[0014] The comprehensiveness of state representation is improved by integrating historical averages and real-time values. At the same time, the control mode is adaptively selected according to the stability of the working conditions, taking into account both system efficiency and dynamic response requirements.
[0015] Preferably, the step S4.4 specifically includes introducing indicators of ambient temperature change, humidity change and heat load change, combining with dynamic adjustment of the control interval, capturing the gradual and sudden fluctuation characteristics of the system, adopting the SAC algorithm, training the intelligent agent based on the Markov decision process, learning the optimal strategy through the state transition probability and reward function, and adaptively adjusting the control interval and parameters.
[0016] By quantifying the changing characteristics of environmental parameters and integrating them into the reward function, the algorithm's sensitivity to dynamic fluctuations of the system is enhanced, thereby improving the adaptability and stability of the control strategy under complex working conditions.
[0017] Preferably, the minimization of energy consumption efficiency achieves energy saving by reducing the ratio of cooling equipment power consumption to IT equipment power consumption, and the avoidance of equipment overheating is achieved by real-time monitoring of the temperature of top and bottom equipment, and an overheating penalty is imposed if the temperature exceeds a threshold.
[0018] PUE optimization directly reduces data center energy consumption costs, while the temperature threshold penalty mechanism prevents equipment from overheating and damage, ensuring safe system operation.
[0019] Preferably, in step S4.2, the basis for determining whether the working conditions are stable is that when the change amplitudes of the heat load, ambient temperature and relative humidity are all less than preset thresholds, it is determined to be a stable state; when the change amplitude of any parameter reaches or exceeds the preset threshold, it is determined to be an unstable state; the preset threshold is dynamically set according to the statistical value of the parameter fluctuation range in the historical data, and is adaptively corrected according to the actual control effect in the closed-loop feedback.
[0020] Through dynamic threshold setting and adaptive correction mechanism, the system's judgment accuracy on changes in working conditions is improved, avoiding misjudgment or missed judgment problems caused by fixed thresholds.
[0021] Preferably, the working conditions include heat load, ambient temperature and relative humidity.
[0022] Preferably, the environmental parameters include temperature, humidity and heat load; the equipment operation parameters include IT equipment power consumption, cooling tower fan frequency and water pump speed; and the data preprocessing includes standardization.
[0023] Comprehensive monitoring is achieved by covering key environmental and equipment parameters, while standardized processing eliminates dimensional differences and improves the convergence speed and generalization ability of model training.
[0024] Compared with the prior art, the present invention has the following beneficial effects.
[0025] 1. The present invention collects environmental parameters and equipment operation data in real time, constructs a state association matrix and state space model based on correlation coefficients, quantifies the dynamic dependency between devices, and achieves accurate prediction of system status. The Markov decision process is combined with the SAC algorithm to train intelligent agents, dynamically adjust control parameters such as the cooling tower fan frequency and water pump speed, and form a closed-loop control mechanism of "perception-decision-execution-feedback". It not only enhances the coordination ability between devices, but also continuously optimizes strategies through real-time feedback, significantly improving fault response speed and operation and maintenance efficiency, and solving the problem of traditional methods relying on historical data and dynamic adjustment lags. It is especially suitable for complex and changeable working environments.
[0026] 2. This invention establishes a multi-objective optimization function that integrates weight factors with the core objectives of minimizing energy efficiency and preventing equipment overheating. It then uses a reward function to balance cooling power consumption and temperature control requirements. A combined state space of integral averages and real-time values is introduced, and the control interval adjustment strategy is optimized based on the dynamic change indicators of ambient temperature, humidity, and heat load. This effectively reduces the energy consumption ratio of cooling equipment to IT equipment, while preventing equipment overheating through a temperature threshold penalty mechanism. This balances energy efficiency optimization with equipment safety, significantly reducing data center operating costs and improving resource utilization.
[0027] 3. This invention uses historical data statistics and closed-loop feedback to adaptively modify thresholds, accurately identifying stable and unstable states. Under unstable conditions, it uses the SAC algorithm to dynamically adjust control intervals, capturing both gradual and sudden fluctuations in the system. It also optimizes control parameters through additional terms in the reward function. This mechanism not only reduces computational load by using long-interval control under stable conditions, but also responds quickly to sudden fluctuations, avoiding misjudgments and oscillations. This significantly improves the system's adaptability and operational stability under extreme loads or sudden environmental changes, ensuring high data center reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is an overall flow chart of a data center proactive preventive operation and maintenance method based on intelligent group control according to the present invention. DETAILED DESCRIPTION
[0029] To make the objectives, technical solutions, and advantages of the present disclosure more apparent, embodiments of the present disclosure are described in further detail below with reference to the accompanying drawings. The proportions of the components herein are not drawn to scale, and the proportions and dimensions shown in the accompanying drawings are not intended to limit the essential technical solutions of the present disclosure. These embodiments do not describe all details in detail, nor do they limit the present disclosure to the specific embodiments described.
[0030] See also Figure 1 As shown, a proactive preventive operation and maintenance method for a data center based on intelligent group control includes the following steps: S1: real-time collection of data center environmental parameters and equipment operating parameters and data preprocessing; S2: Preliminary definition of state space, construction of state correlation matrix and state space model between devices; S3: Define the control action space and establish an optimization function with the goal of minimizing energy efficiency and avoiding equipment overheating; S4: Improve the state space by fusing historical means with real-time states and quantifying system dynamic trends, train intelligent agents based on Markov decision making, and embed dynamic adjustment strategies; S5: The intelligent agent generates control instructions, and the experimental platform executes the instructions and provides feedback, forming a dynamic closed loop.
[0031] like Figure 1 In one embodiment shown, Figure 1 This is an overall flow chart of a data center proactive preventive operation and maintenance method based on intelligent group control according to the present invention.
[0032] The present invention first uses a sensor network deployed within the data center to collect real-time environmental parameters (including temperature, humidity, and heat load) and equipment operating parameters (such as IT equipment power consumption, cooling tower fan frequency, and water pump speed). This data undergoes standardized preprocessing to eliminate numerical discrepancies between parameters of different dimensions, ensuring the convergence speed and accuracy of subsequent model training. This preprocessed data is stored in a central database, providing a unified input foundation for subsequent state modeling and control decisions.
[0033] After completing data collection, the system preliminarily defines key indicators of the state space, including environmental parameters (temperature, humidity, and heat load) and equipment operating parameters (fan frequency, pump speed, etc.). By analyzing the operating data between devices and calculating the correlation coefficients between device parameters (such as the correlation between fan frequency and room temperature, and the coupling relationship between pump speed and cooling efficiency), a state correlation matrix reflecting the dynamic dependencies of the devices is constructed. Based on this matrix, the system establishes a state space model. By analyzing the interaction between the current environmental state and the equipment, it predicts the system state at the next moment (such as temperature change trends and equipment load fluctuations), providing a predictive basis for dynamic control.
[0034] The core parameters of the system-defined control action space are the cooling tower fan frequency and water pump speed. These two parameters directly affect the heat dissipation efficiency and energy consumption level. With the goal of minimizing power usage effectiveness (PUE) and avoiding equipment overheating, a multi-objective optimization function is constructed: energy saving is achieved by reducing the ratio of cooling equipment power consumption to IT equipment power consumption, while monitoring the temperature of top and bottom equipment in real time. If the temperature exceeds the preset threshold, the overheating penalty mechanism is triggered. The optimization function integrates energy consumption and temperature control goals through dynamic weight factors to form a unified reward mechanism. For example, during high-temperature periods, priority is given to ensuring heat dissipation safety, and the weight of temperature control is appropriately increased; during low-load periods, the focus is on reducing energy consumption, and the operating parameters of fans and water pumps are dynamically adjusted.
[0035] To improve the comprehensiveness of state representation, the system integrates the integrated average values of parameters over historical control cycles (such as the average temperature and humidity over the past hour) with the current real-time values to comprehensively represent the system state. For example, when the current temperature suddenly increases, the historical average values are combined to determine whether it is a short-term fluctuation or a sustained warming trend. Based on the stability of operating conditions (whether the changes in heat load, ambient temperature, and relative humidity exceed dynamic thresholds), the system automatically selects a control mode. If the parameter changes are small and stable, a fixed long-interval control strategy (such as adjustment every 10 minutes) is adopted to reduce computing resource consumption. If a sudden change in the parameters is detected (such as a surge in heat load or a sudden drop in humidity), the system switches to dynamic adjustment mode. Using the SAC algorithm, combined with the rate of change indicators of ambient temperature, humidity, and heat load, the control interval is shortened in real time (for example, to 1 minute) to capture the fluctuation characteristics of both gradual and sudden changes in the system. The intelligent agent is trained through a Markov decision process, learning the optimal policy based on state transition probabilities and a reward function, and adaptively adjusting the control interval and parameters.
[0036] The stability of operating conditions is determined based on the following criteria: when the fluctuations in the heat load, ambient temperature, and relative humidity are all less than preset thresholds, the system is considered stable; when the fluctuations in any parameter reach or exceed preset thresholds, the system is considered unstable. These thresholds are dynamically set based on historical data showing the parameter's fluctuation range and adaptively adjusted based on actual control performance during closed-loop feedback. This data-driven dynamic threshold setting mechanism enables the system to adapt to different environmental scenarios and avoids the misjudgment caused by fixed thresholds. In areas with large diurnal temperature differences, ambient temperature fluctuations are naturally large. Using a fixed threshold can frequently trigger unstable mode. However, setting a threshold based on historical fluctuation statistics can distinguish between normal fluctuations and abnormal mutations, reducing the risk of misjudgment. Furthermore, adaptive correction through closed-loop feedback further optimizes threshold accuracy. If the system experiences response delays due to overly loose thresholds, the thresholds can be automatically tightened through reward feedback, enabling iterative parameter optimization. Furthermore, by distinguishing between stable and unstable states and selecting the control mode, energy waste can be reduced (e.g., by reducing the sampling frequency during stable conditions), ensuring control accuracy while minimizing energy waste. This balance between performance and efficiency meets the requirements of green data center operations.
[0037] The trained agent generates control instructions based on the real-time status (such as increasing the fan frequency by 5% and reducing the water pump speed by 3%). After executing the instructions, the experimental platform monitors changes in environmental parameters and equipment status, and feeds back the next-moment status data and reward value (such as the degree of energy consumption reduction and temperature control effect) to the agent. For example, if the device temperature does not drop and energy consumption increases after adjustment, the reward value is reduced, prompting the agent to optimize its action selection in the next decision; if the temperature is stable and energy consumption decreases after adjustment, the reward value is increased, strengthening the current strategy. This closed-loop interaction mechanism enables the control strategy to be continuously iteratively optimized, gradually approaching the global optimal solution, while enhancing the ability to adapt to sudden working conditions (such as a sudden increase in heat load due to server downtime).
[0038] Before executing decisions conveyed by the agent, the experimental platform assesses the health of the equipment and checks the operating status of key equipment, such as cooling tower fans and water pumps. If the current device responds abnormally, it automatically switches to a backup redundant device to execute control instructions. A redundancy switching penalty mechanism is also implemented: When a device switch is triggered, a penalty term is added to the reward function, reducing the priority of subsequent calls to the same policy and marking the anomalous device for maintenance. This mechanism significantly improves the system's fault tolerance and operational continuity by monitoring the status of key equipment in real time and triggering redundancy switching. In data center operations and maintenance, sudden failures of equipment such as cooling tower fans and water pumps can cause localized overheating or even downtime. Automatically switching to a backup device can restore control within milliseconds, avoiding service interruptions. Furthermore, the redundancy switching penalty mechanism dynamically reduces the priority of anomalous policies, guiding the agent to learn more robust control strategies and preventing occasional equipment failures from misleading long-term decision optimization. Furthermore, marking anomalous devices and triggering maintenance alerts reduces manual inspection costs, enables predictive maintenance, and extends equipment life. This mechanism elevates fault handling from reactive to proactive, balancing system security with cost-effective operations.
[0039] The present invention achieves a comprehensive improvement in the efficiency, economy and stability of data center operation and maintenance through intelligent group control and dynamic closed-loop design. First, the state space modeling based on multi-source data fusion and device association analysis can accurately depict the dynamic coupling relationship between complex devices, significantly improve the reliability of system state prediction, and avoid local optimization problems caused by traditional single-device independent control. For example, by identifying the strong correlation between high-density server areas and specific cooling units through the association matrix, the heat dissipation strategy can be adjusted in a targeted manner to avoid "overcooling" or "heat dissipation blind spots". Secondly, the combination of multi-objective optimization mechanism and adaptive control strategy effectively balances the contradiction between energy consumption and equipment safety. The dynamic weight factor design enables the system to prioritize equipment safety during high-temperature peak periods and automatically switch to energy-saving mode during low-load periods. In addition, based on the mode switching mechanism of working condition stability, long interval control is used to reduce computational overhead in a stable state, and the SAC algorithm is used to quickly respond to sudden changes in an unstable state, avoiding response delays or oscillation problems caused by traditional fixed-frequency control. Finally, the closed-loop feedback mechanism continuously optimizes the control strategy, enabling the system to adapt to long-term dynamic factors such as data center load fluctuations and external environmental changes (such as day and night temperature differences and seasonal humidity changes), significantly reducing the frequency of manual intervention in operation and maintenance and improving the overall level of automation.
[0040] In another embodiment, the following method is used to implement the IoT device data center operation data, including temperature, humidity, heat load, device power consumption, etc., and pre-process it: Where X′ is the standardized data, X is the original data, μ is the mean, and σ is the standard deviation. Define the state space of the data center, including key indicators such as temperature, humidity, and heat load, and construct the state correlation matrix between devices to identify the dependencies between devices: Among them, corr ij is the correlation coefficient between device i and device j. Thus, the state space model S can be constructed: Regarding the formulation of active adjustment strategies, the control action space is defined, and a reward function is designed based on the cooling tower fan frequency, water pump speed, etc. to guide the algorithm to learn the optimal control strategy with the goal of reducing PUE and improving system stability.
[0041] In order to minimize the energy consumption efficiency (PUE) and avoid overheating of each device, the core optimization formula is as follows: where [t0, t1] is the optimization time range, [x] + represents max(x, 0), T equip,up Indicates the top device, T equip,down . indicates the bottom device. f ct represents the chiller frequency, and c represents the pump speed. The upper and lower limits of their settings are indicated by the overline and underline, respectively. The subscript 'τ' represents the real-time value of the current time step. The terms in Equation (4) correspond to the penalties for overheating of the top and bottom equipment, where β0 and β1 are the weighting factors of these penalties. The PUE in (4) is converted to P it Indicates the power consumption of IT equipment. Power consumption of cooling equipment P cool .
[0042] The deep learning algorithm is based on Markov decision process, which is defined by the tuple (S, A, p, r), where the state space S includes the heat load Q, the power consumption of the cooling equipment P cool , top equipment temperature T equip,up , bottom equipment temperature T equip,down , ambient temperature T and ambient relative humidity H. The action space A includes the cooling machine frequency f ct,set , water pump speed c set . State transition probability p(s t+1 |s t , a t ) describes the current state of the t ∈S and action a t ∈A to obtain the next state s t+1The reward function is set to r = r0-Ψ, where r0 is a constant to ensure that the reward remains positive, and Ψ is calculated by equation (4), where [t0, t1] is equal to [t s , t e ].
[0043] In the context of a liquid-cooled data center, if the operating conditions (such as heat load, ambient temperature, or relative humidity) remain stable over a period of time, then choosing a relatively long fixed control interval will not cause a problem. However, if the operating conditions change or fluctuate significantly during this period, choosing a relatively long fixed control interval will result in a delay in sensing the system changes. The SAC algorithm is used to train the agent and dynamically adjust the control interval tci. In order to promote the learning of tci, it is considered to add an additional term and integrate it into the reward function. β3 is a weighting factor, and ∈ is a constant used to prevent division by zero. It allows the selection of the optimal control interval for the current state. This approach minimizes the delay in sensing system changes, reduces misjudgments of system states, and prevents frequent oscillations, thereby enhancing the algorithm's control accuracy, energy efficiency, and stability.
[0044] The combined value state space is expressed as where t s is the start of the control interval, t e is the end of the control interval, and μ0 and μ1 represent the weighting factors of the state space. The integrated integral average takes the time aspect into account and provides a more comprehensive and smoother representation of the system behavior covering the entire control interval. The three variables of ambient temperature change (ΔT), ambient relative humidity change (ΔH) and heat load change (ΔQ) are integrated at the same time, along with the difference between the maximum and average values in the last 20% of the current control interval, to resolve gradual and sudden changes in the system dynamics within the specified time range. Compared to the final value state space, which only uses the state value at the end of the current control cycle, the integrated value state space combines the state value at the end of the current control cycle with the integrated average of all state values within that cycle. By considering both the integrated average of the state values within the current control cycle and the state value at the end of that cycle, it can more comprehensively reflect the system state information within the current control cycle. By using the integrated value state space and the combined state space, it is possible to more accurately capture the changing trends of the system state.
[0045] Based on the learned strategy, the intelligent agent dynamically adjusts control parameters such as the cooling tower fan frequency and water pump speed to achieve efficient and energy-saving operation of the data center.
[0046] Deep learning control consists of two parts: a deep learning agent and an experimental platform environment. The deep learning agent receives the current state information of the experimental platform environment, makes decisions based on the trained strategy, and then conveys the decisions to the control strategy execution device. The experimental platform environment executes the decisions conveyed by the agent, and based on the execution results, feeds back the state of the next moment to the agent, while providing immediate rewards. A circular interaction is formed between the agent and the environment. Through continuous learning and decision-making, the agent can find the best strategy to control the experimental platform. The algorithm model execution framework proposed in this invention realizes the dynamic optimization and autonomous decision-making of data center operation and maintenance through a closed-loop interaction mechanism: the deep learning agent continuously learns and adjusts the control strategy based on the real-time feedback state information and multi-dimensional reward function of the experimental platform, breaking through the limitations of traditional fixed logic and accurately responding to complex scenarios such as sudden changes in heat load and environmental fluctuations; at the same time, the hybrid mode of integrating historical data training and online reinforcement learning gives the model powerful generalization capabilities, which can adapt to heterogeneous equipment clusters and seasonal operating conditions. It is seamlessly compatible with existing monitoring systems through modular interfaces, significantly reducing deployment costs. This framework replaces manual experience with data-driven approaches, not only achieving coordinated optimization of cooling equipment and IT loads, but also being extended to emerging needs such as health monitoring and carbon emission management. It reduces operational complexity while ensuring safety, providing an excellent solution for intelligent upgrades of data centers.
[0047] In summary, the present invention proposes a proactive preventive operation and maintenance method for data centers based on intelligent group control, which realizes efficient and safe operation and maintenance of data centers through real-time perception, dynamic modeling and closed-loop optimization. In response to the problems of delayed response and insufficient equipment coordination in traditional passive operation and maintenance, the present invention constructs a multi-dimensional data acquisition network covering environmental parameters and equipment operating status, establishes a state prediction model by analyzing the dynamic correlation between devices, and accurately captures the evolution trend of key indicators such as temperature, humidity, and load. With the dual goals of reducing energy consumption and preventing equipment overheating, the system adopts a reinforcement learning algorithm to dynamically adjust the operating parameters of the cooling equipment, combines the composite state space of the historical data mean and real-time fluctuations, and adaptively selects the control strategy: extending the control interval to reduce computing overhead under stable working conditions, and quickly responding to sudden fluctuations and optimizing heat dissipation efficiency. The innovative closed-loop feedback mechanism transmits the control effect back to the decision-making layer in real time, driving the strategy to continuously iteratively optimize, while introducing equipment redundancy switching and abnormal penalty mechanisms to achieve seamless replacement of faulty equipment within milliseconds, significantly improving the system's fault tolerance. Compared with the existing technology, this solution has broken through three technical bottlenecks: through dynamic dependency modeling between devices, the false alarm rate of faults will be greatly reduced; based on the multi-objective optimization algorithm, the energy efficiency of the data center will be greatly improved; with the help of adaptive control strategies, the average fault repair time and the risk of equipment overheating will be greatly reduced. The present invention is based on the design of an active preventive operation and maintenance system for data centers based on intelligent group control. The system has high flexibility and adaptability, and combines a comprehensive operation and maintenance strategy with dynamic scheduling and real-time feedback to ensure that it can effectively respond to various emergencies in complex environments. Compared with the existing technology, the present invention significantly improves the intelligence level and resource utilization of data center operation and maintenance management. By comprehensively considering the dependency relationship between devices and real-time dynamic adjustment, the present invention can effectively reduce the occurrence rate of faults and significantly improve operation and maintenance efficiency. It has higher application value and promotion prospects.
[0048] The present invention is not limited to the above-mentioned embodiments. Regardless of any changes in shape or material composition, any structural design provided by the present invention is a variation of the present invention and should be considered within the scope of protection of the present invention.
Claims
1. A proactive preventive operation and maintenance method for a data center based on intelligent group control, characterized in that: The following steps are involved: S1: Real-time collection of data center environmental parameters and equipment operating parameters and data preprocessing; S2: Preliminary definition of state space, construction of state correlation matrix and state space model between devices; S3: Define the control action space and establish an optimization function with the goal of minimizing energy efficiency and avoiding equipment overheating; S4: Improve the state space by fusing historical means with real-time states and quantifying system dynamic trends, train intelligent agents based on Markov decision making, and embed dynamic adjustment strategies; S5: The intelligent agent generates control instructions, and the experimental platform executes the instructions and provides feedback, forming a dynamic closed loop.
2. The method for proactive preventive operation and maintenance of a data center based on intelligent group control according to claim 1, characterized in that: The step S2 includes defining key indicators of the state space, which include environmental parameters; constructing a state association matrix reflecting the device dependency by analyzing the operating data between devices and calculating the correlation coefficients of parameters between the devices; establishing a state space model based on the state association matrix, and predicting the system state at the next moment through the current state parameters and the interaction relationship between the devices.
3. The data center proactive preventive operation and maintenance method based on intelligent group control according to claim 1 or 2, characterized in that: In step S3, the control action space is defined by constructing a reward function based on the cooling tower fan frequency and the water pump speed, and establishing an optimization function with the goal of minimizing energy efficiency and avoiding equipment overheating. The optimization function integrates the goals into a unified reward mechanism through weight factors and dynamically adjusts the control strategy.
4. The method for proactive preventive operation and maintenance of a data center based on intelligent group control according to claim 3, characterized in that: The step S5 specifically includes receiving the current status information of the experimental platform environment, making decisions based on the trained strategy, dynamically adjusting control parameters including the cooling tower fan frequency and water pump speed, and then conveying the decision to the control strategy execution device; the experimental platform executes the decision conveyed by the intelligent agent, and feeds back the state and reward value at the next moment to the intelligent agent based on the execution result, forming a closed-loop interactive continuous optimization control strategy.
5. The method for proactive preventive operation and maintenance of a data center based on intelligent group control according to claim 4, characterized in that: The step S4 specifically includes the following steps: S4.1: The integrated average value of the parameters within the historical control cycle is combined with the real-time value at the current moment to comprehensively represent the system status. S4.2: The appropriate control mode is selected based on whether the working conditions are stable. If the working conditions remain stable for a period of time, the process proceeds to step S4.
3. If the working conditions cannot remain stable for a period of time, the process proceeds to step S4.
4. S4.3: Select a relatively long fixed control interval and proceed to step S5; S4.4: Use the SAC algorithm to dynamically adjust the control interval and add an additional term to the reward function to optimize the adjustment.
6. The method for proactive preventive operation and maintenance of a data center based on intelligent group control according to claim 5, characterized in that: The step S4.4 specifically includes introducing indicators of ambient temperature change, humidity change and heat load change, combining with dynamic adjustment of control intervals, capturing the gradual and sudden fluctuation characteristics of the system, using the SAC algorithm, training the intelligent agent based on the Markov decision process, learning the optimal strategy through state transition probability and reward function, and adaptively adjusting the control interval and parameters.
7. The method for proactive preventive operation and maintenance of a data center based on intelligent group control according to claim 1, characterized in that: The minimization of energy efficiency achieves energy saving by reducing the ratio of cooling equipment power consumption to IT equipment power consumption, and the prevention of equipment overheating is achieved by real-time monitoring of the temperature of top and bottom equipment, and overheating penalties are imposed if the temperature exceeds a threshold.
8. The data center proactive preventive operation and maintenance method based on intelligent group control according to claim 5 or 6, characterized in that: In step S4.2, whether the working condition is stable is determined as follows: when the variation of the heat load, ambient temperature, and relative humidity is less than a preset threshold, it is determined to be a stable state; when the variation of any parameter reaches or exceeds a preset threshold, it is determined to be an unstable state; The preset threshold is dynamically set according to the statistical value of the parameter fluctuation range in the historical data, and is adaptively corrected according to the actual control effect in the closed-loop feedback.
9. The method for proactive preventive operation and maintenance of a data center based on intelligent group control according to claim 7, characterized in that: The operating conditions include heat load and ambient temperature and relative humidity.
10. The method for proactive preventive operation and maintenance of a data center based on intelligent group control according to claim 1, characterized in that: The environmental parameters include temperature, humidity and heat load; the equipment operation parameters include IT equipment power consumption, cooling tower fan frequency and water pump speed; and the data preprocessing includes standardization.
Citation Information
Patent Citations
Intelligent operation and maintenance equipment management method, system and equipment
CN118590370A
Cited By
Control method of data center and direct air trapping dynamic coupling system
CN121941020A