Predictive maintenance method and system of data center cooling system and electronic equipment
By leveraging the collaborative work of deep learning agents, mechanistic model agents, and historical experience agents, the operational status of data center cooling systems can be monitored and diagnosed in real time. This solves the problem of not being able to perceive system changes in real time in traditional maintenance methods, thereby improving the operational efficiency and accuracy of cooling systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional data center cooling system maintenance methods mainly rely on fixed-cycle manual inspections, which cannot detect subtle changes in the system's operating status in real time. This leads to potential faults being discovered only after they have evolved into serious problems, resulting in equipment damage and business interruption.
By employing the collaborative work of deep learning agents, mechanistic model agents, and historical experience agents, the system monitors and diagnoses the operating status of the cooling system in real time. It generates maintenance decisions by fusing the outputs of the three, including anomaly type, alarm level, and deviation, and generates alarm messages when necessary.
It enables real-time monitoring and anomaly diagnosis of the cooling system, improving operation and maintenance efficiency and accuracy, and reducing the risk of equipment damage and business interruption.
Smart Images

Figure CN121836684A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data center maintenance, in particular to a predictive maintenance method and system for a data center cooling system and an electronic device BACKGROUND
[0002] With the continuous growth of data center scale and power density, the cooling system has become a key infrastructure to ensure the continuous operation of IT equipment. The traditional data center cooling maintenance generally adopts a passive mode of "periodic inspection + post-failure maintenance": the operation and maintenance personnel on-site copy parameters such as temperature, pressure, and flow rate according to a fixed cycle (such as daily, weekly), or only after the monitoring system issues a threshold overrun alarm to troubleshoot. However, this mode has exposed the following problems in years of application: the traditional maintenance method is mainly based on fixed-cycle manual inspection, which cannot real-time perceive the subtle changes in system running state, resulting in potential faults often being discovered only after they have evolved into serious problems, causing equipment damage and business interruption. SUMMARY
[0003] The embodiments of the present application disclose a predictive maintenance method, system and electronic device for a data center cooling system, which can realize real-time monitoring and abnormal diagnosis of the running state of the cooling system through the collaborative work of a deep learning agent, a mechanism model agent and a historical experience agent, and can improve the operation and maintenance efficiency and accuracy of the cooling system.
[0004] The embodiments of the present application disclose a predictive maintenance method for a data center cooling system, characterized in that it comprises: acquiring running parameters of the cooling system through a network interface according to a first cycle; parallelly inputting the running parameters into a deep learning agent, a mechanism model agent and a historical experience agent to respectively obtain an abnormal pattern recognition result output by the deep learning agent, a thermodynamic first law physical constraint verification result output by the mechanism model agent, and a similar working condition historical optimal strategy matching result output by the historical experience agent; fusing the results output by the deep learning agent, the mechanism model agent and the historical experience agent to output a maintenance decision; the maintenance decision comprises an abnormal type, an alarm level and a deviation degree.
[0005] As an optional implementation, after the output of the maintenance decision, the method further comprises: determine whether human intervention is needed according to multi-dimensional rules and the maintenance decision; the multi-dimensional rules at least include one or more of the following rules: the alarm level is serious and the refrigeration main pipe temperature is greater than 18℃; the alarm level is serious and the refrigeration main pipe temperature is less than 8℃; the pressure difference is less than 80kPa; the outdoor temperature is greater than 38℃ and the refrigeration main pipe temperature is greater than 15℃; the temperature deviation rate is greater than 0.3 and the pressure difference deviation rate is greater than 0.2; generate an alarm message when it is determined that human intervention is needed; the alarm message includes alarm time, alarm type, alarm level, alarm details, current abnormal index, reason for not being suitable for automatic control, and human processing suggestion; send the alarm message to an operation and maintenance terminal device.
[0006] As an optional implementation, the deep learning intelligent agent uses prompt word engineering technology to embed real-time data into a preset natural language template, and then uses a large language model to perform semantic reasoning on the natural language template with embedded real-time data to identify potential abnormal type patterns and obtain the abnormal pattern recognition result.
[0007] As an optional implementation, the machine model intelligent agent establishes a device characteristic curve based on device rated parameters, and calculates whether operating parameters under different working conditions meet physical conservation constraints based on the first law of thermodynamics to obtain the first law of thermodynamics physical constraint verification result.
[0008] As an optional implementation, the historical experience intelligent agent retrieves a historical operation record closest to the operating parameters in a database through a wet bulb temperature matching algorithm to obtain the similar working condition historical optimal strategy matching result.
[0009] As an optional implementation, the alarm level is determined according to a deviation degree.
[0010] As an optional implementation, after the operating parameters of the cooling system are collected according to the first period, and before the operating parameters are input into the deep learning intelligent agent, the machine model intelligent agent, and the historical experience intelligent agent in parallel, the method further includes: determining whether the refrigeration main pipe temperature in the operating parameters is within a normal range, and whether the end pressure difference in the operating parameters is greater than a pressure difference threshold; when it is determined that the refrigeration main pipe temperature is out of the normal range, or the end pressure difference is less than or equal to the pressure difference threshold, performing the step of inputting the operating parameters into the deep learning intelligent agent, the machine model intelligent agent, and the historical experience intelligent agent in parallel.
[0011] Embodiments of the present application disclose a predictive maintenance system of a data center cooling system, the system comprising: The data acquisition layer is configured to acquire operation parameters of the cooling system through a network interface according to a first period; The intelligent analysis layer is configured to input the operation parameters into a deep learning agent, a machine model agent and a historical experience agent in parallel, to obtain an abnormal pattern recognition result output by the deep learning agent, a thermodynamic first law physical constraint verification result output by the machine model agent, and a similar working condition historical optimal strategy matching result output by the historical experience agent. The decision execution layer includes an abnormal diagnosis module, which is configured to fuse the results output by the deep learning agent, the machine model agent and the historical experience agent, and output a maintenance decision, wherein the maintenance decision includes an abnormal type, an alarm level and a deviation degree.
[0012] As an optional implementation, the decision execution layer further includes a manual intervention judgment module. The manual intervention judgment module is configured to determine whether manual intervention is needed according to a multi-dimensional rule and the maintenance decision, wherein the multi-dimensional rule includes at least one or more of the following rules: the alarm level is severe and the refrigeration main pipe temperature is greater than 18℃; the alarm level is severe and the refrigeration main pipe temperature is less than 8℃; the pressure difference is less than 80kPa; the outdoor temperature is greater than 38℃ and the refrigeration main pipe temperature is greater than 15℃; the temperature deviation rate is greater than 0.3 and the pressure difference deviation rate is greater than 0.2. The manual intervention judgment module is further configured to generate an alarm message when it is determined that manual intervention is needed, wherein the alarm message includes an alarm time, an alarm type, an alarm level, alarm details, a current abnormal index, a reason why automatic control is not suitable and a manual processing suggestion. The maintenance system further includes a human-computer interaction layer. The human-computer interaction layer is configured to send the alarm message to an operation and maintenance terminal device.
[0013] Embodiments of the present application disclose an electronic device, the electronic device includes a memory and a processor, the memory has a computer program stored therein, and the computer program is executed by the processor to make the processor implement any one of the predictive maintenance methods of the data center cooling system disclosed by the embodiments of the present application.
[0014] Embodiments of the present application disclose a computer readable storage medium, which has a computer program stored thereon, and the computer program is executed by a processor to implement any one of the predictive maintenance methods of the data center cooling system disclosed by the embodiments of the present application.
[0015] Compared with the related art, the embodiments of the present application have the following beneficial effects: By implementing the embodiments of the present application, the running parameters of each device of the cooling system can be collected in real time, so that through the collaborative work of the deep learning agent, the mechanism model agent and the historical experience agent, the real-time monitoring and abnormal diagnosis of the running state of the cooling system are realized, the outputs of the three are fused to obtain the final maintenance decision, and the operation and maintenance efficiency and accuracy of the cooling system can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 is a flowchart of a predictive maintenance method of a data center cooling system disclosed by the embodiments of the present application; Figure 2 is a flowchart of a predictive maintenance method of a data center cooling system disclosed by the embodiments of the present application; Figure 3 is a structural diagram of a predictive maintenance system of a data center cooling system disclosed by the embodiments of the present application; Figure 4 is a structural diagram of an electronic device disclosed by the embodiments of the present application. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] It should be noted that the terms "include" and "have" and any variations thereof in the embodiments of the present application and the drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed or can optionally include other steps or units inherent to the process, method, product or device.
[0020] The embodiments of the present application disclose a predictive maintenance method, system and electronic device of a data center cooling system, which can realize real-time monitoring, abnormal diagnosis and fault prediction of the running state of the cooling system through the collaborative work of a deep learning agent, a mechanism model agent and a historical experience agent. The following will be described in detail respectively.
[0021] A data center is a dedicated building constructed to centrally house and operate servers, storage, and network equipment. It houses a large number of IT (Information Technology) cabinets, operating continuously year-round and generating high-density heat. Its core objective is to ensure the IT equipment operates reliably and continuously in an environment with constant temperature, humidity, cleanliness, and uninterrupted power supply.
[0022] Data center cooling systems are responsible for removing heat from IT equipment in a timely manner, maintaining the temperature and humidity of the server room within permissible ranges (typically 18–27°C, relative humidity 40–60%). The system uses a combination of chilled water circulation, cooling water circulation, and air circulation to ultimately exhaust heat to the outdoor atmosphere, operating 24 / 7, 365 years a year. Therefore, after a period of operation, data center cooling systems are prone to various anomalies, requiring monitoring of their operational status and timely handling of any abnormalities to ensure the stable operation of the data center.
[0023] Please see Figure 1 , Figure 1 This is a flowchart illustrating a predictive maintenance method for a data center cooling system disclosed in an embodiment of this application. Figure 1 The method shown can be executed by electronic devices with computing capabilities, such as personal computers and industrial computers. Figure 1 As shown, the method may include the following steps: 110. Collect the operating parameters of the cooling system through the network interface according to the first cycle.
[0024] In this embodiment, the first cycle can be set according to actual needs and is not specifically limited. Optionally, considering the rapid and instantaneous changes in the operating status of the cooling center, the first cycle can be set to 1 minute. The network interface refers to the communication interface between the electronic device and the various devices included in the cooling system, such as an API (Application Programming Interface) that can be called via HTTP requests.
[0025] The operating parameters of the cooling system can include any parameter that can characterize the operating status of the cooling system, including but not limited to: Environmental parameters include: outdoor temperature and wet-bulb temperature; where wet-bulb temperature refers to the temperature detected by an ultrasonic humidity probe placed outdoors on a ventilated dry and wet surface. IT load, including: the power output of smart meters for IT equipment; Refrigeration main pipe parameters, including the temperature and pressure differential of the refrigeration main pipe; Parameters of plate heat exchangers, including: the start-up and shutdown status of the plate heat exchanger. Parameters of the chilled pump or cooling pump, including: operating power and number of units in operation; Cooling tower parameters include: fan operating frequency, number of operating units, and return water temperature.
[0026] 102. Input the running parameters in parallel to the deep learning agent, the machine model agent, and the historical experience agent to obtain the abnormal pattern recognition results output by the deep learning agent, the physical constraint verification results of the first law of thermodynamics output by the machine model agent, and the historical optimal strategy matching results for similar working conditions output by the historical experience agent.
[0027] In this embodiment, the deep learning agent is a large language model, such as the commercially available Deepseek large language model or Tongyi Qianwen large language model. The deep learning agent can perform semantic reasoning on the received operating parameters to obtain preliminary abnormal pattern recognition results. Optionally, the deep learning agent can obtain real-time operating parameters of the cooling system every minute, such as chilled water outlet temperature, terminal pressure difference, IT load, outdoor temperature, and configuration information of the chiller, water pump, and cooling tower. Then, using prompt word engineering technology, the real-time data is embedded into a preset natural language template to obtain a segment embedded with the real-time operating parameters, such as: "chilled water outlet temperature ××℃, terminal pressure difference ×× kPa, IT load ×× kW, outdoor temperature ××℃". The deep learning agent can further perform semantic reasoning on this segment to identify potential abnormal type patterns and obtain preliminary abnormal pattern recognition results. Among them, the deep learning agent can use a preliminary threshold judgment method for semantic reasoning. For example, when the temperature of the chilled water outlet exceeds the normal range, the abnormality type can be initially judged as an abnormal model of excessively high or low temperature. The preliminary identification of other operating parameters is similar to that of temperature, and will not be elaborated on in the following content.
[0028] The machine model agent can process operating parameters according to the first law of thermodynamics. First, deviations can be calculated using the following formulas: temperature deviation rate = |actual temperature - 13 °C| / 13, and pressure difference deviation rate = |actual pressure difference - 98 kPa| / 98. It should be noted that, unless otherwise specified, deviations in this application refer to temperature deviations and / or pressure difference deviations. Then, the machine model agent can further substitute the real-time number and frequency of chillers, refrigeration pumps, cooling pumps, and cooling towers into the equipment characteristic curves, and calculate the theoretically required power consumption and cooling capacity based on the first law of thermodynamics. The equipment characteristic curves are established based on the rated parameters of each device included in the cooling system. If the measured values of the cooling system's operating parameters exceed the physical upper and lower limits (i.e., the theoretical power consumption and cooling capacity calculated based on the equipment parameters), or if both deviation rates simultaneously exceed the corresponding preset thresholds, it indicates that the cooling system's operating parameters under the current operating conditions do not meet the physical conservation constraints. If the measured values of the cooling system's operating parameters do not exceed the physical upper and lower limits, and neither of the aforementioned deviation rates exceeds the limits (less than or equal to the preset thresholds), then it is determined that the cooling system's operating parameters under the current operating conditions meet the physical conservation constraints. Therefore, the machine model agent can verify whether the operation of the cooling equipment meets the physical conservation constraints of the first law of thermodynamics based on the operating parameters and equipment characteristic curves, and obtain the verification results.
[0029] The historical experience agent is used for retrieval in the database. Upon receiving new operating parameters, the agent uses these parameters as the search key to retrieve historical operating records of similar conditions. Optionally, the agent can use a wet-bulb temperature matching algorithm to retrieve the historical operating record in the database that most closely matches the operating parameters, thus obtaining the historical optimal strategy matching result for similar operating conditions. The number of closest historical operating records can be one or more, depending on actual needs, and is not limited in specific terms.
[0030] 103. The results output by the deep learning agent, the machine model agent, and the historical experience agent are fused to output maintenance decisions; the maintenance decisions include anomaly type, alarm level, and deviation.
[0031] In this embodiment, the deep learning agent outputs anomaly pattern recognition results, the machine model agent outputs physical constraint verification results based on the first law of thermodynamics, and the historical experience agent outputs historical optimal strategy matching results for similar operating conditions. The anomaly pattern recognition results include the anomaly type. The physical constraint verification results include not only a judgment on whether the physical constraints are met, but also the previously calculated deviation of the temperature-pressure difference. The alarm level can be determined based on the calculated deviation and the specific values of the received operating parameters. The alarm level is an indicator of the urgency level, and can be categorized from mild to severe as abnormal, urgent, and critical; the higher the deviation, the more severe the alarm level. Finally, the output maintenance decision can be a set of anomaly type, alarm level, and deviation.
[0032] As can be seen, by implementing the embodiments of this application, the operating parameters of each device in the cooling system can be collected in real time. Through the collaborative work of deep learning agents, mechanism model agents and historical experience agents, the operating status of the cooling system can be monitored in real time and anomaly diagnosis can be achieved, which can improve the operation and maintenance efficiency and accuracy of the cooling system.
[0033] Please see Figure 2 , Figure 2 This is a flowchart illustrating a predictive maintenance method for a data center cooling system disclosed in an embodiment of this application. Figure 2 The method shown can be executed by electronic devices with computing capabilities, such as personal computers and industrial computers. Figure 2 As shown, the method may include the following steps: 210. Collect the operating parameters of the cooling system through the network interface according to the first cycle.
[0034] 220. Determine whether the temperature of the main refrigeration pipe in the operating parameters is within the normal range, and whether the terminal pressure difference in the operating parameters is greater than the pressure difference threshold; if either is not, proceed to step 230; if both are yes, end the process.
[0035] In this embodiment, the normal range can be determined based on the performance indicators that the cooling main pipe can achieve under normal operating conditions, for example, it can be set to 12-14℃, which is the normal operating temperature of the cooling main pipe. Similarly, the differential pressure threshold can be determined based on the differential pressure value under normal operating conditions, for example, it can be set to 98 kPa. These two indicators are the most critical indicators of the cooling system. If either of these two indicators is abnormal, there is a high probability that there is a problem with the cooling system. The following steps will be performed to further investigate the cause of the abnormality in the cooling system.
[0036] It should be noted that in some other possible embodiments, step 220 is not mandatory but optional. Executing step 220 can pre-determine whether further anomaly identification procedures are necessary, thereby achieving a balance between accuracy and energy consumption.
[0037] 230. Input the running parameters in parallel to the deep learning agent, the machine model agent, and the historical experience agent to obtain the abnormal pattern recognition results output by the deep learning agent, the physical constraint verification results of the first law of thermodynamics output by the machine model agent, and the historical optimal strategy matching results for similar working conditions output by the historical experience agent.
[0038] 240. The results output by the deep learning agent, the machine model agent, and the historical experience agent are fused to output maintenance decisions; the maintenance decisions include anomaly type, alarm level, and deviation.
[0039] For the implementation methods of steps 210 and 230-240, please refer to steps 110-130. The following content will not be repeated.
[0040] 250. Determine whether manual intervention is needed based on multi-dimensional rules and maintenance decisions; if so, proceed to step 260; if not, end the process.
[0041] In this embodiment, the multi-dimensional rules can be a summary of cooling system operation and maintenance experience, and their indicators can be determined based on the rated parameters and / or ideal operating performance of each device. Optionally, the multi-dimensional rules may include at least one or more of the following rules: alarm level is severe and refrigeration main pipe temperature > 18°C; alarm level is severe and refrigeration main pipe temperature < 8°C; pressure difference < 80 kPa; outdoor temperature > 38°C and refrigeration main pipe temperature > 15°C; temperature deviation rate > 0.3 and pressure difference deviation rate > 0.2.
[0042] If any one of the above rules is met, it can be determined that manual intervention is required. Otherwise, if none of the rules are met, it can be determined that no manual intervention is required.
[0043] 260. Generate alarm messages.
[0044] In this embodiment, the alarm message may include any one or more types of information such as text, images, audio, and video. The alarm message includes the alarm time, alarm type, alarm level, alarm details, current abnormal indicators, reasons why automatic control is unsuitable, and manual handling suggestions. Alarm time can be automatically generated based on system time; alarm type can be determined based on the anomaly type in operation and maintenance decisions, and can be the same as the anomaly type; or generated according to the anomaly type according to preset mapping rules. Alarm levels can be derived from operational decisions, and alarm details can include further alarm content descriptions. These can be generated based on template text bound to the alarm level or generated by a deep learning agent through a large language model. The current abnormal indicators refer to the indicators that meet any of the aforementioned rules, such as chilled water main outlet temperature, terminal pressure difference, IT load, outdoor temperature, and outdoor wet-bulb temperature. Reasons for unsuitability for automatic control can be generated based on templates bound to different abnormal indicator types; Manual handling suggestions may include, but are not limited to, one or more of the following combinations: 1) Immediately go to the site to check the equipment's operating status; 2) Determine the root cause of the abnormality based on the actual situation; 3) Take targeted maintenance or adjustment measures; 4) Record the handling results after confirming that the system has returned to normal.
[0045] 270. Send alarm messages to the operation and maintenance terminal equipment.
[0046] In this embodiment, the maintenance terminal device is any type of terminal device that communicates with electronic devices, and it is kept by the maintenance personnel of the cooling system. After an alarm message is sent to the maintenance terminal device, the maintenance personnel can promptly view the anomaly details and suggested handling methods through the terminal device to ensure a timely response.
[0047] As can be seen, by implementing the embodiments of this application, in addition to achieving real-time monitoring and anomaly diagnosis of the cooling system's operating status through the collaborative work of deep learning agents, mechanism model agents, and historical experience agents, the accuracy of anomaly diagnosis and energy consumption can be balanced by using the temperature of the main cooling pipe and the pressure difference at the terminal as preliminary judgment conditions; or, furthermore, the anomaly can be judged based on operation and maintenance decisions to determine whether manual intervention is required. When manual intervention is required, a detailed alarm message is generated and sent to the operation and maintenance terminal device, enabling operation and maintenance personnel to view and respond in a timely manner through the terminal device.
[0048] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of one embodiment of d disclosed in this application. This system can be applied to any of the aforementioned electronic devices. Figure 3 As shown, the system includes: The data acquisition layer 310 is used to acquire the operating parameters of the cooling system through the network interface according to the first cycle; The intelligent analysis layer 320 is used to input the running parameters in parallel to the deep learning agent, the machine model agent, and the historical experience agent, and obtain the abnormal pattern recognition results output by the deep learning agent, the physical constraint verification results of the first law of thermodynamics output by the machine model agent, and the historical optimal strategy matching results for similar working conditions output by the historical experience agent. The decision execution layer 330 includes an anomaly diagnosis module 3310. This module fuses the outputs of the deep learning agent, the machine learning agent, and the historical experience agent to output a maintenance decision. The maintenance decision includes the anomaly type, alarm level, and deviation degree. The selectable alarm level is determined based on the deviation degree.
[0049] In some optional embodiments, the decision execution layer 330 may further include: a manual intervention judgment module 3320, used to determine whether manual intervention is needed based on multi-dimensional rules and maintenance decisions; the multi-dimensional rules include at least one or more of the following rules: alarm level is severe and refrigeration main pipe temperature > 18°C; alarm level is severe and refrigeration main pipe temperature < 8°C; pressure difference < 80 kPa; outdoor temperature > 38°C and refrigeration main pipe temperature > 15°C; temperature deviation rate > 0.3 and pressure difference deviation rate > 0.2; The manual intervention judgment module 3320 is also used to generate an alarm message when it is determined that manual intervention is required; the alarm message includes alarm time, alarm type, alarm level, alarm details, current abnormal indicators, reasons why it is not suitable for automatic control, and manual handling suggestions; Accordingly, the maintenance system also includes: a human-computer interaction layer 340; The human-machine interaction layer 340 is used to send alarm messages to the operation and maintenance terminal equipment.
[0050] Optionally, the aforementioned intelligent analysis layer 320 can be used to embed real-time data into a preset natural language template using prompt word engineering technology through a deep learning agent, and then use a large language model to perform semantic reasoning on the natural language template embedded with real-time data to identify potential abnormal type patterns and obtain abnormal pattern recognition results.
[0051] Optionally, the aforementioned intelligent analysis layer 320 can also be used to establish equipment characteristic curves based on the rated parameters of the equipment through a machine model intelligent agent, and calculate whether the operating parameters under different operating conditions meet the physical conservation constraints based on the first law of thermodynamics, thereby obtaining the physical constraint verification results of the first law of thermodynamics.
[0052] Optionally, the aforementioned intelligent analysis layer 320 can also be used to retrieve historical operating records in the database that are closest to the operating parameters through a wet-bulb temperature matching algorithm using a historical experience agent, so as to obtain the historical optimal strategy matching result for similar operating conditions.
[0053] As an optional implementation, after the data acquisition layer 310 acquires the operating parameters of the cooling system according to the first cycle, and before the intelligent analysis layer 320 inputs the operating parameters in parallel to the deep learning agent, the machine model agent, and the historical experience agent, the intelligent analysis layer 320 may first determine whether the temperature of the main cooling pipe in the operating parameters is within the normal range, and whether the terminal pressure difference in the operating parameters is greater than the pressure difference threshold; and, if it is determined that the temperature of the main cooling pipe is outside the normal range, or the terminal pressure difference is less than or equal to the pressure difference threshold, then the operation of inputting the operating parameters in parallel to the deep learning agent, the machine model agent, and the historical experience agent is executed.
[0054] As can be seen, the predictive maintenance system for the data center cooling system disclosed in this application can achieve real-time monitoring and anomaly diagnosis of the cooling system's operating status through the collaborative work of deep learning agents, mechanism model agents, and historical experience agents. It can also balance the accuracy of anomaly diagnosis and energy consumption by using the temperature of the main cooling pipe and the pressure difference at the terminal as preliminary judgment conditions. Alternatively, it can further determine whether the anomaly requires manual intervention based on operation and maintenance decisions. When manual intervention is required, it can generate detailed alarm messages and send them to the operation and maintenance terminal device, enabling operation and maintenance personnel to view and respond in a timely manner through the terminal device.
[0055] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. For example... Figure 4 As shown, the electronic device may include: Memory 410 containing computer programs; Processor 420 coupled to memory 410; When the computer program stored in the memory 410 is executed by the processor 420, the processor 420 performs any of the identification methods for children with neurodevelopmental disorders disclosed in the embodiments of this application.
[0056] It should be noted that the electronic device shown in the figure may also include components not shown, such as a power supply, input buttons, screen, RF circuit, Wi-Fi module, and Bluetooth module, which will not be described in detail in this embodiment.
[0057] This application discloses a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements any of the predictive maintenance methods for data center cooling systems disclosed in this application.
[0058] This application discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform any of the predictive maintenance methods for data center cooling systems disclosed in this application.
[0059] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Those skilled in the art should also recognize that the embodiments described in the specification are optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0060] In the various embodiments of this application, it should be understood that the sequence number of each process does not necessarily imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0061] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they can be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0062] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0063] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-accessible memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests to cause a computer device (which can be a personal computer, server, or network device, specifically a processor in the computer device) to execute some or all of the steps of the methods described in the various embodiments of this application.
[0064] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compactdisc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0065] The predictive maintenance method, system, and electronic equipment for data center cooling systems disclosed in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A predictive maintenance method for a data center cooling system, characterized in that, include: The operating parameters of the cooling system are collected via the network interface in the first cycle. The operating parameters are input in parallel to the deep learning agent, the machine model agent, and the historical experience agent to obtain the abnormal pattern recognition results output by the deep learning agent, the physical constraint verification results of the first law of thermodynamics output by the machine model agent, and the historical optimal strategy matching results for similar working conditions output by the historical experience agent. The results output by the deep learning agent, the machine model agent, and the historical experience agent are fused to output a maintenance decision. The maintenance decisions include anomaly type, alarm level, and deviation.
2. The method according to claim 1, characterized in that, Following the output maintenance decision, the method further includes: The need for manual intervention is determined based on multi-dimensional rules and the maintenance decision; the multi-dimensional rules include at least one or more of the following: alarm level is severe and refrigeration main pipe temperature > 18℃; alarm level is severe and refrigeration main pipe temperature < 8℃; pressure difference < 80kPa; outdoor temperature > 38℃ and refrigeration main pipe temperature > 15℃; temperature deviation rate > 0.3 and pressure difference deviation rate > 0.2; When it is determined that manual intervention is required, an alarm message is generated; the alarm message includes alarm time, alarm type, alarm level, alarm details, current abnormal indicators, reasons why it is not suitable for automatic control, and manual handling suggestions; The alarm message is sent to the operation and maintenance terminal device.
3. The method according to claim 1 or 2, characterized in that, The deep learning agent uses prompt word engineering technology to embed real-time data into a preset natural language template, and then uses a large language model to perform semantic reasoning on the natural language template embedded with real-time data to identify potential abnormal type patterns and obtain the abnormal pattern identification results.
4. The method according to claim 1 or 2, characterized in that, The machine model agent establishes the equipment characteristic curve based on the equipment's rated parameters, and calculates whether the operating parameters meet the physical conservation constraints under different operating conditions based on the first law of thermodynamics, thereby obtaining the physical constraint verification results of the first law of thermodynamics.
5. The method according to claim 1 or 2, characterized in that, The historical experience agent retrieves the historical operating record in the database that is closest to the operating parameters using a wet-bulb temperature matching algorithm, and obtains the historical optimal strategy matching result for the similar operating conditions.
6. The method according to claim 1 or 2, characterized in that, The alarm level is determined based on the degree of deviation.
7. The method according to claim 1, characterized in that, After collecting the operating parameters of the cooling system according to the first cycle, and before inputting the operating parameters in parallel to the deep learning agent, the machine model agent, and the historical experience agent, the method further includes: Determine whether the temperature of the main refrigeration pipe in the operating parameters is within the normal range, and whether the terminal pressure difference in the operating parameters is greater than the pressure difference threshold; When it is determined that the temperature of the main freezing pipe exceeds the normal range, or the terminal pressure difference is less than or equal to the pressure difference threshold, the step of inputting the operating parameters in parallel to the deep learning agent, the machine model agent, and the historical experience agent is executed.
8. A predictive maintenance system for a data center cooling system, characterized in that, The system includes: The data acquisition layer is used to collect the operating parameters of the cooling system through the network interface according to the first cycle. The intelligent analysis layer is used to input the operating parameters in parallel to the deep learning agent, the machine model agent, and the historical experience agent, and obtain the abnormal pattern recognition results output by the deep learning agent, the physical constraint verification results of the first law of thermodynamics output by the machine model agent, and the historical optimal strategy matching results for similar working conditions output by the historical experience agent, respectively. The decision execution layer includes an anomaly diagnosis module; the anomaly diagnosis module is used to fuse the outputs of the deep learning agent, the machine model agent, and the historical experience agent, and output a maintenance decision; the maintenance decision includes anomaly type, alarm level, and deviation degree.
9. The system according to claim 8, characterized in that, The decision execution layer also includes: a human intervention judgment module; The manual intervention judgment module is used to determine whether manual intervention is needed based on multi-dimensional rules and the maintenance decision; the multi-dimensional rules include at least one or more of the following rules: alarm level is severe and refrigeration main pipe temperature > 18℃; alarm level is severe and refrigeration main pipe temperature < 8℃; pressure difference < 80kPa; outdoor temperature > 38℃ and refrigeration main pipe temperature > 15℃; temperature deviation rate > 0.3 and pressure difference deviation rate > 0.2; The manual intervention judgment module is also used to generate an alarm message when it is determined that manual intervention is required; the alarm message includes alarm time, alarm type, alarm level, alarm details, current abnormal indicators, reasons why it is not suitable for automatic control, and manual handling suggestions; The maintenance system also includes: a human-computer interaction layer; The human-machine interaction layer is used to send the alarm message to the operation and maintenance terminal equipment.
10. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores a computer program, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1-7.