Fault handling methods, apparatus, electronic devices and computer-readable storage media

By acquiring the status parameter information of the cooling equipment, the system can determine host failures and perform intelligent switching, thus solving the problem of excessively high temperatures after a cooling system failure and achieving efficient cooling and equipment reliability in the data center.

CN115451524BActive Publication Date: 2025-10-28SHANGHAI MEICON INTELLIGENT CONSTR CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211065150.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-10-28
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

When the cooling system's refrigeration equipment malfunctions, the delayed start-up of the equipment poses a risk of excessively high temperatures, affecting the reliability of the data center's computing equipment.

Method used

By acquiring the status parameter information of the refrigeration equipment, it is determined whether there is a main unit failure, and the corresponding fault handling strategy is determined according to the number of failures, so as to achieve seamless intelligent switching between the faulty main unit and the normal main unit. At the same time, when the cooling capacity is insufficient, the cold storage device and the cold energy in the cooling pipes are used for cold energy compensation.

Benefits of technology

It improved fault handling efficiency, maintained the cooling effect of the data center, and ensured the reliability and stable operation of computing equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115451524B_ABST
    Figure CN115451524B_ABST
Patent Text Reader

Abstract

This application provides a fault handling method, apparatus, electronic device, and computer-readable storage medium. It collects status parameter information of each cooling device in a data center, then determines whether a host failure has occurred based on the status parameter information. In the event of a host failure, a corresponding fault handling strategy is determined based on the number of cooling devices experiencing host failures. Different fault handling strategies are adopted for different numbers of failures, achieving seamless intelligent switching between failed and normal hosts. Furthermore, in cases where too many failed devices cause insufficient cooling capacity or cooling failure, this solution utilizes continuously running water pumps and a predetermined number of pumps to rationally utilize the cold storage device and the remaining cold energy in the pipes for cooling compensation, thereby providing cooling assurance in extreme situations. Additionally, this solution provides emergency strategies to address the issues of poor stability during equipment signal drops or full host failures, maintaining the stable operation of the data center equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of refrigeration control technology, and more specifically, to a fault handling method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] Data center computing equipment, such as servers and network devices inside the computer room, releases a lot of heat during operation. If heat dissipation is not timely, the equipment may overheat and crash.

[0003] Currently, data centers are generally cooled by cooling systems. Conventional data center cooling systems require emergency activation of backup equipment or cold storage devices after a cooling equipment failure. This approach carries the risk of overheating due to the delayed activation of equipment. Summary of the Invention

[0004] The purpose of this application is to provide a fault handling method, apparatus, electronic device, and computer-readable storage medium to solve the problem of excessively high temperature caused by delayed start-up of refrigeration equipment in current cooling systems after a fault occurs.

[0005] In a first aspect, the present invention provides a fault handling method, the method comprising: acquiring status parameter information of each of a plurality of refrigeration devices; determining, based on the status parameter information of each refrigeration device, whether at least one of the plurality of refrigeration devices has a host failure; if so, determining a corresponding fault handling strategy based on the number of refrigeration devices with host failures, and performing fault handling according to the corresponding fault handling strategy.

[0006] In the fault handling method designed above, this solution collects the status parameter information of each cooling device in the data center, and then determines whether the cooling device has experienced a host failure based on the status parameter information. In the case of a host failure, the corresponding fault handling strategy is determined according to the number of cooling devices with host failures. Thus, different fault handling strategies are adopted for different numbers of failures, thereby achieving seamless intelligent switching between the failed host and the normal host. This solves the problem of excessive temperature caused by the delayed start-up of the cooling devices after the current cooling system fails, improves fault handling efficiency, maintains the cooling effect of the data center, and ensures the reliability of the data center computing equipment.

[0007] In an optional embodiment of the first aspect, the status parameter information includes a device fault signal. Determining whether at least one of the multiple refrigeration devices has a main unit fault based on the status parameter information of each refrigeration device includes: determining whether there is a refrigeration device among the multiple refrigeration devices whose device fault signal is a preset fault signal based on the device fault signal of each refrigeration device; if so, determining that the main unit of the refrigeration device whose device fault signal is a preset fault signal has a fault.

[0008] In an optional implementation of the first aspect, the status parameter information includes local / remote status signals. Determining whether at least one of the multiple refrigeration devices has experienced a host failure based on the status parameter information of each refrigeration device includes: determining whether there is a refrigeration device among the multiple refrigeration devices whose local / remote status signal is a local status signal based on the local / remote status signal of each refrigeration device; if so, determining that the refrigeration device whose local / remote status signal is a local status signal has experienced a communication failure and its host signal has dropped.

[0009] In an optional embodiment of the first aspect, the status parameter information further includes device status information, and a corresponding fault handling strategy is determined based on the number of cooling devices that have experienced host failure, including: determining the running cooling devices based on the device status information of each device to obtain the number of running cooling devices; and determining a corresponding fault handling strategy based on the number of running cooling devices and the number of cooling devices that have experienced host failure.

[0010] In an optional embodiment of the first aspect, a corresponding fault handling strategy is determined based on the number of operating refrigeration units and the number of refrigeration units experiencing main unit failure, including: determining whether the number of operating refrigeration units is the total number of the multiple refrigeration units; if the number of operating refrigeration units is the total number of the multiple refrigeration units, then controlling the main unit of the refrigeration unit experiencing main unit failure to shut down, and controlling the chilled water pump corresponding to the refrigeration unit experiencing main unit failure to keep running, wherein one refrigeration unit and one chilled water pump are connected to each other through a corresponding pipeline.

[0011] In an optional embodiment of the first aspect, a corresponding fault handling strategy is determined based on the number of operating refrigeration equipment and the number of refrigeration equipment experiencing main unit failure, including: determining whether the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment; if the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment, then controlling the main unit of the refrigeration equipment experiencing main unit failure to shut down, and controlling all chilled water pumps to keep running, wherein all refrigeration equipment and all chilled water pumps are connected through a pipeline.

[0012] In an optional embodiment of the first aspect, after determining whether the number of operating refrigeration devices is the total number of the plurality of refrigeration devices, the method further includes: if the number of operating refrigeration devices is not the total number of the plurality of refrigeration devices, then determining whether the number of operating refrigeration devices is 0; if the number of operating refrigeration devices is 0, then determining whether the number of refrigeration devices with host failure is the total number of the plurality of refrigeration devices; if the number of refrigeration devices with host failure is not the total number of the plurality of refrigeration devices, then controlling the start of a preset number of fault-free refrigeration devices; if the number of refrigeration devices with host failure is the total number of the plurality of refrigeration devices, then controlling the chilled water pumps of a target number of refrigeration devices to start.

[0013] In an optional implementation of the first aspect, a corresponding fault handling strategy is determined based on the number of operating refrigeration units and the number of refrigeration units experiencing main unit failures, including: determining whether the number of refrigeration units experiencing main unit failures is less than the number of operating refrigeration units; if the number of refrigeration units experiencing main unit failures is less than the number of operating refrigeration units, determining whether any of the operating refrigeration units experiencing main unit failures are refrigeration units experiencing main unit failures; if any of the operating refrigeration units experiencing main unit failures are refrigeration units experiencing main unit failures, obtaining the number of refrigeration units experiencing main unit failures and currently in operation; controlling the start-up of an equal number of fault-free standby refrigeration units based on the number of failures, and controlling the main unit and corresponding chilled water pump of the refrigeration units experiencing main unit failures and currently in operation to shut down.

[0014] In an optional embodiment of the first aspect, after determining whether the number of refrigeration units experiencing host malfunctions is less than the number of refrigeration units currently in operation, the method further includes: if the number of refrigeration units experiencing host malfunctions is not less than the number of refrigeration units currently in operation, then determining whether the number of refrigeration units experiencing host malfunctions is the total number of multiple refrigeration units; if the number of refrigeration units experiencing host malfunctions is not the total number of multiple refrigeration units, then controlling all standby refrigeration units without malfunctions to start; controlling the host units of the refrigeration units experiencing host malfunctions and currently in operation to stop, and controlling a first number of chilled water pumps to stop, wherein the difference between the total number of multiple refrigeration units and the first number is the number of refrigeration units currently in operation.

[0015] In an optional embodiment of the first aspect, after determining whether the number of refrigeration devices experiencing host failure is the total number of multiple refrigeration devices, the method further includes: if the number of refrigeration devices experiencing host failure is the total number of multiple refrigeration devices, then controlling the host machines of all operating refrigeration devices to shut down, and controlling the chilled water pumps of the target number of refrigeration devices to start.

[0016] In an optional implementation of the first aspect, the method further includes: obtaining the total power of all computing devices currently operating in the data center; and determining the required number of chilled water pumps to be turned on based on the total power of all currently operating computing devices to obtain the target number.

[0017] In the implementation of the above design, this solution compensates for insufficient cooling capacity or cooling failure caused by too many faulty devices by utilizing the remaining cold energy in the cold storage tank and / or cooling pipes of the refrigeration equipment, thereby providing cooling protection in extreme cases. In addition, this solution also provides emergency strategies for the problems of poor stability when equipment signal drops or the entire host fails, to maintain the stable operation of the computer room equipment and thus ensure the operational reliability of the data center computing equipment.

[0018] Secondly, the present invention provides a fault handling device, the device comprising: an acquisition module for acquiring status parameter information of each of a plurality of refrigeration devices; a determination module for determining, based on the status parameter information of each refrigeration device, whether at least one refrigeration device among the plurality of refrigeration devices has a host failure; and a determination module for determining, after determining that at least one refrigeration device among the plurality of refrigeration devices has a host failure, a corresponding fault handling strategy based on the number of refrigeration devices with host failures, so as to perform fault handling according to the corresponding fault handling strategy.

[0019] The fault handling device designed above collects status parameter information of each cooling device in the data center, and then determines whether the cooling device has experienced a host failure based on the status parameter information. In the event of a host failure, the corresponding fault handling strategy is determined according to the number of cooling devices experiencing host failures. Thus, different fault handling strategies are adopted for different numbers of failures, thereby achieving seamless intelligent switching between failed and normal hosts. This solves the problem of excessive temperature caused by delayed equipment startup after a failure in the current cooling system, improves fault handling efficiency, maintains the cooling effect of the data center, and ensures the reliability of the data center computing equipment.

[0020] In an optional embodiment of the second aspect, the status parameter information includes equipment fault signals. The determination module is specifically used to determine whether there is a refrigeration device among the multiple refrigeration devices whose equipment fault signal is a preset fault signal based on the equipment fault signal of each refrigeration device; if so, it is determined that the host device of the refrigeration device whose equipment fault signal is a preset fault signal has failed.

[0021] In an optional implementation of the second aspect, the status parameter information includes local / remote status signals. The determination module is further specifically used to determine whether there is a refrigeration device among the multiple refrigeration devices whose local / remote status signal is a local status signal based on the local / remote status signal of each refrigeration device; if so, it is determined that the communication host signal of the refrigeration device whose local / remote status signal is a local status signal has failed.

[0022] In an optional implementation of the second aspect, the status parameter information further includes device status information. Specifically, the determining module is used to determine the running refrigeration equipment based on the device status information of each device to obtain the number of running refrigeration equipment; and to determine the corresponding fault handling strategy based on the number of running refrigeration equipment and the number of refrigeration equipment that have experienced host failure.

[0023] In an optional embodiment of the second aspect, the determining module is further specifically used to determine whether the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment; if the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment, then the host of the refrigeration equipment with host failure is shut down, and the chilled water pump corresponding to the refrigeration equipment with host failure is kept running, wherein one refrigeration equipment and one chilled water pump are connected to each other through a pipeline.

[0024] In an optional embodiment of the second aspect, the determining module is further specifically used to determine whether the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment; if the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment, then the host of the refrigeration equipment with host failure is shut down, and all chilled water pumps are kept running, wherein all refrigeration equipment and all chilled water pumps are connected through a pipeline.

[0025] In an optional embodiment of the second aspect, the determining module is further specifically configured to: if the number of operating refrigeration devices is not the total number of the multiple refrigeration devices, determine whether the number of operating refrigeration devices is 0; if the number of operating refrigeration devices is 0, determine whether the number of refrigeration devices with host failure is the total number of the multiple refrigeration devices; if the number of refrigeration devices with host failure is not the total number of the multiple refrigeration devices, control the start of a preset number of fault-free refrigeration devices; if the number of refrigeration devices with host failure is the total number of the multiple refrigeration devices, control the start of the chilled water pumps of the target number of refrigeration devices.

[0026] In an optional embodiment of the second aspect, the determining module is further specifically used to determine whether the number of refrigeration units with host failures is less than the number of refrigeration units currently in operation; if the number of refrigeration units with host failures is less than the number of refrigeration units currently in operation, then determine whether there are any refrigeration units with host failures among the currently operating refrigeration units; if there are any refrigeration units with host failures among the currently operating refrigeration units, then obtain the number of refrigeration units with host failures that are currently in operation; control the start of an equal number of standby refrigeration units without failures according to the number of refrigeration units with failures, and control the host units and corresponding chilled water pumps of the refrigeration units with host failures that are currently in operation to stop.

[0027] In an optional embodiment of the second aspect, the determining module is further specifically configured to: if the number of refrigeration devices experiencing host failures is not less than the number of refrigeration devices currently in operation, determine whether the number of refrigeration devices experiencing host failures is the total number of multiple refrigeration devices; if the number of refrigeration devices experiencing host failures is not the total number of multiple refrigeration devices, control all standby refrigeration devices without failures to start; control the host of the refrigeration devices experiencing host failures and currently in operation to stop, and control a first number of chilled water pumps to stop, wherein the difference between the total number of multiple refrigeration devices and the first number is the number of refrigeration devices currently in operation.

[0028] In an optional implementation of the second aspect, the determining module is further specifically configured to, if the number of refrigeration units experiencing host failure is the total number of multiple refrigeration units, control the host units of all operating refrigeration units to shut down, and control the chilled water pumps of the target number of refrigeration units to start.

[0029] In an optional implementation of the second aspect, the acquisition module is further configured to acquire the total power of all computing devices currently operating in the data center; the determination module is further configured to determine the required number of chilled water pumps to be turned on based on the total power of all currently operating computing devices, so as to obtain the target number.

[0030] Thirdly, this application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the method described in the first aspect or any optional implementation thereof.

[0031] Fourthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the method described in the first aspect or any optional implementation thereof.

[0032] Fifthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the method described in the first aspect or any optional implementation thereof.

[0033] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0034] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a first flowchart illustrating the fault handling method provided in an embodiment of this application;

[0036] Figure 2 This is a second flowchart illustrating the fault handling method provided in the embodiments of this application;

[0037] Figure 3 A schematic diagram of the third process of the fault handling method provided in the embodiments of this application;

[0038] Figure 4 A schematic diagram of the fourth process of the fault handling method provided in the embodiments of this application;

[0039] Figure 5 A schematic diagram of the fifth process of the fault handling method provided in the embodiments of this application;

[0040] Figure 6 This is a schematic diagram of the fault handling device provided in the embodiments of this application;

[0041] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0042] Icons: 600 - Acquisition module; 610 - Judgment module; 620 - Confirmation module; 7 - Electronic device; 701 - Processor; 702 - Memory; 703 - Communication bus. Detailed Implementation

[0043] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.

[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0045] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0046] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0047] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0048] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).

[0049] In the description of the embodiments of this application, the technical terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of this application and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this application.

[0050] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.

[0051] Data center computing equipment, such as servers and network devices inside the computer room, releases a lot of heat during operation. If heat dissipation is not timely, the equipment may overheat and crash.

[0052] The inventors have noted that data centers are typically cooled by cooling systems such as air conditioning systems / water cooling systems to prevent equipment from overheating and causing downtime. When the cooling equipment in the cooling system fails, the current practice is to urgently activate backup equipment or cold storage devices to supplement cooling. However, this approach carries the risk of overheating due to the delayed activation of the equipment.

[0053] In response, this applicant proposes a fault handling method. This method collects status parameter information of each refrigeration unit from the data, and then determines whether a main unit failure has occurred based on the status parameter information. In the event of a main unit failure, the corresponding fault handling strategy is determined according to the number of refrigeration units with main unit failures. Thus, different fault handling strategies are adopted for different numbers of failures, thereby achieving seamless intelligent switching between faulty and normal main units. At the same time, this solution also compensates for insufficient cooling capacity or cooling failure caused by too many faulty units by continuously running water pumps and using a predetermined number of units, rationally utilizing the cold storage device and the remaining cold capacity in the pipeline to provide cooling guarantee in extreme cases. In addition, this solution also provides emergency strategies for the problems of poor stability when equipment signal drops or all main units fail, to maintain the stable operation of the computer room equipment.

[0054] Specifically, this application provides a fault handling method that can be applied to computing devices such as servers and controllers. These computing devices can be communicatively or electrically connected to each cooling device in a data center, such as... Figure 1 As shown, the computing device can implement fault handling methods in the following ways, including:

[0055] Step S100: Obtain the status parameter information of each of the multiple refrigeration devices.

[0056] Step S110: Determine whether at least one of the multiple refrigeration devices has a main unit failure based on the status parameter information of each refrigeration device. If so, proceed to step S120.

[0057] Step S120: Determine the corresponding fault handling strategy based on the number of refrigeration units that have experienced main unit failure, and then handle the fault according to the corresponding fault handling strategy.

[0058] In step S100, "multiple cooling devices" refers to all the cooling devices in the data center. For example, if data center A has a total of n cooling devices, then "multiple cooling devices" refers to n cooling devices. Each cooling device may include a chiller unit (main unit), a cooling water pump, valves, and a cold storage tank. Under normal circumstances, the cooling water pump pumps cooling water into the chiller unit through cooling pipes, and the chiller unit works to cool the data center.

[0059] Each cooling unit has its own status parameter information, including unit number, fault status information, operating status information, local / remote status information, and operating time information, etc. The computing device designed in this solution can interact with each cooling unit. For example, the computing device can interact with each cooling unit in the data center through a data interface layer, collecting real-time operating data and issuing control commands, thereby obtaining the status parameter information uploaded by each cooling unit. This data interface layer may include data interface I / O, communication interface I / O, etc.

[0060] Based on the above, this solution can determine whether at least one of the multiple refrigeration devices has experienced a main unit failure based on the status parameter information of each refrigeration device. The main unit failure may include a failure of the main unit of the refrigeration device (i.e., the water chiller unit) or a failure of the communication main unit signal of the refrigeration device to go offline.

[0061] Specifically, this solution can determine whether at least one of the multiple refrigeration units has experienced a main unit failure by using the following method. As one possible implementation method, for example... Figure 2 As shown, this solution can determine whether the main unit of a refrigeration unit is malfunctioning at multiple refrigeration units in the following ways:

[0062] Step S200: Based on the equipment fault signal of each refrigeration unit, determine whether there is a refrigeration unit among the multiple refrigeration units whose equipment fault signal is a preset fault signal. If so, proceed to step S210.

[0063] Step S210: Determine that the main unit of the refrigeration equipment has malfunctioned, and the equipment fault signal is a preset fault signal.

[0064] In the above embodiments, the computing device can receive device fault signals uploaded by each cooling device. The device fault signal generated when the host device of a cooling device malfunctions is different from the device fault signal generated when no fault occurs. For example, the device fault signal generated when the host device of a cooling device malfunctions is 1, and the device fault signal generated when no fault occurs is 0. The computing device can determine whether a cooling device has experienced a host device fault based on whether the received device fault signal is a preset fault signal. For example, cooling devices with a received device fault signal of 1 can be identified as cooling devices with a host device fault.

[0065] As another possible implementation, the host failure described above can also include communication failure, such as host signal disconnection. Figure 3 As shown, this solution can be implemented in the following way:

[0066] Step S300: Determine whether there is a refrigeration device among the multiple refrigeration devices whose local / remote status signal is a local status signal based on the local / remote status signal of each refrigeration device. If so, proceed to step S310.

[0067] Step S310: Determine if the local / remote status signal is a communication failure of the refrigeration equipment whose host signal has dropped.

[0068] In the above embodiments, the computing device can also receive local / remote status signals uploaded by each cooling device. The local / remote status signal generated when the communication host signal of a cooling device is not lost is different from the local / remote status signal generated when the communication host signal is lost due to a communication failure. For example, when the communication host signal of a cooling device is not lost, it generates a remote status signal, which is, for example, 0; when the communication host signal is lost due to a communication failure, it generates a local status signal, which is, for example, 1. The computing device can determine whether a cooling device has experienced a communication failure due to a host signal loss based on whether the received local / remote status signal is a local status signal. For example, cooling devices with a received local / remote status signal of 1 can be identified as cooling devices experiencing a communication failure due to a host signal loss.

[0069] Furthermore, this solution can also determine non-host failures of the cooling equipment based on the signals uploaded by the cooling equipment. In this case, the solution can execute traditional fault handling methods to maintain the cooling of the data center.

[0070] In the case where at least one of the multiple refrigeration units has experienced a main unit failure, this solution determines the corresponding fault handling strategy based on the number of refrigeration units with main unit failures, and then handles the fault according to the fault handling strategy. The fault handling strategy corresponds to different numbers of refrigeration units with different main unit failures.

[0071] Specifically, as described above, the computing device can obtain the operating status information of each refrigeration unit. Therefore, this solution can determine the fault handling strategy in the following ways: Figure 4 As shown, it includes:

[0072] Step S400: Determine the operating refrigeration equipment based on the equipment status information of each device to obtain the number of operating refrigeration equipment.

[0073] Step S410: Determine the corresponding fault handling strategy based on the number of operating refrigeration units and the number of refrigeration units that have experienced main unit failure.

[0074] In the above embodiments, the device status information can indicate whether the refrigeration device is running. For example, when the device status is 1, it means that the refrigeration device is currently running, and when the device status is 0, it means that the refrigeration device is not currently working. Therefore, the computing device can determine all the refrigeration devices that are currently running based on the device status information of each refrigeration device, thereby obtaining the number of refrigeration devices that are running.

[0075] Based on the above, this solution can determine the corresponding fault handling strategy based on the number of refrigeration units that have experienced host failure and the number of refrigeration units that are currently in operation.

[0076] As one possible implementation, this solution can first determine whether the number of operating refrigeration units is the total number of multiple refrigeration units, that is, whether the number of operating refrigeration units is the aforementioned total number n. If so, it means that all refrigeration units are currently operating. Different structures of the cooling system can have different processing methods. The structure of the cooling system can include a one-to-one configuration, where each water-cooled unit of the refrigeration equipment is connected to a cooling water pump through a cooling pipe; it can also include a non-one-to-one configuration, where all cooling units and cooling water pumps of the refrigeration equipment are connected to the main cooling pipe through branches.

[0077] Assuming all refrigeration equipment is currently operating and the cooling system is structured in a one-to-one manner, this solution can control the shutdown of the main unit of the refrigeration equipment experiencing a main unit failure, and keep the corresponding chilled water pump running. For example, if refrigeration equipment A1 and A2 experience a main unit failure, then the cooling units of refrigeration equipment A1 and A2 will be shut down, and the chilled water pumps of refrigeration equipment A1 and A2 will be kept running. This allows the chilled water pumps to transfer coolant to the cooling pipes, thereby compensating for the cooling load through the cooling pipes.

[0078] With all refrigeration equipment currently in operation and the cooling system structure not in a one-to-one manner, this solution can control the shutdown of all refrigeration equipment whose main unit fails. Since there is only one main cooling pipe in this mode, this solution controls all chilled water pumps to start, thereby using the main cooling pipe for cooling capacity compensation.

[0079] As another possible implementation, this solution can also determine whether the number of operating refrigeration devices is 0. If the number of operating refrigeration devices is 0, it means that no refrigeration devices are currently running. Based on this, this solution can execute the following fault handling strategy regardless of whether the cooling system structure is one-to-one or not: determine whether the number of refrigeration devices with host failure is equal to the total number of multiple refrigeration devices, i.e., whether the number of failures is n. If the number of refrigeration devices with host failure is not equal to the total number of multiple refrigeration devices, then control the start of a preset number of fault-free refrigeration devices. If the number of refrigeration devices with host failure is equal to the total number of multiple refrigeration devices, it means that the host of all refrigeration devices is faulty. In this extreme case, this solution controls the chilled water pumps of the target number of refrigeration devices to start, thereby achieving cooling compensation by circulating coolant through the chilled water pumps to the cooling pipes. This provides a chilled water pump start-up scheme that meets the current load requirements in the case of a complete host failure, maintaining the stable operation of the computer room equipment.

[0080] As described above, in the event of a complete failure of the main unit, the chilled water pumps of the target number of refrigeration units will be turned on, such as... Figure 5 As shown, the target quantity can be determined in the following ways, including:

[0081] Step S500: Obtain the total power of all computing devices currently running in the data center.

[0082] Step S510: Determine the required number of chilled water pumps to be turned on based on the total power of all currently operating computing devices, in order to obtain the target number.

[0083] In the above implementation, the cold source for the cooling capacity provided by the data center cooling system in the event of a complete host failure is changed from the host to cooling pipes or cold storage tanks. In this scenario, the number of chilled water pumps turned on will directly affect the heat exchange rate of the cooling system. High heat load and insufficient heat exchange capacity will result in poor heat dissipation of the computing equipment in the data center. Therefore, this solution determines the target number of cooling water pumps to be turned on in the event of a complete host failure in the manner described above.

[0084] Specifically, this solution determines the required number of chilled water pumps to be activated (i.e., the target number) based on the total power of all currently operating computing devices in the data center. For example, the higher the total power of all computing devices, the more chilled water pumps need to be activated. Here, "computing devices" refers to various IT devices within the data center, such as computers and servers. The specific number of chilled water pumps to be activated, i, can be determined using the following formula:

[0085]

[0086] Where k is the total number of chilled water pumps in the system, and P0 is the total power of all heat dissipation equipment when running at full load. P represents the increase in equipment power corresponding to each additional water pump. t The total power of all servers is currently available, expressed via P. t The number of water pumps that need to be turned on can be determined by the range in which the water pump is located.

[0087] As another possible implementation method, in addition to the methods mentioned above, the fault handling strategy can also determine whether the number of refrigeration devices with host failures is less than the number of refrigeration devices in operation, since not all refrigeration devices with host failures are necessarily in operation. If the number of refrigeration devices with host failures is less than the number of refrigeration devices in operation, it means that not all refrigeration devices in operation have failed; some may have failed, or none may have failed. Therefore, this solution can continue to determine whether there are any refrigeration devices with host failures among the refrigeration devices in operation. The specific determination method can be comprehensively determined by combining the operating status of the refrigeration devices and the equipment fault signals.

[0088] If a refrigeration unit experiencing a main unit failure is found among the operating refrigeration units, the number of refrigeration units experiencing main unit failure and currently in operation is obtained; then, the same number of fault-free standby refrigeration units are started, and the main unit and corresponding chilled water pump of the refrigeration unit experiencing main unit failure and currently in operation are stopped, thereby replacing the faulty unit with the standby refrigeration unit to perform refrigeration.

[0089] As another possible implementation, if the number of refrigeration devices with host failures determined by this solution is not less than the number of running refrigeration devices, it indicates that the number of faulty refrigeration devices is excessive. It is possible that all running refrigeration devices are faulty. Therefore, this solution first determines whether the number of refrigeration devices with host failures is the total number n of multiple refrigeration devices, that is, whether it is a full host failure. If the number of refrigeration devices with host failures is not the total number n of multiple refrigeration devices, then start all standby refrigeration devices without failures; stop the hosts of the refrigeration devices with host failures and running, and stop the first number of chilled water pumps. The difference between the total number of multiple refrigeration devices and the first number is the number of running refrigeration devices. Thus, in the case where the number of refrigeration devices with host failures is large, by maintaining a certain number of chilled water pumps closed and opening another batch of chilled water pumps, the refrigeration is achieved by supplying coolant to the chilled water storage tank or cooling pipe through the chilled water pumps, thereby replacing the chiller unit.

[0090] As another possible implementation, when this solution determines that the number of refrigeration devices with host failures is the total number n of multiple refrigeration devices, that is, in the case of full host failure, this solution stops the hosts of all running refrigeration devices, and then starts the chilled water pumps of the target number of refrigeration devices. The target number is the aforementioned number of opened units i. Thus, in the extreme case of full host failure, the refrigeration is achieved by supplying coolant to the chilled water storage tank or cooling pipe through the chilled water pumps, thereby replacing the chiller unit, and further avoiding the downtime of the computing devices in the data center due to excessive temperature and improving the reliability of the data center equipment.

[0091] For the above processing strategy, this solution can be illustrated by the following specific embodiments:

[0092] First category, assume that the total number of refrigeration device hosts in the data center cooling system is n, the number of running hosts is m, and the number of chilled water pumps to be started under full host failure is i. In the case where the chiller unit (host) and the chilled water pump are in the aforementioned one-to-one mode, this solution can enumerate the following 12 failure scenarios and corresponding processing strategies:

[0093] (1) Before the failure occurs, all refrigeration devices (including chiller units and chilled water pumps) are running. If 1 ≤ the number of faulty host units < n - i, stop the faulty units, and the corresponding chilled water pumps keep running.

[0094] (2) Before the failure occurs, all refrigeration devices (including chiller units and chilled water pumps) are running. If n - i ≤ the number of faulty host units < m, stop the faulty units, and the corresponding chilled water pumps keep running.

[0095] (3) All refrigeration equipment (including cooling units and cooling water pumps) was operating before the failure. If all the main units fail, stop the failed units and keep the corresponding chilled water pumps running.

[0096] (4) i ≤ operating refrigeration equipment < n. If 1 ≤ number of failed main units < n - i, when a running device fails, start the standby device first, and then stop the failed unit and the corresponding chilled water pump.

[0097] (5) i ≤ operating refrigeration equipment < n. If n - i ≤ number of failed main units < n, start the standby device first, and then stop the failed units, keeping the total number of running water pumps at m.

[0098] (6) i ≤ operating refrigeration equipment < n. If all the main units fail, stop the failed units and start i chilled water pumps.

[0099] (7) 1 ≤ operating refrigeration equipment < i. If 1 ≤ number of failed main units < n - i, when a running device fails, start the standby device first, and then stop the failed unit and the corresponding chilled water pump.

[0100] (8) 1 ≤ operating refrigeration equipment < i. If n - i ≤ number of failed main units < n, start the standby device first, and then stop the failed units, keeping the total number of running water pumps at m.

[0101] (9) 1 ≤ operating refrigeration equipment < i. If all the main units fail, stop the failed units and start i chilled water pumps.

[0102] (10) Operating refrigeration equipment = 0. If 1 ≤ number of failed main units < n - i, start the non-failed standby device.

[0103] (11) Operating refrigeration equipment = 0. If n - i ≤ number of failed main units < n, start the non-failed standby device.

[0104] (12) Operating refrigeration equipment = 0. If all the main units fail, start i chilled water pumps.

[0105] Second category: In the case where the cooling unit (main unit) and the cooling water pump are in the aforementioned non-one-to-one mode, this solution can enumerate the following 12 failure scenarios and corresponding handling strategies:

[0106] (1) All refrigeration equipment (including cooling units and cooling water pumps) was operating (running) before the failure. If 1 ≤ number of failed main units < n - i, stop the failed units and keep all chilled water pumps running.

[0107] (2) All refrigeration equipment (including cooling units and cooling water pumps) was operating before the failure. If n - i ≤ number of failed main units < m, stop the failed units and keep all chilled water pumps running.

[0108] (3) All refrigeration equipment (including cooling units and cooling water pumps) was operating before the failure occurred. If all the main units failed, the failed units were stopped, and all chilled water pumps remained operating.

[0109] (4) i ≤ operating refrigeration equipment < n. If 1 ≤ number of failed main units < n - i, when a running device fails, standby devices are started first, and then the failed units and the corresponding chilled water pumps are stopped.

[0110] (5) i ≤ operating refrigeration equipment < n. If n - i ≤ number of failed main units < n, standby devices are started first, and then the failed units are stopped, and the total number of pumps in operation is maintained at m.

[0111] (6) i ≤ operating refrigeration equipment < n. If all the main units failed, the failed units are stopped, and i chilled water pumps are started.

[0112] (7) 1 ≤ operating refrigeration equipment < i. If 1 ≤ number of failed main units < n - i, when a running device fails, standby devices are started first, and then the failed units and the corresponding chilled water pumps are stopped.

[0113] (8) 1 ≤ operating refrigeration equipment < i. If n - i ≤ number of failed main units < n, standby devices are started first, and then the failed units are stopped, and the total number of pumps in operation is maintained at m.

[0114] (9) 1 ≤ operating refrigeration equipment < i. If all the main units failed, the failed units are stopped, and i chilled water pumps are started.

[0115] (10) Operating refrigeration equipment = 0. If 1 ≤ number of failed main units < n - i, start the standby devices without failure.

[0116] (11) Operating refrigeration equipment = 0. If n - i ≤ number of failed main units < n, start the standby devices without failure.

[0117] (12) Operating refrigeration equipment = 0. If all the main units failed, i chilled water pumps are started.

[0118] Among them, in the case of opening some of the chilled water pumps as described above, for example, when opening i or m pumps, this solution can be selectively opened according to the running time of the chilled water pumps. For example, i or m chilled water pumps with the shortest running time can be retained, and the chilled water pumps with longer running time can be shut down.

[0119] The fault handling method described above involves collecting status parameter information from each cooling device in the data center. Based on this information, it determines whether a host failure has occurred. In the event of a host failure, the corresponding fault handling strategy is determined based on the number of cooling devices experiencing the failure. This allows for different fault handling strategies to be adopted depending on the number of failures, thus achieving seamless intelligent switching between failed and functioning hosts. Furthermore, in cases where excessive failures lead to insufficient cooling capacity or cooling failure, this solution utilizes the stored cold energy in the cold storage device and the remaining cold energy in the pipes for compensation, providing cooling assurance in extreme situations. Additionally, this solution provides emergency strategies to address issues such as signal dropouts and poor stability during full host failures, maintaining the stable operation of the data center equipment.

[0120] Figure 6 A schematic structural block diagram of a fault handling device provided in this application is presented. It should be understood that this device is related to... Figure 1 and Figure 5 The method embodiments implemented in this paper correspond to the steps involved in the aforementioned method. The specific functions of this device can be found in the description above; to avoid repetition, detailed descriptions are omitted here. This device includes at least one software function module that can be stored in a memory or embedded in the device's operating system (OS) in the form of software or firmware. Specifically, the device includes: an acquisition module 600, used to acquire status parameter information of each of the multiple refrigeration devices; a determination module 610, used to determine, based on the status parameter information of each refrigeration device, whether at least one refrigeration device among the multiple refrigeration devices has experienced a host failure; and a determination module 620, used to, after determining that at least one refrigeration device among the multiple refrigeration devices has experienced a host failure, determine a corresponding fault handling strategy based on the number of refrigeration devices experiencing host failures, and perform fault handling according to the corresponding fault handling strategy.

[0121] The fault handling device designed above collects status parameter information of each cooling device in the data center, and then determines whether the cooling device has experienced a host failure based on the status parameter information. In the event of a host failure, the corresponding fault handling strategy is determined according to the number of cooling devices experiencing host failures. Thus, different fault handling strategies are adopted for different numbers of failures, thereby achieving seamless intelligent switching between failed and normal hosts. This solves the problem of excessive temperature caused by delayed equipment startup after a failure in the current cooling system, improves fault handling efficiency, maintains the cooling effect of the data center, and ensures the reliability of the data center computing equipment.

[0122] In an optional embodiment of this example, the status parameter information includes a device fault signal. The determination module 610 is specifically used to determine whether there is a refrigeration device among the multiple refrigeration devices whose device fault signal is a preset fault signal based on the device fault signal of each refrigeration device; if so, it is determined that the host device of the refrigeration device whose device fault signal is a preset fault signal has malfunctioned.

[0123] In an optional embodiment of this example, the status parameter information includes local / remote status signals. The determination module 610 is further specifically used to determine whether there is a refrigeration device among the multiple refrigeration devices whose local / remote status signal is a local status signal based on the local / remote status signal of each refrigeration device; if so, it is determined that the communication host signal of the refrigeration device whose local / remote status signal is a local status signal has failed.

[0124] In an optional embodiment of this example, the status parameter information further includes device status information. The determining module 620 is specifically used to determine the running refrigeration equipment based on the device status information of each device, so as to obtain the number of running refrigeration equipment; and to determine the corresponding fault handling strategy based on the number of running refrigeration equipment and the number of refrigeration equipment that have experienced host failure.

[0125] In an optional embodiment of this example, the determining module 620 is further specifically used to determine whether the number of operating refrigeration devices is the total number of the multiple refrigeration devices; if the number of operating refrigeration devices is the total number of the multiple refrigeration devices, then the host of the refrigeration device with host failure is shut down, and the chilled water pump corresponding to the refrigeration device with host failure is kept running, wherein one refrigeration device and one chilled water pump are connected to each other through a corresponding pipeline.

[0126] In an optional embodiment of this example, the determining module 620 is further specifically used to determine whether the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment; if the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment, the host of the refrigeration equipment with host failure is shut down, and all chilled water pumps are kept running, wherein all refrigeration equipment and all chilled water pumps are connected through a pipeline.

[0127] In an optional embodiment of this example, the determining module 620 is further specifically configured to: if the number of operating refrigeration devices is not the total number of the multiple refrigeration devices, determine whether the number of operating refrigeration devices is 0; if the number of operating refrigeration devices is 0, determine whether the number of refrigeration devices with host failure is the total number of the multiple refrigeration devices; if the number of refrigeration devices with host failure is not the total number of the multiple refrigeration devices, control the start of a preset number of fault-free refrigeration devices; if the number of refrigeration devices with host failure is the total number of the multiple refrigeration devices, control the start of the chilled water pumps of the target number of refrigeration devices.

[0128] In an optional embodiment of this example, the determining module 620 is further specifically used to determine whether the number of refrigeration devices with host failures is less than the number of refrigeration devices currently in operation; if the number of refrigeration devices with host failures is less than the number of refrigeration devices currently in operation, then determine whether there are any refrigeration devices with host failures among the currently operating refrigeration devices; if there are any refrigeration devices with host failures among the currently operating refrigeration devices, then obtain the number of refrigeration devices with host failures that are currently in operation; control the start of the same number of standby refrigeration devices without failures according to the number of refrigeration failures, and control the host and corresponding chilled water pump of the refrigeration devices with host failures that are currently in operation to stop.

[0129] In an optional embodiment of this example, the determining module 620 is further specifically configured to: determine whether the number of refrigeration devices experiencing host failures is equal to the total number of multiple refrigeration devices if the number of refrigeration devices experiencing host failures is not less than the number of refrigeration devices currently in operation; control all non-faulty standby refrigeration devices to start if the number of refrigeration devices experiencing host failures is not equal to the total number of multiple refrigeration devices; control the host of the refrigeration devices experiencing host failures and currently in operation to stop, and control a first number of chilled water pumps to stop, wherein the difference between the total number of multiple refrigeration devices and the first number is the number of refrigeration devices currently in operation.

[0130] In an optional embodiment of this example, the determining module 620 is further specifically used to control the main unit of all running refrigeration equipment to shut down and control the chilled water pumps of the target number of refrigeration equipment to start if the number of refrigeration equipment experiencing host failure is the total number of multiple refrigeration equipment.

[0131] In an optional embodiment of this example, the acquisition module 600 is further configured to acquire the total power of all computing devices currently running in the data center; the determination module 620 is further configured to determine the required number of chilled water pumps to be turned on based on the total power of all currently running computing devices, so as to obtain the target quantity.

[0132] According to some embodiments of this application, such as Figure 7As shown, this application provides an electronic device 7, including: a processor 701 and a memory 702. The processor 701 and the memory 702 are interconnected and communicate with each other through a communication bus 703 and / or other forms of connection mechanism (not shown). The memory 702 stores a computer program executable by the processor 701. When the computing device is running, the processor 701 executes the computer program to execute the method executed by the external terminal in any optional implementation, such as steps S100 to S120: obtaining the status parameter information of each of the multiple refrigeration devices; determining whether at least one of the multiple refrigeration devices has a host failure based on the status parameter information of each refrigeration device; if so, determining the corresponding fault handling strategy based on the number of refrigeration devices with host failures, and performing fault handling according to the corresponding fault handling strategy.

[0133] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the method in any of the aforementioned optional implementations.

[0134] The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0135] This application provides a computer program product that, when run on a computer, causes the computer to perform a method in any of the optional implementations.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and they should all be covered within the scope of the claims and specification of this application. In particular, as long as there is no structural conflict, the various technical features mentioned in the embodiments can be combined in any way. This application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A fault handling method, characterized in that, The method includes: Obtain the status parameter information of each of the multiple refrigeration devices; Based on the status parameter information of each refrigeration unit, determine whether at least one of the multiple refrigeration units has experienced a main unit failure; If so, the corresponding fault handling strategy is determined based on the number of refrigeration units that have experienced main unit failure, and the fault is handled according to the corresponding fault handling strategy. The status parameter information also includes device status information. The step of determining the corresponding fault handling strategy based on the number of refrigeration units experiencing main unit failure includes: The number of operating refrigeration units is determined based on the equipment status information of each unit. The corresponding fault handling strategy is determined based on the number of operating refrigeration units and the number of refrigeration units experiencing main unit failure. The step of determining the corresponding fault handling strategy based on the number of operating refrigeration units and the number of refrigeration units experiencing main unit failure includes: Determine whether the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment; if the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment, then control the main unit of the refrigeration equipment with the main unit failure to stop, and control the chilled water pump corresponding to the refrigeration equipment with the main unit failure to keep running, wherein one refrigeration equipment and one chilled water pump are connected through a corresponding pipeline. Alternatively, determine whether the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment; if the number of operating refrigeration equipment is the total number of the multiple refrigeration equipment, then control the main unit of the refrigeration equipment with the main unit failure to shut down, and control all chilled water pumps to keep running, wherein all refrigeration equipment and all chilled water pumps are connected through a pipeline. After determining whether the number of operating refrigeration devices is the total number of the multiple refrigeration devices, the method further includes: If the number of operating refrigeration equipment is not the total number of the multiple refrigeration equipment, then determine whether the number of operating refrigeration equipment is 0; If the number of operating refrigeration units is 0, then determine whether the number of refrigeration units that have experienced main unit failure is the total number of multiple refrigeration units. If the number of refrigeration units that experience main unit failure is not greater than the total number of multiple refrigeration units, then the fault-free refrigeration units of the preset control unit will be started. If the number of refrigeration units experiencing main unit failure is equal to the total number of refrigeration units, then the chilled water pumps of the target number of refrigeration units will be turned on.

2. The method according to claim 1, characterized in that, The status parameter information includes equipment fault signals. The step of determining whether at least one of the multiple refrigeration units has experienced a main unit failure based on the status parameter information of each refrigeration unit includes: Based on the equipment fault signals of each refrigeration unit, determine whether there is a refrigeration unit among multiple refrigeration units whose equipment fault signal is a preset fault signal; If so, then the main unit of the refrigeration equipment that is the preset fault signal is determined to have malfunctioned.

3. The method according to claim 1, characterized in that, The status parameter information includes local / remote status signals. Determining whether at least one of the multiple refrigeration devices has experienced a main unit failure based on the status parameter information of each refrigeration device includes: Based on the local / remote status signal of each refrigeration unit, determine whether there is a refrigeration unit among multiple refrigeration units whose local / remote status signal is a local status signal; If so, then the communication failure of the refrigeration equipment whose local status signal is the host signal has been disconnected.

4. The method according to claim 1, characterized in that, The method of determining the corresponding fault handling strategy based on the number of operating refrigeration units and the number of refrigeration units experiencing main unit failure also includes: Determine whether the number of refrigeration units experiencing main unit failures is less than the number of refrigeration units currently in operation; If the number of refrigeration units with main unit failures is less than the number of refrigeration units currently in operation, then determine whether any of the refrigeration units currently in operation have experienced main unit failures. If there is a refrigeration unit with a main unit failure among the refrigeration units that are currently in operation, then obtain the number of refrigeration units with main unit failures that are currently in operation. Based on the number of faults, control the start-up of an equal number of fault-free standby refrigeration units, and control the shutdown of the main unit and corresponding chilled water pump of the refrigeration unit that has experienced a main unit failure and is in operation.

5. The method according to claim 4, characterized in that, After determining whether the number of refrigeration units experiencing host failures is less than the number of refrigeration units currently in operation, the method further includes: If the number of refrigeration units experiencing main unit failure is not less than the number of refrigeration units currently in operation, then determine whether the number of refrigeration units experiencing main unit failure is equal to the total number of multiple refrigeration units. If the number of refrigeration units experiencing main unit failure is not equal to the total number of refrigeration units, then all non-faulty standby refrigeration units are controlled to start; the main units of the refrigeration units experiencing main unit failure and currently in operation are controlled to stop, and a first number of chilled water pumps are controlled to stop, wherein the difference between the total number of the multiple refrigeration units and the first number is the number of refrigeration units currently in operation.

6. The method according to claim 4, characterized in that, After determining whether the number of refrigeration units experiencing host failure is equal to the total number of refrigeration units, the method further includes: If the number of refrigeration units experiencing main unit failure is equal to the total number of refrigeration units, then the main units of all operating refrigeration units will be shut down, and the chilled water pumps of the target number of refrigeration units will be turned on.

7. The method according to claim 6, characterized in that, The method further includes: Obtain the total power of all computing devices currently operating in the data center; The required number of chilled water pumps to be turned on is determined based on the total power of all currently operating computing devices, in order to obtain the target number.

8. A fault handling apparatus for implementing the fault handling method according to any one of claims 1-7, characterized in that, The device includes: The acquisition module is used to acquire the status parameter information of each of the multiple refrigeration devices. The determination module is used to determine whether at least one of the multiple refrigeration devices has experienced a main unit failure based on the status parameter information of each refrigeration device. The determination module is used to determine the corresponding fault handling strategy based on the number of refrigeration devices with main unit failure after determining that at least one refrigeration device among multiple refrigeration devices has a main unit failure, so as to carry out fault handling according to the corresponding fault handling strategy. The determining module is specifically used to determine the running refrigeration equipment based on the equipment status information of each device, so as to obtain the number of running refrigeration equipment; and to determine the corresponding fault handling strategy based on the number of running refrigeration equipment and the number of refrigeration equipment with host failure. In the process of determining the corresponding fault handling strategy based on the number of operating refrigeration units and the number of refrigeration units experiencing main unit failure, the determining module is specifically used to: determine whether the number of operating refrigeration units is the total number of the multiple refrigeration units; if the number of operating refrigeration units is the total number of the multiple refrigeration units, then control the main unit of the refrigeration unit experiencing the main unit failure to shut down, and control the chilled water pump corresponding to the refrigeration unit experiencing the main unit failure to keep running, wherein one refrigeration unit and one chilled water pump are connected through a corresponding pipeline; or determine whether the number of operating refrigeration units is the total number of the multiple refrigeration units; if the number of operating refrigeration units is the total number of the multiple refrigeration units, then control the main unit of the refrigeration unit experiencing the main unit failure to shut down, and control all chilled water pumps to keep running, wherein all refrigeration units and all chilled water pumps are connected through a single pipeline.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Automatic cooling water control system and method for air conditioner in machine room

    CN105571062A

  • Control system, method and device of data center air conditioning unit and storage medium

    CN112212464A