A high-power single-phase immersion liquid-cooled data center cabinet

By integrating internal monitoring modules and adjustment control modules in data center cabinets, and combining machine learning models and multi-dimensional imbalance judgment strategies, the problems of inaccurate thermal management identification, unbalanced resource allocation, and complex fault maintenance are solved, achieving unified operation and maintenance with efficient heat dissipation and high reliability, and adapting to the heat dissipation needs of high-power servers.

CN120529567BActive Publication Date: 2025-09-26TIANJIN TIER TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511020826.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-09-26
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

Existing data center liquid cooling cabinets have inaccurate thermal management identification, unbalanced resource allocation, slow fault maintenance response, and complex operation, making them unable to achieve flexible adaptation and automatic adjustment.

Method used

A high-power single-phase immersion liquid-cooled data center cabinet is used, which integrates an internal monitoring module and an adjustment control module. The temperature monitoring unit and the power monitoring unit are used to obtain data. The trained machine learning model is combined to dynamically match the coolant flow. Through multi-dimensional imbalance judgment and adjustment strategies, including valve adjustment, task scheduling and fault emergency handling, the system redundancy and maintenance convenience are improved.

Benefits of technology

It achieves the unity of efficient heat dissipation, precise operation and maintenance, and high reliability. Through multi-dimensional imbalance judgment and targeted adjustment strategies, it can quickly locate and solve thermal management problems, improve operation and maintenance efficiency and safety redundancy, adapt to the heat dissipation needs of high-power servers, and ensure the stable operation of data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120529567B_ABST
    Figure CN120529567B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of safety equipment for working at heights, and in particular to a high-power single-phase immersion liquid-cooled data center cabinet. The coolant flow rate is determined according to the operating power of the server unit and the inlet and outlet temperature data of the coolant to determine the power of the coolant circulation pump. The presence of a server with thermal management imbalance is determined according to the temperature data of each monitoring point of the central cabinet, and real-time relevant data of the server with thermal management imbalance is obtained to determine the cause of the imbalance of the server, and an imbalance adjustment strategy is determined according to the cause of the imbalance. The present invention solves the problem of inaccurate identification of thermal management imbalance through intelligent monitoring and precise adjustment, and realizes load balancing and efficient heat dissipation. Task migration and valve adjustment improve resource utilization, and the fault emergency mechanism enhances system stability, ensures the stable operation of high-power servers, and improves operation and maintenance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of safety equipment for working at heights, and in particular to a high-power single-phase immersion liquid-cooled data center cabinet. Background Art

[0002] The liquid supply design of liquid cooling cabinets in data centers currently on the market is usually based on specific server designs and cannot be flexibly adapted. The design of liquid cooling cabinets cannot detect and display the internal conditions of the chassis in real time to form automatic adjustments. In addition, it is impossible to flexibly assemble and extract individual services in the cabinet.

[0003] Chinese patent publication number CN115135119B discloses a liquid-cooled cabinet for a data center. The liquid-cooled cabinet includes a cabinet body, an upper cover, and a support frame. The cabinet body has an open top, the upper cover is located on the open top, and the support frame is located outside the cabinet body. The cabinet body is provided with a first side and a second side on either side in the width direction. A line entry area is provided between the first side and the corresponding side of the support frame, and a line management area is provided between the second side and the support frame. A power distribution room and an equipment room are provided in sequence from the first side to the second side of the cabinet body. The power distribution room is provided with at least one power distribution rack, in which power components are installed. The equipment room has a plurality of mounting locations arranged in sequence along the length of the cabinet body, the mounting locations being used to install equipment. An electrical connection structure is provided between the power distribution rack and the mounting location, the electrical connection structure being used to electrically connect the power components to the equipment. The liquid-cooled cabinet has a low total power consumption, which can reduce energy consumption and save the power supply cost of the liquid-cooled cabinet.

[0004] It can be seen that the prior art has the following problems:

[0005] Thermal management imbalances are not accurately identified, with many misjudgments and missed judgments. Unbalanced resource allocation leads to low efficiency, slow fault maintenance response, and complex operations. Summary of the Invention

[0006] To this end, the present invention provides a high-power single-phase immersion liquid-cooled data center cabinet to overcome the problems in the prior art of inaccurate identification of thermal management imbalances, frequent misjudgments and missed judgments, low efficiency due to unbalanced resource allocation, slow fault maintenance response, and complex operations.

[0007] To achieve the above objectives, the present invention provides a high-power single-phase immersion liquid-cooled data center cabinet, comprising a liquid cooling box, a server unit, a plurality of electronically controlled valves, a coolant circulation pump, a server hoisting device, an external coolant pipeline, and a spare liquid baffle;

[0008] Wherein, the server group includes several servers;

[0009] It also includes: an internal monitoring module and a regulation control module;

[0010] The internal monitoring module includes a temperature monitoring unit, which includes a plurality of temperature detectors for obtaining temperature data of each monitoring point of the central cabinet and inlet and outlet temperature data of the coolant;

[0011] The internal monitoring module includes a power monitoring unit for obtaining the operating power of each server;

[0012] The regulating and controlling module is respectively connected to the internal monitoring module, the server unit, each of the electronically controlled valves, the coolant circulation pump, the server hoisting device and the spare liquid baffle;

[0013] The regulating control module is configured to determine the coolant flow rate according to the operating power of the server group and the inlet and outlet temperature data of the coolant to determine the power of the coolant circulation pump;

[0014] The regulation control module is further configured to determine whether there is a server with thermal management imbalance based on the temperature data of each monitoring point in the central cabinet, obtain real-time relevant data of the server with thermal management imbalance to determine the imbalance cause of the server, and determine the imbalance adjustment strategy based on the imbalance cause;

[0015] Wherein, the imbalance adjustment strategy includes adjusting the corresponding valve opening and adjusting the calculation node;

[0016] The real-time related data include operating power and node task volume.

[0017] As an optimal technical solution for high-power single-phase immersion liquid cooling data center cabinets, the internal monitoring module sets the location of the CPU and / or GPU inside each server as the monitoring point;

[0018] The number of monitoring sites is the same as the total number of CPUs and GPUs in the server group.

[0019] As an optimal technical solution for high-power single-phase immersion liquid-cooled data center cabinets, the regulation and control module includes a trained machine learning model to determine the coolant flow rate based on the operating power of the server unit and the inlet and outlet temperature data of the coolant, and to determine the power of the coolant circulation pump based on a pre-stored flow-power comparison table.

[0020] As an optimal technical solution for high-power single-phase immersion liquid cooling of data center cabinets, the machine learning model is trained based on various valid historical data sets;

[0021] The single set of valid historical data includes at least the corresponding server unit operating power, coolant inlet and outlet temperatures, and coolant flow rate when there is no thermal management disordered server.

[0022] As an optimal technical solution for high-power single-phase immersion liquid-cooled data center cabinets, the regulation and control module calculates the standard deviation and the average value based on the temperature data of each monitoring point to determine the coefficient of variation, and determines the temperature characterization state of the corresponding monitoring point based on the coefficient of variation and the temperature data of a single monitoring point to determine whether there is a thermal management imbalance in the corresponding server.

[0023] As an optimal technical solution for high-power single-phase immersion liquid cooling of data center cabinets, the regulation control module determines the temperature characterization state of the corresponding monitoring point based on the coefficient of variation and the temperature data of a single monitoring point, wherein:

[0024] According to the coefficient of variation of a single monitoring site and the determination result that the temperature data meets the abnormal condition, determining that the temperature characteristic state of the corresponding monitoring site is an abnormal characteristic state;

[0025] Determining that the temperature characteristic state of a corresponding monitoring site is a normal characteristic state based on the coefficient of variation of the single monitoring site and the determination result that the temperature data does not meet the abnormal condition;

[0026] Wherein, the abnormal condition is that the above-mentioned coefficient of variation is greater than a preset coefficient and the temperature data is greater than a preset temperature data;

[0027] The temperature characterization state includes a normal characterization state and an abnormal characterization state.

[0028] As an optimal technical solution for high-power single-phase immersion liquid-cooled data center cabinets, the regulation control module determines that a server has a thermal management imbalance based on the determination results of the monitoring points on a single server that have an abnormal characteristic state.

[0029] As a preferred technical solution for high-power single-phase immersion liquid cooling of data center cabinets, the regulation control module responds to a thermal management imbalance in any server by acquiring real-time data related to the server to determine the cause of the imbalance, including:

[0030] According to the result of determining that the node task amount of the server is greater than or equal to the preset task amount, determining that the cause of the server's imbalance is a large amount of calculation;

[0031] According to the result of determining that the operating power of the server is greater than or equal to the preset power, determining that the cause of the imbalance of the server is a heavy computing load;

[0032] Based on the determination result that the operating power of the server is less than the preset power and the node task volume is less than the preset task volume, the cause of the server imbalance is determined to be server abnormality, and the imbalance adjustment strategy is determined to be shutting down the corresponding server, controlling the server lifting device to take out the server for inspection and maintenance, and adding a spare liquid baffle at the corresponding position of the server.

[0033] As an optimal technical solution for high-power single-phase immersion liquid-cooled data center cabinets, the regulation control module responds to the determination result that the cause of the server's imbalance is a large amount of calculation, and performs trial calculations on the queued task in an idle server. Based on the trial calculation results, it is determined whether the queued task is compatible with the idle server, and based on the compatibility determination result, the queued task is automatically transferred to a compatible idle server for calculation.

[0034] As an optimal technical solution for high-power single-phase immersion liquid-cooled data center cabinets, the regulation control module determines that the imbalance adjustment strategy is to adjust the valve opening corresponding to the server so that the adjusted valve opening is greater than the valve opening before adjustment based on the judgment result that the imbalance cause is large computing load.

[0035] Compared with the existing technology, the beneficial effect of the present invention is that the high-power single-phase immersion liquid-cooled data center cabinet provided by the present invention achieves the unity of efficient heat dissipation, precise operation and maintenance, and high reliability by integrating an intelligent monitoring and dynamic adjustment system: relying on the internal monitoring module to comprehensively capture key data such as temperature and power, combined with the trained machine learning model to dynamically match the coolant flow rate, ensuring that the heat dissipation efficiency adapts to the server load; through multi-dimensional imbalance judgment and targeted adjustment strategies (such as valve adjustment, task scheduling, and fault emergency response), thermal management problems can be quickly located and resolved; at the same time, the design of spare liquid baffles and hoisting devices improves system redundancy and maintenance convenience, ultimately meeting the heat dissipation requirements of high-power servers and ensuring the stable operation of the data center;

[0036] In particular, the regulation control module achieves accurate identification and efficient response to thermal management imbalances by integrating the dual judgment logic of the coefficient of variation and temperature data. It uses the coefficient of variation to capture the relative discreteness of temperature distribution, combined with preset temperature thresholds to screen for substantial overheating risks, effectively avoiding misjudgments and missed judgments caused by a single indicator. It locates monitoring points that indicate abnormal conditions and accurately identifies servers with thermal management imbalances, providing a clear basis for subsequent targeted adjustments (such as valve adjustment and task scheduling). Furthermore, relying on multi-dimensional data linkage analysis, it ensures the stable operation of the cooling system in high-power scenarios, taking into account the accuracy and reliability of thermal management, and significantly improving the cabinet's operation and maintenance efficiency and safety redundancy.

[0037] In particular, the regulation and control module adopts precise adjustment strategies for different causes of thermal management imbalance, achieving efficient resolution of thermal management problems and optimal allocation of system resources: by distinguishing three types of imbalance causes, namely large computing volume, heavy computing load, and server abnormality, targeted measures such as task migration, valve adjustment, and equipment maintenance are matched respectively to avoid blind intervention; the task migration mechanism screens compatible idle servers through trial calculations to balance the load and alleviate the risk of overheating; the valve opening is dynamically adjusted to adapt to real-time heat dissipation requirements to ensure stable operation in high-power scenarios; fault emergency handling (such as lifting and maintenance, spare liquid baffles) improves system redundancy, and ultimately maximizes resource utilization and task processing efficiency while ensuring server safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a connection diagram of a high-power single-phase immersion liquid cooling data center cabinet according to an embodiment of the present invention;

[0039] Figure 2 A diagram illustrating the steps for determining an imbalance adjustment strategy according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0041] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0042] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0043] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0044] See also Figure 1The figure shows a connection diagram of a high-power single-phase immersion liquid-cooled data center cabinet according to an embodiment of the present invention. The present invention provides a high-power single-phase immersion liquid-cooled data center cabinet, comprising a liquid cooling cabinet, a server unit, a plurality of electrically controlled valves (equal to the number of servers in the server unit), a coolant circulation pump (typically only one coolant circulation pump per cabinet), a server lifting device (typically only one server lifting device per cabinet, each server lifting device including at least two lifting clamps), an external coolant pipeline (one), and a plurality of spare liquid baffles.

[0045] Wherein, the server group includes several servers;

[0046] It can be understood that the amount of coolant flow is regulated by the power of the coolant circulation pump;

[0047] It also includes: an internal monitoring module and a regulation control module;

[0048] The internal monitoring module includes a temperature monitoring unit, which includes several temperature detectors for obtaining temperature data of each monitoring point of the central cabinet and inlet and outlet temperature data of the coolant;

[0049] The internal monitoring module includes a power monitoring unit to obtain the operating power of each server. In practice, the server motherboard is usually integrated with a power management chip (PMIC) or power sensor to monitor the power consumption of key components within the server (such as the CPU and power module). In addition, the server operating system (such as Windows and Linux) can read the power consumption of the CPU and memory through software (based on the motherboard sensor).

[0050] The regulating and controlling module is respectively connected to the internal monitoring module, the server unit, each of the electronically controlled valves, the coolant circulation pump, the server hoisting device and the spare liquid baffle;

[0051] The regulation and control module is configured to determine the coolant flow rate based on the operating power of the server unit and the inlet and outlet temperature data of the coolant to determine the power of the coolant circulation pump; specifically, the regulation and control module includes a trained machine learning model for determining the coolant flow rate based on the operating power of the server unit and the inlet and outlet temperature data of the coolant, and determining the power of the coolant circulation pump based on a pre-stored flow-power comparison table; specifically, the machine learning model is trained based on each valid historical data group; wherein a single valid historical data group includes at least the server unit operating power, coolant inlet and outlet temperature, and coolant flow rate corresponding to the absence of a thermally imbalanced server;

[0052] It is understood that a single set of valid historical data may also include key data related to server thermal management, operating status, and environmental factors, such as: CPU / GPU utilization, memory usage, task type / priority, attributes of running tasks, number of servers / cluster size, number of servers running in the group, computer room ambient temperature and humidity, coolant medium properties, data collection timestamp, and cumulative server operating time. It should be understood that the larger the data volume of the data set, the higher the training accuracy of the machine learning model and the more accurate the prediction of coolant flow.

[0053] In the implementation, (1) all valid historical data sets are used as input features of the neural network; the coolant flow rate is used as the output of the neural network; (2) the network structure is selected, and the structures such as multi-layer perceptron (MLP), convolutional neural network (CNN) or recurrent neural network (RNN) can be selected, depending on the characteristics of the data and the complexity of the regulation task; (3) the data for training the neural network needs to be preprocessed, that is, the collected data is normalized so that the neural network can learn better; (4) a suitable loss function is selected, such as mean square error (MSE), to measure the difference between the predicted value and the actual value; (5) an optimization algorithm (such as mini-batch stochastic gradient descent, stochastic gradient descent with momentum, etc.) is used to train the neural network;

[0054] See also Figure 2 , which is a diagram showing the steps for determining an imbalance adjustment strategy according to an embodiment of the present invention. The adjustment control module is further configured to determine whether there is a server with thermal management imbalance based on the temperature data of each monitoring point in the central cabinet, obtain real-time relevant data of the server with thermal management imbalance to determine the cause of the imbalance of the server, and determine the imbalance adjustment strategy based on the imbalance cause;

[0055] Wherein, the imbalance adjustment strategy includes adjusting the corresponding valve opening and adjusting the calculation node;

[0056] The real-time related data include operating power and node task volume.

[0057] Specifically, the internal monitoring module sets the location of the CPU and / or GPU inside each server as the monitoring site;

[0058] The number of monitoring sites is the same as the total number of CPUs and GPUs in the server group.

[0059] It should be understood that the cabinet collects data such as server operating power, coolant inlet and outlet temperatures in real time through the internal monitoring module, uses the trained machine learning model to accurately calculate the required coolant flow, and achieves dynamic matching by adjusting the circulating pump power. This mechanism ensures that the heat dissipation capacity is highly adapted to the real-time power consumption of the server unit, avoiding energy waste caused by over-heating or overheating risks caused by under-heating, and is especially suitable for dense deployment scenarios of high-power servers; the internal monitoring module uses the CPU / GPU as the core monitoring site, comprehensively covering the key heat-generating components of the server, and combining temperature data with coefficient of variation analysis to accurately identify servers with thermal management imbalances. The adjustment control module determines the cause of the imbalance (such as large computing workload, server abnormality) based on operating power, task volume and other data, and executes targeted strategies (such as adjusting valve opening, diverting tasks, and activating spare liquid baffles), greatly reducing troubleshooting. time, reducing the risk of downtime due to thermal runaway; equipped with server lifting devices, spare liquid baffles and other hardware, physical replacement or isolation can be carried out quickly when the server malfunctions, avoiding the spread of single-node failures to the entire unit; the coordinated adjustment of electronically controlled valves and circulation pumps provides redundant flow control capabilities. This design takes into account the convenience of daily operation and maintenance and the efficiency of emergency processing, improves the stability and risk resistance of the cabinet in long-term operation, and adapts to the needs of continuous operation of data centers; the machine learning model continuously optimizes the flow regulation strategy based on historical data, and combines task volume prediction and node load analysis to achieve coordinated scheduling of computing resources and cooling resources (such as adjusting computing node allocation), thereby improving overall resource utilization. Compared with the traditional fixed cooling mode, the system can dynamically respond to business load fluctuations, reduce energy consumption while ensuring heat dissipation, and conform to the development trend of green data centers.

[0060] Specifically, the regulation control module calculates the standard deviation and the average value based on the temperature data of each monitoring site at the same time to determine the coefficient of variation (it can be understood that the coefficient of variation is the ratio of the standard deviation to the average value), and determines the temperature characterization state of the corresponding monitoring site based on the coefficient of variation and the temperature data of a single monitoring site to determine whether there is a thermal management disorder in the corresponding server.

[0061] Specifically, the regulation control module determines the temperature characterization state of the corresponding monitoring site according to the coefficient of variation and the temperature data of the single monitoring site, wherein:

[0062] According to the coefficient of variation of a single monitoring site and the determination result that the temperature data meets the abnormal condition, determining that the temperature characteristic state of the corresponding monitoring site is an abnormal characteristic state;

[0063] Determining that the temperature characteristic state of a corresponding monitoring site is a normal characteristic state based on the coefficient of variation of the single monitoring site and the determination result that the temperature data does not meet the abnormal condition;

[0064] Wherein, the abnormal condition is that the above-mentioned coefficient of variation is greater than a preset coefficient and the temperature data is greater than a preset temperature data;

[0065] It is understandable that the coefficient of variation is a dimensionless statistic used to measure the relative dispersion of data. The smaller the coefficient of variation value, the smaller the dispersion of the data relative to the average value, and the more stable and uniform the data. The larger the coefficient of variation value, the greater the dispersion of the data relative to the average value, and the more drastic the data fluctuations. It should be understood that a large standard deviation alone may be due to high overall temperature (for example, the temperature of all servers rises with similar amplitudes during business peaks). In this case, it is not necessarily a thermal management problem. A large coefficient of variation indicates that "the temperature of some sites deviates abnormally from the overall level", which is more likely to be a local heat dissipation failure (for example, the coolant connection of a server is damaged). A small coefficient of variation indicates that the temperature differences between monitoring locations are small and the overall distribution is uniform (i.e., the temperatures in different areas of the same server or multiple servers in the same cabinet are close to the average level), indicating stable operation of the cooling system and normal thermal management. A large coefficient of variation indicates that the temperature of some monitoring locations deviates significantly from the average (possibly due to local overheating or overcooling), and the temperature distribution is uneven (i.e., the temperature of the CPU area of ​​a server is much higher than that of other areas, or the temperature of a server in the cabinet rises sharply while the others are normal), indicating possible cooling imbalance.

[0066] In practice, determining abnormal conditions based solely on the coefficient of variation may misjudge scenarios where the overall temperature is low but the temperature at a certain location is slightly higher, causing the coefficient of variation to exceed the standard, but the temperature does not actually reach a dangerous level. Such situations do not require intervention, so further screening is required in conjunction with the absolute temperature value. The abnormal conditions provided by the present invention are only considered abnormal if they meet two conditions. This eliminates both non-risk scenarios where the temperature fluctuates locally but is safe (e.g., a large coefficient of variation but the temperature does not exceed the threshold), and normal load scenarios where the overall temperature is high but the distribution is normal (e.g., the temperature exceeds the threshold but the coefficient of variation is small). Ultimately, the focus is on the core issue of thermal management imbalance, which is local overheating and significantly different from the overall situation.

[0067] The temperature characterization state includes a normal characterization state and an abnormal characterization state.

[0068] In practice, the preset coefficient is usually less than 0.1, and a coefficient of variation of less than 10% is often considered acceptable in natural science experiments; the preset temperature data is determined based on the average value of the temperature data of each monitoring point in the valid historical data group.

[0069] Specifically, the regulation control module determines that a thermal management disorder exists in a single server based on a determination result of a monitoring location on the single server indicating an abnormal state.

[0070] It should be understood that the regulation control module adopts the dual abnormal conditions of coefficient of variation > preset coefficient and temperature data > preset temperature. It not only uses the coefficient of variation to identify the distribution abnormality of local temperature significantly deviating from the overall temperature (excluding the normal load scenario of overall synchronous temperature rise), but also uses the temperature threshold to screen the risk state of substantial overheating (excluding the non-risk scenario of local fluctuation but safe temperature). This mechanism solves the limitations of single indicator judgment (for example, using only the coefficient of variation may misjudge slight fluctuations in low temperature environment, and using only temperature may miss the hidden danger of overall low temperature but local serious imbalance), which greatly improves the accuracy of identifying thermal management imbalance; the regulation control module uses the monitoring site (CPU / GPU core) on the server as the minimum judgment unit, and can accurately locate the server with thermal management imbalance by identifying the site that represents the abnormal state. This site-server association judgment logic avoids the fuzzy judgment of the overall state of the cabinet, allowing operation and maintenance personnel to directly lock the problem server and its core heat-generating components ( This significantly reduces troubleshooting time for critical issues like clogged coolant interfaces and faulty heat dissipation areas. To address the dense deployment of high-power immersion liquid-cooled cabinets, the regulation control module ensures a sensitive response to local overheating through strict abnormal conditions. Its decision logic is both compatible with the industry standard of an acceptable coefficient of variation of less than 10% in natural science experiments and adapts to server load fluctuations through dynamic thresholds. This ensures stable operation of the cooling system while avoiding excessive interference in normal load scenarios, balancing safety and operational efficiency. The regulation control module's decision logic is based on multi-dimensional data linkage (temperature distribution characteristics, absolute temperature values, and historical baselines), embodying an intelligent risk identification approach. It focuses not only on the absolute value of a single indicator, but also on the relative relationship between data (such as the difference between local and overall temperatures). This design enables the cabinet to proactively avoid the chain reaction risk of local overheating during high-power operation, providing continuous and reliable thermal management for server units, extending hardware life, and reducing the probability of downtime.

[0071] Specifically, in response to a thermal management imbalance in any server, the regulation control module obtains real-time relevant data of the server to determine the cause of the imbalance in the server, including:

[0072] Based on the determination that the server's node task load is greater than or equal to the preset task load, the cause of the server's imbalance is determined to be a large computational load. It should be understood that the node task load reflects the number of tasks currently being processed by the server (such as data calculations, request responses, etc.) and is a direct factor in determining its workload. When the node task load is ≥ the preset task load, the server is in a high-task load state. Excessive task loads can cause hardware such as the CPU and memory to run at full capacity, increasing the heat generated and causing thermal management imbalance (e.g., the heat dissipation rate cannot keep up with the heat generation rate). Therefore, this is attributed to the large computational load.

[0073] Based on the determination that the server's operating power is greater than or equal to the preset power, the cause of the server's imbalance is determined to be heavy computational load. It should be understood that operating power reflects the actual energy consumption of the server hardware (such as the CPU, graphics card, and fan), and is directly related to heat generation (the higher the power, the more heat generated per unit time). When the operating power is ≥ the preset power, the server's energy consumption exceeds the design threshold and generates excessive heat. Even if the workload does not reach the upper limit, a sudden increase in power due to hardware failure (such as abnormal energy consumption of a component) or high load (such as the workload does not exceed the limit but the computational complexity of a single task is extremely high) can also cause thermal management imbalance. Therefore, this is attributed to heavy computational load.

[0074] According to the determination result that the operating power of the server is less than the preset power and the node task volume is less than the preset task volume, it is determined that the cause of the server's imbalance is server abnormality, and the imbalance adjustment strategy is determined to be shutting down the corresponding server, controlling the server lifting device to take out the server for maintenance, and adding a spare liquid baffle at the corresponding position of the server; it should be understood that the operating power is less than the preset power and the node task volume is less than the preset task volume, which means that the server has neither high task requirements nor high energy consumption loads. In theory, thermal management imbalance should not occur. The imbalance at this time can only be caused by the server's own failure (such as damage to the cooling system, abnormal heat generation due to hardware short circuit, etc.). Therefore, it is attributed to the server abnormality and needs to be repaired, that is, shutting down the server, controlling the server lifting device to take out and repair it, and adding a spare liquid baffle at the corresponding position of the server.

[0075] Specifically, in response to the determination result that the cause of the server's imbalance is a large amount of calculation, the adjustment control module performs trial calculations on the queued task in an idle server, and determines whether the queued task is compatible with the idle server based on the trial calculation results. Based on the compatibility determination result, the queued task is automatically transferred to a compatible idle server for calculation.

[0076] In practice, idle servers are servers with no thermal management imbalance and spare computing nodes (usually each server has multiple computing nodes, and administrators or ordinary users usually choose frequently used servers and frequently used computing nodes when submitting computing tasks). This can cause problems such as excessive server workload, a large number of queued tasks, and imbalanced server thermal management.

[0077] In practice, the trial calculation process includes: determining the task type of the queued task (i.e., determining the job command / task submission command when the user submits the computing task); identifying several idle servers that have previously calculated the same task type as standby idle servers; predicting the calculation duration of each queued task based on the task type (methods for predicting task calculation duration are existing technologies, including statistical analysis methods based on historical data, feature-based machine learning prediction methods, and analytical prediction methods based on task structure, and will not be described in detail here); and selecting the queued task with the shortest predicted calculation duration for the same task type to perform trial calculations on the standby idle servers (starting with the idle server with the most available computing nodes and trying them in sequence). It is understandable that when submitting a task, the number of CPU cores used by the task is usually set, and this number of cores affects compatibility with the server's nodes. In addition, different computing tasks may be incompatible with different servers (i.e., the server cannot calculate this type of computing task), so trial calculations are required to determine how to adjust the queued tasks.

[0078] During implementation, if the trial calculation time is less than or equal to the corresponding preset time (usually 1.1 times the predicted time), the corresponding spare idle server is determined to be compatible with this type of queued task, and the queued task corresponding to this task type is transferred to this spare idle server for calculation; it should be understood that different server nodes have different calculation times for different tasks, and new server nodes with more cores may take longer to calculate a certain type of task than old server nodes with fewer cores. Therefore, it is necessary to use the computing market to determine whether the computing task is compatible with the server.

[0079] Specifically, the adjustment control module determines, based on the result that the imbalance cause is a large computational load, that the imbalance adjustment strategy is to adjust the valve opening corresponding to the server so that the adjusted valve opening is greater than the valve opening before adjustment.

[0080] It is understandable that the valve opening of each server should be calculated and determined in the early stage of construction, and then adjusted according to the actual situation. Each adjustment should increase the fixed opening (5% to 10% of the maximum opening, preferably 5% of the maximum opening), and restore to the original opening after there is no thermal management imbalance in the server.

[0081] It should be understood that the regulation and control module subdivides the causes of imbalances according to data such as operating power and task volume, starts task migration for large computing loads, adjusts valve openings for large computing loads, and triggers maintenance processes for server anomalies, thereby achieving accurate correspondence between causes and measures. This differentiated strategy avoids one-size-fits-all treatment (such as shutting down for maintenance for all imbalances), while solving thermal management problems and minimizing interference with normal business (such as migrating tasks instead of shutting down when the computing load is large to ensure task continuity); for imbalances with large computing loads, the module migrates queued tasks to compatible idle servers through a trial calculation mechanism, matches historical processing records based on task types, and selects the optimal target server based on computing time predictions to ensure efficient task transfer. This process balances task allocation between servers, alleviates the problem of overloaded and partially idle resources on some servers, and reduces thermal management pressure caused by load concentration, thereby improving the overall computing power utilization efficiency of the cabinet; for imbalances with large computing loads, the module increases the valve opening in a step-by-step manner and dynamically increases the coolant flow rate to accurately match The server's real-time cooling needs are met to avoid performance throttling or hardware damage caused by insufficient cooling. The design restores the original opening after adjustment, balancing cooling efficiency and energy conservation, preventing resource loss caused by excessive cooling. It is particularly suitable for dynamic load scenarios of high-power servers. In response to server anomalies, the adjustment control module initiates an emergency processing process, shuts down the faulty server, quickly removes it for maintenance via a lifting device, and adds a spare liquid baffle to maintain the sealing of the cabinet's liquid cooling environment. This mechanism significantly shortens the fault handling cycle and reduces the impact of a single server anomaly on the entire cabinet's cooling system. At the same time, through physical isolation and redundant design, the data center's risk resistance under high-power operation is improved. The trial calculation process ensures compatibility between the task and the target server through task type matching and calculation time verification (trial calculation time ≤ preset threshold), avoiding migration failures due to server non-support of the task type or low calculation efficiency. This rigorous matching logic not only ensures the success rate of task migration, but also reserves resource buffer space for subsequent task scheduling by prioritizing servers with more available nodes.

[0082] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

[0083] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A high-power single-phase immersion liquid-cooled data center cabinet, comprising a liquid cooling box, a server unit, several electronically controlled valves, a coolant circulation pump, a server hoisting device, an external coolant pipeline, and a spare liquid baffle; in, The server group includes several servers; It is characterized by further comprising: an internal monitoring module and an adjustment control module; The internal monitoring module includes a temperature monitoring unit, which includes a plurality of temperature detectors for obtaining temperature data of each monitoring point of the central cabinet and inlet and outlet temperature data of the coolant; The internal monitoring module includes a power monitoring unit for obtaining the operating power of each server; The regulating and controlling module is respectively connected to the internal monitoring module, the server unit, each of the electronically controlled valves, the coolant circulation pump, the server hoisting device and the spare liquid baffle; The regulating control module is configured to determine the coolant flow rate according to the operating power of the server group and the inlet and outlet temperature data of the coolant to determine the power of the coolant circulation pump; The regulation control module is further configured to determine whether there is a server with thermal management imbalance based on the temperature data of each monitoring point of the central cabinet, and in response to the presence of thermal management imbalance in any server, obtain real-time relevant data of the server with thermal management imbalance to determine the imbalance cause of the server, and determine the imbalance adjustment strategy based on the imbalance cause, including: According to a result of determining that the node task amount of the server is greater than or equal to a preset task amount, determining that the cause of the server's imbalance is a large amount of calculation, and determining an imbalance adjustment strategy as performing trial calculations on the queued task in an idle server, and determining whether the queued task is compatible with the idle server based on the trial calculation result, and automatically transferring the queued task to the compatible idle server for calculation based on the compatibility determination result; According to the result of determining that the operating power of the server is greater than or equal to the preset power, determining that the cause of the imbalance of the server is a large computing load, and determining that the imbalance adjustment strategy is to adjust the valve opening corresponding to the server so that the valve opening after adjustment is greater than the valve opening before adjustment; Based on the determination result that the operating power of the server is less than the preset power and the node task amount is less than the preset task amount, it is determined that the cause of the server imbalance is a server abnormality, and the imbalance adjustment strategy is determined to be shutting down the corresponding server, controlling the server lifting device to remove the server for maintenance, and adding a spare liquid baffle at the corresponding position of the server; Wherein, the imbalance adjustment strategy includes adjusting the corresponding valve opening and adjusting the calculation node; The real-time related data include operating power and node task volume.

2. A high-power single-phase immersion liquid-cooled data center cabinet according to claim 1, characterized in that: The internal monitoring module sets the location of the CPU and / or GPU inside each server as a monitoring site; The number of monitoring sites is the same as the total number of CPUs and GPUs in the server group.

3. The high-power single-phase immersion liquid cooling data center cabinet according to claim 1, characterized in that: The regulation and control module includes a trained machine learning model, which is used to determine the coolant flow rate based on the operating power of the server unit and the inlet and outlet temperature data of the coolant, and to determine the power of the coolant circulation pump based on a pre-stored flow-power comparison table.

4. The high-power single-phase immersion liquid cooling data center cabinet according to claim 3, characterized in that: The machine learning model is trained based on each valid historical data set; The single set of valid historical data includes at least the corresponding server unit operating power, coolant inlet and outlet temperatures, and coolant flow rate when there is no thermal management disordered server.

5. The high-power single-phase immersion liquid cooling data center cabinet according to claim 1, characterized in that: The regulation control module calculates the standard deviation and the average value based on the temperature data of each monitoring site to determine the coefficient of variation, and determines the temperature characterization state of the corresponding monitoring site based on the coefficient of variation and the temperature data of a single monitoring site to determine whether there is a thermal management disorder in the corresponding server.

6. The high-power single-phase immersion liquid cooling data center cabinet according to claim 5, characterized in that: The regulation control module determines the temperature characterization state of the corresponding monitoring site based on the coefficient of variation and the temperature data of the single monitoring site, wherein: According to the coefficient of variation of a single monitoring site and the determination result that the temperature data meets the abnormal condition, determining that the temperature characteristic state of the corresponding monitoring site is an abnormal characteristic state; Determining that the temperature characteristic state of a corresponding monitoring site is a normal characteristic state based on the coefficient of variation of the single monitoring site and the determination result that the temperature data does not meet the abnormal condition; Wherein, the abnormal condition is that the above-mentioned coefficient of variation is greater than a preset coefficient and the temperature data is greater than a preset temperature data; The temperature characterization state includes a normal characterization state and an abnormal characterization state.

7. The high-power single-phase immersion liquid cooling data center cabinet according to claim 6, characterized in that: The regulation control module determines that a thermal management disorder exists in a single server based on a determination result of a monitoring location on the single server indicating an abnormal state.

Citation Information

Patent Citations

  • Liquid Cooling Cabinets for Data Centers

    CN115135119B

  • Control method and system of data center liquid cooling heat dissipation system

    CN118338630A

  • Immersive liquid-cooled power battery thermal management method and system based on big data

    CN119253148A

  • A data center AI energy consumption analysis and optimization method and system

    CN119739538A

  • Evaporation and condensation type single-phase immersed liquid cooling system

    CN215269314U