Machine room equipment monitoring system and method based on Internet of Things
By collecting and analyzing the thermophysical parameters of the micro-module in real time through the Internet of Things monitoring system, identifying changes in the cold source operating conditions, and calculating the cold source capacity compensation amount in advance, the problem of temperature overshoot caused by the lag in the cooling regulation of the micro-module is solved, and faster and more stable cooling control is achieved.
Patent Information
- Application Number
- CN202511854895.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-20
AI Technical Summary
In existing technologies, the cooling regulation capability inside the micro-module cannot respond to changes in the upstream cold source system in a timely manner, leading to the risk of temperature overshoot. Especially when the cold source system is under load adjustment or running in energy-saving mode, the adjustment behavior of the precision air conditioner is lagging behind and it is difficult to effectively suppress the temperature rise overshoot caused by thermal inertia.
By using an IoT-based monitoring system, the thermophysical parameters inside the micro-module are collected synchronously, sensible heat and cooling capacity are derived, a time evolution trajectory is constructed, the changing operating conditions of the cold source are identified, the compensation amount of the cold source capacity is calculated in advance, a refrigeration control strategy is generated, and the temperature control equipment is adjusted to prevent temperature overshoot.
It enables advance prediction and compensation for changes in cold source capacity, avoids the risk of temperature overshoot, improves the response speed and stability of the refrigeration system, and reduces the occurrence of temperature overshoot.
Smart Images

Figure CN121711952A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data center management technology, and more specifically, relates to an Internet of Things-based monitoring system and method for computer room equipment. Background Technology
[0002] Data center equipment specifically refers to the micro-modular precision infrastructure units (hereinafter referred to as "micro-modules") of a data center. As a complete data center infrastructure integrator capable of independent deployment and rapid delivery, a micro-module typically includes: rack system modules, cold aisle containment and cooling modules, UPS modules, power distribution modules, and environmental monitoring modules. Rack system modules provide standard rack mounting holes for all servers, switches, routers, etc. Key design considerations typically include: air intake at the front of the rack; hot air exhaust at the rear; mesh structures for both front and rear doors to ensure airflow; and internal airflow guidance to reduce turbulence. Cold aisle containment and cooling modules improve cooling efficiency by trapping cold air within the cold aisle, stabilizing the temperature of the airflow in front of the rack and preventing the mixing of hot and cold air—essential for temperature control. UPS modules are the core power backup system for micro-modules, ensuring servers do not crash during power outages and stabilizing and filtering the mains power. Power distribution modules distribute UPS and mains power to various racks, air conditioners, and other equipment, safely distributing power to each device and monitoring the current, voltage, and power of each circuit. The environmental monitoring module is responsible for unified monitoring, including temperature and humidity, UPS status, air conditioning operation status, cold aisle doors, cabinet door status, energy consumption (PUE, etc.), etc., to monitor the environment, power and equipment health within the micro-module in real time.
[0003] Currently, in data center environments employing a combined cooling architecture of "upstream centralized chilled water system + micro-module internal precision air conditioning units," there is an inherent difference in dynamic response between the cooling regulation capabilities of the micro-modules and the upstream chilled water system. When the upstream chilled water system experiences changes in supply and return water temperature, flow rate, or differential pressure due to load adjustments, nighttime energy-saving operation, or centralized chilled water scheduling strategies, its heat exchange capacity shifts relatively smoothly over time. However, the precision air conditioning units within the micro-modules lack the ability to anticipate changes in chilled water conditions and can only perform passive feedback control based on rack inlet air temperature, cold aisle temperature and humidity, and the air conditioning's own heat exchange effect. Due to the significant thermal inertia of the IT equipment within the micro-modules (large servers and storage arrays have continuous heat dissipation capabilities and limited air duct volume), when the upstream chilled water heat exchange capacity gradually decreases over tens of minutes, the air conditioning outlet temperature within the micro-modules will maintain a "seemingly normal" level for a short period. However, as the temperature difference at the heat exchange end continues to shrink, even if the air conditioning airflow remains constant, a significant decrease in cooling capacity and a rapid rise in rack front-end temperature will eventually occur. At this point, the temperature fluctuations detected by the micro-module monitoring system are lagging behind the changes in the upstream cold source's operating conditions, causing a mismatch between the local adjustment actions and the actual capacity of the cold source, which can easily trigger the risk of temperature overshoot.
[0004] For example, in a large financial data center, micro-modules use row-level precision air conditioners as terminal cooling equipment, while the entire park is supplied with 7℃ / 12℃ supply and return water from a chiller station. To achieve an energy-saving strategy, the chiller station enters nighttime optimization mode at 22:00 every day, slowly increasing the supply water temperature from 7℃ to 10℃ over approximately 45 minutes, while simultaneously reducing the pressure difference of the circulating pump to decrease energy consumption. During this process, the precision air conditioning units use electronic expansion valves for regulation, but due to the small initial temperature difference change, the outlet air temperature can still maintain the set value, and the monitoring system shows only a slight increase in the cold aisle temperature. After about 30 minutes, due to the continuous deviation in supply water temperature, the upstream heat exchange capacity is no longer sufficient to support the current heat load, the evaporator heat exchange temperature difference of the precision air conditioner drops sharply, the outlet air temperature rises by 3-5℃ within 10 minutes, and the cold aisle temperature experiences a rapid jump. Because the micro-module monitoring only detects a significant temperature shift after the heat exchange capacity has decreased significantly, its adjustment behavior (such as increasing the air conditioning volume or changing the air supply mode) is significantly delayed, making it difficult to offset the temperature rise overshoot caused by thermal inertia, resulting in a brief overheating of the nodes at the top of the cabinet.
[0005] In another government data center project, the cooling system employs a dual-loop chilled water system with a load priority scheduling mechanism. When other buildings in the park experience a sudden surge in load (such as a large conference center turning on its air conditioning), the main control system prioritizes cooling capacity to that area and temporarily reduces the flow rate of the server room branch. A branch flow rate decreased by approximately 20% within 5 minutes, effectively weakening the heat exchange area of the precision air conditioner evaporator. The cabinets within the micro-modules were under medium to high load (45%–60%), exhibiting high thermal inertia and continuous heat dissipation, easily causing the precision air conditioner outlet temperature to rise from 19°C to 25°C; the cold aisle temperature rose rapidly within 8 minutes; and noticeable localized hotspots appeared on the IT equipment at the top of the cabinets, with the monitoring system only triggering an alarm at 26°C. The micro-module's adjustment action (increasing airflow) was insufficient to suppress the temperature rise due to inadequate heat exchange capacity. Furthermore, because cooling source scheduling is a park-level process, it is difficult for the micro-modules to predict or intervene in advance through local monitoring, resulting in their adjustments consistently lagging behind changes in cooling source capacity and making localized hotspots unavoidable. Summary of the Invention
[0006] To address the shortcomings of existing technologies, the present invention aims to resolve the aforementioned deficiencies and propose an Internet of Things-based data center equipment monitoring system and method.
[0007] The present invention adopts the following technical solution.
[0008] The first aspect of this invention discloses a method for monitoring data center equipment based on the Internet of Things (IoT), the method comprising:
[0009] Synchronously collect the thermophysical parameters inside the micromodule;
[0010] The sensible heat carried away from the air side and the cooling capacity provided by the cold source side are derived based on the aforementioned thermophysical parameters, and the heat exchange capacity evaluation result is output based on the aforementioned sensible heat and cooling capacity.
[0011] Based on the heat exchange capacity assessment results within a set time period, a time evolution trajectory is constructed, and combined with a pre-constructed reference model, the cold source change conditions inside the micro-module are identified.
[0012] When there are signs of impending attenuation in the cold source capacity corresponding to the cold source change condition, calculate the cold source capacity compensation amount and generate a refrigeration control strategy based on the cold source capacity compensation amount.
[0013] The cooling control strategy is executed to adjust the corresponding temperature control equipment according to the compensation amount of the cold source capacity, and the adjusted thermophysical parameters are monitored in real time.
[0014] Furthermore, the synchronous acquisition of thermophysical parameters inside the micromodule includes:
[0015] Multiple temperature sensors are installed on the air intake side of each rack and at different heights in the cold aisle to collect temperature data inside the micro-modules, and a temperature distribution matrix is constructed based on the temperature data.
[0016] The air volume, air temperature and return air temperature provided by the air conditioner fan are obtained, and the instantaneous heat exchange of the air conditioner is approximately calculated based on the air volume, air temperature and return air temperature.
[0017] Based on the current and voltage corresponding to the power consumption of the server, the power consumption of the server is approximately converted into the heat generated by the server, and the heat exchange capacity index of the cold source side is estimated based on the temperature difference between the inlet and outlet of the evaporator.
[0018] The thermophysical parameters are obtained by integrating the temperature distribution matrix, the instantaneous heat exchange of the air conditioner, the heat output of the server, and the heat exchange capacity index of the cold source side after time alignment and correction.
[0019] Furthermore, the step of deriving the sensible heat carried away from the air side and the cooling capacity provided by the cold source side based on the thermophysical parameters, and outputting the heat exchange capacity evaluation result based on the sensible heat and cooling capacity, includes:
[0020] Based on the aforementioned thermophysical parameters, a weighted average of different temperature data from the air intake surface and cold aisle within the same time slice is calculated to obtain the average intake temperature. This average temperature is then combined with the air supply temperature, return air temperature, and air volume of the air conditioner to construct a key air-side quantity group.
[0021] The evaporator inlet and outlet temperatures and chilled water flow rate are determined based on the thermophysical parameters, and a set of key quantities on the cold source side is constructed based on the evaporator inlet and outlet temperatures and chilled water flow rate.
[0022] Furthermore, the step of deriving the sensible heat carried away from the air side and the cooling capacity provided by the cold source side based on the thermophysical parameters, and outputting the heat exchange capacity evaluation result based on the sensible heat and cooling capacity, also includes:
[0023] The sensible heat carried away by the air side is calculated based on the key air-side quantity set, and the air-side thermal balance deviation is defined as the difference between the server's heat output and the sensible heat.
[0024] The cold energy on the cold source side is calculated based on the key quantity set on the cold source side, and the cold-heat imbalance deviation is defined. The cold-heat imbalance deviation is equal to the difference between the cold energy on the cold source side and the sensible heat.
[0025] Define a capacity margin ratio, which is equal to the ratio of the thermal imbalance deviation as a molecule to the sensible heat, and divide the heat transfer capacity region based on the comparison result of the capacity margin ratio and a set threshold range.
[0026] The heat exchange capacity assessment result consists of the capacity margin ratio and the corresponding heat exchange capacity region.
[0027] Furthermore, based on the heat exchange capacity assessment results within a set time period, a time evolution trajectory is constructed, and combined with a pre-built reference model, the changing operating conditions of the cold source inside the micro-module are identified, including:
[0028] The set time period is discretized according to a fixed sampling period to obtain discrete time. At each discrete time, the capacity margin ratio at the corresponding time is extracted from the heat exchange capacity assessment result to construct a discrete sequence of capacity margin.
[0029] The discrete sequence of the capability margin is smoothed over time using a first-order recursive smoothing algorithm to obtain a smoothed capability margin curve, and the rate of change of capability margin corresponding to each sampling point on the capability margin curve is calculated.
[0030] The capacity margin change rate is normalized by introducing a time constant to obtain a capacity change rate index. The capacity change rate index is then compared with a set threshold parameter within a selected time window to identify the cold source change conditions.
[0031] Furthermore, when there are signs of impending attenuation in the cold source capacity corresponding to the changing cold source operating conditions, calculating a cold source capacity compensation amount and generating a refrigeration control strategy based on the cold source capacity compensation amount includes:
[0032] The cold source change conditions are converted into a trend measurement, and when there are signs of cold source capacity decay in the trend measurement, the temperature difference range between the upper limit of the allowable temperature and the target temperature of the cold aisle is calculated.
[0033] Calculate the cooling capacity gap between the sensible heat carried away from the air side and the cooling capacity provided by the cold source side, and determine the required cooling capacity compensation amount according to the changing operating conditions of the cold source, so that when the cooling capacity gap exceeds the corresponding set threshold, the air supply volume of the air conditioner and the number of operating units are adjusted according to the cooling capacity compensation amount.
[0034] Based on the adjusted air conditioning air volume and the number of operating units, combined with the target fan speed and the operating status of each unit, the cooling control strategy is generated.
[0035] Furthermore, the execution of the refrigeration control strategy, which adjusts the corresponding temperature control equipment according to the compensation amount of the cold source capacity, and monitors the adjusted thermophysical parameters in real time, includes:
[0036] The corresponding air conditioning unit is controlled to operate according to the operating status of each unit and the target speed of the fan, and the operating thermophysical parameters within the set operating cycle are collected synchronously.
[0037] Calculate the actual physical deviation between the operating thermophysical parameters and the set target parameters, and determine whether there is temperature overshoot in the current operation based on the actual physical deviation;
[0038] When temperature overshoot occurs during current operation, the overshoot ratio is determined, and the actual physical deviation and overshoot ratio are compared with the set physical deviation threshold and overshoot ratio threshold, respectively, so as to correct the parameters based on the comparison results.
[0039] A second aspect of this invention discloses an Internet of Things (IoT)-based data center equipment monitoring system for implementing the IoT-based data center equipment monitoring method described in any one of the first aspects, the system comprising:
[0040] The Internet of Things (IoT) acquisition module is used to synchronously acquire the thermophysical parameters inside the micro-module;
[0041] The heat exchange capacity assessment module is used to deduce the sensible heat carried away from the air side and the cooling capacity provided by the cold source side based on the thermophysical parameters, and output the heat exchange capacity assessment result based on the sensible heat and cooling capacity.
[0042] The operating condition identification module is used to construct a time evolution trajectory based on the heat exchange capacity evaluation results within a set time period, and identify the cold source change operating conditions inside the micro-module by combining it with a pre-built reference model.
[0043] The control strategy generation module is used to calculate the cold source capacity compensation amount when there are signs of decline in the cold source capacity corresponding to the cold source change condition, and generate a refrigeration control strategy based on the cold source capacity compensation amount.
[0044] The strategy execution module is used to execute the refrigeration control strategy to adjust the corresponding temperature control equipment according to the compensation amount of the cold source capacity, and to monitor the adjusted thermophysical parameters in real time.
[0045] A third aspect of the present invention discloses a terminal, including a processor and a storage medium;
[0046] The storage medium is used to store instructions;
[0047] The processor is configured to operate according to the instructions to perform the steps of the method described in the first aspect.
[0048] A fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.
[0049] The beneficial effects of the present invention are as follows: Compared with the prior art, the present invention has the following advantages:
[0050] (1) This invention characterizes the overall thermal balance of a micromodule within a given time slice by synchronously collecting multiple physical quantities directly related to the thermal process within the micromodule. Then, by comparing the matching degree between the change in heat on the air side and the cooling capacity available from the cold source side, a qualitative assessment of the terminal heat exchange capacity is obtained, clarifying whether the current precision air conditioner is in a state of "sufficient margin," "near critical," or "significantly insufficient." Subsequently, during the assessment process, instead of focusing on a single assessment result, a time evolution trajectory is constructed based on continuous assessment results over a period of time. The first-order inertial response characteristics of the chilled water system and the thermal inertia model within the micromodule are referenced to identify whether there is a continuous, slow downward trend. Therefore, the temperature fluctuations detected by this invention lag behind changes in the upstream cold source's operating conditions, thereby avoiding misalignment between local adjustment actions and the actual capacity of the cold source, and preventing the risk of temperature overshoot.
[0051] (2) When the determination result shows that there are signs of continuous decline in the cold source capacity, this invention no longer uses the existing temperature feedback logic. Instead, based on the thermal inertia model and thermal balance model of the internal equipment of the micro-module, it calculates in advance the compensation measures required to maintain the stability of the cold aisle temperature under the condition of decreased cold source capacity. For example, it increases the air supply volume, optimizes the air conditioning operation mode, puts the backup air conditioning unit into operation in advance, and adjusts the air supply path. At the same time, during the execution of the operation strategy after compensation, thermophysical data such as cabinet inlet air temperature, cold aisle temperature, air conditioning supply air temperature, air supply volume, and equipment power consumption are collected again and compared with the expected temperature change envelope to determine whether the advance adjustment has effectively suppressed the risk of temperature overshoot. Therefore, this invention not only predicts and compensates for relevant operating parameters in advance, but also monitors the operation process after compensation in real time, further avoiding the occurrence of temperature overshoot risk. Attached Figure Description
[0052] Figure 1 This is a flowchart illustrating the IoT-based data center equipment monitoring method provided by the present invention.
[0053] Figure 2 This is a schematic diagram of the structure of the Internet of Things-based data center equipment monitoring system provided by the present invention. Detailed Implementation
[0054] The present application will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and should not be construed as limiting the scope of protection of the present application.
[0055] like Figure 1 As shown, in one embodiment, an IoT-based data center equipment monitoring method, implemented through the aforementioned IoT-based data center equipment monitoring system, includes the following steps:
[0056] Step S110: Synchronously collect the thermophysical parameters inside the micromodule.
[0057] In some embodiments, the IoT-based data center equipment monitoring method provided by the present invention includes the following steps in step S110:
[0058] Step S111: Multiple temperature sensors are installed on the air intake surface of each cabinet and at different heights in the cold aisle to collect temperature data inside the micro-module and construct a temperature distribution matrix based on the temperature data.
[0059] Step S112: Obtain the air volume, air temperature and return air temperature provided by the air conditioner fan, and approximately calculate the instantaneous heat exchange of the air conditioner based on the air volume, air temperature and return air temperature.
[0060] Understandably, the instantaneous heat exchange of an air conditioner is usually calculated by multiplying the temperature difference between the return air and the supply air by the supply air volume, and then combining this with the air density and specific heat.
[0061] Step S113: Based on the current and voltage corresponding to the power consumption of the server, the power consumption of the server is approximately converted into the heat generated by the server, and the heat exchange capacity index of the cold source side is estimated based on the temperature difference between the inlet and outlet of the evaporator.
[0062] Among them, the thermophysical parameters are obtained by integrating the temperature distribution matrix, the instantaneous heat exchange of the air conditioner, the heat output of the server, and the heat exchange capacity index of the cold source side after time alignment and correction (i.e., normalization).
[0063] In a specific embodiment, the IoT-based data center equipment monitoring method provided by the present invention includes steps 1 to 5:
[0064] Step 1: Collect and construct basic thermal state data inside the micro-module.
[0065] Through the IoT acquisition layer, multiple physical quantities directly related to the thermal process within the micro-module are synchronously collected, including at least: multiple point temperatures at the rack air inlet, temperatures at different heights within the cold aisle, temperatures of the air conditioning return and supply air, actual airflow of the air conditioning fan, typical power consumption or current changes of the main servers within the rack, and temperature difference indicators on the refrigerant or chilled water side of the evaporator inlet and outlet. The acquisition system integrates these measurements into a complete set of basic thermal state data at fixed time intervals to characterize the overall thermal balance of the micro-module within a specific time slice. This includes the following sub-steps:
[0066] Sub-step 1.1: Collect the spatial temperature distribution of the server rack and cold aisle.
[0067] Specifically, at least three temperature probes are placed on the air intake side of each rack to form a vertical temperature gradient; monitoring points are arranged at three heights (bottom, middle, and top) in the cold aisle to form a two-dimensional temperature distribution matrix. The temperature value of each cell in this matrix is expressed as:
[0068]
[0069] In the formula, Indicates the first line, number Corrected temperature for column position; The raw temperature value collected by the sensor; The sensor calibration compensation amount, typically within the range of -0.3°C to +0.3°C, is obtained through periodic calibration. These are the rack serial number index and the height index (bottom, middle, top), respectively.
[0070] Sub-step 1.2: Collect the heat exchange parameters of the air supply and return air of the air conditioner.
[0071] Specifically, the instantaneous heat exchange of the air conditioner is approximately calculated by using the real-time air volume (unit: m³ / s) provided by the air conditioner fan, combined with the supply air temperature and return air temperature, through sensible heat transfer. The expression is:
[0072]
[0073] In the formula, For air density, it is taken as 1.2 kg / m³ in engineering. The specific heat capacity of air is taken as 1.005 kJ / (kg·°C) in engineering. The current air volume delivered by the fan can be calculated from the fan characteristic curve; This refers to the return air temperature of the air conditioner. This refers to the air conditioning supply temperature.
[0074] Sub-step 1.3: Collect the approximate equivalent power consumption of the server's heat generation.
[0075] Specifically, almost all of the electrical energy consumed by the server is converted into heat. The instantaneous power can be calculated using current and voltage, and then summed to form the approximate total heat generation of the server. The expression is as follows:
[0076]
[0077] In the formula, The instantaneous equivalent heat output of the server (in W); The operating voltage for the server (usually 220V or 12V bus, depending on the data acquisition point); This refers to the real-time operating current of the server. If there are multiple servers in the rack, then... This is the sum of the instantaneous heat generation of all servers.
[0078] Sub-step 1.4: Collect the temperature difference indication on the refrigerant or chilled water side of the evaporator.
[0079] Specifically, the heat absorbed by the evaporator is estimated by the temperature difference between the evaporator inlet and outlet, and used as an indicator of the heat transfer capacity on the cold source side, using the following heat transfer relationship:
[0080]
[0081] In the formula, This refers to the heat exchange capacity index on the cold source side; The density of the refrigerant or chilled water is approximately 1000 kg / m³. The specific heat capacity of chilled water is approximately 4.2 kJ / (kg·°C); The current flow rate is obtained from the flow meter. These are the evaporator inlet temperature and outlet temperature, respectively.
[0082] Sub-step 1.5: Construct a unified time slice thermal state dataset.
[0083] Specifically, all the physical quantities obtained above ( , , , The data are integrated into the same timestamp according to the same sampling period, and integrity checks and outlier removal are performed to finally construct a complete thermal state basic dataset, in which each quantity corresponds to the physical state at the same time.
[0084] Step S120: Based on the thermophysical parameters, derive the sensible heat carried away from the air side and the cooling capacity provided by the cold source side, and output the heat exchange capacity evaluation result based on the sensible heat and cooling capacity.
[0085] In some embodiments, the IoT-based data center equipment monitoring method provided by the present invention includes the following steps in step S120:
[0086] Step S121: Based on the thermophysical parameters, perform a weighted average of the different temperature data of the air intake surface and cold aisle within the same time slice to obtain the average intake temperature, and combine it with the air supply temperature, return air temperature and air volume of the air conditioner to construct the key quantity group of the air side.
[0087] Step S122: Determine the evaporator inlet and outlet temperatures and chilled water flow rate based on thermophysical parameters, and construct a set of key quantities on the cold source side based on the evaporator inlet and outlet temperatures and chilled water flow rate.
[0088] In some embodiments, the IoT-based data center equipment monitoring method provided by the present invention further includes the following steps in step S120:
[0089] Step S123: Calculate the sensible heat carried away by the air side based on the key quantity group on the air side, and define the air side thermal balance deviation, which is equal to the difference between the server's heat output and sensible heat output.
[0090] Understandably, the calculation method for sensible heat is essentially the same as that for heat exchange; the only difference is the defined scenario.
[0091] Step S124: Calculate the cold energy on the cold source side based on the key quantity group on the cold source side, and define the cold and heat imbalance deviation, which is equal to the difference between the cold energy and the sensible heat on the cold source side.
[0092] Step S125: Define the capacity margin ratio, which is equal to the ratio of the thermal imbalance deviation as the molecular weight to the sensible heat. Based on the comparison between the capacity margin ratio and the set threshold range, divide the heat transfer capacity region.
[0093] The heat exchange capacity assessment result consists of the capacity margin ratio and the corresponding heat exchange capacity zones. Specifically, the heat exchange capacity zones are divided based on the matching degree between the cooling capacity on the cold source side and the heat dissipation demand on the air side. When the cooling capacity provided by the cold source is consistently higher than the sensible heat demand and remains stable, it is classified into the sufficient margin zone, indicating that the system has significant adjustment space. When the cooling capacity is only slightly higher than the demand, with insufficient redundancy and decreased resistance to load disturbances, it is classified into the critical approach zone, indicating that the system is on the verge of risk. When the cooling capacity on the cold source side is consistently lower than the air side demand, and the cooling capacity gap shows an expanding trend, it is classified into the significantly insufficient zone, indicating that the terminal heat exchange capacity can hardly support normal operation and immediate compensation or adjustment is required. The above zone division can serve as a basic reference for subsequent cold source attenuation identification and the generation of early control strategies.
[0094] In a specific embodiment, the IoT-based data center equipment monitoring method provided by this invention, step 2, constructs a terminal heat exchange capacity assessment result based on thermal balance and heat exchange mechanism. The basic thermal state data obtained in step 1 is input into the terminal heat exchange capacity assessment unit. This unit relies on an air-side thermal balance model and a heat exchanger characteristic model: on the one hand, it uses the relationship between the cabinet inlet / outlet temperature difference and the airflow to deduce the sensible heat carried away from the air side; on the other hand, it combines the evaporator inlet / outlet temperature difference indication with chilled water flow information to deduce the cooling capacity level that the cold source side can provide. This step obtains a qualitative terminal heat exchange capacity assessment result by comparing the degree of matching between the air-side heat change and the cooling capacity that the cold source side can provide, clarifying whether the current precision air conditioner is in a state region such as "sufficient margin," "critically close," or "significantly insufficient," including the following sub-steps:
[0095] Sub-step 2.1: Extract key quantities for the air side and cold source side from the thermal state baseline data.
[0096] Specifically, firstly, the rack inlet air temperature matrix is extracted from the thermal state dataset, and the temperatures of all racks at different heights within the same time slice are weighted and averaged to obtain an average temperature representing the overall cold aisle inlet air characteristics. To simplify engineering calculations, the temperatures of each monitoring point can be weighted equally. Then, from... Raw data on air conditioning supply air temperature, return air temperature, and air volume are extracted and combined with the average temperature to form a key air-side quantity set. Simultaneously, from... The evaporator inlet temperature, outlet temperature, and chilled water flow rate are extracted to form a key set of quantities on the cold source side.
[0097] Sub-step 2.2: Calculate the sensible heat carried away based on the air-side thermal balance model.
[0098] Specifically, the sensible heat carried away by the air can be calculated using the supply air volume, the temperature difference between the return air and the supply air. If the air density and specific heat capacity are known, then the sensible heat on the air side is equal to the product of the difference between the return air temperature and the supply air temperature, the air density (which can be taken as 1.2 kg / m³ in engineering), the specific heat capacity of air at constant pressure (which can be taken as 1.005 kJ / (kg·°C) in engineering), and the air conditioning supply air volume (m³ / s).
[0099] Next, to verify whether the sensible heat is basically balanced with the server's heat output, an air-side thermal balance deviation can be defined (if it is close to zero, it means that the heat removed is basically consistent with the server's heat output), equal to the difference between the server's heat output and the aforementioned sensible heat. If the air-side thermal balance deviation fluctuates within a small range, such as within ±10%, the air-side thermal balance can be considered good; if the deviation is too large, it indicates that there may be problems with the sensor data or model assumptions, which need to be corrected or marked in subsequent steps.
[0100] Sub-step 2.3: Calculate the available cooling capacity based on the heat exchange mechanism on the cold source side.
[0101] Specifically, the actual cooling capacity provided by the cold source side at a certain time can be calculated using the evaporator inlet temperature, outlet temperature, and flow rate. If the chilled water density and specific heat capacity are known, the cooling capacity on the cold source side is equal to the product of the inlet / outlet temperature difference, the chilled water density, the chilled water specific heat capacity, and the evaporator chilled water flow rate. To reflect the cooling and heating matching between the cold source side and the air side, the cooling and heating imbalance deviation can be calculated, which is equal to the difference between the cooling capacity on the cold source side and the sensible heat on the air side. When the cooling and heating imbalance deviation is positive, it indicates that the cold source side still has a certain margin; when it is close to zero or negative, it indicates that the cooling capacity provided by the cold source is insufficient to meet the actual heat removed by the air side, and the terminal heat exchange capacity is in a state of strain or inadequacy.
[0102] Sub-step 2.4: Construct terminal heat exchange capacity evaluation indicators and output capacity status.
[0103] Specifically, to quantitatively assess the redundancy capacity of the cold source side to the air side, a capacity margin ratio can be defined, which is measured by the proportion of the extra cooling capacity on the cold source side to the demand on the air side. This capacity margin ratio is equal to the ratio of the aforementioned heat and cold imbalance deviation as the molecular weight to the sensible heat carried away by the air side.
[0104] Subsequently, based on engineering experience, the capabilities can be divided into three areas, for example:
[0105] When the capacity margin ratio is greater than or equal to 0.2, it is marked as "sufficient margin";
[0106] When the capacity margin ratio is between 0 and 0.2, it is marked as "critically close";
[0107] When the capacity margin ratio is less than 0, it is marked as "significantly insufficient".
[0108] Finally, based on the above classification results, the terminal heat exchange capacity assessment results are constructed using the capacity margin ratio and its corresponding labels.
[0109] Step S130: Based on the heat exchange capacity assessment results within a set time period, construct the time evolution trajectory and, in conjunction with the pre-constructed reference model, identify the cold source change conditions inside the micro-module.
[0110] The aforementioned reference model can employ a three-layer LSTM temporal network as its core architecture. Specifically, each layer contains 128 hidden units, and a 20% dropout module is added after each layer to prevent overfitting under specific operating conditions. The training samples consist of chilled water supply and return temperature change sequences, flow regulation processes, and corresponding air conditioning heat exchange capacity change data recorded from multiple engineering sites. The data spans typical operating conditions such as nighttime energy-saving scheduling, partial load regulation, and full load switching. The training process uses a mean squared error loss function, with the Adam optimizer selected. The initial learning rate is set to 0.001, decaying to 90% of its original rate every 20 epochs. The entire model is trained for 200 epochs, and the consistency of the response speed and change curves on the validation set under different regulation events is used for verification. The final reference model can stably generate standard cold source response trajectories, serving as a benchmark for subsequent judgment of whether the cooling capacity exhibits abnormally slow decay.
[0111] Cold source change conditions refer to a sustained, slow capacity shift in the upstream system (e.g., chilled water system or refrigeration unit group) that provides cooling to the micro-modules over a period of time. This includes a gradual increase in supply and return water temperatures, a gradual decrease in chilled water flow rate, and a gradual decrease in the heat exchanger outlet temperature difference. These changes do not instantly trigger terminal temperature anomalies, but they gradually weaken the terminal heat exchange capacity. The key to identifying cold source change conditions lies in observing the evolution trend of the terminal heat exchange capacity over time. Specifically, the system first forms a capacity margin sequence within a fixed sampling period, then removes short-term noise through smoothing, and subsequently calculates the slow rate of capacity decay by combining the first-order inertial characteristics of the cold source system. When this downward trend persists within a time window without causing a significant increase in the cold aisle temperature, it can be determined that a cold source change condition has occurred, serving as a basis for proactive adjustments.
[0112] In some embodiments, the IoT-based data center equipment monitoring method provided by the present invention includes the following steps in step S130:
[0113] Step S131: Discretize the set time period according to a fixed sampling period to obtain discrete time, and extract the corresponding capacity margin ratio from the heat exchange capacity assessment results at each discrete time to construct a discrete capacity margin sequence.
[0114] Understandably, a discrete capacity margin sequence is a digital sequence arranged in chronological order to reflect the changing trend of heat transfer capacity, where each element represents the capacity margin ratio at a certain sampling time.
[0115] Step S132: The discrete sequence of capacity margin is smoothed over time using a first-order recursive smoothing algorithm to obtain the smoothed capacity margin curve, and the rate of change of capacity margin corresponding to each sampling point on the capacity margin curve is calculated.
[0116] The core idea of the first-order recursive smoothing algorithm is that the current smoothed value is obtained by weighting the current original value and the smoothed value from the previous time step according to a fixed ratio. This weighting ratio, i.e., the smoothing coefficient, is the only parameter that needs to be set, and the smoothing coefficient is generally set to 0.2~0.5. The larger the smoothing coefficient, the more sensitive the result; the smaller the coefficient, the more stable the result. This method can weaken short-term fluctuations and highlight the true trend.
[0117] Step S133: Normalize the capacity margin change rate by introducing a time constant to obtain a capacity change rate index, and compare the capacity change rate index with a set threshold parameter within the selected time window to identify the cold source change conditions.
[0118] In a specific embodiment, the IoT-based data center equipment monitoring method provided by this invention, step 3, identifies precursors to the gradual change in cold source operating conditions by combining dynamic response mechanisms. The terminal heat exchange capacity assessment results from step 2 are input into the cold source gradual change operating condition identification unit. This unit no longer only considers a single assessment result, but constructs a time evolution trajectory based on continuous assessment results over a period of time, and identifies whether there is a continuous, slow decreasing trend by referring to the first-order inertial response characteristics of the chilled water system and the internal thermal inertia model of the micro-module. That is, cold source regulation typically manifests as a gradual shift in supply and return water temperatures and flow rates over several minutes to several tens of minutes, and its impact on terminal heat exchange capacity will inevitably exhibit a similar gradual decay curve. Within a sliding time window, the cold source gradual change operating condition identification unit compares the current terminal heat exchange capacity with the baseline level over a past period. If multiple consecutive time windows show a slow decrease in heat exchange capacity without triggering a significant rise in cold aisle temperature, it is determined that there are precursors to a decline in cold source capacity. This includes the following sub-steps:
[0119] Sub-step 3.1: Construct a discrete time series of the end-point heat exchange capacity margin.
[0120] Specifically, the continuous time mentioned above is discretized according to a fixed sampling time interval (10 to 60 seconds in engineering) to obtain discrete time. At each discrete time, the corresponding capacity margin ratio is extracted from the above end heat exchange capacity assessment results. Arranged according to time, the capacity margin data in the form of discrete time series can be obtained, that is, the discrete capacity margin sequence.
[0121] Sub-step 3.2 involves time smoothing of the capacity margin curve to suppress short-term disturbances.
[0122] Specifically, to filter out short-term noise caused by instantaneous fluctuations in server load and short-term adjustments in fan speed, the discrete sequence of capacity margin needs to be smoothed over time. A first-order recursive smoothing method can be used, combining the current discrete sequence of capacity margin with the smoothed value from the previous time step according to weights to obtain the smoothed capacity margin curve, expressed as:
[0123]
[0124] In the formula, Indicates the first The smoothing margin value at each sampling time; Indicates the first The original capability margin value at each sampling time; Indicates the first The smoothing margin value at each sampling time; This is a smoothing coefficient, with a value ranging from 0 to 1, and in engineering applications, it can be taken as 0.2 to 0.5. The sampling point number, . Can be set to This indicates that the smoothed value is equal to the original value during the first sampling.
[0125] Sub-step 3.3 involves calculating and normalizing the capability decay rate using a first-order inertial model.
[0126] Specifically, chilled water systems typically exhibit first-order inertial characteristics when subjected to load regulation or supply / return water temperature adjustments. This means their output (cooling capacity) does not change instantaneously but rather gradually over a time constant. To reflect whether the capacity is slowly decreasing, calculations are needed. The rate of change over time.
[0127] First, calculate the rate of change of the smoothing capability margin between adjacent sampling points, expressed as:
[0128]
[0129] In the formula, For the first The rate of change of raw capacity at each sampling time (unit: per second); This represents the current smoothing capacity margin; This is the smoothing capacity margin from the previous moment; This represents the sampling time interval.
[0130] To link this rate of change with the first-order inertial characteristics of the cold source system, a time constant τ_c is introduced to normalize the rate of change, resulting in the capacity change rate index, which is the product of the aforementioned capacity change rate and the time constant (the first-order inertial time constant of the cold source system, typically ranging from a few minutes to tens of minutes, for example, 300 seconds to 1800 seconds).
[0131] In this embodiment, if the above-mentioned rate of change of capacity is negative for several consecutive sampling periods and the absolute value is not large (e.g., near a certain negative threshold), it indicates that the capacity is decreasing in a "slow but continuous" manner, which is consistent with the typical characteristics of slow-change regulation of cold source; if large positive and negative fluctuations occur in a short period of time, it is more likely to be a momentary disturbance of server load or local operating status.
[0132] Sub-step 3.4: Determine the precursors of the slow change in the cold source operating condition and output the trend results.
[0133] Specifically, to transform the aforementioned capacity change rate index into the conclusion of "whether there are precursors to a slow-changing cold source operating condition," it is necessary to perform an cumulative judgment within a time window. A duration window is selected, for example, 10 to 30 minutes, and the corresponding number of sampling points is calculated. The number of sampling points is equal to the ratio of the selected time window (numerator) to the sampling period. Then, at the current sampling moment, the average of all capacity change rate indices within this window (the moment when the calculated number of sampling points is subtracted from the current sampling moment plus 1, i.e., the average capacity change rate index) is examined. A negative threshold is then set (e.g., between -0.05 and -0.1). When the average capacity change rate index is less than this negative threshold, it indicates that the overall capacity margin has been in a state of continuous slow decline over the recent period. At this point, the trend measure of the slow-changing cold source operating condition (the total decrease in capacity margin from the window start point to the current moment; a positive value indicates a capacity decline) can be equal to the difference between the smoothed capacity margin value at the window start point and the smoothed capacity margin at the current moment. Finally, based on the average capacity change rate index and trend measurement, the trend judgment result of the slow change condition of the cold source can be output. For example, when the average capacity change rate index is less than the negative threshold and the trend measurement is greater than the preset positive value, the judgment result is recorded as "there is a precursor to the slow decrease in cold source capacity"; otherwise, it is recorded as "no obvious precursor to slow change was found".
[0134] Step S140: When there are signs of decline in the cold source capacity corresponding to the cold source change operating condition, calculate the cold source capacity compensation amount and generate a refrigeration control strategy based on the cold source capacity compensation amount.
[0135] Understandably, the cold source capacity compensation refers to the shortfall in cooling capacity that needs to be replenished in advance to maintain stable cold aisle temperatures before the upstream cold source capacity begins to weaken but the terminal temperature has not yet become abnormal. It reflects the additional heat exchange load that the terminal air conditioning system must bear when the cold source continues to weaken in the future, guiding strategies such as increasing air volume, increasing the number of units deployed, and switching operating modes. Specifically, the compensation calculation is based on two core criteria: first, the immediate difference between the current sensible heat demand on the air side and the available cooling capacity on the cold source side; and second, the future decline in cold source capacity given by the trend model. The system first determines whether the current cooling capacity is insufficient, then combines this with the declining trend of the capacity margin within the time window, converting the potentially further weakening portion into an expected shortfall. Subsequently, a safety factor is set based on engineering experience, comprehensively amplifying the immediate cooling capacity shortfall and the trend decline, so that the compensation covers the worst-case scenario in the short term. The final compensation is the additional heat exchange capacity that the terminal must generate in advance, used to drive the subsequent calculation and execution of refrigeration control strategies.
[0136] In some embodiments, the IoT-based data center equipment monitoring method provided by the present invention includes the following steps in step S140:
[0137] Step S141: Convert the cold source change condition into a change trend measure, and when there are signs of cold source capacity decay in the change trend measure, calculate the temperature difference range between the upper limit of the allowable temperature and the target temperature of the cold aisle.
[0138] Step S142: Calculate the cooling capacity gap between the sensible heat carried away from the air side and the cooling capacity provided by the cold source side, and determine the required cooling capacity compensation amount according to the changing operating conditions of the cold source, so that when the cooling capacity gap exceeds the corresponding set threshold, the air supply volume of the air conditioner and the number of operating units are adjusted according to the cooling capacity compensation amount.
[0139] Step S143: Based on the adjusted air conditioning air volume and the number of operating units, combined with the target fan speed and the operating status of each unit, a cooling control strategy is generated.
[0140] In a specific embodiment, the IoT-based data center equipment monitoring method provided by this invention includes step 4, which generates a pre-adjusted cooling control strategy based on the thermal inertia compensation mechanism. The cold source operating condition change trend judgment result obtained in step 3 is input into the cooling control strategy generation unit. When the judgment result shows that there are signs of a continuous decline in cold source capacity, the cooling control strategy generation unit no longer uses the traditional temperature feedback logic. Instead, based on the thermal inertia model and thermal balance model of the equipment inside the micro-module, it pre-calculates the compensation measures needed to maintain a stable cold aisle temperature under the condition of declining cold source capacity. These measures include increasing air supply volume, optimizing air conditioning operation mode, pre-activating backup air conditioning units, and adjusting air supply paths. The core of the thermal inertia compensation mechanism is that the cabinets, equipment, and air inside the micro-module can store a certain amount of heat fluctuation in a short period. If the terminal heat exchange capacity is increased when the cold source just begins to weaken, this available buffer can be used to combat the impending insufficient cooling capacity, thereby reducing and delaying the peak temperature rise. This includes the following sub-steps:
[0141] Sub-step 4.1: Determine the prediction time window and allowable temperature rise range based on the trend results.
[0142] Specifically, first, it is determined whether the above-mentioned cold source gradual change operating condition trend judgment result indicates a precursor to a slow decline in cold source capacity. If so, the advance adjustment logic is initiated; otherwise, the original control logic is maintained, and subsequent compensation calculations are not performed. When there are precursors to attenuation, an allowable upper temperature limit (e.g., 27 degrees Celsius) is set based on the data center's design requirements for cold aisle temperature, and the current target set temperature for the cold aisle (e.g., 24 degrees Celsius) is known. The difference between the two is the allowable temperature rise range. Then, to match the first-order inertial response characteristics of the cold source and the thermal inertia of the micromodule, a prediction time window can be set to ensure that the advance adjustment action is completed before the cold source capacity further declines. The set prediction time window can be obtained by amplifying the time window in step 3, for example, by doubling it.
[0143] Sub-step 4.2: Estimate the future cooling capacity gap and the required compensation cooling capacity based on trends and thermal balance.
[0144] Specifically, step 2 has already obtained the sensible heat on the air side and the cooling capacity on the cold source side at the current moment. If there are signs of a slow decline in the cooling capacity, it generally means that the cooling capacity on the cold source side will continue to decrease in the future. To simplify the engineering calculation, the cooling capacity gap at the current moment can be calculated first (if it is positive, it means that the cooling capacity on the cold source side is already insufficient). This cooling capacity gap is equal to the difference between the sensible heat on the air side and the cooling capacity on the cold source side.
[0145] In a slowly changing scenario, the total decrease in capacity margin from the start of the window to the current time represents the cumulative decrease in capacity margin within a window, which can be considered as the degree of impending further deterioration. To leave a safety margin within the prediction time window, a safety factor (greater than or equal to 1, typically 1.2 to 1.5) can be used to amplify the current cooling capacity gap, resulting in the required compensation cooling capacity. This required compensation cooling capacity is equal to the product of the safety factor and the current cooling capacity gap.
[0146] When the pre-cooling capacity gap is less than or equal to 0, it indicates that the current cooling source can still meet the demand. However, due to the gradual decline trend, the required compensation cooling capacity can be set to a pre-compensation value that is proportional to the total decrease in capacity margin from the start of the window to the current time. The specific proportional coefficient is determined by engineering experience.
[0147] Sub-step 4.3 decomposes the compensation cooling capacity into the increase in air conditioning supply volume and the number of units put into operation.
[0148] Specifically, firstly, the amount of cooling removed by the air conditioner through the air side is equal to the product of air density, air specific heat capacity, supply air volume, and the temperature difference between the return and supply air. When additional cooling capacity is needed, the air-side heat exchange capacity can be increased by increasing the supply air volume while keeping the temperature difference between the return and supply air constant. The increase in supply air volume required for compensation is estimated by the following formula:
[0149]
[0150] In the formula, For the first Increased air supply volume required per cycle (unit: m³ / s). This is to compensate for the increased cooling capacity. These are air density, specific heat capacity, and the temperature difference between return air and supply air, respectively.
[0151] Considering that a single precision air conditioning unit has a maximum air supply capacity and rated cooling capacity, when the air supply volume of the existing operating unit is close to the upper limit, the additional air supply volume alone may not be able to fully absorb the additional cooling capacity required. In this case, it is necessary to calculate the number of units that need to be added. The number of units added is equal to the ratio of the additional cooling capacity to the rated cooling capacity of a single unit under design conditions.
[0152] Finally, by combining the required increase in air volume and the number of additional units, the required air volume and unit compensation strategy set can be obtained, providing a basis for generating specific control commands in the future.
[0153] Sub-step 4.4 generates a thermal inertia compensation control strategy and outputs a control instruction set.
[0154] Specifically, based on the approximate cubic relationship between air volume and fan speed, the air volume can be considered approximately proportional to the fan speed near the current operating conditions. Therefore, the current air volume is equal to the product of the fan flow coefficient (determined by fan selection and commissioning, and considered a constant under current conditions) and the current fan speed. Based on this, the required increase in fan speed is equal to the ratio of the required increase in air volume (numerator) to the fan flow coefficient, thus obtaining the new target fan speed. Furthermore, if this new target fan speed exceeds the set upper limit, it is limited to the set upper limit in the control logic, and the cooling capacity difference that cannot be compensated by air volume is transferred to the corresponding standby unit among the newly added units, allowing the standby unit to share the burden. Simultaneously, for the commissioning of new units, idle precision air conditioning units are selected sequentially according to the aforementioned number of new units, and their operation flag is set to "start." The operation flag indicates the operating status of a unit, and its value can be "stop" or "start."
[0155] Finally, the new target fan speed and the operating indicators of each unit are combined into a set of cooling control strategies, which are then issued and executed.
[0156] Step S150: Execute the refrigeration control strategy to adjust the corresponding temperature control equipment according to the compensation amount of the cold source capacity, and monitor the adjusted thermophysical parameters in real time.
[0157] In some embodiments, the IoT-based data center equipment monitoring method provided by the present invention includes the following steps in step S150:
[0158] Step S151: Control the operation of the corresponding air conditioning unit according to the operating status of each unit and the target speed of the fan, and simultaneously collect the operating thermophysical parameters within the set operating cycle.
[0159] Step S152: Calculate the actual physical deviation between the operating thermophysical parameters and the set target parameters, and determine whether there is temperature overshoot in the current operation based on the actual physical deviation.
[0160] Step S153: When there is temperature overshoot in the current operation, determine the overshoot ratio, and compare the actual physical deviation and overshoot ratio with the set physical deviation threshold and overshoot ratio threshold respectively, so as to correct the parameters according to the comparison results.
[0161] In a specific embodiment, the IoT-based data center equipment monitoring method provided by this invention, in step 5, integrates the execution results and monitoring feedback to form a closed-loop monitoring conclusion. The cooling control strategy generated in step 4 is distributed to the precision air conditioner and related execution units, causing the micro-modules to execute under the new operating conditions for a period of time. During execution, data such as rack inlet air temperature, cold aisle temperature, air conditioning supply air temperature, air volume, and equipment power consumption are re-collected and compared with the expected temperature change envelope in step 4 to determine whether the advance adjustment effectively suppressed temperature overshoot. If the actual temperature change is within the expected range, it indicates that the IoT-based monitoring and advance adjustment strategy has effectively established a stable response mechanism under the slowly changing cold source operating conditions; if a significant temperature overshoot still occurs, this process is recorded as a special operating condition for subsequent correction of the parameters of the thermal inertia model and heat transfer capacity evaluation model, so that the system continuously approaches real engineering behavior during long-term operation. This includes the following sub-steps:
[0162] Sub-step 5.1: Execute the control strategy and collect a new round of thermal state data.
[0163] Specifically, within the current control cycle, the control strategy output from step 4 is sent to the field execution layer, setting the target fan speed to the new target fan speed, and starting and stopping the corresponding precision air conditioning units according to the operation markers. The system maintains the new operating conditions for a period of time (e.g., 5 to 15 minutes) to allow the temperature and airflow to reach a new stable or quasi-stable state. Simultaneously, during this operating period, thermophysical parameters are collected according to the set sampling period, namely the actual temperature sequence at key locations in the cold aisle, the sum of the actual airflow of each operating unit, and the actual power consumption or current changes of the IT equipment in the cabinet, converted into real-time heat generation. Subsequently, all collected thermophysical parameters form a sample point at each discrete moment, and the data over the entire operating time interval constitutes the post-execution thermal state data set.
[0164] Sub-step 5.2 calculates the deviation between the actual temperature and the expected temperature envelope and the overshoot index.
[0165] Specifically, in step 4, the expected envelope of the cold aisle temperature has been given based on the thermal inertia model. Therefore, at each sampling moment, the difference between the actual measured temperature and the expected value is calculated to obtain the corresponding temperature deviation. Next, to assess whether temperature overshoot exists, the peak value of the actual temperature needs to be found within the execution window and compared with the system set temperature and the allowable temperature rise range. The allowable temperature rise range is equal to the difference between the upper limit of the cold aisle's allowable temperature and the set temperature for normal operation. Subsequently, the maximum value of the actual measured temperature is found within the execution window as the actual temperature peak value, and the overshoot ratio is calculated. The overshoot ratio is equal to the ratio of the difference between the actual temperature peak value and the normal operation set temperature (as the numerator) to the allowable temperature rise range. Finally, the temperature deviation between the actual measured temperature and the expected value, combined with the calculated overshoot ratio, is used to output an overshoot index set.
[0166] Sub-step 5.3 provides closed-loop monitoring conclusions based on error indicators.
[0167] Specifically, two thresholds are set according to engineering requirements: a temperature deviation threshold (e.g., 1 to 2 degrees Celsius) and an overshoot ratio threshold (generally not exceeding 1, but can be 0.8 to 1.0 in engineering practice). Within the execution window, the maximum absolute value of the temperature deviation for all sampling points is taken. If this maximum value is less than or equal to the temperature deviation threshold and the overshoot ratio is less than or equal to the overshoot ratio threshold, the monitoring conclusion is recorded as effective control; otherwise, it is recorded as insufficient control. Next, to indicate how much safety margin the current control strategy has from the boundary, a safety margin can be defined as 1 minus the overshoot ratio. The closer it is to 1, the safer it is; closer to 0, the critical; and less than 0, the limit has been exceeded. Finally, the above safety margin and monitoring result labels are integrated and output as a closed-loop monitoring conclusion set.
[0168] Sub-step 5.4 involves modifying the model and updating parameters for special working conditions.
[0169] Specifically, when the monitoring results are marked as "effective control," the parameters of the existing thermal inertia model and heat transfer capacity assessment model can be kept unchanged, and only this execution is recorded as a normal sample. When marked as "insufficient control," it indicates that the current model has a deviation in its characterization of the slowly changing operating conditions or thermal inertia characteristics of the cold source, and correction is required. Therefore, a model error index can be defined to comprehensively consider temperature deviation and overshoot ratio, expressed as:
[0170]
[0171] In the formula, This is a model error index; the larger the value, the less accurate the model is. , These are all weighting coefficients, with values ranging from 0 to 1, used to balance the effects of temperature deviation and overshoot / excessive limits. To determine the maximum absolute value of the temperature deviation within the execution window; This is the temperature overshoot ratio.
[0172] Then, a simple parameter correction strategy is adopted, such as recursively updating a key model parameter (e.g., heat capacity estimation parameter, cold source response time constant, etc.), expressed as:
[0173]
[0174] In the formula, These are the model parameter values before correction; These are the corrected model parameter values; This can be the learning rate or correction step size, and can range from 0.001 to 0.1, depending on engineering experience.
[0175] Finally, all parameters that need to be corrected are updated in a similar manner to form a new set of model parameters.
[0176] like Figure 2 As shown, in one embodiment, a data center equipment monitoring system based on the Internet of Things (IoT) includes an IoT acquisition module, a heat exchange capacity assessment module, an operating condition identification module, a control strategy generation module, and a strategy execution module.
[0177] The IoT acquisition module is used to synchronously acquire the thermophysical parameters inside the micro-module.
[0178] The heat exchange capacity assessment module is used to deduce the sensible heat carried away from the air side and the cooling capacity provided by the cold source side based on thermophysical parameters, and outputs the heat exchange capacity assessment results based on the sensible heat and cooling capacity.
[0179] The operating condition identification module is used to construct a time evolution trajectory based on the heat exchange capacity assessment results within a set time period, and in combination with a pre-built reference model, identify the cold source change conditions inside the micro-module.
[0180] The control strategy generation module is used to calculate the cold source capacity compensation amount when there are signs of impending attenuation in the cold source capacity corresponding to the cold source change operating conditions, and to generate a refrigeration control strategy based on the cold source capacity compensation amount.
[0181] The strategy execution module is used to execute the refrigeration control strategy to adjust the corresponding temperature control equipment according to the compensation amount of the cold source capacity, and to monitor the adjusted thermophysical parameters in real time.
[0182] The applicant of this invention has provided a detailed description of the embodiments of the invention in conjunction with the accompanying drawings. However, those skilled in the art should understand that the above embodiments are merely preferred embodiments of the invention. The detailed description is only intended to help readers better understand the spirit of the invention and is not intended to limit the scope of protection of the invention. On the contrary, any improvements or modifications made based on the inventive spirit of the invention should fall within the scope of protection of the invention.
Claims
1. A method for monitoring data center equipment based on the Internet of Things, characterized in that, The method includes: Synchronously collect the thermophysical parameters inside the micromodule; The sensible heat carried away from the air side and the cooling capacity provided by the cold source side are derived based on the aforementioned thermophysical parameters, and the heat exchange capacity evaluation result is output based on the aforementioned sensible heat and cooling capacity. Based on the heat exchange capacity assessment results within a set time period, a time evolution trajectory is constructed, and combined with a pre-constructed reference model, the cold source change conditions inside the micro-module are identified. When there are signs of impending attenuation in the cold source capacity corresponding to the cold source change condition, calculate the cold source capacity compensation amount and generate a refrigeration control strategy based on the cold source capacity compensation amount. The cooling control strategy is executed to adjust the corresponding temperature control equipment according to the compensation amount of the cold source capacity, and the adjusted thermophysical parameters are monitored in real time.
2. The data center equipment monitoring method based on the Internet of Things according to claim 1, characterized in that, The synchronous acquisition of internal thermophysical parameters of the micromodule includes: Multiple temperature sensors are installed on the air intake side of each rack and at different heights in the cold aisle to collect temperature data inside the micro-modules, and a temperature distribution matrix is constructed based on the temperature data. The air volume, air temperature and return air temperature provided by the air conditioner fan are obtained, and the instantaneous heat exchange of the air conditioner is approximately calculated based on the air volume, air temperature and return air temperature. Based on the current and voltage corresponding to the power consumption of the server, the power consumption of the server is approximately converted into the heat generated by the server, and the heat exchange capacity index of the cold source side is estimated based on the temperature difference between the inlet and outlet of the evaporator. The thermophysical parameters are obtained by integrating the temperature distribution matrix, the instantaneous heat exchange of the air conditioner, the heat output of the server, and the heat exchange capacity index of the cold source side after time alignment and correction.
3. The method for monitoring data center equipment based on the Internet of Things according to claim 2, characterized in that, The process of deriving the sensible heat carried away from the air side and the cooling capacity provided by the cold source side based on the thermophysical parameters, and outputting the heat exchange capacity evaluation result based on the sensible heat and cooling capacity, includes: Based on the aforementioned thermophysical parameters, a weighted average of different temperature data from the air intake surface and cold aisle within the same time slice is calculated to obtain the average intake temperature. This average temperature is then combined with the air supply temperature, return air temperature, and air volume of the air conditioner to construct a key air-side quantity group. The evaporator inlet and outlet temperatures and chilled water flow rate are determined based on the thermophysical parameters, and a set of key quantities on the cold source side is constructed based on the evaporator inlet and outlet temperatures and chilled water flow rate.
4. The data center equipment monitoring method based on the Internet of Things according to claim 3, characterized in that, The step of deriving the sensible heat removed from the air side and the cooling capacity provided by the cold source side based on the thermophysical parameters, and outputting the heat exchange capacity evaluation result based on the sensible heat and cooling capacity, further includes: The sensible heat carried away by the air side is calculated based on the key air-side quantity set, and the air-side thermal balance deviation is defined as the difference between the server's heat output and the sensible heat. The cold energy on the cold source side is calculated based on the key quantity set on the cold source side, and the cold-heat imbalance deviation is defined. The cold-heat imbalance deviation is equal to the difference between the cold energy on the cold source side and the sensible heat. Define a capacity margin ratio, which is equal to the ratio of the thermal imbalance deviation as a molecule to the sensible heat, and divide the heat transfer capacity region based on the comparison result of the capacity margin ratio and a set threshold range. The heat exchange capacity assessment result consists of the capacity margin ratio and the corresponding heat exchange capacity region.
5. The method for monitoring data center equipment based on the Internet of Things according to claim 4, characterized in that, Based on the heat exchange capacity assessment results within a set time period, a time evolution trajectory is constructed, and combined with a pre-constructed reference model, the changing operating conditions of the cold source inside the micro-module are identified, including: The set time period is discretized according to a fixed sampling period to obtain discrete time. At each discrete time, the capacity margin ratio at the corresponding time is extracted from the heat exchange capacity assessment result to construct a discrete sequence of capacity margin. The discrete sequence of the capability margin is smoothed over time using a first-order recursive smoothing algorithm to obtain a smoothed capability margin curve, and the rate of change of capability margin corresponding to each sampling point on the capability margin curve is calculated. The capacity margin change rate is normalized by introducing a time constant to obtain a capacity change rate index. The capacity change rate index is then compared with a set threshold parameter within a selected time window to identify the cold source change conditions.
6. The method for monitoring data center equipment based on the Internet of Things according to claim 1, characterized in that, When there are signs of impending capacity reduction in the cold source capacity corresponding to the changing operating conditions of the cold source, a cold source capacity compensation amount is calculated, and a refrigeration control strategy is generated based on the cold source capacity compensation amount, including: The cold source change conditions are converted into a trend measurement, and when there are signs of cold source capacity decay in the trend measurement, the temperature difference range between the upper limit of the allowable temperature and the target temperature of the cold aisle is calculated. Calculate the cooling capacity gap between the sensible heat carried away from the air side and the cooling capacity provided by the cold source side, and determine the required cooling capacity compensation amount according to the changing operating conditions of the cold source, so that when the cooling capacity gap exceeds the corresponding set threshold, the air supply volume of the air conditioner and the number of operating units are adjusted according to the cooling capacity compensation amount. Based on the adjusted air conditioning air volume and the number of operating units, combined with the target fan speed and the operating status of each unit, the cooling control strategy is generated.
7. The method for monitoring data center equipment based on the Internet of Things according to claim 6, characterized in that, The execution of the refrigeration control strategy, which adjusts the corresponding temperature control equipment according to the cold source capacity compensation amount, and monitors the adjusted thermophysical parameters in real time, includes: The corresponding air conditioning unit is controlled to operate according to the operating status of each unit and the target speed of the fan, and the operating thermophysical parameters within the set operating cycle are collected synchronously. Calculate the actual physical deviation between the operating thermophysical parameters and the set target parameters, and determine whether there is temperature overshoot in the current operation based on the actual physical deviation; When temperature overshoot occurs during current operation, the overshoot ratio is determined, and the actual physical deviation and overshoot ratio are compared with the set physical deviation threshold and overshoot ratio threshold, respectively, so as to correct the parameters based on the comparison results.
8. A data center equipment monitoring system based on the Internet of Things, characterized in that, The system is used to implement the Internet of Things-based data center equipment monitoring method according to any one of claims 1 to 7, the system comprising: The Internet of Things (IoT) acquisition module is used to synchronously acquire the thermophysical parameters inside the micro-module; The heat exchange capacity assessment module is used to deduce the sensible heat carried away from the air side and the cooling capacity provided by the cold source side based on the thermophysical parameters, and output the heat exchange capacity assessment result based on the sensible heat and cooling capacity. The operating condition identification module is used to construct a time evolution trajectory based on the heat exchange capacity evaluation results within a set time period, and identify the cold source change operating conditions inside the micro-module in combination with a pre-built reference model. The control strategy generation module is used to calculate the cold source capacity compensation amount when there are signs of decline in the cold source capacity corresponding to the cold source change operating condition, and generate a refrigeration control strategy based on the cold source capacity compensation amount. The strategy execution module is used to execute the refrigeration control strategy to adjust the corresponding temperature control equipment according to the compensation amount of the cold source capacity, and to monitor the adjusted thermophysical parameters in real time.
9. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-7.