A server cluster radiator energy consumption intelligent monitoring system

By introducing energy consumption acquisition, threshold judgment, heat dissipation efficiency analysis, and regulation correction units into the server cluster heat sink energy consumption monitoring system, the acquisition frequency and regulation intensity are dynamically adjusted, solving the problems of fixed acquisition frequency and insufficient anomaly assessment in the existing technology, and realizing more efficient anomaly capture and system stability monitoring.

CN120891748BActive Publication Date: 2025-12-05CHANGSHU YOUBANG RADIATOR
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511368971.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-12-05
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing server cluster heat sink energy consumption monitoring systems cannot dynamically adjust the acquisition frequency, resulting in delayed capture of abnormal signals or waste of resources. They lack quantitative assessment of the degree of abnormality, making it difficult to predict potential system risks, and their overall stability and adaptability are insufficient.

Method used

The system employs an energy consumption acquisition unit, a threshold judgment unit, a heat dissipation efficiency analysis unit, and a regulation correction unit. By periodically acquiring energy consumption parameters, it generates different types of abnormal signals, calculates parameter deviation and regulation lag ratio, dynamically adjusts the acquisition cycle and regulation intensity, and combines the system stability monitoring unit to statistically analyze the total number of abnormal signals and the frequency of runaway, thereby achieving real-time monitoring and early warning.

Benefits of technology

It improves the accuracy of abnormal signal detection and resource utilization efficiency, enables timely adjustment of control strategies, reduces the difficulty of troubleshooting, and enhances the operational continuity and reliability of server clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120891748B_ABST
    Figure CN120891748B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of server heat dissipation monitoring, and discloses a server cluster radiator energy consumption intelligent monitoring system. An energy consumption acquisition unit of the system periodically acquires radiator energy consumption parameters and sends the parameters to a threshold judgment unit; the threshold judgment unit generates temperature, flow or pressure abnormal signals through threshold comparison; a heat dissipation efficiency analysis unit calculates the deviation amount of the energy consumption parameters from preset thresholds based on the abnormal signals, generates a heat dissipation regulation lag proportion, and calculates the heat dissipation regulation efficiency by counting the regulation time consumption; a regulation correction unit adjusts the acquisition period according to the lag proportion and adjusts the regulation intensity according to the regulation efficiency; a system stability monitoring unit counts the total amount of abnormal signals, calculates the out-of-control frequency in combination with the running time length, and generates a system failure signal when the out-of-control frequency exceeds a threshold. The system realizes dynamic adjustment and stability early warning of energy consumption monitoring, and improves the reliability and adaptability of server cluster radiator operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of server heat dissipation monitoring technology, specifically to an intelligent monitoring system for the energy consumption of server cluster heat sinks. Background Technology

[0002] In the information age, server clusters serve as the core carriers of data processing and storage, and their stable operation directly affects the smooth operation of various business systems. With the explosive growth of data volume and the continuous increase in computing demands, the scale of server clusters continues to expand, and the heat generated during operation also increases significantly. As a key device for maintaining the temperature balance of server clusters, the energy consumption and heat dissipation efficiency of heat sinks are closely related. Once abnormal energy consumption occurs, it will not only increase operating costs, but may also lead to server overheating and shutdown, resulting in data loss or business interruption.

[0003] Server cluster heatsink energy consumption monitoring typically employs a traditional fixed-period data collection mode, where the collection frequency cannot be dynamically adjusted according to actual cooling needs. When the heatsink experiences slight energy consumption fluctuations, the fixed-period collection method may fail to capture abnormal signals in a timely manner, leading to delayed anomaly detection. Conversely, when the cooling status is stable, excessively frequent data collection results in resource waste. Regarding threshold judgment, existing systems usually only perform simple threshold comparisons, generating a single anomaly alert, lacking quantitative analysis of the deviation between energy consumption parameters and preset thresholds, making it difficult to accurately determine the severity and scope of the anomaly.

[0004] Current technologies for assessing heat dissipation efficiency often focus on monitoring the results, neglecting the time consumption and lag issues during the heat dissipation control process. This makes it difficult for maintenance personnel to fully grasp the response speed of the heat dissipation system and to optimize control strategies accordingly. Furthermore, the control mechanism lacks dynamic correction capabilities; the control intensity remains fixed and cannot be adaptively adjusted according to changes in heat dissipation efficiency, leading to instances of insufficient or excessive control in certain scenarios.

[0005] Existing monitoring systems do not pay enough attention to overall stability and lack statistical analysis of the total number of abnormal signals and the frequency of system failures. While a single anomaly may not immediately lead to system failure during long-term operation, the risk of system failure increases significantly after multiple anomalies accumulate. Due to the lack of an effective mechanism for assessing the frequency of system failures, maintenance personnel find it difficult to predict potential risks in advance and often only take remedial action after system failure, increasing the difficulty of troubleshooting and recovery. These problems collectively result in insufficient accuracy, timeliness, and adaptability of current server cluster heatsink energy consumption monitoring, failing to meet the needs of efficient operation and maintenance of large-scale server clusters. Summary of the Invention

[0006] The purpose of this invention is to provide an intelligent monitoring system for the energy consumption of server cluster heat sinks, so as to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides an intelligent monitoring system for the energy consumption of server cluster heat sinks, the system comprising:

[0008] The energy consumption acquisition unit periodically collects the energy consumption parameters of the server cluster heat sink and sends the collection results to the threshold judgment unit.

[0009] The threshold judgment unit compares the energy consumption parameters within the threshold range and generates abnormal temperature signals, abnormal flow signals, or abnormal pressure signals based on the comparison results.

[0010] The heat dissipation efficiency analysis unit obtains abnormal signals and corresponding energy consumption parameters through the threshold judgment unit, calculates the deviation between the energy consumption parameters and the preset threshold, generates the heat dissipation regulation lag ratio based on the deviation, and calculates the time consumption of the heat dissipation regulation process to generate the heat dissipation regulation efficiency.

[0011] The regulation and correction unit obtains the heat dissipation regulation lag ratio and heat dissipation regulation efficiency through the heat dissipation efficiency analysis unit, adjusts the acquisition cycle of the energy consumption acquisition unit according to the heat dissipation regulation lag ratio, and adjusts the heat dissipation regulation intensity according to the heat dissipation regulation efficiency.

[0012] The system stability monitoring unit counts the total number of abnormal signals output by the heat dissipation efficiency analysis unit, calculates the runaway frequency in conjunction with the server cluster runtime, and generates a system failure signal when the runaway frequency exceeds a preset threshold.

[0013] Preferably, the threshold judgment unit compares the heat dissipation temperature, coolant flow rate, and pipe pressure in the energy consumption parameters with the corresponding temperature threshold range, flow rate threshold range, and pressure threshold range, respectively. If the heat dissipation temperature exceeds the temperature threshold range, a temperature abnormality signal is generated; if the coolant flow rate exceeds the flow rate threshold range, a flow rate abnormality signal is generated; and if the pipe pressure exceeds the pressure threshold range, a pressure abnormality signal is generated.

[0014] Preferably, the heat dissipation efficiency analysis unit extracts the heat dissipation temperature value corresponding to the temperature anomaly signal, the coolant flow rate value corresponding to the flow anomaly signal, or the pipe pressure value corresponding to the pressure anomaly signal, and records them as anomaly parameter H, and extracts the endpoint values ​​h0 and h1 of the corresponding threshold range;

[0015] The heat dissipation efficiency analysis unit calculates the difference between abnormal parameters H and h0 and the difference between H and h1, selects the minimum difference as the parameter offset, and generates the heat dissipation control hysteresis ratio based on the ratio between the parameter offset and the preset benchmark value.

[0016] The heat dissipation efficiency analysis unit records the time elapsed from the moment the abnormal signal is triggered until the energy consumption parameter recovers to the threshold range, and compares this time elapsed with a preset standard duration to generate the heat dissipation regulation efficiency.

[0017] Preferably, when the heat dissipation regulation lag ratio exceeds the allowable range, the regulation correction unit generates a sampling cycle shortening instruction and sends it to the energy consumption sampling unit, and the energy consumption sampling unit shortens the sampling cycle based on the heat dissipation regulation lag ratio.

[0018] When the heat dissipation regulation efficiency is lower than the preset standard, the regulation correction unit generates a regulation intensity enhancement command and sends it to the threshold judgment unit. The threshold judgment unit increases the power level of the heat dissipation regulation operation according to the regulation intensity enhancement command.

[0019] Preferably, it also includes a heat dissipation topology modeling unit, which obtains the connection relationship of each heat dissipation node in the server cluster, sorts the nodes according to the hierarchical priority of the heat dissipation nodes, and constructs a heat dissipation node sequence based on the sorting result;

[0020] The heat dissipation topology modeling unit extracts the connection direction and span parameters of adjacent nodes in the heat dissipation node sequence, and generates the heat dissipation path topology structure by combining the hierarchical difference value.

[0021] Preferably, it also includes a path intersection analysis unit, which obtains at least two heat dissipation path topology structures through the heat dissipation topology modeling unit and extracts the node number set of each path;

[0022] The path intersection analysis unit calculates the intersection of different path node sets, counts the frequency of occurrence of nodes in each path, and filters valid intersection paths based on the deviation between the frequency distribution and the preset frequency threshold, generating a heat dissipation path intersection set.

[0023] Preferably, it also includes a thermal label classification unit, which obtains the endpoint nodes in the heat dissipation path intersection set through the path intersection analysis unit and collects the thermal distribution labels corresponding to each endpoint node;

[0024] The thermal label classification unit counts the frequency of each thermal distribution label in the heat dissipation path, sorts them by frequency, matches the most frequent label for each heat dissipation path, and generates a heat dissipation path thermal label group.

[0025] Preferably, it also includes a classification topology generation unit, which obtains heat dissipation path heat dissipation label groups through the heat dissipation label classification unit and merges heat dissipation paths corresponding to the same heat dissipation topology node.

[0026] The classification topology generation unit establishes a mapping relationship between heat dissipation topology nodes and merging paths to generate a server cluster heat dissipation topology classification structure.

[0027] Preferably, the heat dissipation efficiency analysis unit obtains the heat dissipation topology classification structure through the classification topology generation unit and extracts the energy consumption parameters of the high-load heat dissipation topology nodes;

[0028] The heat dissipation efficiency analysis unit dynamically updates the calculation benchmark values ​​of heat dissipation regulation lag ratio and heat dissipation regulation efficiency based on the parameter offset and regulation time of the high-load heat dissipation topology node.

[0029] Preferably, the system stability monitoring unit obtains the total number of heat dissipation paths of the server cluster through the heat dissipation topology modeling unit, and verifies the effectiveness of the system failure signal by combining the ratio of the runaway frequency to the total number of paths.

[0030] Compared with the prior art, the beneficial effects of the present invention are:

[0031] This intelligent energy consumption monitoring system for server cluster heat sinks effectively improves upon the limitations of traditional monitoring methods through the collaborative work of its various units. The energy consumption acquisition unit employs a periodic acquisition method, providing continuous basic data support for subsequent anomaly detection. However, its acquisition cycle is not fixed and can be dynamically adjusted under the control of the regulation and correction unit. This allows for increased acquisition frequency to capture subtle changes when heat dissipation fluctuates significantly, and decreased frequency to reduce resource consumption when the condition is stable, achieving a balance between acquisition efficiency and resource consumption.

[0032] The threshold judgment unit is no longer limited to simple threshold comparison, but generates different types of abnormal signals such as temperature, flow rate, or pressure based on the specific values ​​of energy consumption parameters, making the identification of abnormal types more accurate. This classification and identification method allows maintenance personnel to quickly locate the source of abnormalities, avoiding blind troubleshooting based on traditional single abnormality prompts and reducing the time cost of problem diagnosis.

[0033] The heat dissipation efficiency analysis unit quantifies the degree of anomaly by calculating the deviation between energy consumption parameters and preset thresholds. The generated heat dissipation regulation lag ratio directly reflects the response delay of the heat dissipation system to anomalies. Simultaneously, the statistics on the time consumption of the heat dissipation regulation process provide crucial data for evaluating heat dissipation regulation efficiency, allowing maintenance personnel to clearly understand the dynamic response capability of the heat dissipation system and providing direction for optimizing subsequent regulation strategies.

[0034] The regulation and correction unit makes bidirectional adjustments based on the heat dissipation regulation lag ratio and heat dissipation regulation efficiency. On the one hand, by adjusting the acquisition cycle of the energy consumption acquisition unit, it ensures that abnormal signals can be captured in a timely manner, avoiding the omission of abnormalities due to acquisition lag. On the other hand, by adjusting the heat dissipation regulation intensity, it ensures that the regulation measures can match the actual heat dissipation demand, avoiding the expansion of abnormalities due to insufficient regulation, while also preventing energy waste caused by excessive regulation.

[0035] The system stability monitoring unit calculates the runaway frequency by statistically analyzing the total number of abnormal signals and combining this with runtime, enabling a comprehensive understanding of the long-term operating status of the cooling system. When the runaway frequency exceeds a preset threshold, a system failure signal is generated, providing early warning of potential system risks. This allows maintenance personnel to take preventative measures before a substantial system failure occurs, reducing business interruptions caused by sudden faults and improving the overall continuity and reliability of the server cluster operation. The organic integration of these units gives the entire monitoring system greater adaptability, accuracy, and preventative capabilities, enabling it to better cope with the complex and ever-changing operating environment of server cluster heat sinks. Attached Figure Description

[0036] Figure 1 This is a timing diagram of the intelligent monitoring system for server cluster heat sink energy consumption described in this invention;

[0037] Figure 2 A flowchart for heat dissipation efficiency analysis;

[0038] Figure 3 The flowchart for adjustment and correction;

[0039] Figure 4 This is a flowchart for path intersection analysis. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Please see Figure 1 This invention provides an intelligent monitoring system for the energy consumption of a server cluster heat sink, the system comprising:

[0042] The energy consumption acquisition unit periodically collects energy consumption parameters such as temperature, flow rate, and pressure from the server cluster's heat sink. The collected results are transmitted to the threshold judgment unit for comparison within a threshold range. When a parameter exceeds a preset threshold range, a corresponding temperature anomaly signal, flow rate anomaly signal, or pressure anomaly signal is generated. The heat dissipation efficiency analysis unit receives the anomaly signals and related energy consumption parameters, calculates the deviation between the parameters and the threshold, generates the heat dissipation regulation lag ratio, and calculates the regulation time to generate the heat dissipation regulation efficiency. The regulation correction unit dynamically adjusts the acquisition cycle based on the lag ratio and optimizes the regulation intensity based on the regulation efficiency. The system stability monitoring unit counts the total number of anomaly signals and their duration, calculates the runaway frequency, and generates a system failure signal when the threshold is exceeded, achieving real-time monitoring of the system status.

[0043] Example 1: See Figure 2 The threshold judgment unit receives heat dissipation temperature, coolant flow rate, and pipeline pressure data transmitted from the energy consumption acquisition unit, and compares them in real time with preset temperature threshold ranges, flow rate threshold ranges, and pressure threshold ranges. The temperature threshold range is set to 25-35℃, the flow rate threshold range to 5-8L / min, and the pressure threshold range to 0.3-0.5MPa. When the detected heat dissipation temperature reaches 36℃, a temperature anomaly signal is generated; when the coolant flow rate drops to 4.8L / min, a flow rate anomaly signal is triggered; and when the pipeline pressure exceeds 0.52MPa, a pressure anomaly signal is generated. The heat dissipation efficiency analysis unit extracts the specific parameter values ​​corresponding to the anomaly signals, such as the 36℃ value at the time of the temperature anomaly, and calculates the difference between it and the threshold endpoint of 35℃ to obtain a parameter offset of 1℃. The ratio of this offset to the preset baseline value of 2℃ generates a 50% heat dissipation regulation hysteresis ratio. At the same time, the time taken from the triggering of the temperature anomaly to the recovery to 34℃ is recorded as 180 seconds, and compared with the standard duration of 150 seconds to obtain a heat dissipation regulation efficiency coefficient of 1.2. For abnormal flow rates, the offset of 0.2 L / min between 4.8 L / min and the lower limit threshold of 5 L / min is calculated. Based on the baseline value of 0.5 L / min, a 40% hysteresis ratio is generated, and the control time is recorded to generate the corresponding efficiency coefficient. The specific implementation method is as follows:

[0044] The energy consumption acquisition unit monitors the server cluster's cooling system in real time via a distributed sensor network. Sensor nodes are deployed at key locations such as radiator inlets and outlets, critical nodes in cooling pipes, and the surface of heat exchangers. Each sensor node is equipped with a temperature probe, flow meter, and pressure sensor, enabling simultaneous acquisition of multiple parameters. The acquisition module employs a polling mechanism, reading data from all monitoring points according to a preset cycle, with the initial acquisition frequency set to 60 seconds per cycle. The acquired raw data undergoes filtering and amplification processing by a signal conditioning circuit before being transmitted to the central processing unit via the CAN bus.

[0045] The threshold judgment unit incorporates a multi-stage comparator array to process three parameters: temperature, flow rate, and pressure. The temperature comparator receives temperature data from the CPU heatsink and GPU cooling module, setting a dual-threshold comparison window. When the temperature of a heatsink consistently exceeds the upper threshold, the comparator outputs a high level to trigger a temperature anomaly signal generation circuit. The flow comparator employs a differential input design, comparing the flow rate difference between the main coolant pipe and branch pipes in real time, and making a judgment based on a preset flow fluctuation range. The pressure comparator is equipped with dual overpressure and underpressure detection channels, immediately locking in an abnormal state when the pipe pressure exceeds the safe range.

[0046] The heat dissipation efficiency analysis unit includes a parameter offset calculation module and a control time statistics module. The parameter offset calculation module receives abnormal parameter values ​​transmitted from the threshold judgment unit, first identifies the abnormality type, and retrieves the corresponding threshold baseline data. For abnormal temperature conditions, the calculation module automatically selects the closest threshold boundary value for difference calculation to obtain the absolute offset. This offset is sent to a proportional converter, proportionally converted with a preset standard offset, and outputs a dimensionless hysteresis ratio value. The control time statistics module starts timing from the moment the abnormal signal is triggered, continuously monitors the energy consumption parameter change curve, accurately captures the critical point where the parameters return to the normal range, and records the complete control time span.

[0047] During temperature anomaly handling, when the temperature of a server node's heatsink reaches an abnormal threshold, a sensor data packet is timestamped and transmitted to the central processing unit. The threshold judgment unit parses the data packet content, extracts the temperature measurement value, and compares it with the set safety range. If a sustained exceedance is detected, a digital signal containing the abnormal node ID, anomaly type, and the exceeded value is immediately generated. This signal is transmitted via the data bus to the event processing queue of the heat dissipation efficiency analysis unit.

[0048] The heat dissipation efficiency analysis unit initiates anomaly handling procedures, extracting key parameters from the signal. Taking temperature anomalies as an example, the analysis unit first queries the historical operating data of the node to obtain the temperature change trend curve. Based on the difference between the current temperature value and the upper limit of the threshold, the instantaneous offset is calculated. Simultaneously, a control timer is started to monitor changes in subsequent temperature data points. When the temperature value falls back to a safe range, the timer stops and the total time is recorded. The system automatically compares this time with the standard processing time for similar anomalies to generate an efficiency evaluation coefficient.

[0049] The flow anomaly handling employs a similar mechanism, but adds a flow fluctuation rate analysis function. When coolant flow becomes abnormal, the system not only detects the instantaneous flow value but also calculates the flow change gradient per unit time. This helps distinguish between sudden anomalies and gradual failures. Pressure anomaly handling focuses on the rate of pressure change, calculating potential risk factors in conjunction with pipe volume parameters.

[0050] The calculation of parameter offsets employs a dynamic benchmark adjustment strategy. The system maintains a benchmark database, recording the typical offset ranges of various parameters under different operating conditions. When calculating the current offset ratio, it automatically references the statistical average over a recent period to make the calculation results more consistent with actual operating conditions. For newly deployed nodes or replaced components, the system adopts a progressive learning mechanism, using conservative benchmark values ​​in the initial stage and gradually optimizing the calculation model as operating data accumulates.

[0051] The control time statistics module employs a multi-threaded design, enabling simultaneous tracking of the handling process of multiple abnormal events. Each abnormal event is assigned an independent timing channel, recording the entire process from signal generation to parameter recovery. The system analyzes the time consumption composition, distinguishing between proactive control time and the system's natural response time, providing more refined data support for efficiency evaluation. Statistical results are stored categorized by abnormality type, forming a historical database for trend analysis.

[0052] The abnormal signal processing flow includes multiple verification steps. After the threshold judgment unit generates an abnormal signal, it requires confirmation through three consecutive acquisition cycles before being officially recorded. This effectively avoids false alarms caused by momentary interference. Confirmed abnormal events are assigned a unique event number, including a timestamp, node location, abnormality type, and severity level. Complete event records are written to non-volatile memory, and the real-time status database is updated simultaneously.

[0053] The output interface of the heat dissipation efficiency analysis unit provides a standardized data format, including hysteresis ratio values, regulation efficiency coefficients, and event signature codes. This data is transmitted to the control and correction unit via a high-speed data channel and is also collected by the system stability monitoring unit for long-term trend analysis. All data processing procedures adhere to strict timing requirements to ensure the synchronization of real-time monitoring data with subsequent control commands.

[0054] The system employs a tiered alarm mechanism, automatically determining the anomaly level based on parameter offset and control time. Level 1 anomalies trigger the regular control process, Level 2 anomalies activate an enhanced control mode, and Level 3 anomalies trigger a system-level emergency response. Each level corresponds to a different sampling frequency adjustment strategy and control intensity parameters, forming a tiered response plan. This design ensures both system response speed and avoids resource waste caused by over-control.

[0055] The data acquisition process employs adaptive sampling technology, maintaining a basic sampling frequency under normal operating conditions and automatically increasing the sampling density when increased parameter fluctuations are detected. Each sensor node has local caching capabilities, allowing temporary storage of critical data during communication interruptions. The transmission protocol uses differential coding and CRC checksums to ensure data integrity and real-time performance. The central processing unit is configured with a dual-channel receiving mechanism: the main channel processes regular data, while the emergency channel is dedicated to prioritizing the transmission of abnormal event data.

[0056] The threshold management module supports dynamic adjustment, automatically optimizing the threshold range based on external factors such as ambient temperature and server load. Different temperature threshold curves are used for summer and winter operating conditions, and differentiated pressure warning lines are set for high and low load periods. This dynamic threshold strategy effectively reduces false alarm rates caused by seasonal factors and changes in operating modes. Threshold parameters are stored in programmable memory, supporting remote configuration and batch updates.

[0057] The abnormal signal generation circuit employs opto-isolation to prevent power interference from affecting signal quality. The digital signal output port is equipped with an impedance matching network to ensure signal integrity over long distances. All comparator circuits undergo periodic self-calibration, correcting for zero drift and gain errors using a built-in reference source. The analog signal processing channel is equipped with an anti-aliasing filter to eliminate the impact of high-frequency noise on measurement accuracy.

[0058] The heat dissipation efficiency evaluation algorithm employs a sliding window statistical method, analyzing the processing data of several recent abnormal events to calculate a moving average efficiency value. The evaluation results are standardized and converted into a percentage score for system decision-making reference. Algorithm parameters can be adjusted online based on actual operating performance to maintain the objectivity and applicability of the evaluation results. Historical efficiency data is stored in time series, supporting trend queries and comparative analysis for any time period.

[0059] Example 2: See Figure 3 The control and correction unit receives real-time data streams from the heat dissipation efficiency analysis unit. These data packets contain the heat dissipation control lag ratio and the heat dissipation efficiency coefficient. The system has an internal lag ratio evaluation window that continuously monitors the trend of this ratio. When the lag ratio exceeds the allowable upper limit three times consecutively, the instruction generation module initiates the cycle correction program. This program uses a proportional-differential algorithm to calculate a new acquisition cycle value, dynamically compressing the original 60-second reference cycle according to the current lag ratio. The correction instruction is transmitted to the control register of the energy consumption acquisition unit via a high-speed serial bus. The port mapped to address 0x3F8 receives the 32-bit control word. The high 16 bits of the control word store the correction coefficient, and the low 16 bits define the execution delay parameter. The clock management circuit of the energy consumption acquisition unit adjusts the sampling timer accordingly, shortening the cycle to the target value. The execution result is returned through the status register, completing the closed-loop control.

[0060] For heat dissipation regulation efficiency data, the control and correction unit employs a two-level judgment mechanism. The primary judgment compares the current efficiency coefficient with a preset standard value, while the secondary judgment analyzes the rate of change of this coefficient over time. When the coefficient value is below the standard threshold and the rate of change shows a negative growth trend, the enhanced instruction generation sequence is activated. The instruction code includes a control intensity level identifier; the current version defines four intensity levels from 0 to 3. Instruction transmission uses a reliable transmission protocol with a retransmission mechanism to ensure accurate reception by the threshold judgment unit. The instruction parser of the threshold judgment unit updates the heat dissipation control parameter mapping table based on the level identifier. The table records the fan speed curve coefficient, cooling pump duty cycle parameters, and refrigerant flow valve opening reference value. Taking a two-level to three-level upgrade as an example, the fan speed reference increases by 15%, the cooling pump duty cycle is adjusted from 60% to 75%, and the flow valve opening increases by 8%. The actuator receives the updated parameter set and outputs the corresponding control waveform through a pulse width modulation signal driver.

[0061] Meanwhile, the heat dissipation topology modeling unit initiates the node discovery program. This program polls the connection status registers of each heat dissipation node in the server cluster to obtain the physical connection topology between nodes. The main controller chip traverses the address space of the heat dissipation nodes via the I2C bus, reading the level identifier of each node. The level determination is based on the coolant flow direction: nodes directly connected to the main heatsink are defined as level one, nodes connected via a level one splitter are defined as level two, and the end heatsink belongs to level three. The node relationship matrix construction module converts the obtained connection information into an adjacency matrix for storage. The matrix row numbers represent node numbers, and the column values ​​record the connection direction identifier. The direction encoding uses an 8-bit binary scheme: East 00000001, West 00000010, South 00000100, North 00001000.

[0062] The node sorting algorithm reconstructs the adjacency matrix according to hierarchical priority. The system initializes the sorting index array, placing all first-level nodes at the beginning of the sequence, second-level nodes in the middle, and third-level nodes at the end. Nodes at the same level are arranged in ascending order of node number. After sorting, a node sequence array is generated, with each element containing the node ID, level value, and position index. Based on this sequence, the topology building engine extracts the distance data between adjacent nodes. The distance parameter comes from the physical coordinate values ​​in the node configuration register; the system automatically calculates the Euclidean distance and converts it into a standardized span value. The topology descriptor generation module encapsulates the node sequence, direction encoding, and span parameters into structured data packets. Each data packet contains: a node path identifier (4 bytes), the number of path nodes (2 bytes), a node ID sequence (variable-length array), a connection direction vector set (stored bitwise), and a span parameter array (floating-point sequence). Finally, a heat dissipation path topology database is constructed and stored in non-volatile memory in binary tree format.

[0063] The system periodically executes a topology verification process. The verification program reconstructs a connection graph based on the node sequence and compares it with the actual connection status returned by the physical probes. When a mismatch between the path topology and physical links is detected, a topology database reconstruction process is initiated. The reconstruction process employs an iterative correction algorithm: first, abnormal path data is frozen; then, node connection status is re-collected through a dedicated diagnostic interface; and finally, the confirmed connection information is incrementally updated to the topology database. A version control mechanism ensures that other modules can access historical topology data normally during the reconstruction period.

[0064] The execution status monitoring of control commands employs a dual-channel feedback mechanism. The main channel receives real-time status codes from the actuators, while the secondary channel monitors the actual heat dissipation parameter change curves. The main channel feedback information includes fan drive current, cooling pump operating frequency, and valve position sensor readings. These data are converted into execution percentages and dynamically compared with the command requirements. When the deviation consistently exceeds the tolerance value, the control correction unit triggers a compensation command sequence, automatically increasing the output strength of the control signal. The secondary channel uses temperature and flow rate change gradient data acquired by the energy consumption acquisition unit to verify the actual effectiveness of the control operation. The dual-channel data are aligned and analyzed on the time axis to generate a control quality assessment report, which is used to optimize subsequent command generation parameters.

[0065] The topology data access interface uses a paginated query method. Clients can query complete path information by specifying a path identifier code, or query related paths in reverse order by node number. Query requests are decomposed by the data routing processor and converted into multi-level index retrieval operations. The first-level index locates the path storage page, the second-level index extracts the path node list, and the third-level index obtains the connection direction and span parameters. Query results are returned in a structured data block containing node sequences and auxiliary parameters, prioritizing path completeness.

[0066] The heat dissipation node discovery program has dynamic expansion capabilities. When a new heat dissipation node is added to the server cluster, the node address is automatically added to the polling list. The system assigns a temporary hierarchical identifier to the new node and automatically corrects the hierarchical affiliation during the topology construction phase. The node location index update uses an incremental sorting algorithm, which only reorganizes the index value after the new node is inserted, improving system response efficiency.

[0067] The coolant flow direction detection mechanism employs a thermodynamic marking method. A tracer is injected at the main radiator inlet, and sensors at each node monitor the arrival time of the tracer. The system calculates the flow direction based on the propagation time difference to determine the node connection direction. This method is executed periodically to correct for possible changes in physical connections. Version control is introduced into the direction coding update process, and a complete change record of historical direction data is retained for analysis.

[0068] Ambient temperature compensation is incorporated into the calculation of the span parameter. After the distance sensor collects raw data, it performs linear correction according to a preset formula based on the current ambient temperature. The compensation coefficients are stored in a temperature-distance mapping table, which is generated through experimental calibration and includes compensation parameters ranging from -10℃ to 50℃. The calculation results are rounded to two decimal places before being stored in the topology database.

[0069] A handshake protocol is established for the transmission of control commands. Upon receiving a new command, the threshold judgment unit returns an acknowledgment code; if the adjustment and correction unit does not receive acknowledgment within a timeout period, it initiates a retransmission process. Three consecutive transmission failures trigger a fault alarm, notifying maintenance personnel to check the communication link. Command execution results utilize a feedback receipt mechanism; after completing parameter changes, the actuator sends a status report containing the actual output control signal characteristic values. This data is recorded in the execution log library as the primary basis for system performance evaluation.

[0070] The heat dissipation parameter adjustment operation is set with safety boundary limits. When increasing the intensity of regulation, the system monitors the pressure limit of the cooling system in real time. When the regulation command causes the pressure to approach the safety threshold, an adjustment buffer phase is automatically inserted. This phase adopts a gradual enhancement strategy: the initial adjustment is 60% of the target value, 80% is executed in the next control cycle, and finally the full adjustment is completed. This segmented execution method avoids oscillations caused by sudden changes in system load.

[0071] The topology visualization service module runs on a dedicated processor. This module reads structured data from the topology database and converts it into a graphical description language. Nodes are arranged hierarchically on different horizontal planes, with connecting arrows indicating flow direction, and line segment width positively correlated with span value. The visualization engine supports 3D view rotation, allowing users to interactively examine the details of each heat dissipation path. The topology map is updated synchronously with the modeling units to ensure real-time information presentation.

[0072] Example 3: See Figure 4 The path cross-analysis unit receives path topology data packets generated by the heat dissipation topology modeling unit via a high-speed data channel. Each data packet contains complete node sequence information. The data parser first extracts the path identifier and node count fields, converting the variable-length array format node ID sequence into a standardized set data structure. The system maintains a global node mapping table, recording the index of each node ID's occurrence position in all paths. The cross-detection algorithm employs multi-level hash comparison technology, first filtering potential cross nodes and then verifying actual cross relationships in the second round. For path A (node ​​sequence N1→N3→N5→N7) and path B (node ​​sequence N2→N3→N6→N7), the algorithm detects that nodes 3 and 7 exist in both paths simultaneously and marks them as candidate cross points.

[0073] The frequency statistics algorithm uses a sliding window counting method to analyze the distribution of each candidate node across all paths. The system establishes a node frequency matrix F, where rows represent node IDs and columns record the frequency of occurrence in different paths. The cross-validation module calculates the co-occurrence probability of nodes across multiple paths and filters out significant cross-nodes. Node 3 appears in 5 paths, and node 7 appears in 3 paths; these data are recorded in the cross-feature database. The frequency threshold comparator compares the actual frequencies with a dynamically adjusted benchmark value, which is automatically adjusted according to the system's operational phase: set to 2 times during initialization, increasing to 4 times during stable operation, and decreasing to 3 times during high-load phases. Node 3's 5 occurrences exceed the current phase threshold and are therefore confirmed as a valid cross-node.

[0074] The path cross-set generator uses verified cross nodes as the core of the association and reverse-searches for all paths containing those nodes. The search process employs an inverted index technique, quickly locating relevant path identifiers via node IDs. For node 3, the system finds shared relationships between paths A, B, and C, forming a preliminary cross-set. The set integrity validator checks the associations of these paths outside the cross nodes, confirming that node 7 also exists in the three paths, enhancing the credibility of the cross-set. The final generated cross-set data structure includes: a list of cross node IDs, an array of associated path identifiers, and a cross-strength coefficient. The cross-strength coefficient is calculated using the following formula:

[0075]

[0076] in: Indicates the cross strength coefficient. The number of cross nodes, It is the weight factor of the i-th intersection node (determined according to the node level). This represents the frequency of the i-th node in the path. This represents the maximum frequency of occurrence among all nodes. This coefficient is used to quantify the density of path intersections, and its value ranges from 0 to 1.

[0077] The thermal label classification unit acquires cross-aggregate data through a dedicated interface and initiates the endpoint node analysis process. The endpoint node locator extracts the last node ID from the node sequence of each path and queries the corresponding location coordinates in the physical layout database of the server cluster. The infrared thermal imaging data acquisition module obtains a real-time thermal distribution map from the distributed temperature sensor network based on the coordinate information. The image processing algorithm converts the thermal data into standardized temperature range labels.

[0078] The visualization rendering engine for cross paths transforms the analysis results into graphical elements. In the 3D topology map, cross nodes are displayed as red cubes, and associated paths are represented by curves of different colors. Heat maps are presented through color rings at the ends of the paths: red rings represent high-temperature zones, and blue rings represent low-temperature zones. The visualization system supports interactive queries; clicking on any cross node displays detailed frequency statistics and associated path information.

[0079] The system implements a dynamic label update mechanism. When thermal imaging data changes significantly, the label reassessment trigger is automatically activated. A change detector continuously monitors temperature sensor readings; when the temperature change in a certain area exceeds 15% of the historical baseline value, a local label update process is triggered. The reassessment process prioritizes high-weight nodes and uses incremental updates to correct the label database. The version control system retains records of all label changes, supporting the retrospective analysis of thermal distribution at any given time.

[0080] The path cross-validation unit implements a multi-factor verification mechanism. The raw data validator checks the integrity and continuity of the node sequence to ensure there are no missing or duplicate node IDs. The logic validator confirms that the path direction conforms to the physical constraints of the cooling system, such as prohibiting loop paths. The frequency statistics validator compares real-time calculation results with historical trend data to detect abnormal fluctuations. When data anomalies are detected, the system automatically isolates the problematic data segment and reprocesses it using redundant calculation nodes.

[0081] The heat map tag database adopts a distributed architecture, with the primary replica stored on a central server and multiple read-only replicas deployed on edge computing nodes. The data synchronization mechanism uses incremental replication based on operation logs to ensure eventual consistency across replicas. The query optimizer automatically establishes a cache area for hot data based on request patterns, while frequently accessed tag data is retained in memory to accelerate query response.

[0082] The cross-node importance assessment module periodically performs a global analysis. The assessment algorithm comprehensively considers node cross-node frequency, thermal tag weights, and topological location factors to generate a node criticality score. The score results are used to optimize system resource allocation, with highly critical nodes configured with denser monitoring sampling frequencies and stricter anomaly detection thresholds. Node score data is also provided to the maintenance system to guide the optimization and modification of physical cooling facilities.

[0083] The system implements load balancing monitoring for cross-paths. The flow distribution analyzer calculates the coolant distribution ratio for each path using sensor data from the cross-nodes. When the load on a cross-node exceeds that of its adjacent nodes by 20%, a flow adjustment suggestion is generated. The adjustment algorithm considers path length, node level, and thermal tag factors to calculate the optimal flow redistribution scheme, and the output is the valve opening adjustment parameter.

[0084] The spatiotemporal correlation analysis module of thermal imaging data identifies anomalous thermal patterns. By comparing temperature distribution maps from different time periods, it detects the movement trajectory and diffusion trend of hotspot areas. The anomaly pattern identifier compares current thermal changes with a database of historical typical patterns to provide early warnings of potential heat dissipation bottlenecks. The warning information includes a list of affected paths and an estimated timeline of anomaly development.

[0085] The tag propagation model is used to predict changes in thermal distribution. Based on the current tag distribution and coolant flow direction, the algorithm simulates the heat transfer process in the path network. Model inputs include node heat capacity parameters, coolant flow rate, and ambient temperature; the output is a predicted tag distribution map for future time points. The prediction results are expressed in probabilistic form, assisting maintenance personnel in proactively deploying control measures.

[0086] The system implements multi-granularity cross-analysis capabilities. At the macro level, it analyzes the cross-analysis characteristics of the entire server cluster, identifying system-level thermal distribution patterns. At the meso level, it focuses on cross-analysis patterns in specific heat dissipation areas, assessing local heat dissipation efficiency. At the micro level, it delves into the detailed parameters of individual cross-analysis nodes, diagnosing potential flow resistance or heat exchange problems. The analysis results at each level are integrated and displayed through a unified interface, forming a complete health status report of the cooling system.

[0087] Example 4: The classification topology generation unit receives the heat dissipation path heat label groups processed by the heat label classification unit via the data bus. The data structure contains the mapping relationship between path identifiers and corresponding heat labels. The system initializes the label index table and merges heat dissipation paths with the same heat distribution labels into the same topology node. Taking a heat dissipation system actually deployed in a data center as an example, the system identifies three main types of heat labels: the "high temperature zone" label is associated with 5 paths, the "medium temperature zone" label is associated with 4 paths, and the "low temperature zone" label is associated with 3 paths. The merging process adopts a two-level clustering algorithm: the first round is coarse grouping according to label type, and the second round is fine adjustment based on path similarity.

[0088] Table 1: Topology node structure after heat dissipation path merging.

[0089]

[0090] The topology node relationship builder creates an independent data structure for each merge group, recording the complete node sequence of all paths within the group. The structure includes a path topology fingerprint field, storing the hash value of the path's direction for fast comparison. The mapping relationship generator establishes a bidirectional index between topology nodes and merged paths. The forward index queries for contained paths using node IDs, while the reverse index locates the topology node to which the path belongs using its ID. The data structure uses a B+ tree organization, supporting efficient range queries and batch updates.

[0091] The generation of the server cluster thermal topology classification structure employs an iterative optimization method. The initial version is created based on static label assignment, and the system continuously receives real-time thermal data updates during operation. When a temperature change at the end node of a path causes a change in label type, the dynamic adjuster automatically removes the path from the original topology node and inserts it into the node corresponding to the new label. Migration operations maintain atomicity, ensuring that the system provides a consistent topology view at all times. The version control system records detailed information for each structural adjustment, including the change time, affected paths, and operation type.

[0092] The identification mechanism for high-load heat dissipation topology nodes is based on multi-dimensional evaluation. The node load calculator comprehensively considers three core indicators: duration of sustained temperature exceeding limits, flow rate fluctuation amplitude, and frequency of abnormal pressure. The evaluation process uses a sliding time window, assigning higher weight to indicator data within the most recent 15 minutes. The system calculates a dynamic load coefficient for each topology node, marking it as a high-load state when the coefficient exceeds the category baseline. Taking the TN-2024-01 node in Table 1 as an example, its maximum flow rate deviation of 22.4% triggers the high-load judgment condition.

[0093] After the heat dissipation efficiency analysis unit is connected to the classification topology, a dedicated monitoring channel is established to collect the operating parameters of high-load nodes. The data acquisition strategy distinguishes between two modes: normal sampling and anomaly tracking. In normal mode, basic temperature and flow data are collected at standard cycles; in anomaly tracking mode, the acquisition frequency is increased threefold and pressure gradient monitoring is added when parameter deviations are detected. The parameter processor performs moving average filtering on the raw data to eliminate the influence of transient interference on the analysis results.

[0094] The dynamic baseline update algorithm employs a gradual adjustment strategy. The system maintains two sets of baseline parameters: a long-term baseline calculated based on historical statistical values, reflecting the system's normal operating characteristics; and a short-term baseline derived from the average of the most recent 30 minutes, capturing changes in current operating conditions. When the parameter offset of a high-load node continuously deviates from the baseline, the regulator gradually corrects the baseline value in fixed steps, with each adjustment not exceeding 5% of the original value. This gentle adjustment method avoids drastic fluctuations in system parameters and maintains control stability.

[0095] The timing and statistics module enhances time measurement accuracy by employing a high-resolution clock to record the entire process of abnormal events. A timer is activated when an abnormal signal is generated by the threshold judgment unit, and a callback function mechanism listens for parameter recovery events. Time data is stored with nanosecond-level precision and can be converted to appropriate time units for subsequent analysis based on actual needs. Statistical results are archived by topological node category, and the timing data of multiple abnormal events within the same category form a time-series dataset.

[0096] The calculation of the lag ratio for heat dissipation regulation incorporates a topology weighting factor. The parameter offset calculation results for nodes in the high-temperature zone are multiplied by a correction factor of 1.2, while the values ​​for nodes in the mid-temperature zone remain unchanged, and nodes in the low-temperature zone are reduced by a factor of 0.8. This differentiated processing reflects the heat dissipation sensitivity of different thermal regions, making the generated lag ratio more consistent with actual regulation needs. The ratio value is limited before output to ensure that the result is within the effective range of 0-100%.

[0097] The system provides a topology visualization service, converting classification results into an interactive graphical interface. The main view displays the distribution of three types of topology nodes, using red, yellow, and blue colors to distinguish heat map label types. Clicking on any topology node expands a detailed information panel, displaying path animations and real-time parameter curves. The view supports timeline dragging, allowing users to trace back to any historical point in time regarding topology status and parameter changes.

[0098] The topology optimization suggestion engine periodically analyzes the classification results to identify potential heat dissipation bottlenecks. The analysis algorithm detects three types of abnormal patterns: cross-category path intersections, traffic imbalances within the same category, and blurred label boundaries. For detected problem patterns, the system generates a structured suggestion report, including a problem description, impact assessment, and adjustment plans. For example, when a node in a high-temperature zone is found to contain too many long paths, it is suggested to add intermediate heat dissipation points and reorganize the topology.

[0099] The reliability verification mechanism for thermal tags continuously monitors the match between the tag and the actual temperature. The verifier compares the current temperature of the end node with the range defined by the tag and calculates the tag confidence score. When the score is lower than the verification threshold, a tag reassessment process is triggered. The new tag allocation process incorporates historical path data analysis, referencing the thermal change patterns of the path over the past 24 hours to improve tag accuracy.

[0100] The system implements persistent storage and rapid recovery of the topology structure. The complete classification structure is periodically dumped into a compressed binary format, containing all path data, label mappings, and node relationships. A checksum is appended to the stored files to ensure data integrity, and corrupted files can be automatically recovered from a mirror copy. Upon system restart, the most recent valid topology snapshot is loaded first, and updates to the latest state are gradually performed in a background thread.

[0101] The operations and maintenance interface provides manual adjustment functionality for the topology. Authorized users can temporarily modify automatically generated classification results through the management console, such as forcing a specific path to be assigned to a particular topology node. Detailed operation logs are recorded for manual operations, including the state before modification, the content of the adjustment, and the operator's information. The system will refer to these manual adjustment records during subsequent automatic classification processes to gradually learn the classification preferences of operations and maintenance personnel.

[0102] The topology impact assessment module predicts and categorizes the impact of adjustments on the overall system. Before any structural changes are made, the simulator calculates the potential cascading effects, including changes in load on adjacent nodes and fluctuations in pressure at intersections. The assessment results are presented as a risk matrix to help operations personnel weigh the pros and cons of the adjustment plan. High-risk operations require secondary confirmation before execution to prevent accidental system instability.

[0103] Example 5: The system stability monitoring unit acquires abnormal signal records transmitted by the heat dissipation efficiency analysis unit via the data bus. These records include fields such as generation time, anomaly type, and recovery time. An internal circular buffer stores the most recent abnormal events, designed to store 72 hours of event data. The monitoring algorithm cyclically scans the buffer, counting the total number of abnormal events within the current observation window. The observation window duration is dynamically adjusted, defaulting to 24 consecutive hours, shortening to 8 hours when the average server cluster load exceeds 80% of the design capacity, and extending to 48 hours during low-load operation.

[0104] The runaway frequency calculation module employs a time-weighted statistical method. The module initializes a counter and a time accumulator. When a new abnormal signal is detected, the counter increments, and the time accumulator adds the duration of the event's impact. The clock management circuit triggers a calculation cycle every 300 seconds. The ratio of the number of abnormal events to the service runtime within the statistical cycle generates the raw frequency value. The raw data is smoothed using a second-order Butterworth filter to obtain the final runaway frequency output value. The calculation process considers clock drift correction, and all time parameters use UNIX timestamp format to avoid statistical errors caused by date switching.

[0105] The total number of heat dissipation paths is obtained in real time through the update interface of the heat dissipation topology modeling unit. When the topology changes, the modeling unit actively pushes path addition / reduction notifications. The monitoring unit maintains a mirror copy of the current total number of paths, and the copy update uses a dual-write mechanism to ensure data consistency. The path counter has an anti-jitter design; repeated topology reconstructions within a short period of time only trigger the final state update, avoiding intermediate states from interfering with calculation accuracy.

[0106] The failure verification process employs a three-level confirmation logic. When the out-of-control frequency first exceeds a preset threshold, the initial verification phase is initiated: analyzing the distribution characteristics of the ten most recent anomalies to check if they are concentrated in a specific topology category. The percentage of anomalies in high-temperature zone nodes is queried through the topology classification interface; if this percentage exceeds a warning line, secondary verification begins. Secondary verification focuses on the control response time, extracting the average recovery time for similar anomalies from the historical database and comparing it with the current recovery time. Tertiary verification calculates the distribution density of anomalies along the heat dissipation path to identify any chain reactions of localized path failures.

[0107] Path weight calculation takes topology location factors into account. The primary heat dissipation path is assigned a weight coefficient of 1.2, while secondary paths are assigned a weight of 0.8. Weight data is stored in the path attribute table and automatically assigned based on the hierarchical labeling of the topology modeling unit. During weighted calculation, each anomalous event is multiplied by the weight of its corresponding path, and the accumulated result is divided by the total weight to obtain the weighted runaway frequency. This design ensures that anomalies on the main path have a greater impact, preventing occasional failures on secondary paths from interfering with system-level judgment.

[0108] The system failure signal generation process incorporates multiple protections. The signal triggering circuit connects to three parallel comparator outputs, activating only when all three comparators simultaneously detect an out-of-range condition. The physical signal output interface is equipped with opto-isolation devices to prevent false triggering due to electromagnetic interference. The digital signal encoding uses Manchester code format to enhance anti-interference capabilities. The signal transmission channel uses a dedicated line independent of the data bus to ensure the transmission priority of alarm signals.

[0109] A dynamic tuning mechanism for the runaway frequency threshold operates continuously. The initial threshold baseline is set to 0.8 abnormal events per hour, and the system executes a threshold optimization program every 24 hours. The optimization algorithm analyzes the distribution characteristics of historical abnormal data and selects the 24 hours with the most stable operation over the past 30 days as the gold standard value. Considering the seasonal impact of ambient temperature, the operating threshold is increased by 20% in winter and maintained at the baseline value in summer. When signs of hardware aging are detected, the maintenance factor module automatically relaxes the threshold constraint.

[0110] The event attribution analysis module assists in failure diagnosis. This module labels each abnormal event with a possible cause code, including common fault types such as coolant loss, pump power reduction, and heat exchange blockage. It statistically analyzes the distribution ratio of different causes, generating a root cause analysis report for the failure event. The report data is correlated with corresponding failure signals, providing a reference for subsequent maintenance decisions.

[0111] System status recovery monitoring continues to run after a failure signal is issued. After maintenance personnel intervene, the monitor tracks the decline curve of the frequency of abnormal events. Three recovery judgment conditions are set: the failure frequency is below 75% of the threshold for six consecutive hours; the weighted frequency value returns to the normal range and continues for three statistical periods; the last abnormal event has not been reproduced for 24 hours. If any condition is met, the failure signal is turned off and a system recovery report is generated and archived to the log system.

[0112] Timestamp synchronization ensures the precise timing of event records. The system deploys an NTP server synchronized with satellite clocks, and all monitoring units use the same time source. Abnormal events are recorded with millisecond-level accuracy and time zone information is added for collaborative analysis across multinational data centers. A leap second compensation mechanism is added to time-series data storage to prevent event sequence corruption caused by time jumps.

[0113] The audit trail function for failure signals records the operation history in detail. Each time a signal is triggered or deactivated, an audit entry is generated, containing a complete snapshot of environmental parameters: server load rate, ambient temperature and humidity, cooling system operating mode, and other related parameters. Audit logs are encrypted and stored in tamper-proof storage, supporting full lifecycle accountability. The log analyzer automatically detects abnormal operation patterns to prevent unauthorized signal status changes.

[0114] The maintenance collaboration interface connects failure signals to the work order system. When a signal is activated, a fault work order is automatically generated and dispatched to the on-duty engineer. The work order includes the latest anomaly distribution diagram and root cause analysis summary. Engineers can confirm receipt on their mobile devices and close the work order using a verification code after maintenance is completed. The system automatically detects the work order processing time limit; if the work order is not processed within the time limit, the alarm level is automatically escalated.

[0115] A redundant monitoring mechanism is deployed to provide auxiliary monitoring channels. A backup computing core runs independently outside the main monitoring unit, receiving the same input data but using a simplified algorithm. The results from both channels are compared by an arbitrator; if the difference exceeds the tolerance range, a diagnostic mode is activated. The arbitrator prioritizes the main channel data, but seamlessly switches to the backup system when a hardware malfunction is detected in the main channel. The hot standby system maintains real-time data synchronization, ensuring that the switchover process does not affect continuous monitoring functionality.

[0116] The frequency of out-of-control events is visualized using a waterfall chart. The horizontal axis represents the timeline, and the vertical axis represents the frequency values. Different severity levels of out-of-control areas are rendered in different color zones. The view supports ten levels of time zoom, providing a comprehensive overview from minute-level details to monthly trends. Hovering the cursor displays detailed statistical parameters for the current time point, including a detailed list of abnormal events and a snapshot of the relevant topology.

[0117] The on-site simulation test interface is used for system self-testing. During testing, a simulated abnormal signal sequence is injected to check the response accuracy of the entire monitoring link. The test mode is physically isolated from the normal operation mode to prevent test data from contaminating real monitoring records. A compliance report of the test results is automatically generated, including latency data and calculation error rates for each processing stage, used for periodic calibration of monitoring accuracy.

[0118] The system continuously improves its analysis of historical data on failure signals. Machine learning algorithms identify early warning patterns that exceed frequency limits, such as continuous small anomalies increasing over specific periods. The predictive analytics engine provides early warnings of potential stability deterioration trends, shifting maintenance to a preventative phase. The system generates weekly monitoring performance reports, calculating the effective alarm rate and false alarm ratio, driving continuous optimization and updates to monitoring parameters.

[0119] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0120] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A server cluster radiator energy consumption intelligent monitoring system, characterized in that, The system comprises: an energy consumption acquisition unit, which periodically acquires energy consumption parameters of a server cluster radiator and sends the acquisition results to a threshold judgment unit; the threshold judgment unit compares the energy consumption parameters with threshold ranges, and generates a temperature anomaly signal, a flow anomaly signal or a pressure anomaly signal according to the comparison results; a heat dissipation efficiency analysis unit, which obtains the anomaly signal and the corresponding energy consumption parameter through the threshold judgment unit, calculates the deviation of the energy consumption parameter from the preset threshold, generates a heat dissipation regulation lag ratio according to the deviation, and calculates the heat dissipation regulation efficiency by counting the time consumption of the heat dissipation regulation process; a regulation correction unit, which obtains the heat dissipation regulation lag ratio and the heat dissipation regulation efficiency through the heat dissipation efficiency analysis unit, adjusts the acquisition period of the energy consumption acquisition unit according to the heat dissipation regulation lag ratio, and adjusts the heat dissipation regulation intensity according to the heat dissipation regulation efficiency; a system stability monitoring unit, which counts the total amount of anomaly signals output by the heat dissipation efficiency analysis unit, calculates the out-of-control frequency in combination with the running time of the server cluster, and generates a system failure signal when the out-of-control frequency exceeds a preset threshold.

2. The server cluster heat sink energy consumption intelligent monitoring system of claim 1, wherein, The threshold judgment unit compares the heat dissipation temperature, the cooling liquid flow and the pipeline pressure in the energy consumption parameters with the corresponding temperature threshold range, flow threshold range and pressure threshold range respectively, generates a temperature anomaly signal if the heat dissipation temperature exceeds the temperature threshold range, generates a flow anomaly signal if the cooling liquid flow exceeds the flow threshold range, and generates a pressure anomaly signal if the pipeline pressure exceeds the pressure threshold range.

3. The server cluster heat sink energy consumption intelligent monitoring system of claim 2, wherein, The heat dissipation efficiency analysis unit extracts the heat dissipation temperature value corresponding to the temperature anomaly signal, the cooling liquid flow value corresponding to the flow anomaly signal or the pipeline pressure value corresponding to the pressure anomaly signal, and records them as abnormal parameters H, and extracts the endpoint values h0 and h1 of the corresponding threshold range; The heat dissipation efficiency analysis unit calculates the difference between the abnormal parameter H and h0 and the difference between H and h1, selects the minimum difference as the parameter offset, and generates the heat dissipation regulation lag ratio based on the proportional relationship between the parameter offset and the preset reference value; The heat dissipation efficiency analysis unit records the time consumption from the moment when the anomaly signal is triggered to the moment when the energy consumption parameter returns to the threshold range, and generates the heat dissipation regulation efficiency by comparing the time consumption with a preset standard time length.

4. The server cluster heat sink energy consumption intelligent monitoring system of claim 3, wherein, The regulation correction unit generates an acquisition period shortening instruction and sends it to the energy consumption acquisition unit when the heat dissipation regulation lag ratio exceeds the allowed range, and the energy consumption acquisition unit shortens the acquisition period based on the heat dissipation regulation lag ratio; The regulation correction unit generates a regulation intensity enhancement instruction and sends it to the threshold judgment unit when the heat dissipation regulation efficiency is lower than the preset standard, and the threshold judgment unit enhances the power level of the heat dissipation regulation operation according to the regulation intensity enhancement instruction.

5. The server cluster heat sink energy consumption intelligent monitoring system of claim 1, wherein, The system further comprises a heat dissipation topology modeling unit, which obtains the connection relationship of each heat dissipation node in the server cluster, sorts the nodes according to the hierarchical priority of the heat dissipation nodes, and constructs a heat dissipation node sequence based on the sorting result; The heat dissipation topology modeling unit extracts the connection direction and span parameter of adjacent nodes in the heat dissipation node sequence, and generates a heat dissipation path topology structure in combination with the hierarchical difference value.

6. The server cluster heat sink energy consumption intelligent monitoring system of claim 5, wherein, It also includes a path intersection analysis unit, which obtains at least two heat dissipation path topology structures through the heat dissipation topology modeling unit and extracts the node number set of each path; The path intersection analysis unit calculates the intersection of different path node sets, counts the frequency of occurrence of nodes in each path, and filters valid intersection paths based on the deviation between the frequency distribution and the preset frequency threshold, generating a heat dissipation path intersection set.

7. The server cluster heat sink energy consumption intelligent monitoring system of claim 6, wherein, It also includes a thermal label classification unit, which obtains the endpoint nodes in the heat dissipation path intersection set through the path intersection analysis unit and collects the thermal distribution labels corresponding to each endpoint node; The thermal label classification unit counts the frequency of each thermal distribution label in the heat dissipation path, sorts them by frequency, matches the most frequent label for each heat dissipation path, and generates a heat dissipation path thermal label group.

8. The server cluster heat sink energy consumption intelligent monitoring system of claim 7, wherein, It also includes a classification topology generation unit, which obtains heat dissipation path heat dissipation label groups through the heat label classification unit and merges heat dissipation paths corresponding to the same heat distribution label into the same heat dissipation topology node. The classification topology generation unit establishes a mapping relationship between heat dissipation topology nodes and merging paths to generate a server cluster heat dissipation topology classification structure.

9. The server cluster heat sink energy consumption intelligent monitoring system of claim 8, wherein, The heat dissipation efficiency analysis unit obtains the heat dissipation topology classification structure through the classification topology generation unit and extracts the energy consumption parameters of the high-load heat dissipation topology nodes. The heat dissipation efficiency analysis unit dynamically updates the calculation benchmark values ​​of heat dissipation regulation lag ratio and heat dissipation regulation efficiency based on the parameter offset and regulation time of the high-load heat dissipation topology node.

10. The server cluster heat sink energy consumption intelligent monitoring system of claim 1, wherein, The system stability monitoring unit obtains the total number of heat dissipation paths of the server cluster through the heat dissipation topology modeling unit, and verifies the effectiveness of the system failure signal by combining the ratio of the runaway frequency to the total number of paths.

Citation Information

Patent Citations

  • Refrigeration air conditioner operation efficiency detection system based on Internet of Things

    CN117073154A

  • Circulating heat dissipation fault detection system for immersed liquid cooling server

    CN118626319A