A server temperature sensor disaster control method and device

CN122387284BActive Publication Date: 2026-09-08ZHUOXIN (TIANJIN) INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610834709.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-08
Estimated Expiration
2046-06-10

AI Technical Summary

Technical Problem

[0003]有鉴于此,本申请提供一种服务器温度传感器容灾控制方法和装置,用以解决现有技术中温度传感器异常时容灾控制不连续、不稳定的问题

Benefits of technology

[0019]The server temperature sensor disaster recovery control method and apparatus provided in this application, compared with the prior art that only relies on a single abnormal acquisition to trigger disaster recovery and lacks fine-grained differentiation of sensor states, identifies the presence and temperature acquisition states of multiple temperature sensors and matches a unified state identifier based on the corresponding states. This enables differentiated handling of different abnormal situations, avoiding the use of the same control method for sensor absence and acquisition abnormalities. Furthermore, this application matches the disaster recovery control strategy space based on the state identifier, only entering the disaster recovery control space when the target temperature sensor meets the abnormal triggering conditions, effectively reducing the problem of erroneous switching caused by instantaneous communication fluctuations or short-term sampling abnormalities. In addition, this application determines associated temperature sensors through thermal region correlation and uses the temperature data of associated temperature sensors to generate alternative temperature input data. When the target temperature sensor is abnormal, the temperature data of the corresponding thermal region is compensated, enabling the server to maintain continuous operation of the thermal regulation process even during sensor abnormalities. In the disaster recovery control strategy space, the system uses alternative temperature input data to participate in the generation of thermal regulation strategies, so that the server can still maintain thermal control functions such as fan speed adjustment when the target temperature sensor is abnormal, thereby improving the stability of the server temperature control system and its fault tolerance under abnormal conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122387284B_ABST
    Figure CN122387284B_ABST
Patent Text Reader

Abstract

The application provides a kind of server temperature sensor disaster control method and device.The method is by identifying the in-place state and temperature collection state of multiple temperature sensors and generating a unified state identifier, based on state identifier matching disaster control strategy space, when the in-place state and collection state of target temperature sensor meet abnormal trigger condition, enter disaster control strategy space;Determine the associated temperature sensor based on the thermal region association relationship, generate replacement temperature input data using its temperature data and combining the historical temperature trend of target sensor, generate thermal regulation strategy in disaster control strategy space.The application can maintain the continuous operation of server thermal regulation process when temperature sensor is abnormal by unified state identifier, combined with disaster control strategy space and replacement temperature data generation mechanism, reduce the risk of heat dissipation abnormality caused by temperature data loss, improve the robustness, stability and adaptive ability of server temperature control system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server thermal management technology, and in particular to a server temperature sensor disaster recovery control method and device. Background Technology

[0002] As server computing performance continues to improve, the power consumption of CPUs, GPUs, and storage devices is gradually increasing, leading to a rise in internal server heat density. To ensure stable system operation, temperature sensors are typically used to monitor various hot zones in real time, and management units such as the BMC (Browser Control Center) execute thermal management controls such as fan speed adjustment based on the temperature data. In actual operation, temperature sensors may malfunction due to communication abnormalities, hardware failures, electromagnetic interference, or power supply fluctuations, resulting in the inability to continuously output valid temperature data. Current solutions often rely on a single temperature sensor for thermal control or use a fixed backup sensor as a simple substitute. When a sensor malfunctions, problems such as missing temperature feedback and inaccurate temperature estimation can easily occur, leading to abnormal fan speed adjustment, delayed heat dissipation response, and in severe cases, even excessively high temperatures in localized areas. Furthermore, while some redundant control schemes can compensate for sensor malfunctions, they typically use a fixed backup sensor or a simple averaging method for temperature estimation, without considering the actual thermal correlation between different hot zones. Therefore, under complex load scenarios, the temperature compensation results may still deviate significantly from the actual thermal state. At the same time, the existing system lacks a unified status determination and smooth switching mechanism during the switching between normal control mode and abnormal disaster recovery mode, which easily leads to frequent switching of control mode and affects the stable operation of the server thermal management system. Summary of the Invention

[0003] In view of this, this application provides a server temperature sensor disaster recovery control method and apparatus to solve the problem of discontinuous and unstable disaster recovery control when the temperature sensor is abnormal in the prior art.

[0004] Specifically, this application is implemented through the following technical solution:

[0005] The first aspect of this application provides a disaster recovery control method for a server temperature sensor, the method comprising:

[0006] Identify the presence and temperature acquisition status of multiple temperature sensors;

[0007] Based on the matching status identifiers of the in-situ status and temperature acquisition status, the same status identifier can simultaneously identify both the in-situ status and the temperature acquisition status. Temperature sensors with completely identical in-situ status and temperature acquisition status have the same status identifier.

[0008] Based on the state identifier matching disaster recovery control strategy space, when the in-situ state and acquisition state of the target temperature sensor in the temperature sensor meet the corresponding abnormal triggering conditions, it enters the disaster recovery control strategy space of the target temperature sensor.

[0009] At least one associated temperature sensor is determined based on the thermal region correlation of the target temperature sensor;

[0010] Generate alternative temperature input data based on the temperature data corresponding to the associated temperature sensor;

[0011] The alternative temperature input data is used as the temperature data after the target temperature sensor enters the disaster recovery control strategy space, and a server thermal regulation strategy is generated in the disaster recovery control strategy space.

[0012] A second aspect of this application provides a server temperature sensor disaster recovery control device, the device comprising an identification module, a matching module, a determination module, and a generation module:

[0013] The identification module is used to identify the presence status and temperature acquisition status of multiple temperature sensors.

[0014] The matching module is used to match a status identifier based on the in-situ status and the temperature acquisition status. The same status identifier can identify both the in-situ status and the temperature acquisition status. Temperature sensors with completely identical in-situ status and temperature acquisition status have the same status identifier.

[0015] The matching module is used to match the disaster recovery control strategy space based on the status identifier. When the on-site status and acquisition status of the target temperature sensor in the temperature sensor meet the corresponding abnormal triggering conditions, it enters the disaster recovery control strategy space of the target temperature sensor.

[0016] The determining module is used to determine at least one associated temperature sensor based on the thermal region correlation of the target temperature sensor.

[0017] The determining module is used to generate alternative temperature input data based on the temperature data corresponding to the associated temperature sensor.

[0018] The generation module is used to take the alternative temperature input data as the temperature data after the target temperature sensor enters the disaster recovery control strategy space, and generate a server thermal regulation strategy in the disaster recovery control strategy space.

[0019] The server temperature sensor disaster recovery control method and apparatus provided in this application, compared with the prior art that only relies on a single abnormal acquisition to trigger disaster recovery and lacks fine-grained differentiation of sensor states, identifies the presence and temperature acquisition states of multiple temperature sensors and matches a unified state identifier based on the corresponding states. This enables differentiated handling of different abnormal situations, avoiding the use of the same control method for sensor absence and acquisition abnormalities. Furthermore, this application matches the disaster recovery control strategy space based on the state identifier, only entering the disaster recovery control space when the target temperature sensor meets the abnormal triggering conditions, effectively reducing the problem of erroneous switching caused by instantaneous communication fluctuations or short-term sampling abnormalities. In addition, this application determines associated temperature sensors through thermal region correlation and uses the temperature data of associated temperature sensors to generate alternative temperature input data. When the target temperature sensor is abnormal, the temperature data of the corresponding thermal region is compensated, enabling the server to maintain continuous operation of the thermal regulation process even during sensor abnormalities. In the disaster recovery control strategy space, the system uses alternative temperature input data to participate in the generation of thermal regulation strategies, so that the server can still maintain thermal control functions such as fan speed adjustment when the target temperature sensor is abnormal, thereby improving the stability of the server temperature control system and its fault tolerance under abnormal conditions. Attached Figure Description

[0020] Figure 1 A flowchart of an embodiment of the server temperature sensor disaster recovery control method provided in this application;

[0021] Figure 2 This is a schematic diagram of the structure of a first embodiment of the server temperature sensor disaster recovery control device provided in this application. Detailed Implementation

[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.

[0023] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0024] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0025] The following specific embodiments are given to illustrate the technical solution of this application in detail.

[0026] Figure 1 This is a flowchart of an embodiment of the server temperature sensor disaster recovery control method provided in this application. Please refer to... Figure 1 The method provided in this embodiment may include:

[0027] S101. Identify the presence status and temperature acquisition status of multiple temperature sensors.

[0028] Specifically, the server temperature control system includes a server management and control module (BMC), multiple distributed temperature sensors, a computing processing unit, and a heat dissipation execution unit. The computing processing unit includes, but is not limited to, heat-generating devices such as CPUs, GPUs, or network interface controllers (NICs). The heat dissipation execution unit includes fan assemblies or liquid cooling control assemblies. Each temperature sensor collects temperature data in real time for different hot areas within the server and transmits the collected temperature information to the BMC via a preset communication bus. This communication bus can be an I²C bus, an SMBus bus, or other control bus suitable for communication between sensors within the server. The BMC, as the core control unit of the server temperature control system, receives the temperature data uploaded by each temperature sensor and generates corresponding thermal regulation control commands based on a preset temperature control strategy. These commands control fan speed, heat dissipation power, or other heat dissipation execution parameters to dynamically adjust the internal thermal environment of the server.

[0029] Specifically, each temperature sensor collects temperature information of its corresponding hot area in real time and uploads the temperature data to the BMC; the BMC generates a thermal regulation strategy based on the received temperature data and outputs it to the fan control module to adjust the heat dissipation capacity inside the server and achieve temperature control of high-heat-generating components such as the CPU and GPU.

[0030] Specifically, during server operation, in addition to receiving real-time temperature data from multiple temperature sensors, the BMC also manages the operating status of each temperature sensor in a unified manner and sets up an internal status management module to record the combined results of the temperature sensor's in-situ status and temperature acquisition status.

[0031] Specifically, the "in-situ" status characterizes the physical presence of the temperature sensor in the server hardware system, while the "temperature acquisition" status characterizes whether the temperature sensor can output valid temperature data normally. Based on the combination relationship between these two types of statuses, BMC generates corresponding status combinations and establishes a status identification system in the status management module to uniquely identify and distinguish different status combinations.

[0032] Furthermore, the BMC internally divides the control strategy space into a normal control strategy space and a disaster recovery control strategy space. The normal control strategy space is used to execute thermal regulation strategies based on real-time temperature data output by the temperature sensor, while the disaster recovery control strategy space is used to execute thermal regulation strategies based on alternative temperature input data in the event of a temperature sensor malfunction.

[0033] Specifically, the status identification system and the control strategy space are linked through the strategy scheduling logic within the BMC, enabling different status identifiers to be mapped to the corresponding control strategy space, thereby achieving dynamic switching and unified management of server temperature control strategies.

[0034] Specifically, the BMC periodically or based on a preset trigger mechanism performs status acquisition operations on the temperature sensors to obtain basic operating status information corresponding to each temperature sensor. The basic operating status information includes at least in-situ status information and temperature acquisition status information, wherein the in-situ status information is used to characterize the physical connection status of the temperature sensor in the server hardware system, and the temperature acquisition status information is used to characterize whether the temperature sensor can output valid temperature data normally.

[0035] Specifically, the BMC obtains the presence status information of the temperature sensor through hardware detection signals or communication link handshake signals. For example, it determines whether the temperature sensor is in or out of the system by detecting the presence of I²C or SMBus communication response signals. The presence status indicates whether the temperature sensor is physically present, connected, and recognized by the system, while the absence status indicates whether the temperature sensor is not deployed or recognized by the system.

[0036] Specifically, the BMC obtains temperature acquisition information by reading temperature register data from the temperature sensor or returning values ​​from the acquisition interface, and determines the temperature acquisition status based on the temperature acquisition information. The temperature acquisition status indicates the availability of temperature data and operational constraints, including normal acquisition status, over-threshold acquisition status, and abnormal acquisition status.

[0037] Specifically, when the BMC successfully completes the temperature reading operation of the target temperature sensor, and the read data meets the preset data format, it enters the data validity determination process. If the temperature value is within the preset safety range, it is determined to be in normal acquisition state; if the temperature value exceeds the safety range, it is determined to be in over-threshold acquisition state. The over-threshold acquisition state indicates that data is still obtainable but has exceeded the server's safe operating range. The over-threshold acquisition state further includes recoverable and unrecoverable states; when the temperature data exceeds the preset safety threshold and then recovers to the safety range within a subsequent preset sampling period, it is determined to be in a recoverable state; when the temperature data continuously exceeds the safety threshold range within a consecutive preset time window and does not recover, it is determined to be in an unrecoverable state.

[0038] Specifically, when the BMC fails to complete the temperature reading operation of the target temperature sensor, or when at least one of the following occurs during the reading process: communication timeout, no response, abnormal data format, or abnormal data verification, the temperature acquisition status is determined to be an abnormal acquisition status, which is used to characterize that the target temperature sensor cannot output valid temperature data.

[0039] It should be noted that the preset data format or register encoding rules are used to limit the validity constraints of the temperature sensor output data. In one possible implementation, the data format includes data bit width, byte order, and frame structure; the register encoding rules are used to limit the register address range and value validity corresponding to the temperature data, to avoid reserved registers or abnormally encoded data being misidentified as valid temperature values. In the data validity judgment, when the read temperature data meets the format constraints and does not match the preset abnormal fill value, the data is determined to be valid temperature data; the preset abnormal fill value is used to identify communication abnormalities or invalid output states, including default fill values, communication failure return values, or invalid flag values. When the read data equals the preset abnormal fill value, the temperature data is considered invalid. By combining the determination of the in-situ state and the temperature acquisition state, a unified expression of the temperature sensor's operating state is achieved, and an input basis is provided for subsequent state flag generation and thermal control strategy switching.

[0040] S102. Match the status identifier based on the in-situ status and temperature acquisition status. The same status identifier can identify both the in-situ status and the temperature acquisition status. Temperature sensors with completely identical in-situ status and temperature acquisition status have the same status identifier.

[0041] Specifically, after acquiring the in-situ status and temperature acquisition status of each temperature sensor, the BMC uses a preset state mapping mechanism to uniformly identify different state combinations, generating corresponding state identifiers. These state identifiers encode the operating status of the temperature sensors and serve as input for subsequent thermal control strategy switching and disaster recovery control triggering. They are also used for status recording and interface display. The same in-situ status and temperature acquisition status combination corresponds to the same state identifier, while different combinations correspond to different state identifiers. This state identifier system associates the in-situ status of the temperature sensors with their temperature acquisition status, supporting the differentiated processing of different control logics.

[0042] Optionally, in one possible implementation, the method for establishing and matching the status identifier includes:

[0043] (1) Obtain the preset set of in-situ status types and the set of temperature acquisition status types.

[0044] Specifically, during the initialization phase of the server temperature control system, BMC pre-establishes a set of in-situ status types and a set of temperature acquisition status types to uniformly classify and manage the operating status of different temperature sensors. The in-situ status type set characterizes the physical connection status of the temperature sensor within the server hardware system, and includes at least in-situ and out-of-situ states. The temperature acquisition status type set characterizes the reading status and safety status of the temperature data corresponding to the temperature sensor, and includes at least normal acquisition status, temperature exceeding threshold status, and acquisition anomaly status.

[0045] (2) Based on the set of in-situ state types and the set of temperature acquisition state types, multiple state combinations are generated by arranging and combining them.

[0046] Specifically, BMC combines the two types of states into multiple state combinations based on the correspondence between the in-situ state type set and the temperature acquisition state type set, which are used to characterize the state description of the temperature sensor under different operating scenarios.

[0047] Specifically, the set of in-situ state types is used to characterize the physical presence attributes of the temperature sensor in the server hardware system, including at least in-situ and out-of-situ states; the out-of-situ state can be further divided into: sensor not actually deployed state and sensor not identified state. The sensor not actually deployed state characterizes the situation where the target sensor is not physically installed in the server hardware; the sensor not identified state characterizes the situation where the target sensor has been physically deployed but has not been successfully identified by the BMC.

[0048] Specifically, the set of temperature acquisition state types is used to characterize the temperature sensor's ability to acquire temperature data and its data availability characteristics during current operation, including at least normal acquisition state, temperature exceeding threshold state, and acquisition abnormal state. The temperature exceeding threshold state further includes a recoverable exceeding threshold state and an unrecoverable exceeding threshold state.

[0049] Specifically, BMC performs a Cartesian combination based on the set of in-situ state types and the set of temperature acquisition state types to form a complete state combination space, which is used to describe all possible mapping relationships between in-situ states and acquisition states.

[0050] Specifically, based on the actual operational constraints of the server, the state combination space is engineered and filtered to obtain a set of state combinations for actual control, which includes at least: a combination of in-situ state and normal acquisition state, a combination of in-situ state and temperature exceeding threshold recoverable state, a combination of in-situ state and exceeding threshold non-recoverable state, a combination of in-situ state and acquisition abnormal state, and an out-of-situ state. The states are no longer further divided into sensor not actually deployed state and sensor not identified state.

[0051] (3) Based on the preset mapping rules, generate corresponding state identifiers for each state combination.

[0052] Specifically, based on obtaining multiple state combinations, the BMC performs unique encoding processing on each state combination according to a preset mapping rule to generate a corresponding state identifier. The state identifier is used to uniquely identify the operating state of the temperature sensor, ensuring that different state combinations have a unified encoding expression that is identifiable, matchable, and schedulable within the BMC. Table 1.1 shows the mapping relationship between state combinations and state identifiers as shown in Embodiment 1 of this application:

[0053] Table 1.1 Mapping Relationship between State Combinations and State Identifiers as shown in Example 1

[0054]

[0055] Specifically, through the aforementioned mapping relationship, each state combination can be uniformly converted into a standardized state identifier, facilitating the BMC to perform control policy matching and disaster recovery policy scheduling based on the state identifier in subsequent steps. The introduction of state identifiers compresses the originally multi-dimensional state combinations into a single encoded value, improving the system's processing efficiency and consistency in state identification, policy matching, and log recording.

[0056] (4) Obtain the actual in-situ status and actual temperature acquisition status of the temperature sensor at the current moment.

[0057] Specifically, during server operation, the BMC periodically acquires the current actual presence status and actual temperature acquisition status of each temperature sensor. The actual presence status can be determined based on hardware detection signals, communication link response results, or device identification results; the actual temperature acquisition status can be determined based on temperature readings, temperature data validity, and temperature safety threshold judgment results. Furthermore, the BMC can continuously update the actual status information of each temperature sensor based on a preset sampling period to form dynamic status monitoring results.

[0058] (5) Based on the actual in-situ state and the actual temperature acquisition state, match the multiple state combinations to obtain the corresponding state identifier.

[0059] Specifically, the BMC matches the temperature sensor's current in-situ state and actual temperature acquisition status against multiple pre-established state combinations to determine the target state combination. After determining the target state combination, it further obtains the corresponding state identifier based on the mapping relationship between the target state combination and the state identifier. This state identifier characterizes the temperature sensor's current operating state and serves as the foundational state input for subsequent thermal control strategy selection, disaster recovery control strategy triggering, and log recording processing.

[0060] S103. Based on the status identifier, the disaster recovery control strategy space is matched. When the on-site status and acquisition status of the target temperature sensor in the temperature sensor meet the corresponding abnormal triggering conditions, the disaster recovery control strategy space of the target temperature sensor is entered.

[0061] Specifically, based on the mapping relationship between the state identifier and the preset strategy, the BMC divides the control logic corresponding to the temperature sensor into different control strategy spaces under different state conditions. The control strategy execution space includes a normal control strategy space and a disaster recovery control strategy space. After completing the above control strategy space division, the BMC further dynamically switches and manages the control strategy space according to state changes during operation.

[0062] Optionally, in one possible implementation, when the in-situ status and temperature acquisition status of the target temperature sensor meet the abnormal triggering conditions, the system enters the disaster recovery control strategy space and generates a server thermal control strategy based on alternative temperature input data; the method further includes:

[0063] (1) When the temperature sensor is in the in-situ state and the temperature acquisition state is normal, it enters the normal control strategy space and generates a server thermal control strategy based on the temperature data output by the temperature sensor.

[0064] Specifically, after acquiring the status identifier, the BMC determines the control strategy space corresponding to the current status identifier based on the mapping rules between the status identifier and the preset control strategy space, and selects to enter the corresponding control strategy execution path accordingly, so as to realize differentiated temperature control scheduling under different operating scenarios.

[0065] Specifically, the control strategy space includes a normal control strategy space and a disaster recovery control strategy space. The normal control strategy space is used to execute thermal regulation strategies when the target temperature sensor does not meet the abnormal triggering conditions. When the temperature sensor is in a normal acquisition state, the BMC executes closed-loop fan speed control based on the real-time temperature data output by the target temperature sensor to maintain the server's thermal balance operation. When the target temperature sensor is in an abnormal state but does not meet the abnormal triggering conditions, the BMC maintains the current normal control strategy space without switching and continues to execute thermal regulation strategies based on alternative temperature input data to ensure the continuity and stability of the server's thermal regulation process.

[0066] Specifically, the disaster recovery control strategy space is used to execute a disaster recovery thermal control strategy based on alternative temperature input data when the target temperature sensor meets the abnormal triggering conditions. This strategy replaces the real-time temperature input of the target temperature sensor, thereby achieving redundant temperature regulation of the corresponding thermal region. For the implementation method and principle of the alternative temperature input data, please refer to the relevant description; it will not be repeated here.

[0067] Furthermore, the normal control strategy space and the disaster recovery control strategy space employ different control objectives. The normal control strategy space is used to balance heat dissipation efficiency, system energy consumption, and operating noise while meeting thermal safety constraints; the disaster recovery control strategy space is used to prioritize server thermal safety in the event of a target temperature sensor malfunction, reducing the risk of local thermal runaway by increasing heat dissipation priority.

[0068] Furthermore, when BMC performs a control strategy space switch, it does not immediately switch control logic based on a single state change. Instead, it combines preset anomaly trigger conditions to determine the continuity of state changes. Only when the number of consecutive anomalies reaches a first continuity threshold is a switch from the normal control strategy space to the disaster recovery control strategy space triggered; similarly, only when the number of consecutive normal operations reaches a second continuity threshold is a recovery switch from the disaster recovery control strategy space to the normal control strategy space triggered. This mechanism can avoid frequent control strategy switching due to instantaneous communication fluctuations or short-term sampling anomalies. For the specific implementation methods of the first and second continuity thresholds, please refer to the relevant descriptions, which will not be repeated here.

[0069] (2) When the target temperature sensor is in an abnormal state but the abnormal triggering condition is not met, the normal control strategy space is maintained, and the server thermal control strategy is generated based on the alternative temperature input data.

[0070] Specifically, when the target temperature sensor is in an abnormal state but the preset abnormal trigger condition has not yet been met, the BMC maintains the current normal control strategy space without switching. In the case where the abnormal trigger condition is not met, the BMC assumes that the current target hot area still has the ability to estimate temperature, so it continues to operate the normal thermal speed regulation control process without immediately entering the disaster recovery control strategy space.

[0071] Specifically, BMC determines at least one associated temperature sensor based on the thermal region correlation of the target temperature sensor, and acquires the real-time temperature data corresponding to the associated temperature sensor, as well as the historical temperature data of the target temperature sensor, to generate alternative temperature input data. The specific implementation principle and method of the associated temperature sensor are described in the relevant documentation and will not be repeated here.

[0072] Specifically, the alternative temperature input data is used to supplement the temperature information of the target hot area when the target temperature sensor cannot provide effective temperature information or is in an abnormal state, in order to maintain the continuity and stability of the server's thermal regulation process. The specific calculation method for the alternative temperature input data is described in the relevant section and will not be repeated here.

[0073] (3) Wherein, the thermal regulation strategy in the normal control strategy space is used to perform closed-loop thermal speed regulation control when the abnormal triggering condition is not met, and the thermal regulation strategy in the disaster recovery control strategy space is used to perform redundant thermal speed regulation control on the corresponding thermal area of ​​the target temperature sensor based on the alternative temperature input data when the target temperature sensor meets the abnormal triggering condition.

[0074] Specifically, the closed-loop thermal speed control refers to a control method that constructs a closed-loop control circuit based on temperature feedback and dynamically adjusts the fan PWM parameters according to the error between the real-time temperature and the target threshold. In this control method, the temperature input source can be the sensor's own real-time temperature data or a temperature estimate incorporating alternative data, as long as the control structure is closed-loop. The redundant thermal speed control refers to performing thermal speed control based solely on alternative temperature input data when the target temperature sensor cannot provide reliable real-time temperature data. It employs a more conservative safety strategy than normal closed-loop control, such as increasing the base speed and limiting the rate of speed change, to prioritize thermal safety.

[0075] Specifically, the normal control strategy space is used to execute server thermal speed control when the temperature sensor does not meet the abnormal triggering conditions. In the normal control strategy space, the BMC generates fan PWM control parameters based on the temperature data corresponding to each hot zone to maintain a dynamic balance between server thermal balance and operating energy consumption. Specifically, when the temperature sensor is in a normal acquisition state, the BMC executes closed-loop thermal speed control based on the real-time temperature data output by the sensor; when the target temperature sensor is in an abnormal state but does not meet the abnormal triggering conditions, the BMC maintains the normal control strategy space without switching and compensates for the temperature information of the target hot zone based on alternative temperature input data to generate control input, replacing the real-time temperature data of the target temperature sensor in the thermal speed control.

[0076] Specifically, the disaster recovery control strategy space is used to execute redundant thermal speed control when the target temperature sensor meets the abnormal triggering conditions. In the disaster recovery control strategy space, the BMC no longer relies on the real-time acquisition results of the target temperature sensor, but instead estimates the temperature of the target thermal region based on alternative temperature input data, and generates a corresponding thermal regulation strategy based on the estimation results. At this time, the BMC executes redundant thermal speed control, and its control objective is adjusted from the energy consumption and heat dissipation balance under normal conditions to prioritizing thermal safety constraints.

[0077] Furthermore, after entering the disaster recovery control strategy space, the BMC increases the base PWM output value of the fan corresponding to the target hot area and lowers the fan speed-up trigger threshold. Simultaneously, it applies a slope constraint to the fan speed change rate to improve heat dissipation redundancy under abnormal conditions and suppress control fluctuations. In addition, within the disaster recovery control strategy space, the BMC can perform frequency reduction or current limiting control on the computing resources associated with the target hot area to reduce continuous thermal load.

[0078] It should be noted that for the "out of position" state, the corresponding status identifiers are 0x00 and 0x02, respectively. The BMC does not participate in the thermal speed control of the target hot zone and marks the corresponding temperature input as an invalid input source during the control strategy execution process. For the "over-threshold" state, the corresponding status identifiers are 0x13 and 0x05, respectively. The BMC still performs closed-loop thermal speed control based on the corresponding temperature data under the normal control strategy space, and adds a safety correction coefficient to the fan speed adjustment result to improve heat dissipation redundancy. The safety correction coefficient is used to perform safety redundancy amplification processing on the fan speed output. Essentially, it enhances and corrects the fan PWM control quantity calculated based on the normal control strategy space to improve the system's heat dissipation margin under the "over-threshold" state. In one possible implementation, the safety correction coefficient is dynamically determined based on the magnitude of the current temperature deviating from the safety threshold. The closer the temperature is to or exceeds the safety upper limit, the larger the correction coefficient, so that the fan speed is increased and compensated based on the original control result. In another possible implementation, the safety correction coefficient can also be associated with the server's real-time load, further increasing the correction strength under high-load scenarios to enhance thermal safety protection capabilities.

[0079] Optionally, in one possible implementation, the disaster recovery control strategy space entering the target temperature sensor includes:

[0080] (1) Based on the continuous detection results, construct the state sequence of the target temperature sensor.

[0081] Specifically, the BMC performs periodic status checks on the target temperature sensor according to a preset sampling period, and records the in-situ status detection results corresponding to each sampling period in chronological order to construct a status sequence of the target temperature sensor. The in-situ status includes at least the normal in-situ state and the abnormal in-situ state, and may also include the absent in-situ state and the over-threshold state.

[0082] Specifically, the state sequence is used to characterize the state change process of the target temperature sensor in the time dimension, and is maintained in a sliding window manner. The window length of the sliding window is used to cover the stable determination period of the state change of the target temperature sensor, so as to avoid short-term fluctuations from affecting the continuous determination.

[0083] Specifically, the window length is at least greater than the larger of the first continuity threshold and the second continuity threshold, and redundant sampling periods are reserved to absorb the effects of sensor communication jitter and transient temperature fluctuations. In a preferred embodiment, the sliding window length is 8 to 16 sampling periods, for example, set to 10 sampling periods; when the sampling period is 1 second, it covers a state change observation window of approximately 8 to 16 seconds, which is used to meet the stability determination requirements under the server's thermal inertial response time scale.

[0084] (2) Based on the state sequence, count the number of consecutive occurrences of the target temperature sensor being in an abnormal state and the number of consecutive occurrences of it being in a normal state.

[0085] Specifically, the BMC traverses the state sequence based on time order and performs continuity checks on the in-situ states of adjacent sampling periods to generate continuous count values ​​for the corresponding states. When the in-situ state of the current sampling period is consistent with the in-situ state of the previous sampling period, the continuous count value corresponding to that state is incremented by one; when the in-situ state of the current sampling period is inconsistent with the in-situ state of the previous sampling period, the continuous count value of the current state is reset to 1, and the continuous count values ​​of other states are cleared to zero. In one implementation, in-situ states and states exceeding the threshold can be classified as in-situ abnormal states and included in the continuous counting to enhance the sensitivity of anomaly detection.

[0086] (3) When the number of consecutive occurrences of an in-situ abnormal state reaches the first continuity threshold, it enters the disaster recovery control strategy space.

[0087] Specifically, after each sampling period, the BMC determines whether the number of consecutive occurrences of the in-situ abnormal state has reached a preset first continuity threshold. When the number of consecutive abnormal occurrences is greater than or equal to the first continuity threshold, the abnormal state of the target temperature sensor is determined to have persistent characteristics, triggering a switch from the normal control strategy space to the disaster recovery control strategy space. The first continuity threshold is used to filter out false positives caused by communication jitter, instantaneous sampling failures, or short-term abnormalities, preventing frequent triggering of the disaster recovery strategy space. The first continuity threshold is preferably set to 3 sampling periods, and the specific value can be configured and adjusted according to the server's thermal inertia and the sampling period.

[0088] (4) When the number of consecutive occurrences of the in-place normal state reaches the second continuity threshold, exit the disaster recovery control strategy space and restore the normal control strategy space.

[0089] Specifically, during the operation of the disaster recovery control strategy space, the BMC continuously monitors the status of the target temperature sensor and counts the number of consecutive occurrences of the normal state. When the number of consecutive normal occurrences reaches a preset second continuity threshold, it is determined that the target temperature sensor has recovered to a stable normal state, triggering a switch from the disaster recovery control strategy space to the normal control strategy space. The second continuity threshold is used to avoid frequent reverts due to short-term recovery or occasional normal readings, improving the stability of the control strategy space switching. In one possible implementation, the second continuity threshold is less than or equal to the first continuity threshold, preferably set to 2 or 3 sampling periods.

[0090] Optionally, in one possible implementation, the method for determining the first continuity threshold includes:

[0091] (1) Obtain historical running data, real-time running load data and historical abnormal statistics of the server.

[0092] Specifically, the BMC collects and aggregates multi-source information during server operation, using it as the basic input for dynamically determining the first continuity threshold, so that the disaster recovery triggering condition can adapt to the actual operating environment of the server. The historical operating data includes at least records of temperature sensor status changes, sampling results, and temperature fluctuations over historical operating cycles. In one possible implementation, the statistical time window for the historical operating data can be set to the past 7 days or 30 days to adapt to the stability differences of different server operating environments. This historical operating data is used in subsequent steps to extract the distribution characteristics of the continuous duration of abnormal states. The real-time operating load data is used to characterize the current computing pressure state of the server, including at least one of CPU utilization, GPU power consumption, memory occupancy, and I / O load. This data is used in subsequent steps to perform load-weighted correction on the basic continuity threshold so that the disaster recovery triggering sensitivity matches the current thermal load level. The historical anomaly statistics are used to characterize the statistical characteristics of abnormal states of the temperature sensor during historical operation, including the statistical distribution characteristics of the total number of in-situ anomalies and the duration of continuous anomalies, such as average duration, median duration, or high quantile duration. This data is used in subsequent steps to construct continuous distribution characteristics and calculate fluctuation frequency, avoiding false triggering or missed triggering caused by a single empirical threshold.

[0093] (2) Based on the historical operation data, statistically analyze the continuous distribution characteristics of the target temperature sensor during the historical operation process, and determine the basic continuity threshold based on the continuous distribution characteristics of the continuous in-situ anomalies.

[0094] Specifically, BMC performs time-series analysis on historical operational data, extracting the number of consecutive anomaly samplings corresponding to each historical anomaly event. For example, if an anomaly occurs consecutively for two sampling periods before recovering, a statistical distribution of the number of consecutive anomalies is constructed, such as calculating the frequency or cumulative proportion of each consecutive occurrence. The continuous in-situ anomaly persistence distribution characteristic is used to characterize the persistence pattern of the anomaly state over time. Engineering experience shows that short-term communication jitter or electromagnetic interference typically lasts only 1-2 sampling periods, while real hardware failures or physical damage usually last for multiple sampling periods, such as more than 5 times. Based on this, BMC selects a threshold that covers most short-term anomalies, such as a cumulative proportion reaching a preset range (e.g., 85%-90%), but does not cover persistent faults, as the basic continuity threshold. For example, if statistics show that 85% of anomalies occur no more than twice consecutively, the basic continuity threshold can be set to 3 times.

[0095] (3) The basic continuity threshold is corrected by load weighting based on the real-time operating load data.

[0096] Specifically, the BMC dynamically adjusts the basic continuity threshold based on the server's current real-time operating load level. The adjustment principle is as follows: Under high load scenarios, the server heats up at a higher rate and the temperature rises faster, resulting in a shorter tolerable duration of anomalies; therefore, the continuity threshold should be lowered to improve disaster recovery response speed. Under low load scenarios, the overall thermal load of the server is lower, and temperature changes are relatively slow; the continuity threshold can be appropriately increased to enhance the system's resilience. In one possible implementation, the BMC determines the current server load level based on at least one of CPU utilization, GPU power consumption, and total system power consumption. When the load of any critical heat source reaches a preset high load threshold, it can be determined as a heavy load state. Specifically, it can be set as follows: when CPU utilization ≤ 30% or GPU power consumption ≤ 30% of rated power, it is considered a light load, and the basic threshold is multiplied by 1.2 (or added by 1) and rounded down; when CPU utilization ≥ 80% or GPU power consumption ≥ 80% of rated power, it is considered a heavy load, and the basic threshold is multiplied by 0.8 (or subtracted by 1) and rounded down, with a lower limit not lower than 1; all other cases are considered medium loads, and the basic threshold remains unchanged. The above load level classification thresholds can be adjusted according to the specific server model.

[0097] (4) Based on the fluctuation frequency of the historical abnormal statistical data within a preset time window, the modified basic continuity threshold is adjusted by adding or subtracting to obtain the first continuity threshold.

[0098] Specifically, BMC further utilizes the fluctuation frequency from historical anomaly statistics for secondary correction. The fluctuation frequency characterizes the switching frequency of the temperature sensor between normal and abnormal acquisition states within a preset time window, including at least the switching between a normal in-situ state and an abnormal in-situ state. The preset time window can be set to the past 1 hour or 24 hours. This statistics are used to identify whether the sensor exhibits frequent fluctuations in instability and adjust the sensitivity of disaster recovery triggering. The fluctuation frequency is calculated by dividing the total number of state switching events within the statistical time window by the window duration. For example, if 10 switching events occur in the past 1 hour, the fluctuation frequency is 10 times / hour. There is a positive correlation between the fluctuation frequency and the correction coefficient; that is, the higher the fluctuation frequency, the larger the corresponding threshold correction coefficient. Specifically: If the fluctuation frequency is greater than or equal to the high-frequency threshold, such as ≥10 times / hour, it indicates that the sensor is currently unstable or in a strong interference environment. In this case, the first continuity threshold should be increased, for example, by adding 1 or multiplying by 1.1, to reduce the probability of false triggering of disaster recovery. If the fluctuation frequency is less than or equal to the low-frequency threshold, such as ≤2 times / hour, it indicates that the system is stable. The first continuity threshold can be reduced, for example, by subtracting 1 or multiplying by 0.9, to improve the anomaly response speed. If the fluctuation frequency is between the two, no correction is made or only minor adjustments are made. Finally, the value obtained after load-weighted correction and fluctuation frequency increase / decrease correction is the dynamically determined first continuity threshold.

[0099] It should be noted that, to avoid interference from dynamic threshold changes on the currently ongoing continuous judgment cycle, BMC updates the first continuous threshold only after completing the current judgment cycle, ensuring the atomicity and stability of the judgment process. BMC can update this threshold periodically, for example, every 10 minutes, or trigger an update when a significant change in load status is detected, to ensure that the threshold continuously adapts to the current operating environment.

[0100] Furthermore, the method for determining the second continuity threshold is similar to that for determining the first continuity threshold. Specifically, the BMC performs statistical analysis on the continuous occurrence characteristics of the target temperature sensor being in a normal in-situ state based on historical operating data, real-time operating load data, and historical anomaly statistics, and determines a basic recovery threshold based on this. Further, the BMC performs load-weighted correction on the basic recovery threshold based on current real-time operating load data, and adjusts the corrected basic recovery threshold by adding or subtracting based on historical state fluctuation frequencies to obtain the second continuity threshold. The second continuity threshold is used to control the judgment conditions for recovery from the disaster recovery control strategy space to the normal control strategy space, and to suppress frequent control strategy reverts caused by short-term state recovery fluctuations. In one embodiment, the second continuity threshold is less than or equal to the first continuity threshold, so that the system can revert to the normal control strategy space under more stable recovery conditions after entering the disaster recovery state. In one embodiment, the second continuity threshold can also be directly set to a fixed empirical value, such as 2 or 3 times.

[0101] In summary, this embodiment determines the basic continuity threshold by analyzing historical anomaly distributions and dynamically adjusts the threshold based on the server's real-time load and historical anomaly fluctuations in the temperature sensor, allowing it to adapt to different operating states. For example, when the server is under high load, the continuity threshold can be appropriately lowered to trigger disaster recovery processing more quickly in case of anomalies; conversely, when the system load is low, or when the temperature sensor itself experiences frequent state fluctuations, the continuity threshold can be appropriately relaxed to avoid frequent switching of control strategies due to short-term fluctuations. This reduces the repeated switching of control strategies within a short period, making the server temperature control process more stable.

[0102] S104. Determine at least one associated temperature sensor based on the thermal region correlation of the target temperature sensor.

[0103] Specifically, the BMC first acquires the target thermal region information corresponding to the target temperature sensor, and based on the pre-established thermal region topology relationship within the server, determines at least one associated thermal region that has a thermal correlation with the target thermal region. The target temperature sensor is the temperature sensor currently requiring temperature compensation, disaster recovery analysis, or associated temperature estimation processing; the associated thermal region is used to characterize the thermally affected area that has a thermal conduction relationship, airflow coupling relationship, collaborative load relationship, or shared heat dissipation resource relationship with the target thermal region. In one embodiment, the thermal region topology relationship is pre-established based on the server hardware layout structure, including the spatial adjacency relationship, airflow direction relationship, thermal impact relationship, and heat dissipation resource sharing relationship between the CPU region, GPU region, NIC region, memory region, power supply region, and storage region.

[0104] Specifically, the BMC internally maintains a thermal region association mapping table to record the correspondence and association weights between target thermal regions and associated thermal regions. This mapping table can be pre-configured based on hardware topology during server initialization and can also be dynamically updated based on historical thermal characteristic analysis results during server operation. Based on the associated thermal regions, the BMC further determines the corresponding set of associated temperature sensors. The associated temperature sensors include at least one or more of the following: other temperature sensors located in the same functional component as the target temperature sensor; temperature sensors located in adjacent thermal regions as the target temperature sensor; temperature sensors located on the same airflow cooling path; temperature sensors that have a historical temperature change correlation with the target thermal region; and temperature sensors that share a fan group or cooling module with the target thermal region. For example, when the target temperature sensor is a CPU region temperature sensor, the associated temperature sensors may include CPU adjacent VRM region temperature sensors, memory region temperature sensors, and airflow inlet temperature sensors; when the target temperature sensor is a GPU region temperature sensor, the associated temperature sensors may include GPU adjacent HBM region temperature sensors, PCIe region temperature sensors, and GPU airflow outlet temperature sensors.

[0105] Specifically, to improve the accuracy of alternative temperature input data, BMC can perform correlation degree screening on candidate associated temperature sensors to determine the final set of associated temperature sensors participating in the generation of alternative temperature input data. The correlation degree characterizes the thermal correlation between candidate temperature sensors and the target hot area, and its determining factors include spatial distance, airflow path distance, historical temperature change correlation coefficient, thermal response synchronization, functional component synergy, and the degree of heat dissipation resource sharing. Generally, the closer the spatial distance, the higher the temperature change synchronization, and the larger the historical correlation coefficient, the higher the correlation degree. In one implementation, BMC selects only temperature sensors with a correlation degree greater than a preset correlation threshold as the final associated temperature sensors, or selects the top N associated temperature sensors based on their correlation degree to participate in the subsequent generation of alternative temperature input data. The correlation degree can be dynamically updated according to server operating status, load changes, and historical thermal characteristic changes.

[0106] Specifically, to prevent the spread of abnormal states from causing associated compensation failure, the BMC performs validity checks on candidate associated temperature sensors. When a candidate associated temperature sensor is in an abnormal state (in-situ), an in-situ state, or an unavailable state exceeding a threshold, it is removed from the associated temperature sensor set to prevent abnormal temperature data from propagating during the associated compensation process. Simultaneously, to reduce real-time computational overhead during disaster recovery triggering, the BMC can pre-cache the associated temperature sensor set corresponding to each target thermal region. If the number of remaining associated temperature sensors after screening and validity checks is lower than a preset minimum threshold (e.g., less than one), the BMC will fall back to using historical temperature data from the target temperature sensor as an alternative input source, or will use a system-level default conservative temperature value for subsequent disaster recovery thermal control.

[0107] Specifically, BMC can dynamically determine the set of reliable associated temperature sensors corresponding to the target temperature sensor based on the correlation between thermal regions within the server. This provides a stable and reliable data source for the generation of subsequent alternative temperature input data, improving the accuracy of temperature estimation, the stability of thermal regulation, and the robustness of the system under abnormal scenarios in the server's temperature control disaster recovery mechanism. For details on the specific generation method of alternative temperature input data, please refer to the relevant description; it will not be repeated here.

[0108] S105. Generate alternative temperature input data based on the temperature data corresponding to the associated temperature sensor.

[0109] Specifically, after determining the set of associated temperature sensors corresponding to the target temperature sensor, BMC acquires the temperature data corresponding to the associated temperature sensors and generates alternative temperature input data to replace the output of the target temperature sensor based on the temperature data. This alternative temperature input data is used to estimate and compensate for the temperature state of the target thermal region when the target temperature sensor cannot provide effective temperature information, is in an abnormal state, or the system enters the disaster recovery control strategy space, and serves as the temperature input for the subsequent generation of server thermal control strategies.

[0110] Optionally, in one possible implementation, generating alternative temperature input data based on the temperature data corresponding to the associated temperature sensor includes:

[0111] (1) Obtain the historical temperature data sequence of the target temperature sensor within a preset time window.

[0112] Specifically, when the target temperature sensor cannot provide valid temperature data, the BMC acquires a sequence of historical temperature data from the sensor within a preset time window before the anomaly occurs. This historical temperature data sequence characterizes the temperature change trend of the target thermal region over a period before the anomaly occurs, and includes at least historical temperature values ​​and corresponding timestamps from multiple consecutive sampling periods. The length of the preset time window can be configured based on the server's thermal inertia characteristics and sampling period, for example, set to 30 seconds, 60 seconds, or 5 minutes before the anomaly occurs. In one embodiment, if the target temperature sensor exhibits short-term instability before the anomaly occurs, the BMC can select data from an earlier stable historical window, such as 10 minutes to 5 minutes before the anomaly, to avoid contamination of the historical baseline by recent anomaly fluctuations.

[0113] Furthermore, BMC can perform validity verification and preprocessing on the historical temperature data sequence, including one or more of moving average processing, median filtering processing, and outlier removal processing, to improve the stability of subsequent fusion calculations.

[0114] (2) Based on the temperature data of the associated temperature sensor and the historical temperature data sequence, perform fusion calculation to generate alternative temperature input data.

[0115] Specifically, the BMC acquires temperature data from each associated temperature sensor in the associated temperature sensor set and performs fusion calculations in conjunction with historical temperature data sequences to generate alternative temperature input data. The fusion calculations comprehensively reflect the current thermal state of the target thermal region and its historical temperature change trends.

[0116] Specifically, BMC determines the corresponding fusion weights based on the correlation between associated temperature sensors and the target thermal region; the higher the correlation, the greater the weight. The correlation can be determined based on one or more of spatial proximity, airflow coupling, historical temperature change correlation, and thermal response synchronicity. BMC first performs weighted fusion of temperature data from multiple associated temperature sensors to obtain a fused associated temperature data value. Then, it performs a secondary fusion of this fused value with historical temperature trend values ​​to generate alternative temperature input data. The historical temperature trend values ​​are predicted temperature values ​​calculated based on the changing trends of historical temperature data sequences. The alternative temperature input data can be generated in the following way:

[0117] ;

[0118] in, To replace temperature input data, The target temperature sensor's historical temperature average value within a preset time window. This refers to the real-time temperature data corresponding to the associated temperature sensor, or the weighted fusion result of real-time temperature data from multiple associated temperature sensors. The compensation coefficient α is used to characterize the degree of thermal correlation between the correlated temperature sensor and the target thermal region. The closer the spatial distance between the correlated temperature sensor and the target thermal region, the higher the degree of airflow coupling, or the stronger the correlation with historical temperature changes, the larger the corresponding compensation coefficient α.

[0119] It should be noted that the above method enables the alternative temperature input data to inherit the original historical temperature change trend of the target thermal region, and can also be dynamically corrected by combining the real-time temperature change of the current associated thermal region, so as to improve the accuracy and continuity of temperature estimation under abnormal conditions. In one possible implementation, the alternative temperature input data can also be calculated by weighted averaging, trend prediction, time series fitting or machine learning models, etc., to fuse historical temperature data and associated temperature data. This application does not limit this.

[0120] Specifically, to avoid instability in alternative temperature input data caused by fluctuations in a single associated temperature sensor, BMC can also perform smoothing, outlier filtering, and rate of change limiting on real-time data from multiple associated temperature sensors. If the associated temperature sensor set is empty or all corresponding temperature data are invalid after validity verification, BMC will fall back to using the historical trend prediction value corresponding to the historical temperature data sequence as the alternative temperature input data.

[0121] (3) Use the alternative temperature input data as the temperature input data in the disaster recovery control strategy space.

[0122] Specifically, once the target temperature sensor meets the abnormal triggering conditions and enters the disaster recovery control strategy space, the BMC uses the generated alternative temperature input data as the temperature input data for the corresponding target hot area, participating in the subsequent generation of thermal control strategies. In the disaster recovery control strategy space, the BMC no longer relies on the real-time sampling results of the target temperature sensor, but instead estimates the temperature of the target hot area based on the alternative temperature input data, and generates the corresponding thermal control strategy based on the estimation results. In one possible implementation, the BMC can set an effective time window for the alternative temperature input data. If effective temperature acquisition is not restored after this time window, the BMC further increases the heat dissipation redundancy level or implements functional component operation restrictions, such as performance limitation control, to reduce the thermal risk under prolonged abnormal conditions.

[0123] S106. The alternative temperature input data is used as the temperature data after the target temperature sensor enters the disaster recovery control strategy space, and a server thermal regulation strategy is generated in the disaster recovery control strategy space.

[0124] Specifically, when the target temperature sensor meets the abnormal triggering conditions and enters the disaster recovery control strategy space, the BMC will replace the real-time temperature data of the original target temperature sensor with alternative temperature input data as the main temperature feedback variable for the target hot area. In the disaster recovery control strategy space, the BMC no longer relies on the real-time sampling results of the target temperature sensor, but instead uses the alternative temperature input data as the core feedback source, and combines it with the target temperature setpoints for each hot area of ​​the server, real-time operating load, and heat dissipation resource status to generate corresponding fan speed control strategies. The target temperature setpoint is a pre-configured temperature control target point, such as a CPU core temperature set to 75°C, used to characterize the equilibrium temperature expected to be achieved by the closed-loop control. The BMC determines the thermal regulation error based on the deviation between the alternative temperature input data and the target temperature setpoint, and adjusts the fan PWM output parameters according to the thermal regulation error to achieve closed-loop thermal speed control of the target hot area.

[0125] Specifically, closed-loop thermal speed control can employ a proportional-integral-derivative (PID) control strategy or a preset temperature-speed mapping curve. The BMC can dynamically select the control method based on the fluctuation characteristics of the alternative temperature input data or the current load level of the server. For example, PID can be used to improve response speed under high load or drastic temperature fluctuations, while mapping curves can be used to reduce computational overhead under low load or stable temperature.

[0126] Furthermore, in the disaster recovery control strategy space, BMC prioritizes thermal safety constraints as the primary control objective, giving it a higher priority for heat dissipation compared to the normal control strategy space. Specifically, when generating PWM control parameters, BMC applies a positive bias to the fan base speed, for example, by increasing the base PWM by 10%, while simultaneously reducing the tolerance margin between the target temperature setpoint and the safety upper limit, for example, by lowering the safety upper limit threshold by 5°C, to enhance the heat dissipation redundancy capability under abnormal conditions.

[0127] Specifically, to avoid control output oscillations caused by fluctuations in the alternative temperature input data, the BMC imposes a rate-of-change constraint on the fan PWM output parameters, for example, the change does not exceed 5% per control cycle, and performs first-order low-pass filtering on short-term temperature fluctuations to improve the stability of the disaster recovery control process.

[0128] Specifically, when the disaster recovery control strategy space continues to operate for more than a preset time threshold, and the target temperature sensor still has not recovered its normal data acquisition capability, the BMC can further trigger enhanced thermal protection strategies, including reducing the operating frequency of the computing unit, limiting the power consumption limit, or increasing the forced cooling level. The preset time threshold can be configured by the system designer based on the server's thermal inertia and safety requirements, such as 10 minutes or 30 minutes, or can be dynamically adjusted based on historical anomaly statistics.

[0129] Specifically, under abnormal conditions of the target temperature sensor, BMC uses the alternative temperature input data as the core feedback variable to maintain closed-loop thermal regulation capability in the disaster recovery control strategy space, so that the server thermal management system still has stable, continuous and controllable thermal regulation capability in the case of sensor failure.

[0130] Optionally, in one possible implementation, when the thermal control strategy corresponds to multiple thermal zone control objects, the method further includes:

[0131] (1) Obtain multiple thermal zone speed control objects corresponding to the target temperature sensor.

[0132] Specifically, when executing the thermal regulation strategy, BMC determines multiple thermal zone speed control objects associated with the target thermal zone based on the target thermal zone range mapped by the target temperature sensor. These thermal zone speed control objects represent control execution units within the server that participate in heat dissipation regulation, and include at least one or more of the following: CPU area fan control unit, GPU area fan control unit, power module heat dissipation control unit, storage device heat dissipation control unit, and chassis airflow main fan control unit.

[0133] Specifically, the BMC determines multiple thermal zone speed control objects corresponding to the target temperature sensor based on a pre-configured thermal zone mapping table, and incorporates these control objects into the same thermal regulation strategy for unified control. In one possible implementation, each thermal zone speed control object uses the same type of control strategy interface to receive uniformly issued thermal regulation strategy parameters, ensuring that multiple control objects can execute the speed regulation strategy under the same control mode.

[0134] (2) When the operation mode switching of the thermal regulation strategy is triggered, a unified strategy switching instruction is issued to the multiple thermal zone speed regulation control objects.

[0135] Specifically, when the thermal control strategy switches operating modes, including switching from the normal control strategy space to the disaster recovery control strategy space, or restoring from the disaster recovery control strategy space to the normal control strategy space, the BMC generates a unified strategy switching instruction and synchronously sends it to multiple hot zone speed control objects. This unified strategy switching instruction instructs each hot zone speed control object to switch to the same control strategy space, and includes at least: a control strategy space identifier, a PWM control strategy type identifier, fan base speed correction parameters, fan speed slope limit parameters, and a control cycle synchronization flag. In one implementation, the unified strategy switching instruction is sent to each hot zone speed control object via BMC internal control bus broadcast or register group synchronous writing, ensuring that all control objects receive the same control mode switching signal within the same control scheduling cycle.

[0136] Furthermore, to avoid local asynchrony caused by differences in the execution timing of various controlled objects, the BMC introduces a state latching mechanism after the instruction is issued. This means that the old control strategy output is frozen within the current control cycle, and the new control strategy parameters are loaded uniformly only at the boundary of the next control cycle. This optimizes the control method from independent switching in each region to unified broadcast synchronous switching.

[0137] (3) Based on the unified strategy switching instruction, the operating mode switching of the thermal regulation strategy is completed within the same control cycle for each hot zone speed regulation control object.

[0138] Specifically, the BMC uses a preset control cycle as a unified scheduling benchmark to periodically refresh the control state of each hot zone speed control object. In one implementation, the control cycle is a fixed time interval, such as 50ms, 100ms, or 1s, used to match the refresh frequency of the fan PWM control. Upon receiving a unified strategy switching command, each hot zone speed control object does not immediately perform a mode switch, but instead uniformly enters the state switching process at the end of the current control cycle, and completes the synchronous switch from the old control strategy space to the new control strategy space at the same control cycle boundary. Furthermore, to avoid timing drift caused by differences in the execution response time of different control objects, the BMC adopts a periodic alignment latch mechanism, that is, keeping the old control output unchanged within the current control cycle, uniformly refreshing the control strategy space and control parameters at the cycle boundary, and synchronously enabling the new control strategy in the new cycle.

[0139] Optionally, in one possible implementation, the method further includes:

[0140] (1) Obtain the status change information of the target temperature sensor and the switching information of the disaster recovery control strategy space.

[0141] Specifically, the BMC continuously monitors the state changes of the target temperature sensor during operation and acquires the state change information of the target temperature sensor. This state change information characterizes the state evolution process of the target temperature sensor under different sampling periods, including at least records of changes between the in-situ normal state, in-situ abnormal state, out-of-situ state, and over-threshold state. In one embodiment, the state change information includes a state identifier change sequence and corresponding timestamp information, reflecting the dynamic trajectory of the sensor state over time. Additionally, the BMC acquires switching information of the disaster recovery control strategy space. This switching information characterizes the BMC's switching behavior between the normal control strategy space and the disaster recovery control strategy space, including: trigger events for switching from the normal control strategy space to the disaster recovery control strategy space, recovery events for returning from the disaster recovery control strategy space to the normal control strategy space, and corresponding switching time points and trigger condition identifiers.

[0142] (2) During the disaster recovery control strategy space switching process, the state change information and the control strategy space switching information are sampled synchronously.

[0143] Specifically, during the switching process between the normal control strategy space and the disaster recovery control strategy space by the BMC, a synchronous sampling mechanism is initiated to synchronously acquire state information and control strategy space switching information under a unified time reference. This synchronous sampling is used to simultaneously acquire state change information of the target temperature sensor and switching information of the disaster recovery control strategy space within the same control cycle. In one implementation, the BMC uses a fixed control cycle as the sampling reference, such as 50ms, 100ms, or 1s, and synchronously samples the above information at the boundary of each control cycle to ensure consistency between state changes and control strategy switching in the time dimension.

[0144] Furthermore, to avoid data loss or sampling offset caused by instantaneous policy switching, BMC introduces a sampling window extension mechanism before and after the switching trigger point. That is, it retains the historical sampling results of at least one control cycle before the switching occurs and retains the transition sampling results of at least one control cycle after the switching occurs, so as to fully record the state evolution during the switching process.

[0145] (3) The state change information obtained by synchronous sampling and the control strategy space switching information are associated and stored in chronological order to generate black box record data.

[0146] Specifically, the BMC sorts the synchronously sampled state change information and control strategy space switching information according to timestamps and establishes a one-to-one time correlation. In one implementation, using timestamps as indexes, the state information collected in each control cycle is bound and stored with the strategy switching state to generate structured black-box recording data. The black-box recording data includes at least: timestamp information, target temperature sensor status code, state change event identifier, control strategy space identifier, strategy switching trigger identifier, and control cycle number. The control strategy space identifier is used to identify the normal control strategy space or the disaster recovery control strategy space. Furthermore, the black-box recording data is continuously appended and stored in a time-series manner to form a complete operational trajectory recording chain.

[0147] (4) Store the black box recording data in a storage area independent of the BMC temperature control module.

[0148] Specifically, the BMC stores the generated black-box log data in a non-volatile storage area independent of the temperature control module. This independent storage area includes at least one of the following: an independent Flash storage unit on the server motherboard, an independent BMC log storage partition, or an external storage device.

[0149] Specifically, the black-box recorded data is stored using an append-only method to avoid overwriting historical records and to ensure data integrity and immutability. Furthermore, the independent storage area is physically or logically isolated from the BMC temperature control logic module, ensuring that the black-box recorded data does not participate in real-time thermal regulation calculations and is only used for subsequent diagnostic analysis, fault backtracking, and maintenance auditing.

[0150] Optionally, in one possible implementation, the method further includes:

[0151] (1) Obtain the state change sequence of the target temperature sensor within a preset time window.

[0152] Specifically, the BMC continuously samples and records the state of the target temperature sensor during operation, and forms a state change sequence within a preset time window. This state change sequence characterizes the state evolution of the target temperature sensor within a continuous sampling period, and includes at least the records of changes between the following states: normal in-situ state (0x01), abnormal in-situ state (0x03), absent in-situ state (0x00 / 0x02), and above-threshold state (0x13 / 0x05).

[0153] Specifically, the state change sequence is composed of a state identifier sequence and a timestamp sequence, for example, represented as {( , ), ( , ), …, ( , )},in This serves as the status indicator for the corresponding sampling period. Furthermore, the preset time window can be set as a sliding window, such as the operating interval of the most recent 10 seconds, 30 seconds, or 1 minute, to adapt to the temperature control response period of different servers.

[0154] (2) Based on the state change sequence, count the number of times the state switches between normal and abnormal states.

[0155] Specifically, the BMC sequentially traverses the state change sequence and counts the number of times the state switches between the normal state set and the abnormal state set within a preset time window. The normal state set includes at least the in-place normal state (0x01); the abnormal state set includes at least the in-place abnormal state (0x03). In one possible embodiment, the abnormal state set also includes the out-of-place state (0x00 / 0x02) and the over-threshold state (0x13 / 0x05).

[0156] Specifically, a switch from a normal state to an abnormal state, or vice versa, is counted as a valid switch event. For example, 0x01→0x03 counts as one switch, and 0x03→0x01 also counts as one switch. Furthermore, to avoid duplicate counting caused by sampling jitter, BMC performs deduplication on consecutive identical states, counting only when a substantial change in state occurs, thereby improving the stability and reliability of switch statistics.

[0157] (3) When the number of switching operations meets the preset suppression trigger condition, the switching suppression state is entered. In the switching suppression state, the switching operation of the control strategy space is restricted so that the number of switching operations of the control strategy space within the preset time window meets the preset stability constraint.

[0158] Specifically, when the number of handovers counted within a preset time window reaches or exceeds a preset suppression trigger condition, the BMC enters a handover suppression state. The preset suppression trigger condition can be a fixed threshold or a dynamic threshold; for example, if the number of handovers is greater than or equal to 3 within a 30-second window, the suppression state is triggered.

[0159] Specifically, under the switching suppression state, BMC restricts the switching behavior of the control policy space, including one or more of the following methods: adopting a delayed switching strategy, that is, delaying the entry into the disaster recovery control policy space or restoring the normal control policy space after the triggering conditions are met; introducing a cooldown time mechanism, that is, prohibiting switching again within a preset cooldown period after a switching occurs; enabling a suppression window mechanism, that is, only allowing the current control policy space to remain unchanged during the suppression state; or implementing a filtering short-term triggering strategy, that is, not triggering policy space switching for abnormal state changes that occur within a short period of time.

[0160] (4) After the preset time window ends, the number of switching times is reset to the initial state.

[0161] Specifically, when the preset time window expires, the BMC resets the cumulative number of switching counts within the current window to zero and restores the statistical state to its initial state. Specifically, the preset time window can be a sliding window or a fixed window: when using a sliding window, the window slides forward one sampling period, and a new state change sequence is re-statistically recorded; when using a fixed window, the count is reset uniformly after the window ends and statistics restart.

[0162] Furthermore, after the window is reset, the switching suppression state also exits simultaneously, and the system resumes normal control strategy space switching logic to re-enter the next state monitoring cycle. This mechanism enables the switching suppression control to have periodic reset capabilities, avoiding long-term suppression that could cause the system to be unable to respond to real, persistent faults, thus achieving a balance between system stability and timely disaster recovery response.

[0163] Through the aforementioned switching suppression mechanism, BMC can effectively filter out frequent switching of control strategy space caused by brief fluctuations in sensor status, reduce the probability of falsely triggering disaster recovery control, and improve the long-term operational stability and anti-interference capability of the server thermal management system.

[0164] The server temperature sensor disaster recovery control method provided in this application, compared with the existing technology that uses fixed strategy control based solely on the sampling results of a single temperature sensor, first identifies the in-situ status and temperature acquisition status of multiple temperature sensors, and combines the two types of statuses into a unified status identifier. This identifier simultaneously reflects both the physical in-situ status and data acquisition capability of the sensor, thereby differentiating and managing the sensor's operating status. Based on this, the system matches the disaster recovery control strategy space according to the status identifier. When the target temperature sensor simultaneously meets the abnormal conditions of both its in-situ status and acquisition status, it switches to the disaster recovery control strategy space, allowing the control strategy to adjust according to changes in sensor availability. Furthermore, the system determines associated temperature sensors based on the thermal region correlation of the target temperature sensor and uses their corresponding temperature data to generate alternative temperature input data. When the target sensor fails, this data serves as a temperature feedback source in the generation of the thermal regulation strategy, enabling the server to maintain a closed-loop thermal regulation even in the event of a single sensor failure.

[0165] Corresponding to the aforementioned embodiment of a server temperature sensor disaster recovery control method, this application also provides an embodiment of a server temperature sensor disaster recovery control device.

[0166] Figure 2 This is a schematic diagram of the structure of Embodiment 1 of the server temperature sensor disaster recovery control device provided in this application. Please refer to... Figure 2 The apparatus provided in this embodiment includes an identification module 201, a matching module 202, a determination module 203, and a generation module 204.

[0167] The identification module 201 is used to identify the presence status and temperature acquisition status of multiple temperature sensors.

[0168] The matching module 202 is used to match a status identifier based on the in-situ status and the temperature acquisition status. The same status identifier can identify both the in-situ status and the temperature acquisition status. Temperature sensors with completely identical in-situ status and temperature acquisition status have the same status identifier.

[0169] The matching module 202 is used to match the disaster recovery control strategy space based on the status identifier. When the on-site status and acquisition status of the target temperature sensor in the temperature sensor meet the corresponding abnormal triggering conditions, it enters the disaster recovery control strategy space of the target temperature sensor.

[0170] The determining module 203 is used to determine at least one associated temperature sensor based on the thermal region correlation of the target temperature sensor.

[0171] The determining module 203 is used to generate alternative temperature input data based on the temperature data corresponding to the associated temperature sensor.

[0172] The generation module 204 is used to take the alternative temperature input data as the temperature data after the target temperature sensor enters the disaster recovery control strategy space, and generate a server thermal regulation strategy in the disaster recovery control strategy space.

[0173] The apparatus of this embodiment can be used to perform... Figure 1 The steps of the method embodiment shown are similar in principle and process, and will not be repeated here.

[0174] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0175] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0176] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A disaster recovery control method for a server temperature sensor, characterized in that, include: Identify the presence and temperature acquisition status of multiple temperature sensors; Based on the matching status identifiers of the in-situ status and temperature acquisition status, the same status identifier can simultaneously identify both the in-situ status and the temperature acquisition status. Temperature sensors with completely identical in-situ status and temperature acquisition status have the same status identifier. Based on the state identifier matching disaster recovery control strategy space, when the in-situ state and acquisition state of the target temperature sensor in the temperature sensor meet the corresponding abnormal triggering conditions, it enters the disaster recovery control strategy space of the target temperature sensor. At least one associated temperature sensor is determined based on the thermal region correlation of the target temperature sensor; Generate alternative temperature input data based on the temperature data corresponding to the associated temperature sensor; The alternative temperature input data is used as the temperature data after the target temperature sensor enters the disaster recovery control strategy space, and a server thermal regulation strategy is generated in the disaster recovery control strategy space. When the in-situ status and temperature acquisition status of the target temperature sensor meet the abnormal triggering conditions, the system enters the disaster recovery control strategy space and generates a server thermal control strategy based on the alternative temperature input data. The method further includes: When the temperature sensor is in place and the temperature acquisition status is normal, it enters the normal control strategy space and generates a server thermal control strategy based on the temperature data output by the temperature sensor. When the target temperature sensor is in an abnormal state but the abnormal triggering condition is not met, the normal control strategy space is maintained, and the alternative temperature input data is used as the equivalent input of the target temperature sensor to participate in the generation of the server thermal control strategy. The thermal regulation strategy in the normal control strategy space is used to execute closed-loop thermal speed regulation control when the abnormal triggering conditions are not met. The thermal regulation strategy in the disaster recovery control strategy space is used to execute redundant thermal speed regulation control on the corresponding thermal area of ​​the target temperature sensor based on alternative temperature input data when the target temperature sensor meets the abnormal triggering conditions. The step of generating alternative temperature input data based on the temperature data corresponding to the associated temperature sensor includes: Acquire the historical temperature data sequence of the target temperature sensor within a preset time window; Based on the temperature data from the associated temperature sensor and the historical temperature data sequence, alternative temperature input data is generated through fusion calculation. The alternative temperature input data is used as the temperature input data in the disaster recovery control strategy space.

2. The method according to claim 1, characterized in that, The method for establishing and matching the status identifier includes: Obtain the preset set of in-situ status types and the set of temperature acquisition status types; Multiple state combinations are generated by arranging and combining the in-situ state type set and the temperature acquisition state type set. Based on preset mapping rules, corresponding state identifiers are generated for each state combination. Obtain the current actual location status and actual temperature acquisition status of the temperature sensor; Based on the actual in-situ status and the actual temperature acquisition status, a matching process is performed among the multiple status combinations to obtain the corresponding status identifier.

3. The method according to claim 1, characterized in that, The disaster recovery control strategy space entering the target temperature sensor includes: Based on continuous detection results, a state sequence of the target temperature sensor is constructed; Based on the state sequence, the number of consecutive occurrences of the target temperature sensor being in an abnormal in-situ state and the number of consecutive occurrences of it being in a normal in-situ state are counted. When the number of consecutive occurrences of an in-situ abnormal state reaches the first continuity threshold, it enters the disaster recovery control strategy space; When the number of consecutive occurrences of the in-place normal state reaches the second continuity threshold, exit the disaster recovery control strategy space and restore the normal control strategy space.

4. The method according to claim 3, characterized in that, The method for determining the first continuity threshold includes: Obtain historical server operation data, real-time operational load data, and historical anomaly statistics. Based on the historical operating data, the continuous distribution characteristics of the target temperature sensor’s continuous in-situ anomalies during the historical operating process are statistically analyzed, and a basic continuity threshold is determined based on the continuous distribution characteristics of the continuous in-situ anomalies. The basic continuity threshold is corrected by load weighting based on the real-time operating load data. Based on the fluctuation frequency of the historical abnormal statistical data within a preset time window, the modified basic continuity threshold is adjusted by adding or subtracting to obtain the first continuity threshold.

5. The method according to claim 1, characterized in that, When the thermal control strategy corresponds to multiple thermal zone control objects, the method further includes: Acquire multiple thermal zone speed control objects corresponding to the target temperature sensor; When the operation mode switching of the thermal control strategy is triggered, a unified strategy switching command is issued to the multiple thermal zone speed control objects. Based on the unified strategy switching command, each hot zone speed control object can complete the switching of the operating mode of the thermal regulation strategy within the same control cycle.

6. The method according to claim 1, characterized in that, The method further includes: Acquire information on the status changes of the target temperature sensor and the switching information of the disaster recovery control strategy space; During the disaster recovery control strategy space switching process, the state change information and the control strategy space switching information are sampled synchronously. The state change information obtained by synchronous sampling is associated and stored with the control strategy space switching information in chronological order to generate black box record data; The black box recording data is stored in a storage area independent of the BMC temperature control module.

7. The method according to claim 1, characterized in that, The method further includes: Acquire the state change sequence of the target temperature sensor within a preset time window; The number of times the state switches between normal and abnormal states is counted based on the state change sequence. When the number of switching operations meets the preset suppression trigger condition, the system enters the switching suppression state. In the switching suppression state, the switching operations of the control strategy space are restricted so that the number of switching operations of the control strategy space within the preset time window meets the preset stability constraint. After the preset time window ends, the number of switching times is reset to the initial state.

8. A server temperature sensor disaster recovery control device, characterized in that, The method for performing any one of claims 1 to 7 includes an identification module, a matching module, a determination module, and a generation module, wherein: The identification module is used to identify the presence status and temperature acquisition status of multiple temperature sensors. The matching module is used to match a status identifier based on the in-situ status and the temperature acquisition status. The same status identifier can identify both the in-situ status and the temperature acquisition status. Temperature sensors with completely identical in-situ status and temperature acquisition status have the same status identifier. The matching module is used to match the disaster recovery control strategy space based on the status identifier. When the on-site status and acquisition status of the target temperature sensor in the temperature sensor meet the corresponding abnormal triggering conditions, it enters the disaster recovery control strategy space of the target temperature sensor. The determining module is used to determine at least one associated temperature sensor based on the thermal region correlation of the target temperature sensor. The determining module is used to generate alternative temperature input data based on the temperature data corresponding to the associated temperature sensor. The generation module is used to take the alternative temperature input data as the temperature data after the target temperature sensor enters the disaster recovery control strategy space, and generate a server thermal regulation strategy in the disaster recovery control strategy space.

Citation Information

Patent Citations

  • Server heat dissipation control method, server, storage medium and program product

    CN118409643A

  • Disaster recovery sensor determination method and device, equipment and medium

    CN119249735A