Temperature sensor fault processing method, program product and electronic equipment

By obtaining the backup information and mapping relationship of the temperature sensor, determining the impact level, and adjusting the access cycle of the single board management controller, the system performance degradation caused by temperature sensor failure is solved, and the server's operating efficiency and the sensitivity of the fan speed regulation strategy are improved.

CN120523656AActive Publication Date: 2025-08-22INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511031161.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-08-22
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In the server, a temperature sensor failure causes the single board management controller to frequently access all sensors, reducing the fan speed regulation efficiency and affecting system performance.

Method used

By obtaining the backup information and mapping relationship of the temperature sensor, determining the impact level, and adjusting the access cycle of the single board management controller, optimizing the fan speed regulation strategy.

Benefits of technology

It improves the operating efficiency of the single-board management controller and the sensitivity of fan speed regulation, and optimizes the access efficiency and resource utilization of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523656A_ABST
    Figure CN120523656A_ABST
Patent Text Reader

Abstract

The invention discloses a temperature sensor fault processing method, a program product and an electronic device, the method is applied to a server, the server comprises a bus change-over switch, a single board management controller, a plurality of temperature sensors and a plurality of cooling fans, and the method comprises the following steps: at a target position in the server, the single board management controller is connected with the multiple temperature sensors; acquiring temperature information corresponding to the plurality of temperature sensors at the target position, and determining whether each temperature sensor fails or not based on the plurality of temperature information; under the condition that it is determined that the temperature sensor breaks down, backup information of the temperature sensor is obtained, and a mapping relation when the temperature sensor regulates and controls the rotating speed of the cooling fan is obtained; determining the influence level of the temperature sensor on the server based on the backup information and the mapping relation; and adjusting the access period of the single board management controller for accessing the temperature sensor based on the influence level. According to the method, the operation efficiency of the single-board management controller can be improved, and the sensitivity of fan speed regulation to the temperature sensor can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of temperature sensors, and in particular to a method for handling temperature sensor faults, a program product, and an electronic device. Background Art

[0002] As server equipment becomes increasingly complex, overall energy consumption also increases significantly, posing significant challenges for overall cooling. The industry utilizes two cooling technologies: fan cooling and liquid cooling. Liquid cooling is currently relatively expensive, so air cooling remains the preferred method for most manufacturers. Air cooling relies on temperature sensors. Storage devices are significantly more complex than servers, and the number of temperature sensors increases with board complexity. If a temperature sensor fails, frequent access increases the BMC's (Baseboard Management Controller) access cycle for all temperature sensors, reducing fan speed control efficiency. Summary of the Invention

[0003] The present application provides a temperature sensor fault handling method and program product, storage medium, and electronic device to at least solve the problem in the related art that if a sensor fails and continues to be accessed frequently, the cycle of the single-board management controller accessing all temperature sensors will increase, and the efficiency of fan speed regulation will be reduced. It can improve the operating efficiency of the single-board management controller and increase the sensitivity of fan speed regulation to the temperature sensor.

[0004] The present application provides a temperature sensor fault handling method, characterized in that it is applied in a server, the server including a bus switch, a single board management controller, multiple temperature sensors, and multiple cooling fans, the single board management controller being connected to the multiple temperature sensors via the bus switch, the method comprising: At a target location in the server, acquiring temperature information corresponding to a plurality of the temperature sensors at the target location, and determining whether each of the temperature sensors is faulty based on the plurality of temperature information; When it is determined that the temperature sensor fails, obtaining backup information of the temperature sensor and obtaining a mapping relationship between the temperature sensor and the speed control of the cooling fan; determining an impact level of the temperature sensor on the server based on the backup information and the mapping relationship; An access period for the board management controller to access the temperature sensor is adjusted based on the impact level.

[0005] The present application also provides a computer program product, including a computer program / instruction, which implements the above-mentioned temperature sensor fault processing method when executed by a processor.

[0006] The present application also provides a non-volatile computer-readable storage medium having a program stored thereon, which implements the above-mentioned temperature sensor fault processing method when executed by a processor.

[0007] The present application also provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned temperature sensor fault processing method is implemented.

[0008] Through this application, at a target location in a server, the corresponding temperature information of multiple temperature sensors at the target location is obtained, and based on the multiple temperature information, whether each temperature sensor is faulty is determined. If a temperature sensor is determined to be faulty, the backup information of the temperature sensor is obtained, and the mapping relationship between the temperature sensor and the speed control of the cooling fan is obtained. Based on the backup information and the mapping relationship, the impact level of the temperature sensor on the server is determined, and the access period of the single board management controller to the temperature sensor is adjusted based on the impact level. As a result, this method can improve the operating efficiency of the single board management controller and increase the sensitivity of fan speed control to the temperature sensor. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0010] Figure 1 This is a flow chart of a method for handling a temperature sensor failure according to one embodiment of the present application; Figure 2 Schematic diagram of a hardware control circuit according to one embodiment of the present application; Figure 3 A flow chart of a method for handling a temperature sensor failure according to a specific example of the present application; Figure 4 Schematic diagram of a block diagram of an electronic device according to an embodiment of the present application.

[0011] Reference numerals: 200 - electronic device, 210 - memory, 220 - processor. DETAILED DESCRIPTION

[0012] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0013] The following describes a temperature sensor fault processing method, a computer program product, a non-volatile computer-readable storage medium, and an electronic device proposed in embodiments of the present application with reference to the accompanying drawings.

[0014] Figure 1 Flowchart of a temperature sensor fault handling method according to an embodiment of the present application.

[0015] like Figure 1 As shown, the temperature sensor fault processing method of the embodiment of the present application may include the following steps: S1, at a target location in a server, obtaining temperature information corresponding to a plurality of temperature sensors at the target location, and determining whether each temperature sensor is faulty based on the plurality of temperature information.

[0016] S2: When it is determined that the temperature sensor fails, backup information of the temperature sensor is obtained, and a mapping relationship between the temperature sensor and the speed control of the cooling fan is obtained.

[0017] S3: Determine the impact level of the temperature sensor on the server based on the backup information and the mapping relationship.

[0018] S4: Adjust the access period of the board management controller to the temperature sensor based on the impact level.

[0019] Specifically, in an embodiment of the present application, a server may include a bus switch, a board management controller (BMC), multiple temperature sensors, and multiple cooling fans. Specifically, temperature sensors may be distributed at targeted locations (critical locations) within the server, such as the CPU (Central Processing Unit), memory, hard disk, power module, etc., to monitor the temperatures of these components in real time. A bus switch (e.g., an I2C-switch) connects the BMC and the temperature sensors. Furthermore, each temperature sensor can have a dedicated I2C-switch channel, which improves communication reliability and efficiency. The BMC is responsible for managing and controlling the server's hardware status, such as reading temperature sensors, detecting faults, and adjusting cooling strategies. Cooling fans are used to reduce the internal temperature of the server, ensuring that the server operates within a normal temperature range. Multiple cooling fans provide redundancy; even if a fan fails, the remaining fans can still maintain system cooling.

[0020] At a target location in the server, temperature information corresponding to multiple temperature sensors at the target location is first obtained. Based on the multiple temperature information, a determination is made as to whether each temperature sensor is faulty. For example, a board management controller (BMC) can periodically obtain temperature data from each temperature sensor via an I2C (Inter-Integrated Circuit) bus or other communication interface. These temperature sensors are distributed at target locations in the server, such as the CPU, memory, hard disk, and power supply module, to monitor the temperature of these components in real time. The obtained temperature data is then analyzed to determine whether the temperature sensor is faulty. Fault determination criteria may include, but are not limited to, the following: If temperature data cannot be obtained from a temperature sensor within a specified timeframe, or if the obtained data is invalid (e.g., outside the sensor's measurement range or with an incorrect data format), the sensor is considered faulty. If a temperature sensor repeatedly reads temperature values ​​significantly higher or lower than those of other sensors of the same type, and if the actual temperature changes are not excessively fast, the sensor is considered faulty. Under normal circumstances, the rate of change of internal server temperature is relatively stable. If the temperature change rate detected by a temperature sensor is significantly higher than the normal range, it may be a misreading caused by a sensor failure.

[0021] For example, a server has five temperature sensors, labeled T1, T2, T3, T4, and T5. The BMC collects temperature data from these sensors every one second. During one data collection session, the BMC discovered that it could not obtain temperature data from T3 after three consecutive attempts. Therefore, it determined that T3 was faulty. The BMC also found that T4's temperature was 100°C, while other sensors of the same type were around 50°C. Furthermore, the server's actual operating environment temperature was not high, so it determined that T4 might be faulty.

[0022] If a temperature sensor failure is determined, backup information for the temperature sensor can be obtained, along with the mapping between the temperature sensor and the cooling fan speed control. When obtaining backup information, a backup temperature sensor for the failed temperature sensor can be found based on a pre-defined backup policy. Backup temperature sensors can be of the same or different types, but their temperature data can be cross-referenced and verified. For example, the backup policy can be developed based on factors such as the sensor's location and the object being monitored. The mapping relationship refers to the correspondence between the temperature sensor output (temperature data) and the cooling fan speed. For example, if the temperature sensor reading exceeds a certain threshold, the cooling fan speed will increase; if the temperature sensor reading falls below a certain threshold, the cooling fan speed will decrease. The mapping between the temperature sensor and the cooling fan speed control involves the fan speed control policy. For example, the temperature values ​​of some temperature sensors may directly affect the cooling fan speed, while the temperature values ​​of other sensors may serve only as a reference. Therefore, it is important to clearly define the role and weight of each temperature sensor in the fan speed control policy. Therefore, obtaining backup information for temperature sensors provides more reference data for subsequent troubleshooting, helping to more accurately assess the impact of the fault on the system. And determining the mapping relationship between the temperature sensor and the speed control of the cooling fan will help to reasonably adjust the cooling strategy when a fault occurs to ensure that the cooling needs of the server are met.

[0023] After determining the backup information and mapping relationship, the impact level of the temperature sensor on the server can be determined based on the backup information and mapping relationship. For example, the impact level is determined based on the presence or absence of a backup temperature sensor. If a backup temperature sensor exists and the backup temperature sensor can completely replace the function of the faulty sensor, the impact level of the faulty sensor is low; if there is no backup temperature sensor, or the backup temperature sensor cannot completely replace the function of the faulty sensor, the impact level is high. For another example, the impact level can be determined based on the weight of the speed mapping relationship. For example, the greater the weight of the faulty sensor in the speed control of the cooling fan, the higher its impact level on the cooling system. For example, if the temperature value of a temperature sensor directly affects the full speed operation of the cooling fan, then its impact level will be very high. That is, based on the above factors and combined with pre-set rules, the impact level of each faulty temperature sensor is determined. The level can be determined in the form of tables, formulas, etc.

[0024] For example, after the BMC detects failures in T3 and T4, it searches for backup information. Assume that T3's backup temperature sensors are T1 and T2, both located in the same area of ​​the server and monitoring the temperature of the same component. T4, however, does not have a backup temperature sensor because it monitors a specific component and no other sensor can replace it. The BMC then obtains the mapping between T3 and T4 and the speed control of the cooling fan. According to the fan speed control policy, T3's temperature value accounts for 30% of the fan speed control weight, while T4's temperature value accounts for 20%. Based on the backup information and mapping, the impact levels of T3 and T4 are determined. Because T3 has backup temperature sensors T1 and T2, which provide reliable temperature data, and T3's weight in cooling fan speed control is 30%, the BMC determines T3's impact level as medium. T4, on the other hand, does not have a backup temperature sensor and its weight in cooling fan speed control is 20%. Although the weight is not high, since there is no backup, T4's impact level is determined to be high.

[0025] After determining the impact level, the BMC's access cycle for temperature sensors can be adjusted based on the impact level. For example, for temperature sensors with a high impact level, even if a fault occurs, the access cycle should be kept short to promptly obtain data from the backup temperature sensor and ensure the normal operation of the cooling system. For example, the access cycle could be shortened to 1 / 2 or 1 / 3 of the original time. For temperature sensors with a medium impact level, the access cycle can be appropriately extended, but not excessively, to ensure timely detection of backup temperature sensor failures. For example, the access cycle could be doubled. For temperature sensors with a low impact level, the access cycle can be significantly extended, or even suspended for a certain period of time to conserve system resources. For example, the access cycle could be extended 10 times the original time, or suspended if the backup temperature sensor data is normal for multiple consecutive times. In other words, the access cycle is dynamically adjusted based on real-time system operating status and fault conditions. If the backup temperature sensor data becomes abnormal or the cooling system operating status changes, the BMC can reassess the impact level and adjust the access cycle accordingly.

[0026] Therefore, by reasonably adjusting the access cycle, the number of times the BMC accesses the temperature sensor is reduced, the access efficiency of the system is improved, and the system load is reduced.

[0027] According to one embodiment of the present application, the impact level of the temperature sensor on the server is determined based on the backup information and the mapping relationship, including: when the backup information is that the temperature sensor has no corresponding backup temperature sensor, and if the temperature sensor fails, the cooling fan runs at the maximum speed, the impact level is determined to be the first level; when the backup information is that the temperature sensor has a corresponding backup temperature sensor, and if the temperature sensor and the backup temperature sensor are all faulty, the cooling fan runs at the maximum speed, the impact level is determined to be the second level; when the backup information is that the temperature sensor has a corresponding backup temperature sensor, and if the temperature sensor and the backup temperature sensor are all faulty, the cooling fan increases the operating speed to the first target speed threshold, the impact level is determined to be the third level, wherein the first target speed threshold is less than the maximum speed; when the backup information is that the temperature sensor has no corresponding backup temperature sensor, and if the temperature sensor fails, the cooling fan increases the operating speed to the second target speed threshold, the impact level is determined to be the fourth level, wherein the second target speed threshold is less than the maximum speed.

[0028] Specifically, when determining the impact level of a temperature sensor on a server based on backup information and mapping relationships, if the backup information indicates that there is no corresponding backup temperature sensor, and if the temperature sensor fails, the cooling fan will operate at maximum speed, the impact level can be determined to be Level 1. In this case, the temperature sensor failure will cause the cooling fan to run directly at maximum speed to prevent damage from overheating. Furthermore, since there is no backup temperature sensor, the system's cooling is completely dependent on the normal operation of this sensor, resulting in the highest impact level, Level 1.

[0029] If the backup information indicates that a temperature sensor has a corresponding backup temperature sensor, and if both the primary and backup temperature sensors fail, the cooling fan will continue to operate at maximum speed, the impact level is determined to be Level 2. In this case, even though a backup temperature sensor exists, if both the primary and backup temperature sensors fail, the cooling fan will still need to operate at maximum speed. In this case, the system's cooling relies on the normal operation of both sensors, resulting in a higher impact level of Level 2.

[0030] If the backup information indicates that a temperature sensor has a corresponding backup temperature sensor, and if both the primary and backup temperature sensors fail, the cooling fan will increase its speed to the first target speed threshold, the impact level is determined to be Level 3. In this case, even if both the primary and backup temperature sensors fail, the cooling fan will not directly operate at maximum speed. Instead, it will increase its speed to the first target speed threshold (which is lower than the maximum speed). This indicates that the system has some cooling redundancy, but a higher speed is still required to ensure safety. Therefore, the impact level is Level 3.

[0031] If the backup information indicates that there is no corresponding backup temperature sensor for the temperature sensor, and if the temperature sensor fails, the cooling fan will increase its operating speed to the second target speed threshold, the impact level can be determined to be Level 4. In this case, although there is no backup temperature sensor, the cooling fan will only increase its operating speed to the second target speed threshold (less than the maximum speed) when the sensor fails. This indicates that the sensor has a relatively small impact on system cooling, and therefore its impact level is Level 4.

[0032] Suppose a server has four temperature sensors, labeled T1, T2, T3, and T4. Their backup information and mapping relationships are as follows: For T1, the backup information indicates that there is no backup temperature sensor. Mapping relationship: If T1 fails, the cooling fan runs at maximum speed, and the impact level is determined to be Level 1. For T2, the backup information indicates that there is a backup temperature sensor, T5. Mapping relationship: If both T2 and T5 fail, the cooling fan runs at maximum speed, and the impact level is determined to be Level 2. For T3, the backup information indicates that there is a backup temperature sensor, T6. Mapping relationship: If both T3 and T6 fail, the cooling fan increases to the first target speed threshold (for example, 80% of the maximum speed). The impact level is determined to be Level 3. For T4, the backup information indicates that there is no backup temperature sensor. Mapping relationship: If T4 fails, the cooling fan increases to the second target speed threshold (for example, 60% of the maximum speed). The impact level is determined to be Level 4.

[0033] Therefore, by dividing the impact levels in detail, we can accurately assess the impact of each temperature sensor on the server, providing a scientific basis for subsequent troubleshooting and cooling strategy adjustments.

[0034] According to one embodiment of the present application, the access period of the single board management controller to the temperature sensor is adjusted based on the impact level, including: when the impact level is the first level or the fourth level, keeping the access period of the single board management controller to the temperature sensor unchanged at the initial access period; when the impact level is the second level or the third level, adjusting the access period of the single board management controller to the temperature sensor based on the failure condition of the backup temperature sensor.

[0035] Specifically, when adjusting the board management controller's access cycle to temperature sensors based on the impact level, the current impact level is determined. If the impact level is Level 1 or Level 4, the board management controller's access cycle to the temperature sensors can remain unchanged at the initial access cycle. This means that a Level 1 sensor failure could cause the cooling fan to run at maximum speed, necessitating frequent monitoring of its status to ensure system safety. While Level 4 sensors have less impact on the cooling system, they lack backup temperature points and still require regular monitoring to prevent potential issues.

[0036] If the impact level is Level 2 or Level 3, the BMC's access cycle to the temperature sensor can be adjusted based on the backup temperature sensor's fault condition. For example, if the backup temperature sensor is functioning properly, the access cycle can be appropriately extended, reducing the BMC's access frequency to conserve system resources. If the backup temperature sensor also fails, the access cycle remains unchanged or is shortened to ensure timely detection and resolution of the issue. This allows the BMC's access cycle to the temperature sensor to be dynamically adjusted based on the temperature sensor's impact level and backup information. This approach not only improves system access efficiency but also optimizes cooling control strategies, ensuring stable server operation under various conditions. Furthermore, by properly adjusting the access cycle, system resources can be conserved and overall system performance improved.

[0037] According to one embodiment of the present application, the access period of the single board management controller to the temperature sensor is adjusted based on the fault condition of the backup temperature sensor, including: when the backup temperature sensor is not faulty, closing the single board management controller's access to the temperature sensor; when there are multiple backup temperature sensors and all of the multiple backup temperature sensors are faulty, determining one backup temperature sensor from the multiple backup temperature sensors as the target backup temperature sensor, and keeping the access period of the single board management controller to the target temperature sensor unchanged at the initial access period, and closing the single board management controller's access to temperature sensors other than the target backup temperature sensor.

[0038] Specifically, when adjusting the access cycle of the single board management controller to the temperature sensor according to the fault condition of the backup temperature sensor, if the backup temperature sensor is not faulty, the BMC's access to the original temperature sensor can be closed. That is to say, in this case, the backup temperature sensor can completely replace the function of the original sensor, so there is no need to access the original sensor, thereby saving system resources and reducing access load.

[0039] If all backup temperature sensors fail, one is selected as the target backup temperature sensor. The BMC maintains its initial access cycle for the target backup temperature sensor and disables access to all other temperature sensors. In this scenario, even if all backup temperature sensors fail, selecting a single target backup temperature sensor for monitoring ensures the system can still obtain necessary temperature data while reducing access to other failed sensors and conserving resources.

[0040] Suppose a server has four temperature sensors, labeled T1, T2, T3, and T4. Their backup information is as follows: T1: Backup temperature sensor: T5. Failure scenario: T5 is normal. The access policy is: disable BMC access to T1 and only access T5. T2: Backup temperature sensors: T6 and T7. Failure scenario: T6 and T7 both fail. The access policy is: select T6 as the target backup temperature sensor, maintain the BMC access cycle to T6 at the initial access cycle (for example, once per second), and disable BMC access to T7. T3: Backup temperature sensor: T8. Failure scenario: T8 is normal. The access policy is: disable BMC access to T3 and only access T8. T4: Backup temperature sensors: T9 and T10. Failure scenario: T9 and T10 both fail. The access policy is: select T9 as the target backup temperature sensor, maintain the BMC access cycle to T9 at the initial access cycle (for example, once per second), and disable BMC access to T10.

[0041] Through the above steps, the BMC's access cycle to the temperature sensor can be dynamically adjusted based on the backup temperature sensor's fault status. This method not only saves system resources but also optimizes the access policy, ensuring effective monitoring of the temperature sensor status in all situations, improving overall system performance and reliability.

[0042] According to one embodiment of the present application, the temperature sensor failure handling method further includes: after a preset time, re-determining a backup temperature sensor from the plurality of backup temperature sensors as a target backup temperature sensor, maintaining the single board management controller's access period to the target temperature sensor at the initial access period, and disabling the single board management controller's access to temperature sensors other than the target backup temperature sensor, until all of the plurality of backup temperature sensors fail, and each of the plurality of backup temperature sensors is designated as the target backup temperature sensor. The preset time can be determined based on actual circumstances.

[0043] Specifically, in a server, the fault handling method for temperature sensors not only needs to deal with the current fault situation, but also needs to consider long-term stability and reliability. Therefore, in addition to selecting a target backup temperature sensor for monitoring when multiple backup temperature sensors all fail, it is also necessary to regularly re-evaluate and select the target backup temperature sensor. That is, after a preset time (for example, every 10 minutes, every hour, or a time interval set according to actual needs), re-evaluate the status of multiple backup temperature sensors. A new backup temperature sensor can be selected from multiple backup temperature sensors as the target backup temperature sensor, and the BMC's access cycle to the newly selected target backup temperature sensor remains unchanged at the initial access cycle, and the BMC's access to other backup temperature sensors other than the newly selected target backup temperature sensor is closed until each backup temperature sensor is used as the target backup temperature sensor at least once during the period when multiple backup temperature sensors all fail.

[0044] Suppose a server has four temperature sensors, labeled T1, T2, T3, and T4. Their backup information is as follows: T1: Backup temperature sensor: T5. Failure condition: T5 is normal. Access policy: Disable BMC access to T1 and only access T5. T2: Backup temperature sensors: T6 and T7. Failure condition: Both T6 and T7 fail. Access policy: Initially select T6 as the target backup temperature sensor. Maintain the BMC access cycle to T6 at the initial access cycle (for example, once per second). Disable BMC access to T7. Re-evaluate every 10 minutes and select T7 as the target backup temperature sensor. Maintain the BMC access cycle to T7 at the initial access cycle. Disable BMC access to T6. Repeat this process until both T6 and T7 have been selected as the target backup temperature sensor at least once. T3: Backup temperature sensor: T8. Failure condition: T8 is normal. Access policy: Disable BMC access to T3 and only access T8. T4: Backup temperature sensors: T9 and T10. Failure scenario: Both T9 and T10 fail. Access strategy: Initially select T9 as the target backup temperature sensor. Maintain the BMC's access cycle to T9 at the initial access cycle (for example, once per second). Disable BMC access to T10. Reevaluate every 10 minutes, select T10 as the target backup temperature sensor, maintain the BMC's access cycle to T10 at the initial access cycle, and disable BMC access to T9. Repeat this process until both T9 and T10 have been selected as the target backup temperature sensor at least once.

[0045] This allows the BMC to dynamically adjust its access cycle to the temperature sensor based on the backup temperature sensor's failure status. This approach not only saves system resources but also optimizes access policies, ensuring effective monitoring of temperature sensor status in all situations, improving overall system performance and reliability. Furthermore, by regularly re-evaluating and selecting target backup temperature sensors, system stability and reliability are further enhanced.

[0046] According to an embodiment of the present application, the temperature sensor fault handling method further includes: after the multiple backup temperature sensors return to normal, controlling the access cycle of the single board management controller to access the backup temperature sensors to be an initial access cycle.

[0047] Specifically, in a server, after multiple backup temperature sensors return to normal operation, the normal access policy for these sensors needs to be restored. This step is an important part of the temperature sensor fault handling method, ensuring that the system can effectively monitor temperature data after the sensors return to normal operation. In other words, the BMC can regularly check the status of each backup temperature sensor to determine whether they have returned to normal. This can be achieved by reading temperature data, checking communication status, etc. When it is detected that all backup temperature sensors have returned to normal, the BMC will restore the access cycle to the initial access cycle and re-enable access to all backup temperature sensors, ensuring that the system can fully monitor temperature data.

[0048] For example, the BMC periodically (e.g., once per second) obtains temperature data from each backup temperature sensor, checks whether the data is valid and whether communication is normal. If valid data is successfully obtained multiple times (e.g., three times) in a row, the backup temperature sensor is considered to have recovered. Once all backup temperature sensors have recovered, the BMC restores the access cycle to the initial access cycle and re-enables access to all backup temperature sensors, ensuring that the system can fully monitor temperature data. This allows the BMC to dynamically adjust the access cycle for temperature sensors based on the recovery status of backup temperature sensors. This approach not only improves system reliability but also optimizes resource utilization, ensuring comprehensive temperature data monitoring after the sensors have recovered, thereby improving overall system performance and stability. By promptly restoring the access policy, the system can quickly return to normal operation after the sensors have recovered, reducing misjudgments and erroneous operations caused by sensor failures.

[0049] According to an embodiment of the present application, each channel of the bus switch is individually connected to a temperature sensor, and the temperature sensor fault handling method further includes: controlling access of the board management controller to the temperature sensor based on the bus switch.

[0050] Specifically, if Figure 2As shown, each channel of the bus switch (I2C-SWITCH) is connected to the temperature sensor separately, allowing the BMC (i.e. Figure 2 The BMC chip in the system selectively accesses specific temperature sensors through a bus switch. This bus switch allows the BMC to flexibly control access to individual temperature sensors, reducing communication conflicts and improving communication efficiency.

[0051] That is, under normal circumstances, the BMC regularly accesses all temperature sensors through the bus switch to obtain temperature data. If a temperature sensor fails, the BMC stops accessing the sensor through the bus switch to avoid unnecessary communication attempts. If there is a backup temperature sensor, the BMC switches to the backup temperature sensor through the bus switch and continues to obtain temperature data. When the faulty sensor returns to normal, the BMC re-enables access to the sensor through the bus switch.

[0052] It should be noted that in Figure 2 In the Arduino Uno interface, the RST (Reset) pin is the reset signal or reset pin. It is used to initialize or reset a circuit, microcontroller, chip, or other digital component to a known initial state. A reset can be a power-on reset (occurring when the device is powered on) or a manual reset (triggered by pressing a button or sending a control signal). The RST pin is connected to the I2C-SWITCH and various temperature sensors. At system startup, the I2C-SWITCH can be reset to ensure it is in a known state and ready for communication. If the temperature sensor fails or requires resynchronization, the reset pin can be used to reset the sensor and restore normal operation.

[0053] Therefore, using a bus switch to control BMC access to temperature sensors effectively improves system flexibility and reliability. In the event of a temperature sensor failure, the bus switch quickly switches to the backup temperature sensor, ensuring the system can continuously monitor temperature data. Once the failed sensor returns to normal, the access policy is promptly restored, ensuring optimal utilization of system resources and improving overall system performance and stability. This approach not only improves system reliability but also optimizes resource utilization, ensuring effective monitoring of temperature sensor status in all situations.

[0054] According to an embodiment of the present application, the temperature sensor fault processing method further includes: determining a target processing method when the temperature sensor fault occurs based on the impact level.

[0055] Specifically, if a temperature sensor fails, the temperature of the corresponding monitoring point cannot be read, and the single board where the temperature sensor is located needs to be replaced. Storage devices carry the customer's data access business, and replacing a single board may cause business degradation. The replacement of important components requires suspending business before resuming business, which brings great inconvenience to customers. Currently, after the BMC reports a temperature sensor failure alarm, the single board where this sensor is located will be required to be replaced by the customer, which will affect the customer's business. However, subsequent technical analysis shows that this temperature sensor has little effect on the heat dissipation of the entire system. This is equivalent to the fact that the failure of this sensor has not reached the level of affecting the customer's business, and there is no value in replacing the single board. For the entire system, it is equivalent to a false alarm. For this reason, in the server of the present application, the target processing method for temperature sensor failure can also be determined based on the impact level. This method can take different processing measures according to sensor failures of different impact levels to ensure the stability and reliability of the system.

[0056] For example, for an impact level of level one, the fault can be reported immediately, and the system can be operated at maximum speed to ensure system safety. For an impact level of level two, the fault can be reported, the system can be switched to the backup temperature sensor, and the system can be operated at maximum speed. For an impact level of level three, the fault can be recorded, the system can be switched to the backup temperature sensor, and the system can be operated at the first target speed threshold. For an impact level of level four, the fault can be recorded, the system can be operated at the second target speed threshold without switching to the backup temperature sensor. Therefore, by determining the target handling method for temperature sensor failures based on the impact level, the reliability and stability of the server can be effectively improved. This method adopts different handling measures according to different impact levels, ensuring that the system can adopt the most appropriate response strategy for various fault situations. Through reasonable fault handling methods, not only is the system reliability improved, but resource utilization is also optimized, ensuring that the system can operate efficiently without affecting system safety.

[0057] Further, according to an embodiment of the present application, the target processing method when the temperature sensor fails is determined based on the impact level, including: when the impact level is the first level, reporting the temperature sensor failure; when the impact level is the second level, determining the target processing method based on the failure condition of the backup temperature sensor; when the impact level is the third level, determining the target processing method based on the failure condition of the backup temperature sensor and the speed of the cooling fan; when the impact level is the fourth level, recording the temperature sensor failure.

[0058] Specifically, when determining the target response to a temperature sensor failure based on the impact level, the current impact level is evaluated. For impact level 1, the temperature sensor failure can be directly reported. This means that a level 1 sensor failure has the greatest impact on system cooling, and therefore requires immediate reporting so that emergency measures can be taken. For impact level 2, the target response can be determined based on the failure status of the backup temperature sensor. This means that a level 2 sensor has a backup, but different response measures must still be taken based on the backup temperature sensor's status. For impact level 3, the target response can be determined based on the backup temperature sensor's failure status and the cooling fan speed. This means that a level 3 sensor has a backup, and the cooling fan speed will be increased to a lower target speed threshold when a failure occurs. The response must comprehensively consider the backup temperature sensor's status and the current fan speed. For impact level 4, only the temperature sensor failure needs to be recorded. This means that a level 4 sensor has a minimal impact on system cooling, and therefore, only the failure information needs to be recorded, without requiring immediate reporting or emergency measures.

[0059] For example, a server has four temperature sensors, labeled T1, T2, T3, and T4. Their backup information and impact levels are as follows: T1: Backup Information: No backup temperature sensor, Impact Level: Level 1. The target action is to immediately report the fault and run the cooling fan at maximum speed. T2: Backup Information: A backup temperature sensor T6 is present, Impact Level: Level 2. The target action is to monitor the status of backup temperature sensor T6. If T6 is normal, the fault is switched to T6, but the fault is not reported. If T6 is also faulty, the fault is reported and the cooling fan speed is increased to the maximum speed. T3: Backup Information: A backup temperature sensor T7 is present, Impact Level: Level 3. The target action is to monitor the status of backup temperature sensor T7. If T7 is normal, the fault is switched to T7, but the fault is not reported. If T7 is also faulty, the fault is recorded and the cooling fan speed is increased to the first target speed threshold (for example, 80% of the maximum speed). T4: Backup Information: No backup temperature sensor, Impact Level: Level 4. The target action is to record the fault and take no other measures.

[0060] Therefore, by determining the targeted handling method for temperature sensor failures based on the impact level, server reliability and stability can be effectively improved. This approach adopts different handling measures according to the impact level, ensuring that the system adopts the most appropriate response strategy for various failure scenarios. This rational fault handling method not only improves system reliability but also optimizes resource utilization, ensuring efficient system operation without compromising system security.

[0061] According to one embodiment of the present application, a target processing method is determined based on the fault condition of the backup temperature sensor, including: reporting the temperature sensor fault when all backup temperature sensors fail; and recording the temperature sensor fault when some backup temperature sensors fail.

[0062] Specifically, when determining the target processing method based on the failure status of the backup temperature sensors, the current fault situation is judged. If all backup temperature sensors fail, the temperature sensor failure can be reported. In other words, if all backup temperature sensors fail, the system cannot obtain reliable temperature data through the backup temperature sensors, so the failure needs to be reported immediately so that emergency measures can be taken, such as increasing the speed of the cooling fan to ensure that the system is not damaged by overheating. If there is a partial failure of the backup temperature sensors, the temperature sensor failure can be recorded. In other words, if some backup temperature sensors fail, but there are still other backup temperature sensors that can provide reliable temperature data, the system can continue to monitor the temperature, so only the fault information needs to be recorded and no immediate reporting is required. This can avoid unnecessary alarms and reduce the workload of operation and maintenance personnel.

[0063] Suppose a server has four temperature sensors, labeled T1, T2, T3, and T4. Their backup information is as follows: T2: Backup temperature sensors: T6 and T7. Failure: T6 is faulty, but T7 is normal. The target action is to record the failure of temperature sensor T2 and continue to use backup temperature sensor T7 to obtain temperature data. The cooling fan speed can also be appropriately increased to a lower threshold to ensure system operation within a safe range. T4: Backup temperature sensors: T9 and T10. Failure: T9 is faulty, but T10 is faulty. The target action is to report the failure of temperature sensor T4 and increase the cooling fan speed to its maximum to prevent system damage due to overheating.

[0064] Therefore, by determining targeted handling methods based on backup temperature sensor failure conditions, server reliability and stability can be effectively improved. This approach implements different handling measures based on the specific backup temperature sensor failure conditions, ensuring that the system can adopt the most appropriate response strategy for each failure scenario. This rational fault handling method not only improves system reliability but also optimizes resource utilization, ensuring efficient system operation without compromising system security. Furthermore, this approach reduces unnecessary alerts and improves operational efficiency.

[0065] According to one embodiment of the present application, a target processing method is determined based on the fault condition of the backup temperature sensor and the speed of the cooling fan, including: when there is a partial fault in the backup temperature sensor, recording the temperature sensor fault; when all the backup temperature sensors fail, obtaining the corresponding fan speed value determined based on the backup temperature sensor in the last fan speed regulation cycle, and obtaining the actual fan speed value at the end of the last fan speed regulation cycle; determining the target processing method based on the fan speed value and the actual fan speed value.

[0066] Further, according to an embodiment of the present application, a target processing method is determined based on the fan speed value and the actual fan speed value, including: when the fan speed value is less than the actual fan speed value, recording a temperature sensor failure; when the fan speed value is greater than or equal to the actual fan speed value, reporting a temperature sensor failure.

[0067] Specifically, when determining the target processing method based on the failure condition of the backup temperature sensor and the speed of the cooling fan, if there is a partial failure of the backup temperature sensor, the temperature sensor failure can be recorded. That is to say, if some backup temperature sensors fail, but there are still other backup temperature sensors that can provide reliable temperature data, the system can continue to monitor the temperature, so it is only necessary to record the failure information and no immediate reporting is required. This can avoid unnecessary alarms and reduce the workload of operation and maintenance personnel. If all backup temperature sensors fail, the corresponding fan speed value determined based on the backup temperature sensor in the last fan speed adjustment cycle can be obtained, and the actual fan speed value at the end of the last fan speed adjustment cycle can be obtained. For example, the fan speed value determined based on the backup temperature sensor in the last fan speed adjustment cycle can be obtained from the BMC record, and the actual fan speed value at the end of the last fan speed adjustment cycle can be obtained from the BMC record.

[0068] After obtaining the fan speed value and the actual fan speed value, the target handling method can be determined based on the fan speed value and the actual fan speed value. Specifically, the relationship between the current fan speed value and the actual fan speed value is determined. If the fan speed value is less than the actual fan speed value, a temperature sensor failure can be recorded. This indicates that this temperature sensor is not the critical temperature value for the previous round of fan speed regulation, and even if it fails, it will not trigger full fan speed. Therefore, the failure status can be unreported. If the fan speed value is greater than or equal to the actual fan speed value, a temperature sensor failure can be reported, indicating that the current heat dissipation capacity may not be sufficient to meet system requirements. The failure of the faulty temperature sensor may affect system heat dissipation. Therefore, the fault needs to be reported immediately so that emergency measures can be taken, such as increasing the cooling fan speed to prevent system damage due to overheating.

[0069] For example, assume there are 4 temperature sensors in the server, labeled as T1, T2, T3, and T4 respectively. Their backup information is as follows: T3: Backup temperature sensors: T7, T8. Fault situation: T7 is faulty, T8 is normal. It can be determined that the target handling method is to record the fault of temperature sensor T3 and continue to use the backup temperature sensor T8 to obtain temperature data. If necessary, appropriately increase the rotational speed of the cooling fan to a lower threshold to ensure the system operates within a safe range. T4: Backup temperature sensors: T9, T10. Fault situation: T9 is faulty, T10 is faulty. It can be determined that the target handling method is to record the fault of temperature sensor T4, obtain the fan rotational speed value PWM_last determined based on the backup temperature sensors in the previous fan speed regulation cycle, obtain the actual fan rotational speed value PWM_max at the end of the previous fan speed regulation cycle, compare PWM_last and PWM_max. If PWM_last < PWM_max, record the fault and keep the current fan rotational speed unchanged. If PWM_last >= PWM_max, report the fault and increase the rotational speed of the cooling fan to PWM_last.

[0070] Thus, by determining the target handling method based on the fan rotational speed value and the actual fan rotational speed value, the reliability and stability of the server can be effectively improved. This method takes different handling measures according to different backup temperature sensor fault situations and the current heat dissipation state, ensuring that the system can adopt the most appropriate response strategy in various fault situations. Through reasonable fault handling methods, not only the reliability of the system is improved, but also the resource utilization is optimized, ensuring that the system can operate efficiently without affecting system safety. At the same time, this method can reduce unnecessary alarms and improve the operation and maintenance efficiency.

[0071] In addition, in an embodiment of the present application, when determining the corresponding fan rotational speed value according to the backup temperature sensor, the fan rotational speed value can be determined through a pre-set corresponding relationship. For example, the relationship between the temperature obtained by the backup temperature sensor and the fan rotational speed value is determined in advance. After the temperature is determined, the corresponding relationship can be directly called to obtain the fan rotational speed value. And during the process of the sensor reading the temperature, if the temperature reading fails, clock signal repair (repair of timing error type) and chip reset repair (internal logic fault of the chip) can also be performed.

[0072] The following combines Figure 3 to describe the method of the present application.

[0073] As a specific example, the temperature sensor fault handling method of the present application may include the following steps: S101, at the target location in the server, obtain the temperature information corresponding to multiple temperature sensors at the target location, and determine whether each temperature sensor is faulty based on the multiple temperature information.

[0074] S102 : When it is determined that a temperature sensor fails, backup information of the temperature sensor is obtained, and a mapping relationship between the temperature sensor and the speed control of the cooling fan is obtained.

[0075] S103: Determine the impact level of the temperature sensor on the server based on the backup information and the mapping relationship.

[0076] S104: Determine whether the impact level is level 1 or level 4. If yes, go to step S105; if not, go to step S106.

[0077] S105: Keep the access period of the board management controller to the temperature sensor unchanged at the initial access period.

[0078] S106: Determine whether the impact level is the second level or the third level. If yes, go to step S107; if not, go to step S101.

[0079] S107, determine whether the backup temperature sensor is fault-free. If yes, go to step S108; if not, go to step S109.

[0080] S108: Close the board management controller's access to the temperature sensor.

[0081] S109 , determining a backup temperature sensor as a target backup temperature sensor, maintaining the board management controller's access cycle to the target temperature sensor at the initial access cycle, and closing the board management controller's access to temperature sensors other than the target backup temperature sensor.

[0082] In summary, according to the temperature sensor fault handling method of the embodiment of the present application, at the target location in the server, the temperature information corresponding to multiple temperature sensors at the target location is obtained, and based on the multiple temperature information, it is determined whether each temperature sensor is faulty. In the event that a temperature sensor is determined to be faulty, the backup information of the temperature sensor is obtained, and the mapping relationship between the temperature sensor and the speed control of the cooling fan is obtained. Based on the backup information and the mapping relationship, the impact level of the temperature sensor on the server is determined, and the access cycle of the single board management controller to the temperature sensor is adjusted based on the impact level. Therefore, this method can improve the operating efficiency of the single board management controller and increase the sensitivity of the fan speed control to the temperature sensor.

[0083] Corresponding to the above embodiments, the present application also proposes a computer program product.

[0084] A computer program product according to an embodiment of the present application includes a computer program / instruction, and when the computer program / instruction is executed by a processor, the above-mentioned temperature sensor fault processing method is implemented.

[0085] According to the computer program product of the embodiment of the present application, by executing the above-mentioned temperature sensor fault processing method, the operating efficiency of the single board management controller can be improved and the sensitivity of fan speed regulation to the temperature sensor can be improved.

[0086] Corresponding to the above embodiment, the present application also proposes a non-volatile computer-readable storage medium.

[0087] The non-volatile computer-readable storage medium of the embodiment of the present application stores a program, which, when executed by a processor, implements the above-mentioned temperature sensor fault processing method.

[0088] According to the non-volatile computer-readable storage medium of the embodiment of the present application, by executing the above-mentioned temperature sensor fault processing method, the operating efficiency of the single board management controller can be improved and the sensitivity of fan speed regulation to the temperature sensor can be improved.

[0089] Corresponding to the above embodiment, the present application also proposes an electronic device.

[0090] like Figure 4 As shown, the electronic device 200 of an embodiment of the present application may include: a memory 210, a processor 220, and a program stored in the memory 210 and executable on the processor 220. When the processor 220 executes the program, the above-mentioned temperature sensor fault handling method is implemented.

[0091] According to the electronic device of the embodiment of the present application, by executing the above-mentioned temperature sensor fault processing method, the operating efficiency of the single board management controller can be improved and the sensitivity of fan speed regulation to the temperature sensor can be improved.

[0092] It should be noted that the logic and / or steps represented in flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0093] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0094] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0095] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0096] In this application, unless otherwise specified or limited, the terms "installed," "connected," "connect," "fixed," etc. should be understood in a broad sense. For example, they can refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two elements or interaction between two elements, unless otherwise specified. Those skilled in the art will understand the specific meanings of the above terms in this application based on specific circumstances.

[0097] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A method for handling temperature sensor failure, characterized in that: The method is applied in a server, the server including a bus switch, a single board management controller, multiple temperature sensors, and multiple cooling fans, the single board management controller being connected to the multiple temperature sensors via the bus switch, and the method including: At a target location in the server, acquiring temperature information corresponding to a plurality of the temperature sensors at the target location, and determining whether each of the temperature sensors is faulty based on the plurality of temperature information; When it is determined that the temperature sensor fails, obtaining backup information of the temperature sensor and obtaining a mapping relationship between the temperature sensor and the speed control of the cooling fan; determining an impact level of the temperature sensor on the server based on the backup information and the mapping relationship; An access period for the board management controller to access the temperature sensor is adjusted based on the impact level.

2. The temperature sensor fault processing method according to claim 1, characterized in that: The determining, based on the backup information and the mapping relationship, an impact level of the temperature sensor on the server includes: When the backup information indicates that the temperature sensor has no corresponding backup temperature sensor, and if the temperature sensor fails, the cooling fan runs at a maximum speed, determining that the impact level is the first level; If the backup information indicates that the temperature sensor has a corresponding backup temperature sensor, and if both the temperature sensor and the backup temperature sensor fail, the cooling fan operates at a maximum speed, determining that the impact level is the second level; The impact level is determined to be the third level if the backup information indicates that the temperature sensor has a corresponding backup temperature sensor, and if both the temperature sensor and the backup temperature sensor fail, the cooling fan increases its operating speed to a first target speed threshold, wherein the first target speed threshold is less than the maximum speed; When the backup information indicates that the temperature sensor has no corresponding backup temperature sensor, and if the temperature sensor fails, the cooling fan increases the operating speed to the second target speed threshold, the impact level is determined to be the fourth level, wherein the second target speed threshold is less than the maximum speed.

3. The temperature sensor fault processing method according to claim 2, characterized in that: The adjusting, based on the impact level, an access period of the board management controller to the temperature sensor includes: When the impact level is the first level or the fourth level, maintaining the access period of the board management controller to the temperature sensor at the initial access period; When the impact level is the second level or the third level, an access period of the board management controller to the temperature sensor is adjusted based on the fault condition of the backup temperature sensor.

4. The temperature sensor fault processing method according to claim 3, characterized in that: The adjusting, based on a fault condition of the backup temperature sensor, an access cycle of the board management controller to the temperature sensor includes: When the backup temperature sensor is not faulty, shutting down the board management controller's access to the temperature sensor; In the case that there are multiple backup temperature sensors and all of the backup temperature sensors fail, one backup temperature sensor is determined from the multiple backup temperature sensors as the target backup temperature sensor, and the access cycle of the single board management controller to the target temperature sensor is kept unchanged at the initial access cycle, and the access of the single board management controller to the temperature sensors other than the target backup temperature sensor is closed.

5. The temperature sensor fault processing method according to claim 4, characterized in that: The method further comprises: After a preset time, a backup temperature sensor is re-determined from the multiple backup temperature sensors as the target backup temperature sensor, and the access period of the single board management controller to the target temperature sensor is kept unchanged at the initial access period, and the access of the single board management controller to the temperature sensors other than the target backup temperature sensor is closed until all the backup temperature sensors fail, and each of the multiple backup temperature sensors is used as the target backup temperature sensor.

6. The temperature sensor fault processing method according to claim 5, characterized in that: The method further comprises: After the plurality of backup temperature sensors return to normal, the access period of the board management controller to the backup temperature sensors is controlled to be an initial access period.

7. The temperature sensor fault handling method according to any one of claims 3 to 6, characterized in that: Each channel of the bus switch is individually connected to the temperature sensor, and the method further includes: The access of the board management controller to the temperature sensor is controlled based on the bus switch.

8. The temperature sensor fault processing method according to claim 2, characterized in that: The method further comprises: A target processing method when the temperature sensor fails is determined based on the impact level.

9. The temperature sensor fault processing method according to claim 8, characterized in that: The determining, based on the impact level, a target processing method when the temperature sensor fails, includes: When the impact level is the first level, reporting a temperature sensor failure; When the impact level is the second level, determining a target processing method based on the fault condition of the backup temperature sensor; When the impact level is the third level, determining a target processing method based on the fault condition of the backup temperature sensor and the rotation speed of the cooling fan; When the impact level is the fourth level, a temperature sensor failure is recorded.

10. The temperature sensor fault processing method according to claim 9, characterized in that: The determining of a target processing method based on the fault condition of the backup temperature sensor includes: In the event that all backup temperature sensors fail, reporting a temperature sensor failure; In the event that there is a partial failure of the backup temperature sensor, a temperature sensing failure is recorded.

11. The temperature sensor fault processing method according to claim 9, characterized in that: The determining of a target processing method based on the fault condition of the backup temperature sensor and the rotation speed of the cooling fan includes: In the event that a partial failure occurs in the backup temperature sensor, recording the temperature sensor failure; In the event that all of the backup temperature sensors fail, obtaining the corresponding fan speed value determined based on the backup temperature sensors in the last fan speed adjustment cycle, and obtaining the actual fan speed value at the end of the last fan speed adjustment cycle; The target processing mode is determined based on the fan speed value and the actual fan speed value.

12. The temperature sensor fault processing method according to claim 11, characterized in that: The determining the target processing mode based on the fan speed value and the actual fan speed value includes: When the fan speed value is less than the actual fan speed value, a temperature sensor failure is recorded; When the fan speed value is greater than or equal to the actual fan speed value, a temperature sensor failure is reported.

13. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the temperature sensor fault processing method according to any one of claims 1 to 12 is implemented.

14. A non-volatile computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the temperature sensor fault processing method according to any one of claims 1 to 12 is implemented.

15. An electronic device, characterized in that: include: A memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the temperature sensor fault processing method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Temperature sensor fault control method and device and low-temperature storage equipment

    CN119334067A

  • Server fan fault processing method and device, equipment and medium

    CN119356979A

  • Server cooperative control method, storage medium and electronic equipment

    CN119917350A

  • Storage chip high and low temperature aging test chamber fault self-diagnosis system and method and computer equipment

    CN120353683A

  • Multi-node system-fan-control switch

    EP3442319A1

Cited By

  • A method and system for dynamic control of abnormal temperature components inside a server.

    CN122411727A