A fan fault detection method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202410044282.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-01-11
AI Technical Summary
[0003]当风扇遭遇故障时,由于基板管理控制器驱动风扇的理论脉冲宽度调制值及风扇的反馈转速值均不能详细反映风扇的实际工作情况,因此相关技术仅能对风扇进行下电检测
[0052]可见,本发明首先可利用服务器中的正常风扇构建转换模型,其中该转换模型记录有正常风扇的表面振动值、正常风扇所在风道的风道温度值、服务器所在机房的机房温度值三者与正常风扇所接收的脉冲宽度调制值间的转换关系;随后,在基板管理控制器检测到故障风扇时,本发明可将故障风扇的表面振动值、故障风扇所在风道的风道温度值、故障风扇所属服务器所在机房的机房温度值共同输入转换模型,得到故障风扇所接收的实际脉冲宽度调制值。由于实际脉冲宽度调制值能够详细反映风扇的实际工作情况,因此本发明可根据实际脉冲宽度调制值、基板管理控制器驱动故障风扇的理论脉冲宽度调制值及故障风扇反馈的反馈转速值确定故障风扇的故障原因,进而可在不对风扇进行下电、不破坏故障现场的情况下分析故障风扇的故障原因,从而可提升风扇故障检测的便捷性。本发明还提供一种风扇故障检测装置、电子设备及计算机可读存储介质,具有上述有益效果。
Smart Images

Figure CN117703810B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of servers, and in particular to a fan failure detection method, apparatus, electronic device, and storage medium. Background Technology
[0002] Fans are crucial heat dissipation components in servers, and they are typically controlled by the server's baseboard management controller via pulse width modulation (PWM) signals. Additionally, fans can feed back their rotational speed to the baseboard management controller.
[0003] When a fan malfunctions, the theoretical pulse width modulation value driven by the board management controller and the fan's feedback speed value cannot accurately reflect the actual operating condition of the fan. Therefore, related technologies can only perform power-down testing on the fan. However, power-down testing destroys the fault scene, making it difficult to analyze the specific cause of the fan failure. Summary of the Invention
[0004] The purpose of this invention is to provide a fan fault detection method, device, electronic device, and storage medium, which can determine the actual pulse width modulation value of the faulty fan based on the computer room temperature value, the surface vibration value of the faulty fan, and the air duct temperature value, so that the cause of the fan fault can be analyzed by combining the actual pulse width modulation value without powering off.
[0005] To solve the above-mentioned technical problems, the present invention provides a fan fault detection method, comprising:
[0006] A conversion model is constructed using a normal fan in the server; the conversion model records the conversion relationship between the surface vibration value of the normal fan, the air duct temperature value of the air duct where the normal fan is located, the computer room temperature value of the server room, and the pulse width modulation value received by the normal fan.
[0007] When the baseboard management controller detects a faulty fan, it inputs the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the room temperature value of the server to which the faulty fan belongs into the conversion model to obtain the actual pulse width modulation value received by the faulty fan.
[0008] The cause of the malfunction of the fan is determined based on the actual pulse width modulation value, the theoretical pulse width modulation value driven by the baseboard management controller, and the feedback speed value fed back by the malfunctioning fan.
[0009] Optionally, the step of constructing a conversion model using normal fans in the server includes:
[0010] A vibration sensor is installed on the surface of the normal fan.
[0011] The airflow location of the normal fan inside the server is determined using heat dissipation simulation software, and the temperature sensor located at the airflow location is also identified.
[0012] The computer room temperature value is set to an initial value. The baseboard management controller is controlled to adjust the pulse width modulation value of the normal fan by controlling the server to run stress test software. During the adjustment process, the vibration sensor and the temperature sensor are used to continuously record the surface vibration value and air duct temperature value of the normal fan, so as to obtain the surface vibration value and air duct temperature value corresponding to different pulse width modulation values of the normal fan under the current computer room temperature value.
[0013] The process involves repeatedly adjusting the room temperature value and, after each adjustment, controlling the baseboard management controller to adjust the pulse width modulation value of the normal fan by controlling the server to run stress test software. During the adjustment process, the vibration sensor and the temperature sensor are used to continuously record the surface vibration value and air duct temperature value of the normal fan. This process yields the surface vibration value and air duct temperature value corresponding to different pulse width modulation values of the normal fan under different room temperatures.
[0014] The conversion model is constructed using the surface vibration values and duct temperature values corresponding to different pulse width modulation values of the normal fan under different computer room temperatures.
[0015] Optionally, after installing a vibration sensor on the surface of the normal fan, the system further includes:
[0016] Obtain the bus address of the vibration sensor on the two-wire serial bus of the baseboard management controller, and establish a first correspondence between the bus addresses of the normal fan and the vibration sensor;
[0017] After determining the temperature sensor located in the air duct, the following is also included:
[0018] Obtain the bus address of the temperature sensor on the two-wire serial bus of the baseboard management controller, and establish a second correspondence between the bus address of the normal fan and the temperature sensor;
[0019] Before inputting the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the room temperature value of the server to which the faulty fan belongs into the conversion model, the following steps are also included:
[0020] The bus address of the vibration sensor corresponding to the faulty fan is found according to the first correspondence of the faulty fan, and the surface vibration value of the faulty fan is obtained from the vibration sensor corresponding to the faulty fan according to the bus address of the vibration sensor corresponding to the faulty fan.
[0021] The bus address of the temperature sensor corresponding to the faulty fan is found according to the second correspondence of the faulty fan, and the duct temperature value of the faulty fan is obtained from the temperature sensor corresponding to the faulty fan according to the bus address of the temperature sensor corresponding to the faulty fan.
[0022] Optionally, the temperature sensor used to determine the location of the air duct includes:
[0023] Identify the temperature sensor located in the air duct and installed on the server motherboard;
[0024] And / or, determine the location of the external temperature sensor in the air duct.
[0025] Optionally, the step of constructing a conversion model using normal fans in the server includes:
[0026] The conversion model is constructed by using normal fans from multiple servers.
[0027] Optionally, the faulty fan is a dual-rotor fan, and determining the cause of the faulty fan's failure based on the actual pulse width modulation value, the theoretical pulse width modulation value driven by the board management controller, and the feedback speed value fed back by the faulty fan includes:
[0028] When it is determined that the faulty fan has a single rotor fault, it is determined whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value.
[0029] When it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value, the cause of the fault is determined to be that the faulty rotor is erroneously responding to the theoretical pulse width modulation value.
[0030] When it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value, the baseboard management controller is controlled to increase the theoretical pulse width modulation value, and during the increase process, it is determined whether the feedback speed value of the faulty rotor can be increased.
[0031] When it is determined that the feedback speed value of the faulty rotor can be increased, the cause of the fault is determined to be that the baseboard management controller erroneously interprets the feedback speed value of the faulty rotor.
[0032] When it is determined that the feedback speed value of the faulty rotor cannot be increased, the firmware of the baseboard management controller is replaced. After replacement, the baseboard management controller is re-controlled to increase the theoretical pulse width modulation value, and the feedback speed value is judged during the increase process.
[0033] When it is determined that the feedback speed value of the faulty rotor can be increased after replacing the firmware of the baseboard management controller, the cause of the fault is determined to be a fault in the original firmware of the baseboard management controller.
[0034] When it is determined that the feedback speed value of the faulty rotor cannot be increased after the firmware of the baseboard management controller is replaced, the cause of the fault is determined to be a transmission link failure between the faulty fan and the baseboard management controller.
[0035] Optionally, determining the cause of the malfunction of the fan based on the actual pulse width modulation value, the theoretical pulse width modulation value driven by the board management controller, and the feedback speed value fed back by the malfunctioning fan includes:
[0036] When the baseboard management controller detects a dual-rotor fault in the faulty fan, it determines for each faulty rotor whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value.
[0037] When it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value, the step of controlling the baseboard management controller to increase the theoretical pulse width modulation value is entered, and during the increase process, it is determined whether the feedback speed value of the faulty rotor can be increased.
[0038] When it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value, the baseboard management controller is controlled to increase the theoretical pulse width modulation value, and during the increase process, it is determined whether the actual pulse width modulation value of the faulty rotor can be increased.
[0039] When it is determined that the actual pulse width modulation value of the faulty rotor can be increased, the cause of the fault is determined to be that the faulty rotor is erroneously responding to the theoretical pulse width modulation value.
[0040] When it is determined that the actual pulse width modulation value of the faulty rotor cannot be increased, the firmware of the baseboard management controller is replaced, and after the replacement, it is re-evaluated whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value.
[0041] When it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value after the firmware of the baseboard management controller is replaced, the cause of the fault is determined to be a fault in the original firmware of the baseboard management controller.
[0042] When it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value after the firmware of the baseboard management controller is replaced, the cause of the fault is determined to be a transmission link failure between the faulty fan and the baseboard management controller.
[0043] The present invention also provides a fan fault detection device, comprising:
[0044] The model building module is used to build a conversion model using a normal fan in the server. The conversion model records the conversion relationship between the surface vibration value of the normal fan, the air duct temperature value of the air duct where the normal fan is located, the computer room temperature value of the server room, and the pulse width modulation value received by the normal fan.
[0045] The numerical detection module is used to input the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the computer room temperature value of the computer room where the server to which the faulty fan belongs are all into the conversion model when the baseboard management controller detects a faulty fan, so as to obtain the actual pulse width modulation value received by the faulty fan.
[0046] The fault detection module is used to determine the cause of the fault of the faulty fan based on the actual pulse width modulation value, the theoretical pulse width modulation value driven by the baseboard management controller to drive the faulty fan, and the feedback speed value fed back by the faulty fan.
[0047] The present invention also provides an electronic device, comprising:
[0048] Memory, used to store computer programs;
[0049] A processor is used to implement the fan failure detection method as described above when executing the computer program.
[0050] The present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when loaded and executed by a processor, implement the fan fault detection method described above.
[0051] This invention provides a fan fault detection method, comprising: constructing a conversion model using a normal fan in a server; the conversion model recording the conversion relationship between the surface vibration value of the normal fan, the air duct temperature value of the air duct where the normal fan is located, the server room temperature value of the server room, and the pulse width modulation value received by the normal fan; when the baseboard management controller detects a faulty fan, inputting the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the server room temperature value of the server room where the faulty fan is located into the conversion model to obtain the actual pulse width modulation value received by the faulty fan; determining the cause of the faulty fan based on the actual pulse width modulation value, the theoretical pulse width modulation value driven by the baseboard management controller to the faulty fan, and the feedback speed value fed back by the faulty fan.
[0052] As can be seen, this invention first utilizes a normal fan in a server to construct a conversion model. This model records the conversion relationship between the surface vibration value of the normal fan, the air duct temperature of the normal fan, the server room temperature, and the pulse width modulation (PWM) value received by the normal fan. Subsequently, when the baseboard management controller detects a faulty fan, this invention inputs the surface vibration value of the faulty fan, the air duct temperature of the faulty fan, and the server room temperature of the server to which the faulty fan belongs into the conversion model to obtain the actual PWM value received by the faulty fan. Since the actual PWM value can reflect the actual operating condition of the fan in detail, this invention can determine the cause of the faulty fan based on the actual PWM value, the theoretical PWM value driven by the baseboard management controller, and the feedback speed value of the faulty fan. This allows for analysis of the cause of the faulty fan without powering down the fan or damaging the fault scene, thereby improving the convenience of fan fault detection. This invention also provides a fan fault detection device, electronic device, and computer-readable storage medium, which have the above-mentioned beneficial effects. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0054] Figure 1 A flowchart of a fan fault detection method provided in an embodiment of the present invention;
[0055] Figure 2This is a schematic diagram of a server fan fault detection hardware provided in an embodiment of the present invention;
[0056] Figure 3 This is a schematic diagram illustrating a normal machine fan temperature and vibration parameter acquisition process provided in an embodiment of the present invention.
[0057] Figure 4 This is a schematic diagram of a fault detection and anomaly judgment process for a single rotor of a server fan provided in an embodiment of the present invention;
[0058] Figure 5 This is a schematic diagram of a fault detection and anomaly judgment process for a dual-rotor server fan provided in an embodiment of the present invention;
[0059] Figure 6 This is a structural block diagram of a fan fault detection device provided in an embodiment of the present invention;
[0060] Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] Fans are crucial heat dissipation components in servers, typically controlled by the server's baseboard management controller (BMC) via pulse width modulation (PWM) signals. Additionally, the fan can feed back its rotational speed (Tach) value to the BMC. In related technologies, fan malfunctions are usually detected by the BMC. Specifically, the BMC can identify a fan malfunction when it determines that the fan's feedback rotational speed value does not match the theoretical rotational speed value corresponding to the theoretical PWM signal driving the fan. However, neither the feedback rotational speed value nor the theoretical PWM value, nor the fan itself, can accurately reflect the fan's actual operating condition. For example, the feedback rotational speed value may not match the actual fan speed, and the theoretical PWM value may not match the actual PWM value received by the fan. Therefore, relying solely on the feedback rotational speed value and the theoretical PWM value is insufficient for detailed analysis of fan malfunctions, leading to the reliance on power-down testing. However, power-down testing destroys the fault scene, hindering fault reproduction and further complicating the analysis of the fan's cause. In view of this, the present invention provides a fan fault detection method, which can determine the actual pulse width modulation value of the faulty fan based on the computer room temperature value, the surface vibration value of the faulty fan and the air duct temperature value, so that the cause of the fan fault can be analyzed by combining the actual pulse width modulation value without powering off.
[0063] It should be noted that the embodiments of the present invention do not limit the subject of execution of this method. For example, it can be the server itself, or other electronic devices other than the server, and can be set according to actual application requirements.
[0064] Please refer to Figure 1 , Figure 1 A flowchart of a fan fault detection method provided in an embodiment of the present invention is shown. The method may include:
[0065] S101. Construct a conversion model using a normal fan in the server; the conversion model records the conversion relationship between the surface vibration value of the normal fan, the air duct temperature value of the air duct where the normal fan is located, the server room temperature value of the server room, and the pulse width modulation value received by the normal fan.
[0066] Unlike related technologies that cannot obtain the actual fan speed or pulse width modulation (PWM) value when the fan is powered on, this invention can determine the actual PWM value of the fan by using three factors: the fan's surface vibration value, the air duct temperature, and the room temperature of the server. This allows the actual PWM value to be obtained without affecting the fan's current operating state. In this invention, the air duct is a location within the server whose temperature is significantly affected by the measured fan speed. To achieve this, this invention can pre-construct a conversion model using a normal fan in the server. This model records the conversion relationship between the normal fan's surface vibration value, the air duct temperature, the room temperature, and the PWM value received by the normal fan. After mastering this conversion relationship, the embodiments of the present invention can input the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the computer room temperature value of the server to which the faulty fan belongs into the conversion model when a faulty fan is detected. This allows for quick and convenient acquisition of the actual pulse width modulation value received by the faulty fan. Based on the actual pulse width modulation value, the actual working condition of the faulty fan can be understood, and the cause of the faulty fan can be analyzed in detail by combining the actual pulse width modulation value.
[0067] It is understood that surface vibration values, duct temperature values, and server room temperature values can all be detected using relevant sensors. In other words, to detect the actual pulse width modulation value of the fan, a vibration sensor needs to be installed on the fan surface, a temperature sensor needs to be installed in the fan duct, and a temperature sensor needs to be installed outside the server. It should be noted that this embodiment of the invention does not limit whether these sensors need to be fixed inside the server; they can be fixed inside the server or installed when measurement is required, depending on the actual application requirements. This embodiment of the invention also does not limit the location and number of vibration sensors on the fan surface, which can be set according to actual application requirements. This embodiment of the invention also does not limit the location and number of temperature sensors inside the server, as long as it ensures that the temperature sensors can measure the duct temperature and ambient temperature values. It is worth noting that actual measurements and simulations can be used to determine whether the existing temperature sensors in the server can measure the fan duct temperature value, thus allowing duct temperature measurement based on the existing temperature sensors in the server, avoiding the need for additional temperature sensors. Furthermore, for a normal fan, the actual pulse width modulation value it receives should be equal to the theoretical pulse width modulation value sent to it by the board management controller. Therefore, the theoretical pulse width modulation value can be directly used as the actual pulse width modulation value received by the normal fan.
[0068] Furthermore, the embodiments of the present invention do not limit the construction method of the conversion model. For example, a conversion table can be constructed based on the surface vibration value of a normal fan, the air duct temperature value of the air duct where the normal fan is located, the computer room temperature value of the server where the normal fan is located, and the pulse width modulation value received by the normal fan, and the conversion table can be used as the conversion model. Alternatively, mathematical modeling can be performed based on the surface vibration value of a normal fan, the air duct temperature value of the air duct where the normal fan is located, the computer room temperature value of the server where the normal fan is located, and the pulse width modulation value received by the normal fan, and the mathematical modeling result can be used as the conversion model. It can be set according to the actual application requirements.
[0069] Furthermore, to improve the reliability of measurement data and avoid errors, this embodiment of the invention utilizes normal fans in multiple servers to jointly construct a conversion model. Specifically, the conversion model is constructed based on the surface vibration value of the normal fan measured in different servers, the air duct temperature value of the normal fan, the computer room temperature value of the server to which the normal fan belongs, and the pulse width modulation value received by the normal fan. This can improve the reliability of the conversion model.
[0070] Based on this, building a conversion model using normal fans in the server can include:
[0071] Step 11: Build a conversion model using normal fans from multiple servers.
[0072] S102. When the baseboard management controller detects a faulty fan, it inputs the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the room temperature value of the server to which the faulty fan belongs into the conversion model to obtain the actual pulse width modulation value received by the faulty fan.
[0073] It should be noted that the baseboard management controller will detect fan malfunctions in a conventional manner. Specifically, the baseboard management controller determines a fan malfunction when it finds that the fan's feedback speed value does not correspond to the theoretical speed value corresponding to the theoretical pulse width modulation (PWM) value emitted by the baseboard management controller. After detecting a malfunctioning fan, to conduct a detailed analysis of its cause, this embodiment of the invention can collect the surface vibration value of the malfunctioning fan, the air duct temperature value of the malfunctioning fan's duct, and the server room temperature value of the server to which the malfunctioning fan belongs. These three values are then input into a pre-constructed conversion model to obtain the actual PWM value received by the malfunctioning fan. This allows for further analysis of the cause of the malfunction based on the actual PWM value.
[0074] S103. Determine the cause of the faulty fan based on the actual pulse width modulation value, the theoretical pulse width modulation value of the board management controller driving the faulty fan, and the feedback speed value of the faulty fan.
[0075] Unlike related technologies, in analyzing the cause of a fault, this embodiment of the invention, in addition to using the theoretical pulse width modulation (PWM) value of the faulty fan driven by the baseboard management controller and the feedback speed value of the faulty fan, can also use the actual PWM value of the faulty fan, thereby enabling a more detailed analysis. For example, it is impossible to determine whether the fault lies in the PWM drive section of the fan or the speed feedback section based solely on the theoretical PWM value and the feedback speed value. However, by further combining the actual PWM value, when it is determined that the theoretical PWM value and the actual PWM value are not equal, it can be determined that the PWM drive section of the fan is faulty; when it is determined that the feedback speed value and the actual speed value corresponding to the actual PWM value are not equal, it can be determined that the speed feedback section of the fan is faulty. Furthermore, this embodiment of the invention can also analyze the cause of the fault more meticulously by modifying the theoretical PWM value of the faulty fan and observing whether the actual PWM value of the faulty fan changes. The specific fault analysis process can be set according to the actual application requirements.
[0076] Based on the above embodiments, the present invention first constructs a conversion model using a normal fan in a server. This conversion model records the conversion relationship between the surface vibration value of the normal fan, the air duct temperature of the normal fan, the server room temperature, and the pulse width modulation (PWM) value received by the normal fan. Subsequently, when the baseboard management controller detects a faulty fan, the present invention inputs the surface vibration value of the faulty fan, the air duct temperature of the faulty fan, and the server room temperature of the server to which the faulty fan belongs into the conversion model to obtain the actual PWM value received by the faulty fan. Since the actual PWM value can reflect the actual working condition of the fan in detail, the present invention can determine the cause of the faulty fan based on the actual PWM value, the theoretical PWM value driven by the baseboard management controller, and the feedback speed value of the faulty fan. This allows for analysis of the cause of the faulty fan without powering down the fan or damaging the fault scene, thereby improving the convenience of fan fault detection.
[0077] Based on the above embodiments, the specific construction process of the conversion model will be described in detail below. In one possible scenario, constructing the conversion model using a normal fan in the server may include:
[0078] S201. Install a vibration sensor on the surface of a normal fan.
[0079] In this embodiment of the invention, to facilitate finding the corresponding vibration sensor when the fan malfunctions, a vibration sensor can be installed on the surface of a normal fan and then placed on the communication bus of the server (e.g., BMC I2C bus; BMC stands for Baseboard Management Controller; I2C, also known as IIC, stands for Inter-Integrated Circuit, a two-wire serial bus). The correspondence between the bus addresses of the normal fan and the vibration sensor is recorded so that the corresponding vibration sensor can be directly found on the bus based on this correspondence.
[0080] Based on this, after installing a vibration sensor on the surface of a normal fan, it can also include:
[0081] Step 11: Obtain the bus address of the vibration sensor on the two-wire serial bus of the baseboard management controller, and establish the first correspondence between the bus addresses of the normal fan and the vibration sensor.
[0082] S202. Use heat dissipation simulation software to determine the airflow location of the normal fan inside the server, and determine the temperature sensor located at the airflow location.
[0083] To facilitate the determination of the fan airflow location, embodiments of the present invention can utilize thermal simulation software to identify the airflow location within the server that is most affected by the normal fan speed, thus determining this airflow location at the initial server setup. Subsequently, embodiments of the present invention can either use a temperature sensor located at the airflow location and installed on the server motherboard as the temperature sensor for measuring the normal fan airflow temperature, or add an external temperature sensor (such as a temperature sensing wire) at the airflow location, or use both the motherboard temperature sensor and the external temperature sensor, depending on the actual application requirements.
[0084] Based on this, the temperature sensor located in the air duct can include:
[0085] Step 21: Identify the temperature sensor located in the air duct and installed on the server motherboard; and / or, identify the external temperature sensor located in the air duct.
[0086] Furthermore, in this embodiment of the invention, to facilitate locating the corresponding temperature sensor in the air duct when a fan malfunctions, a temperature sensor can be installed in the air duct of a normal fan, and then the temperature sensor can be set on the server's communication bus (e.g., a BMC I2C bus). The correspondence between the bus addresses of the normal fan and the temperature sensor is recorded, so that the corresponding temperature sensor can be directly located on the bus based on this correspondence later. Please refer to... Figure 2 , Figure 2This is a schematic diagram of server fan fault detection hardware provided in an embodiment of the present invention. As can be seen, each temperature sensor can be connected to the BMC I2C bus; simultaneously, each vibration sensor can also be connected to the BMC I2C bus via an I2C connector.
[0087] Based on this, after determining the temperature sensor located in the air duct, the following may also be included:
[0088] Step 31: Obtain the bus address of the temperature sensor on the two-wire serial bus of the baseboard management controller, and establish a second correspondence between the bus addresses of the normal fan and the temperature sensor.
[0089] S203. Set the computer room temperature value to the initial value, and control the baseboard management controller to adjust the pulse width modulation value of the normal fan by controlling the server to run the stress test software. During the adjustment process, the vibration sensor and temperature sensor are used to continuously record the surface vibration value and air duct temperature value of the normal fan, so as to obtain the surface vibration value and air duct temperature value corresponding to different pulse width modulation values of the normal fan under the current computer room temperature value.
[0090] This invention first sets the control room temperature to a fixed value. Then, to obtain the fan surface vibration value and fan duct temperature value under different server operating conditions, this invention controls the server to operate under different loads by running stress testing software, thereby controlling the baseboard management controller to adjust the pulse width modulation value of the normal fan. Subsequently, during the adjustment process, this invention continuously uses vibration and temperature sensors to record the surface vibration value and duct temperature value of the normal fan, thereby obtaining the surface vibration value and duct temperature value corresponding to different pulse width modulation values of the normal fan under the current data room temperature. Specifically, this invention can control the baseboard management controller to adjust the pulse width modulation value of the normal fan within the range of 20% to 100% by running stress testing software on the server, thereby obtaining the surface vibration value and duct temperature value corresponding to the 20% to 100% pulse width modulation value of the normal fan under the current data room temperature.
[0091] S204. Adjust the room temperature value multiple times, and after each adjustment, control the baseboard management controller to adjust the pulse width modulation value of the normal fan by running the stress test software through the control server. During the adjustment process, continuously use vibration sensors and temperature sensors to record the surface vibration value and air duct temperature value of the normal fan, and obtain the surface vibration value and air duct temperature value corresponding to different pulse width modulation values of the normal fan under different room temperature values.
[0092] Furthermore, in this embodiment of the invention, the computer room temperature value will be adjusted multiple times to measure the surface vibration value and duct temperature value corresponding to different pulse width modulation values of a normal fan under different computer room temperatures. This embodiment of the invention does not limit the adjustment range of the computer room temperature value; for example, it can be adjusted between 15℃ and 27℃.
[0093] For easier understanding, please refer to Figure 3 , Figure 3 This is a schematic diagram illustrating a normal machine fan temperature and vibration parameter acquisition process provided in an embodiment of the present invention.
[0094] S205. Construct a conversion model using the surface vibration values and duct temperature values corresponding to different pulse width modulation values of a normal fan under different machine room temperatures.
[0095] After completing the data measurement, embodiments of the present invention can construct a conversion table using the surface vibration values and duct temperature values corresponding to different pulse width modulation values of a normal fan under different computer room temperatures, and use the conversion table as a conversion model; alternatively, mathematical modeling can be performed based on the above data, and the mathematical model can be used as a conversion model, which can be set according to actual application requirements.
[0096] Based on the above embodiments, since the first correspondence between the fan and the vibration sensor bus address and the second correspondence between the fan and the temperature sensor bus address have been predetermined in the embodiments of the present invention, after a faulty fan is detected, the present invention can locate the sensor corresponding to the fan based on the above two correspondences and use the corresponding sensor to perform data measurement. Therefore, the method may further include:
[0097] S301. Construct a conversion model using a normal fan in the server; the conversion model records the conversion relationship between the surface vibration value of the normal fan, the air duct temperature value of the air duct where the normal fan is located, the server room temperature value of the server room, and the pulse width modulation value received by the normal fan.
[0098] S302. When the baseboard management controller detects a faulty fan, it searches for the bus address of the vibration sensor corresponding to the faulty fan according to the first correspondence of the faulty fan, and obtains the surface vibration value of the faulty fan from the vibration sensor corresponding to the faulty fan according to the bus address of the vibration sensor corresponding to the faulty fan.
[0099] S303. Find the bus address of the temperature sensor corresponding to the faulty fan according to the second correspondence of the faulty fan, and obtain the duct temperature value of the faulty fan from the temperature sensor corresponding to the faulty fan according to the bus address of the temperature sensor corresponding to the faulty fan.
[0100] S304. Input the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the computer room temperature value of the server to which the faulty fan belongs into the conversion model to obtain the actual pulse width modulation value received by the faulty fan.
[0101] S305. Determine the cause of the faulty fan based on the actual pulse width modulation value, the theoretical pulse width modulation value of the board management controller driving the faulty fan, and the feedback speed value of the faulty fan.
[0102] Unlike the previous embodiments, the present invention adds steps S302 and S303. That is, the present invention has pre-determined the first correspondence between the fan and the vibration sensor bus address and the second correspondence between the fan and the temperature sensor bus address. Therefore, after a faulty fan is detected, the present invention can find the sensor corresponding to the fan based on the above two correspondences, and read the sensor measurement value by accessing the two sensors on the bus, thereby facilitating the acquisition of values.
[0103] Based on the above embodiments, the analysis process for fan failure causes is described below based on a specific fan type. In one possible case, this embodiment of the invention can detect the failure of a dual-rotor fan, wherein the dual-rotor fan contains two rotors, and the actual pulse width modulation values of the two rotors can be detected separately. Dual-rotor fans typically encounter two types of failures: single-rotor failure and dual-rotor failure, and the failure analysis processes for these two types of failures are slightly different. This embodiment of the invention will first introduce the failure analysis process for single-rotor failure. In one possible case, the faulty fan is a dual-rotor fan. The cause of the faulty fan is determined based on the actual pulse width modulation value, the theoretical pulse width modulation value driven by the board management controller to the faulty fan, and the feedback speed value fed back by the faulty fan. This can include:
[0104] S401. When it is determined that the faulty fan has a single rotor fault, determine whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value.
[0105] Since the baseboard management controller typically controls both rotors in a fan based on the same theoretical pulse width modulation (PWM) value, when a single rotor failure is determined in the faulty fan, it can be directly concluded that the theoretical PWM value issued by the baseboard management controller is not faulty. In this case, possible causes of the failure include: 1. The faulty rotor's response to the theoretical PWM value is incorrect; 2. The baseboard management controller's interpretation of the feedback speed value is incorrect; 3. The baseboard management controller cannot receive the feedback speed value. To determine the cause of the failure, this embodiment of the invention first determines whether the actual PWM value of the faulty rotor is the same as the theoretical PWM value.
[0106] S402. When it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value, the cause of the fault is determined to be that the faulty rotor is erroneously responding to the theoretical pulse width modulation value.
[0107] In this step, when it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value, the cause of the fault can be directly determined to be that the faulty rotor is erroneously responding to the theoretical pulse width modulation value.
[0108] S403. When it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value, the control board management controller increases the theoretical pulse width modulation value, and during the increase process, it determines whether the feedback speed value of the faulty rotor can be increased.
[0109] In this step, when it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value, it indicates that the faulty rotor's response to the theoretical pulse width modulation value is normal. Therefore, the possible causes of the fault are further narrowed down to: 1. The baseboard management controller misinterprets the feedback speed value; 2. The baseboard management controller cannot receive the feedback speed value. At this time, in order to further narrow down the range of causes of the fault, this embodiment of the invention can control the baseboard management controller to increase the theoretical pulse width modulation value, and determine whether the feedback speed value of the faulty rotor can be increased during the increase process.
[0110] S404. When it is determined that the feedback speed value of the faulty rotor can be increased, the cause of the fault is determined to be that the baseboard management controller erroneously interprets the feedback speed value of the faulty rotor.
[0111] In this step, when it is determined that the feedback speed value of the faulty rotor can be increased, it indicates that the baseboard management controller can receive the feedback speed value normally, but the interpretation of the feedback speed value is incorrect.
[0112] S405. When it is determined that the feedback speed value of the faulty rotor cannot be increased, replace the firmware of the baseboard management controller. After replacement, re-control the baseboard management controller to increase the theoretical pulse width modulation value, and judge whether the feedback speed value is increased during the increase process.
[0113] In this step, if it is determined that the feedback speed value of the faulty rotor can be increased, it indicates that the baseboard management controller cannot receive the feedback speed value. In this case, the fault may also be caused by: 1. a firmware (FM) fault in the baseboard management controller; 2. a transmission link fault between the faulty fan and the baseboard management controller. To further determine the cause of the fault, this embodiment of the invention can replace the baseboard management controller firmware, readjust the theoretical pulse width modulation value of the fan after firmware replacement, and determine whether the fan's feedback speed value increases.
[0114] S406. When it is determined that the feedback speed value of the faulty rotor can be increased after replacing the firmware of the baseboard management controller, the cause of the fault is determined to be a fault in the original firmware of the baseboard management controller.
[0115] In this step, if it is determined that the feedback speed value of the faulty rotor can be increased after replacing the firmware of the baseboard management controller, then the cause of the fault can be directly determined to be a fault in the original firmware of the baseboard management controller.
[0116] S407. When it is determined that the feedback speed value of the faulty rotor cannot be increased after replacing the firmware of the baseboard management controller, the cause of the fault is determined to be a transmission link failure between the faulty fan and the baseboard management controller.
[0117] In this step, if it is determined that the feedback speed value of the faulty rotor cannot be increased after replacing the firmware of the baseboard management controller, the cause of the fault can be determined to be a transmission link failure between the faulty fan and the baseboard management controller.
[0118] For ease of understanding, the fault analysis process for single rotor faults can be found by referring to [reference needed]. Figure 4 , Figure 4 This is a schematic diagram of a fault detection and anomaly judgment process for a single rotor of a server fan provided in an embodiment of the present invention.
[0119] The following describes the fault analysis process for dual-rotor faults. Unlike single-rotor faults, when a dual-rotor fault is determined to occur in the faulty fan, the cause of the fault may also include: the baseboard control controller failing to send pulse width modulation (PWM) values normally. Therefore, for dual-rotor faults, this embodiment of the invention also needs to confirm this additional cause of the fault. This embodiment of the invention can also determine whether the cause of the fault occurs in the baseboard control controller or the transmission link between the baseboard control controller and the faulty rotor based on the feedback speed value of the fan. That is, it can also determine whether the cause of the fault is that the baseboard control controller itself cannot send PWM values, or whether the failure of the transmission link between the baseboard control controller and the faulty rotor causes the inability to send PWM values, based on steps S403 to S407. However, this is easily confused with other causes of faults. Therefore, this embodiment of the invention will analyze whether the cause of the faulty rotor is that the baseboard control controller cannot send PWM values normally based on the actual PWM value of the faulty rotor. Of course, steps S403 to S407 can be executed first to estimate the range of causes of the fault, and then steps S501 to 507 as described below can be executed to further determine the cause of the fault.
[0120] Based on this, the faulty fan is a dual-rotor fan. The cause of the faulty fan is determined by the actual pulse width modulation value, the theoretical pulse width modulation value driven by the board management controller, and the feedback speed value from the faulty fan. This can include:
[0121] S501. When the baseboard management controller detects a dual-rotor fault in the faulty fan, for each faulty rotor, it determines whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value.
[0122] Unlike single-rotor faults, when a dual-rotor fault is determined to occur in the faulty fan, the causes may include: 1. The baseboard management controller cannot send pulse width modulation (PWM) values normally; 2. The faulty rotor responds incorrectly to the theoretical PWM value; 3. The baseboard management controller misinterprets the feedback speed value; 4. The baseboard management controller cannot receive the feedback speed value. To determine the cause of the fault, this embodiment of the invention first determines whether the actual PWM value of the faulty rotor is the same as the theoretical PWM value.
[0123] S502. When it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value, the control board management controller will increase the theoretical pulse width modulation value, and during the increase process, it will determine whether the feedback speed value of the faulty rotor can be increased.
[0124] In this step, when it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value, it indicates that the theoretical pulse width modulation value issued by the board management controller for the faulty rotor is correct. At this time, steps S401 to S407 can be further executed to determine other causes of the faulty rotor.
[0125] S503. When it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value, the control board management controller increases the theoretical pulse width modulation value, and during the increase process, it determines whether the actual pulse width modulation value of the faulty rotor can be increased.
[0126] In this step, when it is determined that the actual pulse width modulation value of the faulty rotor differs from the theoretical pulse width modulation value, the possible causes of the fault are: 1. The faulty rotor incorrectly responds to the theoretical pulse width modulation value issued by the baseboard management controller; 2. The baseboard management controller is unable to issue a pulse width modulation value normally. To further determine the cause of the fault, this embodiment of the invention can control the baseboard management controller to increase the theoretical pulse width modulation value, and during the increase process, determine whether the actual pulse width modulation value of the faulty rotor can be increased.
[0127] S504. When it is determined that the actual pulse width modulation value of the faulty rotor can be increased, the cause of the fault is determined to be the faulty rotor's incorrect response to the theoretical pulse width modulation value.
[0128] In this step, when it is determined that the actual pulse width modulation value of the faulty rotor can be increased, it can be determined that the baseboard control controller can issue the correct theoretical pulse width modulation value, and the cause of the fault can be determined to be that the faulty rotor is responding incorrectly to the theoretical pulse width modulation value.
[0129] S505. When it is determined that the actual pulse width modulation value of the faulty rotor cannot be increased, replace the firmware of the baseboard management controller and re-determine whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value after replacement.
[0130] In this step, when it is determined that the actual pulse width modulation value of the faulty rotor cannot be increased, it can be determined that the baseboard management controller cannot send a pulse width modulation value. The possible causes of the fault are: 1. A fault in the original firmware of the baseboard management controller; 2. A fault in the transmission link between the faulty fan and the baseboard management controller. To further determine the cause of the fault, this embodiment of the invention can replace the firmware of the baseboard management controller, and after the firmware replacement, re-evaluate whether the actual pulse width modulation value of the faulty rotor has returned to normal.
[0131] S506. When it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value after the firmware of the baseboard management controller is replaced, the cause of the fault is determined to be a fault in the original firmware of the baseboard management controller.
[0132] In this step, if the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value after the firmware of the baseboard management controller is replaced, then the cause of the fault can be determined to be a fault in the original firmware of the baseboard management controller.
[0133] S507. When it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value after the firmware of the baseboard management controller is replaced, the cause of the fault is determined to be a transmission link failure between the faulty fan and the baseboard management controller.
[0134] In this step, when it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value after the firmware of the baseboard management controller is replaced, the cause of the fault is determined to be a transmission link failure between the faulty fan and the baseboard management controller.
[0135] As can be seen, the embodiments of the present invention can introduce the actual pulse width modulation value of the faulty fan, thereby effectively determining the specific cause of the single rotor fault of the fan without interrupting the current working state of the faulty fan.
[0136] For ease of understanding, the fault analysis process for dual-rotor faults can be found by referring to [reference needed]. Figure 5 , Figure 5 This is a schematic diagram of a fault detection and anomaly judgment process for a dual-rotor server fan provided in an embodiment of the present invention.
[0137] The fan fault detection device, electronic device, and computer-readable storage medium provided in the embodiments of the present invention are described below. The fan fault detection device, electronic device, and computer-readable storage medium described below can be referred to in correspondence with the fan fault detection method described above.
[0138] Please refer to Figure 6 , Figure 6 This is a structural block diagram of a fan fault detection device provided in an embodiment of the present invention. The device may include:
[0139] The model building module 601 is used to build a conversion model using a normal fan in the server. The conversion model records the conversion relationship between the surface vibration value of the normal fan, the air duct temperature value of the air duct where the normal fan is located, the computer room temperature value of the server room, and the pulse width modulation value received by the normal fan.
[0140] The numerical detection module 602 is used to input the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the computer room temperature value of the server to which the faulty fan belongs into the conversion model when the baseboard management controller detects a faulty fan, so as to obtain the actual pulse width modulation value received by the faulty fan.
[0141] The fault detection module 603 is used to determine the cause of the faulty fan based on the actual pulse width modulation value, the theoretical pulse width modulation value of the faulty fan driven by the baseboard management controller, and the feedback speed value fed back by the faulty fan.
[0142] Optionally, the model building module 601 may include:
[0143] The vibration sensor setting submodule is used to set vibration sensors on the surface of a normal fan;
[0144] The temperature sensor setting submodule is used to determine the airflow location of the normal fan inside the server using heat dissipation simulation software, and to identify the temperature sensor located in the airflow location.
[0145] The measurement submodule is used to set the computer room temperature value to the initial value, and control the baseboard management controller to adjust the pulse width modulation value of the normal fan by controlling the server to run the stress test software. During the adjustment process, the vibration sensor and temperature sensor are used to record the surface vibration value and air duct temperature value of the normal fan, so as to obtain the surface vibration value and air duct temperature value corresponding to different pulse width modulation values of the normal fan under the current computer room temperature value.
[0146] The computer room temperature adjustment submodule is used to adjust the computer room temperature value multiple times. After each adjustment, it controls the baseboard management controller to adjust the pulse width modulation value of the normal fan by running stress test software through the control server. During the adjustment process, it continuously uses vibration sensors and temperature sensors to record the surface vibration value and air duct temperature value of the normal fan. This yields the surface vibration value and air duct temperature value corresponding to different pulse width modulation values of the normal fan under different computer room temperatures.
[0147] The model building submodule is used to build a conversion model using the surface vibration values and duct temperature values corresponding to different pulse width modulation values of a normal fan under different computer room temperatures.
[0148] Optionally, the vibration sensor setting submodule may also include:
[0149] The first correspondence establishment unit is used to obtain the bus address of the vibration sensor on the two-wire serial bus of the baseboard management controller, and establish the first correspondence between the bus addresses of the normal fan and the vibration sensor.
[0150] The temperature sensor setting submodule may also include:
[0151] The second correspondence establishment unit is used to obtain the bus address of the temperature sensor on the two-wire serial bus of the baseboard management controller, and establish a second correspondence between the bus addresses of the normal fan and the temperature sensor.
[0152] The device may further include:
[0153] The surface vibration value acquisition module is used to find the bus address of the vibration sensor corresponding to the faulty fan according to the first correspondence of the faulty fan, and to obtain the surface vibration value of the faulty fan from the vibration sensor corresponding to the faulty fan according to the bus address of the vibration sensor corresponding to the faulty fan.
[0154] The duct temperature value acquisition module is used to find the bus address of the temperature sensor corresponding to the faulty fan according to the second correspondence of the faulty fan, and to obtain the duct temperature value of the faulty fan from the temperature sensor corresponding to the faulty fan according to the bus address of the temperature sensor corresponding to the faulty fan.
[0155] Optionally, the temperature sensor setting submodule can be used for:
[0156] Identify the temperature sensor located in the air duct and positioned on the server motherboard;
[0157] And / or, determine the location of the external temperature sensor in the air duct.
[0158] Optionally, the model building module 601 can be used for:
[0159] The conversion model is constructed by using normal fans from multiple servers.
[0160] Optionally, the faulty fan is a dual-rotor fan, and the fault detection module 603 may include:
[0161] The first judgment submodule is used to determine whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value when it is determined that the faulty fan has a single rotor fault.
[0162] The first determination submodule is used to determine the cause of the fault as the faulty rotor erroneously responding to the theoretical pulse width modulation value when the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value.
[0163] The second judgment submodule is used to control the baseboard management controller to increase the theoretical pulse width modulation value when it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value, and to judge whether the feedback speed value of the faulty rotor can be increased during the increase process.
[0164] The second determination submodule is used to determine the cause of the fault as the baseboard management controller erroneously parsing the feedback speed value of the faulty rotor when it is determined that the feedback speed value of the faulty rotor can be increased.
[0165] The third judgment submodule is used to replace the firmware of the baseboard management controller when it is determined that the feedback speed value of the faulty rotor cannot be increased. After replacement, the baseboard management controller is re-controlled to increase the theoretical pulse width modulation value, and the feedback speed value is judged during the increase process.
[0166] The third determination submodule is used to determine that the cause of the fault is a fault in the original firmware of the baseboard management controller when it is determined that the feedback speed value of the faulty rotor can be increased after replacing the firmware of the baseboard management controller.
[0167] The fourth determination submodule is used to determine the cause of the fault as a transmission link failure between the faulty fan and the baseboard management controller when the feedback speed value of the faulty rotor cannot be increased after the firmware of the baseboard management controller is replaced.
[0168] Optionally, the fault detection module 603 may further include:
[0169] The fourth judgment submodule is used to determine whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value for each faulty rotor when the baseboard management controller detects a dual-rotor fault in the faulty fan.
[0170] The process control submodule is used to call the first judgment module when it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value.
[0171] The fifth judgment submodule is used to control the baseboard management controller to increase the theoretical pulse width modulation value when the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value, and to judge whether the actual pulse width modulation value of the faulty rotor can be increased during the increase process.
[0172] The fifth determination submodule is used to determine the cause of the fault as the faulty rotor's erroneous response to the theoretical pulse width modulation value when it is determined that the actual pulse width modulation value of the faulty rotor can be increased.
[0173] The sixth judgment submodule is used to replace the baseboard management controller firmware when it is determined that the actual pulse width modulation value of the faulty rotor cannot be increased, and then re-determine whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value after replacement.
[0174] The sixth determination submodule is used to determine the cause of the fault as a fault in the original firmware of the baseboard management controller when the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value after the baseboard management controller firmware is replaced.
[0175] The seventh determination submodule is used to determine the cause of the fault as a transmission link failure between the faulty fan and the baseboard management controller when the actual pulse width modulation value of the faulty rotor differs from the theoretical pulse width modulation value after the baseboard management controller firmware is replaced.
[0176] Please refer to Figure 7 , Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. The present invention provides an electronic device 70, including a processor 71 and a memory 72; wherein, the memory 72 is used to store a computer program; the processor 71 is used to execute the fan fault detection method provided in the foregoing embodiment when executing the computer program.
[0177] The specific process of the above-mentioned fan fault detection method can be found in the corresponding content provided in the foregoing embodiments, and will not be repeated here.
[0178] Furthermore, the memory 72, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, and the storage method can be temporary storage or permanent storage.
[0179] In addition, the electronic device 70 also includes a power supply 73, a communication interface 74, an input / output interface 75, and a communication bus 76; wherein, the power supply 73 is used to provide operating voltage for the various hardware devices on the electronic device 70; the communication interface 74 can create a data transmission channel between the electronic device 70 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this invention, and is not specifically limited here; the input / output interface 75 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0180] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the fan fault detection method described in any of the above embodiments.
[0181] Since the embodiments of the computer-readable storage medium portion correspond to the embodiments of the fan failure detection method portion, the embodiments of the storage medium portion are described in the description of the embodiments of the fan failure detection method portion, and will not be repeated here.
[0182] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0183] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0184] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0185] The present invention has provided a detailed description of a fan fault detection method, apparatus, electronic device, and storage medium. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of these embodiments are merely illustrative of the method and its core concepts. It should be noted that those skilled in the art can make various improvements and modifications to the invention without departing from its principles, and these improvements and modifications also fall within the scope of protection of the claims.
Claims
1. A fan fault detection method, characterized in that, include: A conversion model is constructed using a normal fan in the server; the conversion model records the conversion relationship between the surface vibration value of the normal fan, the air duct temperature value of the air duct where the normal fan is located, the computer room temperature value of the server room, and the pulse width modulation value received by the normal fan; When the baseboard management controller detects a faulty fan, it inputs the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the room temperature value of the server to which the faulty fan belongs into the conversion model to obtain the actual pulse width modulation value received by the faulty fan. The cause of the malfunction of the fan is determined based on the actual pulse width modulation value, the theoretical pulse width modulation value driven by the baseboard management controller, and the feedback speed value fed back by the malfunctioning fan. The faulty fan is a dual-rotor fan. Determining the cause of the faulty fan's failure based on the actual pulse width modulation (PWM) value, the theoretical PWM value driven by the baseboard management controller, and the feedback speed value from the faulty fan includes: When it is determined that the faulty fan has a single rotor fault, it is determined whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value. When it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value, the cause of the fault is determined to be that the faulty rotor is erroneously responding to the theoretical pulse width modulation value. When it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value, the baseboard management controller is controlled to increase the theoretical pulse width modulation value, and during the increase process, it is determined whether the feedback speed value of the faulty rotor can be increased. When it is determined that the feedback speed value of the faulty rotor can be increased, the cause of the fault is determined to be that the baseboard management controller erroneously interprets the feedback speed value of the faulty rotor. When it is determined that the feedback speed value of the faulty rotor cannot be increased, the firmware of the baseboard management controller is replaced. After replacement, the baseboard management controller is re-controlled to increase the theoretical pulse width modulation value, and the feedback speed value is judged during the increase process. When it is determined that the feedback speed value of the faulty rotor can be increased after replacing the firmware of the baseboard management controller, the cause of the fault is determined to be a fault in the original firmware of the baseboard management controller. When it is determined that the feedback speed value of the faulty rotor cannot be increased after replacing the firmware of the baseboard management controller, the cause of the fault is determined to be a transmission link failure between the faulty fan and the baseboard management controller.
2. The fan fault detection method according to claim 1, characterized in that, The method of constructing a conversion model using normal fans in the server includes: A vibration sensor is installed on the surface of the normal fan. The airflow location of the normal fan inside the server is determined using thermal simulation software, and the temperature sensor located at the airflow location is also identified. The computer room temperature value is set to an initial value. The baseboard management controller is controlled to adjust the pulse width modulation value of the normal fan by controlling the server to run stress test software. During the adjustment process, the vibration sensor and the temperature sensor are used to continuously record the surface vibration value and air duct temperature value of the normal fan, so as to obtain the surface vibration value and air duct temperature value corresponding to different pulse width modulation values of the normal fan under the current computer room temperature value. The process involves repeatedly adjusting the room temperature value and, after each adjustment, controlling the baseboard management controller to adjust the pulse width modulation value of the normal fan by controlling the server to run stress test software. During the adjustment process, the vibration sensor and the temperature sensor are used to continuously record the surface vibration value and air duct temperature value of the normal fan. This process yields the surface vibration value and air duct temperature value corresponding to different pulse width modulation values of the normal fan under different room temperatures. The conversion model is constructed using the surface vibration values and duct temperature values corresponding to different pulse width modulation values of the normal fan under different computer room temperatures.
3. The fan fault detection method according to claim 2, characterized in that, After installing a vibration sensor on the surface of the normal fan, the system further includes: Obtain the bus address of the vibration sensor on the two-wire serial bus of the baseboard management controller, and establish a first correspondence between the bus addresses of the normal fan and the vibration sensor; After determining the temperature sensor located in the air duct, the following is also included: Obtain the bus address of the temperature sensor on the two-wire serial bus of the baseboard management controller, and establish a second correspondence between the bus address of the normal fan and the temperature sensor; Before inputting the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the room temperature value of the server to which the faulty fan belongs into the conversion model, the following steps are also included: The bus address of the vibration sensor corresponding to the faulty fan is found according to the first correspondence of the faulty fan, and the surface vibration value of the faulty fan is obtained from the vibration sensor corresponding to the faulty fan according to the bus address of the vibration sensor corresponding to the faulty fan. The bus address of the temperature sensor corresponding to the faulty fan is found according to the second correspondence of the faulty fan, and the duct temperature value of the faulty fan is obtained from the temperature sensor corresponding to the faulty fan according to the bus address of the temperature sensor corresponding to the faulty fan.
4. The fan fault detection method according to claim 2, characterized in that, The temperature sensor used to determine the location of the air duct includes: Identify the temperature sensor located in the air duct and installed on the server motherboard; And / or, determine the location of the external temperature sensor in the air duct.
5. The fan fault detection method according to claim 1, characterized in that, The method of constructing a conversion model using normal fans in the server includes: The conversion model is constructed by using normal fans from multiple servers.
6. The fan fault detection method according to claim 1, characterized in that, include: When the baseboard management controller detects a dual-rotor fault in the faulty fan, it determines for each faulty rotor whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value. When it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value, the step of controlling the baseboard management controller to increase the theoretical pulse width modulation value is entered, and during the increase process, it is determined whether the feedback speed value of the faulty rotor can be increased. When it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value, the baseboard management controller is controlled to increase the theoretical pulse width modulation value, and during the increase process, it is determined whether the actual pulse width modulation value of the faulty rotor can be increased. When it is determined that the actual pulse width modulation value of the faulty rotor can be increased, the cause of the fault is determined to be that the faulty rotor is erroneously responding to the theoretical pulse width modulation value. When it is determined that the actual pulse width modulation value of the faulty rotor cannot be increased, the firmware of the baseboard management controller is replaced, and after the replacement, it is re-evaluated whether the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value. When it is determined that the actual pulse width modulation value of the faulty rotor is the same as the theoretical pulse width modulation value after the firmware of the baseboard management controller is replaced, the cause of the fault is determined to be a fault in the original firmware of the baseboard management controller. When it is determined that the actual pulse width modulation value of the faulty rotor is different from the theoretical pulse width modulation value after the firmware of the baseboard management controller is replaced, the cause of the fault is determined to be a transmission link failure between the faulty fan and the baseboard management controller.
7. A fan fault detection device, characterized in that, Based on the fan fault detection method according to claim 1 or claim 6, the device includes: The model building module is used to build a conversion model using a normal fan in the server. The conversion model records the conversion relationship between the surface vibration value of the normal fan, the air duct temperature value of the air duct where the normal fan is located, the computer room temperature value of the server room, and the pulse width modulation value received by the normal fan. The numerical detection module is used to input the surface vibration value of the faulty fan, the air duct temperature value of the air duct where the faulty fan is located, and the computer room temperature value of the computer room where the server to which the faulty fan belongs are all into the conversion model when the baseboard management controller detects a faulty fan, so as to obtain the actual pulse width modulation value received by the faulty fan. The fault detection module is used to determine the cause of the fault of the faulty fan based on the actual pulse width modulation value, the theoretical pulse width modulation value driven by the baseboard management controller to drive the faulty fan, and the feedback speed value fed back by the faulty fan.
8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the fan failure detection method as described in any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when loaded and executed by a processor, implement the fan fault detection method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Reducing acoustic tone excitation in computer system
CN103148025A
Real-time fan state monitoring system
CN107061338A