Fault detection method and device and electronic equipment
By pressurizing the electronic equipment, real-time temperature and error times are collected, and whether the equipment is a faulty equipment is solved, the problem of difficulty in accurately determining potential faults in the prior art is improved, the accuracy of fault detection is improved and the failure rate is reduced.
Patent Information
- Application Number
- CN202510541787.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-26
AI Technical Summary
It is difficult for the prior art to accurately determine whether there is a potential failure of electronic devices, resulting in the continuous increase in the failure rate of electronic devices after they are launched.
By pressurizing the electronic device, the internal temperature of the equipment will rise, and the real-time temperature and error times are collected. If the protection temperature does not exceed and the number of errors does not reach the preset number, it will be determined to be a normal device, otherwise it will be a faulty device.
The interception rate of electronic devices before they are officially put into use has been improved, and the failure rate of electronic devices after they are put into use has been reduced.
Smart Images

Figure CN120540883A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of electronic equipment, and in particular to a fault detection method, device and electronic equipment. Background Art
[0002] Errors that occur during the production process of electronic devices can affect their performance, so errors that occur during production represent potential problems in actual use. As the computing power of electronic devices increases, managing errors that occur during production becomes increasingly complex, making it difficult to determine whether an electronic device has a potential fault. This leads to a continuously increasing failure rate for electronic devices after they are put into production.
[0003] How to improve the accuracy of detecting whether electronic equipment has potential faults, so as to improve the accuracy of identifying faulty electronic equipment and thereby reduce the failure rate of electronic equipment after it is put into use is an urgent problem to be solved in this field. Summary of the Invention
[0004] The present application provides a fault detection method, device, and electronic device, which can improve the accuracy of detecting potential faults in electronic equipment, thereby reducing the failure rate of electronic equipment after it is put into use.
[0005] This application provides a fault detection method, including:
[0006] responding to a fault detection instruction for a target device to determine a protection temperature;
[0007] Collect the real-time temperature and number of errors generated by the target device during the stress test;
[0008] If the target device's real-time temperature does not exceed the protection temperature within the preset time, and the number of errors generated does not reach the preset number, the target device is determined to be normal;
[0009] If the real-time temperature exceeds the protection temperature, or the number of errors generated reaches the preset number, the target device is determined to be a faulty device.
[0010] The present application also provides a fault detection device, comprising:
[0011] A response module, configured to respond to a fault detection instruction for a target device to determine a protection temperature;
[0012] A processing module is used to collect the real-time temperature and the number of errors generated by the target device during the stress test operation;
[0013] The processing module is further configured to determine that the target device is a normal device if the real-time temperature of the target device does not exceed the protection temperature within a preset time and the number of errors generated does not reach a preset number;
[0014] The processing module is further configured to determine that the target device is a faulty device if the real-time temperature exceeds the protection temperature or the number of errors generated reaches a preset number.
[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned fault detection methods when executing the computer program.
[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned fault detection methods are implemented.
[0017] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned fault detection methods when executed by a processor.
[0018] Through the fault detection method, device, and electronic equipment provided by this application, before the electronic equipment is officially put into use, the electronic equipment is pressurized, causing the temperature of the electronic equipment to rise. Since high temperature will accelerate the aging of electronic components and material degradation, it can accelerate the exposure of possible faults in the electronic equipment, thereby screening electronic equipment that has not failed under the pressurization treatment, and officially putting the electronic equipment that has not failed under the pressurization treatment into use. In this way, the accuracy of detecting potential faults in electronic equipment can be improved, thereby improving the interception rate of electronic equipment before it is officially put into use, and thereby reducing the failure rate of electronic equipment after it is put into use. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 This is a schematic diagram of the scenario for this application example;
[0021] Figure 2 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 1 ;
[0022] Figure 3 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 2 ;
[0023] Figure 4 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 3 ;
[0024] Figure 5 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 4 ;
[0025] Figure 6 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 5 ;
[0026] Figure 7 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 6 ;
[0027] Figure 8 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 7 ;
[0028] Figure 9 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 8 ;
[0029] Figure 10 The overall process of fault detection for example;
[0030] Figure 11 A schematic diagram of the structure of a fault detection device provided in an embodiment of the present application;
[0031] Figure 12 This is a schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION
[0032] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0033] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0034] The embodiments of the present application pressurize electronic devices to increase the internal temperature of the electronic devices, thereby accelerating the exposure of potential faults in the electronic devices. Electronic devices that fail during the pressurization process are determined as faulty devices, and electronic devices that fail during the pressurization process are determined as normal devices. Only normal devices are officially put into use, which can improve the interception rate of electronic devices before they are officially put into use, and thus reduce the failure rate of electronic devices after they are put into use.
[0035] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0036] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the fault detection method depends, the specific application environment architecture or specific hardware architecture is described here. Figure 1 , Figure 1 This is a schematic diagram of the scenario of this application example. Figure 1 In the embodiment of the present application, the execution subject may be a detection device with computing capabilities, which is used to detect electronic equipment in a pressurized state to determine whether the electronic equipment has potential faults.
[0037] Figure 2 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 1 ,like Figure 2 As shown, the embodiment of the present application provides a fault detection method, which is described in detail as follows:
[0038] S201: Responding to a fault detection instruction for a target device to determine a protection temperature.
[0039] In this scenario, the target device is the electronic device to be tested. The fault detection command can be the ipmitoolsensor list command, which lists various sensor information on the server, including parameters such as temperature, voltage, and fan speed. The protection temperature is a preset fuse temperature. When the actual temperature of the target device reaches the preset fuse temperature, the preset fuse mechanism is triggered.
[0040] S202: Collect the real-time temperature and the number of errors generated by the target device during the stress test.
[0041] Based on the scenario example, the target device is subjected to stress testing. Specifically, the target device can be put into a stress test state by executing a preset stress test script. When the target device is in the stress test state, the internal real-time temperature of the target device begins to rise. Errors generated by the target device can be divided into correctable errors (CE) and uncorrectable errors (UCE).
[0042] S203: If the real-time temperature of the target device does not exceed the protection temperature within the preset time, and the number of errors generated does not reach the preset number, the target device is determined to be a normal device.
[0043] In this scenario example, the preset time can be determined based on actual conditions. For example, if the preset time is set to 10 hours, the target device must maintain a stress test state for 10 hours. During this preset time, the number of CE and UCE events generated is determined. The preset number can be determined based on actual conditions and can be one or more. If the number of CE and UCE events generated does not reach the preset number, and the real-time temperature does not reach the protection temperature, it indicates that the electronic device has remained stable and has not experienced any failures during the stress test.
[0044] S204: If the real-time temperature exceeds the protection temperature, or the number of errors generated reaches a preset number, the target device is determined to be a faulty device.
[0045] In this scenario, if the target device's real-time temperature exceeds the protection temperature, it indicates that the target device has reached the fuse temperature and the fuse mechanism has been triggered, and the target device can be determined to be faulty. Alternatively, if the number of CE and UCE events generated during a stress test reaches a preset number, it indicates that the target device has a high error rate and can also be determined to be faulty. The preset number is determined based on actual conditions and can be one or more.
[0046] This example applies pressure to the target device to increase the internal temperature of the electronic device, thereby accelerating the exposure of potential failures in the electronic device. This can improve the interception rate of the electronic device before it is officially put into use, thereby reducing the failure rate of the electronic device after it is put into use.
[0047] Optional, Figure 3 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 2 ,like Figure 3 As shown, in S201, determining the protection temperature includes:
[0048] S301: Collect the ambient temperature corresponding to the target device.
[0049] In combination with the scenario example, the ambient temperature of the environment can be collected through a temperature sensor.
[0050] S302: Obtain a threshold temperature of a target device.
[0051] In this scenario example, the target device's threshold temperature is a hardware specification attribute of the target device, representing the upper temperature limit the target device can reach. This can be a preset value. For example, the threshold temperature of a central processing unit (CPU) is typically around 100°C.
[0052] S303: Obtaining a target temperature based on the threshold temperature and the ambient temperature.
[0053] In conjunction with the scenario example, the target temperature is the maximum temperature that can be reached during the test of the target device without damaging the target device.
[0054] S304: Determine a protection temperature based on the target temperature.
[0055] In the scenario example, the protection temperature means that below the protection temperature, the target device can operate normally, and above the protection temperature, the target device will trigger the fuse mechanism.
[0056] Based on the method provided in this example, it can be determined that the protection temperature is affected by the threshold temperature and the ambient temperature. The purpose of determining the protection temperature can be achieved based on the actual threshold temperature and the actual ambient temperature of the target device, and the protection temperature achieved can be made more accurate.
[0057] Optional, Figure 4 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 3 ,like Figure 4 As shown, S303 includes:
[0058] S401: Subtract the threshold temperature from the first preset temperature to obtain a result temperature.
[0059] Combined with the scenario example, the first preset temperature can be determined according to the actual situation, and is generally selected as the interval of [6-8°C]. For example, if the first preset temperature is determined to be 7°C, the result temperature is obtained by the threshold temperature of -7°C.
[0060] S402: Compensating the result temperature based on the ambient temperature to obtain a target temperature.
[0061] In this scenario example, since the ambient temperature affects the actual temperature of the electronic device, the target temperature can be adaptively adjusted based on the ambient temperature. For example, if the ambient temperature is high, the resulting temperature can be appropriately lowered to achieve the target temperature. If the ambient temperature is high, the resulting temperature can be appropriately raised to achieve the target temperature.
[0062] Based on the method provided in this example, the mutual influence among the threshold temperature, the ambient temperature and the first preset temperature can be combined to make the obtained target temperature more accurate.
[0063] Optional, Figure 5 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 4 ,like Figure 5 As shown, S402 includes:
[0064] S501: Determine whether the ambient temperature is within a target temperature range;
[0065] In this scenario example, two temperature nodes can be preset. The specific temperature values of the temperature nodes can be determined based on actual conditions. For example, if the first temperature node is set to 25°C and the second temperature node is set to 18°C, three temperature intervals can be obtained: the first temperature interval corresponding to below 18°C, the second temperature interval corresponding to above 25°C, and the third temperature interval corresponding to 18°C-25°C. Based on the actual ambient temperature, the temperature interval in which the ambient temperature falls is determined as the target temperature interval. For example, if the ambient temperature is 20°C, the target temperature interval can be determined as the third temperature interval.
[0066] S502: Determine a target adjustment strategy corresponding to the target temperature range;
[0067] Based on the scenario example, each temperature range has a corresponding adjustment strategy. For example, the adjustment strategy for the first temperature range is to increase the temperature by 0-2°C, the adjustment strategy for the second temperature range is to decrease the temperature by 1-3°C, and the adjustment strategy for the third temperature range is to keep the temperature unchanged. Therefore, when the ambient temperature is 20°C, the corresponding target adjustment strategy is to keep the temperature unchanged.
[0068] S503: Adjust the result temperature based on the adjustment strategy to obtain the target temperature.
[0069] In combination with the scenario example, since the target adjustment strategy is to keep the temperature unchanged, the obtained target temperature is the result temperature obtained by subtracting the threshold temperature from the first preset temperature.
[0070] Based on the method provided in this example, by dividing the ambient temperature into multiple intervals, the adjustment of the result temperature can be made more accurate, so as to obtain a more accurate target temperature.
[0071] Optionally, S304 includes:
[0072] The target temperature and the second preset temperature are superimposed to obtain the protection temperature.
[0073] Combined with the scenario example, the protection temperature includes the first-level protection temperature and the second-level protection temperature, so the corresponding second preset temperature can also be divided into two. For example, when the second preset temperature is 3-4°C, the target temperature is superimposed by 3-4°C, and the resulting protection temperature is the first-level protection temperature. If the actual temperature of the electronic device reaches the first-level protection temperature, the first-level fuse mechanism is triggered. At this time, the fan speed can be forced to adjust to 100%. For example, when the second preset temperature is 5-6°C, the target temperature is superimposed by 5-6°C, and the resulting protection temperature is the second-level protection temperature. If the actual temperature of the electronic device reaches the second-level protection temperature, the second-level fuse mechanism is triggered. At this time, the test biscuit can be immediately terminated to start emergency cooling.
[0074] Based on the method provided in this example, a second preset temperature is superimposed on the target temperature to obtain a protection temperature. When the actual temperature of the electronic device reaches the protection temperature, a fuse mechanism is adopted to protect the safety of the electronic device.
[0075] Optional, Figure 6 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 5 ,like Figure 6 As shown, it also includes:
[0076] S601: Based on the detection instruction, the initial temperature of the target device is collected.
[0077] In conjunction with the scenario example, the initial temperature of the target device is the temperature of the target device before pressurization of the target device.
[0078] S602: Determine a fan speed adjustment strategy based on the initial temperature of the target device and the real-time temperature during the stress test.
[0079] In combination with the scenario example, based on the initial temperature and real-time temperature of the target device, the temperature rise variation can be determined, and the fan speed adjustment strategy can be determined based on the temperature rise variation.
[0080] S603: Control the operating speed of the fan based on the fan speed adjustment strategy.
[0081] In this scenario example, if the fan speed adjustment policy is to reduce by 5%, the fan speed is reduced to 95% of the original speed. This is how the fan speed is adjusted in real time.
[0082] This example method adjusts the fan speed in a timely manner to ensure that the target device's real-time temperature rises steadily, rather than rising too quickly and triggering a fuse, thereby ensuring the safety of the target device.
[0083] Optionally, the target device includes a first device, a second device, and a third device;
[0084] Accordingly, Figure 7 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 6 ,like Figure 7 As shown, S601 includes:
[0085] S701: Based on the detection instruction, a first initial temperature of the first device, a second initial temperature of the second device, and a third initial temperature of the third device are collected.
[0086] In this scenario example, the target device typically includes multiple components. This example uses three components requiring quality inspection: the first, second, and third devices. The first device can be a graphics processing unit (GPU), the second device can be high-bandwidth memory (HBM), and the third device can be a universal baseband board (UBM).
[0087] The temperature sensor can measure a first initial temperature of the first device, a second initial temperature of the second device, and a third initial temperature of the third device. For example, the first initial temperature can be 65°C, the second initial temperature can be 60°C, and the third initial temperature can be 55°C.
[0088] Accordingly, S202 includes:
[0089] S702: Collecting the real-time temperatures of the first device, the second device, and the third device at preset time intervals to obtain a first real-time temperature corresponding to the first device, a second real-time temperature corresponding to the second device, and a third real-time temperature corresponding to the third device.
[0090] In combination with the scenario example, the preset time can be determined according to the actual situation. For example, the preset time can be determined as 15 minutes. That is, before the target device is stress tested, after the first initial temperature, the second initial temperature, and the third initial temperature are measured, the target device starts to run the stress test. After the first 15 minutes, the first real-time temperature corresponding to the first device, the second real-time temperature corresponding to the second device, and the third real-time temperature corresponding to the third device are measured for the first time. Then, after the second 15 minutes, the first real-time temperature corresponding to the first device, the second real-time temperature corresponding to the second device, and the third real-time temperature corresponding to the third device are measured for the second time. And so on, the first real-time temperature corresponding to the first device, the second real-time temperature corresponding to the second device, and the third real-time temperature corresponding to the third device are measured every 15 minutes.
[0091] Based on the method provided in this example, by detecting the temperature changes of the first device, the second device, and the third device, the temperature of the target device is detected more accurately.
[0092] Optional, Figure 8 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 7 ,like Figure 8 As shown, S602 includes:
[0093] S801: Monitors the real-time speed of the fan.
[0094] In the scenario example, the real-time fan speed can be expressed as a percentage of the maximum speed. For example, the current real-time fan speed is 80% of the maximum speed.
[0095] S802: Determine whether the collection operation is the first operation.
[0096] In this scenario example, the first real-time temperature collection operation for the first, second, and third devices occurs 15 minutes after the first collection. This can be determined by checking whether there are historically collected real-time temperatures. If so, this indicates that this is not the first real-time temperature collection operation. If not, this indicates that this is the first real-time temperature collection operation.
[0097] S803: If the acquisition operation is the first operation, the difference between the first real-time temperature obtained in this acquisition and the first initial temperature is determined as the first difference temperature, the difference between the second real-time temperature obtained in this acquisition and the second initial temperature is determined as the second difference temperature, and the difference between the third real-time temperature obtained in this acquisition and the third initial temperature is determined as the third difference temperature.
[0098] In this scenario example, if this is the first real-time temperature collection operation, there is no historical real-time temperature data. Therefore, the target device's temperature rise calculation must be compared with the initial temperature. For example, if the real-time temperatures collected this time are: the first real-time temperature for the GPU is 76°C, the second real-time temperature for the HBM is 67°C, and the third real-time temperature for the UBB is 63°C. The first temperature difference is 10°C, the second temperature difference is 7°C, and the third temperature difference is 8°C. This indicates that the GPU temperature rose by 11°C, the HBM temperature rose by 7°C, and the UBB temperature rose by 8°C in the first 15 minutes.
[0099] S804: If the collection operation is not the first operation, the difference between the first real-time temperature obtained in this collection and the first real-time temperature obtained in the previous collection is determined as the first difference temperature, the difference between the second real-time temperature obtained in this collection and the second real-time temperature obtained in the previous collection is determined as the second difference temperature, and the difference between the third real-time temperature obtained in this collection and the third real-time temperature obtained in the previous collection is determined as the third difference temperature.
[0100] In this scenario, if this is not the first acquisition, for example, the second acquisition, the real-time temperatures collected this time are: the first real-time temperature for the GPU is 82°C, the second real-time temperature for the HBM is 72°C, and the third real-time temperature for the UBB is 70°C. The second real-time temperature obtained in the previous acquisition is the same as the real-time temperature obtained in the first acquisition. Therefore, the first temperature difference is 6°C, the second temperature difference is 5°C, and the third temperature difference is 7°C.
[0101] S805: Determine a maximum temperature difference and a minimum temperature difference among the first temperature difference, the second temperature difference, and the third temperature difference.
[0102] In this scenario example, if this is the first collection operation, the first temperature difference is 11°C, the second temperature difference is 7°C, and the third temperature difference is 8°C. The maximum temperature difference is 11°C, and the minimum temperature difference is 7°C. If this is not the first collection operation, the first temperature difference is 6°C, the second temperature difference is 5°C, and the third temperature difference is 7°C. The maximum temperature difference is 7°C, and the minimum temperature difference is 5°C.
[0103] S806: If the maximum temperature difference is greater than a third preset temperature, the speed of the fan is reduced by a first preset ratio based on the real-time speed of the fan.
[0104] Based on the scenario example, the third preset temperature and the first preset ratio can be determined based on actual conditions. For example, if the third preset temperature is 10°C, the first preset ratio can be selected as 10%. When the maximum temperature difference is greater than 10°C, that is, after the first real-time temperature acquisition, the first temperature difference corresponding to the GPU is 11°C, so after the first real-time temperature acquisition, the real-time fan speed is reduced by 10%. Since the fan speed was previously 80%, after the fan speed is reduced by 10%, the current real-time fan speed is reduced to 70%.
[0105] S807: If the maximum temperature difference is less than or equal to the third preset temperature and greater than the fourth preset temperature, the speed of the fan is reduced by a second preset ratio based on the real-time speed of the fan.
[0106] Based on the scenario example, the fourth preset temperature and the second preset ratio can be determined based on actual conditions. For example, if the fourth preset temperature is 5°C, the second preset ratio can be 5%. When the maximum temperature difference is less than or equal to 10% and greater than 5°C, that is, after the second real-time temperature acquisition, the corresponding third temperature difference of UBB is 7°C. At this time, the fan speed is reduced by 5% from 70%, resulting in a new speed of 65%.
[0107] S808: If the maximum temperature difference is less than or equal to the fourth preset temperature value, maintain the real-time speed of the fan.
[0108] In this scenario, let's assume that this is the third real-time temperature acquisition. The first real-time temperature acquired is 85°C for the GPU, 75°C for the HBM, and 72°C for the UBB. Therefore, the temperature differences from the second real-time temperature acquisition are 3°C for the first, 3°C for the second, and 2°C for the third. The maximum temperature difference is 3°C, and the minimum temperature difference is 3°C. In this case, the maximum temperature difference is less than or equal to 5°C. The fan speed remains unchanged at 65%.
[0109] S809: If the minimum temperature difference is less than or equal to the fifth preset temperature, the speed of the fan is increased by a second preset ratio based on the real-time speed of the fan.
[0110] In this scenario example, the fifth preset temperature can be determined based on actual conditions. For example, if the fifth preset temperature is 0°C and this is the third acquisition, the real-time temperatures collected this time are: the first real-time temperature for the GPU is 87°C, the second real-time temperature for the HBM is 75°C, and the third real-time temperature for the UBB is 73°C. Therefore, the temperature differences from the real-time temperatures collected in the second acquisition are: the first temperature difference is 2°C, the second temperature difference is 0°C, and the third temperature difference is 1°C. The maximum temperature difference is 2°C, and the minimum temperature difference is 0°C. At this time, the fan speed is increased by 5%, resulting in a new fan speed of 70%.
[0111] Based on the method provided in this example, the fan speed can be adjusted in real time according to the real-time temperature changes of the target device to keep the real-time temperature of the target device stable and ensure the safety of the target device.
[0112] Optional, Figure 9 Schematic diagram of the process of the fault detection method provided in the embodiment of the application Figure 8 ,like Figure 9 As shown, it also includes:
[0113] S901: Monitor the temperature change rate of the target device.
[0114] In this scenario example, the temperature change rate is the rate of change between the current real-time temperature and the previous real-time temperature. For example, for the first real-time temperature, the GPU's first real-time temperature is 76°C, the HBM's second real-time temperature is 67°C, and the UBB's third real-time temperature is 63°C. The GPU's first initial temperature can be 65°C, the HBM's second initial temperature can be 60°C, and the UBB's third initial temperature can be 55°C. Therefore, the GPU's temperature change rate in the first 15 minutes is 0.73°C / min, the HBM's temperature change rate is 0.47°C / min, and the UBB's temperature change rate is 0.53°C / min.
[0115] S902: If the temperature change rate exceeds a preset value, the target device is determined to be a faulty device.
[0116] Based on the scenario example, the preset value can be determined based on actual conditions. For example, the preset value can be set to 1°C / min. If the temperature change rate of any device (GPU, HBM, or UBB) exceeds 1°C / min, the target device's temperature is considered to be rising too quickly, and the target device can be identified as faulty. The method provided in this example can improve the accuracy of determining whether a target device is faulty.
[0117] Figure 10 The overall process of fault detection is shown as follows: Figure 10 As shown, the target device can be initialized and safety checked. If initialization fails, the target device is directly judged to be a faulty device, the test is terminated, and a faulty device warning is triggered. If initialization fails successfully, data such as ambient temperature, initial temperature, and threshold temperature are collected, and the target temperature is obtained by combining the dynamic threshold value of the threshold temperature and performing temperature compensation based on the ambient temperature. Then, the target device stress test is started, and the target device is monitored in real time to detect whether the temperature change of the target device triggers the fuse mechanism and the number of CE or CUE errors. When the preset number of CE or CUE errors is 1, the test will be terminated as long as a CE or CUE error occurs. If no CE or CUE error occurs, the monitoring will continue until the preset time is reached. For example, if no CE or CUE error occurs within 10 hours, it can be determined that the target device is not a faulty device.
[0118] This embodiment applies pressure to the target device to increase the internal temperature of the electronic device, thereby accelerating the exposure of potential faults in the electronic device. This can improve the interception rate of the electronic device before it is officially put into use, thereby reducing the failure rate of the electronic device after it is put into use.
[0119] Figure 11 This is a schematic diagram of the structure of the fault detection device provided in the embodiment of the present application. Figure 11 As shown, an embodiment of the present application further provides a fault detection device, comprising:
[0120] The response module 111 is used to respond to the fault detection instruction of the target device to determine the protection temperature;
[0121] Processing module 112, for collecting the real-time temperature and number of errors generated by the target device during the stress test operation;
[0122] The processing module 112 is further configured to determine that the target device is a normal device if the real-time temperature of the target device does not exceed the protection temperature within a preset time and the number of errors generated does not reach a preset number;
[0123] The processing module 112 is further configured to determine that the target device is a faulty device if the real-time temperature exceeds the protection temperature or the number of errors generated reaches a preset number.
[0124] Optionally, the response module 111 is specifically configured to collect the ambient temperature corresponding to the target device;
[0125] The response module 111 is further configured to obtain a threshold temperature of a target device;
[0126] The response module 111 is further configured to obtain a target temperature based on the threshold temperature and the ambient temperature;
[0127] The response module 111 is further configured to determine a protection temperature based on the target temperature.
[0128] Optionally, the response module 111 is further configured to subtract the threshold temperature from the first preset temperature to obtain a result temperature;
[0129] The response module 111 is further configured to compensate the result temperature based on the ambient temperature to obtain a target temperature.
[0130] Optionally, the response module 111 is further configured to determine the target temperature range in which the ambient temperature is located;
[0131] The response module 111 is further configured to determine a target adjustment strategy corresponding to a target temperature range;
[0132] The response module 111 is further configured to adjust the result temperature based on the adjustment strategy to obtain a target temperature.
[0133] Optionally, the response module 111 superimposes the target temperature and the second preset temperature to obtain a protection temperature.
[0134] Optionally, the processing module 112 is further configured to collect the initial temperature of the target device based on the detection instruction;
[0135] The processing module 112 is further configured to determine a fan speed adjustment strategy based on the initial temperature of the target device and the real-time temperature during the stress test operation;
[0136] The processing module 112 is further configured to control the operating speed of the fan based on the fan speed adjustment strategy.
[0137] Optionally, the processing module 112 is specifically configured to collect a first initial temperature of the first device, a second initial temperature of the second device, and a third initial temperature of the third device based on the detection instruction;
[0138] The processing module 112 is further configured to collect the real-time temperatures of the first device, the second device, and the third device at preset time intervals to obtain a first real-time temperature corresponding to the first device, a second real-time temperature corresponding to the second device, and a third real-time temperature corresponding to the third device.
[0139] Optionally, the processing module 112 is further configured to monitor the real-time speed of the fan;
[0140] The processing module 112 is further configured to determine whether the acquisition operation is a first operation;
[0141] The processing module 112 is further configured to, if the acquisition operation is the first operation, determine the difference between the first real-time temperature acquired in the current acquisition and the first initial temperature as the first difference temperature, determine the difference between the second real-time temperature acquired in the current acquisition and the second initial temperature as the second difference temperature, and determine the difference between the third real-time temperature acquired in the current acquisition and the third initial temperature as the third difference temperature;
[0142] The processing module 112 is further configured to, if the acquisition operation is not the first operation, determine the difference between the first real-time temperature acquired in the current acquisition and the first real-time temperature acquired in the previous acquisition as a first difference temperature, determine the difference between the second real-time temperature acquired in the current acquisition and the second real-time temperature acquired in the previous acquisition as a second difference temperature, and determine the difference between the third real-time temperature acquired in the current acquisition and the third real-time temperature acquired in the previous acquisition as a third difference temperature;
[0143] The processing module 112 is further configured to determine a maximum temperature difference and a minimum temperature difference among the first temperature difference, the second temperature difference, and the third temperature difference;
[0144] The processing module 112 is further configured to reduce the speed of the fan by a first preset ratio based on the real-time speed of the fan if the maximum temperature difference is greater than a third preset temperature;
[0145] The processing module 112 is further configured to reduce the fan speed by a second preset ratio based on the real-time fan speed if the maximum temperature difference is less than or equal to the third preset temperature and greater than a fourth preset temperature;
[0146] The processing module 112 is further configured to maintain the real-time speed of the fan if the maximum temperature difference is less than or equal to a fourth preset temperature value;
[0147] The processing module 112 is further configured to increase the speed of the fan by a second preset ratio based on the real-time speed of the fan if the minimum temperature difference is less than or equal to the fifth preset temperature.
[0148] For the description of the features in the embodiment corresponding to the fault detection device, reference can be made to the relevant description of the embodiment corresponding to the fault detection method, which will not be repeated here.
[0149] Figure 12 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 12 As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the electronic device 50 further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus.
[0150] During the specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502 , so that the at least one processor 501 executes the above-mentioned fault detection method embodiment.
[0151] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0152] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0153] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.
[0154] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0155] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned fault detection method embodiments when running.
[0156] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0157] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned fault detection method embodiments are implemented.
[0158] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned fault detection method embodiments are implemented.
[0159] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0160] The above is a detailed introduction to a fault detection method, device, and electronic device provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A fault detection method, characterized in that: include: responding to a fault detection instruction for a target device to determine a protection temperature; Collecting the real-time temperature and number of errors generated by the target device during the stress test operation; If the real-time temperature of the target device does not exceed the protection temperature within a preset time, and the number of errors generated does not reach a preset number, the target device is determined to be a normal device; If the real-time temperature exceeds the protection temperature, or the number of errors generated reaches a preset number, the target device is determined to be a faulty device.
2. The method according to claim 1, characterized in that Determining the protection temperature includes: Collecting the ambient temperature corresponding to the target device; Obtaining a threshold temperature of the target device; Obtaining a target temperature based on the threshold temperature and the ambient temperature; A protection temperature is determined based on the target temperature.
3. The method according to claim 2, characterized in that The obtaining of a target temperature based on the threshold temperature and the ambient temperature includes: Subtracting the threshold temperature from the first preset temperature to obtain a result temperature; The resulting temperature is compensated based on the ambient temperature to obtain the target temperature.
4. The method according to claim 3, characterized in that The compensating the result temperature based on the ambient temperature to obtain the target temperature includes: Determining the target temperature range within which the ambient temperature lies; Determining a target adjustment strategy corresponding to the target temperature range; The result temperature is adjusted based on the adjustment strategy to obtain the target temperature.
5. The method according to claim 2, characterized in that The determining of the protection temperature based on the target temperature includes: The target temperature and the second preset temperature are superimposed to obtain the protection temperature.
6. The method according to claim 1, characterized in that Also includes: Based on the detection instruction, collecting the initial temperature of the target device; Determining a fan speed adjustment strategy based on the initial temperature of the target device and the real-time temperature during the stress test operation; The operating speed of the fan is controlled based on the fan speed adjustment strategy.
7. The method according to claim 6, characterized in that The target device includes a first device, a second device and a third device; Accordingly, collecting the initial temperature of the target device based on the detection instruction includes: Based on the detection instruction, collecting a first initial temperature of the first device, a second initial temperature of the second device, and a third initial temperature of the third device; Accordingly, collecting the real-time temperature of the target device during the stress test includes: At each preset time interval, the real-time temperatures of the first device, the second device and the third device are collected respectively to obtain a first real-time temperature corresponding to the first device, a second real-time temperature corresponding to the second device and a third real-time temperature corresponding to the third device.
8. The method according to claim 7, characterized in that The determining of the fan speed adjustment strategy based on the initial temperature of the target device and the real-time temperature during the stress test includes: Monitor the real-time speed of the fan; Determining whether the acquisition operation is a first operation; If the acquisition operation is the first operation, the difference between the first real-time temperature obtained in the current acquisition and the first initial temperature is determined as the first difference temperature, the difference between the second real-time temperature obtained in the current acquisition and the second initial temperature is determined as the second difference temperature, and the difference between the third real-time temperature obtained in the current acquisition and the third initial temperature is determined as the third difference temperature; If the acquisition operation is not the first operation, the difference between the first real-time temperature obtained in the current acquisition and the first real-time temperature obtained in the previous acquisition is determined as the first difference temperature, the difference between the second real-time temperature obtained in the current acquisition and the second real-time temperature obtained in the previous acquisition is determined as the second difference temperature, and the difference between the third real-time temperature obtained in the current acquisition and the third real-time temperature obtained in the previous acquisition is determined as the third difference temperature; determining a maximum temperature difference and a minimum temperature difference among the first temperature difference, the second temperature difference, and the third temperature difference; If the maximum temperature difference is greater than a third preset temperature, reducing the speed of the fan by a first preset ratio based on the real-time speed of the fan; If the maximum temperature difference is less than or equal to the third preset temperature and greater than a fourth preset temperature, the speed of the fan is reduced by a second preset ratio based on the real-time speed of the fan; If the maximum temperature difference is less than or equal to a fourth preset temperature value, maintaining the real-time speed of the fan; If the minimum temperature difference is less than or equal to a fifth preset temperature, the speed of the fan is increased by a second preset ratio based on the real-time speed of the fan.
9. A fault detection device, characterized in that: include: A response module, configured to respond to a fault detection instruction for a target device to determine a protection temperature; A processing module, configured to collect the real-time temperature and the number of errors generated by the target device during the stress test operation; The processing module is further configured to determine that the target device is a normal device if the real-time temperature of the target device does not exceed the protection temperature within a preset time and the number of errors generated does not reach a preset number; The processing module is further configured to determine that the target device is a faulty device if the real-time temperature exceeds the protection temperature or the number of errors generated reaches a preset number.
10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the fault detection method according to any one of claims 1 to 8 when executing the computer program.