Fault detection method and device, electronic equipment and storage medium

By receiving enhanced error reports from the server and machine inspection architecture register data, accurately locate the specific internal equipment for the fault of the graphics processing unit, the problem that fault diagnosis in the prior art cannot cover all scenarios and improves the efficiency of fault detection.

CN120086049APending Publication Date: 2025-06-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510233541.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art cannot accurately determine the specific internal equipment in which the graphic processing unit fails, resulting in failure diagnosis being unable to cover all scenarios.

Method used

By receiving enhanced error report register data from the server and machine inspection architecture register data, it is determined whether the graphics processing unit is in a healthy state, and determines the fault type and fault location based on the processing core information and memory information.

Benefits of technology

Accurate positioning of faults of graphics processing unit is achieved, and the efficiency and coverage of fault detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086049A_ABST
    Figure CN120086049A_ABST
Patent Text Reader

Abstract

The invention discloses a fault detection method and device, electronic equipment and a storage medium, and relates to the technical field of computers.The method comprises the steps that enhanced error report register data and machine inspection architecture register data from a server are received, the enhanced error report register data and the machine inspection architecture register data are sent when the server detects that the graphics processing unit has uncorrectable errors and the number of times of correctable errors reaches a preset threshold value; under the condition that the enhanced error report register data and the machine check architecture register data are valid, determining whether the graphics processing unit is in a healthy state according to the machine check architecture register data; if not, the fault type and the fault position of the graphic processing unit are determined according to at least one of the processing core information and the memory information included in the enhanced error report register data, so that specific internal equipment with faults of the graphic processing unit can be accurately positioned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a fault detection method, apparatus, electronic device, and storage medium. Background Art

[0002] During the use of a server, various device failures will inevitably occur, such as the hang of a peripheral component interconnect express (PCIE) bus, the failure of a Central Processing Unit (CPU), the failure of a Graphics Processing Unit (GPU) card, etc. These device failures may cause the server system to crash, which has already caused inconvenience to users in using the server.

[0003] In related technologies, generally, a Baseboard Management Controller (BMC) can obtain the health status information of internal devices of a server and monitor a GPU card based on the health status information of the internal devices of the server. However, this method can only detect that the GPU card has failed, but cannot accurately locate the specific internal device where the GPU card has failed, resulting in the failure diagnosis of the GPU card not covering all scenarios. Summary of the Invention

[0004] This application provides a fault detection method, apparatus, electronic device, and storage medium to at least solve the problem in related technologies that the specific internal device where a graphics processing unit fails cannot be accurately located, resulting in the failure diagnosis of the graphics processing unit not covering all scenarios.

[0005] This application provides a fault detection method applied to a controller. The method includes: receiving enhanced error report register data and machine check architecture register data from a server, where the enhanced error report register data includes at least one of processing core information and memory information of a graphics processing unit, and the machine check architecture register data is used to indicate multiple performance metrics of the graphics processing unit, and the enhanced error report register data and the machine check architecture register data are sent when the number of uncorrectable errors and correctable errors detected by the server for the graphics processing unit reaches a preset threshold; when the enhanced error report register data and the machine check architecture register data are valid, determining whether the graphics processing unit is in a healthy state according to the machine check architecture register data; if not, determining the fault type and fault location of the graphics processing unit according to at least one of the processing core information and the memory information, and the fault location includes one or more of a processing core or memory.

[0006] The present application provides a fault detection method, which is applied to a server. The method includes: executing a fault diagnosis script to detect uncorrectable errors and correctable errors of a graphics processing unit through the fault diagnosis script; when the number of uncorrectable errors and correctable errors detected in the graphics processing unit reaches a preset threshold, obtaining enhanced error report register data and machine check architecture register data of the graphics processing unit, where the enhanced error report register data includes at least one of processing core information and memory information of the graphics processing unit, and the machine check architecture register data is used to indicate multiple performance metrics of the graphics processing unit; and sending the enhanced error report register data and the machine check architecture register data to a controller, so that the controller determines the fault type and fault location of the graphics processing unit based on the enhanced error report register data and the machine check architecture register data.

[0007] The present application further provides a fault detection device, which includes: a transceiver module, configured to receive enhanced error report register data and machine check architecture register data from a server, where the enhanced error report register data includes at least one of processing core information and memory information of a graphics processing unit, the machine check architecture register data is used to indicate multiple performance metrics of the graphics processing unit, and the enhanced error report register data and the machine check architecture register data are sent when the number of uncorrectable errors and correctable errors detected in the graphics processing unit by the server reaches a preset threshold; a processing module, configured to determine whether the graphics processing unit is in a healthy state according to the machine check architecture register data when the enhanced error report register data and the machine check architecture register data are valid; and the processing module is further configured to, if not, determine the fault type and fault location of the graphics processing unit according to at least one of the processing core information and the memory information, and the fault location includes one or more of the processing core or the memory.

[0008] The present application further provides a fault detection device, which includes: a script collection module, configured to execute a fault diagnosis script to detect uncorrectable errors and correctable errors of a graphics processing unit through the fault diagnosis script. The script collection module is further configured to, when the number of uncorrectable errors and correctable errors detected in the graphics processing unit reaches a preset threshold, obtain enhanced error report register data and machine check architecture register data of the graphics processing unit, where the enhanced error report register data includes at least one of processing core information and memory information of the graphics processing unit, and the machine check architecture register data is used to indicate multiple performance metrics of the graphics processing unit. A script sending module is further configured to send the enhanced error report register data and the machine check architecture register data to a controller, so that the controller determines the fault type and fault location of the graphics processing unit based on the enhanced error report register data and the machine check architecture register data.

[0009] The present application also provides a fault detection system, including: a controller and a server.

[0010] The server is configured to execute a fault diagnosis script to detect uncorrectable errors and correctable errors of the graphics processing unit through the fault diagnosis script.

[0011] The server is further configured to, when the number of uncorrectable errors and correctable errors detected in the graphics processing unit reaches a preset threshold, obtain enhanced error report register data and machine check architecture register data of the graphics processing unit. The enhanced error report register data includes at least one of processing core information and memory information of the graphics processing unit, and the machine check architecture register data is used to indicate multiple performance metrics of the graphics processing unit.

[0012] The server is further configured to send the enhanced error report register data and the machine check architecture register data to the controller.

[0013] The controller receives the enhanced error report register data and the machine check architecture register data from the server.

[0014] When the enhanced error report register data and the machine check architecture register data are valid, the controller determines whether the graphics processing unit is in a healthy state according to the machine check architecture register data.

[0015] If not, the controller determines the fault type and fault location of the graphics processing unit according to at least one of the processing core information and the memory information. The fault type includes one or more of processing core faults or memory faults.

[0016] The present application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any one of the above fault detection methods when executing the computer program.

[0017] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any one of the above fault detection methods when executed by a processor.

[0018] The present application also provides a computer program product including a computer program, and the computer program implements the steps of any one of the above fault detection methods when executed by a processor.

[0019] Through this application, the controller can receive enhanced error report register data and machine check architecture register data from the server. Since the enhanced error report register data and the machine check architecture register data are sent by the server when the number of uncorrectable errors and correctable errors detected in the graphics processing unit reaches a preset threshold, that is, when there may be a risk of the graphics processing unit crashing, the controller can receive the enhanced error report register data and the machine check architecture register data.

[0020] Furthermore, the controller can determine whether the graphics processing unit is in a healthy state according to the machine check architecture register data, and when the graphics processing unit is in an unhealthy state, based on at least one of the processing core information and memory information included in the enhanced error report register data, determine the fault type and fault location of the graphics processing unit, so that the controller can accurately locate the specific internal device (processing core or memory) where the GPU card fails, solve the problem that the fault diagnosis of the graphics processing unit cannot cover all scenarios, and improve the fault detection efficiency of the graphics processing unit. Brief Description of the Drawings

[0021] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0022] Figure 1 It is a topology diagram of a fault detection system provided by an embodiment of the present application;

[0023] Figure 2 It is a flowchart of a fault detection method provided by an embodiment of the present application;

[0024] Figure 3 It is a flowchart of another fault detection method provided by an embodiment of the present application;

[0025] Figure 4 It is a structural block diagram of a fault detection device provided by an embodiment of the present application;

[0026] Figure 5 It is a structural block diagram of another fault detection device provided by an embodiment of the present application;

[0027] Figure 6 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed Description of the Embodiments

[0028] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0029] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0030] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the fault detection method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0032] The embodiments of the present application are applied to the scenario where a baseboard management controller performs fault detection on a graphics processing unit.

[0033] In the related art, the Baseboard Management Controller (BMC) supports obtaining the health status information of internal devices of the server through the management component transport protocol. However, the BMC cannot collect and analyze the uncorrectable errors (UCE) and FAULT class faults of the internal devices of the graphics processing unit. Therefore, it is impossible to accurately locate the specific internal device where the GPU card fails, resulting in the out-of-band fault diagnosis of the GPU card not covering all scenarios.

[0034] To solve the above technical problems, the embodiments of the present application provide a fault detection method. The method receives enhanced error report register data and machine check architecture register data from the server. When the enhanced error report register data and the machine check architecture register data are valid, according to the machine check architecture register data, it is determined whether the graphics processing unit is in a healthy state. If not, according to at least one of the processing core information and the memory information, the fault type and fault location of the graphics processing unit are determined to accurately locate the specific internal device where the graphics processing unit fails and improve the efficiency of fault detection.

[0035] Taking the following Figure 1 shown fault detection system 100 as an example, the method provided by the embodiments of the present application will be described. Figure 1 It is only a schematic diagram and does not constitute a limitation on the applicable scenarios of the technical solutions provided by the present application.

[0036] As Figure 1 shown, Figure 1 it is a topology diagram of a fault detection system provided by the embodiments of the present application. Figure 1 In it, the fault detection system 100 may include a controller 101 and a server 102.

[0037] The controller 101 in the embodiments of the present invention may be a Baseboard Management Controller (BMC) integrated in a server, a network device, and other computer systems. The controller is also integrated with a client, and the client has a display interface.

[0038] The server 102 in the embodiments of the present invention may be any device with communication functions and computing functions. For example, the server 102 may be a server, a cloud server, or a virtual machine. A Graphics Processing Unit (GPU) is deployed on the server 102. The Graphics Processing Unit is also called a GPU card. There are registers on the GPU card, and these registers are used to store the enhanced error reporting register data and machine check architecture register data of the GPU card.

[0039] Figure 1 The shown fault detection system 100 is only for illustration and does not limit the technical solutions of the present application. Those skilled in the art should understand that in the specific implementation process, the fault detection system 100 may further include other devices, which are not limited.

[0040] According to the embodiments of the present application, an embodiment of a fault detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0041] In this embodiment, a fault detection method is provided, which can be used for the above-mentioned server, Figure 2 It is a flowchart of a fault detection method provided by the embodiments of the present application. As Figure 2 shown, this process includes the following steps:

[0042] S201: Execute a fault diagnosis script to detect uncorrectable errors and correctable errors of the Graphics Processing Unit through the fault diagnosis script.

[0043] Among them, an uncorrectable error refers to an error that occurs in the GPU and cannot be repaired through its own mechanism or software means. Uncorrectable errors usually lead to serious failures or data loss in the GPU system. For example, uncorrectable errors are memory failures and physical damage to the GPU hardware.

[0044] A correctable error refers to an error that can be repaired by the GPU's own error correction mechanism, software driver adjustment, or specific operations of the system during the operation of the GPU. For example, correctable errors are graphic rendering lags and driver conflicts.

[0045] Exemplarily, the server receives a detection instruction from the controller instructing the server to start executing a fault diagnosis script, executes the fault diagnosis script, and detects uncorrectable errors and correctable errors of the GPU card through the fault diagnosis script.

[0046] Optionally, the server can also detect the temperature, voltage, current, and presence status of the GPU card.

[0047] S202: When the number of uncorrectable errors and correctable errors detected in the graphics processing unit reaches a preset threshold, obtain the enhanced error report register data and machine check architecture register data of the graphics processing unit.

[0048] Among them, the enhanced error report register data includes at least one of the processing core information and memory information of the graphics processing unit.

[0049] The processing core information includes the first occurrence numbers of uncorrectable errors and correctable errors in the first-level cache (also known as L1 cache) and second-level cache (also known as L2 cache) of the processing core of the graphics processing unit, arithmetic logic unit failure information, and floating-point unit failure information.

[0050] The arithmetic logic unit (ALU) failure information and floating-point unit (FPU) failure information are used to indicate various fault class failures. For example, various fault class failures of the GPU card can include: failures caused by overheating (too high a temperature of the GPU will affect the electron mobility of the ALU and FPU, extend the switching time of logic gates, and cause errors in the operation results), electrical failures (voltage fluctuations and current overloads of the GPU will prevent the ALU and FPU from obtaining stable working voltages and currents, affect the accuracy and stability of the operation, and may also burn out the hardware), physical damage failures, and aging failures.

[0051] The fault information of the logic operation unit can also be used to indicate operation result errors, abnormal data output, abnormalities of other devices connected to the logic operation unit, and signal transmission failures. The fault information of the floating-point operation unit can also be used to indicate abnormal operation results (random errors in results, inaccurate results), program freezes, or slow operation, etc.

[0052] The memory information includes the second number of uncorrectable errors and correctable errors that occur in the memory of the graphics processing unit, and the occupancy information of at least one memory-resident program of the graphics processing unit.

[0053] The machine check architecture register data is used to indicate multiple performance metrics of the graphics processing unit. For example, the multiple performance metrics can be temperature, voltage, current, and in-position status.

[0054] In a possible design, the preset threshold can be set according to actual needs without limitation. For example, the preset threshold is 10.

[0055] Optionally, when the server detects that the number of uncorrectable errors and correctable errors that occur in the graphics processing unit reaches the preset threshold, the first register value of the first register data bit of the graphics processing unit is set to a preset value. For example, the preset value can be 1. When the server detects that the number of uncorrectable errors and correctable errors that occur in the graphics processing unit does not reach the preset threshold, the first register value of the first register data bit of the graphics processing unit is set to a non-preset value. For example, the non-preset value can be 0.

[0056] The server detects the first number of uncorrectable errors and correctable errors that occur in the level-1 cache and level-2 cache of the processing core of the graphics processing unit. If the first number is greater than or equal to the first threshold, the second register value of the second register data bit of the graphics processing unit is set to a preset value. If the first number is less than the first threshold, the second register value of the second register data bit of the graphics processing unit is set to a non-preset value.

[0057] The server detects the second number of uncorrectable errors and correctable errors that occur in the memory of the graphics processing unit. If the second number is greater than or equal to the second threshold, the third register value of the third register data bit of the graphics processing unit is set to a preset value. If the second number is less than the second threshold, the third register value of the third register data bit of the graphics processing unit is set to a non-preset value.

[0058] Among them, the first register data bit, the second register data bit, and the third register data bit are all data bits of a preset register. For example, the preset register can be the MCerrlogReg register.

[0059] The server detects at least one occupancy information of at least one memory resident program of the graphics processing unit and sets it to the corresponding register; similarly, the server detects the logical operation unit fault information and floating-point operation unit fault information of the graphics processing unit and sets them to the corresponding registers.

[0060] It can be understood that both the enhanced error report register data and the machine check architecture register data are stored in the data bits of the register, facilitating subsequent determination of the fault type and fault location of the graphics processing unit based on the values stored in the data bits of the register.

[0061] S203: Send the enhanced error report register data and the machine check architecture register data to the controller.

[0062] In one example, the server sends the enhanced error report register data and the machine check architecture register data to the controller based on a preset protocol.

[0063] Among them, the preset protocol can be a data transmission protocol pre-agreed between the controller and the fault diagnosis script executed by the server. For example, the preset protocol can be the Standard Intelligent Platform Management Interface protocol (Intelligent Platform Management Interface, IPMI protocol).

[0064] It can be understood that the server can detect the uncorrectable errors and correctable errors of the graphics processing unit through the fault diagnosis script, and when the number of uncorrectable errors and correctable errors detected in the graphics processing unit reaches a preset threshold, that is, when the number of uncorrectable errors and correctable errors detected in the graphics processing unit is relatively large, indicating that there is a risk of the graphics processing unit crashing, obtain the enhanced error report register data and the machine check architecture register data of the graphics processing unit and send them to the controller, so that the BMC can further determine the fault location and fault type of the graphics processing unit based on the enhanced error report register data and the machine check architecture register data.

[0065] In this embodiment, another fault detection method is provided, which can be used for the above-mentioned controller. Figure 3 It is a flowchart of another fault detection method provided by the embodiments of the present application. As Figure 3 shown, the process includes the following steps:

[0066] S301: Receive the enhanced error report register data and the machine check architecture register data from the server.

[0067] Exemplarily, the controller receives the enhanced error report register data and the machine check architecture register data from the server based on the IPMI protocol.

[0068] S302: When the enhanced error report register data and the machine check architecture register data are valid, determine whether the graphics processing unit is in a healthy state according to the machine check architecture register data.

[0069] In some alternative embodiments, the controller detects whether the in-position state is a normal state; if so, detects whether the temperature is less than or equal to a preset temperature, whether the voltage is less than or equal to a preset voltage, and whether the current is less than or equal to a preset current; if it is detected that the temperature is less than or equal to the preset temperature, the voltage is less than or equal to the preset voltage, and the current is less than or equal to the preset current, determine that the graphics processing unit is in a healthy state.

[0070] Among them, the preset temperature, preset voltage, and preset current can all be set according to the type of the graphics processing unit, the number of processing cores, or the memory size, without limitation.

[0071] In one example, if it is detected that the temperature is greater than the preset temperature, or the voltage is greater than the preset voltage, or the current is greater than the preset current, the controller determines that the graphics processing unit is not in a healthy state.

[0072] Optionally, if the controller detects that the in-position state is not a normal state, determine that the graphics processing unit is not in a healthy state.

[0073] It can be understood that if the controller detects that the in-position state of the graphics processing unit is not a normal state, or there is overheating of the temperature, excessive voltage, or excessive current, it can be determined that the graphics processing unit is not in a healthy state.

[0074] In some alternative embodiments, the controller detects whether the in-position state is a normal state; if so, determine a first deviation value corresponding to the temperature according to the temperature and a preset temperature range; determine a second deviation value corresponding to the voltage according to the voltage and a preset voltage range; and determine a third deviation value corresponding to the current according to the current and a preset current range; obtain a first weight value corresponding to the temperature, a second weight value corresponding to the voltage, and a third weight value corresponding to the current; determine a target deviation value of the graphics processing unit according to the first deviation value and the first weight value, the second deviation value and the second weight value, and the third deviation value and the third weight value; if the target deviation value is less than a preset deviation value, determine that the graphics processing unit is in a healthy state.

[0075] Among them, the minimum temperature in the preset temperature range is T1, the maximum temperature is T2, the temperature is T, and the first deviation value is H1.

[0076] When T ≤ T1, H1 = (T1 - T) / (T2 - T1).

[0077] When T1 ≤ T ≤ T2, H1 = 0.

[0078] When T2 ≤ T, H1 = (T - T2) / (T2 - T1).

[0079] The minimum voltage in the preset voltage range is U1, the maximum voltage is U2, the voltage is U, and the second deviation value is H2.

[0080] When U ≤ U1, H2 = (U1 - U) / (U2 - U1).

[0081] When U1 ≤ U ≤ U2, H2 = 0.

[0082] When U2 ≤ U, H2 = (U - U2) / (U2 - U1).

[0083] The minimum current in the preset current range is I1, the maximum current is I2, the current is I, and the third deviation value is H3.

[0084] When I ≤ I1, H3 = (I1 - I) / (I2 - I1).

[0085] When I1 ≤ I ≤ I2, H3 = 0.

[0086] When I2 ≤ I, H3 = (I - I2) / (I2 - I1).

[0087] In one example, the controller determines the target deviation value of the graphics processing unit as the sum of the product of the first deviation value and the first weight value, the product of the second deviation value and the second weight value, and the product of the third deviation value and the third weight value.

[0088] It can be understood that when the target deviation value is less than the preset deviation value, the controller determines that the graphics processing unit is in a healthy state. When the target deviation value is greater than or equal to the preset deviation value, the controller determines that the graphics processing unit is in a healthy state. That is, it indicates that the smaller the deviation of the voltage, current, and temperature of the graphics processing unit, the healthier the graphics processing unit is. The greater the deviation of the voltage, current, and temperature of the graphics processing unit, the more likely it is in an unhealthy state or a faulty state.

[0089] In some alternative embodiments, before determining whether the graphics processing unit is in a healthy state according to the machine check architecture register data, the controller reads the first register value of the first register bit of the graphics processing unit and detects whether the first register value is a preset value. If so, it determines that the enhanced error reporting register data and the machine check architecture register data are valid.

[0090] It can be understood that since the first register value of the first register data bit is set based on the number of uncorrectable errors and correctable errors that occur in the graphics processing unit, and the enhanced error report register data and machine check architecture register data are sent when the number of uncorrectable errors and correctable errors that occur in the graphics processing unit is greater than a preset threshold, therefore, the controller can verify the validity of the received enhanced error report register data and machine check architecture register data by reading the data of the first register data bit again, avoid the possibility of the server sending the enhanced error report register data and machine check architecture register data incorrectly, and improve the reliability of fault detection.

[0091] S303: If not, determine the fault type and fault location of the graphics processing unit according to at least one of the processing core information and the memory information.

[0092] Among them, the fault location includes one or more of the processing core or the memory.

[0093] In some optional embodiments, the controller compares the first number with the first threshold; if the first number is greater than or equal to the first threshold, determine that the fault location of the graphics processing unit is the processing core; according to the logical operation unit fault information, the floating-point operation unit fault information and the preset rule table, determine the first fault type of the graphics processing unit.

[0094] Among them, the preset rule table includes the corresponding relationship between the logical operation unit fault information, the floating-point operation unit fault information and the first fault type.

[0095] It can be understood that if the first number is less than the first threshold, the controller determines that the processing core of the graphics processing unit has no fault.

[0096] In some optional embodiments, the controller compares the second number with the second threshold; if the second number is greater than or equal to the second threshold, determine that the fault location of the graphics processing unit is the memory; according to at least one occupancy information and the preset rule table, determine the second fault type of the graphics processing unit.

[0097] The preset rule table further includes the corresponding relationship between at least one occupancy information and the second fault type.

[0098] It can be understood that if the second number is less than the second threshold, the controller determines that the processing core of the graphics processing unit has no fault.

[0099] Further, the controller can also display the enhanced error report register data and machine check architecture register data on the client of the controller, so that the user can input the manual fault analysis conclusion of the graphics processing unit based on the enhanced error report register data and machine check architecture register data. The manual fault analysis conclusion includes the preset fault location and preset fault type of the graphics processing unit. Compare the preset fault location with the fault location of the graphics processing unit, and compare the preset fault type with the fault type of the graphics processing unit. If the preset fault location is consistent with the fault location, and the preset fault type is consistent with the fault type, then output the fault type and fault location.

[0100] It can be understood that comparing the manual fault analysis conclusion with the fault location and fault type determined by the controller can further determine whether the fault location and fault type determined by the controller are accurate, and improve the accuracy rate of the determined fault location and fault type.

[0101] Based on the above Figure 3 According to the method shown, the controller can receive the enhanced error report register data and machine check architecture register data from the server. Since the enhanced error report register data and machine check architecture register data are sent by the server when the number of uncorrectable errors and correctable errors detected in the graphics processing unit reaches the preset threshold, that is, it indicates that when there may be a risk of the graphics processing unit crashing, the controller can receive the enhanced error report register data and machine check architecture register data.

[0102] Furthermore, the controller can determine whether the graphics processing unit is in a healthy state according to the machine check architecture register data, and when the graphics processing unit is in an unhealthy state, determine the fault type and fault location of the graphics processing unit based on at least one of the processing core information and memory information included in the enhanced error report register data, so that the controller can accurately locate the specific internal device (processing core or memory) where the GPU card fails, solve the problem that the fault diagnosis of the graphics processing unit cannot cover all scenarios, and improve the fault detection efficiency of the graphics processing unit.

[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0104] The embodiment of the present application also provides a fault detection device, which is applied to the controller, as Figure 4 shown, Figure 4 is the structural block diagram of a fault detection device provided by the embodiment of the present application; the device includes:

[0105] A transceiver module 401, configured to receive enhanced error report register data and machine check architecture register data from a server. The enhanced error report register data includes at least one of processing core information and memory information of a graphics processing unit. The machine check architecture register data is used to indicate multiple performance metrics of the graphics processing unit. The enhanced error report register data and the machine check architecture register data are sent when the number of uncorrectable errors and correctable errors detected by the server in the graphics processing unit reaches a preset threshold.

[0106] A processing module 402, configured to determine whether the graphics processing unit is in a healthy state according to the machine check architecture register data when the enhanced error report register data and the machine check architecture register data are valid.

[0107] The processing module 402 is further configured to, if not, determine the fault type and fault location of the graphics processing unit according to at least one of the processing core information and the memory information. The fault location includes one or more of a processing core or a memory.

[0108] In some alternative embodiments, the machine check architecture register data includes the presence status, temperature, voltage, and current of the graphics processing unit. The processing module 402 is specifically configured to detect whether the presence status is a normal status. If so, it detects whether the temperature is less than or equal to a preset temperature, whether the voltage is less than or equal to a preset voltage, and whether the current is less than or equal to a preset current. If it is detected that the temperature is less than or equal to the preset temperature, the voltage is less than or equal to the preset voltage, and the current is less than or equal to the preset current, it is determined that the graphics processing unit is in a healthy state. If it is detected that the temperature is greater than the preset temperature, or the voltage is greater than the preset voltage, or the current is greater than the preset current, it is determined that the graphics processing unit is not in a healthy state.

[0109] In some alternative embodiments, before determining whether the graphics processing unit is in a healthy state according to the machine check architecture register data, the processing module 402 is further configured to read a first register value of a first register data bit of the graphics processing unit and detect whether the first register value is a preset value. The first register value is set by the server based on the number of uncorrectable errors and correctable errors detected in the graphics processing unit. If so, it is determined that the enhanced error report register data and the machine check architecture register data are valid.

[0110] In some alternative embodiments, processing core information includes the number of first uncorrectable errors and correctable errors that occur in the level-1 cache and level-2 cache of the processing cores of the graphics processing unit, logic operation unit failure information, and floating-point operation unit failure information; the processing module 402 is further specifically configured to compare the number of first times with a first threshold; if the number of first times is greater than or equal to the first threshold, determine that the failure location of the graphics processing unit is the processing core; and determine a first failure type of the graphics processing unit according to the logic operation unit failure information, the floating-point operation unit failure information, and a preset rule table, where the preset rule table includes the correspondence between the logic operation unit failure information, the floating-point operation unit failure information, and the first failure type.

[0111] In some alternative embodiments, memory information includes the number of second uncorrectable errors and correctable errors that occur in the memory of the graphics processing unit, and occupancy information of at least one memory resident program of the graphics processing unit; the processing module 402 is further specifically configured to compare the number of second times with a second threshold; if the number of second times is greater than or equal to the second threshold, determine that the failure location of the graphics processing unit is the memory; and determine a second failure type of the graphics processing unit according to the at least one occupancy information and the preset rule table, where the preset rule table further includes the correspondence between the at least one occupancy information and the second failure type.

[0112] In some alternative embodiments, the processing module 402 is further configured to display enhanced error report register data and machine check architecture register data on the client of the controller, so that the user can input a manual failure analysis conclusion of the graphics processing unit based on the enhanced error report register data and the machine check architecture register data, where the manual failure analysis conclusion includes a preset failure location and a preset failure type of the graphics processing unit; compare the preset failure location with the failure location of the graphics processing unit, and compare the preset failure type with the failure type of the graphics processing unit; the transceiver module 401 is further configured to output the failure type and the failure location if the preset failure location is consistent with the failure location, and the preset failure type is consistent with the failure type.

[0113] An embodiment of the present application further provides another failure detection device, which is applied to a server, as Figure 5 shown Figure 5 is a structural block diagram of another failure detection device provided by an embodiment of the present application; the device includes:

[0114] A script collection module 501 is configured to execute a failure diagnosis script to detect uncorrectable errors and correctable errors of the graphics processing unit through the failure diagnosis script.

[0115] The script collection module 501 is further configured to obtain the enhanced error report register data and machine check architecture register data of the graphics processing unit when the number of uncorrectable errors and correctable errors detected in the graphics processing unit reaches a preset threshold. The enhanced error report register data includes at least one of the processing core information and memory information of the graphics processing unit, and the machine check architecture register data is used to indicate multiple performance metrics of the graphics processing unit.

[0116] The script sending module 502 is further configured to send the enhanced error report register data and machine check architecture register data to the controller, so that the controller determines the fault type and fault location of the graphics processing unit based on the enhanced error report register data and machine check architecture register data.

[0117] For the description of the features in the corresponding embodiment of the fault detection device, reference may be made to the relevant description in the corresponding embodiment of the fault detection method, which will not be elaborated here one by one.

[0118] An embodiment of the present application further provides an electronic device, such as Figure 6 shown Figure 6 is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present application. The electronic device includes a processor 10 and a memory 20. A computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in any of the above-mentioned fault detection method embodiments.

[0119] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-mentioned fault detection method embodiments when running.

[0120] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc and other various media that can store computer programs.

[0121] An embodiment of the present application further provides a computer program product. The above-mentioned computer program product includes a computer program, and the computer program implements the steps in any of the above-mentioned fault detection method embodiments when executed by a processor.

[0122] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the computer program implements the steps in any of the above-mentioned fault detection method embodiments when executed by a processor.

[0123] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0124] The above has introduced in detail a fault detection method, device, electronic device, and storage medium provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A fault detection method, characterized in that: Applied to a controller, the method comprises: receiving enhanced error reporting register data and machine check architecture register data from a server, wherein the enhanced error reporting register data includes at least one of processing core information and memory information of a graphics processing unit, the machine check architecture register data is used to indicate multiple performance indicators of the graphics processing unit, and the enhanced error reporting register data and the machine check architecture register data are sent when the server detects that the number of uncorrectable errors and correctable errors occurring in the graphics processing unit reaches a preset threshold; In a case where the enhanced error reporting register data and the machine check architecture register data are valid, determining whether the graphics processing unit is in a healthy state according to the machine check architecture register data; If not, determine the fault type and fault location of the graphics processing unit according to at least one of the processing core information and the memory information, where the fault location includes one or more of the processing core or the memory.

2. The fault detection method according to claim 1, characterized in that: The machine check architecture register data includes the in-place status, temperature, voltage and current of the graphics processing unit; The determining whether the graphics processing unit is in a healthy state according to the machine check architecture register data comprises: Detecting whether the in-place state is a normal state; If so, detecting whether the temperature is less than or equal to a preset temperature, detecting whether the voltage is less than or equal to a preset voltage, and detecting whether the current is less than or equal to a preset current; If it is detected that the temperature is less than or equal to the preset temperature, the voltage is less than or equal to the preset voltage, and the current is less than or equal to the preset current, it is determined that the graphics processing unit is in the healthy state; If it is detected that the temperature is greater than the preset temperature, or the voltage is greater than the preset voltage, or the current is greater than the preset current, it is determined that the graphics processing unit is not in the healthy state.

3. The fault detection method according to claim 2, characterized in that: Before determining whether the graphics processing unit is in a healthy state according to the machine check architecture register data, the method further includes: Reading a first register value of a first register data bit of the graphics processing unit, and detecting whether the first register value is a preset value, wherein the first register value is set by the server based on the number of uncorrectable errors and correctable errors detected in the graphics processing unit; If so, it is determined that the enhanced error reporting register data and the machine check architecture register data are valid.

4. The fault detection method according to claim 3, characterized in that: The processing core information includes the first counts of uncorrectable errors and correctable errors occurring in the first-level cache and the second-level cache of the processing core of the graphics processing unit, logic operation unit fault information, and floating-point operation unit fault information; The determining, according to at least one of the processing core information and the memory information, a fault type and a fault location of the graphics processing unit includes: comparing the first number of times with a first threshold; If the first number is greater than or equal to the first threshold, determining that the fault location of the graphics processing unit is the processing core; A first fault type of the graphics processing unit is determined according to the logic operation unit fault information, the floating point operation unit fault information and a preset rule table, wherein the preset rule table includes a correspondence between the logic operation unit fault information, the floating point operation unit fault information and the first fault type.

5. The fault detection method according to claim 4, characterized in that: The memory information includes a second number of uncorrectable errors and correctable errors occurring in the memory of the graphics processing unit, and at least one occupancy information of at least one memory-resident program of the graphics processing unit; The determining the fault type and fault location of the graphics processing unit according to at least one of the processing core information and the memory information further includes: comparing the second number with a second threshold; If the second number is greater than or equal to the second threshold, determining that the fault location of the graphics processing unit is the memory; A second fault type of the graphics processing unit is determined according to the at least one occupancy information and the preset rule table, wherein the preset rule table further includes a corresponding relationship between the at least one occupancy information and the second fault type.

6. The fault detection method according to claim 5, characterized in that: The method further comprises: Displaying the enhanced error reporting register data and the machine check architecture register data on a client of the controller, so that a user can input an artificial fault analysis conclusion of the graphics processing unit on the client based on the enhanced error reporting register data and the machine check architecture register data, wherein the artificial fault analysis conclusion includes a preset fault location and a preset fault type of the graphics processing unit; comparing the preset fault position with the fault position of the graphics processing unit, and comparing the preset fault type with the fault type of the graphics processing unit; If the preset fault location is consistent with the fault location, and the preset fault type is consistent with the fault type, then the fault type and the fault location are output.

7. A fault detection method, characterized in that: Applied to a server, the method comprises: Executing a fault diagnosis script to detect uncorrectable errors and correctable errors of a graphics processing unit through the fault diagnosis script; When it is detected that the number of uncorrectable errors and correctable errors occurring in the graphics processing unit reaches a preset threshold, acquiring enhanced error reporting register data and machine check architecture register data of the graphics processing unit, wherein the enhanced error reporting register data includes at least one of processing core information and memory information of the graphics processing unit, and the machine check architecture register data is used to indicate multiple performance indicators of the graphics processing unit; The enhanced error reporting register data and the machine check architecture register data are sent to a controller so that the controller determines a fault type and a fault location of the graphics processing unit based on the enhanced error reporting register data and the machine check architecture register data.

8. A fault detection device, characterized in that: The fault detection device comprises: a transceiver module, configured to receive enhanced error reporting register data and machine check architecture register data from a server, wherein the enhanced error reporting register data includes at least one of processing core information and memory information of a graphics processing unit, and the machine check architecture register data is used to indicate a plurality of performance indicators of the graphics processing unit, and the enhanced error reporting register data and the machine check architecture register data are sent when the server detects that the number of uncorrectable errors and correctable errors occurring in the graphics processing unit reaches a preset threshold; a processing module, configured to determine whether the graphics processing unit is in a healthy state according to the machine check architecture register data if the enhanced error reporting register data and the machine check architecture register data are valid; The processing module is further used to determine the fault type and fault location of the graphics processing unit based on at least one of the processing core information and the memory information, if not, the fault type including one or more of a processing core fault or a memory fault.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the fault detection method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the fault detection method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Graphics processor module anomaly detection method and electronic equipment

    CN120821624A

  • A graphics processor module exception detection method and electronic device

    CN120821624B