Method and device for collecting black box data of GPU (Graphics Processing Unit) module of server

By monitoring the black box data collection instructions of the server GPU module, data collection is collected using Redfish, SMBPBI, SMBus or SMBPBI, which solves the problem of low efficiency in black box data collection for server GPU modules in the prior art, and achieves efficient and low invasive data collection.

CN120104426APending Publication Date: 2025-06-06NINGCHANG INFORMATION TECH (HANGZHOU) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510181570.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, the server GPU module black box data collection efficiency is low, and it is necessary to have a deep understanding of the internal structure of the GPU, which poses system performance and security risks.

Method used

By monitoring the black box data collection instructions of the server GPU module, if it is the preset first collection instruction, Redfish, SMBPBI and SMBus are used for data collection; if it is the preset second collection instruction, SMBPBI is used for data collection, avoiding direct access to hardware or deep dependence on drivers.

Benefits of technology

It realizes efficient collection of server GPU module black box data, reduces dependence on internal details of GPU, has low invasiveness and high compatibility, and improves data collection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104426A_ABST
    Figure CN120104426A_ABST
Patent Text Reader

Abstract

The invention provides a server GPU module black box data collection method and device, and relates to the technical field of computers. The method for collecting the black box data of the GPU module of the server comprises the following steps: monitoring a black box data collection instruction of the GPU module of the server; if the monitored black box data collection instruction is a preset first collection instruction, black box data of a server GPU module is collected through Redfish, SMBPBI and SMBus; the first collection instruction is an instruction which is triggered by an external request and indicates to collect black box data; if the monitored black box data collection instruction is a preset second collection instruction, black box data of a server GPU module is collected through SMBPBI; the second collection instruction is an instruction for collecting the black box data according to an instruction triggered by the interruption alarm. According to the method, the efficiency of collecting the black box data of the GPU module of the server can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a method and device for collecting black box data of a server GPU module. Background Art

[0002] In modern data centers and high-performance computing environments, server GPU modules play a core role, especially in applications such as artificial intelligence, deep learning, and massively parallel computing. As GPU computing power continues to improve, the demand for GPU performance monitoring and fault diagnosis is also growing.

[0003] Traditional data collection methods often rely on API calls at the operating system level or direct access to hardware registers, which requires a deep understanding of the internal structure of the GPU and may affect system performance or security. As the complexity of GPU modules increases, the process of collecting black box data from server GPU modules is more cumbersome, so the efficiency of collecting black box data from server GPU modules is not high. Therefore, how to provide a method for collecting black box data from server GPU modules to solve the problem of low efficiency of collecting black box data from server GPU modules is of great practical significance. Summary of the invention

[0004] The embodiments of the present application provide a method and device for collecting black box data of a server GPU module, which can improve the efficiency of collecting black box data of a server GPU module.

[0005] To achieve the above purpose, the technical solution of the embodiment of the present application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for collecting black box data of a server GPU module, the method comprising:

[0007] Monitoring black box data collection instructions of the server GPU module;

[0008] If the monitored black box data collection instruction is a preset first collection instruction, the black box data of the server GPU module is collected through Redfish, SMBPBI and SMBus; the first collection instruction is an instruction triggered by an external request to instruct the collection of black box data;

[0009] If the black box data collection instruction monitored is a preset second collection instruction, the black box data of the server GPU module is collected through SMBPBI; the second collection instruction is an instruction triggered by an interrupt alarm to collect black box data.

[0010] The method for collecting black box data of a server GPU module provided in an embodiment of the present application monitors the black box data collection instruction of the server GPU module during the process of collecting black box data of the server GPU module; if the monitored black box data collection instruction is a preset first collection instruction, the black box data of the server GPU module is collected through Redfish, SMBPBI and SMBus; the first collection instruction is an instruction for collecting black box data triggered by an external request; if the monitored black box data collection instruction is a preset second collection instruction, the black box data of the server GPU module is collected through SMBPBI; the second collection instruction is an instruction for collecting black box data triggered by an interrupt alarm. This method can avoid direct access to hardware or deep reliance on drivers, and realize efficient collection of black box data of server GPU modules. The black box data collection process does not require an in-depth understanding of the internal details of the GPU, has low invasiveness and high compatibility, and can effectively improve the efficiency of black box data collection of server GPU modules.

[0011] In an optional embodiment, collecting the black box data of the server GPU module through Redfish, SMBPBI and SMBus includes:

[0012] Communicate and connect with the module device of the server GPU module through Redfish, SMBPBI and SMBus;

[0013] Based on the communication connection, black box data of the server GPU module is collected.

[0014] The method of this embodiment is to establish a communication connection with the module device of the server GPU module through Redfish, SMBPBI and SMBus; based on the communication connection, collect the black box data of the server GPU module. This method can establish a communication connection with the module device of the server GPU module through Redfish, SMBPBI and SMBus when collecting the black box data of the server GPU module, and then collect the black box data of the server GPU module based on the communication connection, so as to realize the comprehensive data collection of the GPU module, and can effectively improve the efficiency of the black box data collection of the server GPU module.

[0015] In an optional embodiment, the module device includes an HGX platform management controller HMC, a field programmable gate array FPGA, and a graphics processor GPU; the communication connection with the module device of the server GPU module through Redfish, SMBPBI and SMBus includes:

[0016] Establishing a first type of communication connection with the HMC via Redfish;

[0017] Performing a second type of communication connection with the FPGA and the GPU respectively through SMBPBI;

[0018] The third type of communication connection is respectively established with the HMC, the FPGA and the GPU through the SMBus.

[0019] In the method of this embodiment, the module device includes an HGX platform management controller HMC, a field programmable gate array FPGA, and a graphics processor GPU; a first type of communication connection is established with the HMC through Redfish; a second type of communication connection is established with the FPGA and the GPU respectively through SMBPBI; and a third type of communication connection is established with the HMC, the FPGA, and the GPU respectively through SMBus, thereby providing a mechanism for communicating with the module devices of the server GPU module through Redfish, SMBPBI, and SMBus, further improving the efficiency of black box data collection of the server GPU module.

[0020] In an optional embodiment, collecting the black box data of the server GPU module based on the communication connection includes:

[0021] If it is determined that the HMC is not abnormal based on the first type of communication connection, then collecting black box data of the server GPU module through the first type of communication connection;

[0022] If it is determined that the HMC is abnormal based on the first type of communication connection, the black box data of the server GPU module is collected through the second type of communication connection and the third type of communication connection.

[0023] The method of this embodiment, if based on the first type of communication connection, it is judged that the HMC has no abnormality, then the black box data of the server GPU module is collected through the first type of communication connection; if based on the first type of communication connection, it is judged that the HMC has an abnormality, then the black box data of the server GPU module is collected through the second type of communication connection and the third type of communication connection. The black box data of the server GPU module can be efficiently collected based on the communication connection, and the black box data can be collected from three methods of Redfish, SMBPBI, and SMBus. The three methods can be redundant with each other, further improving the real-time performance of data collection and the stability of system operation, and can effectively improve the efficiency of collecting black box data of the server GPU module.

[0024] In an optional embodiment, judging that the HMC is not abnormal based on the first type of communication connection includes:

[0025] If the state of the first type of communication connection is normal communication, it is determined that the HMC is not abnormal;

[0026] The determining, based on the first type of communication connection, that the HMC is abnormal, includes:

[0027] If the state of the first type of communication connection is abnormal communication, it is determined that the HMC is abnormal.

[0028] The method of this embodiment determines that the HMC is normal if the state of the first-type communication connection is normal communication; and determines that the HMC is abnormal if the state of the first-type communication connection is abnormal communication, thereby providing a mechanism for identifying whether the HMC is abnormal, and can accurately and efficiently determine whether the HMC is abnormal, thereby improving the efficiency of collecting black box data of the server GPU module.

[0029] In an optional embodiment, collecting the black box data of the server GPU module through the first type of communication connection includes:

[0030] Obtaining first black box information from the HMC through the first type of communication connection; the first black box information includes Redfish interface information, GPU module self-test report, FPGA dump information, and HMC log;

[0031] The first black box information is added to the black box data of the server GPU module.

[0032] The method of this embodiment includes: obtaining first black box information from the HMC through the first type of communication connection; the first black box information includes Redfish interface information, GPU module self-test report, FPGA dump information, and HMC log; adding the first black box information to the black box data of the server GPU module, thereby providing a mechanism for collecting the black box data of the server GPU module through the first type of communication connection, thereby more efficiently improving the efficiency of collecting the black box data of the server GPU module.

[0033] In an optional embodiment, collecting the black box data of the server GPU module through the second type of communication connection and the third type of communication connection includes:

[0034] Obtaining second black box information from the FPGA and the GPU through the second type of communication connection; the second black box information includes hardware status, firmware version, temperature and power consumption monitoring, and PCIe connection status;

[0035] Obtaining third black box information from the HMC, the GPU, and the FPGA through the third type of communication connection; the third black box information includes component temperature, component power consumption, and register information;

[0036] The second black box information and the third black box information are added to the black box data of the server GPU module.

[0037] The method of this embodiment includes: obtaining second black box information from the FPGA and the GPU through the second type of communication connection; the second black box information includes hardware status, firmware version, temperature and power consumption monitoring, and PCIe connection status; obtaining third black box information from the HMC, the GPU, and the FPGA through the third type of communication connection; the third black box information includes component temperature, component power consumption, and register information; adding the second black box information and the third black box information to the black box data of the server GPU module, thereby providing a mechanism for collecting the black box data of the server GPU module through the second type of communication connection and the third type of communication connection, thereby more efficiently improving the efficiency of collecting the black box data of the server GPU module.

[0038] In a second aspect, an embodiment of the present application further provides a device for collecting black box data of a server GPU module, comprising:

[0039] An instruction monitoring module, used to monitor the black box data collection instructions of the server GPU module;

[0040] A first acquisition module is used to collect the black box data of the server GPU module through Redfish, SMBPBI and SMBus if the monitored black box data collection instruction is a preset first collection instruction; the first collection instruction is an instruction triggered by an external request to instruct to collect black box data;

[0041] The second acquisition module is used to collect the black box data of the server GPU module through SMBPBI if the black box data collection instruction monitored is a preset second collection instruction; the second collection instruction is an instruction triggered by an interrupt alarm to collect black box data.

[0042] In an optional embodiment, the first acquisition module is specifically configured to:

[0043] Communicate and connect with the module device of the server GPU module through Redfish, SMBPBI and SMBus;

[0044] Based on the communication connection, black box data of the server GPU module is collected.

[0045] In an optional embodiment, the module device includes an HGX platform management controller HMC, a field programmable gate array FPGA, and a graphics processor GPU; the first acquisition module is specifically used to:

[0046] Establishing a first type of communication connection with the HMC via Redfish;

[0047] Performing a second type of communication connection with the FPGA and the GPU respectively through SMBPBI;

[0048] The third type of communication connection is respectively established with the HMC, the FPGA and the GPU through the SMBus.

[0049] In an optional embodiment, the first acquisition module is specifically configured to:

[0050] If it is determined that the HMC is not abnormal based on the first type of communication connection, then collecting black box data of the server GPU module through the first type of communication connection;

[0051] If it is determined that the HMC is abnormal based on the first type of communication connection, the black box data of the server GPU module is collected through the second type of communication connection and the third type of communication connection.

[0052] In an optional embodiment, the first acquisition module is specifically configured to:

[0053] If the state of the first type of communication connection is normal communication, it is determined that the HMC is not abnormal;

[0054] The first acquisition module is specifically used to:

[0055] If the state of the first type of communication connection is abnormal communication, it is determined that the HMC is abnormal.

[0056] In an optional embodiment, the first acquisition module is specifically configured to:

[0057] Obtaining first black box information from the HMC through the first type of communication connection; the first black box information includes Redfish interface information, GPU module self-test report, FPGA dump information, and HMC log;

[0058] The first black box information is added to the black box data of the server GPU module.

[0059] In an optional embodiment, the first acquisition module is specifically configured to:

[0060] Obtaining second black box information from the FPGA and the GPU through the second type of communication connection; the second black box information includes hardware status, firmware version, temperature and power consumption monitoring, and PCIe connection status;

[0061] Obtaining third black box information from the HMC, the GPU, and the FPGA through the third type of communication connection; the third black box information includes component temperature, component power consumption, and register information;

[0062] The second black box information and the third black box information are added to the black box data of the server GPU module.

[0063] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for collecting black box data of a server GPU module of the first aspect is implemented.

[0064] In a fourth aspect, an embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the computer program is executed by the processor, the processor implements the method for collecting black box data of a server GPU module of the first aspect.

[0065] The technical effects brought about by any one of the implementation methods in the second to fourth aspects can refer to the technical effects brought about by the corresponding implementation method in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0067] Figure 1 A flowchart of a method for collecting black box data of a server GPU module provided in an embodiment of the present application;

[0068] Figure 2 A schematic diagram of a process of collecting black box data of a server GPU module through Redfish, SMBPBI and SMBus in a method for collecting black box data of a server GPU module provided in an embodiment of the present application;

[0069] Figure 3A schematic flow chart of a method for collecting black box data of a server GPU module provided in an embodiment of the present application, which is to communicate and connect with a module device of a server GPU module through Redfish, SMBPBI and SMBus;

[0070] Figure 4 A schematic diagram of a flow chart of a method for collecting black box data of a server GPU module based on a communication connection according to an embodiment of the present application;

[0071] Figure 5 A schematic diagram of a flow chart of a method for collecting black box data of a server GPU module through a first type of communication connection provided in an embodiment of the present application;

[0072] Figure 6 A schematic diagram of a flow chart of a method for collecting black box data of a server GPU module through a second type of communication connection and a third type of communication connection provided in an embodiment of the present application;

[0073] Figure 7 A schematic diagram of the structure of a device for collecting black box data of a server GPU module provided in an embodiment of the present application;

[0074] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0076] It should be noted that the terms "including" and "having" and their variations involved in the documents of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0077] The following are explanations of some of the terms that appear in the text:

[0078] (1) GPU (Graphics Processing Unit): GPU is a microprocessor specially designed for efficient processing of graphics and image data.

[0079] (2) GPU module: GPU modules, especially in the field of high-performance computing and deep learning, physically combine multiple GPUs together and connect them through efficient data exchange mechanisms (such as NVLink or PCIe high-speed interfaces) to achieve a higher level of parallel computing capabilities. This configuration is designed to address application scenarios that require a large number of parallel processing resources, such as large-scale machine learning training, complex physical simulation, and graphics rendering.

[0080] (3) BMC (Baseboard Management Controller): BMC is a dedicated microcontroller integrated on the server motherboard, used to monitor and manage the server. It plays a core role in server hardware management.

[0081] (4) HMC (HGX management controller): HMC is used to strengthen the monitoring and control of hardware in the HGX platform. The HGX platform is an AI and HPC accelerated computing platform designed specifically for cloud data centers. HMC is essentially a baseboard management controller.

[0082] (5) FPGA (Field-Programmable Gate Array): FPGA is a highly flexible integrated circuit.

[0083] (6) Redfish: Redfish is a modern device management standard that aims to simplify and unify the management of servers, storage, and network devices in data centers. It uses HTTPs-based RESTful APIs and JSON-formatted data exchange to achieve efficient, secure, and scalable management of hardware resources.

[0084] (7) SMBus (System Management Bus): SMBus is a low-speed, low-cost serial communication protocol designed specifically for system management and monitoring.

[0085] (8) SMBPBI (SMBus Post-Box Interface): SMBPBI is a software protocol implemented over SMBus, specifically designed for out-of-band management of NVIDIA data center GPUs. This technology allows system administrators or management software to set and monitor specific parameters of the GPU without directly interacting with the operating system kernel and driver.

[0086] (9) GPIO (General Purpose Input Output): GPIO is widely used in microcontrollers and embedded systems. It allows users to control specific pins on the chip through software to implement the function of input or output electrical signals.

[0087] In modern data centers and high-performance computing environments, server GPU modules play a core role, especially in applications such as artificial intelligence, deep learning, and massively parallel computing. As GPU computing power continues to improve, the demand for GPU performance monitoring and fault diagnosis is also growing.

[0088] Traditional data collection methods often rely on API calls at the operating system level or direct access to hardware registers, which requires a deep understanding of the internal structure of the GPU and may affect system performance or security. As the complexity of GPU modules increases, the process of collecting black box data from server GPU modules is more cumbersome, so the efficiency of collecting black box data from server GPU modules is not high. Therefore, how to provide a method for collecting black box data from server GPU modules to solve the problem of low efficiency of collecting black box data from server GPU modules is of great practical significance.

[0089] In order to solve the existing technical problems, the embodiment of the present application provides a method and device for collecting black box data of a server GPU module. In the process of collecting black box data of a server GPU module, the black box data collection instruction of the server GPU module is monitored; if the monitored black box data collection instruction is a preset first collection instruction, the black box data of the server GPU module is collected through Redfish, SMBPBI and SMBus; the first collection instruction is an instruction for collecting black box data triggered by an external request; if the monitored black box data collection instruction is a preset second collection instruction, the black box data of the server GPU module is collected through SMBPBI; the second collection instruction is an instruction for collecting black box data triggered by an interrupt alarm. This method can avoid direct access to hardware or deep dependence on drivers, and realize efficient collection of black box data of server GPU modules. The black box data collection process does not require an in-depth understanding of the internal details of the GPU, has low invasiveness and high compatibility, and can effectively improve the efficiency of black box data collection of server GPU modules.

[0090] In order to make the invention purpose, technical scheme and advantages of the embodiments of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.

[0091] The technical solution provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0092] The method for collecting black box data of a server GPU module provided in the embodiments of the present application may be applied to a server in some embodiments, and may also be applied to a terminal device in other embodiments. The following embodiments of the present application are described by taking the method for collecting black box data of a server GPU module applied to a server as an example.

[0093] The present application embodiment provides a method for collecting black box data of a server GPU module, such as Figure 1 As shown, the following steps are included:

[0094] Step S101, monitoring the black box data collection instructions of the server GPU module.

[0095] During specific implementation, the BMC of the server monitors the black box data collection instructions of the GPU module of the server.

[0096] In the embodiments of the present application, the black box data collection action may be actively triggered by a user or automatically triggered by the BMC.

[0097] In some embodiments of the present application, when there is an external request to actively collect black box data, the BMC will communicate with the module device to collect all black box data; when there is no external request, the BMC will continuously monitor the Alert GPIO status of the GPU module, and when an interrupt alarm is generated, the BMC collects the preset black box data for fault analysis.

[0098] Step S102, if the monitored black box data collection instruction is a preset first collection instruction, the black box data of the server GPU module is collected through Redfish, SMBPBI and SMBus; the first collection instruction is an instruction triggered by an external request to instruct the collection of black box data.

[0099] In specific implementation, if the monitored black box data collection instruction is the preset first collection instruction, the server's BMC collects the black box data of the server GPU module through Redfish, SMBPBI and SMBus; the first collection instruction is an instruction triggered by an external request to collect black box data.

[0100] In some optional embodiments, the process of collecting black box data of the server GPU module through Redfish, SMBPBI and SMBus in the above step S102 is as follows: Figure 2 As shown, this can be achieved by following the steps below:

[0101] Step S201, communicating with the module device of the server GPU module through Redfish, SMBPBI and SMBus.

[0102] During specific implementation, the BMC of the server communicates with the module device of the GPU module of the server through Redfish, SMBPBI and SMBus.

[0103] In some optional embodiments, the module device includes an HGX platform management controller HMC, a field programmable gate array FPGA, and a graphics processor GPU; in step S201, the process of communicating and connecting with the module device of the server GPU module through Redfish, SMBPBI and SMBus is as follows: Figure 3 As shown, this can be achieved by following the steps below:

[0104] Step S301: establish a first type of communication connection with the HMC through Redfish.

[0105] In specific implementation, the BMC of the server establishes a first-class communication connection with the HMC through Redfish.

[0106] Step S302, establish a second type of communication connection with the FPGA and GPU respectively through the SMBPBI.

[0107] During specific implementation, the BMC of the server respectively establishes the second type of communication connection with the FPGA and the GPU through SMBPBI.

[0108] Step S303: Establish the third type of communication connection with the HMC, FPGA and GPU respectively through the SMBus.

[0109] In specific implementation, the BMC of the server respectively establishes the third type of communication connection with the HMC, FPGA and GPU through the SMBus.

[0110] In some embodiments of the present application, the BMC directly communicates with the HMC using the SMBus, specifically to obtain SMBus debugging information.

[0111] In the method of this embodiment, the module devices include an HGX platform management controller HMC, a field programmable gate array FPGA, and a graphics processor GPU; a first type of communication connection is established with the HMC through Redfish; a second type of communication connection is established with the FPGA and GPU respectively through SMBPBI; and a third type of communication connection is established with the HMC, FPGA, and GPU respectively through SMBus, thereby providing a mechanism for communicating with the module devices of the server GPU module through Redfish, SMBPBI, and SMBus, further improving the efficiency of black box data collection of the server GPU module.

[0112] Step S202, based on the communication connection, collect black box data of the server GPU module.

[0113] During specific implementation, the BMC of the server collects black box data of the GPU module of the server based on the communication connection.

[0114] The method of this embodiment can establish a communication connection with the module device of the server GPU module through Redfish, SMBPBI and SMBus when collecting the black box data of the server GPU module, and then collect the black box data of the server GPU module based on the communication connection, thereby realizing comprehensive data collection of the GPU module, and effectively improving the efficiency of collecting the black box data of the server GPU module.

[0115] In some optional embodiments, in step S202, the process of collecting black box data of the server GPU module based on the communication connection is as follows: Figure 4 As shown, this can be achieved by following the steps below:

[0116] Step S401: If it is determined that the HMC is not abnormal based on the first type of communication connection, the black box data of the server GPU module is collected through the first type of communication connection.

[0117] In a specific implementation, if the BMC of the server determines that the HMC is not abnormal based on the first type of communication connection, the black box data of the GPU module of the server is collected through the first type of communication connection.

[0118] In some optional embodiments, in step S401, the process of collecting black box data of the server GPU module through the first type of communication connection is as follows: Figure 5 As shown, this can be achieved by following the steps below:

[0119] Step S501, obtaining first black box information from the HMC through a first type of communication connection; the first black box information includes Redfish interface information, GPU module self-test report, FPGA dump information, and HMC log.

[0120] In the embodiment of the present application, the first type of communication connection is that the BMC is connected to the HMC through Redfish.

[0121] In the embodiment of the present application, the FPGA dump information is data dumped from the FPGA to the HMC. The FPGA dump information is data that needs to be collected to determine if there is a problem with the GPU module hardware, and is provided to the GPU module manufacturer for analysis.

[0122] In some embodiments of the present application, the GPU module has a self-test function, such as SelfTest, which allows the user to test all devices and connection status on the GPU module, and can be tested from the following six aspects: test all power-related signals and device status around the target device; check whether the target device firmware is correctly booted and verified; check whether the pin status of the target device is normal; check all hardware connection status around the target device, such as I2C, SPI, PCIe, etc.; check whether the protocol of the target device is normal, such as SMBPBI; check whether the target device firmware is working properly. When the six-layer test is completed, a test result report for each device will be generated. Further, the test result report of each device is stored in the HMC partition as a GPU module self-test report, and the BMC can download the GPU module self-test report through Redfish. According to the GPU module self-test report, the abnormal device is preliminarily judged, and then the specific cause of the failure is determined through the Redfish interface information.

[0123] Step S502: adding the first black box information to the black box data of the server GPU module.

[0124] The method of this embodiment provides a mechanism for collecting black box data of a server GPU module through a first type of communication connection, thereby more efficiently improving the efficiency of collecting black box data of the server GPU module.

[0125] Step S402: If it is determined that the HMC is abnormal based on the first type of communication connection, the black box data of the server GPU module is collected through the second type of communication connection and the third type of communication connection.

[0126] In a specific implementation, if the BMC of the server determines that the HMC is abnormal based on the first type of communication connection, the black box data of the GPU module of the server is collected through the second type of communication connection and the third type of communication connection.

[0127] The method of this embodiment, if based on the first type of communication connection, it is judged that the HMC has no abnormality, then the black box data of the server GPU module is collected through the first type of communication connection; if based on the first type of communication connection, it is judged that the HMC has an abnormality, then the black box data of the server GPU module is collected through the second type of communication connection and the third type of communication connection. The black box data of the server GPU module can be efficiently collected based on the communication connection, and the black box data can be collected from three methods of Redfish, SMBPBI, and SMBus. The three methods can be redundant with each other, further improving the real-time nature of data collection and the stability of system operation, and can effectively improve the efficiency of collecting black box data of the server GPU module.

[0128] In some optional embodiments, in step S402, the process of collecting black box data of the server GPU module through the second type of communication connection and the third type of communication connection is as follows: Figure 6 As shown, this can be achieved by following the steps below:

[0129] Step S601, obtaining second black box information from FPGA and GPU through a second type of communication connection; the second black box information includes hardware status, firmware version, temperature and power consumption monitoring, and PCIe connection status.

[0130] In the embodiment of the present application, the second type of communication connection is the connection between the BMC and the FPGA and GPU respectively through the SMBPBI.

[0131] In some embodiments of the present application, the hardware status includes: GPU alarm information, onboard device alarm information, and device power supply status information; the firmware version refers to the firmware version of the GPU module component; temperature and power consumption monitoring is to collect temperature data and power consumption information of all module components; the PCIe connection status is the collected rate, bandwidth and error information of each PCIe connection of the GPU module.

[0132] Step S602: Obtain third black box information from the HMC, GPU, and FPGA through a third type of communication connection; the third black box information includes component temperature, component power consumption, and register information.

[0133] In the embodiment of the present application, the third type of communication connection is that the BMC is connected to the HMC, FPGA and GPU respectively through the SMBus.

[0134] In some embodiments of the present application, the FPGA, GPU, and HMC in the GPU module also support SMBus access, and the BMC can obtain component temperature, component power consumption, and register information through the SMBus.

[0135] In some embodiments of the present application, the register information includes FPGA register data and HMC register data.

[0136] Step S603: Add the second black box information and the third black box information to the black box data of the server GPU module.

[0137] In some embodiments of the present application, the BMC communicates directly with the GPU module device through SMBPBI and SMBus to collect black box data. The collected black box data is less than the black box data collected through Redfish. However, when collecting black box data directly with the GPU module device through SMBPBI and SMBus, it does not rely on the HMC. Therefore, when Redfish cannot collect information, it can be used as a redundant solution to ensure that sufficient black box data of the GPU module can be collected for fault analysis and location.

[0138] The method of this embodiment provides a mechanism for collecting black box data of a server GPU module through a second type of communication connection and a third type of communication connection, thereby more efficiently improving the efficiency of collecting black box data of the server GPU module.

[0139] In some embodiments of the present application, whether the HMC is abnormal is determined based on the state of the communication connection. In a specific implementation, whether the HMC is abnormal is identified by whether the state of the communication connection is normal communication.

[0140] In some optional embodiments, judging that the HMC is not abnormal based on the first type of communication connection includes:

[0141] If the status of the first type of communication connection is normal communication, it is determined that the HMC is not abnormal;

[0142] Based on the first type of communication connection, it is determined that the HMC is abnormal, including:

[0143] If the state of the first type of communication connection is abnormal communication, it is determined that the HMC is abnormal.

[0144] The method of this embodiment determines that the HMC is normal if the status of the first type of communication connection is normal communication; and determines that the HMC is abnormal if the status of the first type of communication connection is abnormal communication, thereby providing a mechanism for identifying whether the HMC is abnormal, and can accurately and efficiently determine whether the HMC is abnormal, thereby improving the efficiency of collecting black box data of the server GPU module.

[0145] Step S103, if the monitored black box data collection instruction is a preset second collection instruction, the black box data of the server GPU module is collected through SMBPBI; the second collection instruction is an instruction triggered by an interrupt alarm to collect black box data.

[0146] In specific implementation, if the monitored black box data collection instruction is the preset second collection instruction, the server's BMC collects the black box data of the server GPU module through SMBPBI; the second collection instruction is an instruction triggered by an interrupt alarm to collect black box data.

[0147] In some embodiments of the present application, when an alarm interrupt is detected in the Alert GPIO of the GPU module, a second collection instruction is generated, thereby triggering the BMC to automatically collect the GPU module black box information action. When an alarm interrupt occurs, the GPU module will be powered off in a short time, and the SMBPBI method is used to quickly collect black box data.

[0148] In some optional embodiments, the black box data of the server GPU module is collected through SMBPBI, specifically, the interrupt status black box information is obtained through SMBPBI; the interrupt status black box information includes device over-temperature information, GPU alarm information, onboard device alarm information, device power supply status information, and device PCIe connection information.

[0149] In some embodiments of the present application, the automatic collection black box process first initializes the Alert GPIO of the GPU module to the interrupt mode, and starts to monitor the GPIO in real time to see if an interrupt occurs; if no interrupt occurs, continue to monitor; if an interrupt occurs, start to collect the interrupt status black box information through SMBPBI as the collected black box data; save the collected black box data to a file, wait for the user to call the external interface to download the black box data file for problem location and repair, and then continue to monitor the AlertGPIO status for the next round of collection. The interrupt status black box information includes: device overtemperature information, used to determine whether there is an overtemperature alarm on a device; GPU alarm information, used to determine whether there is a GPU in-place abnormality, temperature alarm, PCIe abnormality; onboard device alarm information, used to determine whether there is an onboard device in-place abnormality, hard interrupt, peripheral abnormality, temperature alarm; device power supply status information, used to determine whether the device peripheral power supply device has a communication failure, overcurrent and overvoltage failure, or insufficient input voltage; device PCIe connection information, used to determine the device PCIe connection status and error information. Device PCIe connection information, including PCIe connection rate, bandwidth, PCIe protocol errors, and error count of each GPU module.

[0150] The method for collecting black box data of a server GPU module provided in an embodiment of the present application monitors the black box data collection instruction of the server GPU module during the process of collecting the black box data of the server GPU module; if the monitored black box data collection instruction is a preset first collection instruction, the black box data of the server GPU module is collected through Redfish, SMBPBI and SMBus; the first collection instruction is an instruction for collecting black box data triggered by an external request; if the monitored black box data collection instruction is a preset second collection instruction, the black box data of the server GPU module is collected through SMBPBI; the second collection instruction is an instruction for collecting black box data triggered by an interrupt alarm. This method can avoid direct access to hardware or deep reliance on drivers, and realize efficient collection of black box data of server GPU modules. The black box data collection process does not require an in-depth understanding of the internal details of the GPU, has low invasiveness and high compatibility, and can effectively improve the efficiency of black box data collection of server GPU modules.

[0151] Based on the same inventive concept, a device for collecting black box data of a server GPU module is also provided in the embodiment of the present application. Since the device corresponds to the method for collecting black box data of a server GPU module provided in the embodiment of the present application, and the principle of solving the problem by the device is similar to that of the method, the implementation of the device can refer to the implementation of the above method, and the repeated parts will not be repeated.

[0152] Figure 7 A schematic diagram of the structure of a device for collecting black box data of a server GPU module provided in an embodiment of the present application is shown. Figure 7 As shown, the device for collecting black box data of the server GPU module includes an instruction monitoring module 701 , a first acquisition module 702 and a second acquisition module 703 .

[0153] Among them, the instruction monitoring module 701 is used to monitor the black box data collection instructions of the server GPU module;

[0154] The first acquisition module 702 is used to collect the black box data of the server GPU module through Redfish, SMBPBI and SMBus if the monitored black box data collection instruction is a preset first collection instruction; the first collection instruction is an instruction triggered by an external request to instruct to collect black box data;

[0155] The second acquisition module 703 is used to collect the black box data of the server GPU module through SMBPBI if the monitored black box data collection instruction is a preset second collection instruction; the second collection instruction is an instruction to collect black box data triggered by an interrupt alarm.

[0156] In an optional embodiment, the first acquisition module 702 is specifically configured to:

[0157] Communicate with the module devices of the server GPU module through Redfish, SMBPBI and SMBus;

[0158] Based on the communication connection, collect the black box data of the server GPU module.

[0159] In an optional embodiment, the module device includes an HGX platform management controller HMC, a field programmable gate array FPGA, and a graphics processor GPU; the first acquisition module 702 is specifically used to:

[0160] First class communication connection with HMC via Redfish;

[0161] The second type of communication connection is respectively performed with the FPGA and the GPU through the SMBPBI;

[0162] The third type of communication connection is carried out with HMC, FPGA and GPU respectively through SMBus.

[0163] In an optional embodiment, the first acquisition module 702 is specifically configured to:

[0164] If it is determined that the HMC is not abnormal based on the first type of communication connection, the black box data of the server GPU module is collected through the first type of communication connection;

[0165] If the HMC is judged to be abnormal based on the first type of communication connection, the black box data of the server GPU module is collected through the second type of communication connection and the third type of communication connection.

[0166] In an optional embodiment, the first acquisition module 702 is specifically configured to:

[0167] If the status of the first type of communication connection is normal communication, it is determined that the HMC is not abnormal;

[0168] The first acquisition module 702 is specifically used for:

[0169] If the state of the first type of communication connection is abnormal communication, it is determined that the HMC is abnormal.

[0170] In an optional embodiment, the first acquisition module 702 is specifically configured to:

[0171] Obtain first black box information from the HMC through the first type of communication connection; the first black box information includes Redfish interface information, GPU module self-test report, FPGA dump information, and HMC log;

[0172] The first black box information is added to the black box data of the server GPU module.

[0173] In an optional embodiment, the first acquisition module 702 is specifically configured to:

[0174] Obtain second black box information from FPGA and GPU through the second type of communication connection; the second black box information includes hardware status, firmware version, temperature and power consumption monitoring, and PCIe connection status;

[0175] Obtain third black box information from the HMC, GPU, and FPGA through a third type of communication connection; the third black box information includes component temperature, component power consumption, and register information;

[0176] The second black box information and the third black box information are added to the black box data of the server GPU module.

[0177] In an optional embodiment, the second acquisition module 703 is specifically used to obtain interrupt status black box information through SMBPBI; the interrupt status black box information includes device over-temperature information, GPU alarm information, onboard device alarm information, device power supply status information, and device PCIe connection information.

[0178] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application. The electronic device can be used to collect black box data of a server GPU module. In the embodiment of the present application, the electronic device can be a server or a terminal device. In one embodiment, the electronic device can be a server. In this embodiment, the structure of the electronic device can be as follows: Figure 8 As shown, it includes a memory 801 , a communication module 803 and one or more processors 802 .

[0179] The memory 801 is used to store computer programs executed by the processor 802. The memory 801 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and programs required for running the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0180] The memory 801 may be a volatile memory, such as a random-access memory (RAM); the memory 801 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), or the memory 801 may be any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 801 may be a combination of the above memories.

[0181] The processor 802 may include one or more central processing units (CPU) or a digital processing unit, etc. The processor 802 is used to implement the above-mentioned method for collecting black box data of the server GPU module when calling the computer program stored in the memory 801.

[0182] The communication module 803 is used to communicate with the server and other terminal devices.

[0183] The specific connection medium between the memory 801, the communication module 803 and the processor 802 is not limited in the embodiment of the present application. Figure 8In the embodiment, the memory 801 and the processor 802 are connected via a bus 804. The bus 804 is Figure 8 The connections between other components are shown in bold lines, and are not intended to be limiting. Bus 804 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0184] According to one aspect of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for collecting black box data of the server GPU module in the above embodiment. The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0185] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A method for collecting black box data of a server GPU module, characterized in that: The method comprises: Monitoring black box data collection instructions of the server GPU module; If the monitored black box data collection instruction is a preset first collection instruction, the black box data of the server GPU module is collected through the Redfish standard Redfish, the system management bus mailbox interface SMBPBI and the system management bus SMBus; the first collection instruction is an instruction triggered by an external request to instruct the collection of black box data; If the black box data collection instruction monitored is a preset second collection instruction, the black box data of the server GPU module is collected through SMBPBI; the second collection instruction is an instruction triggered by an interrupt alarm to collect black box data.

2. The method according to claim 1, characterized in that: The black box data of the server GPU module is collected through Redfish, SMBPBI and SMBus, including: Communicate and connect with the module device of the server GPU module through Redfish, SMBPBI and SMBus; Based on the communication connection, black box data of the server GPU module is collected.

3. The method according to claim 2, characterized in that The module device includes an HGX platform management controller HMC, a field programmable gate array FPGA, and a graphics processor GPU; the module device of the server GPU module is communicated and connected through Redfish, SMBPBI and SMBus, including: Establishing a first type of communication connection with the HMC via Redfish; Performing a second type of communication connection with the FPGA and the GPU respectively through SMBPBI; The third type of communication connection is respectively established with the HMC, the FPGA and the GPU through the SMBus.

4. The method according to claim 3, characterized in that The collecting black box data of the server GPU module based on the communication connection includes: If it is determined that the HMC is not abnormal based on the first type of communication connection, then collecting black box data of the server GPU module through the first type of communication connection; If it is determined that the HMC is abnormal based on the first type of communication connection, the black box data of the server GPU module is collected through the second type of communication connection and the third type of communication connection.

5. The method according to claim 4, characterized in that The determining that the HMC is not abnormal based on the first type of communication connection includes: If the state of the first type of communication connection is normal communication, it is determined that the HMC is not abnormal; The determining, based on the first type of communication connection, that the HMC is abnormal, includes: If the state of the first type of communication connection is abnormal communication, it is determined that the HMC is abnormal.

6. The method according to claim 4, characterized in that The collecting the black box data of the server GPU module through the first type of communication connection includes: Obtaining first black box information from the HMC through the first type of communication connection; the first black box information includes Redfish interface information, GPU module self-test report, FPGA dump information, and HMC log; The first black box information is added to the black box data of the server GPU module.

7. The method according to claim 4, characterized in that The collecting the black box data of the server GPU module through the second type of communication connection and the third type of communication connection includes: Obtaining second black box information from the FPGA and the GPU through the second type of communication connection; the second black box information includes hardware status, firmware version, temperature and power consumption monitoring, and PCIe connection status; Obtaining third black box information from the HMC, the GPU, and the FPGA through the third type of communication connection; the third black box information includes component temperature, component power consumption, and register information; The second black box information and the third black box information are added to the black box data of the server GPU module.

8. A device for collecting black box data of a server GPU module, characterized in that: The device comprises: An instruction monitoring module, used to monitor the black box data collection instructions of the server GPU module; A first acquisition module is used to collect the black box data of the server GPU module through Redfish, SMBPBI and SMBus if the monitored black box data collection instruction is a preset first collection instruction; the first collection instruction is an instruction triggered by an external request to instruct to collect black box data; The second acquisition module is used to collect the black box data of the server GPU module through SMBPBI if the black box data collection instruction monitored is a preset second collection instruction; the second collection instruction is an instruction triggered by an interrupt alarm to collect black box data.

9. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.