An abnormality positioning system, method, electronic device, and storage medium
By introducing a high-speed serial switching module and a conversion module into the baseboard management controller, parallel diagnostic information acquisition from multiple graphics processors is achieved, solving the problem of low fault location efficiency in multi-graphics processor environments and improving the real-time performance and accuracy of anomaly location.
Patent Information
- Application Number
- CN202411534263.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-10-30
AI Technical Summary
In a multi-GPU environment, the board management controller cannot simultaneously capture diagnostic information from multiple GPUs, resulting in low fault location efficiency. Traditional debugging methods cannot support high-concurrency signal transmission, limiting the real-time performance and accuracy of anomaly location.
By employing a high-speed serial exchange module, a conversion module, a serial port interface module, and a joint test working group interface module, the baseboard management controller can acquire register information and status information of multiple graphics processors in parallel. The high-speed serial exchange module extends the control commands and sends them to the conversion module in parallel, and integrates the diagnostic information returned in parallel for anomaly localization.
It enables centralized management and control of multiple graphics processors, quickly locates the source of faults, reduces troubleshooting time and manpower costs, and improves system operating efficiency and economic benefits.
Smart Images

Figure CN119046053B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of server anomaly detection, and particularly relates to an anomaly positioning system and method, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development of artificial intelligence, big data processing and high-performance computing, graphics processors play an increasingly important role in data processing and computing tasks. Modern graphics processor servers are usually configured with multiple graphics processors to meet the needs of high concurrency and high performance. However, as the number of graphics processors increases, the complexity of the system also rises, making it more difficult to locate problems when a graphics processor fails or malfunctions.
[0003] In related technologies, the baseboard management controller is usually responsible for monitoring and managing the hardware state of the server. However, in a multi-graphics processor environment, the baseboard management controller can only obtain the state information of the graphics processors one by one and cannot simultaneously capture the diagnostic information of multiple graphics processors, making it necessary for engineers to obtain fault information by repeatedly reproducing the problem, shutting down and restarting, and other tedious steps when a graphics processor fails, resulting in low debugging efficiency and thus prolonging the troubleshooting time. In addition, traditional debugging methods often rely on traditional serial communication and joint test action group interfaces, which cannot effectively support high-concurrency signal transmission, limiting the real-time and accuracy of anomaly positioning. SUMMARY
[0004] The embodiments of the present disclosure provide an anomaly positioning system, method, electronic device and storage medium, aiming to solve the problems in the background art.
[0005] To solve the above technical problems, the present disclosure is implemented as follows:
[0006] In a first aspect, the embodiments of the present disclosure provide an anomaly positioning system, which comprises:
[0007] a baseboard management controller, configured to connect a controller device;
[0008] a plurality of graphics processors, each graphics processor having a joint test action group interface and a serial interface;
[0009] the controller device comprising a high-speed serial switching module, a first conversion module, a second conversion module, a serial interface module and a joint test action group interface module;
[0010] an input end of the high-speed serial switching module is connected with the baseboard management controller, and output ends of the high-speed serial switching module are connected with the first conversion module and the second conversion module respectively;
[0011] The serial interface module is configured to connect the serial interface of each of the plurality of graphic processors; and the joint test working group interface module is configured to connect the joint test working group interface of each of the plurality of graphic processors.
[0012] The baseboard management controller is configured to acquire, in parallel, the register information from the joint test working group interface of the plurality of graphic processors and the state information of the serial interface returned by the high-speed serial switching module, to perform abnormality positioning.
[0013] Optionally, the baseboard management controller has a first high-speed serial interface connected to a controller device through a first high-speed serial signal line; the input end of the high-speed serial switching module has a second high-speed serial interface, the output end has a third high-speed serial interface and a fourth high-speed serial interface, the first high-speed serial signal line is connected between the second high-speed serial interface and the first high-speed serial interface of the baseboard management controller, the third high-speed serial interface is connected to the first conversion module, and the fourth high-speed serial interface is connected to the second conversion module.
[0014] The joint test working group interfaces on the plurality of graphic processors include high-potential joint test working group interfaces and low-potential joint test working group interfaces.
[0015] The joint test working group interface module includes a first joint test working group interface module and a second joint test working group interface module.
[0016] The first joint test working group interface module is connected in parallel to the low-potential joint test working group interfaces on the plurality of graphic processors, and the second joint test working group interface module is configured to be connected in parallel to the high-potential joint test working group interfaces on the plurality of graphic processors.
[0017] Optionally, the input end of the first conversion module has a fifth high-speed serial interface, and the output end has a first joint test working group interface and a second joint test working group interface.
[0018] The fifth high-speed serial interface of the first conversion module is connected to the third high-speed serial interface of the high-speed serial switching module through a second high-speed serial signal line.
[0019] The first joint test working group interface of the first conversion module is connected to the first joint test working group interface module through a first joint test working group signal line, and the second joint test working group interface of the first conversion module is connected to the second joint test working group interface module through a second joint test working group signal line.
[0020] Optionally, the input end of the second conversion module has a sixth high-speed serial interface, and the output end has a serial interface.
[0021] The sixth high-speed serial interface of the second conversion module is connected with the fourth high-speed serial interface of the high-speed serial exchange module through a third high-speed serial signal line, and the serial interface of the second conversion module is connected with the serial interface module through a serial signal line.
[0022] Optionally, the baseboard management controller is configured with a first storage space, and a diagnostic firmware is stored in the first storage space, the diagnostic firmware being used to perform abnormality positioning according to register information of a joint test action group interface and state information of a serial interface of the plurality of graphic processors.
[0023] The baseboard management controller is configured with a second storage space, and a baseboard management controller firmware is stored in the second storage space, the baseboard management controller firmware having a priority different from that of the diagnostic firmware.
[0024] In a second aspect, the embodiments of the present disclosure provide an abnormality positioning method applied to an abnormality positioning system, and the method comprises the following steps.
[0025] Register information and state information of at least one first graphic processor in which an abnormality occurs in a plurality of graphic processors of a server are received in parallel through a joint test action group interface module and a serial interface module, respectively, and the register information and the state information are returned to a first conversion module and a second conversion module in parallel;
[0026] The register information and the state information are converted into a protocol and a data format compatible with a high-speed serial conversion module through the first conversion module and the second conversion module, respectively, and are returned to the high-speed serial conversion module;
[0027] The register information and the state information are integrated into diagnostic information through the high-speed serial exchange module, and the diagnostic information is returned to a baseboard management controller;
[0028] The baseboard management controller performs fault positioning according to the diagnostic information to determine an abnormality cause of the first graphic processor.
[0029] Optionally, the baseboard management controller performs fault positioning according to the diagnostic information to determine an abnormality cause of the first graphic processor, and the method comprises the following steps.
[0030] The baseboard management controller determines a register value of an over-temperature register used to monitor a temperature state according to the register information of the first graphic processor.
[0031] It is determined whether the register value of the over-temperature register is set, and in a case where the register value of the over-temperature register is set, a current temperature of the first graphic processor returned through the high-speed serial exchange module is acquired.
[0032] In a case where the current temperature of the first graphics processor is greater than a preset temperature threshold, it is determined that the abnormal cause of the first graphics processor is an over-temperature abnormality.
[0033] Optionally, in a case where the register value of the over-temperature register is not set, it is determined that the first graphics processor does not have an over-temperature abnormality.
[0034] Optionally, the method further comprises:
[0035] In a case where the current temperature of the first graphics processor is greater than a preset temperature threshold, the current temperature of the second graphics processor returned via the high-speed serial switching module is acquired by the baseboard management controller, the second graphics processor being a graphics processor deployed in a position adjacent to the first graphics processor.
[0036] In a case where the current temperature of the second graphics processor is greater than the preset temperature threshold, it is determined that the abnormal cause of the first graphics processor is an ambient temperature abnormality.
[0037] In a case where the current temperature of the second graphics processor is less than or equal to the preset temperature threshold, it is determined that the abnormal cause of the first graphics processor is a self-temperature abnormality.
[0038] Optionally, the fault location according to the diagnostic information by the baseboard management controller to determine the abnormal cause of the first graphics processor comprises:
[0039] The register value of a logic device used for logical control of the first graphics processor is determined by the baseboard management controller according to the register information of the first graphics processor.
[0040] It is determined whether the register value of the logic device is set, and in a case where the register value of the logic device is set, it is determined that the abnormal cause of the first graphics processor is a power supply abnormality.
[0041] Optionally, the method further comprises:
[0042] In a case where the register value of the logic device is set, the register value of a key alarm register used for monitoring the power-off protection state of the first graphics processor returned via the high-speed serial switching module is acquired by the baseboard management controller.
[0043] It is determined whether the register value of the key alarm register is set, and in a case where the register value of the key alarm register is set, it is determined that the abnormal cause of the first graphics processor is a power supply abnormality caused by a power-off protection abnormality of the first graphics processor.
[0044] In a case where the register value of the critical alarm register is not set, it is determined that the abnormality of the first graphics processor is a power supply abnormality caused by a non-self abnormality of the first graphics processor.
[0045] Optionally, the method further comprises:
[0046] In a case where the register value of the critical alarm register is not set, it is determined that the abnormality of the first graphics processor is a power supply abnormality caused by a non-self abnormality of the first graphics processor.
[0047] In a case where the register value of the critical alarm register is not set, it is determined that the abnormality of the first graphics processor is a power supply abnormality caused by a non-self abnormality of the first graphics processor.
[0048] In a case where the register value of the critical alarm register is not set, it is determined that the abnormality of the first graphics processor is a power supply abnormality caused by a non-self abnormality of the first graphics processor.
[0049] Optionally, the method further comprises:
[0050] In a case where the register value of the critical alarm register is not set, it is determined that the abnormality of the first graphics processor is a power supply abnormality caused by a non-self abnormality of the first graphics processor.
[0051] In a case where the register value of the critical alarm register is not set, it is determined that the abnormality of the first graphics processor is a power supply abnormality caused by a non-self abnormality of the first graphics processor.
[0052] In a case where the register value of the critical alarm register is not set, it is determined that the abnormality of the first graphics processor is a power supply abnormality caused by a non-self abnormality of the first graphics processor.
[0053] In the case that the communication link of the first graphics processor is abnormal, all graphics processors with abnormal communication links are determined; all graphics processors with abnormal communication links are exchanged with the first graphics processor position one by one, and it is checked whether the communication link is normal after the position exchange; in the case that the communication link is normal after the position exchange, it is determined that the abnormal reason of the first graphics processor is the communication link abnormality caused by self-abnormality; in the case that the communication link is not normal after the position exchange, it is determined that the abnormal reason of the first graphics processor is the communication link abnormality caused by link configuration abnormality.
[0054] In a third aspect, the embodiments of the present disclosure provide an electronic device, comprising a processor, a memory, and a computer program stored in the memory and capable of running on the processor, when the computer program is executed by the processor, the steps of an abnormality positioning method are implemented.
[0055] In a fourth aspect, the embodiments of the present disclosure provide a computer readable storage medium, the computer readable storage medium stores a computer program, when the computer program is executed by a processor, the steps of an abnormality positioning method are implemented.
[0056] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0057] The present disclosure can obtain the register information and state information of multiple graphics processors in parallel, avoiding the inefficient process of obtaining data one by one in the traditional method. When a fault occurs, the state information of multiple graphics processors can be monitored and analyzed at the same time, the fault source can be quickly located, and the troubleshooting time is reduced. In addition, the high-speed serial exchange module, the conversion module, and the serial and joint test working group interface module built-in the control device realize the centralized management and control of multiple graphics processors. Based on the integrated design of the present disclosure, not only the hardware connection is simplified, but also the downtime and labor cost caused by fault troubleshooting are reduced through fast and accurate abnormality positioning, thereby improving the operation efficiency and economic benefit of the overall system. BRIEF DESCRIPTION OF DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0059] Figure 1 It is a hardware topology schematic diagram of a graphics processor substrate based on a serial interface in related technologies;
[0060] Figure 2 is a hardware topology diagram of a graphics processor substrate based on a joint test action group interface in the related art;
[0061] Figure 3 is a hardware topology diagram of an abnormality positioning system provided by one embodiment of the present disclosure;
[0062] Figure 4 is a step diagram of an abnormality positioning method provided by one embodiment of the present disclosure;
[0063] Figure 5 is an abnormality positioning flowchart of an over-temperature abnormality in one embodiment of the present disclosure;
[0064] Figure 6 is an abnormality positioning flowchart of an unrecognized abnormality in one embodiment of the present disclosure. DETAILED DESCRIPTION
[0065] Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present disclosure. In the description of the embodiments of the present disclosure, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; "and / or" herein is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In the present disclosure, "at least one" means one or more, and "multiple" means two or more. "At least one" or the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0066] Fault diagnosis and positioning of a graphics processor is a challenge that the growing server industry must face. Although traditional graphics processor abnormality diagnosis methods provide basic functions, in a high-density and high-performance server environment, the diagnosis and management are inefficient. Figure 1 is a hardware topology diagram of a graphics processor substrate based on a serial interface in the related art, as shown in Figure 1As shown, the solution has multiple operation, administration, and maintenance card slots, each of which can be inserted with a graphics processor. Each graphics processor has its own serial interface for outputting diagnostic information. Since the number of interfaces of the baseboard management controller is limited, it is not possible to provide each graphics processor with an independent serial connection to the baseboard management controller when there are multiple graphics processors in the server. Therefore, a serial multiplexer is used to switch between multiple serial signals, allowing the baseboard management controller to access the diagnostic information of any graphics processor. The baseboard management controller is responsible for monitoring all graphics processor diagnostic information output through the serial port. When a problem occurs, the baseboard management controller can switch different serial lines as needed to capture error information from a specific graphics processor. However, when multiple graphics processors have problems at the same time, the baseboard management controller cannot simultaneously capture the serial information of all graphics processors due to hardware and software limitations. This results in the inability to monitor the status of all graphics processors simultaneously in an emergency. In a multi-graphics processor system, capturing and analyzing diagnostic information requires switching the serial output of each graphics processor one by one, increasing the diagnostic time. In particular, when multiple errors occur, multiple switching and testing are required to locate the problem. If it is necessary to simultaneously connect the serial information of multiple graphics processors, it is necessary to shut down the server, open the case, and physically connect the serial signal lines of each graphics processor, which is not only time-consuming but also inefficient. And the use of traditional serial multiplexer and serial converter solutions is limited by hardware resources, especially in high-density graphics processor deployment environments, the application limitations are more obvious.
[0067] Figure 2 A hardware topology diagram of a graphics processor baseboard based on the Joint Test Action Group interface in related art is shown in FIG. 1. Figure 2 As shown, in a multi-graphics processor environment, the baseboard management controller can only access a single graphics processor or its CPLD (Complex Programmable Logic Device) / FPGA (Field Programmable Gate Array) through one Joint Test Action Group connection. If you want to access information from another graphics processor, you need to reconfigure the connection or physically switch to another Joint Test Action Group chain, limiting the ability to debug multiple graphics processors simultaneously. Since only one graphics processor can be connected at a time, when multiple graphics processors have problems or need to be debugged in parallel, the efficiency of debugging will be significantly reduced, and the time to locate and solve problems will be extended, especially when complex or cross-impact faults occur. In addition, Joint Test Action Group debugging requires physical interfaces such as specific Joint Test Action Group headers or adapters. In a dense server environment, frequent access and modification of physical connections not only takes time, but also increases the complexity of operations and the likelihood of errors.
[0068] In summary, the embodiment of the present disclosure provides an abnormal positioning system, which aims to capture relevant fault information and determine the problem of the corresponding graphics processor abnormality through logical analysis and graphic processing. Figure 3 is a hardware topology diagram of an abnormal positioning system provided by one embodiment of the present disclosure, as shown in Figure 3 The system comprises:
[0069] The system comprises:
[0070] The substrate management controller is configured to connect the controller device.
[0071] The plurality of graphics processors each have a joint test action group interface and a serial interface.
[0072] The controller device comprises a high-speed serial switching module, a first conversion module, a second conversion module, a serial interface module, and a joint test action group interface module.
[0073] The input end of the high-speed serial switching module is connected with the substrate management controller, and the output end of the high-speed serial switching module is connected with the first conversion module and the second conversion module respectively.
[0074] The serial interface module is configured to connect the serial interfaces of the plurality of graphics processors respectively, and the joint test action group interface module is configured to connect the joint test action group interfaces of the plurality of graphics processors respectively.
[0075] The substrate management controller is configured to acquire the register information from the joint test action group interfaces and the state information from the serial interfaces of the plurality of graphics processors returned by the high-speed serial switching module in parallel, so as to perform abnormal positioning.
[0076] It should be noted that the abnormal positioning system of the present disclosure realizes bidirectional signal interaction, that is, the baseboard management controller can simultaneously issue relevant control instructions to the specified abnormal graphics processor, and can also simultaneously receive the diagnostic information (register information and state information) returned by the multiple abnormal graphics processors to perform abnormal positioning. In the embodiment of the present disclosure, the baseboard management controller is the core control unit of the entire system, responsible for monitoring and managing the diagnostic information of all graphics processors, including the register information from the joint test action group interface of the multiple graphics processors and the state information from the serial interface. The baseboard management controller communicates with other components through a high-speed serial bus to realize parallel monitoring of the state of the graphics processor. The core idea of the present disclosure is to expand the control instructions from the baseboard management controller through a high-speed serial exchange module, and issue them to the first conversion module and the second conversion module in parallel. Or integrate the register information and state information from the first conversion module and the second conversion module and return them to the baseboard management controller. Parallel acquisition of the diagnostic information of multiple graphics processors avoids the single-threaded monitoring problem caused by the serial port multiplexer switching in the traditional scheme, thereby improving the efficiency of fault diagnosis.
[0077] For multiple graphics processors, each graphics processor has a serial interface and a joint test action group interface, Figure 3 The number of graphics processors shown is 8, which is a typical configuration for a graphics processor server. In the abnormal positioning process, the two interfaces are used for different diagnosis and positioning tasks. The serial interface is used to output the state information of the graphics processor, such as temperature, fault state, etc., while the joint test action group interface is used to obtain the information of the related registers of the graphics processor. In the traditional scheme, the baseboard management controller cannot simultaneously monitor the serial port information of multiple graphics processors or perform joint test action group debugging due to interface limitations, but this limitation is overcome in the present disclosure.
[0078] The control device is the core of the abnormal positioning system in the embodiment of the present disclosure, responsible for coordinating and converting data communication between the baseboard management controller and the plurality of graphic processors. The control device can be an FPGA. The modules deployed on the control device include a high-speed serial switching module, a first conversion module, a second conversion module, a serial interface module, and a joint test working group interface module. The high-speed serial switching module is used to realize high-speed data exchange in the system. The input end thereof is connected with the baseboard management controller, and the output end thereof is connected to the first conversion module and the second conversion module respectively, so as to convert different types of interface data into high-speed serial data. The first conversion module and the second conversion module are respectively responsible for converting different types of data into a format that can be processed by the high-speed serial switching module, so as to uniformly manage and transmit. The serial interface module is used to connect the serial interfaces of the plurality of graphic processors to collect state information. The joint test working group interface module is used to connect the joint test working group interfaces of the plurality of graphic processors to collect register information.
[0079] The present disclosure can obtain joint test working group and state information of the plurality of graphic processors in parallel, avoiding the inefficient process of obtaining data one by one in the traditional method. When a fault occurs, the state information of the plurality of graphic processors can be monitored and analyzed simultaneously, the fault source can be quickly located, and the troubleshooting time can be reduced. In addition, the high-speed serial switching module, the conversion module, and the serial and joint test working group interface module built-in the control device realize centralized management and control of the plurality of graphic processors. Based on the integrated design of the present disclosure, not only the hardware connection is simplified, but also the downtime and labor cost caused by troubleshooting are reduced through fast and accurate abnormal positioning, thereby improving the operation efficiency and economic benefit of the overall system.
[0080] Exemplarily, the baseboard management controller has a first high-speed serial interface connected with the control device through a first high-speed serial signal line; the input end of the high-speed serial switching module has a second high-speed serial interface, and the output end thereof has a third high-speed serial interface and a fourth high-speed serial interface; the first high-speed serial signal line is connected between the second high-speed serial interface and the first high-speed serial interface of the baseboard management controller; the third high-speed serial interface is connected with the first conversion module; and the fourth high-speed serial interface is connected with the second conversion module.
[0081] The baseboard management controller has a high speed serial interface for efficient data transfer with the high speed serial switch module. The high speed serial switch module has multiple high speed serial interfaces, including a second high speed serial interface for receiving high speed data requests and commands from the baseboard management controller; a third high speed serial interface (output) connected to the first conversion module for sending commands from the baseboard management controller to the first conversion module, or receiving register information returned from the first conversion module; a fourth high speed serial interface (output) connected to the second conversion module for performing similar functions, responsible for another parallel path for sending commands from the baseboard management controller to the second conversion module, or receiving status information returned from the second conversion module. The first high speed serial signal line connects the first high speed serial interface of the baseboard management controller and the second high speed serial interface of the high speed serial switch module, and this transmission path is used to transmit commands from the baseboard management controller, or to transmit integrated register information and status information. The third high speed serial interface is connected to the first conversion module, and this transmission path is used to convert high speed serial data into Joint Test Action Group interface compatible protocols and formats, or to convert register information from the Joint Test Action Group interface module into protocols and formats compatible with the high speed serial switch module. The fourth high speed serial interface is connected to the second conversion module for serial interface status information. The baseboard management controller is connected to the high speed serial switch module through its high speed serial interface, responsible for issuing commands and receiving data from different modules. The high speed serial switch module distributes data from the baseboard management controller to two different conversion modules (first and second conversion modules) for processing different formats of data.
[0082] Exemplarily, the Joint Test Action Group interfaces on the plurality of graphic processors include high potential Joint Test Action Group interfaces and low potential Joint Test Action Group interfaces;
[0083] The Joint Test Action Group interface module includes a first Joint Test Action Group interface module and a second Joint Test Action Group interface module;
[0084] The first Joint Test Action Group interface module is connected in parallel to the low potential Joint Test Action Group interfaces on the plurality of graphic processors, and the second Joint Test Action Group interface module is used to connect in parallel to the high potential Joint Test Action Group interfaces on the plurality of graphic processors.
[0085] On multiple graphics processors, the joint test action group interface is divided into two types of high potential and low potential, and the two interfaces differ in electrical characteristics, signal transmission rate and power consumption. The high potential joint test action group interface is used for debugging operations that require higher voltage signals, provides stronger signal integrity, and is suitable for high-speed data transmission. The low potential joint test action group interface is used for low power operation and is suitable for basic debugging tasks. The second joint test action group interface module and the first joint test action group interface module are responsible for connecting the high potential and low potential joint test action group interfaces respectively. Among them, the second joint test action group interface is used to connect the high potential joint test action group interfaces of multiple graphics processors in parallel, and the baseboard management controller can simultaneously access the high potential joint test action group interfaces of multiple graphics processors to perform efficient debugging and data collection, significantly improving the debugging efficiency, especially when rapid fault location is required, the register information of multiple graphics processors can be obtained simultaneously. The first joint test action group interface module is used to connect the low potential joint test action group interfaces of multiple graphics processors in parallel. Similarly, the first joint test action group module also accesses multiple graphics processors in parallel, allowing the baseboard management controller to simultaneously access multiple graphics processors during low potential abnormal positioning debugging. Efficient resource utilization is achieved during abnormal positioning, and the baseboard management controller can process register information from multiple graphics processors at the same time, avoiding the inefficient process of accessing one by one.
[0086] The present disclosure not only improves the response speed of debugging, but also reduces the complexity of physical connection and reduces the operation errors that may be caused by frequent access and modification of connection. In the entire abnormal positioning system, the joint test action group interface module is an important component of the controller device, which cooperates with the high-speed serial switching module and the serial interface module to work together to realize centralized management and control of multiple graphics processors, so that the baseboard management controller can obtain the register information of the graphics processors in parallel, quickly analyze and locate the abnormal source, and improve the overall abnormal processing capability and efficiency.
[0087] Exemplarily, the input end of the first conversion module has a fifth high-speed serial interface, and the output end has a first joint test action group interface and a second joint test action group interface;
[0088] The fifth high-speed serial interface of the first conversion module is connected with the third high-speed serial interface of the high-speed serial switching module through a second high-speed serial signal line;
[0089] The first joint test action group interface of the first conversion module is connected with the first joint test action group interface module through a first joint test action group signal line, and the second joint test action group interface of the first conversion module is connected with the second joint test action group interface module through a second joint test action group signal line.
[0090] The first conversion module is responsible for converting data from the high-speed serial switching module into a data format suitable for the joint test action group interface, aiming to achieve efficient transmission and format conversion of data to facilitate subsequent abnormal diagnosis and positioning. The fifth high-speed serial interface is the input end of the first conversion module and is specially used to receive data streams from the high-speed serial switching module. Through the fifth high-speed serial interface, the first conversion module obtains diagnostic data of multiple graphics processors. The fifth high-speed serial interface is connected with the third high-speed serial interface of the high-speed serial switching module through a second high-speed serial signal line. It should be noted that the high-speed serial signal lines in the abnormal positioning system provided by the embodiments of the present disclosure all support high-bandwidth and low-latency data transmission, and can effectively transmit control signals and data high-speed serial bus.
[0091] The output end of the first conversion module includes a first joint test action group interface and a second joint test action group interface. The first joint test action group interface is connected to the first joint test action group interface module through a first joint test action group signal line. The first joint test action group interface is responsible for transmitting the converted joint test action group signals to the first joint test action group interface module for debugging and diagnosing a specific graphics processor. The second joint test action group interface is connected to the second joint test action group interface module through a second joint test action group signal line. The second joint test action group interface is responsible for transmitting another group of converted joint test action group signals to the second joint test action group interface module to support debugging and diagnosing tasks for other graphics processors.
[0092] The first conversion module is used to convert the input high-speed serial signals into a data format suitable for the joint test action group interface. Specifically, the first conversion module parses the high-speed serial signals into a format required by the joint test action group protocol, ensuring that the data can be correctly parsed and used. After receiving the data from the high-speed serial switching module, the data is processed according to the preset protocol and format to ensure that the output signal meets the requirements of the joint test action group interface. The first conversion module can simultaneously receive register information from multiple graphics processors. These data streams are processed collectively, rather than relying on a switching connection one by one. That is, at the same time, the baseboard management controller can route the first conversion module to obtain state information and diagnostic data of multiple graphics processors, thereby improving the efficiency of monitoring.
[0093] Exemplarily, the input end of the second conversion module has a sixth high-speed serial interface, and the output end has a serial interface;
[0094] The sixth high-speed serial interface of the second conversion module is connected with the fourth high-speed serial interface of the high-speed serial switching module through a third high-speed serial signal line, and the serial interface of the second conversion module is connected with the serial interface module through a serial signal line.
[0095] Similar to the first conversion module, the second conversion module is mainly responsible for converting data from the high-speed serial switching module into a data format suitable for the serial interface, aiming to realize efficient transmission and format conversion of data, so as to facilitate subsequent abnormal diagnosis and state monitoring. The sixth high-speed serial interface is the input end of the second conversion module, used to receive data streams from the high-speed serial switching module. Through the sixth high-speed serial interface, the second conversion module can obtain serial diagnostic data of multiple graphic processors. The sixth high-speed serial interface is connected with the fourth high-speed serial interface of the high-speed serial switching module through a third high-speed serial signal line. The third high-speed serial signal line is also part of the high-speed serial bus, supporting high-bandwidth and low-latency data transmission. The output end of the second conversion module includes a serial interface, which is responsible for transmitting the converted serial signals to the serial interface module. Through the serial interface, the second conversion module can transmit state information (such as temperature, abnormal state, etc.) from multiple graphic processors to the serial interface module for centralized processing and monitoring.
[0096] The second conversion module closely cooperates with the high-speed serial switching module, the serial interface module and the first conversion module, ensuring that the state information from multiple graphic processors can be effectively collected and transmitted, supporting parallel abnormal diagnosis and state monitoring. By delivering signals from the high-speed serial switching module to the serial interface module, parallel monitoring of multiple graphic processors is realized, avoiding the low efficiency caused by acquiring data one by one in traditional methods.
[0097] Exemplarily, the baseboard management controller is configured with a first storage space storing a diagnostic firmware, the diagnostic firmware being used for abnormal positioning according to register information of a joint test action group interface of the multiple graphic processors and state information of a serial interface; the baseboard management controller is configured with a second storage space storing a baseboard management controller firmware, the priority of the baseboard management controller firmware being different from the priority of the diagnostic firmware.
[0098] As mentioned above, the baseboard management controller is the core control unit of the entire abnormal positioning system, responsible for monitoring and managing the state information of multiple graphics processors. Through communication with the joint test working group and serial interface of the graphics processor, diagnostic data is obtained for abnormal positioning and system management. The first storage space is part of the internal configuration of the baseboard management controller, dedicated to storing diagnostic firmware, the main function of which is to perform abnormal positioning based on data obtained from the joint test working group interface and serial interface of multiple graphics processors. The diagnostic firmware analyzes the received register information (such as register state) and state information (such as temperature, abnormal state, etc.), extracting information that is helpful for abnormal diagnosis.
[0099] The second storage space is also part of the internal configuration of the baseboard management controller, used to store the baseboard management controller firmware responsible for the basic operation and functions of the baseboard management controller, including system initialization, resource management, communication protocol processing, etc. The priority of the baseboard management controller firmware is different from that of the diagnostic firmware, that is, during system operation, the baseboard management controller will schedule and execute tasks according to the priority of the firmware, ensuring effective switching between normal operation and abnormal processing.
[0100] Due to the different priorities of the baseboard management controller firmware and the diagnostic firmware, the system can quickly switch to abnormal processing mode when an abnormality occurs. Even in a normal running state, the diagnostic firmware runs in the background and does not affect the operation of the baseboard management controller firmware, and the baseboard management controller can maintain the stability of its basic functions. In a complex multi-graphics processor environment, efficient monitoring and management are achieved while ensuring the stability and performance of the overall system.
[0101] Figure 4 is a step diagram of an abnormal positioning method provided by one embodiment of the present disclosure, applied to a baseboard management controller in an abnormal positioning system, the method comprising:
[0102] Step S101, respectively through the joint test working group interface module and the serial interface module, parallelly receive register information and state information of at least one first graphics processor which occurs abnormality in multiple graphics processors of a server, and return the register information and the state information to the first conversion module and the second conversion module in parallel.
[0103] Step S102, through the first conversion module and the second conversion module, respectively convert the register information and the state information into a protocol and data format compatible with the high-speed serial conversion module, and return to the high-speed serial conversion module.
[0104] Step S103, integrating the register information and the state information into diagnostic information by the high-speed serial exchange module, and returning the diagnostic information to the baseboard management controller.
[0105] Step S104, fault locating according to the diagnostic information by the baseboard management controller to determine the abnormal reason of the first graphic processor.
[0106] Involving step S101, in the abnormal positioning system as mentioned above, the baseboard management controller is configured with a monitoring mechanism for graphic processor abnormality, which monitors the running state of the graphic processor in real time. When there is an abnormality in a graphic processor or its associated hardware (such as memory, data bus, cooling system, etc.) in multiple graphic processors on a server, a corresponding abnormal alarm signal is generated, which is returned to the baseboard management controller through the interface module, the conversion module and the high-speed serial exchange module, respectively. For example, the temperature of the graphic processor is too high, the memory access is wrong, the data processing is abnormal (such as execution error, calculation error) or the graphic processor is not recognizable. Once the baseboard management controller receives the abnormal alarm signal, it can quickly locate the first graphic processor that has an abnormality in multiple graphic processors on the server. In the embodiment of the present disclosure, each graphic processor is pre-assigned with its own encoding information, which can be a combination of serial number and manufacturer number, and the specific graphic processor can be locked through the encoding information. Specifically, the register information and the state information of at least one first graphic processor are received in parallel through the joint test action group interface module and the serial interface module. The joint test action group interface module and the serial interface module are responsible for communication with the first graphic processor and obtain the internal register information and state information, respectively. In the embodiment of the present disclosure, the joint test action group interface module and the serial interface module obtain information from multiple graphic processors at the same time, rather than processing one by one. The obtained register information and state information are returned to the first conversion module and the second conversion module in parallel, ready for subsequent data conversion and processing.
[0107] Involving step S102, the register information and status information obtained from the at least one first graphics processor exist in the format and protocol of Joint Test Action Group and serial port respectively, thus conversion is needed to ensure that the subsequent processing can understand and use these diagnostic data. The first conversion module is responsible for processing the register information from the Joint Test Action Group interface module, converting it into a protocol and data format required by the high-speed serial conversion module. The second conversion module is responsible for processing the status information from the serial port interface module, also converting it into a format compatible with the high-speed serial conversion module. Optionally, in the first conversion module and the second conversion module, units for encoding, formatting and structuring data can be configured. After the conversion is completed, the first conversion module and the second conversion module return the processed register information and status information to the high-speed serial conversion module, making preparations for subsequent integration and diagnostic information generation.
[0108] Involving step S103, in step S102, the register information and status information have been converted into a protocol and data format compatible with the high-speed serial conversion module. In this step, the converted register information and status information are summarized and integrated into unified diagnostic information to facilitate subsequent fault location and analysis. Specifically, the high-speed serial exchange module is responsible for receiving data from the first conversion module and the second conversion module and integrating it into diagnostic information. After the integration is completed, the diagnostic information will be returned to the baseboard management controller. At this time, the baseboard management controller will be able to obtain detailed information about the abnormal graphics processor, providing a basis for subsequent fault analysis and location.
[0109] Involving step S104, through the baseboard management controller, the diagnostic information (register information and status information) is used for fault analysis to find out the specific reason for the abnormal graphics processor. The baseboard management controller determines the nature of the abnormality according to the preset rules and logic through the built-in diagnostic firmware. For example, the temperature of the graphics processor, the memory state, the error code, etc. can be checked to determine whether there is overheating, memory failure or other hardware problems. Through analysis, the baseboard management controller can gradually obtain the required diagnostic information to determine the specific abnormal reason of the first graphics processor.
[0110] The present disclosure improves the efficiency of fault detection and positioning by managing the register information and state information of multiple abnormal graphic processors in parallel through the baseboard management controller. Compared with the traditional method of obtaining information one by one, the present disclosure can quickly integrate diagnostic data from multiple graphic processors, reducing the time and labor cost required for troubleshooting. In addition, the baseboard management controller uses the built-in monitoring mechanism and diagnostic firmware to analyze and judge the abnormal reason in real time, so as to realize fast response and accurate positioning. It can be seen that the integrated design of the present disclosure not only simplifies the hardware connection, improves the reliability and stability of the system, but also effectively reduces the downtime caused by faults, improves the overall operation efficiency and economic benefits.
[0111] Exemplarily, the fault positioning according to the diagnostic information by the baseboard management controller to determine the abnormal reason of the first graphic processor includes: determining the register value of an over-temperature register for monitoring the temperature state according to the register information of the first graphic processor by the baseboard management controller; judging whether the register value of the over-temperature register is set, and in the case that the register value of the over-temperature register is set, obtaining the current temperature of the first graphic processor returned via the high-speed serial switching module; and in the case that the current temperature of the first graphic processor is greater than a preset temperature threshold, determining that the abnormal reason of the first graphic processor is over-temperature abnormality.
[0112] Each graphic processor is configured with a dedicated register for monitoring the temperature state, i.e. an over-temperature register. When the temperature of the graphic processor exceeds a certain safety threshold, the over-temperature register will be triggered, and the register value will be set (set to a certain state), indicating that the graphic processor has a temperature abnormality. Figure 5 is an abnormal positioning flowchart of over-temperature abnormality in an embodiment of the present disclosure, please see Figure 5 After the baseboard management controller receives the abnormal alarm signal, first, according to the register information of the first graphic processor, check the current state of the over-temperature register. If the register shows set, it means that the graphic processor has detected that the temperature exceeds the safety limit. If it is confirmed that the register value of the over-temperature register has been set, further obtain the current temperature data of the graphic processor, which can be obtained by accessing a specific hardware interface or monitoring system.
[0113] To determine whether the temperature is an abnormal situation, the baseboard management controller obtains the current temperature of the first graphics processor returned via the high-speed serial switching module. In the embodiment of the present disclosure, the baseboard management controller first sends a command indicating to obtain the current temperature to the first graphics processor, which will be extended via the high-speed serial switching module, distributed to the first conversion module and the second conversion module respectively, and then sent to at least one graphics processor connected to the joint test working group interface module and the serial interface module via the first conversion module and the second conversion module. The baseboard management controller compares the current temperature of the first graphics processor with the preset temperature threshold. The preset temperature threshold is preset during system design and configuration, based on the safe working range of the graphics processor and the temperature standard recommended by the manufacturer. For example, if the preset temperature threshold is 80°C, then when the current temperature of the first graphics processor exceeds this value, it will be considered as an over-temperature abnormality. If it is detected that the current temperature value is greater than the preset temperature threshold, it is determined that the abnormality of the graphics processor is an over-temperature abnormality, indicating that there may be a problem with the heat dissipation system, such as fan failure or heat sink blockage, or the environmental temperature is too high to cause the system to be unable to effectively dissipate heat.
[0114] Through the above steps, the first graphics processor can be quickly and accurately located and confirmed as an over-temperature abnormality. Not only can the real-time response of the abnormality be realized through the detection of the register information, but also the problem root cause can be further analyzed through the state information, and the temperature abnormality can be found and solved in time, which can avoid hardware damage caused by overheating, thereby prolonging the service life of the system and improving the reliability of the high-performance computing system.
[0115] For example, in the case where the register value of the over-temperature register is not set, it is determined that the first graphics processor does not have an over-temperature abnormality.
[0116] See Figure 5 If the register value of the over-temperature register is not set (i.e., remains "0" or "false"), it indicates that during the monitoring process, the temperature of the first graphics processor has always been within the normal range and has not exceeded the safety threshold. It can be judged that the current temperature state of the first graphics processor is normal, and it is further confirmed that the first graphics processor does not have a temperature-related abnormality. That is, it indicates that the current alarm signal is not caused by a temperature problem. In summary, this judgment logic helps the system to quickly confirm whether the graphics processor has a temperature abnormality, and if the register state is normal, it can exclude temperature as the cause of the abnormality, focus on troubleshooting other types of abnormalities, improve the efficiency of abnormal positioning, and enable the system to determine the root cause of the problem more quickly.
[0117] Exemplarily, the method further comprises: in a case where the current temperature of the first graphics processor is greater than a preset temperature threshold, acquiring, by the baseboard management controller, the current temperature of the second graphics processor returned via the high-speed serial switching module, the second graphics processor being a graphics processor deployed in a position adjacent to the first graphics processor;
[0118] In a case where the current temperature of the second graphics processor is greater than the preset temperature threshold, determining that the abnormal cause of the first graphics processor is an environmental temperature abnormality; in a case where the current temperature of the second graphics processor is less than or equal to the preset temperature threshold, determining that the abnormal cause of the first graphics processor is a self temperature abnormality.
[0119] Please refer to Figure 5 When the temperature of the first graphics processor exceeds a preset safety threshold, the temperature status of other graphics processors adjacent to the first graphics processor (i.e., the second graphics processor) via the high-speed serial switching module is acquired from the serial interface module. The graphics processors in the adjacent position refer to the graphics processors that are physically close to or adjacent to the first graphics processor in the server, share the same cooling facilities, or run in a similar physical environment. By comparing the current temperature of the second graphics processor with the preset temperature threshold, the temperature condition of the surrounding environment is determined. If the temperature of the second graphics processor also exceeds the safety threshold, it indicates that there is a temperature abnormality problem in the entire physical environment (such as the cabinet, the server), not just a single graphics processor problem.
[0120] If the temperature of the second graphics processor is also higher than the preset temperature threshold, it means that the system is in an environmental temperature abnormality situation. The environmental temperature abnormality may be caused by external cooling system failure, high temperature of the computer room, insufficient heat dissipation, etc. At this time, it can be determined that the temperature abnormality cause of the first graphics processor is the environmental temperature abnormality, i.e., the high temperature of the first graphics processor is not a hardware failure of a single graphics processor, but is caused by the overall temperature environment problem in or outside the server.
[0121] If the temperature of the second graphics processor is within the normal range, i.e., its current temperature is less than or equal to the preset temperature threshold, the possibility of environmental temperature abnormality can be excluded. At this time, the temperature abnormality of the first graphics processor is determined to be caused by its own hardware problem, including internal heat dissipation system failure of the graphics processor (such as fan failure or heat sink blockage), excessive heat inside the chip, excessive load of the graphics processor, etc. In this case, the abnormal cause is classified as a self temperature abnormality, indicating that the first graphics processor needs to be independently checked and repaired.
[0122] The above steps can determine the real cause of the abnormality by obtaining the state information of multiple graphic processors, especially the temperature information of the adjacent second graphic processor. In this way, it can be determined whether the abnormality is caused by the environmental temperature problem or the hardware problem of the graphic processor itself, so that a more targeted solution can be taken.
[0123] Exemplarily, the fault positioning performed by the baseboard management controller according to the diagnostic information to determine the abnormality cause of the first graphic processor includes: determining, by the baseboard management controller, a register value of a logic device for logically controlling the first graphic processor according to the register information of the first graphic processor; and determining, by the baseboard management controller, that the abnormality cause of the first graphic processor is a power supply abnormality when the register value of the logic device is set.
[0124] In addition to the over-temperature abnormality, the embodiment of the present disclosure also provides a diagnostic method for non-identification abnormality, which is a special type of abnormality, indicating that the system cannot identify a certain graphic processor (graphic processor), which may be caused by hardware failure, communication error or power supply problem, etc. Figure 6 is an abnormality positioning flowchart for non-identification abnormality in an embodiment of the present disclosure, as Figure 6 When the non-identification abnormality occurs, the baseboard management controller first extracts the register value from the register of the first graphic processor. In the embodiment of the present disclosure, the register value of the register (such as a complex programmable device CPLD) of the logic device responsible for the logical control of the graphic processor is extracted through the joint test action group interface module, and the register information is returned via the high-speed serial exchange module. The register of the logic device records the state of the hardware logical control, and if the register information is abnormal, it will cause the graphic processor to be unable to be normally identified.
[0125] In the case where the register value of the logic device is set, it indicates that the power module of the graphic processor has failed to work normally, causing the graphic processor to be unable to be identified by the system. Therefore, it is preliminarily determined that the abnormality cause is the power supply abnormality, which may be manifested as that the graphic processor fails to obtain sufficient voltage or current, causing it to be unable to start normally or maintain the working state. The power supply problem can be caused by multiple factors, such as power module failure, poor line connection, unstable voltage, etc. These problems will cause the graphic processor to be unable to be identified by the system, thereby triggering the non-identification abnormality. It can be seen that the present disclosure provides an effective way to quickly and accurately locate the non-identification problem of the graphic processor caused by the power supply abnormality.
[0126] Exemplarily, the method further comprises: in the case that the register value of the logic device is set, obtaining, by the baseboard management controller, the register value of a key alarm register returned via the high-speed serial switching module for monitoring the power-off protection state of the first graphics processor; determining whether the register value of the key alarm register is set, and in the case that the register value of the key alarm register is set, determining that the abnormality of the first graphics processor is caused by the power supply abnormality caused by the power-off protection abnormality of the first graphics processor; and in the case that the register value of the key alarm register is not set, determining that the abnormality of the first graphics processor is caused by the power supply abnormality caused by the non-self abnormality of the first graphics processor.
[0127] Please refer to Figure 6 If the register value of the logic device is set, further, the register information of a key alarm register related to power supply extracted from the first graphics processor via the high-speed serial switching module is obtained. The register is used to monitor the power-off protection state of the graphics processor. The power-off protection state is a protection mechanism to prevent the graphics processor from being damaged or causing system failure when the power supply is unstable or interrupted. When the power-off protection mechanism is triggered, the graphics processor enters a protection mode and cannot be normally started or run, resulting in being unable to be recognized by the system. The register value of the key alarm register is checked to determine whether it is set. If the value of the key alarm register is set, it indicates that the power-off protection mechanism has been started and there is a risk of unstable power supply or power-off for the graphics processor. If the value of the key alarm register is set, it indicates that the power-off protection mechanism of the graphics processor has been started, indicating that the power supply abnormality is caused by the power-off protection mechanism. This situation may be caused by insufficient supply voltage, transient power failure or unstable power module.
[0128] If the register value of the key alarm register is not set when checking the key alarm register, it indicates that the power-off protection mechanism has not been triggered. At this time, although the graphics processor has a power supply problem, it is not caused by its own power-off protection state abnormality, but by external causes of power supply abnormality. Non-self abnormality may include server power module failure, poor main power connection or problems with certain elements on the power supply circuit. Such problems often not only affect a single graphics processor, but also affect the power supply stability of other hardware devices and even the entire system.
[0129] By analyzing the logic control register and the key alarm register, the power supply problem caused by the power-off protection state abnormality can be quickly identified and targeted measures can be taken. Not only can the power-off protection problem inside the graphics processor be identified, but it can also be distinguished whether it is caused by the power supply problem outside the system, thereby improving the fault positioning accuracy and efficiency in the multi-graphics processor system at the same time.
[0130] Exemplarily, the method further comprises: in the case that the register value of the critical alarm register is not set, determining, by the baseboard management controller, whether the power supply and timing of the logic chip on the first graphics processor are normal according to the state information of the first graphics processor; in the case that the power supply and timing of the logic chip on the first graphics processor are normal, determining that the abnormal cause of the first graphics processor is power supply abnormality caused by self-abnormality; in the case that the power supply and timing of the logic chip on the first graphics processor are not normal, determining that the abnormal cause of the first graphics processor is power supply abnormality caused by non-self-abnormality.
[0131] Please refer to Figure 6 When the register value of the critical alarm register is not set, it indicates that the power-off protection mechanism of the first graphics processor is not triggered, and the first graphics processor does not detect obvious power supply interruption or power-off problem at the register level. Therefore, the power-off protection is not triggered at this time. Since the alarm register is not set, it can be inferred that the power supply problem has not been directly detected by the first graphics processor internally, and it is necessary to further check the internal state of the first graphics processor, especially whether the power supply and timing of the logic chip on the first graphics processor are normal.
[0132] The baseboard management controller determines the power supply and timing state of the logic chip of the first graphics processor according to the state information of the first graphics processor. Specifically, it is determined whether the logic chip can obtain stable current and voltage. If the power supply is unstable or insufficient, the operation of the graphics processor will be affected, and it may not be able to start normally or maintain a normal working state. Timing check refers to whether the clock synchronization mechanism inside the logic chip is working normally. Timing exception will cause abnormal data processing and instruction execution, and even may cause the entire graphics processor to fail to work. By checking the power supply and timing of the logic chip, it can be determined whether the first graphics processor is abnormal due to these two problems. If the check shows that the power supply and timing of the logic chip are normal, it can be excluded that the abnormality is caused by power supply or timing problem. At this time, it is necessary to further confirm whether the first graphics processor itself has other problems. It is determined that the problem is not caused by external factors, but by some hardware failure inside the first graphics processor. For example, some execution units or memory controllers may have problems, causing the first graphics processor to fail to operate normally. At this time, although the power supply and timing of the logic chip are normal, the graphics processor may still be unable to be recognized by the system or work normally due to hardware failure. In this case, it is determined that the power supply problem of the graphics processor is caused by hardware failure of the graphics processor itself, such as some internal components failing to start or maintain a working state.
[0133] If the check result shows that neither the power supply nor the timing of the logic chip is abnormal, it indicates that the problem of the graphics processor is not caused by internal failure, but by external causes. For example, the external power module fails to provide sufficient voltage or current for the graphics processor, causing the graphics processor to fail to work normally. In this case, it is determined that the power supply problem of the graphics processor is caused by external power supply problem, not by self-abnormality.
[0134] The present disclosure provides a refined diagnosis step for power supply abnormality of a graphics processor, which relies on register information and logic chip state check to determine whether the power supply problem is caused by self-abnormality or external power supply problem, so that targeted fault repair measures can be taken.
[0135] Exemplarily, according to the state information of the first graphics processor, the baseboard management controller determines whether the communication link of each of the plurality of graphics processors of the server is connected normally; in the case that the communication link of each of the plurality of graphics processors is connected normally, the state information of a third graphics processor returned via the high-speed serial switching module is obtained, the state information of the third graphics processor is compared with the state information of the first graphics processor item by item, and it is determined whether there is a difference item, the third graphics processor being a graphics processor that has not occurred abnormality; in the case that there is a difference item, it is determined that the abnormality cause of the first graphics processor is the hardware failure or configuration error corresponding to the difference item; in the case that there is no difference item, the abnormality related information of the first graphics processor and the third graphics processor is recorded into the operation log of the baseboard management controller, the abnormality related information being used for failure analysis of the first graphics processor and the third graphics processor; in the case that there is a graphics processor with communication link connection abnormality, all graphics processors with communication link connection abnormality are determined; the first graphics processor is exchanged with each of the graphics processors with communication link connection abnormality in position, and it is checked whether the communication link connection is restored to normal after the position exchange; in the case that the communication link connection is restored to normal after the position exchange is checked, it is determined that the abnormality cause of the first graphics processor is communication link abnormality caused by self-abnormality; in the case that the communication link connection is not restored to normal after the position exchange is checked, it is determined that the abnormality cause of the first graphics processor is communication link abnormality caused by link configuration abnormality.
[0136] Please refer to Figure 6When the check finds that the register value of the logic device (such as a CPLD) responsible for logical control is not set, that is, the current exception is not a power supply problem. In this case, the exception can be related to the communication link of the graphics processor or other hardware configuration. The communication link between the multiple graphics processors in the server is checked. In a multi-graphics processor system, the graphics processor interacts with other processors or system components through a communication link. If there is a connection problem with the communication link, it can cause some graphics processors to be unable to be recognized or to communicate normally with the system. By checking the communication link of each graphics processor one by one, it is confirmed whether the connection is normal. If the check finds that the communication link of all graphics processors is connected normally, it indicates that the problem is not a physical connection error of the link. At this time, it is necessary to further compare the state information of the first graphics processor and the third graphics processor that does not have an exception. The state information of the first graphics processor and the third graphics processor (normally working graphics processor) is compared item by item, such as register information, hardware configuration, running state, etc., to determine whether there are differences. If it is found through state comparison that there are differences between the first graphics processor and the third graphics processor, it indicates that the exception is caused by some hardware failure or configuration error. At this time, it can be determined that the cause of the exception of the first graphics processor is due to the problem corresponding to these differences. If it is found that there is no obvious difference between the state of the first graphics processor and the third graphics processor after the state information comparison, it indicates that the exception can be difficult to directly locate through the existing state information. At this time, the exception-related information (such as register value, hardware state, fault information) of the two is recorded in the running log for subsequent detailed failure analysis. Failure analysis can help find more exception clues in the future, thereby further improving the diagnosis and recovery capabilities of the system.
[0137] If it is found in the initial check that there is a connection problem (i.e., communication link exception) with the communication link of one or more graphics processors, it can be determined that this can be the direct cause of the exception. Next, these link exception graphics processors need to be checked step by step, and the problem is confirmed whether it comes from the specific graphics processor itself or the link configuration exception through position interchanging. The graphics processor with a communication link exception is interchanged with the first graphics processor, the communication link after the position interchanging is rechecked whether it returns to normal by changing the physical position of the graphics processor in the system. If it is found that the communication link returns to normal after the position interchanging, it indicates that the original communication link exception is caused by the problem of the first graphics processor itself. That is, some hardware failure or internal problem of the first graphics processor affects its communication with the system. At this time, it can be determined that the cause of the exception of the first graphics processor is due to the communication link problem caused by its own exception. If the communication link still does not return to normal after the position interchanging, it can be ruled out that the first graphics processor itself is faulty, indicating that the cause of the exception is a communication link exception caused by a link configuration exception, such as a slot failure.
[0138] Through the above series of steps, the causes leading to the failure of the graphics processor to be recognized are effectively located and ruled out, whether caused by hardware failure, configuration error, communication link problem or system link configuration anomaly, thereby improving the fault diagnosis and repair efficiency in the multi-graphics processor system.
[0139] Exemplarily, the historical fault data of the plurality of graphics processors in the server are analyzed by a machine learning model to generate a fault mode database; in response to the abnormal alarm signal, based on the fault mode database and the register information, state information and environment data of the first graphics processor, a potential abnormal type that the first graphics processor is likely to occur is predicted; in the case that the predicted potential abnormal type is inconsistent with the actually detected abnormal type, an intelligent repair scheme is generated, and the first graphics processor is dynamically configured and adjusted according to the intelligent repair scheme.
[0140] In the historical running process of the plurality of graphics processors in the server, various different abnormal situations (such as temperature being too high, power supply problem, communication link problem, etc.) will be encountered. The occurrence of these abnormalities will record relevant data, such as register information, state information and environment data (such as temperature, working condition of other hardware around, etc.). In a period of time, these historical fault data will accumulate to form a larger fault data set. By analyzing these historical fault data, a machine learning model (such as decision tree, random forest, support vector machine or deep learning model, etc.) is used to learn the feature patterns of different fault types. The machine learning model can extract the typical features of each fault, for example, a certain type of abnormality may correspond to a specific register value abnormality, state information abnormality or temperature being too high, etc.
[0141] Through the trained machine learning model, a fault mode database is formed, that is, a mapping relationship between each fault type and its corresponding feature pattern, which can help the system to quickly match and identify the fault type occurring in the future. In response to the abnormal alarm signal, the current register information, state information and related environment data of the first graphics processor are immediately collected. Based on the fault mode database generated by the machine learning, the real-time information currently acquired is compared with the fault features in the database, so as to predict the abnormal type that the graphics processor is likely to occur. The prediction ability of machine learning is used to guess the most possible abnormal reason.
[0142] The potential abnormal type given by the machine learning model is a prediction based on historical data and patterns. However, the system can obtain the actual detected abnormal type through direct detection, such as directly reading the temperature value through monitoring the register to determine the over-temperature abnormality. If the potential abnormal type predicted by the machine learning and the actual detected result are inconsistent, for example, the machine learning model predicts that it may be a power problem, but the detected result is a temperature problem, which indicates that the problem is more complex than the directly matched pattern.
[0143] In the case of inconsistency between prediction and actuality, an intelligent repair scheme is generated, which is generated according to the current graphics processor information, historical fault data and existing best practice methods. It can include adjusting the running state of the graphics processor, resetting part of the functional module, or even making system-level configuration changes. For example, the working frequency of the graphics processor can directly affect its heat generation and power consumption, and by reducing the working frequency, the heat generation of the processor can be reduced to avoid further damage.
[0144] In the process of each fault handling and repair, new data is continuously accumulated. If the currently generated intelligent repair scheme is successfully applied, it means that the model prediction and the repair scheme are effective, and the system can input these new data into the machine learning model to further optimize the fault pattern database. Through iteration and circulation, the future abnormal detection, prediction and repair will be more accurate, and the self-healing ability of the entire system will also be continuously improved. If the abnormality cannot be repaired after dynamic configuration adjustment, a fault report is automatically generated, which contains the process of repair attempt, failure cause analysis and suggested further operation scheme. It helps administrators quickly locate the problem and is used for continuous improvement of the fault handling process.
[0145] The present disclosure combines the analysis of historical data and the intelligent processing of real-time data, uses machine learning for prediction and generates intelligent repair schemes in complex scenarios, significantly improving the automation and efficiency of graphics processor abnormality processing. Through dynamic configuration adjustment and state monitoring, human intervention can be minimized, the self-healing ability of the server can be improved, and the efficient and stable operation of the server can be ensured.
[0146] The embodiment of the present disclosure also provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, it implements the processes of the above-mentioned embodiment of the abnormal positioning method and achieves the same technical effects. To avoid repetition, this will not be repeated here.
[0147] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement each process of the above-described embodiment of the abnormal positioning method, and achieves the same technical effects. Each embodiment in the present specification is described in a progressive manner, and each embodiment mainly describes differences from other embodiments. The same or similar parts of each embodiment are cross-referenced.
[0148] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, device, electronic device and storage medium. Therefore, the embodiments of the present disclosure can adopt a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present disclosure can adopt the form of a computer program product implemented on one or more computer readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer usable program codes.
[0149] The embodiments of the present disclosure are described with reference to flowcharts and / or block diagrams of the methods and devices according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal equipment to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal equipment produce a device for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks Figure 1 Figure 1 The functions specified in one or more flows and / or blocks
[0150] While preferred embodiments of the present disclosure have been described, those skilled in the art will appreciate that other modifications and changes can be made thereto without departing from the scope of the present disclosure. It is therefore intended that the appended claims encompass all such modifications and changes as fall within the scope of the present disclosure.
[0151] Finally, it should be noted that, in the description of the present disclosure, the terms such as first and second, etc. are merely intended to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or terminal device. Without more limitations, the elements defined by the statement "comprising" do not exclude the presence of additional identical elements in the process, method, article or terminal device including the elements. The above has described in detail an abnormal positioning system, method, electronic device and storage medium provided by the present disclosure, specific examples are applied herein to set forth the principles and implementation manners of the present disclosure, the above description of the embodiments is only for the purpose of helping to understand the method of the present disclosure and its core idea; at the same time, for those skilled in the art, according to the idea of the present disclosure, the specific implementation manners and application range will be changed, and the above description of the present disclosure should not be understood as a limitation.
Claims
1. An anomaly location system, characterized in that, The system includes: A baseboard management controller is used to connect control devices; Multiple graphics processors, each with a joint test workgroup interface and a serial port interface; The control device includes: a high-speed serial switching module, a first conversion module, a second conversion module, a serial port interface module, and a joint test working group interface module; The input terminal of the high-speed serial switching module is connected to the baseboard management controller, and the output terminal of the high-speed serial switching module is connected to the first conversion module and the second conversion module respectively. The high-speed serial switching module is used to extend the control commands from the baseboard management controller, and send the control commands to the first conversion module and the second conversion module in parallel, or to integrate the register information from the first conversion module and the status information from the second conversion module, and return it to the baseboard management controller. The first conversion module is used to connect to the joint test workgroup interface module, and the second conversion module is used to connect to the serial port interface module. The serial port interface module is used to connect to the serial port interfaces of the multiple graphics processors; the joint test workgroup interface module is used to connect to the joint test workgroup interfaces of the multiple graphics processors. The baseboard management controller is used to acquire, in parallel, register information and serial port interface status information returned by the high-speed serial exchange module from the joint test workgroup interface of the multiple graphics processors, in order to locate anomalies.
2. The system according to claim 1, characterized in that, The substrate management controller has a first high-speed serial interface, which is connected to the controller device via a first high-speed serial signal line; the input end of the high-speed serial switching module has a second high-speed serial interface, and the output end has a third high-speed serial interface and a fourth high-speed serial interface. The second high-speed serial interface is connected to the first high-speed serial interface of the substrate management controller via a first high-speed serial signal line. The third high-speed serial interface is connected to the first conversion module, and the fourth high-speed serial interface is connected to the second conversion module. The joint test workgroup interface on the multiple graphics processors includes a high-potential joint test workgroup interface and a low-potential joint test workgroup interface. The joint test working group interface module includes a first joint test working group interface module and a second joint test working group interface module; The first joint test workgroup interface module is connected in parallel to the low-potential joint test workgroup interfaces on the multiple graphics processors, and the second joint test workgroup interface module is used to connect in parallel to the high-potential joint test workgroup interfaces on the multiple graphics processors.
3. The system according to claim 2, characterized in that, The first conversion module has a fifth high-speed serial interface at its input end and a first joint test workgroup interface and a second joint test workgroup interface at its output end. The fifth high-speed serial interface of the first conversion module is connected to the third high-speed serial interface of the high-speed serial switching module through the second high-speed serial signal line. The first joint test workgroup interface of the first conversion module is connected to the first joint test workgroup interface module via a first joint test workgroup signal line, and the second joint test workgroup interface of the first conversion module is connected to the second joint test workgroup interface module via a second joint test workgroup signal line.
4. The system according to claim 1, characterized in that, The second conversion module has a sixth high-speed serial interface at its input and a serial port interface at its output. The sixth high-speed serial interface of the second conversion module is connected to the fourth high-speed serial interface of the high-speed serial switching module through the third high-speed serial signal line, and the serial port interface of the second conversion module is connected to the serial port interface module through the serial port signal line.
5. The system according to any one of claims 1-4, characterized in that, The baseboard management controller is configured with a first storage space to store diagnostic firmware, which is used to locate anomalies based on the register information of the joint test workgroup interface of the multiple graphics processors and the status information of the serial port interface. The baseboard management controller is equipped with a second storage space for storing baseboard management controller firmware. The priority of the baseboard management controller firmware is different from that of the diagnostic firmware.
6. An anomaly localization method, characterized in that, Applied to the system as described in any one of claims 1-5, the method comprises: The system receives register information and status information of at least one first graphics processor that has malfunctioned from among the multiple graphics processors of the server in parallel through the joint test working group interface module and the serial port interface module, and returns the register information and status information in parallel to the first conversion module and the second conversion module. The first conversion module and the second conversion module convert the register information and the status information into a protocol and data format compatible with the high-speed serial switching module, respectively, and return them to the high-speed serial switching module. The high-speed serial switching module is used to extend the control instructions from the baseboard management controller, and send the control instructions to the first conversion module and the second conversion module in parallel, or to integrate the register information from the first conversion module and the status information from the second conversion module, and return them to the baseboard management controller. The high-speed serial switching module integrates the register information and the status information into diagnostic information, and returns the diagnostic information to the baseboard management controller. The baseboard management controller locates the fault based on the diagnostic information to determine the cause of the abnormality in the first graphics processor.
7. The method according to claim 6, characterized in that, The step of using the baseboard management controller to locate faults based on the diagnostic information to determine the cause of the abnormality in the first graphics processor includes: The baseboard management controller determines the register value of the over-temperature register used to monitor the temperature status based on the register information of the first graphics processor. Determine whether the register value of the over-temperature register is set. If the register value of the over-temperature register is set, obtain the current temperature of the first graphics processor returned by the high-speed serial exchange module. If the current temperature of the first graphics processor is greater than a preset temperature threshold, the cause of the abnormality of the first graphics processor is determined to be an over-temperature abnormality.
8. The method according to claim 7, characterized in that, If the register value of the over-temperature register is not set, it is determined that the first graphics processor has not experienced an over-temperature anomaly.
9. The method according to claim 7, characterized in that, The method further includes: If the current temperature of the first graphics processor is greater than a preset temperature threshold, the baseboard management controller obtains the current temperature of the second graphics processor returned via the high-speed serial exchange module. The second graphics processor is a graphics processor deployed in a location adjacent to the first graphics processor. The graphics processor in the adjacent location is a graphics processor that is physically adjacent to the first graphics processor in the server and shares the same heat dissipation facilities. If the current temperature of the second graphics processor is greater than the preset temperature threshold, the cause of the abnormality of the first graphics processor is determined to be an abnormal ambient temperature. If the current temperature of the second graphics processor is less than or equal to the preset temperature threshold, the cause of the abnormality of the first graphics processor is determined to be its own temperature abnormality.
10. The method according to claim 6, characterized in that, The step of using the baseboard management controller to locate faults based on the diagnostic information to determine the cause of the abnormality in the first graphics processor includes: The baseboard management controller determines the register values of the logic devices used to perform logic control on the first graphics processor based on the register information of the first graphics processor. Determine whether the register value of the logic device is set. If the register value of the logic device is set, determine that the cause of the abnormality of the first graphics processor is an abnormal power supply.
11. The method according to claim 10, characterized in that, The method further includes: When the register value of the logic device is set, the register value of the key alarm register used to monitor the power-down protection status of the first graphics processor is obtained by the baseboard management controller via the high-speed serial exchange module. Determine whether the register value of the critical alarm register is set. If the register value of the critical alarm register is set, determine that the cause of the abnormality of the first graphics processor is an abnormal power supply caused by the power failure protection of the first graphics processor. If the register value of the critical alarm register is not set, the cause of the abnormality of the first graphics processor is determined to be a power supply abnormality caused by a non-inherent abnormality of the first graphics processor.
12. The method according to claim 11, characterized in that, The method further includes: If the register value of the critical alarm register is not set, the baseboard management controller determines whether the power supply and timing of the logic chip on the first graphics processor are normal based on the status information of the first graphics processor. If the power supply and timing of the logic chip on the first graphics processor are normal, the cause of the abnormality of the first graphics processor is determined to be an abnormal power supply caused by its own abnormality. If the power supply and timing of the logic chip on the first graphics processor are abnormal, it is determined that the abnormality of the first graphics processor is caused by a power supply abnormality that is not due to its own abnormality.
13. The method according to claim 10, characterized in that, The method further includes: When the register value of the logic device is not set, the baseboard management controller determines whether the communication links of the multiple graphics processors of the server are connected normally based on the status information of the first graphics processor. When the communication links of the multiple graphics processors are all connected normally, the status information of the third graphics processor returned by the high-speed serial exchange module is obtained, and the status information of the third graphics processor is compared with the status information of the first graphics processor item by item to determine whether there are any differences. The third graphics processor is a graphics processor that has not experienced any abnormalities. If a discrepancy exists, the cause of the anomaly in the first graphics processor is determined to be a hardware failure or configuration error corresponding to the discrepancy. If no discrepancy exists, the anomaly-related information of the first graphics processor and the third graphics processor is recorded in the operation log of the baseboard management controller. The anomaly-related information is used to perform failure analysis on the first graphics processor and the third graphics processor. In the case of a graphics processor with an abnormal communication link connection, identify all graphics processors with abnormal communication link connections; swap the positions of all graphics processors with abnormal communication link connections with the first graphics processor one by one, and check whether the communication link connection is restored after the position swap; if the communication link connection is restored after the position swap, determine that the abnormality of the first graphics processor is caused by its own abnormality leading to the communication link abnormality; if the communication link connection is not restored after the position swap, determine that the abnormality of the first graphics processor is caused by the link configuration abnormality leading to the communication link abnormality.
14. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of the method as described in any one of claims 6-13.
15. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the method as described in any one of claims 6-13.
Citation Information
Patent Citations
Graphic processor board card
CN109408445A
Board card management method, device and system
CN117493255A