Abnormal monitoring method and device, server, computer equipment and medium

The GPU running logs are automatically collected through the baseboard management controller by calling the log collector, which solves the problems of high GPU failure rate and low log collection efficiency, and realizes fast and effective abnormal log collection and failure analysis, improving the reliability and stability of the server.

CN119961034APending Publication Date: 2025-05-09INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411971879.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The GPU failure rate is high, especially in large-scale application scenarios, which leads to a sharp increase in the failure risk and affects business stability. It is difficult for the existing technology to quickly and effectively collect abnormal logs.

Method used

The running status of the graphics processor is obtained through the substrate management controller to evaluate whether it is in an abnormal state. If it is abnormal, the log collector will be called to automatically collect the running log to improve the fineness and richness of log collection.

Benefits of technology

Improves the efficiency of abnormal log collection, can analyze the causes of graphics processor abnormalities more quickly, reduces troubleshooting time, and improves the reliability and stability of the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961034A_ABST
    Figure CN119961034A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to an exception monitoring method of a graphics processor, an exception monitoring device of the graphics processor, a server, computer equipment and a computer readable storage medium. The abnormity monitoring method of the graphics processor comprises the following steps: a baseboard management controller acquires the running state of the graphics processor connected with the baseboard management controller; evaluating whether the running state feeds back that the graphics processor is in an abnormal state currently; in response to the judgment that the graphics processor is in the abnormal state currently, calling a log collector of the graphics processor, and collecting running logs of the graphics processor by using the log collector; wherein at least one of the fineness and the richness of the running logs collected by the log collector is higher than that of the running logs collected by the baseboard management controller. By adopting the method, the log collector of the graphics processor can be automatically called, so that the log collection fineness and / or richness can be improved, and the abnormal log collection efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method for monitoring an abnormality of a graphics processor, an apparatus for monitoring an abnormality of a graphics processor, a server, a computer device, and a computer-readable storage medium. Background Art

[0002] As some institutions summarize the application layer trends of AIGC (Artificial Intelligence Generated Content), innovations in the AIGC application layer are still developing. The implementation of related AI (Artificial Intelligence) technologies usually depends on the computing power of hardware.

[0003] Compared with traditional hardware such as CPU (Central Processing Unit), the failure rate of GPU (Graphics Processing Unit) is relatively high, especially in large-scale application scenarios. Due to the large number of GPUs, the risk of GPU failure also increases dramatically. When GPUs are used to build super computing clusters, single-point failures of GPUs can easily spread rapidly and cause large-scale chain reactions, and even affect business stability. Therefore, how to quickly and effectively collect exception logs has become an urgent problem to be solved. Quickly and effectively collecting exception logs is conducive to timely solving GPU failures. Summary of the invention

[0004] Based on this, it is necessary to provide a graphics processor anomaly monitoring method, a graphics processor anomaly monitoring device, a server, a computer device and a computer-readable storage medium to address the above-mentioned technical problems, which can automatically call the log collector of the graphics processor, thereby improving the fineness and / or richness of log collection and improving the efficiency of abnormal log collection.

[0005] On the one hand, a method for monitoring an abnormality of a graphics processor is provided, the method comprising: a baseboard management controller acquiring the operating status of a graphics processor connected thereto; evaluating whether the operating status indicates that the graphics processor is currently in an abnormal state; in response to determining that the graphics processor is currently in an abnormal state, calling a log collector of the graphics processor, and using the log collector to collect operating logs of the graphics processor; wherein at least one of the precision and richness of the operating logs collected by the log collector is higher than that of the operating logs collected by the baseboard management controller.

[0006] In one embodiment of the present application, calling the log collector of the graphics processor includes: triggering the monitoring prompt log of the baseboard management controller; wherein the monitoring prompt log indicates monitoring the running status of the graphics processor and a prompt log that the graphics processor is in an abnormal state; sending a management signal to the management flag of the log collector to enable the log collector to activate the log collection function of the log collector; using the log collector to collect the running log of the graphics processor includes: in response to the management flag to enable the log collector, controlling the platform management interface process to search for the script of the log collector; wherein the script of the log collector is a control code that is pre-compatible with the baseboard management controller; in response to the search to locate the script of the log collector, controlling the platform management interface to execute the script of the log collector to enable the log collector to automatically perform the task of collecting the graphics processor log; obtaining the running log of the graphics processor collected by the log collector; and storing the running log in a preset storage module.

[0007] In one embodiment of the present application, storing the operation log to a preset storage module includes: obtaining a log download instruction; in response to obtaining the log download instruction, calling a collection logger of a baseboard management controller; using the collection logger to download the operation log collected by the log collector, and saving the downloaded operation log to a local storage module, so as to perform graphics processor fault analysis through the operation log stored in the local storage module.

[0008] In one embodiment of the present application, after the baseboard management controller obtains the operating status of the graphics processor connected to it, it also includes: in response to determining that the graphics processor is in a normal state based on the operating status, determining whether the graphics processor has an error; in response to the graphics processor reporting an error, generating a prompt message; wherein the prompt message is used to feedback the situation that the graphics processor reports an error and the baseboard management controller has not detected it; obtaining a trigger instruction to actively trigger the log collection function of the log collector based on the trigger instruction, and using the log collector to collect the operating log of the graphics processor.

[0009] In one embodiment of the present application, the exception monitoring method also includes: obtaining an exception log of a graphics processor; analyzing the exception log to identify the exception type and the cause of the exception represented by the exception log; evaluating the exception level of the graphics processor exception reported by the exception log, and providing feedback on the exception message using a preset feedback method that matches the exception level; wherein the exception message includes at least one of the exception type, the cause of the exception, and the exception level.

[0010] In one embodiment of the present application, obtaining an exception log of a graphics processor includes: monitoring a hardware control manager of the graphics processor to collect exception information of the graphics processor; converting the format of the exception information to obtain an exception log in a target format and / or containing target information; and sending the exception log to a preset storage module for storage.

[0011] On the other hand, a graphics processor abnormality monitoring device is provided, the abnormality monitoring device comprising: a baseboard management controller, an abnormality monitoring module and a storage module; the abnormality monitoring module is arranged in the baseboard management controller, and is used to perform abnormality monitoring on the graphics processor connected to the baseboard management controller using any of the above-mentioned graphics processor abnormality monitoring methods; the storage module is used to store the collected operation logs.

[0012] On the other hand, a server is provided, comprising a server body, a graphics processor and an abnormality monitoring device for the graphics processor as in the above embodiment; the graphics processor is arranged in the server body; the abnormality monitoring device is arranged in the server body for performing abnormality monitoring on the graphics processor.

[0013] On the other hand, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: a baseboard management controller obtains the operating status of a graphics processor connected thereto; the operating status is evaluated to determine whether the graphics processor is currently in an abnormal state; in response to determining that the graphics processor is currently in an abnormal state, a log collector of the graphics processor is called, and the operating log of the graphics processor is collected by the log collector; wherein at least one of the fineness and richness of the operating logs collected by the log collector is higher than the operating logs collected by the baseboard management controller.

[0014] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: a baseboard management controller obtains the operating status of a graphics processor connected thereto; an evaluation is made as to whether the operating status indicates that the graphics processor is currently in an abnormal state; in response to determining that the graphics processor is currently in an abnormal state, a log collector of the graphics processor is called, and the operating log of the graphics processor is collected by the log collector; wherein at least one of the precision and richness of the operating logs collected by the log collector is higher than the operating logs collected by the baseboard management controller.

[0015] The above-mentioned graphics processor abnormality monitoring method, graphics processor abnormality monitoring device, server, computer equipment and computer-readable storage medium, the baseboard management controller can call the graphics processor log collector, the log collector collects the graphics processing operation log with higher precision and / or richness than the operation log collected by the baseboard management controller, and compared with the method of waiting for the operation and maintenance personnel to manually use the log collection tool to obtain the deep log, it can improve the efficiency of operation log collection, so as to improve the efficiency of analyzing the cause of the graphics processor abnormality based on the operation log and solving the abnormality, thereby improving the reliability and stability of the server. At the same time, this also means that when the graphics processor fails, the operation log can be collected by a log collector that is more compatible with the graphics processor to improve the effectiveness of the operation log used to analyze the cause of the graphics processor abnormality. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of a first embodiment of the abnormality monitoring method of a graphics processor of the present application;

[0017] Figure 2 It is a schematic diagram of an application scenario of an embodiment of a method for monitoring abnormality of a graphics processor of the present application;

[0018] Figure 3 It is a flowchart of a second embodiment of the abnormality monitoring method of a graphics processor of the present application;

[0019] Figure 4 is a flowchart of a third embodiment of the abnormality monitoring method for a graphics processor of the present application;

[0020] Figure 5 is a flowchart of a fourth embodiment of the abnormality monitoring method for a graphics processor of the present application;

[0021] Figure 6 It is a structural schematic diagram of an embodiment of an abnormality monitoring device for a graphics processor;

[0022] Figure 7 It is a structural diagram of an embodiment of the server of the present application;

[0023] Figure 8 It is a structural diagram of an embodiment of a computer device of the present application. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0025] In order to solve the technical problems in the related art that the failure rate of the graphics processor is relatively high and the efficiency of solving the failure is relatively low, the present application provides a method for monitoring the abnormality of the graphics processor, a device for monitoring the abnormality of the graphics processor, a server, a computer device and a computer-readable storage medium. The abnormality monitoring method includes: the baseboard management controller obtains the operating status of the graphics processor connected to it; the evaluation of the operating status indicates whether the graphics processor is currently in an abnormal state; in response to determining that the graphics processor is currently in an abnormal state, the log collector of the graphics processor is called, and the operation log of the graphics processor is collected by the log collector; wherein, at least one of the precision and richness of the operation log collected by the log collector is higher than the operation log collected by the baseboard management controller. The specific working principle of the present application is described in detail below.

[0026] See also Figure 1 , Figure 1 It is a flowchart of the first embodiment of the abnormality monitoring method of the graphics processor of the present application.

[0027] S101: The baseboard management controller obtains the operating status of the graphics processor connected thereto.

[0028] In this embodiment, the abnormality monitoring function of the graphics manager is integrated into the baseboard management controller such as BMC (baseboard management controller) outside the server band. The baseboard management controller is connected to the graphics processor to obtain the operating status of the graphics processor connected thereto, thereby performing abnormality monitoring on the graphics processor based on the operating status.

[0029] S102: Evaluate the running status to see whether the graphics processor is currently in an abnormal state.

[0030] In this embodiment, the baseboard management control responds to obtaining the operating status of the graphics processor and can evaluate whether the graphics processor is in an abnormal state based on the operating status, thereby automating the abnormal monitoring of the graphics processor and timely discovering that the graphics processor is in an abnormal state.

[0031] S103: In response to determining that the graphics processor is currently in an abnormal state, calling a log collector of the graphics processor, and using the log collector to collect operation logs of the graphics processor; wherein at least one of the precision and richness of the operation logs collected by the log collector is higher than the operation logs collected by the baseboard management controller.

[0032] In this embodiment, when the baseboard management controller determines that the graphics processor is in an abnormal state, the log collector of the graphics processor is called to collect the operation log of the graphics processor by using the log collector.

[0033] It is easy to understand that the baseboard management controller itself can obtain the operation log of the graphics processor. In this embodiment, a log collector is additionally introduced. When the graphics processor is in an abnormal state, the baseboard management controller calls the log collector to collect the operation log of the graphics processor by using the log collector, so as to analyze the cause of the abnormality of the graphics processor by using the collected operation log, so as to perform abnormal recovery processing such as fault recovery on the graphics processor based on this.

[0034] Furthermore, the log collector in this embodiment can collect operation logs with a higher degree of precision and / or richness than the operation logs collected by the baseboard management controller, so as to further improve the effectiveness of the operation logs, thereby facilitating improving the efficiency of resolving graphics processor anomalies.

[0035] Optionally, the log collector can be a log collection tool provided by the graphics processor manufacturer. In this embodiment, the log collection tool script can be integrated into the code, so that the control function of the log collection function of the log collector is integrated into the baseboard management controller, so that the baseboard management controller can directly call the log collection tool. Compared with the method in which relevant operation and maintenance personnel use the log collection tool to collect logs after discovering the graphics processor failure, this embodiment can realize automatic detection and automatic driving of the log collector, thereby improving the efficiency of running log collection when the graphics processor is in an abnormal state. Alternatively, the log collector can be self-developed, and the manufacturer is confirmed and authorized to allow in-depth interaction with the graphics processor and collect relatively complete running logs, which is not limited here.

[0036] That is to say, the baseboard management controller can call the log collector of the graphics processor, and the log collector collects the operation logs of the graphics processing with higher precision and / or richness than the operation logs collected by the baseboard management controller. Compared with waiting for the operation and maintenance personnel to manually use the log collection tool to obtain the deep logs, it can improve the efficiency of operation log collection, so as to improve the efficiency of analyzing the cause of the graphics processor abnormality based on the operation log and solving the abnormality, thereby improving the reliability and stability of the server. At the same time, this also means that when the graphics processor fails, the operation log can be collected by a log collector that is more compatible with the graphics processor to improve the effectiveness of the operation log used to analyze the cause of the graphics processor abnormality.

[0037] Please refer to Figure 2 as well as Figure 3 , Figure 2 is a schematic diagram of an application scenario of an embodiment of a method for monitoring abnormalities of a graphics processor of the present application, Figure 3 It is a flowchart of the second embodiment of the abnormality monitoring method of the graphics processor of the present application.

[0038] In one embodiment, the application scenario of the abnormality monitoring method of the graphics processor may include a personal computer (PC), a baseboard management controller, and a graphics processor to be monitored for abnormality. The graphics processor may include a hardware control manager such as an HMC (a management controller of a GPU module), which can manage the graphics processor, and the baseboard manager can be connected to the HMC to obtain the operating status of the graphics processor by polling the HMC, real-time monitoring of the HMC, etc.

[0039] S201: Obtaining the operating status of the graphics processor.

[0040] In this embodiment, the baseboard management controller can obtain the operating status of the graphics processor.

[0041] S202: Evaluate whether the graphics processor is currently in an abnormal state.

[0042] In this embodiment, when it is determined that the graphics processor is in an abnormal state, step S206 and step S208 are executed; when it is determined that the graphics processor is not in an abnormal state, step S203 is executed.

[0043] S203: Determine whether the graphics processor reports an error.

[0044] In this embodiment, when it is determined that the graphics processor has not reported an error, it can be considered that the graphics processor is operating normally and the process ends; when it is determined that the graphics processor has reported an error, step S204 can be executed.

[0045] That is to say, when monitoring the graphics processor in this embodiment, the baseboard management controller can identify whether the graphics processor is in an abnormal state, and the baseboard management controller can also identify whether the graphics processor reports an error. In this way, there can be two situations for the graphics processor to be in an abnormal state, one is that the baseboard management controller identifies that it is in an abnormal state, and the other is that the baseboard management controller does not identify that the graphics processor is in an abnormal state, but the graphics processor reports an error. Different solution strategies can exist for different situations, thereby improving the precision of abnormal monitoring in this embodiment, and respectively configuring response strategies to help improve the reliability of abnormal monitoring.

[0046] S204: Generate prompt information.

[0047] In this embodiment, the prompt information is used to feedback the situation that the graphics processor reports an error and the baseboard management controller does not detect it.

[0048] That is, in response to determining that the graphics processor is in a normal state based on the running state, but the graphics processor reports an error, a prompt message is generated, so that the user receives the prompt message to select a solution.

[0049] Optionally, the prompt information may be output to a display device, or the graphics processor status may be marked on the display device, or related text and / or image messages may be sent to a user terminal, which are not limited here.

[0050] S205: Obtain a trigger instruction to call a log collector.

[0051] In this embodiment, the user can input a start instruction through the interactive terminal. If so, the baseboard management controller can obtain the trigger instruction to actively trigger the log collection function of the log collector based on the trigger instruction, and use the log collector to collect the operation log of the graphics processor.

[0052] S206: Calling the log collector of the graphics processor.

[0053] In this embodiment, in response to the baseboard management controller evaluating and determining that the graphics processor is in an abnormal state, the log collector of the graphics processor may be called.

[0054] The monitoring prompt log of the baseboard management controller can be triggered; wherein the monitoring prompt log indicates monitoring of the running status of the graphics processor and a prompt log that detects that the graphics processor is in an abnormal state; a management signal is sent to the management flag of the log collector to enable the log collector to activate the log collection function of the log collector.

[0055] As such, in this embodiment, the baseboard manager and the log collector are connected and configured in an associated manner. The baseboard management controller can actively trigger the log collector so that the baseboard management controller can control the log collector, thereby enabling the baseboard management controller to drive the log collector to collect the operation logs of the graphics processor.

[0056] S207: Collect the operation log of the graphics processor using a log collector.

[0057] In this embodiment, in response to the management flag indicating that the log collector is enabled, the baseboard management controller can control the platform management interface process to search for the log collector script; wherein the log collector script is a control code pre-compatible with the baseboard management controller. The control code can represent a code related to the server code and the baseboard management controller, and can be considered as a code in a broad sense, rather than a code that necessarily controls the baseboard controller.

[0058] In response to searching to locate the script of the log collector, the control platform management interface executes the script of the log collector to enable the log collector to automatically perform the task of collecting the graphics processor log; obtain the operation log of the graphics processor collected by the log collector; and store the operation log in a preset storage module.

[0059] In this way, the log collector is triggered by adding a management flag, and the log collector script can be executed through a platform management interface such as IPMI (Intelligent Platform Management Interface) to collect the operation log of the graphics processor. The operation log collection time can be preset, for example, it can last for 20 minutes, 30 minutes, 40 minutes and other preset collection times, to obtain the operation log of the graphics processor, and store the operation log in a preset storage module to reliably implement the driving of the log collector.

[0060] Furthermore, a log download instruction can also be obtained. The log download instruction is sent externally to the baseboard management controller, and the log download instruction identifies whether to download the operation log, thereby improving the reasonable planning of the operation log download timing. In response to obtaining the log download instruction, the collection logger of the baseboard management controller is called; the collection logger is used to download the operation log collected by the log collector, and the downloaded operation log is saved to the local storage module, so as to perform graphics processor fault analysis through the operation log stored in the local storage module.

[0061] That is to say, in this embodiment, when the operation log of the graphics processor is collected by the log collector, the operation log of the collected graphics processor can be stored inside the baseboard management controller, so that the user can decide whether to download the operation log, and when there is no operation log downloaded for fault analysis, it is beneficial to reduce the occupation of local storage resources, thereby facilitating service reliability. In other words, in this embodiment, two preset storage modules for storing operation logs can be included, one of which is a storage unit provided inside the baseboard management controller, which stores the operation logs collected by the log collector and its own log collector; the other is a local storage module, which is used by the user to use the operation log downloaded therein to perform fault analysis of the graphics processor, so as to improve the rationality of the storage module setting.

[0062] In this embodiment, the baseboard management controller may not perform abnormality analysis on the operation logs collected by the log collector, so as to reduce the calculation and analysis burden of the baseboard management control.

[0063] Furthermore, the operation logs obtained by the log collector can be analyzed, and the current abnormality type of the graphics processor and the details of the operation log can be analyzed to match the abnormality of the graphics processor with the corresponding log type and log expression that can reflect the abnormality. In this way, when the baseboard management controller obtains the log type and log expression later, it can judge whether the graphics processor is in an abnormal state through another dimension, thereby improving the reliability of graphics processor monitoring.

[0064] Furthermore, when a new log collector script is obtained, it is possible to evaluate whether the log collector needs to be updated at present. For example, the currently loaded version of the log collector script and the target version can be compared, and the target version is the new script version to be loaded. The difference code between the two is analyzed to analyze whether the difference code has a substantial impact on the collection of operation logs, such as whether to collect new types of operation logs, whether to reduce the collection of some types of operation logs, and if there are increases or decreases, it can be determined whether it can be implemented alternatively through the baseboard management controller. If the baseboard management controller can be implemented alternatively, it is determined that there is no need to update the log collector script.

[0065] S208: Obtaining an abnormal log of the graphics processor.

[0066] In this embodiment, the hardware control manager of the graphics processor can be monitored to collect exception information of the graphics processor; the exception information is formatted to obtain an exception log in a target format and / or containing target information; and the exception log is sent to a preset storage module for storage.

[0067] S209: Analyze the exception log and feedback the exception message.

[0068] In this embodiment, the exception log may be analyzed to identify the exception type and the cause of the exception represented by the exception log.

[0069] Evaluate the abnormality level of the graphics processor abnormality reported by the abnormality log, and use a preset feedback method matching the abnormality level to feedback the abnormality message; wherein the abnormality message includes at least one of the abnormality type, the abnormality cause, and the abnormality level.

[0070] That is to say, in this embodiment, the storage strategy and prompt mechanism can be flexibly configured to facilitate the user to flexibly choose whether to handle the exception of the graphics processor according to the feedback exception message, thereby meeting the exception monitoring needs and log collection needs of different scenarios, improving the reliability and availability of the graphics processor and the server, and also improving the generalizability of the exception monitoring method of the graphics processor, which can enrich the application scenarios and provide effective guarantee for the stable operation of servers such as AI servers, thereby improving the user experience.

[0071] In layman's terms, in this embodiment, the baseboard management controller can monitor the graphics processor for abnormalities through the abnormality detection module. The abnormality detection module can be refined to include a collection unit, an analysis unit, a prompt unit and a storage unit. As the name implies, the collection unit is used to communicate with the HMC to obtain the operating status and abnormal information of the GPU through regular polling, real-time monitoring, event triggering, etc. The collection unit can also organize the collected error information into a format such as json (a data exchange format) according to a preset format, so that the analysis unit on the downstream side can analyze the operation log under the abnormal state, identify the abnormality type, the cause of the abnormality, etc., and then make the prompt unit form a corresponding prompt action. In addition, the collection unit can send the collected operation log under the abnormal state to the storage unit, because the storage unit stores the collected operation log.

[0072] Specifically, the BMC prompt unit can determine whether it is necessary to issue an abnormal message for prompting based on the output results of the analysis unit. When it is identified that the graphics processor is currently in a serious error or failure, a standard SEL (System Event Log) will be generated. Among them, SEL can be a log of BMC recording system events, which may include hardware events, sensor status changes, error messages, etc. This log can effectively assist in diagnosing server hardware problems. And the prompt unit can be selected to promptly notify the management and operation and maintenance personnel through emails, text messages, etc. to ensure that the current failure of the graphics processor is handled in a timely manner.

[0073] The storage module mainly involves temporarily storing the logs dumped from the GPU module in the current file system (i.e., storage unit) of the BMC. Different storage strategies can also be adopted according to the importance and timeliness of the logs, such as regular backup and compressed storage, to save storage space and improve data access efficiency.

[0074] In summary, in this embodiment, the BMC that can be deployed on the AI ​​server can realize the functions of automatically collecting, storing, analyzing and alerting GPU abnormal information such as errors, exceptions, and failures by configuring corresponding parameters and strategies, and the operation log of the abnormal state collected by the log collector is delivered to the maintenance personnel. In other words, by developing the logic of the log collection module on the AI ​​server, the frequency and trigger conditions of log collection are set. The storage module (including the local storage module and / or the storage unit in the baseboard management controller) is configured to specify the location and strategy of log storage. The prompt unit can also be configured to set the trigger conditions and notification methods of prompts such as alarms. In this way, when the abnormal monitoring device of the graphics processor and the server are started, the operating status of the graphics processor can be monitored in real time, and the collected error logs can be automatically processed. When the graphics processor is abnormal, the maintenance personnel can also decide whether to directly collect the logs collected by the baseboard management controller, or to actively send IPMI commands to trigger the log collector to collect the operation logs of the graphics processor, and then collect the logs in the baseboard management controller.

[0075] Please refer to Figure 2 as well as Figure 4 , Figure 4 It is a flowchart of the third embodiment of the abnormality monitoring method of the graphics processor of the present application.

[0076] S301: poll the HMC to obtain the running status of the GPU.

[0077] In this embodiment, the BMC may poll the HMC of the GPU through thread_A (a monitoring thread), that is, poll whether the running status of the GPU is in a normal state.

[0078] S302: Evaluate whether the GPU is in a normal state.

[0079] In this embodiment, when it is determined that the GPU is in a normal state, step S301 is executed; when it is determined that the GPU is not in a normal state, step S303 is executed.

[0080] S303: triggering a GPU_Status alarm.

[0081] In this embodiment, GPU_Status indicates the GPU status.

[0082] When the BMC detects that the GPU is in an abnormal state such as an alarm or error, the BMC can trigger the alarm log of the GPU_Status sensor.

[0083] S304: Call OOB_dump_tool.sh to collect GPU logs.

[0084] In this embodiment, OOB_dump_tool.sh is a script of a GPU log collector.

[0085] BMC can also call OOB_dump_tool.sh to collect GPU operation logs and store the operation logs collected by OOB_dump_tool.sh in BMC. In this way, when users, operation and maintenance personnel find GPU_Status alarms, they can call BMC's BMC_OneKeylog.sh to download the operation logs to the local computer for fault analysis. BMC_OneKeylog.sh represents the one-key log collection tool script of BMC.

[0086] It can be seen that in this embodiment, an example is given to illustrate the strategy after the BMC detects that the GPU is in an abnormal state. The following example is given to illustrate the strategy when the GPU fails but the BMC does not detect the abnormal state.

[0087] See also Figure 5 , Figure 5 It is a flowchart of the fourth embodiment of the abnormality monitoring method of the graphics processor of the present application.

[0088] S401: Obtain the IPMI command for collecting OOB logs sent by onekeyLog.

[0089] In one embodiment, the external party may execute BMC_OneKeylog.sh to trigger OOB (log collection tool) through onekeyLog (one-key log collection). The external party refers to the BMC, which may be a user, operation and maintenance personnel, etc.

[0090] S402: Call OOB_dump_tool.sh to collect GPU logs.

[0091] In this embodiment, in response to the execution of BMC_OneKeylog.sh, the log collection function of OOB_dump_tool.sh is actively triggered.

[0092] S403: Store the GPU log in the / var directory.

[0093] In this embodiment, the / var directory indicates the location of the storage module.

[0094] You can download the operation logs collected by OOB_dump_tool.sh to your local computer for fault analysis.

[0095] In summary, the present application can improve the efficiency and accuracy of log collection, reduce the cost and error rate of manual collection. Automated log analysis can also be integrated to quickly locate the cause of the fault and shorten the troubleshooting time. Alternatively, the log is not analyzed for fault location to reduce the burden on the baseboard management controller. For example, for the various error types of the GPU, when the BMC cannot fully define and alarm, the BMC can define a dynamic interface to integrate the manufacturer's log collection tools, and finally present the log to the GPU manufacturer and R&D personnel for analysis, which is conducive to reducing the dilemma of maintenance personnel plugging in on-site to collect logs, and can also improve the efficiency of problem analysis. Flexible storage strategies and alarm mechanisms can meet the needs of different scenarios, and are also conducive to improving the availability and reliability of the system. It can also ensure the operating stability of the AI ​​server, which is conducive to improving the user experience.

[0096] It should be understood that although Figure 1 , Figures 3 to 5 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 , Figures 3 to 5 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0097] See also Figure 6 , Figure 6 It is a structural diagram of an embodiment of an abnormality monitoring device for a graphics processor.

[0098] In one embodiment, the abnormality monitoring device for a graphics processor includes: a baseboard management controller 10 , an abnormality monitoring module 11 and a storage module 12 .

[0099] The abnormality monitoring module 11 is disposed in the baseboard management controller 10 and is used to perform abnormality monitoring on the graphics processor connected to the baseboard management controller 10 by using any of the above-mentioned abnormality monitoring methods for the graphics processor.

[0100] The storage module 12 is used to store the collected operation logs.

[0101] Specifically, the abnormality monitoring module 11 is used to obtain the operating status of the graphics processor connected thereto; evaluate whether the operating status indicates that the graphics processor is currently in an abnormal state; in response to determining that the graphics processor is currently in an abnormal state, call the log collector of the graphics processor, and use the log collector to collect the operating logs of the graphics processor; wherein at least one of the fineness and richness of the operating logs collected by the log collector is higher than the operating logs collected by the baseboard management controller 10.

[0102] In this way, the abnormality monitoring device of the graphics processor can monitor the abnormality of the graphics processor, and can improve the efficiency of collecting operation logs, so as to improve the efficiency of the abnormality cause of the graphics processor and solve the abnormality, thereby improving the reliability and stability of the server. And by collecting the operation logs with a log collector that is more compatible with the graphics processor, the effectiveness of the operation log used to analyze the abnormality cause of the graphics processor can be improved.

[0103] For the specific definition of the abnormality monitoring device of the graphics processor, please refer to the definition of the abnormality monitoring method of the graphics processor mentioned above, which will not be repeated here. Each module in the above-mentioned abnormality monitoring device of the graphics processor can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0104] See also Figure 7 , Figure 7 It is a structural diagram of an embodiment of the server of this application.

[0105] In one embodiment, the server includes a server body 21, a graphics processor 22, and an abnormality monitoring device 23 of the graphics processor as in the above-mentioned embodiment.

[0106] The graphics processor 22 is disposed in the server body 21 .

[0107] The abnormality monitoring device 23 is disposed on the server body 21 and is used to monitor the graphics processor 22 for abnormalities.

[0108] That is to say, in this embodiment, the abnormality monitoring device 23 of the graphics processor mentioned above is used to monitor the graphics processor 22, which can improve the efficiency of collecting operation logs, so as to improve the efficiency of analyzing the abnormality cause of the graphics processor 22 based on the operation log and solving the abnormality, thereby improving the reliability and stability of the server. At the same time, when the graphics processor 22 fails, the operation log can also be collected by a log collector with a higher degree of adaptability to the graphics processor 22, so as to improve the effectiveness of the operation log used to analyze the abnormality cause of the graphics processor 22.

[0109] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the method for monitoring abnormalities of a graphics processor described above are implemented, which will not be described in detail here.

[0110] See also Figure 8 , Figure 8 It is a structural diagram of an embodiment of a computer device of the present application.

[0111] In one embodiment, the computer device may be a server, and its internal structure diagram may be as follows: Figure 8 As shown in the example.

[0112] The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for monitoring abnormalities of a graphics processor is implemented.

[0113] Those skilled in the art will understand that Figure 8 The structure shown as an example is merely a block diagram of a portion of the structure related to the present application solution, and does not constitute a limitation on the computer device to which the present application solution is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0114] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps can be implemented: a baseboard management controller obtains the operating status of a graphics processor connected thereto; an evaluation is made as to whether the operating status indicates that the graphics processor is currently in an abnormal state; in response to determining that the graphics processor is currently in an abnormal state, a log collector of the graphics processor is called, and the log collector is used to collect operating logs of the graphics processor; wherein at least one of the precision and richness of the operating logs collected by the log collector is higher than the operating logs collected by the baseboard management controller.

[0115] In one embodiment, when the log collector of the graphics processor is called, the processor can also implement the following steps when executing the computer program: triggering the monitoring prompt log of the baseboard management controller; wherein the monitoring prompt log indicates monitoring the running status of the graphics processor and a prompt log that the graphics processor is in an abnormal state; sending a management signal to the management flag of the log collector to enable the log collector to activate the log collection function of the log collector; using the log collector to collect the running log of the graphics processor includes: in response to the management flag to enable the log collector, controlling the platform management interface process to search for the script of the log collector; wherein the script of the log collector is a control code pre-compatible with the baseboard management controller; in response to the search to locate the script of the log collector, controlling the platform management interface to execute the script of the log collector to enable the log collector to automatically perform the task of collecting the graphics processor log; obtaining the running log of the graphics processor collected by the log collector; and storing the running log in a preset storage module.

[0116] In one embodiment, when the operation log is stored in a preset storage module, the processor can also implement the following steps when executing the computer program: obtain a log download instruction; in response to obtaining the log download instruction, call a collection logger of the baseboard management controller; use the collection logger to download the operation log collected by the log collector, and save the downloaded operation log to the local storage module, so as to perform graphics processor fault analysis through the operation log stored in the local storage module.

[0117] In one embodiment, after the baseboard management controller obtains the operating status of the graphics processor connected to it, the processor can also implement the following steps when executing the computer program: in response to determining that the graphics processor is in a normal state based on the operating status, determine whether the graphics processor reports an error; in response to the graphics processor reporting an error, generate prompt information; wherein the prompt information is used to feedback the situation that the graphics processor reports an error and the baseboard management controller has not detected it; obtain a trigger instruction to actively trigger the log collection function of the log collector based on the trigger instruction, and use the log collector to collect the operating log of the graphics processor.

[0118] In one embodiment, when the processor executes the computer program, the following steps may also be implemented: obtaining an exception log of the graphics processor; analyzing the exception log to identify the exception type and the cause of the exception represented by the exception log; evaluating the exception level of the graphics processor exception reported by the exception log, and using a preset feedback method matching the exception level to provide feedback on the exception message; wherein the exception message includes at least one of the exception type, the cause of the exception, and the exception level.

[0119] In one embodiment, when obtaining the exception log of the graphics processor, the processor may further implement the following steps when executing the computer program: monitor the hardware control manager of the graphics processor to collect exception information of the graphics processor; convert the format of the exception information to obtain an exception log in a target format and / or containing target information; and send the exception log to a preset storage module for storage.

[0120] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps can be implemented:

[0121] The baseboard management controller obtains the operating status of the graphics processor connected thereto; evaluates whether the operating status feedback indicates that the graphics processor is currently in an abnormal state; in response to determining that the graphics processor is currently in an abnormal state, calls a log collector of the graphics processor, and uses the log collector to collect operating logs of the graphics processor; wherein at least one of the precision and richness of the operating logs collected by the log collector is higher than the operating logs collected by the baseboard management controller.

[0122] In one embodiment, when the log collector of the graphics processor is called, the following steps may also be implemented when the computer program is executed by the processor: triggering a monitoring prompt log of a baseboard management controller; wherein the monitoring prompt log indicates monitoring of the running state of the graphics processor and a prompt log that detects that the graphics processor is in an abnormal state; sending a management signal to the management flag of the log collector to enable the log collector to activate the log collection function of the log collector; using the log collector to collect the running log of the graphics processor includes: in response to the management flag enabling the log collector, controlling the platform management interface process to search for the script of the log collector; wherein the script of the log collector is a control code pre-compatible with the baseboard management controller; in response to the search to locate the script of the log collector, controlling the platform management interface to execute the script of the log collector to enable the log collector to automatically perform the task of collecting the graphics processor log; obtaining the running log of the graphics processor collected by the log collector; and storing the running log in a preset storage module.

[0123] In one embodiment, when the operation log is stored in a preset storage module, the following steps may also be implemented when the computer program is executed by the processor: obtaining a log download instruction; in response to obtaining the log download instruction, calling a collection logger of the baseboard management controller; using the collection logger to download the operation log collected by the log collector, and saving the downloaded operation log to the local storage module, so as to perform graphics processor fault analysis through the operation log stored in the local storage module.

[0124] In one embodiment, after the baseboard management controller obtains the operating status of the graphics processor connected to it, the computer program can also implement the following steps when being executed by the processor: in response to determining that the graphics processor is in a normal state based on the operating status, determining whether the graphics processor reports an error; in response to the graphics processor reporting an error, generating prompt information; wherein the prompt information is used to feedback the situation that the graphics processor reports an error and the baseboard management controller has not detected it; obtaining a trigger instruction to actively trigger the log collection function of the log collector based on the trigger instruction, and using the log collector to collect the operating log of the graphics processor.

[0125] In one embodiment, when the computer program is executed by the processor, the following steps may also be implemented: obtaining an exception log of the graphics processor; analyzing the exception log, identifying the exception type and the cause of the exception represented by the exception log; evaluating the exception level of the graphics processor exception reported by the exception log, and providing feedback on the exception message using a preset feedback method that matches the exception level; wherein the exception message includes at least one of the exception type, the cause of the exception, and the exception level.

[0126] In one embodiment, when obtaining the exception log of the graphics processor, the computer program can also implement the following steps when executed by the processor: monitor the hardware control manager of the graphics processor to collect exception information of the graphics processor; convert the format of the exception information to obtain an exception log in a target format and / or containing target information; and send the exception log to a preset storage module for storage.

[0127] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0128] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0129] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application.

Claims

1. A method for monitoring abnormality of a graphics processor, characterized in that: The abnormality monitoring method comprises: The baseboard management controller obtains the operating status of the graphics processor connected thereto; evaluating whether the operating status provides feedback that the graphics processor is currently in an abnormal state; In response to determining that the graphics processor is currently in an abnormal state, a log collector of the graphics processor is called, and the log collector is used to collect operation logs of the graphics processor; wherein at least one of the precision and richness of the operation logs collected by the log collector is higher than the operation logs collected by the baseboard management controller.

2. The abnormality monitoring method according to claim 1, characterized in that: The log collector calling the graphics processor includes: triggering a monitoring prompt log of the baseboard management controller; wherein the monitoring prompt log indicates monitoring of the running state of the graphics processor and a prompt log that detects that the graphics processor is in an abnormal state; Sending a management signal to the management flag of the log collector to enable the log collector so as to activate the log collection function of the log collector; The collecting of the operation log of the graphics processor by using the log collector includes: In response to the management flag indicating that the log collector is enabled, the control platform management interface process searches for a script of the log collector; wherein the script of the log collector is a control code pre-compatible with the baseboard management controller; In response to searching to locate the script of the log collector, controlling the platform management interface to execute the script of the log collector so as to enable the log collector to automatically perform the task of collecting the graphics processor log; Obtaining the operation log of the graphics processor collected by the log collector; The operation log is stored in a preset storage module.

3. The abnormality monitoring method according to claim 2, characterized in that: The step of storing the operation log in a preset storage module includes: Get log download instructions; In response to obtaining the log download instruction, calling a collection logger of the baseboard management controller; The operation log collected by the log collector is downloaded by using the log collection device, and the downloaded operation log is saved to a local storage module, so as to perform a graphics processor fault analysis through the operation log stored in the local storage module.

4. The abnormality monitoring method according to claim 1, characterized in that: After the baseboard management controller obtains the running status of the graphics processor connected thereto, the baseboard management controller further includes: In response to determining that the graphics processor is in a normal state based on the operating state, determining whether the graphics processor reports an error; In response to the graphics processor reporting an error, generating prompt information; wherein the prompt information is used to feedback the situation that the graphics processor reports an error and the baseboard management controller has not detected it; A trigger instruction is obtained to actively trigger the log collection function of the log collector based on the trigger instruction, and the log collector is used to collect the operation log of the graphics processor.

5. The abnormality monitoring method according to claim 1, characterized in that: The abnormality monitoring method further comprises: Obtaining an exception log of the graphics processor; Analyze the abnormality log to identify the abnormality type and abnormality cause represented by the abnormality log; The abnormality level of the graphics processor abnormality reported by the abnormality log is evaluated, and an abnormality message is fed back using a preset feedback method matching the abnormality level; wherein the abnormality message includes at least one of the abnormality type, the abnormality cause, and the abnormality level.

6. The abnormality monitoring method according to claim 5, characterized in that: The obtaining of the abnormal log of the graphics processor comprises: Monitoring a hardware control manager of the graphics processor to collect abnormal information of the graphics processor; Convert the format of the exception information to obtain the exception log in a target format and / or containing the target information; The abnormal log is sent to a preset storage module for storage.

7. A graphics processor abnormality monitoring device, characterized in that: The abnormality monitoring device comprises: Baseboard management controller; An abnormality monitoring module, provided in the baseboard management controller, for performing abnormality monitoring on the graphics processor connected to the baseboard management controller using the abnormality monitoring method for the graphics processor according to any one of claims 1 to 6; The storage module is used to store the collected operation logs.

8. A server, characterized in that: The server comprises: Server body; A graphics processor, disposed in the server body; The abnormality monitoring device for a graphics processor as described in claim 7, wherein the abnormality monitoring device is arranged in the server body and is used to perform abnormality monitoring on the graphics processor.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the abnormality monitoring method for a graphics processor according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the abnormality monitoring method for a graphics processor according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Monitoring method and device of graphics processor, baseboard management controller and medium

    CN120687327A