Graphics processor monitoring method and device, electronic equipment and storage medium

By establishing multiple integrated circuit bus connections in the server to read graphics processor data in parallel, the problem of data acquisition delay in the prior art is solved, enabling timely reading of graphics processor information and stable operation of the server.

CN120743689BActive Publication Date: 2025-11-21INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511261653.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-21
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

In existing technologies, integrated circuit switching reads data from the graphics processor in a certain order, which is slow and causes information acquisition delays. This makes it difficult to adjust the server's operating parameters in a timely manner, thus affecting the server's operating efficiency.

Method used

By establishing multiple integrated circuit bus connections between multiple graphics processors and the baseboard management controller in the target server, and utilizing multiple threads to read the graphics processor's running data in parallel, parallel data processing under independent communication conditions is achieved.

Benefits of technology

It improves the timeliness of graphics processor information reading, enabling timely capture of rapidly changing data, ensuring stable server operation and rapid fault diagnosis, and enhancing server reliability and availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743689B_ABST
    Figure CN120743689B_ABST
Patent Text Reader

Abstract

The application discloses a kind of monitoring method, device, electronic equipment and storage medium of graphic processor, it is related to electric digital data processing technical field, including in target server, between multiple graphic processors and baseboard management controller, through the connection of multiple integrated circuit buses that baseboard management controller leads out, to make that baseboard management controller and any described graphic processor between satisfy preset independent communication condition, further when needing to carry out graphic processor data acquisition, parallel data processing can be implemented, in related art, integrated circuit switching switch reads the data of graphic processor according to certain order, the speed of data acquisition is slower, it is difficult to adjust the operating parameter of server in time, it is not conducive to maintaining the technical problem of server operating efficiency, reaches the technical effect of guaranteeing graphic processor information reading timeliness, it is advantageous to capture instantaneous change data, further so that server can be adjusted in time operation, guarantee the technical effect of server operating efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electric digital data processing, and in particular to a graphics processor monitoring method and device, electronic equipment and a storage medium. BACKGROUND

[0002] When GPU (Graphics Processing Unit) data is collected, BMC (Baseboard Management Controller) can be used to capture various information of all GPUs, such as temperature, memory temperature, power consumption, manufacturer, model, serial number, part number and other static information, and speed, GPU occupancy, link state, various fault information and other dynamic information.

[0003] In related technologies, an integrated circuit switching switch can be used to select the information of the target GPU to be read, so as to read the data of the GPU in a certain order. For a large AI server with a large number of GPUs or a large number of read information, the speed of reading GPU information by BMC is relatively slow, which causes a delay in obtaining GPU information. For some real-time information such as temperature, power consumption and fault information, if the integrated circuit switching switch is switched one by one and the GPU information is read one by one in a sequential polling manner, the information acquisition will be slow and not timely, which needs to be improved. SUMMARY

[0004] The present application provides a graphics processor monitoring method and device, electronic equipment and a storage medium to at least solve the technical problem that the integrated circuit switching switch reads the data of the graphics processor in a certain order in related technologies, the speed of data acquisition is slow, which causes information acquisition delay, and it is difficult to adjust the running parameters of the server in time, which is not conducive to maintaining the running efficiency of the server.

[0005] The present application provides a graphics processor monitoring method, a plurality of graphics processors in a target server are connected to a baseboard management controller through a plurality of integrated circuit buses led out by the baseboard management controller, so that the baseboard management controller and any graphics processor satisfy a preset independent communication condition, wherein the method comprises the following steps: receiving a data acquisition instruction of at least part of the graphics processors of the target server; in response to the data acquisition instruction, controlling the baseboard management controller to create a corresponding thread for at least part of the graphics processors; using a plurality of threads to read the running data of the corresponding graphics processors in parallel, and generating a control action of the target server according to the information of the graphics processors, so as to adjust the running state of the target server to a preset stable running state by using the control action.

[0006] The application further provides a monitoring device of a graphics processor, a plurality of graphics processors in a target server are connected with a baseboard management controller through a plurality of integrated circuit buses led out by the baseboard management controller, so that a preset independent communication condition is met between the baseboard management controller and any graphics processor, wherein the device comprises: a receiving module configured to receive a data acquisition instruction of at least part of the graphics processors of the target server; a control module configured to create a corresponding thread for the at least part of the graphics processors in response to the data acquisition instruction; and a monitoring module configured to read running data of the corresponding graphics processors in parallel by using the plurality of threads, and generate a regulation and control action of the target server according to information of the graphics processors, so as to adjust a running state of the target server to a preset stable running state by using the regulation and control action.

[0007] The application further provides an electronic device, comprising: a memory configured to store a computer program; and a processor configured to implement the steps of the monitoring method of any of the graphics processors when executing the computer program.

[0008] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program implements the steps of the monitoring method of any of the graphics processors when executed by a processor.

[0009] The application further provides a computer program product, comprising a computer program, and the computer program implements the steps of the monitoring method of any of the graphics processors when executed by a processor.

[0010] According to the application, the plurality of graphics processors and the baseboard management controller in the target server are connected through the plurality of integrated circuit buses led out by the baseboard management controller, so that the preset independent communication condition is met between the baseboard management controller and any graphics processor, and then parallel data processing can be implemented when graphics processor data acquisition is needed, thereby solving the technical problem in the related art that the integrated circuit switching switch reads data of the graphics processor according to a certain order, the data acquisition speed is slow, information acquisition delay is caused, and it is difficult to adjust running parameters of the server in time, which is not conducive to maintaining the running efficiency of the server, and the technical effects of guaranteeing the timeliness of information reading of the graphics processor, being conducive to capturing instantaneous change data, and enabling the server to adjust running in time and guarantee the running efficiency of the server are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort.

[0012] Figure 1 A schematic diagram of a monitoring method of a graphics processor provided for an embodiment of the present application is shown in FIG. 1.

[0013] Figure 2 A flowchart of a monitoring method of a graphics processor provided for an embodiment of the present application is shown in FIG. 2.

[0014] Figure 3 A flowchart of a monitoring method of a graphics processor provided for an embodiment of the present application is shown in FIG. 3.

[0015] Figure 4 A structural diagram of a monitoring device of a graphics processor provided for an embodiment of the present application is shown in FIG. 4. DETAILED DESCRIPTION

[0016] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, any other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0017] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0018] In order for those skilled in the art to better understand the technical solutions of the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0019] With the development of artificial intelligence, deep learning and high-performance computing, GPUs are increasingly widely used in modern servers. In order to increase the GPU computing power of AI servers, more and more manufacturers now begin to use GPU modules or GPU BOX to capture various information of all GPUs through BMC, such as temperature, memory temperature, power consumption, manufacturer, model, serial number, part number and other static information, as well as speed, GPU occupancy, link state, various fault information and other dynamic information. The BMC provides the captured GPU information to the Host BMC through the ipmi or redfish interface.

[0020] In the related art, GPUs are all placed on an integrated circuit bus of a BMC, and an integrated circuit switch is used to select which GPU information to read, which is a sequential reading of GPU method. For large AI servers with a large number of GPUs or a large amount of information to read, due to the large number of GPUs and the large amount of information to read, reading all GPU information through sequential polling is slow, affects the timeliness of GPU information reading, and it is difficult to capture some rapidly changing temperature, power consumption values and quickly capture GPU failures that may occur at any time.

[0021] To solve the above technical problems, the embodiment of the present application constructs a system architecture as shown in Figure 1

[0022] As shown in Figure 1 , it is a schematic diagram of the connection relationship between the GPU and the BMC on the universal substrate (UBB) of the target server. The embodiment of the present application can hang each GPU on an integrated circuit bus (i2c bus) of the BMC, for example, GPU0 is hung on i2c bus0, GPU1 is hung on i2c bus1, GPU2 is hung on i2c bus2, …, and GPUX is hung on i2c busX. Because the GPU is directly hung on a bus of the BMC, there is no integrated circuit switch, so there is only GPU on this i2c bus, and the BMC only reads the information of this GPU through this i2c bus, without switching the integrated circuit switch or reading the information of other GPUs or other devices, so there is no i2c bus to reduce the speed of processing and responding to GPU data because of the need to respond to the data of other devices.

[0023] In actual execution, a thread can be created for the information reading of each graphics processor, Thread0 reads GPU0 information, Thread1 reads GPU1 information, Thread2 reads GPU2 information, …, and ThreadX reads GPUX information. Multiple threads (or processes) can read multiple image processors simultaneously without affecting each other.

[0024] The embodiment of the present application can be based on physical parallelism, so that the BMC has the ability to simultaneously communicate with multiple GPUs, and no longer needs to wait for a switch, thereby solving the problem of bus conflict and switching delay.

[0025] Based on the above architecture, as shown in Figure 2 , the embodiment of the present application provides a graphics processor monitoring method, comprising the following steps:

[0026] ​In step S201, a data collection instruction of at least part of the graphic processors of the target server is received.

[0027] In actual implementation, the embodiment of the present application can receive the data collection instruction, which can be automatically generated based on a timer at a fixed time interval, or initiated through a remote management interface. The data collection instruction can also be triggered due to the satisfaction of some other conditions (e.g., a sensor reading exceeding a threshold).

[0028] In the data collection instruction, an instruction for collecting data of all graphic processors can be included to realize full-coverage graphic processor monitoring, or an instruction for collecting data of specific graphic processors (e.g., graphic processors suspected of having problems) can be included to realize more targeted data collection and reduce energy consumption.

[0029] Optionally, in an embodiment of the present application, before receiving the data collection instruction of at least part of the graphic processors of the target server, the method further includes: obtaining current state data of the target server; selecting at least part of the graphic processors satisfying a preset running condition from a plurality of graphic processors of the target server based on the current state data; and generating the data collection instruction based on the at least part of the graphic processors.

[0030] In some embodiments, data reflecting real-time state of the server, historical data, and current execution process, etc. collected by each component of the target server can be obtained, such as overall power consumption, computer room environment temperature, global rotation speed of a cooling fan, cooling liquid flow / temperature, average temperature and utilization rate of a central processing unit, memory error count, network interface flow, historical trend of the last collected graphic processor data, etc.

[0031] In combination with the current state data, the embodiment of the present application can determine which or all of the graphic processors need to be collected according to the current state data. For example, if an error related to a PCIe link appears in the system log, all graphic processors connected to the link are selected; if the computer room environment temperature sensor reading exceeds a threshold, all graphic processors located on the upper part of the case are selected; if the overall power consumption sharply rises, the graphic processors currently performing a computing task are selected; if the temperature of a graphic processor continuously rises in the previous three monitoring, even if it does not exceed an absolute threshold, it is selected for intensive monitoring; in the absence of abnormalities, all can be collected at a fixed time interval, or part of the graphic processors can be randomly selected for detailed inspection, etc.

[0032] Based on the screening result, the embodiment of the present application can determine corresponding data collection instructions to pay more attention to the target graphic processor, save a large amount of computing resources for other management tasks or perform more complex analysis, reduce unnecessary communication, reduce bus load, ensure timely and reliable transmission of important data, and reduce electromagnetic interference.

[0033] In order to avoid missing detection, the embodiment of the present application can collect data from the graphic processors selected according to the current state data and the historical data in the fixed time interval while collecting data from all graphic processors in the fixed time, thereby effectively ensuring the normal operation of the server.

[0034] In step S202, in response to the data collection instruction, the baseboard management controller creates corresponding threads for at least part of the graphic processors.

[0035] It can be understood that a thread is a task that can be independently run in the operating system of the baseboard management controller.

[0036] After receiving the data collection instruction, the embodiment of the present application can create corresponding independent execution threads for the graphic processors that need to collect data according to the requirements of the data collection instruction. For example, if 8 graphic processors need to be read, 8 threads are created, each thread corresponds to each graphic processor, and software parallelization is realized on the basis of hardware parallelization. Each thread works independently and does not block each other, which is the key software mechanism to realize high-speed concurrent reading.

[0037] Optionally, in an embodiment of the present application, before responding to the data collection instruction, it further includes: obtaining the number of cores of the baseboard management controller; based on the number of cores, allocating a corresponding number of threads to each core of the baseboard management controller; based on the number of graphic processors, determining whether the target server meets the preset under-provisioning condition; if the preset under-provisioning condition is met, creating a corresponding number of threads based on the number.

[0038] In actual execution, when the number of cores of the baseboard management controller chip is sufficient, the multiple threads (processes) created by the embodiment of the application can ensure that each thread (process) is in a core. However, for the baseboard management controller chip with insufficient cores, the kernel resources need to be reasonably allocated. For the baseboard management controller chip with only one core, all the graphic processor monitoring threads (processes) can only be placed in the one kernel resource, but the multi-thread monitoring method can still improve the monitoring efficiency of the graphic processor information, because the multi-thread can make more full use of the chip core. For the baseboard management controller chip with multiple cores such as 2 cores, 4 cores and 8 cores, the graphic processor monitoring threads need to be evenly allocated to each core according to the number of cores. For example, there are 8 graphic processor monitoring threads in total, for the baseboard management controller chip with 2 cores, 4 graphic processor monitoring threads are allocated to each core, and for the baseboard management controller chip with 4 cores, 2 graphic processor monitoring threads are allocated to each core.

[0039] Further, there can be a case that the graphic processors are not fully equipped, for example, the full equipment on one baseboard in the server is 8 graphic processors, but it is not necessarily possible to plug in 8 graphic processors. When creating the graphic processor monitoring threads, the baseboard management controller can create 8 graphic processor monitoring threads according to the full equipment, resulting in the situation that some threads are empty and the chip core resources are wasted. Therefore, before creating the threads, the baseboard management controller first detects whether the graphic processors at the position are in place, and then creates the monitoring threads corresponding to the graphic processors at the position if the graphic processors are in place. For example, if GPU1 is not in place, Thread1 is not created to save the chip core resources.

[0040] Optionally, in an embodiment of the application, before responding to the data collection instruction, the method further comprises: setting the priority of the multiple threads to the target priority to read the running data in parallel.

[0041] In some embodiments, the priority of each graphic processor monitoring thread (process) can be set to be the same, to prevent some graphic processor information from being polled faster and some from being polled slower, and to make the priority of the graphic processor monitoring thread (process) higher than the priority of other device monitoring threads (processes) so that the baseboard management controller can read the data of the graphic processors more quickly.

[0042] In step S203, the running data of the corresponding graphic processors is read in parallel by using the multiple threads, and the control action of the target server is generated according to the information of the graphic processors, so as to adjust the running state of the target server to the preset stable running state by using the control action.

[0043] As a possible implementation, the embodiment of the present application can send read commands to the corresponding graphic processors through the exclusive I2C bus of each thread simultaneously and independently before data reading, and receive returned data. The running data can include temperature, memory temperature, power consumption, manufacturer, model, serial number, part number, and other static data, as well as speed, GPU occupancy, link state, various fault information, and other dynamic data.

[0044] The embodiment of the present application can quickly analyze all the collected image processor data, and make decisions according to the preset strategy, so that the target server can maintain efficient and stable operation.

[0045] For example, if the temperature of GPU2 exceeds a certain threshold, a control action of increasing the speed of the system fan is generated; if the power consumption of GPU5 is abnormal, a control action of trying to reduce the frequency of the GPU is generated, and a log is recorded and an alarm is sent to the administrator; if GPU6 does not respond, a control action of resetting or isolating is generated, and GPU6 is marked as a fault to notify the system, etc.

[0046] The embodiment of the present application can read in parallel, thereby effectively reducing the time of the state information of all graphic processors, and ensuring the timeliness of the data. Due to the extremely high reading frequency and extremely low delay, temperature spikes or power consumption glitches that occur instantaneously can be captured, as well as transient faults, providing a data basis for precise control and rapid fault diagnosis. The server can also take proactive and intelligent actions (control actions) based on real-time data, and kill the problem in the cradle, thereby greatly improving the reliability, availability and stability of the entire server.

[0047] Optionally, in an embodiment of the present application, the running data of the corresponding graphic processor is read in parallel by multiple threads, including: determining whether any thread meets the preset normal running condition; if the preset normal running condition is not met, increasing the failure count until the failure count is greater than the preset threshold, or until any thread meets the preset normal running condition, and clearing the failure count.

[0048] Further, the embodiment of the present application can determine whether the monitoring process is normal before data reading, for example, determining whether the thread is still running, whether there is a crash or exit, determining whether the thread can read data from the target graphic processor within a specified time, determining whether the data read by the thread is valid and conforms to the protocol, etc. For example, the present application can set a software watchdog for each thread (process), and when the thread (process) is normally polled, each polling will clear the watchdog counter to 0, and if the thread (process) is stuck due to various reasons, i.e. the failure count of the watchdog is accumulated when the normal running condition is not met.

[0049] Through the mechanism of failure count and threshold, the embodiment of the present application can effectively distinguish transient error and permanent failure. Transient error is automatically ignored and recovered, avoiding disturbance caused by unnecessary recovery operation, preventing the whole monitoring unit from being misjudged as failure due to single and random communication error, and greatly enhancing the stability of the server.

[0050] Optionally, in an embodiment of the present application, after the failure count is greater than the preset threshold, it further comprises: creating a new thread of the graphic processor corresponding to any thread; writing a processing function for the new thread, so that the processing function of the new thread is consistent with the processing function of any thread, so that the new thread and the graphic processor corresponding to any thread satisfy the preset connection relationship.

[0051] Further, when the watchdog program corresponding to the thread detects that the value of the counter reaches a certain value, the thread is destroyed and a new thread (process) is created. The thread processing function of the newly created thread (process) still uses the processing function of the destroyed thread, so that the new thread can be used to monitor the corresponding graphic processor monitored by the original thread, realizing self-repair of a single monitoring unit, reducing the influence range of failure, reducing the number of manual intervention, and reducing the cost.

[0052] Optionally, in an embodiment of the present application, the running data of the corresponding graphic processor is read by multiple threads in parallel, comprising: obtaining the type of at least part of the graphic processor; reading the running data based on the type, and storing the running data and the type based on the identification of at least part of the graphic processor, so as to query the corresponding running data with the identification and / or type as an index after receiving the data calling instruction of any module of the baseboard management controller.

[0053] In some embodiments, the embodiment of the present application can first detect the type of the graphic processor in each thread, and then read the graphic processor data according to the type of the graphic processor, and store the read graphic processor data according to the serial number and type of the graphic processor.

[0054] For example, the type of GPU0 detected in Thread0 is A, and the read data is stored in the variable GPU_A_0. The type of GPU1 detected in Thread1 is B, and the read data is stored in the variable GPU_B_1. For the graphic processor not in the bit, for example, GPU_Y is not in the bit, because the thread is not created, so the variable GPU_*_Y of the graphic processor with serial number Y has no data. The above variables are allocated according to the type of the graphic processor and whether it is dynamically created. For the non-existent graphic processor, the corresponding variable is not allocated memory.

[0055] After the storage by classification, when the data is called subsequently, the running data of the corresponding graphic processor can be directly retrieved according to the type, serial number and the like as indexes, and then the fast calling of the data is realized.

[0056] Optionally, in an embodiment of the present application, the running data of the corresponding graphic processor is read in parallel by the multiple threads, comprising: mapping the variable for saving the running data of the corresponding graphic processor in the multiple threads into a preset graphic processor data memory area of the global variable of the main process of the baseboard management controller, so as to call the corresponding running data from the preset graphic processor data memory area after receiving the data calling instruction of any module of the baseboard management controller.

[0057] In other embodiments, the variable for saving the graphic processor data is mapped to the memory area of the global variable of the main process of the baseboard management controller in the monitoring thread (process) of the graphic processor, so that the monitoring thread (process) of the graphic processor directly saves the data of the graphic processor into the global variable, and the data storage is faster and more timely.

[0058] Optionally, in an embodiment of the present application, it further comprises: judging whether the multiple threads meet the preset collection completion condition based on the running data; if the multiple threads all meet the preset collection completion condition, generating a reading completion signal; and pushing the reading completion signal to the baseboard management controller to control the baseboard management controller to copy the running data into any global variable.

[0059] As a possible implementation, in an embodiment of the present application, each thread (process) can read all the data of one graphic processor, and finally all the data of the graphic processors are summarized together to be provided to other modules of the baseboard management controller, and the method is that each thread (process) sends a reading completion signal signal to the main process of the baseboard management controller after polling and reading the data of the corresponding graphic processor once, and the main process of the baseboard management controller copies the data of the graphic processor into a global variable after receiving the signal, so that the data integrity and real-time performance can be guaranteed. If the main process of the baseboard management controller directly copies the data of a graphic processor into a global variable, the polling of the graphic processor may not be completed, and the data of the graphic processor is the data of the last time or some data has not been read out.

[0060] In combination with Figure 3 The working principle of the graphic processor monitoring method of an embodiment of the present application is described in detail with reference to the system architecture shown in the accompanying drawings.

[0061] Based on the system architecture shown in the accompanying drawings, the graphic processor monitoring method of an embodiment of the present application comprises the following steps. Figure 1 Based on the system architecture shown in the accompanying drawings, the graphic processor monitoring method of an embodiment of the present application comprises the following steps.Figure 3 As shown, the method can comprise the following steps:

[0062] Step S301, creating a graphics processor monitoring thread. The embodiment of the present application can create a thread for each graphics processor data collection, for example, Thread0 reads the data of GPU0, Thread1 reads the data of GPU1, Thread2 reads the data of GPU2, …, ThreadX reads the data of GPUX. The embodiment of the present application can simultaneously read multiple GPU information in multiple threads (or processes) without affecting each other. For the case of not enough, no thread is established for the graphics processor not in the position.

[0063] Step S302, assigning core resources and setting priority for the graphics processor monitoring thread. When the substrate management controller chip core is enough, ensure that each thread (process) is in a core, and when it is not enough, it can be evenly distributed.

[0064] The embodiment of the present application can set the priority of the monitoring thread of each graphics processor to be consistent, guarantee the polling speed to be consistent, and in the case of needing to acquire graphics processor information more quickly, the priority of the graphics processor monitoring thread can be set to be higher than that of other device monitoring threads.

[0065] Step S303, judging whether the graphics processor monitoring thread is normal. The embodiment of the present application can judge the graphics processor monitoring thread, if it is abnormal, such as stuck, return to step S301, and increase the failure count once, until the graphics processor monitoring thread is normal, clear the failure count, or until the failure count exceeds a certain value, at this time, a new thread can be created to replace the abnormal thread.

[0066] Step S304, reading the data of the graphics processor. The embodiment of the present application can read the data, and store the read data according to the type of the graphics processor, so as to be called subsequently.

[0067] Step S305, transferring the data of the graphics processor to the substrate management controller main process. The substrate management controller main process can copy the collected information to the global variable after confirming that the data collection of all graphics processors is completed.

[0068] In summary, the embodiment of the present application can improve the speed of the substrate management controller reading the data of the graphics processor, especially the reading speed of the temperature related to the graphics processor can be more timely, so that the substrate management controller can obtain the temperature related to the graphics processor in time and then perform heat dissipation regulation and control, so that the heat dissipation regulation and control can be quickly responded and accurately regulated, which is very helpful for reducing the power consumption of the server and can prevent the high temperature of the graphics processor in time and ensure the normal work of the graphics processor. Some dynamic information and fault information of the graphics processor can also be obtained in time, the running time parameters of the graphics processor can be accurately obtained, the fault of the graphics processor can be quickly captured, and the fault of the graphics processor can be handled in time, so that more accurate data is provided for the maintenance of the graphics processor server.

[0069] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment.

[0070] As shown in Figure 4 The embodiment of the present application also provides a graphics processor monitoring device 10, which comprises a receiving module 100, a control module 200 and a monitoring module 300.

[0071] Specifically, the receiving module 100 is used for receiving a data acquisition instruction of at least part of the graphics processor of a target server.

[0072] The control module 200 is used for creating a corresponding thread for the at least part of the graphics processor by the substrate management controller in response to the data acquisition instruction.

[0073] The monitoring module 300 is used for reading the running data of the corresponding graphics processor in parallel by using the plurality of threads, and generating a regulation and control action of the target server according to the information of the graphics processor, so as to adjust the running state of the target server to a preset stable running state by using the regulation and control action.

[0074] Optionally, in an embodiment of the present application, the monitoring module 300 comprises a judgment unit and a second control module.

[0075] The judgment unit is used for judging whether any thread meets a preset normal running condition.

[0076] The control unit is used for increasing a failure count in the case that the preset normal running condition is not met, until the failure count is greater than a preset failure count threshold or until any thread meets the preset normal running condition, and clearing the failure count.

[0077] Optionally, in an embodiment of the present application, the monitoring module 300 further comprises a creating unit and a writing unit.

[0078] The creating unit is configured to create a new thread of the graphic processor corresponding to any thread.

[0079] The writing unit is configured to write the processing function for the new thread, so that the processing function of the new thread is consistent with the processing function of any thread, and the new thread and the graphic processor corresponding to any thread satisfy the preset connection relationship.

[0080] Optionally, in an embodiment of the present application, the monitoring device 10 of the graphic processor further comprises a first obtaining module, a screening module and a first generating module.

[0081] The first obtaining module is configured to obtain the current state data of the target server.

[0082] The screening module is configured to screen at least part of the graphic processors satisfying the preset running condition from the plurality of graphic processors of the target server based on the current state data.

[0083] The first generating module is configured to generate the data acquisition instruction based on the at least part of the graphic processors.

[0084] Optionally, in an embodiment of the present application, the monitoring device 10 of the graphic processor further comprises a second obtaining module, an allocating module, a first judging module and a creating module.

[0085] The second obtaining module is configured to obtain the number of cores of the baseboard management controller.

[0086] The allocating module is configured to allocate a corresponding number of threads for each of the plurality of cores of the baseboard management controller based on the number of cores.

[0087] The first judging module is configured to judge whether the target server satisfies the preset graphic processor under-provisioning condition based on the number of graphic processors.

[0088] The creating module is configured to create a corresponding number of threads based on the number in the case of satisfying the preset under-provisioning condition.

[0089] Optionally, in an embodiment of the present application, the monitoring device 10 of the graphic processor further comprises a setting module.

[0090] The setting module is configured to set the priority of the plurality of threads as a target priority to read the running data in parallel.

[0091] Optionally, in an embodiment of the present application, the monitoring module 300 comprises an obtaining unit and a storage unit.

[0092] The obtaining unit is configured to obtain the type of at least part of the graphic processors.

[0093] The storage unit is configured to read the running data based on the type and store the running data and the type based on the identification of at least part of the graphic processors, so as to query the corresponding running data based on the identification and / or the type as an index after receiving the data calling instruction of any module of the baseboard management controller.

[0094] Optionally, in an embodiment of the present application, the monitoring module 300 comprises a mapping unit.

[0095] The mapping unit is configured to map the variable in the plurality of threads for saving the running data of the corresponding graphic processor into a preset graphic processor data memory area of the global variable of the main process of the baseboard management controller, so as to call the corresponding running data from the preset graphic processor data memory area after receiving the data calling instruction of any module of the baseboard management controller.

[0096] Optionally, in an embodiment of the present application, the monitoring device 10 of the graphic processor further comprises a second judging module, a second generating module and a pushing module.

[0097] The second judging module is configured to judge whether the plurality of threads satisfy a preset collection completion condition based on the running data.

[0098] The second generating module is configured to generate a reading completion signal in the case that the plurality of threads all satisfy the preset collection completion condition.

[0099] The pushing module is configured to push the reading completion signal to the baseboard management controller, so as to control the baseboard management controller to copy the running data into any global variable.

[0100] The features of the embodiments of the monitoring device of the graphic processor can be referred to the related descriptions of the embodiments of the monitoring method of the graphic processor, which will not be repeated here.

[0101] The embodiments of the present application further provide an electronic device comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the embodiments of the monitoring method of the graphic processor.

[0102] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the embodiments of the monitoring method of the graphic processor when running.

[0103] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0104] Embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program, when executed by a processor, implements the steps in any of the above-mentioned embodiments of the monitoring method of the graphic processor.

[0105] Embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps in any of the above-mentioned embodiments of the monitoring method of the graphic processor.

[0106] The skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0107] The above describes in detail the monitoring method, device, electronic equipment and storage medium of the graphic processor provided by the present application. The principles and implementation manners of the present application are described by applying specific examples in this paper, and the above description of the examples is only applicable to help understand the method of the present application and its core idea. It should be pointed out that for the ordinary skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for monitoring a graphics processor, characterized in that, Multiple graphics processors in the target server are respectively connected to a baseboard management controller via multiple integrated circuit buses derived from the baseboard management controller, so that the baseboard management controller and any of the graphics processors meet preset independent communication conditions. The method includes the following steps: Receive data acquisition instructions from at least a portion of the graphics processor of the target server; In response to the data acquisition command, the baseboard management controller is controlled to create corresponding threads for at least a portion of the graphics processors; Multiple threads are used to read the corresponding graphics processor's running data in parallel, and control actions for the target server are generated based on the graphics processor's information, so as to adjust the target server's running state to a preset stable running state using the control actions. The step of using multiple threads to read the corresponding graphics processor's running data in parallel includes: determining whether any thread meets a preset normal operating condition; if the preset normal operating condition is not met, then increasing the failure count until the failure count is greater than a preset count threshold, or until any thread meets the preset normal operating condition, and then clearing the failure count. The step of using multiple threads to read the corresponding graphics processor's running data in parallel includes: obtaining at least a portion of the graphics processor's type; reading the running data based on the type; and storing the running data and the type based on the identifier of at least a portion of the graphics processor, so that upon receiving a data call instruction from any module of the baseboard management controller, the corresponding running data can be queried using the identifier and / or the type as an index. The step of using multiple threads to read the corresponding graphics processor's runtime data in parallel includes: The variables used to store the corresponding graphics processor's running data in the multiple threads are mapped to the preset graphics processor data memory area of ​​the global variables of the main process of the baseboard management controller, so that after receiving a data call instruction from any module of the baseboard management controller, the corresponding running data can be called from the preset graphics processor data memory area.

2. The monitoring method for a graphics processor according to claim 1, characterized in that, After the failure count exceeds a preset count threshold, the method further includes: Create a new thread for the graphics processor corresponding to any of the described threads; Write a processing function for the new thread so that the processing function of the new thread is consistent with the processing function of any of the threads, and so that the new thread and the graphics processor corresponding to any of the threads satisfy a preset connection relationship.

3. The monitoring method for a graphics processor according to claim 1, characterized in that, Before receiving data acquisition instructions from at least a portion of the graphics processor of the target server, the process also includes: Obtain the current status data of the target server; Based on the current status data, at least a portion of the graphics processors that meet the preset operating conditions are selected from the plurality of graphics processors of the target server; The data acquisition instructions are generated based on at least a portion of the graphics processor.

4. The monitoring method for a graphics processor according to claim 1, characterized in that, Prior to responding to the data acquisition command, it also includes: Obtain the number of cores of the baseboard management controller; Based on the number of cores, a corresponding number of threads are allocated to each of the multiple cores of the baseboard management controller; Based on the number of graphics processors, determine whether the target server meets the preset condition of insufficient graphics processor configuration; If the preset condition of insufficient graphics processor configuration is met, then the corresponding number of threads are created based on the specified quantity.

5. The monitoring method for a graphics processor according to claim 1, characterized in that, Prior to responding to the data acquisition command, it also includes: The priorities of multiple threads are set to a target priority to read the running data in parallel.

6. The monitoring method for a graphics processor according to claim 1, characterized in that, Also includes: Based on the running data, determine whether the multiple threads meet the preset data collection completion conditions; If all of the aforementioned threads meet the preset acquisition completion condition, a read completion signal is generated; The read completion signal is pushed to the baseboard management controller to control the baseboard management controller to copy the running data to any global variable.

7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for implementing the steps of the monitoring method for a graphics processor as described in any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Gene data analysis method and heterogeneous scheduling platform

    CN110427262A

  • Baseboard management controller system operation method and apparatus, device, and non-volatile readable storage medium

    WO2024230401A1