Monitoring method and device of graphics processor, electronic equipment and storage medium
By establishing multiple integrated circuit bus connections and multi-threaded parallel data reading in the server, the problem of slow graphics processor data reading speed is solved, and timely acquisition of graphics processor information and efficient and stable operation of the server are achieved.
Patent Information
- Application Number
- CN202511261653.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-04
AI Technical Summary
In the prior art, the integrated circuit switching switch reads data from the graphics processor in a certain order at a slow speed, resulting in information acquisition delays, making it difficult to adjust the server's operating parameters in a timely manner, and affecting the server's operating efficiency.
By establishing multiple integrated circuit bus connections between multiple graphics processors and baseboard management controllers in the target server, and using multiple threads to read the operating data of the graphics processors in parallel, parallel data processing under independent communication conditions is achieved.
It improves the timeliness of graphics processor information reading and can capture instantaneous changes in data in a timely manner, ensuring that the server can make timely operational adjustments and guaranteeing efficient and stable operation of the server.
Smart Images

Figure CN120743689A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a monitoring method, device, electronic equipment and storage medium for a graphics processor. Background Art
[0002] When collecting GPU (Graphics Processing Unit) data, you can use the Baseboard Management Controller (BMC) to capture various information about all GPUs, such as temperature, memory temperature, power consumption, manufacturer, model, serial number, part number, and other static information, as well as speed, GPU utilization, link status, and various fault information.
[0003] In related technologies, the information of the target GPU to be read can be selected through an integrated circuit switching switch to read the GPU data in a certain order. For large AI servers with a large number of GPUs or a large amount of information to be read, the BMC reads GPU information relatively slowly, which will cause delays in obtaining GPU information. For some information with high real-time requirements such as temperature, power consumption, fault information, etc., if the integrated circuit switching switches are switched one by one through sequential polling, and the GPU information is read one by one, it will cause slow and untimely information acquisition, which urgently needs to be improved. Summary of the Invention
[0004] The present invention provides a method, device, electronic device, and storage medium for monitoring a graphics processor, aiming to at least address the technical problem in related technologies where an integrated circuit switching switch reads graphics processor data in a certain order, resulting in a slow data acquisition speed and thus causing information collection delays. This makes it difficult to adjust server operating parameters in a timely manner, hindering the maintenance of server operating efficiency.
[0005] The present invention provides a method for monitoring a graphics processor, wherein multiple graphics processors in a target server are respectively connected to a baseboard management controller through multiple integrated circuit buses led out of the baseboard management controller, so that preset independent communication conditions are met between the baseboard management controller and any of the graphics processors. The method includes the following steps: receiving data acquisition instructions from at least some of the graphics processors of the target server; in response to the data acquisition instructions, controlling the baseboard management controller to create corresponding threads for at least some of the graphics processors; using multiple threads to read operating data of the corresponding graphics processors in parallel, and generating a control action for the target server based on the information of the graphics processors, so as to adjust the operating state of the target server to a preset stable operating state using the control action.
[0006] The present invention also provides a monitoring device for a graphics processor, wherein multiple graphics processors in a target server are respectively connected to a baseboard management controller through multiple integrated circuit buses led out by the baseboard management controller, so that preset independent communication conditions are met between the baseboard management controller and any of the graphics processors, wherein the device includes: a receiving module for receiving data acquisition instructions of at least some of the graphics processors of the target server; a control module for controlling the baseboard management controller to create corresponding threads for at least some of the graphics processors in response to the data acquisition instructions; and a monitoring module for using multiple threads to read the operating data of the corresponding graphics processors in parallel, and generating a control action of the target server based on the information of the graphics processor, so as to adjust the operating state of the target server to a preset stable operating state by using the control action.
[0007] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned graphics processor monitoring methods when executing the computer program.
[0008] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for monitoring a graphics processor are implemented.
[0009] The present invention also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned graphics processor monitoring methods when executed by a processor.
[0010] Through the present invention, in a target server, multiple graphics processors and a baseboard management controller can be connected through multiple integrated circuit buses led out of the baseboard management controller, so that the baseboard management controller and any of the graphics processors meet preset independent communication conditions, and thus parallel data processing can be achieved when graphics processor data collection is required. This solves the technical problem in the related art that an integrated circuit switching switch reads graphics processor data in a certain order, the data acquisition speed is slow, and thus information collection delays occur, making it difficult to adjust the server's operating parameters in a timely manner, which is not conducive to maintaining the server's operating efficiency. The present invention achieves the technical effect of ensuring the timeliness of graphics processor information reading, being conducive to capturing instantaneous change data, and thus enabling the server to make timely operation adjustments, thereby ensuring the server's operating efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 A schematic diagram illustrating the principle of a method for monitoring a graphics processor according to an embodiment of the present invention; Figure 2 A flowchart of a method for monitoring a graphics processor provided by an embodiment of the present invention; Figure 3 A flowchart of a method for monitoring a graphics processor according to an embodiment of the present invention; Figure 4 A schematic diagram of the structure of a monitoring device for a graphics processor provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0014] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.
[0015] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0016] With the development of artificial intelligence, deep learning, and high-performance computing, GPUs are increasingly being used in modern servers. To increase the GPU computing power of AI servers, more and more manufacturers are using GPU modules or GPU boxes. These cards use the BMC to capture various GPU information, including static information such as temperature, memory temperature, power consumption, manufacturer, model, serial number, and part number, as well as dynamic information such as speed, GPU utilization, link status, and various fault information. The BMC then provides this captured GPU information to the host BMC via interfaces such as IPMI or Redfish.
[0017] In related technologies, all GPUs are placed on a BMC integrated circuit bus, and an integrated circuit switch is used to select which GPU to read. This is a sequential way of reading GPUs. For large AI servers with a large number of GPUs or a large amount of information to read, reading all GPU information through sequential polling is slow due to the large number of GPUs and the large amount of information that needs to be read. This affects the GPU information reading timeliness, makes it difficult to capture some rapidly changing temperature and power consumption values, and makes it difficult to quickly capture GPU failures that may occur at any time.
[0018] In order to solve the above technical problems, the embodiment of the present invention constructs the following Figure 1 The system architecture shown.
[0019] Among them, such as Figure 1 Figure 2 shows the connection relationship between the GPU and BMC on a universal baseboard (UBB) of a target server. In this embodiment of the present invention, each GPU can be connected to a separate integrated circuit bus (IC bus) of the BMC. For example, GPU0 is connected to I2C bus0, GPU1 to I2C bus1, GPU2 to I2C bus2, and so on. GPUX is connected to I2C busX. Because the GPUs are directly connected to a BMC bus without an IC switch, only the GPU is connected to this I2C bus. The BMC reads information from only this one GPU via this I2C bus, without switching the IC switch or reading information from other GPUs or other devices. Therefore, the I2C bus does not slow down processing and responding to GPU data due to the need to respond to data from other devices.
[0020] In practice, a separate thread can be created for each GPU to read information: Thread0 reads GPU0, Thread1 reads GPU1, Thread2 reads GPU2, and so on, until ThreadX reads GPUX. This allows multiple threads (or processes) to read data from multiple GPUs simultaneously without interfering with each other.
[0021] The embodiment of the present invention can be based on physical parallelism, so that the BMC has the ability to physically communicate with multiple GPUs at the same time without waiting for a switch, thereby solving the problems of bus conflict and switching delay.
[0022] Based on the above architecture, Figure 2 As shown, an embodiment of the present invention provides a method for monitoring a graphics processor, comprising the following steps: In step S201, a data collection instruction is received from at least part of a graphics processor of a target server.
[0023] During actual execution, embodiments of the present invention may receive data collection instructions, wherein the data collection instructions may be automatically generated at fixed time intervals based on a timer, may be initiated through a remote management interface, or may be triggered because certain other conditions are met (e.g., a sensor reading exceeds a threshold).
[0024] The data collection instructions may include instructions for collecting data from all graphics processors to achieve full coverage of graphics processor monitoring, or instructions for collecting data from specific graphics processors (for example, graphics processors suspected of having problems) to achieve more targeted data collection and reduce energy consumption.
[0025] Optionally, in one embodiment of the present invention, before receiving the data collection instruction of at least part of the graphics processors of the target server, it also includes: obtaining the current status data of the target server; filtering out at least part of the graphics processors that meet the preset operating conditions from the multiple graphics processors of the target server based on the current status data; and generating the data collection instruction based on at least part of the graphics processors.
[0026] In some embodiments, data collected by various components of the target server that can reflect the real-time status of the server, historical data, and current execution processes can be obtained, such as the power consumption of the entire machine, the ambient temperature of the computer room, the global speed of the cooling fan, the coolant flow / temperature, the average temperature and utilization of the central processing unit, the memory error count, the network interface traffic, the historical trend of the graphics processor data collected last time, etc.
[0027] In combination with current status data, embodiments of the present invention can determine which or all graphics processors require data collection based on the current status data. For example, if a PCIe link-related error appears in the system log, all graphics processors connected to that link are screened out; if the room ambient temperature sensor reading exceeds a threshold, all graphics processors located in the upper chassis are screened out; if the overall power consumption of the entire machine rises sharply, the graphics processor currently executing computing tasks is screened out; if the temperature of a graphics processor continues to rise during the first three monitoring sessions, even if it does not exceed the absolute threshold, it is screened out for key monitoring; in the absence of abnormalities, all data can be collected regularly, or some graphics processors can be randomly screened for detailed inspection, etc.
[0028] Based on the screening results, the embodiment of the present invention can determine the corresponding data acquisition instructions to focus on the target graphics processor in a more targeted manner, saving a large amount of computing resources for other management tasks or performing more complex analysis, reducing unnecessary communication, reducing bus load, ensuring timely and reliable transmission of important data, and also helping to reduce electromagnetic interference.
[0029] To avoid missed detection, the embodiment of the present invention can collect data from all graphics processors at a fixed time, and at a fixed time interval, select graphics processors that need to be paid special attention to for data collection based on current status data and historical data, thereby effectively ensuring the normal operation of the server.
[0030] In step S202 , in response to the data acquisition instruction, the baseboard management controller is controlled to create corresponding threads for at least part of the graphics processor.
[0031] It is understandable that a thread is a task that can be run independently in the operating system of the baseboard management controller.
[0032] After receiving a data acquisition instruction, an embodiment of the present invention can create corresponding independent execution threads for each graphics processor that needs to perform data acquisition according to the requirements of the data acquisition instruction. For example, if eight graphics processors need to be read, eight threads are created, each thread corresponding to each graphics processor, so as to achieve software parallelism based on hardware parallelism. In particular, each thread works independently and does not block each other. This is a key software mechanism for achieving high-speed concurrent reading.
[0033] Optionally, in one embodiment of the present invention, before responding to the data acquisition instruction, it also includes: obtaining the number of cores of the baseboard management controller; based on the number of cores, allocating a corresponding number of threads to the multiple cores of the baseboard management controller; based on the number of graphics processors, judging whether the target server meets the preset graphics processor under-matching condition; if the preset under-matching condition is met, creating a corresponding number of threads based on the quantity.
[0034] In actual implementation, the multiple threads (processes) created by the embodiments of the present invention can ensure that each thread (process) resides within a core when the baseboard management controller chip has sufficient cores. However, for baseboard management controller chips with fewer cores, core resources must be properly allocated. For baseboard management controller chips with only one core, all GPU monitoring threads (processes) can only be placed within this single core. However, a multi-threaded monitoring method can still improve GPU information monitoring efficiency because multi-threading can more fully utilize chip cores. For baseboard management controller chips with multiple cores (such as 2-core, 4-core, or 8-core), the GPU monitoring threads need to be evenly distributed to each core based on the number of cores. For example, if there are a total of 8 GPU monitoring threads, for a 2-core baseboard management controller chip, each core is allocated 4 GPU monitoring threads, and for a 4-core baseboard management controller chip, each core is allocated 2 GPU monitoring threads.
[0035] Furthermore, there may be a situation where the graphics processor is not fully configured. For example, one of the baseboards in the server is fully configured with 8 graphics processors, but it is not necessarily possible to fully insert 8 graphics processors. When creating the graphics processor monitoring thread, the baseboard management controller may create 8 graphics processor monitoring threads according to the full configuration, resulting in some threads running idle, thereby wasting chip core resources. Therefore, before creating the thread, the baseboard management controller first detects whether the graphics processor at this location is in place, and then creates the monitoring thread corresponding to the graphics processor here if it is in place. For example, if GPU1 is not in place, thread Thread1 will not be created to save chip core resources.
[0036] Optionally, in one embodiment of the present invention, before responding to the data collection instruction, the method further includes: setting the priorities of multiple threads to a target priority to read the running data in parallel.
[0037] In some embodiments, the priority of each GPU monitoring thread (process) can be set to the same high to prevent some GPUs from being polled faster than others. In addition, to allow the baseboard management controller to read GPU data more quickly, the priority of the GPU monitoring thread (process) is set higher than that of other device monitoring threads (processes).
[0038] In step S203, multiple threads are used to read the operating data of the corresponding graphics processor in parallel, and a control action of the target server is generated according to the information of the graphics processor, so as to adjust the operating state of the target server to a preset stable operating state by using the control action.
[0039] As a possible implementation, before reading data, embodiments of the present invention can utilize all previously created threads to simultaneously and independently send read commands to the corresponding graphics processor via their dedicated I2C bus and receive return data. This operational data can include static data such as temperature, memory temperature, power consumption, manufacturer, model, serial number, and part number, as well as dynamic data such as speed, GPU occupancy, link status, and various fault information.
[0040] The embodiment of the present invention can quickly analyze all collected image processor data and make decisions according to preset strategies, thereby enabling the target server to maintain efficient and stable operation.
[0041] For example, if the temperature of GPU2 exceeds a certain threshold, a control action is generated to increase the system fan speed; if the power consumption of GPU5 is abnormal, a control action is generated to try to reduce the frequency of the GPU, and a log is recorded and an alarm is sent to the administrator; if GPU6 is unresponsive, a reset or isolation control action is generated, and GPU6 is marked as a fault to notify the system, etc.
[0042] This embodiment of the present invention enables parallel reading, effectively reducing the time required to read status information from all graphics processors and ensuring data timeliness. Due to its extremely high read frequency and extremely low latency, it can capture instantaneous temperature spikes, power consumption glitches, and transient faults, providing a data foundation for precise control and rapid fault diagnosis. The server can also proactively and intelligently take action (control actions) based on real-time data to address problems before they occur, significantly improving the reliability, availability, and stability of the entire server.
[0043] Optionally, in one embodiment of the present invention, multiple threads are used to read the operating data of the corresponding graphics processor in parallel, including: determining whether any thread meets the preset normal operating conditions; if the preset normal operating conditions are not met, increasing the failure count until the failure count is greater than a preset count threshold or until any thread meets the preset normal operating conditions, and clearing the failure count.
[0044] Furthermore, before reading data, embodiments of the present invention can determine whether the monitoring process is normal, for example, whether the thread is still running, whether it has crashed or exited, whether the thread can read data from the target graphics processor within the specified time, and whether the data read back by the thread is valid and complies with the protocol. Taking the acquisition process being stuck as an example, embodiments of the present invention can set a software watchdog for each thread (process). When the thread (process) is polling normally, the watchdog counter will be cleared to 0 each time it is polled. If the thread (process) is stuck for various reasons, that is, if the normal operating conditions are not met, the watchdog failure count will be accumulated.
[0045] By using a failure count and threshold mechanism, the present invention effectively distinguishes between transient errors and permanent faults. Transient errors are automatically ignored and recovered, avoiding unnecessary recovery operations and preventing the entire monitoring unit from being misidentified as faulty due to a single, random communication error, thus significantly enhancing server stability.
[0046] Optionally, in one embodiment of the present invention, after the failure count is greater than a preset count threshold, the method further includes: creating a new thread of the graphics processor corresponding to any thread; writing a processing function for the new thread so that the processing function of the new thread is consistent with the processing function of any thread, so that a preset connection relationship is satisfied between the new thread and the graphics processor corresponding to any thread.
[0047] Furthermore, when the watchdog program of the corresponding thread detects that the counter value reaches a certain value, it destroys the thread and then creates a new thread (process). The thread processing function of the newly created thread (process) still uses the processing function of the destroyed thread, so that the new thread can be used to monitor the corresponding graphics processor monitored by the original thread, realizing self-repair of a single monitoring unit, reducing the scope of fault impact, reducing the number of manual interventions, and reducing costs.
[0048] Optionally, in one embodiment of the present invention, multiple threads are used to read the operating data of the corresponding graphics processor in parallel, including: obtaining the type of at least part of the graphics processor; reading the operating data based on the type, and storing the operating data and type based on the identifier of at least part of the graphics processor, so that after receiving a data call instruction from any module of the baseboard management controller, the corresponding operating data can be queried using the identifier and / or type as an index.
[0049] In some embodiments, the present invention may first detect the type of the graphics processor in each thread, then read the graphics processor data according to the graphics processor type, and store the read graphics processor data according to the sequence number and type of the graphics processor.
[0050] For example, if Thread0 detects that GPU0 is type A, the data read from it is stored in the variable GPU_A_0. If Thread1 detects that GPU1 is type B, the data read from it is stored in the variable GPU_B_1. For a non-existent GPU, such as GPU_Y, the corresponding GPU_*_Y variables will contain no data because the thread has not been created. These variables are dynamically created based on the GPU type and whether they exist. Memory is not allocated for non-existent GPUs.
[0051] After classified storage, when subsequently calling data, the embodiment of the present invention can directly retrieve the corresponding graphics processor's operating data based on the type, sequence number, etc. as an index, thereby realizing rapid data calling.
[0052] Optionally, in one embodiment of the present invention, multiple threads are used to read the operating data of the corresponding graphics processor in parallel, including: mapping the variables used to store the operating data of the corresponding graphics processor in the multiple threads to the preset graphics processor data memory area of the global variables of the main process of the baseboard management controller, so as to call the corresponding operating data from the preset graphics processor data memory area after receiving the data call instruction of any module of the baseboard management controller.
[0053] In other embodiments, in the monitoring thread (process) of the graphics processor, the variable used to store the graphics processor data is mapped to the memory area of the global variable in the main process of the baseboard management controller that stores the corresponding graphics processor data. In this way, the graphics processor monitoring thread (process) directly saves the graphics processor data to the global variable. The advantage of this is that data storage is faster and more timely.
[0054] Optionally, in one embodiment of the present invention, it also includes: judging whether multiple threads meet preset acquisition completion conditions based on the running data; if multiple threads all meet the preset acquisition completion conditions, generating a read completion signal; pushing the read completion signal to the baseboard management controller to control the baseboard management controller to copy the running data to any global variable.
[0055] As a possible implementation, in an embodiment of the present invention, each thread (process) can read all the data from a graphics processor (GPU). Ultimately, the data from all GPUs is aggregated and provided to other modules of the baseboard management controller (BMC). This is accomplished by sending a read completion signal to the BMC's main process after completing a poll to read the corresponding GPU's data. Upon receiving this signal, the BMC's main process copies the GPU's data to a global variable, ensuring data integrity and real-time performance. If the BMC's main process directly copies a GPU's data to a global variable, polling for the current GPU may not be complete, and the GPU's data may still be the previous data, or some data may not yet be read.
[0056] Combine Figure 3 As shown, the working principle of the method for monitoring a graphics processor according to an embodiment of the present invention is described in detail using an embodiment.
[0057] by Figure 1 Based on the system architecture shown, the present invention implements Figure 3 As shown, the following steps may be included: Step S301: Create a GPU monitoring thread. This embodiment of the present invention can create a thread for each GPU data collection. For example, Thread 0 reads data from GPU 0, Thread 1 reads data from GPU 1, Thread 2 reads data from GPU 2, and so on, and Thread X reads data from GPU X. This embodiment of the present invention allows multiple threads (or processes) to simultaneously read information from multiple GPUs without interfering with each other. If the GPU is not fully configured, no thread is created for the missing GPU.
[0058] Step S302: allocate core resources and set priorities for the GPU monitoring thread. When the BMC chip has enough cores, each thread (process) is guaranteed to be in one core. If not, they can be evenly distributed.
[0059] The embodiment of the present invention can set the priority of each GPU monitoring thread to be consistent to ensure consistent polling speed. In addition, when it is necessary to obtain GPU information more quickly, the priority of the GPU monitoring thread can be set higher than that of other device monitoring threads.
[0060] Step S303 determines whether the GPU monitoring thread is normal. This embodiment of the present invention can determine whether the GPU monitoring thread is normal. If it is abnormal, such as stuck, the process returns to step S301 and increments the failure count until the GPU monitoring thread is normal and the failure count is cleared, or until the failure count exceeds a certain value. At this point, a new thread can be created to replace the abnormal thread.
[0061] Step S304: Reading data from the graphics processor. This embodiment of the present invention can read data and classify and store the read data according to the type of the graphics processor for subsequent use.
[0062] Step S305: The data of the graphics processor is transmitted to the baseboard management controller main process. After confirming that the data collection of all graphics processors is completed, the baseboard management controller main process can copy the collected information into the global variable.
[0063] In summary, embodiments of the present invention can improve the speed at which a baseboard management controller reads graphics processor data, particularly enabling more timely reading of graphics processor-related temperatures. This allows the baseboard management controller to promptly obtain graphics processor-related temperatures and perform heat dissipation control, enabling rapid response and precise heat dissipation control. This significantly helps reduce server power consumption and more promptly prevents the graphics processor from overheating, ensuring its normal operation. Dynamic and fault information about the graphics processor can also be obtained more promptly, enabling more accurate acquisition of graphics processor runtime parameters, rapid detection of graphics processor faults, and timely resolution of these faults, providing more accurate data for maintaining graphics processor servers.
[0064] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0065] like Figure 4 As shown, an embodiment of the present invention further provides a monitoring device 10 for a graphics processor, comprising: a receiving module 100 , a control module 200 and a monitoring module 300 .
[0066] Specifically, the receiving module 100 is configured to receive data collection instructions from at least part of the graphics processor of the target server.
[0067] The control module 200 is configured to control the baseboard management controller to create corresponding threads for at least part of the graphics processor in response to the data acquisition instruction.
[0068] The monitoring module 300 is used to read the operating data of the corresponding graphics processor in parallel using multiple threads, and generate a control action for the target server based on the information of the graphics processor, so as to adjust the operating state of the target server to a preset stable operating state using the control action.
[0069] Optionally, in one embodiment of the present invention, the monitoring module 300 includes: a judgment unit and a second control module.
[0070] The judging unit is used to judge whether any thread meets the preset normal running conditions.
[0071] The control unit is configured to increase the failure count when the preset normal operating conditions are not met until the failure count is greater than a preset count threshold or until any thread meets the preset normal operating conditions, and clear the failure count.
[0072] Optionally, in one embodiment of the present invention, the monitoring module 300 further includes: a creating unit and a writing unit.
[0073] The creation unit is used to create a new thread of the graphics processor corresponding to any thread.
[0074] The writing unit is used to write a processing function for the new thread so that the processing function of the new thread is consistent with the processing function of any thread, so that the new thread and the graphics processor corresponding to any thread meet the preset connection relationship.
[0075] Optionally, in one embodiment of the present invention, the monitoring device 10 for a graphics processor further includes: a first acquisition module, a screening module, and a first generation module.
[0076] The first acquisition module is used to acquire the current status data of the target server.
[0077] The screening module is configured to screen out at least some of the graphics processors that meet preset operating conditions from the multiple graphics processors of the target server based on the current status data.
[0078] The first generating module is configured to generate a data acquisition instruction based on at least part of the graphics processor.
[0079] Optionally, in one embodiment of the present invention, the monitoring device 10 for a graphics processor further includes: a second acquisition module, an allocation module, a first judgment module, and a creation module.
[0080] The second acquisition module is used to acquire the number of cores of the baseboard management controller.
[0081] The allocation module is used to allocate a corresponding number of threads to the multiple cores of the baseboard management controller based on the number of cores.
[0082] The first judgment module is configured to judge whether the target server meets a preset graphics processor under-allocation condition based on the number of graphics processors.
[0083] The creation module is used to create a corresponding number of threads based on the quantity when the preset non-full allocation conditions are met.
[0084] Optionally, in one embodiment of the present invention, the monitoring device 10 for a graphics processor further includes: a setting module.
[0085] The setting module is used to set the priorities of multiple threads to the target priority level so as to read the running data in parallel.
[0086] Optionally, in one embodiment of the present invention, the monitoring module 300 includes: an acquisition unit and a storage unit.
[0087] The acquiring unit is configured to acquire the types of at least part of the graphics processor.
[0088] The storage unit is configured to read the operating data based on the type and store the operating data and the type based on at least part of the identification of the graphics processor, so that upon receiving a data call instruction from any module of the baseboard management controller, the corresponding operating data can be queried using the identification and / or type as an index.
[0089] Optionally, in one embodiment of the present invention, the monitoring module 300 includes: a mapping unit.
[0090] The mapping unit is configured to map variables in multiple threads for storing operating data of corresponding graphics processors to a preset graphics processor data memory area of global variables of a main process of a baseboard management controller, so as to call corresponding operating data from the preset graphics processor data memory area upon receiving a data call instruction from any module of the baseboard management controller.
[0091] Optionally, in one embodiment of the present invention, the monitoring device 10 for a graphics processor further includes: a second judgment module, a second generation module, and a push module.
[0092] The second judgment module is used to judge whether the multiple threads meet the preset collection completion conditions based on the running data.
[0093] The second generating module is used to generate a reading completion signal when multiple threads all meet the preset acquisition completion condition.
[0094] The push module is used to push a read completion signal to the baseboard management controller to control the baseboard management controller to copy the operation data to any global variable.
[0095] For descriptions of features in the embodiments corresponding to the graphics processor monitoring device, reference can be made to the descriptions of the embodiments corresponding to the graphics processor monitoring method, which will not be detailed here.
[0096] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any of the above-mentioned graphics processor monitoring method embodiments.
[0097] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned graphics processor monitoring method embodiments when running.
[0098] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0099] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned graphics processor monitoring method embodiments are implemented.
[0100] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned graphics processor monitoring method embodiments are implemented.
[0101] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0102] The above describes in detail the graphics processor monitoring method, device, electronic device, and storage medium provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core concept of the present invention. It should be noted that for those skilled in the art, without departing from the principles of the present invention, various improvements and modifications can be made to the present invention, and such improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A method for monitoring a graphics processor, characterized in that: Multiple graphics processors in a target server are respectively connected to a baseboard management controller via multiple integrated circuit buses derived from the baseboard management controller, so that a preset independent communication condition is met between the baseboard management controller and any of the graphics processors, wherein the method includes the following steps: receiving a data collection instruction of at least a portion of a graphics processor of a target server; In response to the data acquisition instruction, controlling the baseboard management controller to create corresponding threads for at least part of the graphics processors; The plurality of threads are used to read the operating data of the corresponding graphics processor in parallel, and a control action of the target server is generated according to the information of the graphics processor, so as to adjust the operating state of the target server to a preset stable operating state by using the control action.
2. The method for monitoring a graphics processor according to claim 1, wherein: The using the plurality of threads to read the corresponding graphics processor's operating data in parallel includes: Determining whether any of the threads meets a preset normal operating condition; If the preset normal operation condition is not met, the failure count is increased until the failure count is greater than a preset count threshold or until any one of the threads meets the preset normal operation condition, and the failure count is cleared.
3. The method for monitoring a graphics processor according to claim 2, wherein: After the failure count exceeds a preset count threshold, the method further includes: Creating a new thread of the graphics processor corresponding to any of the threads; A processing function is written for the new thread, so that the processing function of the new thread is consistent with the processing function of any of the threads, and a preset connection relationship is satisfied between the new thread and the graphics processor corresponding to any of the threads.
4. The method for monitoring a graphics processor according to claim 1, wherein: Before receiving the data collection instruction of at least part of the graphics processors of the target server, the method further includes: Obtaining current status data of the target server; Filtering at least some of the graphics processors that meet preset operating conditions from the multiple graphics processors of the target server based on the current state data; The data acquisition instructions are generated based at least in part on the graphics processor.
5. The method for monitoring a graphics processor according to claim 1, wherein: Before responding to the data collection instruction, the method further includes: Obtaining the number of cores of the baseboard management controller; Based on the number of cores, allocating a corresponding number of threads to each of the multiple cores of the baseboard management controller; Based on the number of the graphics processors, determining whether the target server meets a preset graphics processor under-allocation condition; If the preset under-matching condition is met, a corresponding number of threads are created based on the number.
6. The method for monitoring a graphics processor according to claim 1, wherein: Before responding to the data collection instruction, the method further includes: The priorities of the plurality of threads are set to a target priority so as to read the running data in parallel.
7. The method for monitoring a graphics processor according to claim 1, wherein: The using the plurality of threads to read the corresponding graphics processor's operating data in parallel includes: Obtaining at least part of the type of the graphics processor; The operating data is read based on the type, and the operating data and the type are stored based on at least part of the identifier of the graphics processor, so that after receiving a data call instruction from any module of the baseboard management controller, the corresponding operating data can be queried using the identifier and / or the type as an index.
8. The method for monitoring a graphics processor according to claim 1, wherein: The using the plurality of threads to read the corresponding graphics processor's operating data in parallel includes: Variables in the plurality of threads for storing operating data of corresponding graphics processors are mapped to a preset graphics processor data memory area of global variables of a main process of the baseboard management controller, so that upon receiving a data call instruction from any module of the baseboard management controller, corresponding operating data can be called from the preset graphics processor data memory area.
9. The method for monitoring a graphics processor according to claim 1, wherein: Also includes: Determining whether the plurality of threads meet a preset acquisition completion condition based on the running data; If the plurality of threads all meet the preset acquisition completion condition, a read completion signal is generated; The read completion signal is pushed to the baseboard management controller to control the baseboard management controller to copy the operation data to any global variable.
10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for monitoring a graphics processor according to any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Graphic processor board card
CN109408445A
Onboard graphics processor control method and device
CN110399328A
Gene data analysis method and heterogeneous scheduling platform
CN110427262A
Server memory recovery method, device and equipment and readable storage medium
CN113672390A
Method, device and equipment for acquiring information of graphics processor, and storage medium
CN115599561A