Management method for graphics processing unit, and server
By instructing the CPU to set the GPU's persistent mode through an out-of-band controller, the problems of low efficiency and low success rate in existing technologies are solved, achieving efficient and reliable GPU management.
Patent Information
- Application Number
- PCT/CN2025/081643
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-30
- Filing Date
- 2025-03-10
- Publication Date
- 2025-12-04
AI Technical Summary
In existing technologies, setting the persistent mode state of a graphics processing unit (GPU) is inefficient and has a low success rate, mainly because it requires users to manually enter code commands, which is prone to errors and requires a high level of expertise.
The update request is obtained through an out-of-band controller (such as BMC), which instructs the processor (such as CPU) to perform the operation of setting the GPU's persistent mode to the target state, thus avoiding manual intervention and directly controlling the GPU's persistent mode to be on or off.
It improves the efficiency and success rate of setting GPU persistent mode, reduces errors caused by user misoperation, and enhances the user experience.
Smart Images

Figure CN2025081643_04122025_PF_FP_ABST
Abstract
Description
A method for managing a graphics processor and a server
[0001] This application claims priority to Chinese patent application No. 202410694545.0, filed on May 30, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of servers, and more particularly to a method for managing a graphics processor and a server. Background Technology
[0003] As is well known, when a graphics processing unit (GPU) server is under low or no load, the GPUs in the GPU server will enter a sleep mode and will not process any business on the GPU server. To prevent the GPU from entering sleep mode, users need to manually enter commands in the GPU server's operating system (OS) to enable persistent mode for the GPUs; this ensures that the GPUs will not enter sleep mode (i.e., remain in running mode) even under low or no load.
[0004] However, since the above method sets the state of the GPU's persistent mode by having the user manually enter code, it results in low efficiency in setting the state of the GPU's persistent mode. Summary of the Invention
[0005] This application provides a method for managing a graphics processor and a server to improve the efficiency of setting the persistent mode state of the GPU.
[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0007] In a first aspect, embodiments of this application provide a method for managing a graphics processor, the method comprising: an out-of-band controller acquiring an update request for a target graphics processor GPU in a computing device; the update request being used to instruct the persistent mode of the target GPU to be set to a target state, the target state including: an on state or an off state; wherein, when the persistent mode is set to the on state, the GPU is always in a running state; based on the above update request, the out-of-band controller instructs the processor CPU to execute the operation indicated by the update request.
[0008] The above embodiment obtains an update request through the out-of-band controller (BMC) to instruct the persistent mode of the target GPU in the computing device to be set to the target state. Then, based on the update request, the out-of-band controller instructs the processor (CPU) to perform the operation indicated by the update request, thereby setting the persistent mode of the target GPU to the target state. Compared with the prior art, it does not require the user to manually enter code commands on the CPU, thus improving the efficiency of setting the persistent mode of the GPU.
[0009] Furthermore, since the above embodiments control the CPU to execute the operation indicated by the update request by instructing the CPU through an out-of-band controller, and since this process does not involve human intervention, the problem of low success rate in setting the GPU's persistent mode due to user error is solved.
[0010] In one possible implementation, the out-of-band controller acquires an update request for a target graphics processing unit (GPU) in a computing device; this includes: the out-of-band controller generating display information for displaying a management interface of the out-of-band controller, the display information including identifiers of multiple GPUs in the computing device; the out-of-band controller receiving operation information for the management interface, the operation information including the identifier of the target GPU among the multiple GPUs that needs to have its persistent mode set to the target state; and the out-of-band controller generating the update request based on the operation information.
[0011] In the above embodiment, the BMC generates display information for displaying the BMC's management interface. This display information includes identifiers of multiple GPUs in the computing device. Then, the BMC receives operation information for the management interface, which includes the identifier of the target GPU among the multiple GPUs that needs to have its persistent mode set to the target state. Based on this operation information, the BMC generates an update request instructing the target GPU to have its persistent mode set to the target state and sends the update request to the CPU in the computing device, causing the CPU to execute the operation indicated by the update request. As can be seen, the above embodiment does not require the user to manually input code commands, thus improving the efficiency of setting the persistent mode of the GPU.
[0012] In one possible implementation, the out-of-band controller acquires an update request for a target graphics processing unit (GPU) in a computing device; this includes: the out-of-band controller receiving an update request sent by a management device; the management device is used to manage multiple computing devices, including the aforementioned computing device.
[0013] The above embodiment sends an update request to the BMC of the computing device, instructing it to set the persistent mode of the target GPU in the computing device to the target state. Based on this update request, the BMC sends an instruction to the CPU, instructing it to set the persistent mode of the target GPU to the target state. This causes the CPU to execute the operation indicated by the instruction on the target GPU, thereby setting the persistent mode of the target GPU to the target state. Since the management device manages multiple computing devices, it can simultaneously set the persistent mode of target GPUs in multiple computing devices without requiring the user to manually input commands into the CPU of each computing device. Therefore, the efficiency of setting the persistent mode of GPUs in multiple computing devices is improved.
[0014] In one possible implementation, where the update request is used to instruct the persistent mode of the target GPU to be enabled each time the operating system (OS) of the computing device starts, the method further includes, after the out-of-band controller obtains the update request for the target graphics processor (GPU) in the computing device, the out-of-band controller storing the update operation indicated by the update request; the update operation is the operation of setting the persistent mode of the target GPU to be enabled.
[0015] In one possible implementation, the method further includes: an out-of-band controller receiving an acquisition request sent by the CPU after the OS starts; the out-of-band controller responding to the acquisition request sending the update operation to the CPU so that the CPU sets the persistent mode of the target GPU to the enabled state.
[0016] In the above embodiments, after the BMC obtains an update request instructing the OS of the computing device to set the persistent mode of the target GPU to an enabled state each time it starts, the BMC stores the update operation indicated by the update request and sends the update request to the CPU in the computing device, so that the CPU sets the persistent mode of the target GPU to an enabled state. Subsequently, after each OS startup, the CPU sends an acquisition request to the BMC, and the BMC, in response to the acquisition request, sends the stored update operation to the CPU, so that the CPU executes the update operation, which is the operation of setting the persistent mode of the target GPU to an enabled state. It can be seen that the above embodiments, by executing the above update operation after each OS restart, set the persistent mode of the target GPU to an enabled state, thus realizing the function of setting the default state of the persistent mode of the target GPU to an enabled state, without requiring the user to manually enter code commands after each OS restart. Therefore, the efficiency of setting the persistent mode state of the target GPU is improved.
[0017] In one possible implementation, the method further includes: an out-of-band controller receiving the execution result of an operation indicated by an execution update request sent by the CPU, the execution result including success and failure.
[0018] In the above embodiment, after the CPU completes the operation indicated by the update request (hereinafter referred to as the update operation), it sends the execution result of the update operation to the BMC so that the user can know in a timely manner whether the update operation has been successfully executed through the BMC, thereby improving the user experience.
[0019] In one possible implementation, the method further includes: an out-of-band controller acquiring a query request for a target GPU; the query request instructing the query to retrieve attribute information of the target GPU, the attribute information including at least: the current state of the persistent mode of the target GPU, the persistent mode state including: an on state and a off state; the out-of-band controller sending the query request to the CPU to cause the CPU to execute the operation indicated by the query request.
[0020] The above embodiments send a query request to the CPU of the computing device through the BMC of the computing device, so that the CPU queries the current state of the persistent mode of the target GPU in the computing device and sends the query result to the BMC. Compared with the method of entering a query command in the OS of the computing device, the above embodiments improve the efficiency of querying the current state of the GPU's persistent mode.
[0021] Secondly, embodiments of this application provide another method for managing a graphics processor, the method comprising: the processor receiving an instruction message sent by an out-of-band controller of a computing device, the instruction message being used to instruct the persistent mode of a target GPU in the computing device to be set to a target state, the target state including: an on state or an off state; the GPU being in a running mode when the persistent mode is on; and the processor, in response to the instruction message, setting the persistent mode of the target GPU to the target state.
[0022] In the above embodiments, the processor responds to the instruction sent by the out-of-band controller to set the persistent mode of the target GPU in the computing device to the target state, and sets the persistent mode of the target GPU to the target state. Compared with the prior art, it does not require the user to manually enter code commands on the CPU, thus improving the efficiency of setting the persistent mode of the GPU.
[0023] In one possible implementation, before setting the persistent mode of the target GPU to the target state, the method further includes: the processor obtaining the on-state of the target GPU based on the identifier of the target GPU included in the above-mentioned indication information, wherein the on-state of the GPU includes: a startup state and a shutdown state; the above-mentioned setting the persistent mode of the target GPU to the target state includes: when the target GPU is in the startup state, the processor sets the persistent mode of the target GPU to the target state.
[0024] In the above embodiment, after receiving an update request indicating that the persistent mode of the target GPU be set to the target state, the CPU obtains the on state of the target GPU. If the on state of the target GPU is the startup state, the CPU executes the operation indicated by the update request. This avoids the problem that the CPU cannot operate on the target GPU when the on state of the target GPU is the shutdown state, which would lead to the failure of setting the persistent mode of the target GPU. Therefore, the success rate of setting the persistent mode of the GPU is improved.
[0025] Thirdly, embodiments of this application provide an out-of-band controller, which includes a transceiver unit and a processing unit. The transceiver unit is used to acquire an update request for a target graphics processing unit (GPU) in a computing device. The update request is used to instruct the persistent mode of the target GPU to be set to a target state, which includes an on state or a off state. When the persistent mode is set to the on state, the GPU is always in a running state. The processing unit is used to instruct the processor CPU to execute the operation indicated by the update request based on the update request.
[0026] In one possible implementation, the processing unit is used to generate display information for displaying the management interface of the out-of-band controller, the display information including the identifiers of multiple GPUs in the computing device; the transceiver unit is used to receive operation information for the management interface, the operation information including the identifier of the target GPU among the multiple GPUs that needs to have its persistent mode set to the target state; the processing unit is also used to generate an update request based on the operation information.
[0027] In one possible implementation, the transceiver unit is also used to receive update requests sent by the management device; the management device is used to manage multiple computing devices, including computing devices.
[0028] In one possible implementation, the out-of-band controller further includes: a storage unit; the storage unit is used to store the update operation indicated by the update request; the update operation is the operation of setting the persistent mode of the target GPU to an enabled state.
[0029] In one possible implementation, the transceiver unit is used to receive an acquisition request sent by the CPU after the OS starts; the transceiver unit is also used to send an update operation to the CPU in response to the acquisition request, so that the CPU sets the persistent mode of the target GPU to the enabled state.
[0030] In one possible implementation, the transceiver unit is used to receive the execution result of the operation indicated by the execution update request sent by the CPU, and the execution result includes success and failure.
[0031] In one possible implementation, the transceiver unit is used to obtain a query request for the target GPU; the query request is used to instruct the querying of attribute information of the target GPU, the attribute information including at least: the current state of the persistent mode of the target GPU, the state of the persistent mode including: an on state and a off state; the transceiver unit is also used to send the query request to the CPU so that the CPU executes the operation indicated by the query request.
[0032] Fourthly, embodiments of this application provide a processor, which includes: a transceiver unit and a setting unit; the transceiver unit is used to receive indication information sent by an out-of-band controller of a computing device, the indication information being used to indicate that the persistent mode of a target GPU in the computing device is set to a target state, the target state including: an on state or an off state; when the GPU is in the persistent mode on state, the GPU is always in running mode; the setting unit is used to set the persistent mode of the target GPU to the target state in response to the indication information.
[0033] In one possible implementation, the transceiver unit is used to obtain the power-on state of the target GPU based on the identifier of the target GPU included in the indication information, wherein the power-on state of the GPU includes: a startup state and a shutdown state; the setting unit is used to set the persistent mode of the target GPU to the target state when the target GPU is in the startup state.
[0034] Fifthly, embodiments of this application provide a server, which includes a memory, an out-of-band controller, and a processor. The memory is coupled to both the processor and the out-of-band controller. The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the out-of-band controller, a method is provided to implement the first aspect and its possible implementations. When the computer instructions are executed by the processor, a method is provided to implement the second aspect and its possible implementations.
[0035] In a sixth aspect, a computing system is provided, comprising: a management device and a server as described in the fifth aspect; the management device is connected to the server via a network; the management device is configured to send an update request to the server, the update request being configured to instruct the persistent mode of a target GPU in the server to be set to a target state, the target state including: an on state or an off state.
[0036] In a seventh aspect, a computer-readable storage medium is provided, which stores computer instructions that, when executed by a computing device, cause the computing device to perform the functions of the means of the first aspect and the second aspect and their possible implementations.
[0037] Eighthly, a computer program product containing instructions is provided, which, when run on a computing device, causes the computing device to perform the functions of the means of the first and second aspects and their possible implementations described above.
[0038] It should be understood that the beneficial effects achieved by the third to eighth aspects of the technical solutions and their corresponding possible implementations in the embodiments of this application can be referred to the above-described technical effects of the first and second aspects and their corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0039] Figure 1 is a schematic diagram of a GPU management system provided in an embodiment of this application;
[0040] Figure 2 is a schematic diagram of another GPU management system provided in an embodiment of this application;
[0041] Figure 3 is a schematic diagram of the hardware structure of a computing device provided in an embodiment of this application;
[0042] Figure 4 is a schematic flowchart of a graphics processor management method according to an embodiment of this application;
[0043] Figure 5 is a schematic diagram of a graphics processor management method according to an embodiment of this application;
[0044] Figure 6 is a schematic flowchart of a graphics processor management method according to an embodiment of this application;
[0045] Figure 7 is a schematic diagram of a GPU management interface provided in an embodiment of this application;
[0046] Figure 8 is a schematic diagram of a GPU management interface provided in an embodiment of this application;
[0047] Figure 9 is a schematic diagram of a GPU management interface provided in an embodiment of this application;
[0048] Figure 10 is a schematic flowchart of a graphics processor management method according to an embodiment of this application;
[0049] Figure 11 is a schematic diagram of a device management interface provided in an embodiment of this application;
[0050] Figure 12 is a schematic flowchart of a graphics processor management method according to an embodiment of this application;
[0051] Figure 13 is a schematic diagram of an out-of-band controller provided in an embodiment of this application;
[0052] Figure 14 is a schematic diagram of the structure of a processor provided in an embodiment of this application. Detailed Implementation
[0053] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0054] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first update request" and "second update request," etc., are used to distinguish different update requests, not to describe a specific order of update requests.
[0055] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0056] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple GPUs means two or more GPUs.
[0057] First, the management method for a graphics processor provided in the embodiments of this application, as well as some concepts involved in the server, will be explained in detail below:
[0058] GPU: Also known as a graphics processing unit, visual processor, or display chip, it is a microprocessor specifically designed to perform image and graphics-related calculations in personal computers, workstations, game consoles, and servers.
[0059] In common GPU management methods, users first log into the GPU server's operating system (OS) and then manually enter commands within that OS to set the persistent mode status (i.e., enabled or disabled) of the GPUs on the GPU server. When a GPU's persistent mode is enabled, the GPU will not enter sleep mode under low or no load, and it will not crash under prolonged high load operation; instead, it will remain in running mode.
[0060] Specifically, when the GPU in the GPU server is powered on, a code command is manually entered into the OS of the GPU server (e.g., to set the persistent mode of GPU_1 in the GPU server to the enabled state). After receiving the code command, the OS calls the management driver of GPU_1 in the OS to control GPU_1 to set its persistent mode to the enabled state.
[0061] As can be seen, the above GPU management method requires the user to manually input a command in the GPU server's OS to set the persistent mode status of a specific GPU (referred to as the "status setting command"). Since the OS runs on the central processing unit (CPU) of the computing device, the above GPU management method requires the user to manually input the status setting command on the CPU. When it is necessary to set the persistent mode status of multiple GPUs, the user needs to manually input multiple status setting commands on the CPU. It is evident that the efficiency of the user manually inputting the above status setting commands on the CPU is low, which in turn leads to low efficiency in setting the persistent mode status of GPUs.
[0062] Furthermore, the success rate of setting the state of the GPU in persistent mode is low because the error rate is high when manually entering state setting commands in the CPU and the user's expertise is required.
[0063] In view of this, embodiments of this application provide a GPU management method. This method controls the processor (e.g., CPU) to set the persistent mode of the target GPU to a target state, including an on or off state, by instructing the processor (e.g., a BMC) through an out-of-band controller (e.g., a BMC). It is evident that the above GPU management method does not require the user to manually input code commands on the CPU, thus improving the efficiency of setting the persistent mode of the GPU.
[0064] In some embodiments, this application provides a GPU management method, which includes: an out-of-band controller acquiring an update request for a target GPU in a computing device; the update request instructing the target GPU to set a persistent mode to a target state including an on state or a off state, wherein when the persistent mode is on, the GPU is always in a running mode; and then, based on the update request, the out-of-band controller instructs the CPU to perform the operation indicated by the update request.
[0065] The above embodiment obtains an update request from an out-of-band controller, which instructs the CPU to perform the operation indicated by the update request, thereby setting the persistent mode of the target GPU to the target state. Compared with the prior art, it does not require the user to manually enter code commands on the CPU, thus improving the efficiency of setting the persistent mode of the GPU.
[0066] Furthermore, since the above embodiments control the CPU to execute the operation indicated by the update request by instructing the CPU through an out-of-band controller, and since this process does not involve human intervention, the problem of low success rate in setting the GPU's persistent mode due to user error is solved.
[0067] In one implementation, the GPU management method provided in this application embodiment is applied in the GPU management system shown in Figure 1. The management system is applied in a computing device and includes: BMC, CPU and multiple GPUs (i.e., GPU_1 to GPU_N); wherein, the BMC is physically connected to the CPU, and the CPU is physically connected or network connected to GPU_1 to GPU_N; as an example, the network can be a public network, the intranet of a data center, an enterprise intranet, or a school intranet, etc., and this application embodiment does not limit it.
[0068] The aforementioned BMC is used to generate display information for the management interface, which includes identifiers for GPU_1 to GPU_N; it is also used to obtain user operation information within the management interface, including the identifier of the target GPU among GPU_1 to GPU_N that needs to have its persistent mode set to the target state; then, the BMC instructs the CPU to set the persistent mode of the target GPU to the target state based on this operation information. The BMC is also used to receive the execution result from the CPU of the operation indicated by the update request.
[0069] The CPU is used to receive the instruction information sent by the BMC and execute the operation indicated by the instruction information, thereby setting the persistent mode of the target GPU to the target state; then, the CPU sends the execution result to the BMC.
[0070] The target GPUs in GPU_1 to GPU_N mentioned above are used to receive operations executed by the CPU and send the execution results of the operations to the CPU.
[0071] It should be understood that the above management system may include multiple GPUs or a single GPU, and the specific embodiments of this application do not limit the number of GPUs in the management system.
[0072] In another implementation, the GPU management method provided in this application embodiment is applied in the GPU management system shown in Figure 2. The method includes: management device 101, GPU server 102 and GPU server 103, wherein the management device 101 is connected to GPU server 102 and GPU server 103 through a network. As an example, the network can be a public network, the intranet where the data center is located, an enterprise intranet, or a school intranet, etc., and this application embodiment does not limit it.
[0073] Management device 101 is used to manage GPU servers 102 and 103. Specifically, management device 101 displays all managed GPU servers and the identifier of each GPU in each GPU server on the server management interface. When a user selects to enable persistent mode for GPU_A in GPU server 102 on the server management interface, management device 101 sends an update request to GPU server 102 based on the selection information. This update request instructs that persistent mode for GPU_A be set to enabled. Management device 101 is also used to receive the execution result sent by GPU server 102 of the operation indicated by the update request.
[0074] The GPU server 102 includes the BMC, CPU, GPU_A and GPU_B shown in Figure 1 above. Their connection relationship is consistent with the connection relationship in the GPU management system shown in Figure 1, and will not be described again here.
[0075] The BMC is used to receive an update request sent by the management device 101, which indicates that the persistent mode of GPU_A is set to the enabled state; and based on the update request, instruct the CPU to perform the operation indicated by the update request. The CPU responds to the instruction information sent by the BMC and controls GPU_A to set the persistent mode to the enabled state. The CPU is also used to send the execution result to the management device 101 through the BMC.
[0076] GPU server 103 is similar to GPU server 102 described above. For details, please refer to the relevant description of GPU server 102 above. It will not be repeated here.
[0077] It should be noted that the GPU management system described above may include one GPU server or multiple GPU servers. In specific embodiments of this application, the number of GPU servers is not limited.
[0078] For example, Figure 3 is a schematic diagram of the hardware structure of a computing device, such as the computing device shown in Figure 1 and the GPU server shown in Figure 2. Taking this computing device as an example, in terms of form, the server can be a high-density server, a rack server, or a full-rack server; in terms of performance, it can be a general-purpose server, a GPU (graphics processing unit) server, or an artificial intelligence (AI) server.
[0079] The hardware of this computing device includes a processor, an out-of-band controller, and memory. The software includes an out-of-band management module, processor firmware, and an operating system (OS) management unit.
[0080] The aforementioned out-of-band management module runs within the out-of-band controller, while the OS management unit runs on the processor (as shown in Figure 3). The processor firmware can also reside within the processor (as shown in Figure 3). The out-of-band management module can be a management unit for non-business modules. For example, the out-of-band management module can remotely maintain and manage the computing device through a dedicated data channel. This out-of-band management module is completely independent of the computing device's operating system and can communicate with the OS through the computing device's out-of-band management interface.
[0081] For example, an out-of-band management module may include a management unit for the operating mode of a computing device, a management system in a management chip outside the processor, a baseboard management controller (BMC), a system management mode (SMM), etc. It should be noted that the specific form of the out-of-band management module is not limited in the embodiments of this application; the above is merely illustrative. In the following embodiments, only a BMC is used as an example of an out-of-band management module for explanation.
[0082] It should be noted that, in the case where the computing device is a GPU server, the computing device also includes at least one GPU (not shown in Figure 3), which is used for image processing or training of data models.
[0083] The memory, also known as internal memory or main memory, is installed in memory slots on the motherboard of a computing device. The memory communicates with the memory controller through memory channels.
[0084] It should be noted that the system architecture and application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0085] It should be noted that, for ease of description, this application uses the out-of-band controller (BMC) and the processor (CPU) as an example for illustration, and will not be repeated hereafter.
[0086] This application provides a GPU management method, which is applied to the computing device shown in FIG3; as shown in FIG4, the method includes S110-S130.
[0087] S110, The BMC of the computing device obtains an update request for the target GPU in the computing device.
[0088] The aforementioned computing device includes at least one GPU, and the at least one GPU includes the aforementioned target GPU, wherein the target GPU is the GPU whose persistent mode needs to be set to a target state, the target state including: an on state or a off state.
[0089] It should be noted that when a GPU's persistent mode is disabled, it will enter sleep mode under low or no load, and will no longer process tasks. When a GPU's persistent mode is enabled, it will not enter sleep mode under low or no load, but will remain in running mode.
[0090] It should be understood that the target GPU may be one of the multiple GPUs in the computing device or one of the multiple GPUs in the computing device. In particular, the number of target GPUs is not limited in the embodiments of this application.
[0091] The above update request is used to instruct the persistent mode state of the target GPU to be set to the target state.
[0092] The above-mentioned S110 can be implemented in the following ways: the BMC receives the update request from the management device, which is used to manage the computing device; or the BMC obtains the update request triggered by the user on the BMC's management interface; or the BMC obtains the update request imported by the user or other devices from the BMC's configuration file. The specific implementation of the above-mentioned S110 is not limited in this application embodiment.
[0093] Specifically, when S110 is implemented by the BMC obtaining the update request triggered by the user on the BMC's management interface, its specific implementation is described in S310-S340 below, and will not be repeated here. When S110 is implemented by the BMC receiving the update request from the management device, its specific implementation is described in S410-S430 below, and will not be repeated here.
[0094] S120. Based on the update request, the BMC instructs the CPU to perform the operation indicated by the update request.
[0095] It should be understood that the aforementioned BMC and CPU are components in the same computing device; the BMC and CPU are physically connected.
[0096] The above-mentioned S120 is implemented by the BMC sending instruction information to the BMC agent module in the CPU through a physical connection or network connection with the CPU. The instruction information may be the above-mentioned update request or the operation indicated by the update request (hereinafter referred to as the update operation). The update operation is used to set the persistent mode of the target GPU to the target state. The specific embodiments of this application do not limit it.
[0097] The aforementioned BMC agent module can be either BMC agent software or smart provisioning (SP) software; however, this application does not specifically limit it in its embodiments.
[0098] S130, The CPU of the computing device executes the operation indicated by the update request.
[0099] The implementation of S130 includes: the CPU responding to the above indication information sets the persistent mode of the target GPU to the target state; wherein, when the indication information is the above update request, the CPU obtains the operation instruction of the update operation based on the operation indicated by the update request (hereinafter referred to as: update operation); then, the CPU sets the persistent mode of the target GPU to the target state by executing the operation instruction.
[0100] It should be understood that the CPU mentioned above includes a management driver module for the target GPU. Based on this, the specific implementation of S130 includes: the CPU calls the management driver module to execute the operation instructions for the above update operation, thereby controlling the target GPU to set its persistent mode to the target state.
[0101] The above embodiment obtains an update request through the out-of-band controller (BMC) to instruct the persistent mode of the target GPU in the computing device to be set to the target state. Then, based on the update request, the out-of-band controller instructs the processor (CPU) to perform the operation indicated by the update request, thereby setting the persistent mode of the target GPU to the target state. Compared with the prior art, it does not require the user to manually enter code commands on the CPU, thus improving the efficiency of setting the persistent mode of the GPU.
[0102] Furthermore, since the above embodiments control the CPU to execute the operation indicated by the update request by instructing the CPU through an out-of-band controller, and since this process does not involve human intervention, the problem of low success rate in setting the GPU's persistent mode due to user error is solved.
[0103] Based on the GPU management method shown in Figure 4, when it is necessary to set the default state of the persistent mode of the target GPU to the target state (e.g., the enabled state), this application embodiment provides a specific implementation method, as shown in Figure 5, which includes: S210-S270.
[0104] S210, The BMC of the computing device obtains an update request for the target GPU in the computing device.
[0105] The update request is intended to instruct the persistent mode of the target GPU to be enabled each time the OS of the aforementioned computing device starts.
[0106] It should be noted that the implementation of S210 is similar to that of S110. For a detailed description of S210, please refer to the relevant description of S110 above. It will not be repeated here.
[0107] S220, The update operation indicated by the BMC storage update request of the computing device.
[0108] The above update operation sets the persistent mode of the target GPU to the enabled state.
[0109] S230, the BMC of the computing device sends an update request to the CPU of the computing device.
[0110] It should be understood that the BMC can execute S220 first and then S230, or it can execute S230 first and then S220; or it can execute S220 and S230 simultaneously.
[0111] S240, The CPU of the computing device executes the update operation indicated by the update request.
[0112] It should be noted that the implementation methods of S230-S240 are similar to those of S120-S130. For a detailed description of S230-S240, please refer to the relevant descriptions of S120-S130 above. They will not be repeated here.
[0113] S250. After the OS of the computing device starts, the CPU of the computing device sends an acquisition request to the BMC of the computing device.
[0114] The above request is used to retrieve update operations stored in the BMC.
[0115] It should be understood that the above update request includes: the update operation to be performed (i.e., setting the persistent mode of the target GPU to the above-mentioned enabled state) and the execution conditions (i.e., performing the update operation after each boot of the OS of the above-mentioned computing device).
[0116] The above-mentioned S250 is implemented as follows: the CPU responds to the above-mentioned execution conditions and sends the above-mentioned acquisition request to the BMC.
[0117] S260. In response to the acquisition request, the BMC of the computing device sends an update operation to the CPU of the computing device.
[0118] S270, The CPU of the computing device performs an update operation.
[0119] The above update operation sets the persistent mode of the target GPU in the computing device to the enabled state.
[0120] The implementation of S270 is similar to that of S130. For a detailed description of S270, please refer to the relevant description of S130 above. It will not be repeated here.
[0121] It should be noted that the default state of GPU persistent mode is off; that is, if a GPU's persistent mode is on, when the computing device containing that GPU restarts, the GPU's persistent mode will switch back to the default state (i.e., off). In other words, when a computing device restarts, the persistent mode of the GPU in that computing device is off. Subsequently, after S270, the OS of the aforementioned computing device executes S250-S270 after each boot.
[0122] In the above embodiments, after the BMC obtains an update request instructing the OS of the computing device to set the persistent mode of the target GPU to an enabled state each time it starts, the BMC stores the update operation indicated by the update request and sends the update request to the CPU in the computing device, so that the CPU sets the persistent mode of the target GPU to an enabled state. Subsequently, after each OS startup, the CPU sends an acquisition request to the BMC, and the BMC, in response to the acquisition request, sends the stored update operation to the CPU, so that the CPU executes the update operation, which is the operation of setting the persistent mode of the target GPU to an enabled state. It can be seen that the above embodiments, by executing the above update operation after each OS restart, set the persistent mode of the target GPU to an enabled state, thus realizing the function of setting the default state of the persistent mode of the target GPU to an enabled state, without requiring the user to manually enter code commands after each OS restart. Therefore, the efficiency of setting the persistent mode state of the target GPU is improved.
[0123] When the GPU management method shown in Figure 4 is applied to the GPU management system shown in Figure 1, this application embodiment provides a specific implementation method, as shown in Figure 6, which includes: S310-S380.
[0124] S310, The BMC of the computing device generates display information for displaying the management interface of the BMC.
[0125] The information displayed above includes the identifiers of multiple GPUs in the aforementioned computing device; wherein, the management interface is used to display the identifiers of these multiple GPUs.
[0126] The specific implementation of S310 includes: the BMC obtaining the identifiers of multiple GPUs in the computing device, and then the BMC encapsulating the identifiers of the multiple GPUs into the display information.
[0127] For example, assuming that the computing device mentioned above includes GPU_1 to GPU_4, then the BMC management interface used to display the display information generated by the BMC is shown in Figure 7. The BMC management interface includes the identifiers of GPU_1 to GPU_4.
[0128] S320, the BMC of the computing device receives operation information for the management interface.
[0129] The above operation information includes the identifier of the target GPU among the multiple GPUs in the above computing device that needs to have its persistent mode set to the target state.
[0130] For example, as shown in Figure 8, after the user selects GPU_1 in the BMC management interface, they click the button to enable persistent mode. At this time, the BMC obtains the operation information of this operation from the BMC management interface, which includes the identifier of GPU_1.
[0131] S330, the BMC of the computing device generates an update request based on the operation information.
[0132] The implementation of S330 includes BMC encapsulating the above operation information in the form of a request to obtain the above update request. The update request is used to indicate that the persistent mode of the target GPU is set to the target state, which includes an on state or a off state.
[0133] S340, the BMC of the computing device sends an update request to the CPU of the computing device.
[0134] It should be noted that the implementation method of S340 is the same as that of S120. For a detailed description of S340, please refer to the relevant description of S120 above. It will not be repeated here.
[0135] S350: The CPU of the computing device obtains the startup status of the target GPU.
[0136] The aforementioned GPU's on / off states include: startup state and shutdown state. When the GPU is in the shutdown state, it does not receive operation commands, and its state is equivalent to being powered off. When the GPU is in the startup state, it can be in hibernation mode or running mode. In this state, the GPU can receive operation commands, and its state is equivalent to being powered on.
[0137] It should be understood that the aforementioned GPU's on state is not related to the GPU's persistent mode state. When the GPU's on state is in the startup state, the GPU's persistent mode can be either on or off. This application does not specifically limit this.
[0138] The above-mentioned S350 is implemented as follows: the CPU obtains the startup status of the target GPU based on the identifier of the target GPU in the update request; the specific implementation includes: the CPU obtains the startup status of the target GPU through the management driver module of the target GPU, and the specific implementation is described in the prior art, which will not be repeated here.
[0139] S360, the CPU of the computing device determines whether the target GPU is in the startup state.
[0140] If the target GPU is in a disabled state, the CPU performs a termination action, ending the method flow.
[0141] When the target GPU is in the startup state, the CPU executes the following S370.
[0142] S370, the CPU of the computing device sets the persistent mode of the target GPU to the target state.
[0143] It should be noted that the implementation of S370 is the same as that of S130. For a detailed description of S370, please refer to the relevant description of S130 above. It will not be repeated here.
[0144] In the above embodiment, after receiving an update request indicating that the persistent mode of the target GPU be set to the target state, the CPU obtains the on state of the target GPU. If the on state of the target GPU is the startup state, the CPU executes the operation indicated by the update request. This avoids the problem that the CPU cannot operate on the target GPU when the on state of the target GPU is the shutdown state, which would lead to the failure of setting the persistent mode of the target GPU. Therefore, the success rate of setting the persistent mode of the GPU is improved.
[0145] S380, The CPU of the computing device sends the execution result of the operation indicated by the execution update request to the BMC.
[0146] The above execution results are used to display in the BMC management interface, and the execution results include: success and failure.
[0147] It should be noted that the implementation of S380 is similar to that of S120. For a detailed description of S380, please refer to the relevant description of S120 above. It will not be repeated here.
[0148] For example, assuming that the CPU successfully executes the operation indicated by the update request, the CPU sends the execution result to the BMC. When the BMC receives the execution result, as shown in Figure 9, the BMC's management interface displays "GPU_1's persistent mode has been successfully enabled!!!" (i.e., the execution result).
[0149] In the above embodiment, the BMC generates display information for displaying the BMC's management interface. This display information includes identifiers of multiple GPUs in the computing device. Then, the BMC receives operation information for the management interface, which includes the identifier of the target GPU among the multiple GPUs that needs to have its persistent mode set to the target state. Based on this operation information, the BMC generates an update request instructing the target GPU to have its persistent mode set to the target state and sends the update request to the CPU in the computing device, causing the CPU to execute the operation indicated by the update request. As can be seen, the above embodiment does not require the user to manually input code commands, thus improving the efficiency of setting the persistent mode of the GPU.
[0150] Furthermore, in the above embodiments, after the CPU completes the operation indicated by the update request (hereinafter referred to as the update operation), it sends the execution result of the update operation to the BMC, so that the user can know in a timely manner whether the update operation has been successfully executed through the BMC, thereby improving the user experience.
[0151] When the GPU management method shown in Figure 4 is applied to the GPU management system shown in Figure 2, this application embodiment provides a specific implementation method, as shown in Figure 10, which includes: S410-S460.
[0152] S410, Manage devices to obtain user operation information in the device management interface.
[0153] The aforementioned management device is used to manage multiple computing devices; wherein, there is a connection between the management device and the managed device, which can be a network connection or a physical connection, and the specific embodiments of this application do not specifically limit it.
[0154] The aforementioned device management interface includes: the identifiers of multiple devices managed by the management device, and the identifiers of the GPUs in each of these multiple devices. For example, as shown in Figure 11, the device management interface includes device A and device B managed by the management device; device A includes GPU_1 to GPU_4; device B also includes GPU_1 to GPU_4.
[0155] The aforementioned operation information includes: the identifier of the target GPU whose persistent mode needs to be set to the target state, and the identifier of the target device where the target GPU is located. Specifically, this operation information instructs that the persistent mode of the target GPU on the target device be set to the target state, which can be either enabled or disabled.
[0156] For example, as shown in Figure 11, when a user selects GPU_1 and GPU_4 in device A in the device management interface and clicks the button to enable persistent mode, the operation information obtained by the managed device from the device management interface includes the identifiers of device A, GPU_1, and GPU_4.
[0157] It should be noted that the above-mentioned S410 can be implemented by the CPU in the management device obtaining the above-mentioned operation information, or by the BMC in the management device obtaining the above-mentioned operation information. The specific implementation of this application does not limit it.
[0158] S420: The management device generates an update request based on the operation information.
[0159] The above update request is used to instruct the target GPU to be set to the target state in persistent mode.
[0160] It should be noted that the implementation of S420 is similar to that of S330. For a detailed description of S420, please refer to the relevant description of S330 above. It will not be repeated here.
[0161] S430, the management device sends an update request to the target device's BMC.
[0162] The specific implementation of S430 above includes: the management device sending the update request to the BMC of the target device based on the MAC address of the BMC of the target device.
[0163] S440. Based on the update request, the BMC of the target device sends instruction information to the CPU of the target device.
[0164] The above instruction information is used to instruct the persistent mode state of the target GPU in the target device to be set to the target state.
[0165] S450: The CPU of the target device responds to the instruction information and performs the operation indicated by the instruction information on the target GPU.
[0166] It should be noted that the implementation methods of S440-S450 are similar to those of S120-S130. For a detailed description of S440-S450, please refer to the relevant descriptions of S120-S130 above. They will not be repeated here.
[0167] S460, The CPU of the target device sends the execution result of the operation indicated by the execution instruction information to the management device through the BMC of the target device.
[0168] The above execution results include: success and failure.
[0169] The above embodiment sends an update request to the BMC of the computing device, instructing it to set the persistent mode of the target GPU in the computing device to the target state. Based on this update request, the BMC sends an instruction to the CPU, instructing it to set the persistent mode of the target GPU to the target state. This causes the CPU to execute the operation indicated by the instruction on the target GPU, thereby setting the persistent mode of the target GPU to the target state. Since the management device manages multiple computing devices, it can simultaneously set the persistent mode of target GPUs in multiple computing devices without requiring the user to manually input commands into the CPU of each computing device. Therefore, the efficiency of setting the persistent mode of GPUs in multiple computing devices is improved.
[0170] Optionally, the GPU management method shown in Figure 4 above also includes a method for querying the current state of the GPU's persistent mode, as shown in Figure 12, which includes: S510-S540.
[0171] S510, the BMC of the computing device obtains a query request for the target GPU.
[0172] The query request is used to instruct the query for the attribute information of the target GPU. The attribute information of the target GPU includes at least the current state of the persistent mode of the target GPU, wherein the state of the persistent mode of the GPU includes: on and off.
[0173] It should be noted that the GPU attribute information mentioned above, in addition to the current state of the GPU in continuous mode, may also include attributes such as GPU temperature and GPU resource utilization.
[0174] The aforementioned query request can be a request sent by the management device to the BMC, or a request triggered by the user on the BMC's management interface. When the query request is sent by the management device to the BMC, the implementation of S510 is similar to that of S410-S430, as detailed in the descriptions of S410-S430, and will not be repeated here. When the query request is triggered by the user on the BMC's management interface, the implementation of S510 is similar to that of S310-S340, as detailed in the descriptions of S310-S340, and will not be repeated here.
[0175] S520, the BMC of the computing device sends a query request to the CPU of the computing device.
[0176] S530: The CPU of the computing device responds to the query request and executes the operation indicated by the query request.
[0177] It should be noted that the implementation methods of S520-S530 are similar to those of S120-S130. For a detailed description of S520-S530, please refer to the relevant descriptions of S120-S130 above. They will not be repeated here.
[0178] S540, The CPU of the computing device sends the query result of the query operation to the BMC of the computing device.
[0179] The above query operation is the operation indicated by the above query request; that is, the query operation is to query the current state of the persistent mode of the target GPU.
[0180] The query results above represent the current state of the persistent mode of the target GPU, that is, the query results include: on and off states.
[0181] The above embodiments send a query request to the CPU of the computing device through the BMC of the computing device, so that the CPU queries the current state of the persistent mode of the target GPU in the computing device and sends the query result to the BMC. Compared with the method of entering a query command in the OS of the computing device, the above embodiments improve the efficiency of querying the current state of the GPU's persistent mode.
[0182] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0183] This application embodiment can divide the out-of-band controller into functional modules according to the above method example. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0184] With each functional module divided according to its corresponding function, Figure 13 shows a possible structural schematic diagram of the out-of-band controller involved in the above embodiments. As shown in Figure 13, the out-of-band controller includes: a transceiver unit 1301 and a processing unit 1302.
[0185] The transceiver unit 1301 is used to obtain update requests for the target graphics processing unit (GPU) in the computing device; for example, to execute steps S110 or S210 in the above method embodiments.
[0186] The processing unit 1302 is used to instruct the processor CPU to perform the operation indicated by the update request based on the update request; for example, to perform step S120 in the above method embodiment.
[0187] Optionally, the processing unit 1302 is used to generate display information for displaying the management interface of the out-of-band controller; for example, performing step S310 in the above method embodiment.
[0188] The transceiver unit 1301 is used to receive operation information for the management interface; for example, to execute step S320 in the above method embodiment.
[0189] The processing unit 1302 is used to generate an update request based on the operation information; for example, by executing step S330 in the above method embodiment.
[0190] Optionally, the transceiver unit 1301 is used to receive update requests sent by the management device; for example, to perform step S430 in the above method embodiment.
[0191] Optionally, the out-of-band controller also includes a storage unit 1303.
[0192] Storage unit 1303 is used to store the update operation indicated by the update request; for example, to perform step S220 in the above method embodiment.
[0193] Optionally, the transceiver unit 1301 is used to receive an acquisition request sent by the CPU after the OS starts; for example, to execute step S250 in the above method embodiment.
[0194] The transceiver unit 1301 is also configured to send an update operation to the CPU of the computing device in response to an acquisition request; for example, to perform step S260 in the above method embodiment.
[0195] Optionally, the transceiver unit 1301 is used to receive the execution result of the operation indicated by the execution update request sent by the CPU; for example, to execute step S380 in the above method embodiment.
[0196] Optionally, the transceiver unit 1301 is used to obtain a query request for the target GPU; for example, to execute step S510 in the above method embodiment.
[0197] The transceiver unit 1301 is also used to send a query request to the CPU; for example, to execute step S520 in the above method embodiment.
[0198] Each unit of the above-mentioned out-of-band controller can also be used to perform other actions in the above method embodiments. All relevant content of each step involved in the above method embodiments can be referred to the functional description of the corresponding functional module, and will not be repeated here.
[0199] This application embodiment can divide the processor into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0200] Figure 14 shows a possible structural diagram of the processor involved in the above embodiments, where each functional module is divided according to its corresponding function. As shown in Figure 14, the processor includes a transceiver unit 1401 and a setting unit 1402.
[0201] The transceiver unit is used to receive instruction information sent by the out-of-band controller of the computing device; for example, to perform step S440 in the above method embodiment.
[0202] The setting unit 1402 is used to set the persistent mode of the target GPU to the target state in response to the indication information; for example, to perform step S450 in the above method embodiment.
[0203] Optionally, the transceiver unit 1401 is used to obtain the on-state of the target GPU based on the identifier of the target GPU included in the indication information; for example, by performing step S350 in the above method embodiment.
[0204] The setting unit 1402 is used to set the persistent mode of the target GPU to the target state when the target GPU is in the startup state; for example, to execute step S370 in the above method embodiment.
[0205] Each unit of the processor described above can also be used to perform other actions in the above method embodiments. All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0206] This application embodiment also provides a server, including a memory, an out-of-band controller, and a processor. The memory is coupled to the processor and the out-of-band controller. The memory is used to store computer program code, which includes computer instructions. When the computer instructions are executed by the out-of-band controller, they implement any of the methods executed by the out-of-band controller. When the computer instructions are executed by the processor, they implement any of the methods executed by the processor.
[0207] This application also provides a computing system, which includes: a management device and a server; the management device and the server are connected via a network; the server is used to execute any of the methods executed by the out-of-band controller and the processor described above; the management device is used to send an application update request to the server, the update request being used to instruct the persistent mode of a target GPU in the server to be set to a target state, the target state including: an on state or an off state.
[0208] This application also provides a computer-readable storage medium storing computer instructions that, when executed on a computing device, cause the computing device to perform any of the methods described above, executed by the out-of-band controller and the processor.
[0209] For explanations of the relevant content and descriptions of the beneficial effects in any of the computer-readable storage media provided above, please refer to the corresponding embodiments described above, which will not be repeated here.
[0210] This application provides a computer program product that, when run on a computer, causes the computer to execute any of the methods described above, executed by the out-of-band controller and the processor.
[0211] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).
[0212] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0213] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0214] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0215] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0216] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.
[0217] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for managing a graphics processor, characterized in that, The method includes: An out-of-band controller acquires an update request for a target graphics processing unit (GPU) in a computing device; the update request is used to instruct the persistent mode of the target GPU to be set to a target state, the target state including: an on state or an off state; wherein, when the persistent mode is set to the on state, the target GPU is always in a running state; Based on the update request, the out-of-band controller instructs the processor CPU to perform the operation indicated by the update request.
2. The method according to claim 1, characterized in that, The out-of-band controller acquires an update request for the target graphics processing unit (GPU) in the computing device; including: The out-of-band controller generates display information for displaying the management interface of the out-of-band controller, and the display information includes the identifiers of multiple GPUs in the computing device; The out-of-band controller receives operation information for the management interface, the operation information including the identifier of the target GPU among the plurality of GPUs that needs to have its persistent mode set to the target state; The out-of-band controller generates the update request based on the operation information.
3. The method according to claim 1, characterized in that, The out-of-band controller acquires an update request for the target graphics processing unit (GPU) in the computing device; including: The out-of-band controller receives the update request sent by the management device; the management device is used to manage multiple computing devices, including the computing device.
4. The method according to any one of claims 1-3, characterized in that, When the update request is used to instruct the persistent mode of the target GPU to be enabled each time the operating system (OS) of the computing device starts, the method further includes, after the out-of-band controller obtains the update request for the target graphics processing unit (GPU) in the computing device: The out-of-band controller stores the update operation indicated by the update request; the update operation is to set the persistent mode of the target GPU to an enabled state.
5. The method according to claim 4, characterized in that, The method further includes: The out-of-band controller receives the acquisition request sent by the CPU after the OS starts; In response to the acquisition request, the out-of-band controller sends the update operation to the CPU, so that the CPU sets the persistent mode of the target GPU to the enabled state.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: The out-of-band controller receives the execution result of the operation indicated by the update request sent by the CPU, and the execution result includes: success and failure.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: The out-of-band controller acquires a query request for the target GPU; the query request is used to instruct the query of the target GPU's attribute information, the attribute information including at least: the current state of the target GPU's persistent mode, the persistent mode state including: an on state and a off state; The out-of-band controller sends the query request to the CPU so that the CPU performs the operation indicated by the query request.
8. A method for managing a graphics processor, characterized in that, The method includes: The processor receives instruction information sent by the out-of-band controller of the computing device. The instruction information is used to instruct the persistent mode of the target GPU in the computing device to be set to a target state. The target state includes: an on state or an off state. When the persistent mode of the GPU is on, the GPU is always in running mode. In response to the indication information, the processor sets the persistent mode of the target GPU to the target state.
9. The method according to claim 8, characterized in that, Before setting the persistent mode of the target GPU to the target state, the method further includes: The processor obtains the power-on state of the target GPU based on the identifier of the target GPU included in the indication information, wherein the power-on state of the GPU includes: startup state and shutdown state; Setting the persistent mode of the target GPU to the target state includes: When the target GPU is in a powered-on state, the processor sets the persistent mode of the target GPU to the target state.
10. A server, characterized in that, The device includes a memory, an out-of-band controller, and a processor, wherein the memory is coupled to the processor and the out-of-band controller respectively; the memory is used to store computer program code, the computer program code including computer instructions; when the computer instructions are executed by the out-of-band controller, the out-of-band controller performs the method as described in any one of claims 1 to 7; when the computer instructions are executed by the processor, the processor performs the method as described in any one of claims 8 to 9.
Citation Information
Patent Citations
Working mode switching method and device of PCIE switching chip
CN110780932A
Server power consumption intelligent optimization method and system
CN111857319A
System and method for supporting runtime SMM update and telemetry for bare computer deployment
CN115955394A
Graphic processor monitoring method and device, medium and server monitoring system
CN117931581A
Graphic processor management method and server
CN118691456A