Server, server system, virtual machine creation method and cloud management platform
By introducing front-end and back-end driver interception modules into the server, applications in the virtual machine can flexibly access part of the GPU's computing power, solving the problem of poor flexibility in accessing GPUs in the existing technology, realizing virtualization segmentation of GPUs and multiple scheduling strategies, and improving access flexibility.
Patent Information
- Application Number
- CN202410559048.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-04-30
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, virtual machines access GPUs have poor flexibility and rely on the hardware information and implementation details of the GPU.
By introducing front-end driver interception modules and back-end driver interception modules into the server, applications in the virtual machine can access some of the GPU's computing power through these modules, reducing their dependence on GPU hardware information.
It improves the flexibility of GPU access by applications in virtual machines, realizes virtualization segmentation of GPUs, supports multiple scheduling strategies, and reduces dependence on GPU hardware information.
Smart Images

Figure CN120216092A_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202311819308.4 and the invention title "Cloud Service Provision Method and System" filed on December 25, 2023, the entire content of which is incorporated herein by reference. Technical Field
[0002] This application relates to the technical field of cloud services, and particularly to a server, a server system, a virtual machine creation method, and a cloud management platform based on cloud computing technology. Background Art
[0003] Virtualization technology can abstract and transform various physical resources of a host, such as computing resources, network resources, and storage resources, etc., and present them, so as to break the non-separable barriers between the physical structures of the host, enabling tenants to apply these resources in a better way than the original configuration. In this way, virtualization technology can split the server GPU for use by multiple virtual machines.
[0004] Generally, when a virtual machine has a need to use the GPU, it needs to access the driver file of the GPU. After the driver file system of the server detects this access operation, it will forward the access request of the virtual machine to the GPU driver module. After receiving the access request, the GPU driver module will access the GPU based on this access request and provide the processing result to the virtual machine.
[0005] However, the flexibility of this access method is poor. Summary of the Invention
[0006] This application provides a server, a server system, a virtual machine creation method, and a cloud management platform based on cloud computing technology. This application reduces the degree of dependence on implementation details such as the hardware information of the GPU in the process of invoking the GPU to process access requests. The technical solutions provided by this application are as follows:
[0007] In a first aspect, the present application provides a server based on cloud computing technology. A first virtual machine and a second virtual machine are running on the server. The server is further provided with a backend driver interception module, a graphics processing unit (GPU) driver module, and a GPU. Among them, the first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on a first part of the computing power of the GPU to generate a first virtual graphics processing unit (VGPU) accessible to the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU. The second virtual machine is provided with a second front-end driver interception module and a second application. The second front-end driver interception module is used to perform device simulation on a second part of the computing power of the GPU to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is further used to obtain a second access request of the second application for the second VGPU. The backend driver interception module is used to obtain the first access request from the first front-end driver interception module and / or the second access request from the second front-end driver interception module, and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions, and / or send the second access request to the GPU driver module when it is determined that the second access request meets the preset conditions. The GPU driver module is used to call the first part of the computing power of the GPU to process the first access request, and / or call the second part of the computing power of the GPU to process the second access request.
[0008] In the present application, since the front-end driver interception module can obtain the access request of the application in the virtual machine for the VGPU and provide the access request to the GPU driver module through the backend driver interception module, the GPU driver module calls a part of the computing power of the GPU to process the access request. In this way, the operation of the application in the virtual machine to access the GPU is realized in a software manner, reducing the degree of dependence on the implementation details such as the hardware information of the GPU in the process of calling the GPU to process the access request, and improving the flexibility of the application in the virtual machine to access the GPU.
[0009] In the present application, there are multiple implementation manners of the server. The following takes three implementation manners as examples to illustrate it.
[0010] In the first implementation manner, the first application runs directly in the first virtual machine, and the second application runs directly in the second virtual machine.
[0011] In the second implementation, a first container is set up in the first virtual machine, and a first application runs in the first container. The first container is used to mount a first VGPU for the first application to access the first VGPU in the first container. A second container is set up in the second virtual machine, and a second application runs in the second container. The second container is used to mount a second VGPU for the second application to access the second VGPU in the second container.
[0012] Optionally, the number of containers set up in the virtual machine can be adjusted according to application requirements. For example, one container can be set up in the virtual machine, or multiple containers can be set up in the virtual machine. Exemplarily, a third container is further set up in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU for the third application to access the third VGPU in the third container, where the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the GPU to generate a third VGPU accessible to the third application. And / or, a fourth container is further set up in the second virtual machine, and a fourth application runs in the fourth container. The fourth container is used to mount a fourth VGPU for the fourth application to access the fourth VGPU in the fourth container, where the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the GPU to generate a fourth VGPU accessible to the fourth application.
[0013] After the GPU driver module calls the computing power of the GPU to process an access request, if it receives the processing result generated by the GPU processing the access request, the GPU driver module is further used to provide the processing result to the back-end driver interception module, so as to return the processing result to the application that triggered the access request through the back-end driver interception module. Then, in a possible implementation, the GPU driver module is further used to obtain the first processing result generated by the first part of the computing and processing power of the GPU processing the first access request, and / or obtain the second processing result generated by the second part of the computing and processing power of the GPU processing the second access request, and send the first processing result and / or the second processing result to the back-end driver interception module; the back-end driver interception module is further used to obtain the first processing result and / or the second processing result from the GPU driver module, send the first processing result to the first front-end driver interception module, and / or send the second processing result to the second front-end driver interception module; the first front-end driver interception module is further used to obtain the first processing result from the back-end driver interception module and provide the first processing result to the first application; the second front-end driver interception module is further used to obtain the second processing result from the back-end driver interception module and provide the second processing result to the second application.
[0014] It can be seen from this that the access requests triggered by the application program in the virtual machine to access the VGPU can be sequentially transmitted to the GPU driver module in the server via the front-end driver interception module in the virtual machine and the back-end driver interception module in the server. The processing results generated by the GPU in processing the access requests can be sequentially transmitted to the application program in the virtual machine via the GPU driver module in the server, the back-end driver interception module in the server, and the front-end driver interception module in the virtual machine. Since both of these transmission processes are implemented in software, the dependence on implementation details such as the hardware information of the GPU in the process of calling the GPU to process the access requests can be reduced, and even the dependence on implementation details such as the hardware information of the GPU is not required. For example, without depending on the hardware information of the GPU manufacturer, decoupling from the GPU can be achieved, and the flexibility of the application program in the virtual machine to access the GPU is improved. Moreover, when the virtual machine uses part of the computing power of the GPU, it is equivalent to virtualizing and partitioning the GPU. The software implementation method for transmitting access requests and processing results in this application enables the partitioning specifications for virtualizing and partitioning the GPU to break through the hardware control of the GPU, supports multiple scheduling policies for scheduling the GPU, and improves the flexibility of scheduling the GPU.
[0015] In a possible implementation manner, the server further includes: a virtual machine manager, configured to obtain a first access request from the first front-end driver interception module and send the first access request to the back-end driver interception module, obtain a second access request from the second front-end driver interception module, and send the second access request to the back-end driver interception module.
[0016] When the back-end driver interception module determines that the access request meets the preset conditions, it forwards the access request to the GPU driver module. When the back-end driver interception module determines that the access request does not meet the preset conditions, it does not forward the access request to the GPU driver module. At this time, the back-end driver interception module can return an error message to the virtual machine to which the access request belongs to indicate that the virtual machine cannot access the GPU. In a possible implementation manner, the back-end driver interception module determining that the first access request meets the preset conditions includes: the back-end driver interception module determining that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or, the back-end driver interception module determining that the computing power of the GPU required by the first access request is not greater than a first partial computing power or a first computing power threshold; the back-end driver interception module determining that the second access request meets the preset conditions includes: the back-end driver interception module determining that the video memory of the GPU required by the second access request is not greater than a second video memory threshold; and / or, the back-end driver interception module determining that the computing power of the GPU required by the second access request is not greater than a second partial computing power or a second computing power threshold.
[0017] Similarly, when there are also containers running in the virtual machine, after the front-end driver interception module set in the virtual machine receives an access request, it may first determine whether the access request meets the preset conditions. If the access request meets the preset conditions, the access request is then provided to the back-end driver interception module. If the access request does not meet the preset conditions, the virtual machine may return an error message to the container to indicate that the container cannot access the GPU. In a possible implementation, the virtual machine may, on a per-container basis, determine whether the video memory of the GPU required by the access request triggered by the application running in the container is greater than the video memory threshold; and / or, whether the computing power of the GPU required by the access request is greater than a partial computing power or computing power threshold of the GPU available for the application to access the VGPU. Here, the video memory threshold is the amount of video memory configured by the tenant for the container, and the computing power threshold is the amount of computing power configured by the tenant for the container.
[0018] In a second aspect, the present application provides a server system based on cloud computing technology. The server system includes a first server and a second server. The first server is provided with a first back-end driver interception module, a first GPU driver module, and a first GPU. A first virtual machine and a second virtual machine are running on the second server, and a connection channel is established between the first server and the second server.
[0019] Among them, the first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on a first partial computing power of the first GPU to generate a first VGPU accessible by the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU and send the first access request to the connection channel; the second virtual machine is provided with a second front-end driver interception module and a second application. The second front-end driver interception module is used to perform device simulation on a second partial computing power of the first GPU to generate a second VGPU accessible by the second application. The second application is used to access the second VGPU. The second front-end driver interception module is further used to obtain a second access request of the second application for the second VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module and / or the second access request from the second front-end driver interception module from the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions, and / or, send the second access request to the first GPU driver module when it is determined that the second access request meets the preset conditions; the first GPU driver module is used to call the first partial computing power of the first GPU to process the first access request, and / or, call the second partial computing power of the first GPU to process the second access request.
[0020] In a possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further configured to perform device simulation on the first part of the computing power of the second GPU to generate a third VGPU accessible to the first application. The first application is used to access the third VGPU, and the first front-end driver interception module is further configured to obtain a third access request of the first application for the third VGPU. The second front-end driver interception module of the second virtual machine is further configured to perform device simulation on the second part of the computing power of the second GPU to generate a fourth VGPU accessible to the second application. The second application is used to access the fourth VGPU, and the second front-end driver interception module is further configured to obtain a fourth access request of the second application for the fourth VGPU. The second back-end driver interception module is configured to obtain the third access request from the first front-end driver interception module and / or the fourth access request from the second front-end driver interception module, and send the third access request to the second GPU driver module when it is determined that the third access request meets the preset conditions, and / or send the fourth access request to the second GPU driver module when it is determined that the fourth access request meets the preset conditions. The second GPU driver module is configured to call the first part of the computing power of the second GPU to process the third access request, and / or call the second part of the computing power of the second GPU to process the fourth access request.
[0021] In a possible implementation, a first container is provided in the first virtual machine, and the first application runs in the first container. The first container is used to mount the first VGPU for the first application to access the first VGPU in the first container. A second container is provided in the second virtual machine, and the second application runs in the second container. The second container is used to mount the second VGPU for the second application to access the second VGPU in the second container.
[0022] In a possible implementation, a third container is further provided in the first virtual machine, and the third application runs in the third container. The third container is used to mount the third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is configured to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application.
[0023] And / or, a fourth container is further provided in the second virtual machine, and the fourth application runs in the fourth container. The fourth container is used to mount the fourth VGPU for the fourth application to access the fourth VGPU in the fourth container. Among them, the second front-end driver interception module is configured to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application.
[0024] In a possible implementation, the first application runs directly in the first virtual machine, and the second application runs directly in the second virtual machine.
[0025] In a possible implementation, the first GPU driver module is further configured to obtain a first processing result generated by processing a first access request using a first part of the computing processing power of the first GPU, and / or obtain a second processing result generated by processing a second access request using a second part of the computing processing power of the first GPU, and send the first processing result and / or the second processing result to the first back-end driver interception module; the first back-end driver interception module is further configured to obtain the first processing result and / or the second processing result from the first GPU driver module, send the first processing result to the connection channel, and / or send the second processing result to the connection channel; the first front-end driver interception module is further configured to obtain the first processing result from the connection channel from the first back-end driver interception module and provide the first processing result to the first application; the second front-end driver interception module is further configured to obtain the second processing result from the connection channel from the first back-end driver interception module and provide the second processing result to the second application.
[0026] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain a first access request from the first front-end driver interception module and send the first access request to the connection channel, obtain a second access request from the second front-end driver interception module, and send the second access request to the connection channel.
[0027] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, and the high-speed interconnection protocol includes Compute Express Link (CXL) or Lingqu Universal Bus (UB) protocol.
[0028] In a possible implementation, the first back-end driver interception module determines that the first access request meets a preset condition, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than a first part of the computing power or a first computing power threshold.
[0029] The second back-end driver interception module determines that the second access request meets a preset condition, including: the second back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than a second video memory threshold; and / or the second back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than a second part of the computing power or a second computing power threshold.
[0030] In a third aspect, the present application provides a method for creating a virtual machine based on cloud computing technology. The method is applied to a cloud management platform. The cloud management platform is used to manage the infrastructure. The infrastructure includes multiple servers. The method includes: obtaining a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of the first VGPU of the first virtual machine to be created; selecting a target server that can provide the specification of the first VGPU from the multiple servers; and creating the first virtual machine on the target server.
[0031] Among them, the target server is provided with a backend driver interception module, a GPU driver module, and a GPU. The first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on the first part of the computing power of the GPU to generate a first VGPU that can be accessed by the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU; the backend driver interception module is used to obtain the first access request from the first front-end driver interception module and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions; the GPU driver module is used to call the first part of the computing power of the GPU to process the first access request.
[0032] In a possible implementation, a first container is set in the first virtual machine. The first application runs in the first container. The first container is used to mount the first VGPU for the first application to access the first VGPU in the first container.
[0033] In a possible implementation, a third container is also set in the first virtual machine. The third application runs in the third container. The third container is used to mount the third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the GPU to generate a third VGPU that can be accessed by the third application.
[0034] In a possible implementation, the first application runs directly in the first virtual machine.
[0035] In a possible implementation, the GPU driver module is further used to obtain the first processing result generated by processing the first access request with the first part of the computing processing power of the GPU, and send the first processing result to the backend driver interception module; the backend driver interception module is further used to obtain the first processing result from the GPU driver module and send the first processing result to the first front-end driver interception module; the first front-end driver interception module is further used to obtain the first processing result from the backend driver interception module and provide the first processing result to the first application.
[0036] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain a first access request from the first front-end driver interception module and send the first access request to the back-end driver interception module.
[0037] In a possible implementation, the back-end driver interception module determines that the first access request meets a preset condition, including: the back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than a first partial computing power or a first computing power threshold.
[0038] In a possible implementation, a second virtual machine is further set up on the target server. A second front-end driver interception module and a second application are set up in the second virtual machine. The second front-end driver interception module is configured to perform device simulation on a second partial computing power of the GPU to generate a second VGPU accessible to the second application. The second application is configured to access the second VGPU. The second front-end driver interception module is further configured to obtain a second access request of the second application for the second VGPU; the back-end driver interception module is configured to obtain the second access request from the second front-end driver interception module and send the second access request to the GPU driver module when it determines that the second access request meets the preset condition; the GPU driver module is configured to call the second partial computing power of the GPU to process the second access request.
[0039] In a possible implementation, a second container is set up in the second virtual machine. The second application runs in the second container. The second container is configured to mount the second VGPU for the second application to access the second VGPU in the second container.
[0040] In a possible implementation, the second application runs directly in the second virtual machine.
[0041] In a possible implementation, the GPU driver module is further configured to obtain a second processing result generated by processing the second access request with the second partial computing power of the GPU and send the second processing result to the back-end driver interception module; the back-end driver interception module is further configured to obtain the second processing result from the GPU driver module and send the second processing result to the second front-end driver interception module; the second front-end driver interception module is further configured to obtain the second processing result from the back-end driver interception module and provide the second processing result to the second application.
[0042] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain a second access request from the second front-end driver interception module and send the second access request to the back-end driver interception module.
[0043] In a possible implementation, the backend driver interception module determines that the second access request meets the preset conditions, including: the backend driver interception module determines that the video memory of the GPU required by the second access request is not greater than the second video memory threshold; and / or, the backend driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second partial computing power or the second computing power threshold.
[0044] In a fourth aspect, the present application provides a cloud management platform. The cloud management platform is used to manage infrastructure. The infrastructure includes multiple servers. The cloud management platform includes: an obtaining unit, configured to obtain a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of the first VGPU of the first virtual machine to be created; a selection unit, configured to select a target server that can provide the specification of the first VGPU from multiple servers; and a creation unit, configured to create the first virtual machine on the target server.
[0045] Among them, the target server is provided with a backend driver interception module, a GPU driver module, and a GPU. The first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is configured to perform device simulation on the first partial computing power of the GPU to generate a first VGPU accessible by the first application. The first application is configured to access the first VGPU. The first front-end driver interception module is further configured to obtain a first access request of the first application for the first VGPU. The backend driver interception module is configured to obtain the first access request from the first front-end driver interception module and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions. The GPU driver module is configured to call the first partial computing power of the GPU to process the first access request.
[0046] In a possible implementation, a first container is provided in the first virtual machine. The first application runs in the first container. The first container is configured to mount the first VGPU for the first application to access the first VGPU in the first container.
[0047] In a possible implementation, a third container is further provided in the first virtual machine. The third application runs in the third container. The third container is configured to mount the third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is configured to perform device simulation on the third partial computing power of the GPU to generate a third VGPU accessible by the third application.
[0048] In a possible implementation, the first application runs directly in the first virtual machine.
[0049] In a possible implementation, the GPU driver module is further configured to obtain a first processing result generated by processing a first access request using a first part of the computing processing power of the GPU, and send the first processing result to the back-end driver interception module; the back-end driver interception module is further configured to obtain the first processing result from the GPU driver module and send the first processing result to the first front-end driver interception module; the first front-end driver interception module is further configured to obtain the first processing result from the back-end driver interception module and provide the first processing result to the first application program.
[0050] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the back-end driver interception module.
[0051] In a possible implementation, the back-end driver interception module determines that the first access request meets a preset condition, including: the back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first part of the computing power or the first computing power threshold.
[0052] In a possible implementation, a second virtual machine is further set on the target server. A second front-end driver interception module and a second application program are set in the second virtual machine. The second front-end driver interception module is configured to perform device emulation on a second part of the computing power of the GPU to generate a second VGPU accessible by the second application program. The second application program is configured to access the second VGPU. The second front-end driver interception module is further configured to obtain a second access request of the second application program for the second VGPU; the back-end driver interception module is configured to obtain the second access request from the second front-end driver interception module and send the second access request to the GPU driver module when determining that the second access request meets the preset condition; the GPU driver module is configured to call the second part of the computing power of the GPU to process the second access request.
[0053] In a possible implementation, a second container is set in the second virtual machine. The second application program runs in the second container. The second container is configured to mount the second VGPU for the second application program to access the second VGPU in the second container.
[0054] In a possible implementation, the second application program runs directly in the second virtual machine.
[0055] In a possible implementation, the GPU driver module is further configured to obtain a second processing result generated by processing a second access request with a second part of the computing processing power of the GPU, and send the second processing result to the back-end driver interception module; the back-end driver interception module is further configured to obtain the second processing result from the GPU driver module and send the second processing result to the second front-end driver interception module; the second front-end driver interception module is further configured to obtain the second processing result from the back-end driver interception module and provide the second processing result to the second application program.
[0056] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain a second access request from the second front-end driver interception module and send the second access request to the back-end driver interception module.
[0057] In a possible implementation, the back-end driver interception module determines that the second access request meets a preset condition, including: the back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than a second video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0058] In a fifth aspect, the present application provides a virtual machine creation method based on cloud computing technology. The method is applied to a cloud management platform. The cloud management platform is used to manage infrastructure. The infrastructure includes multiple servers. The method includes: obtaining a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of a first VGPU of a first virtual machine to be created; selecting a second server that can provide the specification of the first VGPU from the multiple servers; and creating the first virtual machine on the second server.
[0059] Among them, the multiple servers further include a first server, the first server is provided with a first back-end driver interception module, a first GPU driver module, and a first GPU. A connection channel is established between the first server and the second server. The first virtual machine is provided with a first front-end driver interception module and a first application program. The first front-end driver interception module is configured to perform device simulation on a part of the computing power of the first GPU to generate a first VGPU accessible to the first application program. The first application program is configured to access the first VGPU. The first front-end driver interception module is further configured to obtain a first access request of the first application program for the first VGPU and send the first access request to the connection channel; the first back-end driver interception module is configured to obtain the first access request from the first front-end driver interception module through the connection channel and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset condition; the first GPU driver module is configured to call a part of the computing power of the first GPU to process the first access request.
[0060] In a possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further configured to perform device simulation on the first part of the computing power of the second GPU to generate a third VGPU accessible to the first application, the first application is used to access the third VGPU, and the first front-end driver interception module is further configured to obtain a third access request of the first application for the third VGPU; the second back-end driver interception module is configured to obtain the third access request from the first front-end driver interception module and send the third access request to the second GPU driver module when it is determined that the third access request meets a preset condition; the second GPU driver module is configured to call the first part of the computing power of the second GPU to process the third access request.
[0061] In a possible implementation, a first container is set up in the first virtual machine, the first application runs in the first container, and the first container is used to mount the first VGPU for the first application to access the first VGPU in the first container.
[0062] In a possible implementation, a third container is further set up in the first virtual machine, the third application runs in the third container, and the third container is used to mount the third VGPU for the third application to access the third VGPU in the third container, wherein the first front-end driver interception module is configured to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application.
[0063] In a possible implementation, the first application runs directly in the first virtual machine.
[0064] In a possible implementation, the first GPU driver module is further configured to obtain the first processing result generated by processing the first access request with the first part of the computing and processing power of the first GPU, and send the first processing result to the first back-end driver interception module; the first back-end driver interception module is further configured to obtain the first processing result from the first GPU driver module and send the first processing result to the connection channel; the first front-end driver interception module is further configured to obtain the first processing result from the connection channel from the first back-end driver interception module and provide the first processing result to the first application.
[0065] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel.
[0066] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, and the high-speed interconnection protocol includes the Compute Express Link (CXL) or the Lingqu Universal Bus (UB) protocol.
[0067] In a possible implementation, the first back-end driver interception module determines that the first access request meets the preset conditions, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0068] In a possible implementation, a second virtual machine is further set up on the second server. A second front-end driver interception module and a second application are set up in the second virtual machine. The second front-end driver interception module is used to perform device simulation on the second partial computing power of the first GPU to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is further used to obtain a second access request of the second application for the second VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the second access request from the second front-end driver interception module from the connection channel and send the second access request to the first GPU driver module when it is determined that the second access request meets the preset conditions; the first GPU driver module is used to call the second partial computing power of the first GPU to process the second access request.
[0069] In a possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The second front-end driver interception module of the second virtual machine is further used to perform device simulation on the second partial computing power of the second GPU to generate a fourth VGPU accessible to the second application. The second application is used to access the fourth VGPU. The second front-end driver interception module is further used to obtain a fourth access request of the second application for the fourth VGPU; the second back-end driver interception module is used to obtain the fourth access request from the second front-end driver interception module and send the fourth access request to the second GPU driver module when it is determined that the fourth access request meets the preset conditions; the second GPU driver module is used to call the second partial computing power of the second GPU to process the fourth access request.
[0070] In a possible implementation, a second container is set up in the second virtual machine. The second application runs in the second container. The second container is used to mount the second VGPU for the second application to access the second VGPU in the second container.
[0071] In a possible implementation, a fourth container is further set up in the second virtual machine. The fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU for the fourth application to access the fourth VGPU in the fourth container. Among them, the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application.
[0072] In a possible implementation, the second application runs directly in the second virtual machine.
[0073] In a possible implementation, the first GPU driver module is further used to obtain the second processing result generated by processing the second access request with the second part of the computing and processing power of the first GPU, and send the second processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the second processing result from the first GPU driver module and send the second processing result to the connection channel; the second front-end driver interception module is further used to obtain the second processing result from the first back-end driver interception module through the connection channel and provide the second processing result to the second application.
[0074] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module and send the second access request to the connection channel.
[0075] In a possible implementation, the second back-end driver interception module determines that the second access request meets the preset conditions, including: the second back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than the second video memory threshold; and / or, the second back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0076] In a sixth aspect, the present application provides a cloud management platform. The cloud management platform is used to manage the infrastructure. The infrastructure includes multiple servers. The cloud management platform includes: an acquisition unit, configured to acquire a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of the first VGPU of the first virtual machine to be created; a selection unit, configured to select a second server that can provide the specification of the first VGPU from multiple servers; a creation unit, configured to create the first virtual machine on the second server.
[0077] Among them, the multiple servers further include a first server, which is provided with a first back-end driver interception module, a first GPU driver module, and a first GPU. A connection channel is established between the first server and the second server. A first front-end driver interception module and a first application are set in the first virtual machine. The first front-end driver interception module is used to perform device simulation on a part of the computing power of the first GPU to generate a first VGPU accessible to the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU and send the first access request to the connection channel. The first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions. The first GPU driver module is used to call a part of the computing power of the first GPU to process the first access request.
[0078] In a possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further used to perform device simulation on the first part of the computing power of the second GPU to generate a third VGPU accessible to the first application. The first application is used to access the third VGPU. The first front-end driver interception module is further used to obtain a third access request of the first application for the third VGPU. The second back-end driver interception module is used to obtain the third access request from the first front-end driver interception module and send the third access request to the second GPU driver module when it is determined that the third access request meets the preset conditions. The second GPU driver module is used to call the first part of the computing power of the second GPU to process the third access request.
[0079] In a possible implementation, a first container is set in the first virtual machine. The first application runs in the first container. The first container is used to mount the first VGPU for the first application to access the first VGPU in the first container.
[0080] In a possible implementation, a third container is further set in the first virtual machine. The third application runs in the third container. The third container is used to mount the third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application.
[0081] In a possible implementation, the first application runs directly in the first virtual machine.
[0082] In a possible implementation, the first GPU driver module is further configured to obtain a first processing result generated by processing a first access request using a first part of the computing processing power of the first GPU, and send the first processing result to the first back-end driver interception module; the first back-end driver interception module is further configured to obtain the first processing result from the first GPU driver module and send the first processing result to the connection channel; the first front-end driver interception module is further configured to obtain the first processing result from the connection channel that comes from the first back-end driver interception module and provide the first processing result to the first application.
[0083] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel.
[0084] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, and the high-speed interconnection protocol includes Compute Express Link (CXL) or Lingqu Universal Bus (UB) protocol.
[0085] In a possible implementation, the first back-end driver interception module determines that the first access request meets a preset condition, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first part of the computing power or a first computing power threshold.
[0086] In a possible implementation, a second virtual machine is further set up on the second server. A second front-end driver interception module and a second application are set up in the second virtual machine. The second front-end driver interception module is configured to perform device emulation on a second part of the computing power of the first GPU to generate a second virtual GPU (VGPU) accessible to the second application. The second application is configured to access the second VGPU. The second front-end driver interception module is further configured to obtain a second access request of the second application for the second VGPU and send the first access request to the connection channel; the first back-end driver interception module is configured to obtain the second access request from the connection channel that comes from the second front-end driver interception module and send the second access request to the first GPU driver module when determining that the second access request meets the preset condition; the first GPU driver module is configured to call a second part of the computing power of the first GPU to process the second access request.
[0087] In a possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The second front-end driver interception module of the second virtual machine is further configured to perform device simulation on the second part of the computing power of the second GPU to generate a fourth VGPU accessible to the second application. The second application is used to access the fourth VGPU. The second front-end driver interception module is further configured to obtain a fourth access request of the second application for the fourth VGPU. The second back-end driver interception module is configured to obtain the fourth access request from the second front-end driver interception module and send the fourth access request to the second GPU driver module when it is determined that the fourth access request meets a preset condition. The second GPU driver module is configured to call the second part of the computing power of the second GPU to process the fourth access request.
[0088] In a possible implementation, a second container is provided in the second virtual machine. The second application runs in the second container. The second container is used to mount the second VGPU for the second application to access the second VGPU in the second container.
[0089] In a possible implementation, a fourth container is further provided in the second virtual machine. The fourth application runs in the fourth container. The fourth container is used to mount the fourth VGPU for the fourth application to access the fourth VGPU in the fourth container. Among them, the second front-end driver interception module is configured to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application.
[0090] In a possible implementation, the second application runs directly in the second virtual machine.
[0091] In a possible implementation, the first GPU driver module is further configured to obtain the second processing result generated by processing the second access request with the second part of the computing and processing power of the first GPU, and send the second processing result to the first back-end driver interception module. The first back-end driver interception module is further configured to obtain the second processing result from the first GPU driver module and send the second processing result to the connection channel. The second front-end driver interception module is further configured to obtain the second processing result from the connection channel from the first back-end driver interception module and provide the second processing result to the second application.
[0092] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module and send the second access request to the connection channel.
[0093] In a possible implementation, the second back-end driver interception module determines that the second access request meets the preset conditions, including: the second back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than the second video memory threshold; and / or, the second back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second partial computing power or the second computing power threshold.
[0094] In a seventh aspect, the present application provides a computing device, including a memory and a processor. The memory stores program instructions, and the processor runs the program instructions to implement the server provided in the first aspect of the present application and any of its possible implementation manners, and to implement the server system provided in the first aspect of the present application and any of its possible implementation manners.
[0095] In an eighth aspect, the present application provides a computing device cluster, including a plurality of computing devices. The plurality of computing devices include a plurality of processors and a plurality of memories. Program instructions are stored in the plurality of memories, and the plurality of processors run the program instructions, so that the computing device cluster implements the server provided in the first aspect of the present application and any of its possible implementation manners, and implements the server system provided in the first aspect of the present application and any of its possible implementation manners.
[0096] In a ninth aspect, the present application provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium. The computer-readable storage medium includes program instructions. When the program instructions run on a computing device, the computing device implements the server provided in the first aspect of the present application and any of its possible implementation manners, and implements the server system provided in the first aspect of the present application and any of its possible implementation manners.
[0097] In a tenth aspect, the present application provides a computer program product containing instructions. When the computer program product runs on a computer, the computer implements the server provided in the first aspect of the present application and any of its possible implementation manners, and implements the server system provided in the first aspect of the present application and any of its possible implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Figure 1 is a schematic structural diagram of an implementation scenario provided by an embodiment of the present application;
[0099] Figure 2 is a schematic deployment diagram of a basic resource provided by an embodiment of the present application;
[0100] Figure 3 is a schematic structural diagram of a server based on cloud computing technology provided by an embodiment of the present application;
[0101] Figure 4 It is a schematic structural diagram of another server based on cloud computing technology provided by an embodiment of the present application;
[0102] Figure 5 It is a schematic structural diagram of another server based on cloud computing technology provided by an embodiment of the present application;
[0103] Figure 6 It is a schematic structural diagram of yet another server based on cloud computing technology provided by an embodiment of the present application;
[0104] Figure 7 It is a schematic structural diagram of another server based on cloud computing technology provided by an embodiment of the present application;
[0105] Figure 8 It is a schematic structural diagram of yet another server based on cloud computing technology provided by an embodiment of the present application;
[0106] Figure 9 It is a schematic structural diagram of a server system based on cloud computing technology provided by an embodiment of the present application;
[0107] Figure 10 It is a schematic structural diagram of another server system based on cloud computing technology provided by an embodiment of the present application;
[0108] Figure 11 It is a schematic structural diagram of another server system based on cloud computing technology provided by an embodiment of the present application;
[0109] Figure 12 It is a schematic structural diagram of yet another server based on cloud computing technology provided by an embodiment of the present application;
[0110] Figure 13 It is a schematic diagram showing that the functions of a backend-driven interception module and a frontend-driven interception module can be implemented by multiple functional units;
[0111] Figure 14 It is a flowchart of a method for creating a virtual machine based on cloud computing technology provided by an embodiment of the present application;
[0112] Figure 15 It is a schematic diagram of a cloud management platform provided by an embodiment of the present application;
[0113] Figure 16 It is a flowchart of another method for creating a virtual machine based on cloud computing technology provided by an embodiment of the present application;
[0114] Figure 17 It is a schematic diagram of a cloud management platform provided by an embodiment of the present application;
[0115] Figure 18It is a schematic structural diagram of a computing device provided by an embodiment of the present application;
[0116] Figure 19 It is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application;
[0117] Figure 20 It is a schematic structural diagram of another computing device cluster provided by an embodiment of the present application. Detailed implementation manners
[0118] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0119] For ease of understanding, the technologies and backgrounds involved in the embodiments of the present application will be introduced first below.
[0120] An Internet Data Center (IDC) is a facility and related service system for the operation and maintenance of devices that centrally collect, store, process, and send data based on the Internet network. Conceptually, it can be understood as a public commercial Internet "computer room", and at the same time, it is also a type of IT professional service and an important infrastructure of the IT industry. An IDC is not only a service concept but also a network concept. It constitutes a part of the network basic resources and provides a high-end data delivery service and high-speed access service like a backbone network and an access network. Generally, the offline IDC of a tenant can be understood as the tenant's offline computer room, where the tenant uses existing Internet communication lines and bandwidth resources to establish a standardized telecommunications professional-level computer room environment for providing all-round services such as server hosting, leasing, and related value-added services. A cloud data center is an Internet data center deployed using the basic resources owned by cloud providers.
[0121] A resource pool is a collection of various hardware resources and software resources involved in a cloud data center. Generally, according to the type of resources, the resources in the resource pool can be divided into computing resources, storage resources, network resources, etc.
[0122] Host (Physical Machine, PM): A physical resource for carrying virtualization technology. A host is also called a physical machine. Generally, the host used to deploy virtual instances is a physical server. A physical machine has multiple physical devices. For example, a physical server has physical devices such as a processor and a memory. Multiple virtual instances can be deployed on one host, and the multiple virtual instances deployed on the same host share the physical resources of the host. According to different usage situations, the multiple virtual instances deployed on one host can optionally belong to the same tenant or different tenants respectively.
[0123] Virtualization is a resource management technology. Virtualization can abstract and transform various physical resources of a host, such as computing resources, network resources, and storage resources, etc., and present them, so as to break the non-separable barriers between the physical structures of the host, enabling tenants to apply these resources in a better way than the original configuration. The resources obtained through virtualization are called virtualized resources, and the virtualized resources are not restricted by the installation method, setting location, or physical configuration of the existing physical resources.
[0124] Virtualized resources are usually provided to tenants in the form of virtual instances. The virtual instance can use the hardware resources of the host and run on the operating system of the host. An application program runs in the virtual instance, and this application program is used to implement the business of the tenant. The hardware resources of the host can be used by one or more tenants at the granularity of virtual instances. Different virtual instances are isolated from each other, enabling tenants to conveniently and flexibly use physical resources on the premise of security isolation, and can greatly improve the utilization rate of physical resources. Generally, the virtual instance can be a virtual machine, a container, or an independent process (such as a function), etc. The virtual instance can also be called an elastic compute service (ECS), an elastic instance (different cloud service providers have different names).
[0125] Virtual machine (VM): It refers to a complete computer system with the functions of a complete hardware system obtained through virtualization technology and running in a completely isolated environment. A partial subset of the instructions of the virtual machine can be processed in the host machine, and the other part of the instructions can be executed in an emulated manner. A virtual machine is also called a virtual server. A virtual machine can be regarded as a collection of several virtual devices. The collection of these several virtual devices has the functions of a complete hardware system and runs in a completely isolated environment. The virtual devices are virtually obtained through virtualization technology based on physical devices that can share resources. For example, based on virtualization technology, a virtual processor virtually obtained based on a processor is a virtual device. Another example is that based on virtualization technology, a training card virtually obtained based on a field-programmable gate array (FPGA) is also a virtual device. By way of example, the virtual machine in the present application can be a kernel-based virtual machine (KVM). What can be accomplished in a server can all be achieved in a virtual machine. When creating a virtual machine in a server, it is necessary to use a part of the hard disk and memory capacity of the physical machine as the hard disk and memory capacity of the virtual machine. Each virtual machine has an independent hard disk and operating system. The tenant of the virtual machine can operate the virtual machine just like using a server. The running environments in different virtual machines (such as virtual machine applications, operating systems, and virtual hardware) are completely isolated. To communicate between different virtual machines, it is necessary to forward network packets through a virtual manager.
[0126] The container utilizes the namespace and cgroup technologies supported by the Linux kernel to isolate the application APP process and its dependent packages (runtime environment bins / libs, specifically all files required to run the APP) in an independent runtime environment. The container provides a lightweight virtual runtime environment. The container can be obtained by packaging all the code, libraries, and dependencies of the tenant's application into an image. When the image is executed, the image runs in the virtual runtime environment. At this time, the container is the runtime instance of the image, similar to a lightweight sandbox, which can be started, stopped, and deleted. The infrastructure of the container can be the hardware of the server or a virtual machine on the cloud (i.e., containers can also be deployed in virtual machines). The operating system uses the Linux kernel, which supports namespace and cgroup. Among them, namespace is used to isolate processes, and cgroup is used to allocate process resources. The resources are specifically the virtual processors and memory allocated to the process. The container engine is similar to a virtual machine manager and runs in the operating system to manage containers. Compared with the characteristics of virtual machines with built-in operating systems, containers do not have an operating system. Containers run as processes in the host operating system, so the startup speed of containers is faster than that of virtual machines. They are particularly suitable for lightweight applications, and a single host can run thousands of containers (processes) simultaneously.
[0127] Resource pooling refers to integrating various computing and storage resources together to form a unified resource library for unified dynamic allocation and management. Resource pooling can achieve highly shared resources, improve resource utilization, simplify resource management, and provide users with flexible on-demand allocation services.
[0128] In view of this, the embodiments of the present application provide a server based on cloud computing technology, a server system, and a corresponding virtual machine creation method and cloud management platform based on cloud computing technology. In the present application, a virtual machine runs in the server. A front-end driver interception module and an application are set in the virtual machine. The front-end driver interception module can provide a virtual graphics processing unit (VGPU) accessible to the application in the virtual machine based on partial computing capabilities of the graphics processing unit (GPU). After the application accesses the VGPU, the front-end driver interception module can obtain the access request of the application for the VGPU and provide the access request to the back-end driver interception module. After obtaining the access request from the front-end driver interception module, the back-end driver interception module can send the access request to the GPU driver module when it determines that the access request meets the preset conditions. The GPU driver module can call partial computing capabilities of the GPU to process the access request.
[0129] In this application, since the front-end driver interception module can obtain the access request of the application program in the virtual machine for the VGPU and provide the access request to the GPU driver module through the back-end driver interception module, the GPU driver module can call part of the computing power of the GPU to process the access request. In this way, the operation of the application program in the virtual machine to access the GPU is realized in a software manner, reducing the degree of dependence on the implementation details such as the hardware information of the GPU in the process of calling the GPU to process the access request, and improving the flexibility of the application program in the virtual machine to access the GPU.
[0130] It should be noted that although this application is described by taking the access of the application program in the virtual machine to the GPU as an example, it does not exclude that the implementation method of accessing the GPU in this application can also be applied to the access of the application program running in other types of virtual instances to the GPU, and can also be applied to the access of other types of hardware. For example, the implementation method of accessing the GPU in this application can also be applied to the response of the access request of the application program running in the container to the GPU. Another example is that the implementation method of accessing the GPU in this application can also be applied to the access of the application program running in the virtual instance to other hardware such as hard disks, disks, graphics cards, intelligent processing units (IPUs), and data processing units (DPUs). And when the implementation method of accessing the hardware in this application is applied to other types of virtual instances and / or other types of hardware, the implementation method can refer to the implementation method of the application program running in the virtual machine accessing the GPU accordingly, and will not be elaborated herein.
[0131] This article introduces the technical solution of this application in detail from multiple perspectives such as implementation scenarios, method flows, hardware devices, and software devices.
[0132] First, an example of the implementation scenario of the embodiments of this application will be described below.
[0133] Figure 1 It is a schematic structural diagram of an implementation scenario provided by the embodiments of this application. As Figure 1 shown, this implementation scenario includes: a data center 1 and a client 2. A communication connection can be established between the data center 1 and the client 2 through a network. Optionally, the network can be the Internet or other networks, which is not limited in the embodiments of this application. A tenant can interact with the data center 1 through the client 2. For example, a tenant can send information such as a cloud service request to the data center 1 through the client 2. The data center 1 is used to respond based on the information sent by the client 2.
[0134] A large number of infrastructures owned by cloud service providers are deployed in Data Center 1, such as computing resources, storage resources, and network resources. For example, computing resources can be computing devices (such as servers, etc.) that can provide computing power. As Figure 1 shown, Data Center 1 includes a cloud management platform and infrastructure ( Figure 1 not shown in the figure). The cloud management platform and the infrastructure are connected through the internal network of the data center. The cloud management platform is used to manage the infrastructure. The infrastructure is used to provide public cloud services. The infrastructure includes multiple servers. Cloud services can be optionally deployed in the servers. Cloud services are implemented by running virtual instances, so they are also called virtual instances for implementing tenant services deployed in the servers. Tenants can send cloud service requests and related information to the servers through the client 2 they use. The servers can process the cloud service requests and related information, and provide cloud services to the tenants based on the processed cloud service requests and related information. For example, the servers can receive virtual instance creation requests provided by tenants and create virtual instances for tenants in the servers according to the virtual instance creation requests.
[0135] Logically, the cloud management platform can be divided into: tenant console, computing management service, network management service, storage management service, authentication service, and image management service. The tenant console provides an interface or application program interface (API) to interact with tenants. The computing management service is used to manage the servers running virtual instances and bare metal servers. The network management service is used to manage network services (such as gateways, firewalls, etc.). The storage management service is used to manage storage services (such as data bucket services). The authentication service is used to manage tenant accounts and passwords. The image management service is used to manage images of virtual instances.
[0136] In Figure 1 the shown implementation scenario, multiple servers are set in a data center. The servers include a hardware layer and a software layer. The hardware layer is the conventional configuration of the servers. Hardware devices such as processors, memories, network cards, disks, and buses are deployed in the hardware layer. The software layer includes an operating system installed and running on the servers. The operating system relative to the virtual machine can be called the host operating system. A virtual machine manager (also called Hypervisor) runs in the host operating system. The role of the virtual machine manager is to implement computing virtualization, network virtualization, and storage virtualization of virtual machines and be responsible for managing virtual machines.
[0137] A cloud management platform client runs in the virtual machine manager. The cloud management platform client can receive control plane commands sent by the cloud management platform, create virtual instances on the server according to the control plane control commands, and perform full-life cycle management on the virtual instances. For example, the cloud management platform client can detect the usage of the hardware resources of the server where it is located in real time and report it to the cloud management platform. When the cloud management platform confirms to create a virtual instance on a certain server, it will send a virtual instance creation command to the cloud management platform client on that server. After receiving the command, the cloud management platform client creates a virtual instance on that server. In this way, tenants can create, manage, log in to, and operate virtual instances in the data center through the cloud management platform.
[0138] The server can be used to run virtual machines of different specifications. The specifications of virtual machines are divided into: general computing type, memory optimized type, extra-large memory type, etc. There are specific specifications under each type. After the tenant selects the specification of the virtual machine, the cloud management platform selects a server that supports this specification in the data center, and determines that there is enough free hardware resource on this server, and then creates a virtual machine with this specification set on this server. By configuring the server through the cloud management platform, it is possible to analyze and plan the hardware resources of the server, plan the corresponding computing products of the physical hardware according to the hardware performance of the server, such as planning virtual machines of different specifications, to meet the different demands of different tenants. Moreover, according to the performance differences of virtual machines of different specifications, a differential pricing strategy can be implemented. For example, virtual instances with high-performance specifications are sold at a higher price, and virtual instances with ordinary performance specifications are sold at a lower price, so that tenants can purchase virtual instances according to their needs.
[0139] In one implementation, such as Figure 2As shown, the location of the basic resources in the data center can be described by the cloud resource deployment region (region) and the availability zone (AZ). Tenants can choose to deploy cloud services based on the resources in a specific region and AZ. Among them, regions are divided from the dimensions of geographical location and network latency. The same resource pool is used within the same region, which can be understood as sharing common services such as elastic computing, block storage, object storage, virtual private cloud (VPC) network, elastic internet protocol (EIP) address, and images. Regions are divided into general regions and dedicated regions. A general region refers to a region that provides general cloud services for public tenants. A dedicated region refers to a dedicated region that hosts the same type of business or provides business services for specific tenants. A region usually includes multiple AZs. Multiple AZs in a region are connected by high-speed optical fibers to meet the needs of tenants to build highly available systems across AZs. An AZ is a set of one or more Figure 2 data centers as shown. The resources such as computing, network, and storage within an AZ are logically divided into multiple clusters.
[0140] Tenants can send instructions to the cloud management platform through the used client 2 to create, manage, log in to, and operate virtual instances in the server, and use the cloud services provided by the virtual instances. Exemplarily, the cloud management platform can provide access interfaces. The access interfaces can optionally be provided in the form of an interface or an API. Tenants can operate the client to remotely access the access interface to register cloud accounts and passwords on the cloud management platform and use the cloud accounts and passwords to log in to the cloud management platform. The cloud management platform can also authenticate the cloud accounts and passwords. After successful authentication, tenants can further select and pay for virtual instances of specific specifications (processor, memory, disk) on the cloud management platform. After tenants successfully pay for the virtual instances, the cloud management platform provides the remote login accounts and passwords of the purchased virtual instances to the tenants. Tenants can use the remote login accounts and passwords to remotely log in to the virtual instances on the client, install and run the tenants' applications in the virtual instances, so as to implement the tenants' business through the applications.
[0141] The client 2 can optionally be a computer, a personal computer, a laptop, a mobile phone, a smartphone, a tablet computer, a cloud host, a portable mobile terminal, a multimedia player, an e-book reader, a wearable device, a smart home appliance, an artificial intelligence device, a smart wearable device, a smart vehicle-mounted device, or an Internet of Things device, etc.
[0142] In one implementation, the functions of the server, server system, corresponding virtual machine creation method based on cloud computing technology, and cloud management platform provided by the embodiments of the present application can be implemented by a computing device in data center 1 running an executable program. Moreover, the executable program for implementing this function can optionally be presented in the form of an application installation package. After the server installs this application installation package, it can implement this function by running the executable program therein.
[0143] It should be understood that the above content is an exemplary description of the implementation scenarios involved in the embodiments of the present application, and does not constitute a limitation on the implementation scenarios involved in the embodiments of the present application. Those of ordinary skill in the art know that with the change of business requirements, its implementation scenarios can be adjusted according to application requirements, and the embodiments of the present application do not make specific limitations thereto.
[0144] First, the implementation of the server based on cloud computing technology provided by the present application will be introduced.
[0145] Figure 3 It is a schematic structural diagram of a server based on cloud computing technology provided by the embodiments of the present application. As Figure 3 shown, a first virtual machine and a second virtual machine are running on the server. The server is further provided with a backend driver interception module, a GPU driver module, and a GPU.
[0146] The first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on the first part of the computing power of the GPU to generate a first VGPU that can be accessed by the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU. Among them, the first application is used to implement the user's business.
[0147] The second virtual machine is provided with a second front-end driver interception module and a second application. The second front-end driver interception module is used to perform device simulation on the second part of the computing power of the GPU to generate a second VGPU that can be accessed by the second application. The second application is used to access the second VGPU. The second front-end driver interception module is further used to obtain a second access request of the second application for the second VGPU. Among them, the second application is used to implement the user's business.
[0148] The backend driver interception module is used to obtain the first access request from the first front-end driver interception module and / or the second access request from the second front-end driver interception module, and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions, and / or send the second access request to the GPU driver module when it is determined that the second access request meets the preset conditions.
[0149] The GPU driver module is used to call the first part of the computing power of the GPU to process the first access request, and / or call the second part of the computing power of the GPU to process the second access request.
[0150] It should be noted that the fact that the first virtual machine and the second virtual machine are running on the server is only an example. The number of virtual machines deployed in the server can be adjusted based on application requirements, and the embodiments of the present application do not make specific limitations on this. For example, fewer or more servers can be optionally deployed in the server according to application requirements. For example, at least one virtual machine is deployed in the server. When at least one virtual machine is deployed in the server, for the access request of the virtual hardware obtained based on the hardware by the application program set in any virtual machine, it can be transmitted to the backend driver interception module according to the method provided in the present application, and transmitted to the driver of the hardware through the backend driver module, so that the driver of the hardware calls the capabilities of the hardware to process the access request. Among them, Figure 4 FIG. is a schematic diagram of deploying one virtual machine in the server.
[0151] After the GPU driver module calls the computing power of the GPU to process the access request, if it receives the processing result generated by the GPU processing the access request, the GPU driver module is further used to provide the processing result to the backend driver interception module, so as to return the processing result to the application program that triggers the access request through the backend driver interception module. The roles played by each module in the server during the process of returning the processing result are as follows:
[0152] The GPU driver module is used to obtain the first processing result generated by the first part of the computing and processing power of the GPU to process the first access request, and / or obtain the second processing result generated by the second part of the computing and processing power of the GPU to process the second access request, and send the first processing result and / or the second processing result to the backend driver interception module.
[0153] The backend driver interception module is used to obtain the first processing result and / or the second processing result from the GPU driver module, and send the first processing result to the first front-end driver interception module, and / or send the second processing result to the second front-end driver interception module.
[0154] The first front-end driver interception module is used to obtain the first processing result from the backend driver interception module and provide the first processing result to the first application program.
[0155] The second front-end driver interception module is used to obtain the second processing result from the backend driver interception module and provide the second processing result to the second application program.
[0156] Exemplarily, when the first application in the first virtual machine needs to use the GPU for rendering operations, the first application will access the first VGPU, and this access operation will trigger a first access request indicating the access to the GPU. The first access request carries relevant data of the image to be rendered. The first front-end driver interception module can obtain the first access request and forward the first access request to the back-end driver interception module. After receiving the first access request, if the first access request meets the preset conditions, the back-end driver interception module forwards the first access request to the GPU driver module. After receiving the first access request, the GPU driver module calls the first part of the computing power of the GPU to process the first access request. After obtaining the processing result of the GPU for the first access request, the GPU driver module forwards the processing result to the back-end driver interception module. The processing result carries the video memory address of the rendering result, and the rendering result is the result of the GPU rendering based on the relevant data carried by the first access request. After receiving the processing result, the back-end driver interception module forwards the processing result to the first front-end driver interception module. After receiving the processing result, the first front-end driver interception module sends the processing result to the first application, so that the first application can obtain the rendering result from the video memory address based on the video memory address carried by the processing result.
[0157] It can be seen from this that the access request triggered by the application in the virtual machine accessing the VGPU can be transmitted to the GPU driver module in the server through the front-end driver interception module in the virtual machine and the back-end driver interception module in the server in sequence. The processing result generated by the GPU processing the access request can be transmitted to the application in the virtual machine through the GPU driver module in the server, the back-end driver interception module in the server, and the front-end driver interception module in the virtual machine in sequence. Since both of these transmission processes are implemented in software, the dependence on implementation details such as the hardware information of the GPU in the process of calling the GPU to process the access request can be reduced, and even the implementation details such as the hardware information of the GPU do not need to be relied on. For example, it does not rely on the hardware information of the GPU manufacturer, and the decoupling from the GPU can be achieved, improving the flexibility of the application in the virtual machine to access the GPU. And when the virtual machine uses part of the computing power of the GPU, it is equivalent to virtualizing and partitioning the GPU. The software implementation method of transmitting the access request and the processing result in this application enables the partitioning specification of virtualizing and partitioning the GPU to break through the hardware control of the GPU, and can support various scheduling strategies for scheduling the GPU, improving the flexibility of scheduling the GPU.
[0158] In this application, there are multiple implementation methods for the server. The following takes the following three implementation methods as examples to illustrate it.
[0159] In the first implementation, the first application runs directly in the first virtual machine, and the second application runs directly in the second virtual machine. As Figure 3 shown, the scenario presented by this implementation is that the first application running directly in the virtual machine deployed inside the server accesses the GPU available on the server. Since the operations of the virtual machine are managed by a hypervisor, and the GPU, GPU driver module, and backend driver interception module are managed by the operating system (OS) of the server, this scenario can be referred to as a scenario for GPU access implemented through the Hypervisor and OS.
[0160] In this first implementation, the VGPU used by the virtual machine deployed in the server is obtained based on a portion of the computing power of the GPU in the server, which is equivalent to virtualizing and partitioning the server's GPU. In this application, the backend driver interception module and the frontend driver interception module forward access requests and processing results between the GPU driver module and the application, reducing the degree of dependence on implementation details such as the hardware information of the GPU. Even without relying on implementation details such as the hardware information of the GPU, the flexibility of the application in the virtual machine to access the GPU is improved. Moreover, the server of this application supports multiple scheduling policies for GPU scheduling, enhancing the flexibility of GPU scheduling.
[0161] In the second implementation, a first container is set up in the first virtual machine, the first application runs in the first container, and the first container is used to mount a first VGPU for the first application to access the first VGPU in the first container. A second container is set up in the second virtual machine, the second application runs in the second container, and the second container is used to mount a second VGPU for the second application to access the second VGPU in the second container. Figure 5 is a schematic diagram of a server provided by an embodiment of this application. As Figure 5 shown, the scenario presented by this implementation is that a virtual machine is deployed in the server, a container is deployed in the virtual machine, and the application in the container has a need to access the GPU of the VGPU provided for it. Since the operations of both the virtual machine and the container are managed by the hypervisor, and the GPU, GPU driver module, and backend driver interception module are managed by the OS of the server, this scenario can be referred to as a GPU access scenario implemented through the Hypervisor, OS, and container.
[0162] Optionally, the number of containers set up in the virtual machine can be adjusted according to application requirements. For example, one container can be set up in the virtual machine, or multiple containers can be set up in the virtual machine. By way of example, as Figure 6As shown, a third container is also set in the first virtual machine. The third application runs in the third container, and the third container is used to mount a third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the GPU to generate a third VGPU that can be accessed by the third application. For the implementation method of the third application accessing the third VGPU, please refer to the implementation methods of the first application accessing the first VGPU and the second application accessing the second VGPU respectively, which will not be elaborated here. Similarly, a fourth container ( Figure 6 not shown in the figure) is also set in the second virtual machine. The fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU for the fourth application to access the fourth VGPU in the fourth container. Among them, the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the GPU to generate a fourth VGPU that can be accessed by the fourth application. For the implementation method of the fourth application accessing the fourth VGPU, please refer to the implementation methods of the first application accessing the first VGPU and the second application accessing the second VGPU respectively, which will not be elaborated here.
[0163] This second implementation method not only has the effects of the first implementation method, but also further divides the virtual machine on the basis of virtualizing and dividing the GPU, realizing the secondary division and comprehensive scheduling of the virtual machine after GPU resource division, and further improving the flexibility of accessing the GPU and the flexibility of scheduling the GPU. It should be noted that the container running in the virtual machine is only an example, and other types of virtual instances can also run in the virtual machine. As is known to those of ordinary skill in the art, with the change of business requirements, the type of virtual instance running in the virtual machine can be adjusted according to application requirements, and the embodiments of the present application do not make specific limitations on it. For example, the way of deploying virtual instances in the server can be further evolved on this basis.
[0164] It should be noted that the above two implementation methods of the server can be used alone or in combination. For example, when multiple virtual machines are running in the server, each virtual machine in the multiple virtual machines can be deployed according to the above first implementation method, that is, the application runs directly in the virtual machine. Or, each virtual machine in the multiple virtual machines can be deployed according to the above second implementation method, that is, containers are set in the virtual machines, and applications run in the containers. Or, some of the multiple virtual machines are deployed according to the above first implementation method, and some of the virtual machines are deployed according to the above second implementation method. By way of example, Figure 7 is a schematic diagram of a server provided by an embodiment of the present application. As Figure 7As shown, the second application program runs directly in the second virtual machine in the server. The first application program runs in the first container in the first virtual machine in the server, and the third application program runs in the third container.
[0165] When the backend driver interception module determines that an access request meets the preset conditions, it forwards the access request to the GPU driver module. When the backend driver interception module determines that the access request does not meet the preset conditions, it does not forward the access request to the GPU driver module. At this time, the backend driver interception module can return an error message to the virtual machine to which the access request belongs to indicate that the virtual machine cannot access the GPU. In a possible implementation, the backend driver interception module determines that the access request meets the preset conditions, including: the backend driver interception module determines that the video memory of the GPU required by the access request is not greater than the video memory threshold; and / or, the backend driver interception module determines that the computing power of the GPU required by the access request is not greater than the specified computing power or the first computing power threshold. For example, the backend driver interception module determines that the first access request meets the preset conditions, including: the backend driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or, the backend driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold. The backend driver interception module determines that the second access request meets the preset conditions, including: the backend driver interception module determines that the video memory of the GPU required by the second access request is not greater than the second video memory threshold; and / or, the backend driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second partial computing power or the second computing power threshold.
[0166] In a possible implementation, the backend driver interception module receives an access request, obtains the virtual machine that initiates the access request, and determines whether the virtual machine has a preset condition for the GPU. There are multiple ways to express the preset condition. For example, when the tenant does not configure the GPU capability for the virtual machine, the access request from the virtual machine does not meet the preset condition. Or, when the usage quota of the virtual machine for the GPU exceeds the quota configured by the tenant for the virtual machine, the access request from the virtual machine does not meet the preset condition. Among them, the access request from the virtual machine refers to the access request generated by the application running in the virtual machine to access the VGPU. For example, when the usage quota configured by the tenant for the virtual machine is the video memory quota that can be used by the virtual machine, the video memory quota configured by the tenant for the virtual machine can be regarded as the maximum usage of the video memory that can be used by the virtual machine, and the video memory quota can be used as the video memory threshold. When the video memory required by the virtual machine is not greater than the video memory threshold, it is determined that the access request from the virtual machine meets the preset condition. For example, the backend driver interception module determines that the first access request meets the preset condition, including: the backend driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold. The back-end driver interception module determines that the second access request meets the preset conditions, including: the back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold. For another example, when the usage quota configured by the tenant for the virtual machine is the quota of computing power that the virtual machine can use, the quota of computing power configured by the tenant for the virtual machine can be regarded as the maximum usage of the computing power that the virtual machine can use, and the quota of the computing power can be used as the computing power threshold. When the computing power required by the virtual machine is not greater than the computing power threshold, it is determined that the access request from the virtual machine meets the preset conditions. For example, the back-end driver interception module determines that the first access request meets the preset conditions, including: the back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first computing power threshold. The back-end driver interception module determines that the second access request meets the preset conditions, including: the back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second computing power threshold. For another example, since the VGPU is generated by the front-end driver interception module in the virtual machine by simulating part of the GPU's computing power, when judging whether the access request meets the preset conditions, it can also be judged whether the GPU's computing power required by the access request is greater than this part of the computing power. When the GPU's computing power required by the access request is not greater than this part of the computing power, it is determined that the access request meets the preset conditions. For example, the back-end driver interception module determines that the first access request meets the preset conditions, including: the back-end driver interception module determines that the GPU's computing power required by the first access request is not greater than the first part of the computing power. The back-end driver interception module determines that the second access request meets the preset conditions, including: the back-end driver interception module determines that the GPU's computing power required by the second access request is not greater than the second part of the computing power.Among them, the server records the GPUs configured by the tenant for the virtual machines, as well as specifications such as the computing power of the GPUs, the computing power threshold, and the video memory threshold. Therefore, the backend driver interception module can obtain this information and determine whether the access request from the virtual machine meets the preset conditions based on this information.
[0167] Similarly, when there are also containers running in the virtual machine, after obtaining the access request, the front-end driver interception module set in the virtual machine can first determine whether the access request meets the preset conditions. If the access request meets the preset conditions, the access request is then provided to the backend driver interception module. If the access request does not meet the preset conditions, the virtual machine can return an error message to the container to indicate that the container cannot access the GPU. In a possible implementation, the virtual machine can determine, on a container-by-container basis, whether the video memory of the GPU required by the access request triggered by the application running in the container is greater than the video memory threshold; and / or whether the computing power of the GPU required by the access request is greater than the partial computing power or the computing power threshold of the GPU available for the VGPU accessed by the application. Here, the video memory threshold is the quota of the video memory configured by the tenant for the container, and the computing power threshold is the quota of the computing power configured by the tenant for the container.
[0168] First, the working principle of the front-end driver interception module in the virtual machine obtaining the access request triggered by the application running in the virtual machine and providing the access request to the backend driver interception module will be described below.
[0169] Generally, when an application needs to access a certain GPU, it can call the interface provided by the application middleware to access the GPU to trigger an access request for the GPU. In this application, the front-end driver interception module in the virtual machine performs device simulation on a part of the computing power of the GPU to generate a VGPU that can be accessed by the application running in the virtual machine. When the application calls the interface provided by the application middleware to access the GPU, it is actually accessing the VGPU simulated based on the GPU. Then, when the application calls the interface provided by the application middleware to access the VGPU, the access request is triggered. The VGPU can be regarded as a virtual driver file of the GPU set in the virtual machine. When the application calls the interface provided by the application middleware to access the VGPU, it is actually accessing the virtual driver file of the GPU. Then, when the application calls the interface provided by the application middleware to access the virtual driver file, the access request is triggered. In a possible implementation, the virtual driver file of the GPU is obtained by the front-end driver interception module using virtualization technology to simulate the GPU. Among them, when the virtual machine directly runs the application, the virtual driver file is actually set in the virtual machine. When the application runs in a container, the virtual driver file is actually set in the container.
[0170] To respond to virtual drive files, a virtual access interface for the drive file system is also set up in the virtual machine. This virtual access interface is used to respond to accesses to virtual drive files. After an application accesses a virtual drive file and triggers an access request, according to the file access logic, this access request will be processed by the drive file system. However, in this application, the front-end drive interception module is configured to intercept the virtual access interface of the drive file system. Therefore, after an application accesses the virtual drive file of the GPU, the front-end drive interception module can intercept the triggered access request. In a possible implementation, the virtual access interface of the drive file system is virtualized from the access interface of the drive file system using virtualization technology. When the virtual machine directly runs an application, the virtual access interface is actually set in the virtual machine. When the application runs in a container, the virtual access interface is actually set in the container. For example, when an application needs to access a VGPU, the application can call the interface provided by the application middleware. For example, for a certain manufacturer's GPU, the application can call cuda-related interfaces. By calling this interface, the application can open the virtual drive file of the corresponding GPU and perform file system access operations on it, such as ioctl operations or mmap operations, etc. When the application performs an ioctl operation on the virtual drive file of the GPU, an ioctl access request will be triggered, and the virtual access interface that responds to this ioctl access request is the virtual ioctl interface. The front-end drive interception module can intercept the ioctl access request by intercepting this virtual ioctl interface. When the application performs an mmap operation on the virtual drive file of the GPU, an mmap access request will be triggered, and the virtual access interface that responds to this mmap access request is the virtual mmap interface. The front-end drive interception module can intercept the mmap access request by intercepting this virtual mmap interface. Among them, the mmap access request is used to indicate mapping data to the video memory of the GPU. The ioctl access request is used to send control and configuration commands to the GPU.
[0171] After the front-end drive interception module intercepts the access request sent by the application, it needs to provide the access request to the back-end drive interception module. In a possible implementation, information is transmitted between the back-end drive interception module and the front-end drive interception module through one or more of the following methods: interrupt call or network connection.
[0172] The transmission of information from the sending end to the receiving end in the back-end driver interception module and the front-end driver interception module through interrupt calls means that the sending end sends an interrupt to the receiving end and carries the information to be transmitted in the interrupt. Taking the front-end driver interception module as the sending end and the back-end driver interception module as the receiving end as an example, since the virtual machines in the server are managed by the virtual machine manager in the server, the implementation process of the front-end driver interception module sending an access request to the back-end driver interception module through an interrupt call is as follows: the front-end driver interception module sends an interrupt to the virtual machine manager, and the virtual machine manager sends an interrupt to the back-end driver interception module. When the virtual machine manager is a virtual machine monitor (VMM), the interrupt call is a hypercall of the VMM. It should be noted that when transmitting information between the back-end driver interception module and the front-end driver interception module through hypercall, a new custom hypercall type needs to be added to support this transmission. When the back-end driver interception module is the sending end and the front-end driver interception module is the receiving end, for the implementation method of transmitting information between the two through interrupt calls, please refer to the implementation method when the front-end driver interception module is the sending end and the back-end driver interception module is the receiving end accordingly, which will not be elaborated here.
[0173] The transmission of information between the back-end driver interception module and the front-end driver interception module through a network connection means that a network connection is established between the back-end driver interception module and the front-end driver interception module, and the two transmit the information to be transmitted through the network. The type of network connection between the two can be determined according to application requirements, and the embodiments of the present application do not make specific limitations on it.
[0174] The working principle of the back-end driver interception module receiving an access request and forwarding the access request to the GPU driver module will be described below.
[0175] After the back-end driver interception module obtains the access request provided by the front-end driver interception module, it needs to send the access request to the real driver of the GPU that provides the VGPU for the application program, that is, the GPU driver module, so as to realize the access to the GPU. After the front-end driver interception module provides the access request to the back-end driver interception module through a specified transmission method, the back-end driver interception module can obtain the access request using the transmission method corresponding to the specified transmission method. For example, after the front-end driver interception module provides the access request to the back-end driver interception module through an interrupt call, the back-end driver interception module can obtain the access request by receiving the interrupt sent by the virtual management component. Another example is that after the front-end driver interception module sends the access request to the back-end driver interception module through the network, the back-end driver interception module can receive the access request through the network. It should be noted that the information transmission method used between the back-end driver interception module and the front-end driver interception module can be selected as pre-configured or determined through negotiation between the two.
[0176] After receiving the access request provided by the front-end driver interception module, when determining that the second access request meets the preset conditions, the back-end driver interception module may also first process the access request and then forward the access request to the GPU driver module. For example, after obtaining the access request provided by the front-end driver interception module, the back-end driver interception module schedules GPU resources for the virtual machine that initiated the access request based on the access request, and sends the access request to the GPU driver module according to the scheduled GPU resources. For example, when the access request of the virtual machine indicates accessing the GPU, after obtaining the access request, the back-end driver interception module schedules the video memory block to be allocated for the virtual machine in the video memory of the GPU, and then sends an access request carrying information indicating the video memory block to the GPU driver module. In another possible implementation, after obtaining the access request, the back-end driver interception module can also schedule the access request from the virtual machine to determine the timing of providing the access request to the GPU driver module. For example, the back-end driver interception module can schedule the access request according to scheduling methods such as task preemption, queuing, or a hybrid scheduling of multiple types to determine the order of sending multiple obtained access requests to the GPU driver module, and send the access requests to the GPU driver module in sequence according to this order. It should be noted that the back-end driver interception module can also perform other processing on the access request, and the embodiments of the present application do not make specific limitations on this.
[0177] There are multiple communication methods between the back-end driver interception module and the GPU driver module. By way of example, the back-end driver interception module and the GPU driver module may communicate via an interrupt. For example, after obtaining the access request, the back-end driver interception module sends an interrupt to the GPU driver module indicated by the access request to notify the GPU driver module of the access requirements indicated by the access request. It should be noted that the back-end driver interception module and the GPU driver module may also communicate in other ways, and communicating via an interrupt is an example, and the embodiments of the present application do not make specific limitations on their communication methods.
[0178] After the back-end driver interception module forwards the access request to the GPU driver module, the GPU driver module can then call the computing power of the GPU to process the access request. After the GPU driver module obtains the processing result of the access request by the GPU, it can transmit the processing result to the virtual machine along the reverse path of the path from the virtual machine to the GPU driver module for the access request. For example, the GPU driver module uses the communication method between it and the back-end driver interception module to provide the processing result to the back-end driver interception module. After the back-end driver interception module obtains the processing result, it uses the communication method between the back-end driver interception module and the front-end driver interception module to provide the processing result to the front-end driver interception module. After the front-end driver interception module obtains the processing result, the front-end driver interception module provides the processing result to the application program. Additionally, as Figures 3 to 7 shown, when the server includes a virtual machine manager, if information is transmitted between the back-end driver interception module and the front-end driver interception module through the virtual machine manager, the process for the back-end driver interception module to send the processing result to the front-end driver interception module is as follows: The back-end driver interception module provides the processing result to the virtual machine manager, and the virtual machine manager provides the processing result to the front-end driver interception module. For example, in response to information being transmitted between the back-end driver interception module and the front-end driver interception module through an interrupt call, the back-end driver interception module needs to send an interrupt indicating the processing result to the virtual machine manager, and the virtual machine manager sends an interrupt indicating the processing result to the front-end driver interception module. Among them, for the communication methods between the GPU driver module and the back-end driver interception module, between the back-end driver interception module and the front-end driver interception module, between the back-end driver interception module and the virtual machine manager, and between the virtual machine manager and the front-end driver interception module, please refer to the relevant descriptions in the previous content accordingly, and will not be elaborated here.
[0179] When the processing result includes a storage address, the back-end driver interception module and the front-end driver interception module also need to perform address conversion on the storage address so that the application can recognize the storage address. For example, in response to the processing result including the host physical address (HPA) of the video memory block allocated to the virtual machine in the video memory, the back-end driver interception module is also used to convert the host physical address into a guest physical address (GPA) and provide the guest physical address to the front-end driver interception module. The front-end driver interception module is also used to convert the guest physical address into a guest virtual address (GVA) and provide the guest virtual address to the application. For example, when the operation triggering the access request is an mmap operation, the processing result is the host physical address of the video memory block allocated to the virtual machine. However, since the address space of the video memory is different from the address space of the virtual machine, the host physical address in the GPU cannot be directly mmap-mapped to the guest virtual address in the virtual machine. Therefore, after obtaining the processing result, the back-end driver interception module needs to map the host physical address in the processing result to the host virtual address (HVA), then convert the host virtual address into a guest physical address, and provide the guest physical address to the front-end driver interception module. After obtaining the guest physical address, the front-end driver interception module also needs to convert the guest physical address into a guest virtual address to achieve shared access to the video memory of the virtual machine and the server. Among them, the virtual address is the address in the virtual address space used to load program data during the program running process. That is to say, the virtual address is the address allocated to the process of the application during the running process of the application. The virtual address can be mapped to a physical storage block. The data indicated by the virtual address is recorded on the physical storage block to which it is mapped. The physical address is the address of the physical storage block.
[0180] When the GPU allocates video memory blocks to virtual machines, the virtual machine manager also needs to provide virtual video memory to the virtual machines based on virtualization technology. The purpose of this virtualization technology is to provide a continuous physical video memory space starting from address 0 to the virtual machines, and to effectively isolate and schedule video memory resources among virtual machines. This virtualization technology mainly involves the conversion of guest virtual addresses -> guest physical addresses -> host virtual addresses -> host physical addresses. In virtualization technology, multiple virtual machines often run on a physical host, and each virtual machine believes that it exclusively occupies the video memory space of the physical host. Therefore, the virtual machine represents the video memory space it owns with a guest physical address, where this video memory space is considered continuous by the virtual machine (i.e., it can be understood that the virtual machine believes it has a complete physical video memory bar). The guest virtual address is an address formed by the operating system of the virtual machine mapping the guest physical address. The operating system of the virtual machine provides the guest virtual address to processes or application software set on the operating system of the virtual machine for use. The operating system of the virtual machine can obtain the mapping relationship from the guest virtual address to the guest physical address based on the mapping situation. In a possible implementation, the conversion from the guest virtual address to the guest physical address is implemented by the page table of the operating system of the virtual machine. The host physical address is the actual physical video memory address, and the host virtual address is an address formed by the operating system of the host mapping the host physical address. The operating system of the host provides the host virtual address to processes (such as virtual machines) on the operating system for use. The operating system of the host can obtain the mapping relationship from the host virtual address to the host physical address based on the mapping situation. In a possible implementation, the conversion from the host virtual address to the host physical address is implemented by the page table of the operating system of the host. Before the virtual machine manager simulates virtual video memory for a virtual machine, it needs to first determine the video memory blocks allocated to the virtual machine based on the specifications of the virtual machine, and then perform device simulation based on the video memory blocks allocated to the virtual machine to obtain virtual video memory. Therefore, the virtual machine manager can obtain the guest physical address of the virtual video memory and the host virtual address of the video memory block allocated to it, and obtain the mapping relationship between the host virtual address and the guest physical address based on this. In a possible implementation, the conversion from the host virtual address to the guest physical address is implemented by the page table of the operating system of the virtual machine. Therefore, this application can achieve the above conversion between the host physical address and the guest virtual address.
[0181] In Figure 8In this case, virtual machine 1 and virtual machine 2 are set in the same server (hereinafter referred to as the host). The virtual machine manager of the host sets the address range of the client physical address of virtual machine 1 to 0 - 5GB, which corresponds to the address ranges of the host physical addresses on the physical video memory of 1.5GB - 4.5GB and 6.5GB - 8.5GB. Moreover, the virtual machine manager of the host sets the address range of the client physical address of virtual machine 2 to 0 - 4GB, corresponding to the address ranges of the host physical addresses on the physical video memory of 9GB - 11GB and 13GB - 15GB. Therefore, virtual machine 1 exclusively uses the address range of 0 - 5GB of the client physical address, and virtual machine 2 exclusively uses the address range of 0 - 4GB of the client physical address. The address range of 0 - 5GB of the client physical address and the address range of 0 - 4GB of the client physical address can both be mapped to different address ranges of the host physical addresses on the physical video memory, thereby realizing the isolation of the virtual machine video memory.
[0182] The implementation manner of the server system provided by this application based on cloud computing technology will be introduced below.
[0183] Figure 9 It is a schematic structural diagram of a server system provided by an embodiment of this application. As Figure 9 shown, the server system includes a first server and a second server. The first server is provided with a first back - end driver interception module, a first GPU driver module, and a first GPU. The first virtual machine and the second virtual machine are running on the second server. A connection channel is established between the first server and the second server. In a possible implementation manner, the connection channel is implemented through a high - speed interconnection protocol, and the high - speed interconnection protocol includes a Compute Express Link (CXL) protocol or a Lingqu bus (also known as UB bus) protocol. It should be noted that different servers can also communicate and connect through other implementation manners, and the embodiments of this application do not make specific limitations on it.
[0184] The first virtual machine is provided with a first front - end driver interception module and a first application program. The first front - end driver interception module is used to perform device simulation on a first part of the computing power of the first GPU to generate a first VGPU that can be accessed by the first application program. The first application program is used to access the first VGPU. The first front - end driver interception module is also used to obtain a first access request of the first application program for the first VGPU and send the first access request to the connection channel. Among them, the first application program is used to implement the user's business.
[0185] The second virtual machine is provided with a second front-end driver interception module and a second application. The second front-end driver interception module is used to perform device simulation on the second part of the computing power of the first GPU to generate a second VGPU that can be accessed by the second application. The second application is used to access the second VGPU. The second front-end driver interception module is further used to obtain a second first access request of the second application for the second VGPU and send the first access request to the connection channel. Among them, the second application is used to implement the user's service.
[0186] The first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module and / or the second first access request from the second front-end driver interception module from the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions, and / or, send the second first access request to the first GPU driver module when it is determined that the second first access request meets the preset conditions.
[0187] The first GPU driver module is used to call the first part of the computing power of the first GPU to process the first access request, and / or, call the second part of the computing power of the first GPU to process the second first access request.
[0188] It should be noted that the fact that the first virtual machine and the second virtual machine are running on the second server is only an example. The number of virtual machines deployed in the second server can be determined based on application requirements, and the embodiments of the present application do not make specific limitations on this. For example, fewer or more virtual machines can be deployed in the second server, such as at least one virtual machine is deployed in the second server. When at least one virtual machine is deployed in the second server, for the access request of the application program set in any virtual machine to access the virtual hardware obtained based on the hardware, it can be transmitted to the first back-end driver interception module with reference to the method provided in the present application, and transmitted to the driver of the hardware through the first back-end driver module, so that the driver of the hardware calls the capabilities of the hardware to process the access request.
[0189] After the first GPU driver module calls the computing power of the first GPU to process the access request, if it receives the processing result generated by the first GPU processing the access request, the first GPU driver module is further used to provide the processing result to the first back-end driver interception module, so as to return the processing result to the application program that triggered the access request through the first back-end driver interception module. The roles played by each module in the server during the process of returning the processing result are as follows:
[0190] The first GPU driver module is used to obtain the first part of the computing and processing capabilities of the first GPU to process the first processing result generated by the first access request, and / or obtain the second part of the computing and processing capabilities of the first GPU to process the second first access request to generate the second processing result, and send the first processing result and / or the second processing result to the first back-end driver interception module.
[0191] The first back-end driver interception module is used to obtain the first processing result and / or the second processing result from the first GPU driver module, send the first processing result to the connection channel, and / or send the second processing result to the connection channel.
[0192] The first front-end driver interception module is used to obtain the first processing result from the connection channel from the first back-end driver interception module and provide the first processing result to the first application.
[0193] The second front-end driver interception module is used to obtain the second processing result from the connection channel from the first back-end driver interception module and provide the second processing result to the second application.
[0194] Exemplarily, when the first application in the first virtual machine needs to use the first GPU for rendering operations, the first application will access the first VGPU, and this access operation will trigger a first access request indicating access to the first GPU. The first access request carries relevant data of the image to be rendered. The first front-end driver interception module can obtain the first access request and forward the first access request to the first back-end driver interception module. After receiving the first access request, if the first access request meets the preset conditions, the first back-end driver interception module forwards the first access request to the first GPU driver module. After receiving the first access request, the first GPU driver module calls the first part of the computing capabilities of the first GPU to process the first access request. After obtaining the processing result of the first GPU for the first access request, the first GPU driver module forwards the processing result to the first back-end driver interception module. The processing result carries the video memory address of the rendering result, and the rendering result is the result of the first GPU rendering based on the relevant data carried by the first access request. After receiving the processing result, the first back-end driver interception module forwards the processing result to the first front-end driver interception module. After receiving the processing result, the first front-end driver interception module sends the processing result to the first application so that the first application can obtain the rendering result from the video memory address based on the video memory address carried by the processing result.
[0195] It can be seen from this that the access request triggered by the application program running in the virtual machine in the second server to access the VGPU can be sequentially transmitted to the first GPU driver module in the first server via the front-end driver interception module in the virtual machine and the first back-end driver interception module in the first server. The processing result generated by the first GPU in the first server for processing the access request can be sequentially transmitted to the application program in the virtual machine via the first GPU driver module, the first back-end driver interception module, and the front-end driver interception module in the virtual machine. Since both of these transmission processes are implemented in software, the dependence of the process of the virtual machine in the second server calling the first GPU in the first server to process the access request on implementation details such as the hardware information of the GPU can be reduced, and even the implementation details such as the hardware information of the GPU do not need to be relied on. For example, it does not rely on the hardware information of the GPU manufacturer, and the decoupling from the GPU can be achieved, improving the flexibility of the application program in the virtual machine to access the GPU. Moreover, when the virtual machine uses part of the computing power of the GPU, it is equivalent to virtualizing and partitioning the GPU. The software implementation method of transmitting the access request and the processing result in this application enables the partitioning specification of virtualizing and partitioning the GPU to break through the hardware control of the GPU and support various scheduling strategies for scheduling the GPU, improving the flexibility of scheduling the GPU.
[0196] In this application, there are multiple implementation methods for the server system. The following takes the following three implementation methods as examples to illustrate it.
[0197] In the first implementation method, the second server is not provided with a back-end driver interception module, a GPU driver module, and a GPU. As Figure 9 shown, the scenario presented by this implementation method is that the virtual machine deployed inside the second server accesses the GPU owned by the first server. Exemplarily, the scenario of this implementation method can be the scenario where the first server provides hardware resources for the first server. For example, this scenario is a pooling scenario for pooling GPUs. In this pooling scenario, multiple servers are deployed in the resource pool, each server has a GPU, a GPU driver module, and a back-end driver interception module, and the multiple servers include the first server. After the second server applies for resources from the resource pool, the first server in the resource pool provides the GPU for the second server. This implementation method can break through the limitation that the virtual machine deployed in the current server can only access the GPU owned by the server itself, and further improve the flexibility of accessing the GPU and the flexibility of scheduling the GPU.
[0198] This first implementation method can be further divided into multiple cases. The following takes the following several cases as examples to illustrate it.
[0199] In the first case, the first application runs directly in the first virtual machine, and the second application runs directly in the second virtual machine. As Figure 9 shown, the scenario presented by this implementation method is that the first application running directly in the virtual machine deployed inside the second server accesses the first GPU of the first server. Since the operations executed by the virtual machine in the second server are managed by the hypervisor, and the first GPU, the first GPU driver module, and the first back-end driver interception module in the first server are managed by the operating system (OS) of the first server, this scenario can be called a scenario for GPU access implemented through cross-server Hypervisor and OS.
[0200] In this first implementation method, the VGPU used by the virtual machine deployed in the second server is obtained based on a partial computing power of the first GPU in the first server, which is equivalent to virtualizing and partitioning the first GPU of the first server. In this application, through the first back-end driver interception module and the front-end driver interception module in the second server, the first access request and the processing result are forwarded between the first GPU driver module and the application running in the virtual machine in the second server, reducing the degree of dependence on implementation details such as the hardware information of the GPU. Even without relying on implementation details such as the hardware information of the GPU, the flexibility of the application in the virtual machine to access the GPU is improved, and the server of this application supports multiple scheduling strategies for GPU scheduling, improving the flexibility of GPU scheduling.
[0201] In the second case, a first container is set up in the first virtual machine, the first application runs in the first container, and the first container is used to mount the first VGPU for the first application to access the first VGPU in the first container. A second container is set up in the second virtual machine, the second application runs in the second container, and the second container is used to mount the second VGPU for the second application to access the second VGPU in the second container. Figure 10 It is a schematic diagram of a server system provided by an embodiment of this application. As Figure 10 shown, the scenario presented by this implementation method is that a virtual machine is deployed in the second server, a container is deployed in the virtual machine, and the application in the container has a need to access the first GPU in the first server that provides the VGPU for it. Since the operations executed by the virtual machine in the second server are managed by the hypervisor, and the first GPU, the first GPU driver module, and the first back-end driver interception module in the first server are managed by the OS of the first server, this scenario can be called a GPU access scenario implemented through cross-server Hypervisor, OS, and container.
[0202] Optionally, the number of containers set in the virtual machine can be adjusted according to application requirements. For example, one container can be set in the virtual machine, or multiple containers can be set in the virtual machine. By way of example, as Figure 11 shown, a third container is further set in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU that can be accessed by the third application. For the implementation method of the third application accessing the third VGPU, please refer to the implementation methods of the first application accessing the first VGPU and the second application accessing the second VGPU respectively, which will not be elaborated here. Similarly, a fourth container is further set in the second virtual machine, and a fourth application runs in the fourth container. The fourth container is used to mount a fourth VGPU for the fourth application to access the fourth VGPU in the fourth container. Among them, the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU that can be accessed by the fourth application. For the implementation method of the fourth application accessing the fourth VGPU, please refer to the implementation methods of the first application accessing the first VGPU and the second application accessing the second VGPU respectively, which will not be elaborated here.
[0203] This second case not only has the effects of the first case, but also further partitions the virtual machines in the second server on the basis of virtualizing and partitioning the first GPU in the first server, realizing the secondary partitioning and comprehensive scheduling of the virtual machines in the second server after the first GPU resource is partitioned, and further improving the flexibility of accessing the GPU and the flexibility of scheduling the GPU. It should be noted that the container running in the virtual machine is only an example, and other types of virtual instances can also run in the virtual machine. As is known to those of ordinary skill in the art, with the change of business requirements, the type of virtual instance running in the virtual machine can be adjusted according to application requirements, and the embodiments of the present application do not make specific limitations on this. For example, the way of deploying virtual instances in the server can be further evolved on this basis.
[0204] In the second implementation manner, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. At this time, the access of the application program running in the virtual machine in the second server to the VGPU obtained based on the second GPU is actually the access of the application program running in the virtual machine in the server to the VGPU obtained from the GPU on the server itself. For the implementation process, please refer to the implementation process in the foregoing server accordingly. For example, the first front-end driver interception module of the first virtual machine is used to perform device simulation on the first part of the computing power of the second GPU to generate a third VGPU accessible to the first application program. The first application program is used to access the third VGPU. The first front-end driver interception module is further used to obtain a third first access request of the first application program for the third VGPU. The second front-end driver interception module of the second virtual machine is further used to perform device simulation on the second part of the computing power of the second GPU to generate a fourth VGPU accessible to the second application program. The second application program is used to access the fourth VGPU. The second front-end driver interception module is further used to obtain a fourth first access request of the second application program for the fourth VGPU. The second back-end driver interception module is used to obtain the third first access request from the first front-end driver interception module and / or the fourth first access request from the second front-end driver interception module, and send the third first access request to the second GPU driver module when it is determined that the third first access request meets the preset conditions, and / or send the fourth first access request to the second GPU driver module when it is determined that the fourth first access request meets the preset conditions. The second GPU driver module is used to call the first part of the computing power of the second GPU to process the third first access request, and / or call the second part of the computing power of the second GPU to process the fourth first access request.
[0205] In this second implementation manner, the deployment manners of virtual machines and application programs in the second server can also be divided into multiple cases. This application takes the following cases as examples to illustrate it. In the first case, the first application program runs directly in the first virtual machine, and the second application program runs directly in the second virtual machine. In the second case, a first container is set in the first virtual machine, and the first application program runs in the first container. The first container is used to mount a first VGPU for the first application program to access the first VGPU in the first container. A second container is set in the second virtual machine, and the second application program runs in the second container. The second container is used to mount a second VGPU for the second application program to access the second VGPU in the second container. Moreover, the number of containers set in the virtual machine can be adjusted according to application requirements. For example, one container can be set in the virtual machine, or multiple containers can be set in the virtual machine. Exemplarily, a third container is further set in the first virtual machine, and the third application program runs in the third container. The third container is used to mount a third VGPU for the third application program to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application program. Similarly, a fourth container is further set in the second virtual machine, and the fourth application program runs in the fourth container. The fourth container is used to mount a fourth VGPU for the fourth application program to access the fourth VGPU in the fourth container. Among them, the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application program. For the implementation manners of multiple cases in this second implementation manner, please refer to the relevant descriptions above accordingly, and details are not described here again. In addition, the first server can also deploy virtual machines, and the application programs running in the virtual machines can also access the first GPU or the second GPU. Moreover, the deployment manners of virtual machines and application programs in the first server and the manner of accessing the GPU can all refer to the relevant descriptions of the deployment manners of virtual machines and application programs in the second server and the manner of accessing the GPU accordingly, and details are not described here either.
[0206] It should be noted that the above two implementation manners of the server can be used alone or in combination. For example, when the server system includes multiple servers, each server in the multiple servers can be deployed according to the first implementation manner described above, or each server in the multiple servers can be deployed according to the second implementation manner described above, or some servers in the multiple servers are deployed according to the first implementation manner and some servers are deployed according to the second implementation manner described above. Similarly, when there are multiple virtual machines running in any server, the virtual machines in the multiple virtual machines can be deployed according to any one of multiple situations in the above two implementation manners. For example, an application program is directly run in a virtual machine, and containers are set in some virtual machines, and application programs are run in the containers. Figure 12 is a schematic diagram of a server system provided by an embodiment of the present application. As Figure 12 shown, the server system includes a first server and a second server having a connection channel. The first server is provided with a first back-end driver interception module, a first GPU driver module, and a first GPU. The second server is provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first virtual machine in the second server directly runs a first application program, a second container is set in the second virtual machine in the second server, and a second application program is run in the second container. The fifth virtual machine in the first server directly runs a fifth application program, and a fifth front-end driver interception module is set in the fifth virtual machine.
[0207] In this server system, the way for the backend driver interception module to forward an access request to the GPU driver module can also refer to the way for the candidate driver interception module in the aforementioned server to forward an access request to the GPU driver. For example, when the backend driver interception module determines that the access request meets the preset conditions, it forwards the access request to the GPU driver module. When the backend driver interception module determines that the access request does not meet the preset conditions, it does not forward the access request to the GPU driver module. At this time, the backend driver interception module can return an error message to the virtual machine to which the access request belongs to indicate that the virtual machine cannot access the GPU. In a possible implementation, the first backend driver interception module determines that the first access request meets the preset conditions, including: the first backend driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or the first backend driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold. The second backend driver interception module determines that the second first access request meets the preset conditions, including: the second backend driver interception module determines that the video memory of the GPU required by the second first access request is not greater than the second video memory threshold; and / or the second backend driver interception module determines that the computing power of the GPU required by the second first access request is not greater than the second partial computing power or the second computing power threshold. For the implementation process of this process, please refer to the relevant description in the aforementioned server accordingly, and it will not be elaborated here.
[0208] Similarly, when there are also containers running in the virtual machine, after obtaining the access request, the front-end driver interception module set in the virtual machine can first optionally determine whether the access request meets the preset conditions. If the access request meets the preset conditions, it then provides the access request to the backend driver interception module. If the access request does not meet the preset conditions, the virtual machine can return an error message to the container to indicate that the container cannot access the GPU. In a possible implementation, the virtual machine can optionally determine, in units of containers, whether the video memory of the GPU required by the access request triggered by the application running in the container is greater than the video memory threshold; and / or whether the computing power of the GPU required by the access request is greater than the partial computing power or the computing power threshold of the GPU available for the VGPU accessed by the application. Here, the video memory threshold is the quota of the video memory configured by the tenant for the container, and the computing power threshold is the quota of the computing power configured by the tenant for the container.
[0209] In this server system, an application triggers an access request. The front-end driver interception module obtains the access request. The back-end driver interception module processes the received access request and forwards the access request to the GPU driver module. The GPU driver module calls the computing power of the GPU to process the access request. The GPU driver module transfers the processing result to the virtual machine along the reverse path of the path from the virtual machine to the GPU driver module for the access request. When the processing result includes a storage address, the implementation process of performing address conversion on the storage address can refer to the relevant description in the previous server, and it will not be elaborated here.
[0210] Different from the implementation method of the previous server, information is transmitted between the first back-end driver interception module and the front-end driver interception module in the second server through one or more of the following methods: shared memory or network connection. Here, an example is given to illustrate the information transmission between the first back-end driver interception module and the first front-end driver interception module in the second server.
[0211] The first back-end driver interception module and the sending end in the first front-end driver interception module transmit information to the receiving end therein through shared memory, which means that a shared memory for transmitting information is pre-configured for the sending end and the receiving end. When the sending end needs to transmit information to the receiving end, it first stores the information to be transmitted in the shared memory. When the receiving end determines that there is data update in the shared memory, it obtains the updated data from the shared memory to obtain the information to be transmitted. Among them, the shared memory can be configured on the server where any one of the first back-end driver interception module and the first front-end driver interception module is located, or configured on a remote server other than the servers where the first back-end driver interception module and the first front-end driver interception module are located. It should be noted that there are various implementation methods for the shared memory, and those skilled in the art of this application can select the implementation method to be used according to needs, and no specific examples will be given here.
[0212] The first back-end driver interception module and the first front-end driver interception module transmit information through a network connection, which means that a network connection is established between the first back-end driver interception module and the first front-end driver interception module, and the two transmit the information to be transmitted through the network. The type of network connection between the two can be determined according to application requirements, and no specific limitation is made in this embodiment of the application.
[0213] In addition, the second server further includes: a virtual machine manager. The first back-end driver interception module and the front-end driver interception module in the second server can transmit information through the virtual machine manager. For example, the virtual machine manager is used to obtain the first access request from the first front-end driver interception module, send the first access request to the connection channel, obtain the second access request from the second front-end driver interception module, and send the second access request to the connection channel. The virtual machine manager is used to obtain the first processing result from the first back-end driver interception module from the connection channel, and send the first processing result to the first front-end driver interception module, so that the first front-end driver interception module provides the first processing result to the first application program, and obtain the second processing result from the first back-end driver interception module from the connection channel, and send the second processing result to the second front-end driver interception module, so that the second front-end driver interception module provides the second processing result to the second application program.
[0214] The following uses three examples to illustrate the implementation process of an application accessing the GPU.
[0215] As Figure 3 shown, a first virtual machine is deployed in the server, and the GPU to be accessed is the GPU in the server. The implementation process of the first application program in the first virtual machine accessing the GPU includes the following steps:
[0216] S11. The first front-end driver interception module in the first virtual machine virtualizes the access interface of the driver file system in the first virtual machine to obtain a virtual access interface of the driver file system, and performs device simulation on the first part of the computing power of the GPU to generate a virtual driver file (i.e., the first VGPU) that can be accessed by the first application program.
[0217] S12. The first application program runs in the first virtual machine. The first application program accesses the virtual driver file of the GPU by calling the interface provided by the application middleware, triggering an access request indicating access to the GPU.
[0218] S13. The first front-end driver interception module intercepts the virtual access interface, obtains the access request, and sends the access request to the back-end driver interception module by means of hypercall.
[0219] S14. After receiving the access request passed by hypercall, the Hypervisor in the server sends the access request to the back-end driver interception module.
[0220] S15. After receiving the access request, the back-end driver interception module first performs GPU global resource monitoring and scheduling based on the access request, and then forwards the access request to the GPU driver module in the server.
[0221] S16. The GPU driver module in the server accesses the GPU based on this access request to invoke some computing capabilities of the GPU to process the access request. After the GPU driver module obtains the processing results generated by using some computing and processing capabilities of the GPU to process the access request, it transmits the processing results to the first application program along the reverse path of the above path.
[0222] As Figure 5 shown, a first virtual machine and a second virtual machine are deployed in the server, a first container is deployed in the first virtual machine, a second container is deployed in the second virtual machine, and the accessed GPU is the GPU in the server. The implementation process of the first application program in the first container accessing the GPU includes the following steps:
[0223] S21. The first front-end driver interception module in the first virtual machine virtualizes the access interface of the drive file system in the first virtual machine to obtain a virtual access interface of the drive file system.
[0224] S22. When the container runtime creates the first container, it performs device simulation on the first part of the computing capabilities of the GPU to generate a virtual drive file of the GPU (i.e., the first VGPU) that can be accessed by the first application program, and embeds (mounts) this virtual drive file into the first container.
[0225] S23. The first application program runs in the first container. The first application program makes a system access to the virtual drive file of the GPU in the container by calling the interface provided by the application middleware, triggering an access request indicating access to the GPU.
[0226] S24. The first front-end driver interception module intercepts the virtual access interface, obtains the access request, performs resource monitoring and scheduling in the GPU virtual machine, and then sends this access request to the back-end driver interception module through the hypercall method.
[0227] S25. After receiving the access request passed by the hypercall, the Hypervisor in the server sends the access request to the back-end driver interception module.
[0228] S26. After receiving this access request, the back-end driver interception module first performs global resource monitoring and scheduling of the GPU based on this access request, and then forwards this access request to the GPU driver module in the server.
[0229] S27. The GPU driver module in the server accesses the GPU based on this access request to invoke some computing capabilities of the GPU to process the access request. After the GPU driver module obtains the processing results generated by using some computing and processing capabilities of the GPU to process the access request, it transmits the processing results to the first application program along the reverse path of the above path.
[0230] As Figure 9 shown, the server system includes a first server and a second server. The first server and the second server are connected through a high-speed interconnection protocol such as CXL / UB. A first virtual machine is deployed in the second server, and the accessed GPU is the first GPU in the first server. The implementation process of the first virtual machine accessing the first GPU includes the following steps:
[0231] S31. The first front-end driver interception module in the first virtual machine virtualizes the access interface of the drive file system in the first virtual machine to obtain a virtual access interface of the drive file system, and performs device simulation on the first part of the computing power of the first GPU to generate a virtual drive file of the first GPU (i.e., the first VGPU) that can be accessed by the first application.
[0232] S32. The first application runs in the first virtual machine. The first application accesses the virtual drive file of the first GPU through the interface provided by the application middleware, triggering an access request indicating access to the first GPU.
[0233] S33. The first front-end driver interception module intercepts the virtual access interface, obtains the access request, and uses the CXL bus connected to the second server to access the memory shared by the first server to the second server, and stores the access request in the memory.
[0234] S34. The first back-end driver interception module in the first server accesses the access request obtained by accessing the memory. After the first back-end driver interception module performs GPU global resource monitoring and scheduling based on the access request, it forwards the access request to the first GPU driver module in the first server.
[0235] S35. The first GPU driver module accesses the first GPU based on the access request to call a part of the computing power of the first GPU to process the access request. After the first GPU driver module obtains the processing result generated by processing the access request with a part of the computing processing power of the first GPU, it transmits the processing result to the first application along the reverse path of the above path.
[0236] In an implementation manner, the functions of the back-end driver interception module and the front-end driver interception module can be implemented by multiple functional units. Exemplarily, such as Figure 13As shown, the functions of the front-end driver interception module are implemented through the device simulation and interception system access logic unit, the front-end data transmission logic unit, and the front-end video memory mapping logic unit. The device simulation and interception system access logic unit is used to obtain the virtual access interface of the driver file system and the virtual driver file of the GPU by using virtualization technology, intercept access requests, provide the access requests to the front-end data transmission logic unit, and provide the processing results to the application program. The front-end data transmission logic unit is used to transmit the access requests to the back-end data transmission logic unit and provide the processing results obtained from the back-end data transmission logic unit to the device simulation and interception system access logic unit. The front-end video memory mapping logic unit is used to convert the client physical address into the client virtual address and provide the client virtual address to the application program. The functions of the back-end driver interception module are implemented through the back-end data transmission logic unit, the back-end video memory mapping logic unit, and the resource access monitoring and scheduling logic unit. The back-end data transmission logic unit is used to receive the access requests transmitted by the front-end data transmission logic unit, forward the access requests to the resource access monitoring and scheduling logic unit, receive the processing results provided by the resource access monitoring and scheduling logic unit, and provide the processing results to the front-end data transmission logic unit. The back-end video memory mapping logic unit is used to convert the host physical address into the client physical address and provide the client physical address to the front-end driver interception module. The resource access monitoring and scheduling logic unit processes the access requests, and then, after the processing results indicate that the access requests need to be forwarded to the GPU driver module, forwards the access requests to the GPU driver module, and after receiving the processing results provided by the GPU driver module, provides the processing results to the back-end data transmission logic unit.
[0237] It should be noted that compared with the existing method of intercepting the API in the industry, the present application intercepts the file system access interface at the driver level, and the interception levels are different. Moreover, the access requests intercepted by the present application are transmitted from the virtual machine to the accessed hardware, which does not occupy the tenant network and will not affect the network used by the tenant. In addition, the present application can also be applied to the access of other types of resources. For example, the accessed object can be not only physical resources but also virtual resources (such as VGPUs). At the same time, the charging method for resources can also be flexibly selected. For example, it can be charged according to the segmented resource volume, or it can be selected to be charged according to the actual usage amount of the resources. The embodiments of the present application do not make specific limitations on this.
[0238] As can be seen from the above, the improvement of the server system of the present application is mainly reflected in the access process of the GPU by the virtual machines in the server and the server system. Corresponding to the foregoing server, the embodiment of the present application further provides a method for creating a virtual machine based on cloud computing technology. This method is applied to a cloud management platform, which is used to manage the infrastructure, and the infrastructure includes multiple servers. The implementation process of this method for creating a virtual machine based on cloud computing technology will be described below.
[0239] Figure 14 It is a flowchart of a method for creating a virtual machine based on cloud computing technology provided by an embodiment of the present application. As Figure 14 shown, this method for creating a virtual machine based on cloud computing technology includes the following steps:
[0240] Step 1401: Obtain a virtual machine creation request input by a tenant. The virtual machine creation request carries the specification of the first VGPU of the first virtual machine to be created.
[0241] When a tenant needs to create a virtual machine based on the infrastructure managed by the cloud management platform, the tenant can perform a specified operation on the client used by the tenant to trigger a virtual machine creation request, so that the cloud management platform creates a virtual machine for the tenant under the indication of this virtual machine creation request. In a possible implementation manner, the cloud management platform can provide a virtual machine creation interface to the tenant, and the tenant can trigger a virtual machine creation request based on this virtual machine creation interface. The virtual machine creation request carries the specification of the first virtual machine to be created. After the tenant triggers the virtual machine creation request, the cloud management platform can obtain the virtual machine creation request through this virtual machine creation interface and obtain the specification of the first virtual machine to be created from this virtual machine creation request.
[0242] Exemplarily, the virtual machine creation interface is implemented through one or more of the following: application programming interface (API), interaction template, and configuration interface. Among them, the interaction template is a template provided by the cloud management platform to the tenant for implementing different functions. When a tenant needs to use a certain function, the tenant can download the template for implementing this function, add the relevant information of the tenant in this template, and then feedback the template added with the relevant information of the tenant to the cloud management platform. After receiving the template added with the relevant information of the tenant, the cloud management platform can obtain the function that this template needs to implement and customize the implementation of this function according to the information of this tenant. The configuration interface means that the tenant can operate in this configuration interface to indicate the function that the tenant needs to implement.
[0243] Step 1402: Select a target server from multiple servers that can provide the specification of the first VGPU.
[0244] After the cloud management platform obtains the specification of the first VGPU of the first virtual machine to be created, it can select a server in the infrastructure that can provide the specification of the first VGPU to obtain the target server.
[0245] Step 1403: Create the first virtual machine on the target server.
[0246] After the cloud management platform selects the target server that can provide the specification of the first VGPU in the infrastructure, it can create a virtual machine that conforms to the specification of the first VGPU in the target server. Moreover, in order to ensure that the application program running on the first virtual machine can access the first VGPU in the manner provided by this application, it is also necessary to set a first front-end driver interception module and a first application program in the first virtual machine, and use the first front-end driver interception module to virtualize the access interface of the drive file system in the first virtual machine to obtain a virtual access interface of the drive file system, and perform device simulation on the first part of the computing power of the GPU of the target server to generate a virtual drive file (i.e., the first VGPU) that can be accessed by the first application program. The parameters of the first VGPU match the specification of the first VGPU. For example, when the specification of the first VGPU indicates multiple parameters (such as computing power, video memory size, video memory bit width, and video memory bandwidth), the multiple parameters of the first VGPU obtained by device virtualization correspond one by one to the multiple parameters indicated by the specification, and any one of the multiple parameters of the first VGPU is equal to or slightly greater than the corresponding parameter indicated by the specification. At the same time, if the target server is not provided with a back-end driver interception module, it is also necessary to configure a back-end driver interception module in the target server. Among them, the first front-end driver interception module is also used to obtain the first access request of the first application program for the first VGPU. The back-end driver interception module is used to obtain the first access request from the first front-end driver interception module and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions. Correspondingly, the GPU driver module is used to call the first part of the computing power of the GPU to process the first access request.
[0247] In a possible implementation manner, a first container is set in the first virtual machine, the first application program runs in the first container, and the first container is used to mount the first VGPU for the first application program to access the first VGPU in the first container.
[0248] In a possible implementation manner, a third container is also set in the first virtual machine, the third application program runs in the third container, and the third container is used to mount the third VGPU for the third application program to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the GPU to generate a third VGPU that can be accessed by the third application program.
[0249] In a possible implementation, the first application runs directly in the first virtual machine.
[0250] In a possible implementation, the GPU driver module is further configured to obtain a first processing result generated by processing a first access request using a first part of the computing processing power of the GPU, and send the first processing result to the backend driver interception module; the backend driver interception module is further configured to obtain the first processing result from the GPU driver module and send the first processing result to the first front-end driver interception module; the first front-end driver interception module is further configured to obtain the first processing result from the backend driver interception module and provide the first processing result to the first application.
[0251] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the backend driver interception module.
[0252] In a possible implementation, the backend driver interception module determines that the first access request meets a preset condition, including: the backend driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or, the backend driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first part of the computing power or the first computing power threshold.
[0253] In a possible implementation, a second virtual machine is further set on the target server. A second front-end driver interception module and a second application are set in the second virtual machine. The second front-end driver interception module is configured to perform device simulation on a second part of the computing power of the GPU to generate a second VGPU accessible by the second application. The second application is configured to access the second VGPU. The second front-end driver interception module is further configured to obtain a second access request of the second application for the second VGPU; the backend driver interception module is configured to obtain the second access request from the second front-end driver interception module and send the second access request to the GPU driver module when it determines that the second access request meets the preset condition; the GPU driver module is configured to call a second part of the computing power of the GPU to process the second access request.
[0254] In a possible implementation, a second container is set in the second virtual machine. The second application runs in the second container. The second container is configured to mount the second VGPU for the second application to access the second VGPU in the second container.
[0255] In a possible implementation, the second application runs directly in the second virtual machine.
[0256] In a possible implementation, the GPU driver module is further configured to obtain a second processing result generated by processing a second access request using a second part of the computing processing capability of the GPU, and send the second processing result to the backend driver interception module; the backend driver interception module is further configured to obtain the second processing result from the GPU driver module and send the second processing result to the second front-end driver interception module; the second front-end driver interception module is further configured to obtain the second processing result from the backend driver interception module and provide the second processing result to the second application.
[0257] In a possible implementation, the target server further includes:
[0258] The virtual machine manager is configured to obtain the second access request from the second front-end driver interception module and send the second access request to the backend driver interception module.
[0259] In a possible implementation, the backend driver interception module determines that the second access request meets the preset conditions, including:
[0260] The backend driver interception module determines that the video memory of the GPU required by the second access request is not greater than the second video memory threshold; and / or, the backend driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0261] The method for creating a virtual machine based on cloud computing technology in the embodiments of the present application is introduced above. Corresponding to the above method, the embodiments of the present application also provide a cloud management platform. The cloud management platform is used to manage the infrastructure. The infrastructure includes multiple servers. Figure 15 It is a schematic structural diagram of a cloud management platform provided by an embodiment of the present application. Based on Figure 15 the following multiple components shown, the Figure 15 shown cloud management platform can perform all or part of the operations shown above. Figure 14 It should be understood that the device may include more additional components than those shown or omit some of the components shown. The embodiments of the present application do not limit this. As Figure 15 shown, the cloud management platform 150 may include:
[0262] An obtaining unit 1501, configured to obtain a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of the first VGPU of the first virtual machine to be created.
[0263] A selection unit 1502, configured to select a target server that can provide the specification of the first VGPU from multiple servers.
[0264] A creation unit 1503, configured to create the first virtual machine on the target server.
[0265] Among them, the target server is provided with a backend driver interception module, a GPU driver module, and a GPU. The first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on the first part of the computing power of the GPU to generate a first VGPU that can be accessed by the first application. The first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request from the first application for the first VGPU. The backend driver interception module is used to obtain the first access request from the first front-end driver interception module and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions. The GPU driver module is used to call the first part of the computing power of the GPU to process the first access request.
[0266] In a possible implementation, the first virtual machine is provided with a first container, and the first application runs in the first container. The first container is used to mount the first VGPU for the first application to access the first VGPU in the first container.
[0267] In a possible implementation, the first virtual machine is further provided with a third container, and a third application runs in the third container. The third container is used to mount a third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the GPU to generate a third VGPU that can be accessed by the third application.
[0268] In a possible implementation, the first application runs directly in the first virtual machine.
[0269] In a possible implementation, the GPU driver module is further used to obtain a first processing result generated by processing the first access request with the first part of the computing processing power of the GPU, and send the first processing result to the backend driver interception module. The backend driver interception module is further used to obtain the first processing result from the GPU driver module and send the first processing result to the first front-end driver interception module. The first front-end driver interception module is further used to obtain the first processing result from the backend driver interception module and provide the first processing result to the first application.
[0270] In a possible implementation, the target server further includes: a virtual machine manager, which is used to obtain the first access request from the first front-end driver interception module and send the first access request to the backend driver interception module.
[0271] In a possible implementation, the backend driver interception module determines that the first access request meets the preset conditions, including: the backend driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or, the backend driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0272] In a possible implementation, a second virtual machine is also set up on the target server. A second front-end driver interception module and a second application are set up in the second virtual machine. The second front-end driver interception module is used to perform device simulation on the second partial computing power of the GPU to generate a second VGPU for the second application to access. The second application is used to access the second VGPU. The second front-end driver interception module is also used to obtain a second access request from the second application for the second VGPU; the backend driver interception module is used to obtain the second access request from the second front-end driver interception module and send the second access request to the GPU driver module when it determines that the second access request meets the preset conditions; the GPU driver module is used to call the second partial computing power of the GPU to process the second access request.
[0273] In a possible implementation, a second container is set up in the second virtual machine. The second application runs in the second container. The second container is used to mount the second VGPU for the second application to access the second VGPU in the second container.
[0274] In a possible implementation, the second application runs directly in the second virtual machine.
[0275] In a possible implementation, the GPU driver module is also used to obtain a second processing result generated by processing the second access request with the second partial computing power of the GPU, and send the second processing result to the backend driver interception module; the backend driver interception module is also used to obtain the second processing result from the GPU driver module and send the second processing result to the second front-end driver interception module; the second front-end driver interception module is also used to obtain the second processing result from the backend driver interception module and provide the second processing result to the second application.
[0276] In a possible implementation, the target server further includes: a virtual machine manager, which is used to obtain the second access request from the second front-end driver interception module and send the second access request to the backend driver interception module.
[0277] In a possible implementation, the back-end driver interception module determines that the second access request meets the preset conditions, including: the back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than the second video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second partial computing power or the second computing power threshold.
[0278] Here, for the detailed working processes of the obtaining unit 1501, the selecting unit 1502, and the creating unit 1503, please refer to the descriptions in the foregoing method embodiments. For example, the obtaining unit 1501 obtains the virtual machine creation request input by the tenant by using the foregoing step 1401. The selecting unit 1502 selects the target server that can provide the specification of the first VGPU from multiple servers by using the foregoing step 1402. The creating unit 1503 creates the first virtual machine on the target server by using the foregoing step 1403. The embodiments of the present application will not repeat the description herein.
[0279] Corresponding to the foregoing server system, an embodiment of the present application further provides a virtual machine creation method based on cloud computing technology. This method is applied to a cloud management platform, and the cloud management platform is used to manage the infrastructure, and the infrastructure includes multiple servers. The implementation process of the virtual machine creation method based on cloud computing technology will be described below.
[0280] Figure 16 is a flowchart of a virtual machine creation method based on cloud computing technology provided by an embodiment of the present application. As Figure 16 shown, the virtual machine creation method based on cloud computing technology includes the following steps:
[0281] Step 1601: Obtain the virtual machine creation request input by the tenant, where the virtual machine creation request carries the specification of the first VGPU of the first virtual machine to be created.
[0282] For the implementation process of this step 1601, please refer to the implementation process of the foregoing step 1401 accordingly, and details will not be described here.
[0283] Step 1602: Select a second server that can provide the specification of the first VGPU from multiple servers.
[0284] After the cloud management platform obtains the specification of the first VGPU of the first virtual machine to be created, it can select a server that can provide the specification of the first VGPU from the infrastructure to obtain the second server.
[0285] Step 1603: Create the first virtual machine on the second server.
[0286] After the cloud management platform selects a second server that can provide the specifications of the first VGPU in the infrastructure, it can create a virtual machine that conforms to the specifications of the first VGPU in the second server. Moreover, in order to ensure that the application running on the first virtual machine can access the first VGPU in the manner provided by this application, it is also necessary to set up a first front-end driver interception module and a first application in the first virtual machine, and use the first front-end driver interception module to virtualize the access interface of the drive file system in the first virtual machine to obtain a virtual access interface of the drive file system. In addition, select a first server in the cloud management platform, perform device simulation on the first part of the computing power of the GPU of the first server to generate a virtual drive file (i.e., the first VGPU) that can be accessed by the first application. The parameters of the first VGPU match the specifications of the first VGPU. For example, when the specifications of the first VGPU indicate multiple parameters (such as computing power, video memory size, video memory bit width, and video memory bandwidth), the multiple parameters of the first VGPU obtained by device virtualization correspond one by one to the multiple parameters indicated by the specifications, and any one of the multiple parameters of the first VGPU is equal to or slightly greater than the corresponding parameter indicated by the specifications. At the same time, if the first server does not have a back-end driver interception module, it is also necessary to configure a first back-end driver interception module in the first server. Among them, the first front-end driver interception module is used to perform device simulation on the first part of the computing power of the first GPU to generate a first VGPU that can be accessed by the first application, the first application is used to access the first VGPU, and the first front-end driver interception module is also used to obtain a first access request of the first application for the first VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions. Correspondingly, the first GPU driver module is used to call the first part of the computing power of the first GPU to process the first access request.
[0287] In a possible implementation manner, the second server is also provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is also used to perform device simulation on the first part of the computing power of the second GPU to generate a third VGPU that can be accessed by the first application, the first application is used to access the third VGPU, and the first front-end driver interception module is also used to obtain a third access request of the first application for the third VGPU; the second back-end driver interception module is used to obtain the third access request from the first front-end driver interception module and send the third access request to the second GPU driver module when it is determined that the third access request meets the preset conditions; the second GPU driver module is used to call the first part of the computing power of the second GPU to process the third access request.
[0288] In a possible implementation, a first container is set up in the first virtual machine, and a first application runs in the first container. The first container is used to mount a first VGPU for the first application to access the first VGPU in the first container.
[0289] In a possible implementation, a third container is also set up in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application.
[0290] In a possible implementation, the first application runs directly in the first virtual machine.
[0291] In a possible implementation, the first GPU driver module is further used to obtain the first processing result generated by processing the first access request with the first part of the computing and processing power of the first GPU, and send the first processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the first processing result from the first GPU driver module and send the first processing result to the connection channel; the first front-end driver interception module is further used to obtain the first processing result from the connection channel from the first back-end driver interception module and provide the first processing result to the first application.
[0292] In a possible implementation, the second server further includes: a virtual machine manager, which is used to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel.
[0293] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, and the high-speed interconnection protocol includes the Compute Express Link (CXL) or the Lingqu Universal Bus (UB) protocol.
[0294] In a possible implementation, the first back-end driver interception module determines that the first access request meets the preset conditions, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first part of the computing power or the first computing power threshold.
[0295] In a possible implementation, a second virtual machine is further set up on the second server. A second front-end driver interception module and a second application are set up in the second virtual machine. The second front-end driver interception module is used to perform device simulation on the second part of the computing power of the first GPU to generate a second VGPU that can be accessed by the second application. The second application is used to access the second VGPU. The second front-end driver interception module is further used to obtain a second access request of the second application for the second VGPU and send a first access request to the connection channel. The first back-end driver interception module is used to obtain the second access request from the second front-end driver interception module from the connection channel, and send the second access request to the first GPU driver module when it is determined that the second access request meets the preset conditions. The first GPU driver module is used to call the second part of the computing power of the first GPU to process the second access request.
[0296] In a possible implementation, a second back-end driver interception module, a second GPU driver module, and a second GPU are further set up on the second server. The second front-end driver interception module of the second virtual machine is further used to perform device simulation on the second part of the computing power of the second GPU to generate a fourth VGPU that can be accessed by the second application. The second application is used to access the fourth VGPU. The second front-end driver interception module is further used to obtain a fourth access request of the second application for the fourth VGPU. The second back-end driver interception module is used to obtain the fourth access request from the second front-end driver interception module and send the fourth access request to the second GPU driver module when it is determined that the fourth access request meets the preset conditions. The second GPU driver module is used to call the second part of the computing power of the second GPU to process the fourth access request.
[0297] In a possible implementation, a second container is set up in the second virtual machine. The second application runs in the second container. The second container is used to mount the second VGPU for the second application to access the second VGPU in the second container.
[0298] In a possible implementation, a fourth container is further set up in the second virtual machine. The fourth application runs in the fourth container. The fourth container is used to mount the fourth VGPU for the fourth application to access the fourth VGPU in the fourth container. Among them, the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU that can be accessed by the fourth application.
[0299] In a possible implementation, the second application runs directly in the second virtual machine.
[0300] In a possible implementation, the first GPU driver module is further configured to obtain a second processing result generated by processing a second access request using a second part of the computing processing capability of the first GPU, and send the second processing result to the first back-end driver interception module; the first back-end driver interception module is further configured to obtain the second processing result from the first GPU driver module and send the second processing result to the connection channel; the second front-end driver interception module is further configured to obtain the second processing result from the connection channel that comes from the first back-end driver interception module and provide the second processing result to the second application.
[0301] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain a second access request from the second front-end driver interception module and send the second access request to the connection channel.
[0302] In a possible implementation, the second back-end driver interception module determines that the second access request meets a preset condition, including: the second back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than a second video memory threshold; and / or, the second back-end driver interception module determines that the computing capability of the GPU required by the second access request is not greater than a second part of the computing capability or a second computing capability threshold.
[0303] The foregoing has introduced the virtual machine creation method based on cloud computing technology in the embodiments of the present application. Corresponding to the foregoing method, the embodiments of the present application further provide a cloud management platform. The cloud management platform is used to manage infrastructure. The infrastructure includes multiple servers. Figure 17 It is a schematic structural diagram of a cloud management platform provided by an embodiment of the present application. Based on Figure 17 the following multiple components shown, the Figure 17 cloud management platform shown can perform all or part of the operations shown above. It should be understood that the device may include more additional components than those shown or omit some of the components shown. The embodiments of the present application do not limit this. As Figure 16 shown, the cloud management platform 170 may include: Figure 17
[0304] An obtaining unit 1701, configured to obtain a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of a first VGPU of a first virtual machine to be created.
[0305] A selection unit 1702, configured to select a second server that can provide the specification of the first VGPU from multiple servers.
[0306] A creation unit 1703, configured to create a first virtual machine on the second server.
[0307] Among them, the multiple servers further include a first server, which is provided with a first back-end driver interception module, a first GPU driver module, and a first GPU. A connection channel is established between the first server and the second server. A first front-end driver interception module and a first application are set in the first virtual machine. The first front-end driver interception module is used to perform device simulation on a part of the computing power of the first GPU to generate a first VGPU that can be accessed by the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU, and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions; the first GPU driver module is used to call a part of the computing power of the first GPU to process the first access request.
[0308] In a possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further used to perform device simulation on the first part of the computing power of the second GPU to generate a third VGPU that can be accessed by the first application. The first application is used to access the third VGPU. The first front-end driver interception module is further used to obtain a third access request of the first application for the third VGPU; the second back-end driver interception module is used to obtain the third access request from the first front-end driver interception module, and send the third access request to the second GPU driver module when it is determined that the third access request meets the preset conditions; the second GPU driver module is used to call the first part of the computing power of the second GPU to process the third access request.
[0309] In a possible implementation, a first container is set in the first virtual machine. The first application runs in the first container. The first container is used to mount the first VGPU for the first application to access the first VGPU in the first container.
[0310] In a possible implementation, a third container is further set in the first virtual machine. A third application runs in the third container. The third container is used to mount the third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU that can be accessed by the third application.
[0311] In a possible implementation, the first application runs directly in the first virtual machine.
[0312] In a possible implementation, the first GPU driver module is further configured to obtain a first processing result generated by processing a first access request using a first part of the computing processing power of the first GPU, and send the first processing result to the first back-end driver interception module; the first back-end driver interception module is further configured to obtain the first processing result from the first GPU driver module and send the first processing result to the connection channel; the first front-end driver interception module is further configured to obtain the first processing result from the connection channel that comes from the first back-end driver interception module and provide the first processing result to the first application.
[0313] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel.
[0314] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, and the high-speed interconnection protocol includes Compute Express Link (CXL) or Lingqu Universal Bus (UB) protocol.
[0315] In a possible implementation, the first back-end driver interception module determines that the first access request meets a preset condition, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first part of the computing power or the first computing power threshold.
[0316] In a possible implementation, a second virtual machine is further set up on the second server. A second front-end driver interception module and a second application are set up in the second virtual machine. The second front-end driver interception module is configured to perform device emulation on a second part of the computing power of the first GPU to generate a second virtual GPU (VGPU) accessible to the second application. The second application is configured to access the second VGPU. The second front-end driver interception module is further configured to obtain a second access request of the second application for the second VGPU and send the first access request to the connection channel; the first back-end driver interception module is configured to obtain the second access request from the connection channel that comes from the second front-end driver interception module and send the second access request to the first GPU driver module when determining that the second access request meets the preset condition; the first GPU driver module is configured to call a second part of the computing power of the first GPU to process the second access request.
[0317] In a possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The second front-end driver interception module of the second virtual machine is further configured to perform device simulation on the second part of the computing power of the second GPU to generate a fourth VGPU accessible to the second application. The second application is used to access the fourth VGPU. The second front-end driver interception module is further configured to obtain a fourth access request of the second application for the fourth VGPU. The second back-end driver interception module is configured to obtain the fourth access request from the second front-end driver interception module and send the fourth access request to the second GPU driver module when it is determined that the fourth access request meets a preset condition. The second GPU driver module is configured to call the second part of the computing power of the second GPU to process the fourth access request.
[0318] In a possible implementation, a second container is provided in the second virtual machine. The second application runs in the second container. The second container is used to mount the second VGPU for the second application to access the second VGPU in the second container.
[0319] In a possible implementation, a fourth container is further provided in the second virtual machine. The fourth application runs in the fourth container. The fourth container is used to mount the fourth VGPU for the fourth application to access the fourth VGPU in the fourth container. Among them, the second front-end driver interception module is configured to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application.
[0320] In a possible implementation, the second application runs directly in the second virtual machine.
[0321] In a possible implementation, the first GPU driver module is further configured to obtain the second processing result generated by processing the second access request with the second part of the computing and processing power of the first GPU, and send the second processing result to the first back-end driver interception module. The first back-end driver interception module is further configured to obtain the second processing result from the first GPU driver module and send the second processing result to the connection channel. The second front-end driver interception module is further configured to obtain the second processing result from the connection channel from the first back-end driver interception module and provide the second processing result to the second application.
[0322] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module and send the second access request to the connection channel.
[0323] In a possible implementation, the second back-end driver interception module determines that the second access request meets a preset condition, including: the second back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than a second video memory threshold; and / or, the second back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than a second partial computing power or a second computing power threshold.
[0324] Here, for the detailed working processes of the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703, please refer to the descriptions in the foregoing method embodiments. For example, the obtaining unit 1701 obtains the virtual machine creation request input by the tenant by using the foregoing step 1601. The selecting unit 1702 selects a second server that can provide the specification of the first VGPU from multiple servers by using the foregoing step 1602. The creating unit 1703 creates a first virtual machine on the second server by using the foregoing step 1603. The embodiments of the present application will not be described repeatedly herein.
[0325] Among them, the obtaining unit 1501, the selecting unit 1502, and the creating unit 1503, the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703 can all be implemented by software or can be implemented by a GPU. Exemplarily, next, taking the obtaining unit 1501 as an example, the implementation manner of the obtaining unit 1501 will be introduced. Similarly, the implementation manners of the selecting unit 1502 and the creating unit 1503, the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703 can refer to the implementation manner of the obtaining unit 1501.
[0326] As an example of a software functional unit, the obtaining unit 1501 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the foregoing computing instance may be one or more. For example, the obtaining unit 1501 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ) or may be distributed in different AZs, and each AZ includes one cloud data center or multiple geographically adjacent cloud data centers. Among them, generally, one region may include multiple AZs.
[0327] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Usually, one VPC is set up within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, communication gateways need to be set up within each VPC, and the interconnection between VPCs is achieved through the communication gateways.
[0328] As an example of a hardware functional unit, the obtaining unit 1501 may include at least one computing device, such as a server. Alternatively, the obtaining unit 1501 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0329] The multiple computing devices included in the obtaining unit 1501 may be distributed in the same region or in different regions. The multiple computing devices included in the obtaining unit 1501 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the obtaining unit 1501 may be distributed within the same VPC or across multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0330] It should be noted that in other embodiments, any one of the obtaining unit 1501, the selecting unit 1502, and the creating unit 1503, the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703 may be used to execute any step in the virtual machine creation method based on cloud computing technology. The steps to be implemented by the obtaining unit 1501, the selecting unit 1502, and the creating unit 1503, the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703 can be specified as needed. The full functions of the cloud management platform are realized by separately implementing different steps in the virtual machine creation method based on cloud computing technology through the obtaining unit 1501, the selecting unit 1502, and the creating unit 1503, the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703.
[0331] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described respective components can refer to the corresponding content in the foregoing method embodiments and will not be elaborated herein.
[0332] The following gives an example of the basic hardware structure involved in the embodiments of the present application.
[0333] The present application also provides a computing device 1800. As Figure 18 shown, the computing device 1800 includes: a bus 1802, a processor 1804, a memory 1806, and a communication interface 1808. The processor 1804, the memory 1806, and the communication interface 1808 communicate with each other through the bus 1802. The computing device 1800 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1800.
[0334] The bus 1802 may be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 18 only one line is shown herein, but it does not mean that there is only one bus or one type of bus. The bus 1802 may include a path for transmitting information between various components of the computing device 1800 (for example, the memory 1806, the processor 1804, the communication interface 1808).
[0335] The processor 1804 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0336] The memory 1806 is used to store computer programs, which include an operating system and executable code (i.e., program instructions). The memory 1806 may include volatile memory, such as random access memory (RAM). The processor 1804 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0337] The executable program code is stored in the memory 1806, and the processor 1804 executes the executable program code to respectively implement the functions of the aforementioned obtaining unit 1501, selecting unit 1502, and creating unit 1503, obtaining unit 1701, selecting unit 1702, and creating unit 1703, so as to implement any server based on cloud computing technology in this application. That is to say, the memory 1806 stores instructions for implementing any server based on cloud computing technology in this application.
[0338] The communication interface 1808 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 1800 and other devices or communication networks.
[0339] The embodiment of this application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0340] As Figure 19 shown, the computing device cluster includes at least one computing device 1800. The same instructions for implementing any server based on cloud computing technology in this application may be stored in the memory 1806 of one or more computing devices 1800 in the computing device cluster.
[0341] In some possible implementation manners, the memory 1806 of one or more computing devices 1800 in the computing device cluster may also respectively store partial instructions for implementing any server based on cloud computing technology in this application. In other words, the combination of one or more computing devices 1800 may jointly execute the instructions for implementing any server based on cloud computing technology in this application.
[0342] It should be noted that the memories 1806 in different computing devices 1800 in the computing device cluster may store different instructions, respectively used to execute some functions of the cloud management platform. That is, the instructions stored in the memories 1806 in different computing devices 1800 can implement the functions of one or more modules among the obtaining unit 1501, the selecting unit 1502, the creating unit 1503, the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703.
[0343] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 20 A possible implementation manner is shown. As Figure 20 shown, two computing devices 1800A and 1800B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device.
[0344] It should be understood that Figure 20 the functions of the computing device 1800A shown in
[0345] can also be completed by multiple computing devices 1800. Similarly, the functions of the computing device 1800B can also be completed by multiple computing devices 1800. Figure 19 and Figure 20 the connection manner of the computing device cluster. The difference is that the memories 1806 in one or more computing devices 1800 in this computing device cluster may store the same instructions for implementing any server based on cloud computing technology in this application.
[0346] In some possible implementation manners, the memories 1806 in one or more computing devices 1800 in this computing device cluster may also respectively store some instructions for implementing any server based on cloud computing technology in this application. In other words, the combination of one or more computing devices 1800 can jointly implement the instructions of any server based on cloud computing technology in this application.
[0347] The embodiments of the present application also provide a computer program product including instructions. The computer program product may be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it enables at least one computing device to implement any server based on cloud computing technology in this application.
[0348] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to implement any server based on cloud computing technology in the present application.
[0349] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0350] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions. For example, the original data and executable code involved in the present application are obtained under full authorization.
[0351] In the embodiments of the present application, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. The term "at least one" means one or more, and the term "a plurality" means two or more, unless otherwise clearly defined.
[0352] The term "and / or" in the present application is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0353] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A server based on cloud computing technology, characterized in that: The server runs a first virtual machine and a second virtual machine. The server is also provided with a back-end driver interception module, an image processing unit GPU driver module and a GPU, wherein: The first virtual machine is provided with a first front-end driver interception module and a first application program, wherein the first front-end driver interception module is used to perform device simulation on a first part of the computing power of the GPU to generate a first virtual image processing unit VGPU accessible to the first application program, the first application program is used to access the first VGPU, and the first front-end driver interception module is further used to obtain a first access request of the first application program to the first VGPU; The second virtual machine is provided with a second front-end driver interception module and a second application, the second front-end driver interception module is used to perform device simulation on the second part of the computing power of the GPU to generate a second VGPU accessible to the second application, the second application is used to access the second VGPU, and the second front-end driver interception module is further used to obtain a second access request of the second application for the second VGPU; The back-end driver interception module is used to obtain the first access request from the first front-end driver interception module and / or the second access request from the second front-end driver interception module, and send the first access request to the GPU driver module if it is determined that the first access request meets a preset condition, and / or send the second access request to the GPU driver module if it is determined that the second access request meets the preset condition; The GPU driver module is used to call the first part of the computing power of the GPU to process the first access request, and / or call the second part of the computing power of the GPU to process the second access request.
2. The server according to claim 1, characterized in that: A first container is provided in the first virtual machine, the first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container; A second container is provided in the second virtual machine, the second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
3. The server according to claim 2, characterized in that: The first virtual machine is further provided with a third container, in which a third application is run, and the third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the GPU to generate the third VGPU accessible to the third application; And / or, a fourth container is also provided in the second virtual machine, a fourth application runs in the fourth container, the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the GPU to generate the fourth VGPU accessible to the fourth application.
4. The server according to claim 1, characterized in that: The first application program runs directly in the first virtual machine, and the second application program runs directly in the second virtual machine.
5. The server according to any one of claims 1 to 4, characterized in that: The GPU driver module is further configured to obtain a first processing result generated by a first part of the computing processing capability of the GPU to process the first access request, and / or obtain a second processing result generated by a second part of the computing processing capability of the GPU to process the second access request, and send the first processing result and / or the second processing result to the back-end driver interception module; The back-end driver interception module is further used to obtain the first processing result and / or the second processing result from the GPU driver module, send the first processing result to the first front-end driver interception module, and / or send the second processing result to the second front-end driver interception module; The first front-end driver interception module is further used to obtain the first processing result from the back-end driver interception module and provide the first processing result to the first application program; The second front-end driver interception module is further used to obtain the second processing result from the back-end driver interception module, and provide the second processing result to the second application.
6. The server according to any one of claims 1 to 4, characterized in that: The server also includes: The virtual machine manager is used to obtain the first access request from the first front-end driver interception module and send the first access request to the back-end driver interception module, obtain the second access request from the second front-end driver interception module, and send the second access request to the back-end driver interception module.
7. The server according to any one of claims 1 to 6, characterized in that: The backend driver interception module determines that the first access request meets a preset condition, including: The back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or the back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold; The backend driver interception module determines that the second access request meets the preset condition, including: The back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than a second video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
8. A server system based on cloud computing technology, characterized in that: The server system includes a first server and a second server, wherein the first server is provided with a first back-end driver interception module, a first GPU driver module and a first GPU, a first virtual machine and a second virtual machine are running on the second server, and a connection channel is established between the first server and the second server, wherein: The first virtual machine is provided with a first front-end driver interception module and a first application program, wherein the first front-end driver interception module is used to perform device simulation on a first part of the computing power of the first GPU to generate a first VGPU accessible to the first application program, the first application program is used to access the first VGPU, and the first front-end driver interception module is further used to obtain a first access request of the first application program for the first VGPU and send the first access request to the connection channel; The second virtual machine is provided with a second front-end driver interception module and a second application, the second front-end driver interception module is used to perform device simulation on the second part of the computing power of the first GPU to generate a second VGPU accessible to the second application, the second application is used to access the second VGPU, and the second front-end driver interception module is further used to obtain a second access request of the second application for the second VGPU and send the first access request to the connection channel; The first back-end driver interception module is configured to obtain the first access request from the first front-end driver interception module and / or the second access request from the second front-end driver interception module from the connection channel, and send the first access request to the first GPU driver module if it is determined that the first access request meets a preset condition, and / or send the second access request to the first GPU driver module if it is determined that the second access request meets the preset condition; The first GPU driver module is used to call a first part of the computing power of the first GPU to process the first access request, and / or call a second part of the computing power of the first GPU to process the second access request.
9. The server system according to claim 8, characterized in that: The second server is also provided with a second back-end driver interception module, a second GPU driver module and a second GPU. The first front-end driver interception module of the first virtual machine is further used to perform device simulation on the first part of the computing power of the second GPU to generate a third VGPU accessible to the first application, the first application is used to access the third VGPU, and the first front-end driver interception module is further used to obtain a third access request of the first application for the third VGPU; The second front-end driver interception module of the second virtual machine is further used to perform device simulation on the second part of the computing power of the second GPU to generate a fourth VGPU accessible to the second application, the second application is used to access the fourth VGPU, and the second front-end driver interception module is further used to obtain a fourth access request of the second application for the fourth VGPU; the second back-end driver interception module is configured to obtain the third access request from the first front-end driver interception module and / or the fourth access request from the second front-end driver interception module, and send the third access request to the second GPU driver module if it is determined that the third access request meets the preset condition, and / or send the fourth access request to the second GPU driver module if it is determined that the fourth access request meets the preset condition; The second GPU driver module is used to call the first part of the computing power of the second GPU to process the third access request, and / or call the second part of the computing power of the second GPU to process the fourth access request.
10. The server system according to claim 8 or 9, characterized in that: A first container is provided in the first virtual machine, the first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container; A second container is provided in the second virtual machine, the second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
11. The server system according to claim 10, characterized in that: The first virtual machine is further provided with a third container, a third application is run in the third container, the third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate the third VGPU accessible to the third application; And / or, a fourth container is also provided in the second virtual machine, a fourth application runs in the fourth container, the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate the fourth VGPU accessible to the fourth application.
12. The server system according to claim 8 or 9, characterized in that: The first application program runs directly in the first virtual machine, and the second application program runs directly in the second virtual machine.
13. The server system according to any one of claims 8 to 12, characterized in that: The first GPU driver module is further configured to obtain a first processing result generated by a first part of the computing processing capability of the first GPU processing the first access request, and / or obtain a second processing result generated by a second part of the computing processing capability of the first GPU processing the second access request, and send the first processing result and / or the second processing result to the first back-end driver interception module; The first back-end driver interception module is further used to obtain the first processing result and / or the second processing result from the first GPU driver module, send the first processing result to the connection channel, and / or send the second processing result to the connection channel; The first front-end driver interception module is further used to obtain the first processing result from the first back-end driver interception module from the connection channel, and provide the first processing result to the first application; The second front-end driver interception module is further used to obtain the second processing result from the first back-end driver interception module through the connection channel, and provide the second processing result to the second application.
14. The server system according to any one of claims 8 to 13, characterized in that: The second server also includes: The virtual machine manager is used to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel, obtain the second access request from the second front-end driver interception module, and send the second access request to the connection channel.
15. The server system according to any one of claims 8 to 14, characterized in that: The connection channel is implemented through a high-speed interconnection protocol, which includes a computing high-speed link CXL or a Lingqu UB bus protocol.
16. The server system according to any one of claims 8 to 15, characterized in that: The first backend driver interception module determines that the first access request meets a preset condition, including: The first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold; The second backend driver interception module determines that the second access request meets the preset condition, including: The second back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than a second video memory threshold; and / or, the second back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
17. A method for creating a virtual machine based on cloud computing technology, characterized in that: The method is applied to a cloud management platform, the cloud management platform is used to manage infrastructure, the infrastructure includes multiple servers, and the method includes: Obtaining a virtual machine creation request input by a tenant, where the virtual machine creation request carries a specification of a first VGPU of a first virtual machine to be created; Selecting a target server from the plurality of servers that can provide the specifications of the first VGPU; Creating the first virtual machine on the target server; Among them, the target server is provided with a back-end driver interception module, a GPU driver module and a GPU, and the first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on the first part of the computing power of the GPU to generate the first VGPU accessible to the first application. The first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request of the first application for the first VGPU; the back-end driver interception module is used to obtain the first access request from the first front-end driver interception module, and send the first access request to the GPU driver module when it is determined that the first access request meets a preset condition; the GPU driver module is used to call the first part of the computing power of the GPU to process the first access request.
18. A cloud management platform, characterized in that: The cloud management platform is used to manage infrastructure, the infrastructure includes multiple servers, and the cloud management platform includes: An acquisition unit, configured to acquire a virtual machine creation request input by a tenant, wherein the virtual machine creation request carries a specification of a first VGPU of a first virtual machine to be created; A selection unit, configured to select a target server from among the multiple servers that can provide the specification of the first VGPU; A creating unit, configured to create the first virtual machine on the target server; Among them, the target server is provided with a back-end driver interception module, a GPU driver module and a GPU, and the first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on the first part of the computing power of the GPU to generate the first VGPU accessible to the first application. The first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request of the first application for the first VGPU; the back-end driver interception module is used to obtain the first access request from the first front-end driver interception module, and send the first access request to the GPU driver module when it is determined that the first access request meets a preset condition; the GPU driver module is used to call the first part of the computing power of the GPU to process the first access request.
19. A method for creating a virtual machine based on cloud computing technology, characterized in that: The method is applied to a cloud management platform, the cloud management platform is used to manage infrastructure, the infrastructure includes multiple servers, and the method includes: Obtaining a virtual machine creation request input by a tenant, where the virtual machine creation request carries a specification of a first VGPU of a first virtual machine to be created; Selecting a second server from the plurality of servers that can provide the specifications of the first VGPU; Creating the first virtual machine on the second server; Among them, the multiple servers also include a first server, the first server is provided with a first back-end driver interception module, a first GPU driver module and a first GPU, a connection channel is established between the first server and the second server, the first virtual machine is provided with a first front-end driver interception module and a first application, the first front-end driver interception module is used to perform device simulation on part of the computing power of the first GPU to generate the first VGPU accessible to the first application, the first application is used to access the first VGPU, the first front-end driver interception module is also used to obtain a first access request of the first application for the first VGPU, and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets a preset condition; the first GPU driver module is used to call part of the computing power of the first GPU to process the first access request.
20. A cloud management platform, characterized in that: The cloud management platform is used to manage infrastructure, the infrastructure includes multiple servers, and the cloud management platform includes: An acquisition unit, configured to acquire a virtual machine creation request input by a tenant, wherein the virtual machine creation request carries a specification of a first VGPU of a first virtual machine to be created; A selection unit, configured to select, from among the plurality of servers, a second server that can provide the specification of the first VGPU; A creating unit, configured to create the first virtual machine on the second server; Among them, the multiple servers also include a first server, the first server is provided with a first back-end driver interception module, a first GPU driver module and a first GPU, a connection channel is established between the first server and the second server, the first virtual machine is provided with a first front-end driver interception module and a first application, the first front-end driver interception module is used to perform device simulation on part of the computing power of the first GPU to generate the first VGPU accessible to the first application, the first application is used to access the first VGPU, the first front-end driver interception module is also used to obtain a first access request of the first application for the first VGPU, and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets a preset condition; the first GPU driver module is used to call part of the computing power of the first GPU to process the first access request.
21. A computing device, characterized in that It comprises a processor and a memory, wherein the memory stores program instructions, and the processor executes the program instructions so that the computing device implements the server described in any one of claims 1 to 16.
22. A computer-readable storage medium, characterized in that: The method comprises program instructions, which, when executed on a computing device, enable the computing device to implement the server according to any one of claims 1 to 16.
23. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device, the computing device implements the server described in any one of claims 1 to 16.