Server, server system, virtual machine creation method, and cloud management platform
By introducing front-end and back-end driver interception modules into the virtual machine to simulate GPU computing capabilities, the problem of poor flexibility in accessing GPUs by virtual machines is solved, and flexibility in virtualization segmentation and scheduling of GPUs is achieved, and the dependence on GPU hardware information is reduced.
Patent Information
- Application Number
- PCT/CN2024/142246
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-30
- Filing Date
- 2024-12-25
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the way virtual machines access GPUs is less flexible and cannot effectively break through GPU hardware control, resulting in excessive dependence on GPU hardware information.
By introducing front-end and back-end driver interception modules into the virtual machine, some of the computing power of the GPU is simulated to provide VGPUs to the virtual machine, and access requests are sent to the GPU driver module under preset conditions, so as to realize the software delivery of access requests and processing results, and reduce dependence on GPU hardware information.
It improves the flexibility of GPU access by applications in virtual machines, supports multiple scheduling strategies, realizes flexibility in virtualization segmentation and scheduling of GPUs, and reduces dependence on GPU hardware information.
Smart Images

Figure CN2024142246_03072025_PF_FP_ABST
Abstract
Description
Server, server system, virtual machine creation method and cloud management platform
[0001] This application claims priority to Chinese patent application No. 202311819308.4 filed on December 25, 2023, with invention name “Cloud Service Providing Method and System”, and priority to Chinese patent application No. 202410559048.X filed on April 30, 2024, with invention name “Server, Server System, Virtual Machine Creation Method and Cloud Management Platform”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of cloud service technology, and in particular to a server, a server system, a virtual machine creation method, and a cloud management platform based on cloud computing technology. Background Art
[0003] Virtualization technology abstracts and transforms a host's physical resources, such as computing, networking, and storage, to create a more tangible representation. This breaks down the barriers between the host's physical structure and allows tenants to utilize these resources in a more efficient manner than their original configuration. This allows virtualization to divide a server's GPU among multiple virtual machines.
[0004] Typically, when a virtual machine needs to use a GPU, it needs to access the GPU's driver files. Upon detecting this access, the server's driver file system forwards the virtual machine's access request to the GPU driver module. Upon receiving this access request, the GPU driver module accesses the GPU based on the request and provides the virtual machine with the results.
[0005] However, this access method is less flexible. Summary of the Invention
[0006] This application provides a server, server system, virtual machine creation method, and cloud management platform based on cloud computing technology. This application reduces the reliance of the process of calling a GPU to process access requests on GPU hardware information and other implementation details. The technical solutions provided by this application are as follows:
[0007] In a first aspect, the present application provides a server based on cloud computing technology, on which a first virtual machine and a second virtual machine run. The server is also provided with a back-end driver interception module, an image processing unit GPU driver module and a GPU. Among them, the first virtual machine is provided with a first front-end driver interception module and a first application, the first front-end driver interception module is used to simulate the first part of the GPU's computing power to generate a first virtual image processing unit VGPU accessible to the first application, the first application is used to access the first VGPU, and the first front-end driver interception module is further used to obtain a first access request from the first application to the first VGPU; the second virtual machine is provided with a second front-end driver interception module and a second application, the second front-end driver interception module is used to simulate the second part of the GPU's computing power to generate a second VGPU accessible to the second application, the second application is used to access the second VGPU, and the second front-end driver interception module is further used to obtain a second access request from the second application to the second VGPU; the back-end driver interception module is used to obtain the first access request from the first front-end driver interception module and / or the second access request from the second front-end driver interception module, and send the first access request to the GPU driver module if it is determined that the first access request meets the preset conditions, and / or, send the second access request to the GPU driver module if it is determined that the second access request meets the preset conditions; the GPU driver module is used to call the first part of the GPU's computing power to process the first access request, and / or call the second part of the GPU's computing power to process the second access request.
[0008] In this application, the front-end driver interception module can capture access requests for the VGPU from applications in the virtual machine and, through the back-end driver interception module, provide these access requests to the GPU driver module. This allows the GPU driver module to utilize some of the GPU's computing power to process the access requests. This allows applications in the virtual machine to access the GPU through software, reducing the reliance on GPU hardware information and other implementation details when invoking the GPU to process access requests, and increasing the flexibility of GPU access for applications in the virtual machine.
[0009] In this application, there are multiple ways to implement the server. The following three implementations are used as examples to illustrate them.
[0010] In a first implementation, the first application program runs directly in the first virtual machine, and the second application program runs directly in the second virtual machine.
[0011] In a second implementation, a first container is provided in a first virtual machine, a first application runs in the first container, and the first container is used to mount a first VGPU so that the first application can access the first VGPU in the first container. A second container is provided in a second virtual machine, a second application runs in the second container, and the second container is used to mount a second VGPU so that the second application can access the second VGPU in the second container.
[0012] Optionally, the number of containers set in the virtual machine can be adjusted according to application requirements. For example, one container can be set in the virtual machine, or multiple containers can be set in the virtual machine. For example, the first virtual machine is further provided with a third container, a third application runs in the third container, and the third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the GPU's computing power to generate a third VGPU accessible to the third application. And / or, the second virtual machine is further provided with a fourth container, a fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth part of the GPU's computing power to generate a fourth VGPU accessible to the fourth application.
[0013] After the GPU driver module calls the GPU's computing power to process the access request, if it receives the processing result generated by the GPU processing the access request, the GPU driver module is also used to provide the processing result to the back-end driver interception module, so that the processing result is returned to the application that triggered the access request through the back-end driver interception module. In one possible implementation, the GPU driver module is also used to obtain the first processing result generated by the first part of the GPU's computing power to process the first access request, and / or obtain the second processing result generated by the second part of the GPU's computing power to process the second access request, and send the first processing result and / or the second processing result to the back-end driver interception module; the back-end driver interception module is also used to obtain the first processing result and / or the second processing result from the GPU driver module, send the first processing result to the first front-end driver interception module, and / or send the second processing result to the second front-end driver interception module; the first front-end driver interception module is also used to obtain the first processing result from the back-end driver interception module and provide the first processing result to the first application; the second front-end driver interception module is also used to obtain the second processing result from the back-end driver interception module and provide the second processing result to the second application.
[0014] As can be seen, the access request triggered by the application in the virtual machine accessing the VGPU is sequentially transmitted to the GPU driver module in the server via the front-end driver interception module in the virtual machine and the back-end driver interception module in the server. The processing result generated by the GPU processing the access request is sequentially transmitted to the application in the virtual machine via the GPU driver module in the server, the back-end driver interception module in the server, and the front-end driver interception module in the virtual machine. Because both transmission processes are implemented in software, the process of calling the GPU to process the access request is less dependent on GPU hardware information and other implementation details, and even eliminates the need to rely on GPU hardware information and other implementation details. For example, it does not rely on GPU manufacturer hardware information, achieving decoupling from the GPU and improving the flexibility of GPU access for applications in the virtual machine. Furthermore, when the virtual machine uses part of the GPU's computing power, it is equivalent to virtualizing and partitioning the GPU. The software implementation of transmitting access requests and processing results in this application allows the virtualization of GPU partitioning to be divided beyond the GPU hardware control, supporting multiple scheduling strategies for GPU scheduling, and improving GPU scheduling flexibility.
[0015] In one possible implementation, the server also includes: a virtual machine manager, used to obtain a first access request from a first front-end driver interception module and send the first access request to the back-end driver interception module, obtain a second access request from a second front-end driver interception module, and send the second access request to the back-end driver interception module.
[0016] When the back-end driver interception module determines that the access request meets the preset conditions, the back-end driver interception module forwards the access request to the GPU driver module. When the back-end driver interception module determines that the access request does not meet the preset conditions, the access request is not forwarded to the GPU driver module. At this time, the back-end driver interception module may return an error message to the virtual machine to which the access request belongs to indicate that the virtual machine cannot access the GPU. In one possible implementation, the back-end driver interception module determines that the first access request meets the preset conditions, including: the back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first part of the computing power or the first computing power threshold; the back-end driver interception module determines that the second access request meets the preset conditions, including: the back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than the second video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0017] Similarly, when a container is also running in the virtual machine, the front-end driver interception module set in the virtual machine may optionally determine whether the access request meets the preset conditions after receiving the access request, and then provide the access request to the back-end driver interception module if the access request meets the preset conditions. If the access request does not meet the preset conditions, the virtual machine may return an error message to the container to indicate that the container cannot access the GPU. In one possible implementation, the virtual machine may optionally determine, on a container basis, whether the GPU memory required by the access request triggered by the application running in the container is greater than the memory threshold; and / or whether the GPU computing power required by the access request is greater than the partial computing power or computing power threshold of the GPU available for use by the VGPU accessible by the application. Here, the memory threshold is the amount of video memory configured by the tenant for the container, and the computing power threshold is the amount of computing power configured by the tenant for the container.
[0018] In the second aspect, the present application provides a server system based on cloud computing technology, the server system including a first server and a second server, the first server being provided with a first back-end driver interception module, a first GPU driver module and a first GPU, the second server running a first virtual machine and a second virtual machine, and a connection channel being established between the first server and the second server.
[0019] Among them, the first virtual machine is provided with a first front-end driver interception module and a first application, the first front-end driver interception module is used to simulate the first part of the computing power of the first GPU to generate a first VGPU accessible to the first application, the first application is used to access the first VGPU, and the first front-end driver interception module is also used to obtain the first access request of the first application for the first VGPU and send the first access request to the connection channel; the second virtual machine is provided with a second front-end driver interception module and a second application, the second front-end driver interception module is used to simulate the second part of the computing power of the first GPU to generate a second VGPU accessible to the second application, the second application is used to access the second VGPU, and the second front-end driver interception module The interception module is also used to obtain a second access request from the second application for the second VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module and / or the second access request from the second front-end driver interception module from the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions, and / or send the second access request to the first GPU driver module when it is determined that the second access request meets the preset conditions; the first GPU driver module is used to call the first part of the computing power of the first GPU to process the first access request, and / or call the second part of the computing power of the first GPU to process the second access request.
[0020] In a possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further configured to perform device simulation on the first portion of the computing power of the second GPU to generate a third VGPU accessible to the first application, the first application being configured to access the third VGPU, and the first front-end driver interception module is further configured to obtain a third access request from the first application for the third VGPU; the second front-end driver interception module of the second virtual machine is further configured to perform device simulation on the second portion of the computing power of the second GPU to generate a fourth VGPU accessible to the second application, the second application being configured to access the fourth VGPU, and the second front-end driver interception module is further configured to obtain a fourth access request from the second application for the fourth VGPU; the second back-end driver interception module is configured to obtain the third access request from the first front-end driver interception module and / or the fourth access request from the second front-end driver interception module, and send the third access request to the second GPU driver module if it is determined that the third access request meets a preset condition, and / or send the fourth access request to the second GPU driver module if it is determined that the fourth access request meets the preset condition; the second GPU driver module is configured to call the first portion of the computing power of the second GPU to process the third access request, and / or call the second portion of the computing power of the second GPU to process the fourth access request.
[0021] In one possible implementation, a first container is provided in a first virtual machine, a first application runs in the first container, and the first container is used to mount a first VGPU so that the first application can access the first VGPU in the first container; a second container is provided in a second virtual machine, a second application runs in the second container, and the second container is used to mount a second VGPU so that the second application can access the second VGPU in the second container.
[0022] In one possible implementation, a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application.
[0023] And / or, a fourth container is also provided in the second virtual machine, a fourth application runs in the fourth container, the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application.
[0024] In a possible implementation, the first application program runs directly in the first virtual machine, and the second application program runs directly in the second virtual machine.
[0025] In one possible implementation, the first GPU driver module is also used to obtain a first processing result generated by the first part of the computing processing power of the first GPU to process a first access request, and / or to obtain a second processing result generated by the second part of the computing processing power of the first GPU to process a second access request, and send the first processing result and / or the second processing result to the first back-end driver interception module; the first back-end driver interception module is also used to obtain the first processing result and / or the second processing result from the first GPU driver module, send the first processing result to the connection channel, and / or send the second processing result to the connection channel; the first front-end driver interception module is also used to obtain the first processing result from the first back-end driver interception module from the connection channel, and provide the first processing result to the first application; the second front-end driver interception module is also used to obtain the second processing result from the first back-end driver interception module from the connection channel, and provide the second processing result to the second application.
[0026] In one possible implementation, the second server also includes: a virtual machine manager, used to obtain a first access request from the first front-end driver interception module and send the first access request to the connection channel, obtain a second access request from the second front-end driver interception module, and send the second access request to the connection channel.
[0027] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, which includes a computing express link (CXL) or a Lingqu UB bus protocol.
[0028] In one possible implementation, the first back-end driver interception module determines that the first access request meets preset conditions, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0029] The second back-end driver interception module determines that the second access request meets the preset conditions, including: the second back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than the second video memory threshold; and / or, the second back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0030] In a third aspect, the present application provides a method for creating a virtual machine based on cloud computing technology. The method is applied to a cloud management platform. The cloud management platform is used to manage infrastructure. The infrastructure includes multiple servers. The method comprises: receiving a virtual machine creation request input by a tenant, the virtual machine creation request including specifications for a first VGPU of a first virtual machine to be created; selecting a target server from the multiple servers that can provide the specifications for the first VGPU; and creating the first virtual machine on the target server.
[0031] Among them, the target server is provided with a back-end driver interception module, a GPU driver module and a GPU, and the first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to simulate the first part of the GPU's computing power to generate a first VGPU accessible to the first application. The first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request from the first application for the first VGPU; the back-end driver interception module is used to obtain the first access request from the first front-end driver interception module, and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions; the GPU driver module is used to call the first part of the GPU's computing power to process the first access request.
[0032] In a possible implementation, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container.
[0033] In one possible implementation, a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the GPU's computing power to generate a third VGPU that can be accessed by the third application.
[0034] In a possible implementation, the first application program runs directly in the first virtual machine.
[0035] In one possible implementation, the GPU driver module is also used to obtain the first processing result generated by the first part of the GPU's computing processing power to process the first access request, and send the first processing result to the back-end driver interception module; the back-end driver interception module is also used to obtain the first processing result from the GPU driver module, and send the first processing result to the first front-end driver interception module; the first front-end driver interception module is also used to obtain the first processing result from the back-end driver interception module, and provide the first processing result to the first application.
[0036] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module, and send the first access request to the back-end driver interception module.
[0037] In one possible implementation, the back-end driver interception module determines that the first access request meets preset conditions, including: the back-end driver interception module determines that the GPU video memory required by the first access request is not greater than the first video memory threshold; and / or, the back-end driver interception module determines that the GPU computing power required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0038] In one possible implementation, a second virtual machine is also provided on the target server, and a second front-end driver interception module and a second application are provided in the second virtual machine. The second front-end driver interception module is used to simulate the second part of the GPU's computing power to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is also used to obtain a second access request from the second application for the second VGPU; the back-end driver interception module is used to obtain the second access request from the second front-end driver interception module, and send the second access request to the GPU driver module when it is determined that the second access request meets preset conditions; the GPU driver module is used to call the second part of the GPU's computing power to process the second access request.
[0039] In one possible implementation, a second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
[0040] In a possible implementation, the second application program runs directly in the second virtual machine.
[0041] In one possible implementation, the GPU driver module is also used to obtain the second processing result generated by the second part of the GPU's computing processing power to process the second access request, and send the second processing result to the back-end driver interception module; the back-end driver interception module is also used to obtain the second processing result from the GPU driver module, and send the second processing result to the second front-end driver interception module; the second front-end driver interception module is also used to obtain the second processing result from the back-end driver interception module, and provide the second processing result to the second application.
[0042] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module, and send the second access request to the back-end driver interception module.
[0043] In one possible implementation, the back-end driver interception module determines that the second access request meets preset conditions, including: the back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold; and / or the back-end driver interception module determines that the GPU computing power required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0044] In a fourth aspect, the present application provides a cloud management platform. The cloud management platform is used to manage infrastructure. The infrastructure includes multiple servers. The cloud management platform includes: an acquisition unit for acquiring a virtual machine creation request input by a tenant, the virtual machine creation request carrying specifications of a first VGPU of a first virtual machine to be created; a selection unit for selecting a target server from multiple servers that can provide the specifications of the first VGPU; and a creation unit for creating the first virtual machine on the target server.
[0045] Among them, the target server is provided with a back-end driver interception module, a GPU driver module and a GPU, and the first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to simulate the first part of the GPU's computing power to generate a first VGPU accessible to the first application. The first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request from the first application for the first VGPU; the back-end driver interception module is used to obtain the first access request from the first front-end driver interception module, and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions; the GPU driver module is used to call the first part of the GPU's computing power to process the first access request.
[0046] In a possible implementation, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container.
[0047] In one possible implementation, a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the GPU's computing power to generate a third VGPU that can be accessed by the third application.
[0048] In a possible implementation, the first application program runs directly in the first virtual machine.
[0049] In one possible implementation, the GPU driver module is also used to obtain the first processing result generated by the first part of the GPU's computing processing power to process the first access request, and send the first processing result to the back-end driver interception module; the back-end driver interception module is also used to obtain the first processing result from the GPU driver module, and send the first processing result to the first front-end driver interception module; the first front-end driver interception module is also used to obtain the first processing result from the back-end driver interception module, and provide the first processing result to the first application.
[0050] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module, and send the first access request to the back-end driver interception module.
[0051] In one possible implementation, the back-end driver interception module determines that the first access request meets preset conditions, including: the back-end driver interception module determines that the GPU video memory required by the first access request is not greater than the first video memory threshold; and / or, the back-end driver interception module determines that the GPU computing power required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0052] In one possible implementation, a second virtual machine is also provided on the target server, and a second front-end driver interception module and a second application are provided in the second virtual machine. The second front-end driver interception module is used to simulate the second part of the GPU's computing power to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is also used to obtain a second access request from the second application for the second VGPU; the back-end driver interception module is used to obtain the second access request from the second front-end driver interception module, and send the second access request to the GPU driver module when it is determined that the second access request meets preset conditions; the GPU driver module is used to call the second part of the GPU's computing power to process the second access request.
[0053] In one possible implementation, a second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
[0054] In a possible implementation, the second application program runs directly in the second virtual machine.
[0055] In one possible implementation, the GPU driver module is also used to obtain the second processing result generated by the second part of the GPU's computing processing power to process the second access request, and send the second processing result to the back-end driver interception module; the back-end driver interception module is also used to obtain the second processing result from the GPU driver module, and send the second processing result to the second front-end driver interception module; the second front-end driver interception module is also used to obtain the second processing result from the back-end driver interception module, and provide the second processing result to the second application.
[0056] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module, and send the second access request to the back-end driver interception module.
[0057] In one possible implementation, the back-end driver interception module determines that the second access request meets preset conditions, including: the back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold; and / or the back-end driver interception module determines that the GPU computing power required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0058] In a fifth aspect, the present application provides a method for creating a virtual machine based on cloud computing technology. This method is applied to a cloud management platform. The cloud management platform is used to manage infrastructure. The infrastructure includes multiple servers. The method comprises: receiving a virtual machine creation request input by a tenant, the virtual machine creation request including specifications for a first VGPU of a first virtual machine to be created; selecting a second server from the multiple servers that can provide the specifications for the first VGPU; and creating the first virtual machine on the second server.
[0059] Among them, the multiple servers also include a first server, the first server is provided with a first back-end driver interception module, a first GPU driver module and a first GPU, a connection channel is established between the first server and the second server, and a first front-end driver interception module and a first application are provided in the first virtual machine. The first front-end driver interception module is used to simulate part of the computing power of the first GPU to generate a first VGPU accessible to the first application, and the first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request from the first application for the first VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions; the first GPU driver module is used to call part of the computing power of the first GPU to process the first access request.
[0060] In one possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further configured to simulate the first portion of the computing power of the second GPU to generate a third VGPU accessible to the first application, the first application being configured to access the third VGPU, and the first front-end driver interception module being configured to obtain a third access request from the first application for the third VGPU; the second back-end driver interception module is configured to obtain the third access request from the first front-end driver interception module, and upon determining that the third access request meets preset conditions, the third access request is sent to the second GPU driver module; and the second GPU driver module is configured to invoke the first portion of the computing power of the second GPU to process the third access request.
[0061] In a possible implementation, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container.
[0062] In one possible implementation, a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application.
[0063] In a possible implementation, the first application program runs directly in the first virtual machine.
[0064] In one possible implementation, the first GPU driver module is further used to obtain a first processing result generated by the first part of the computing processing power of the first GPU to process the first access request, and send the first processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the first processing result from the first GPU driver module, and send the first processing result to the connection channel; the first front-end driver interception module is further used to obtain the first processing result from the first back-end driver interception module from the connection channel, and provide the first processing result to the first application.
[0065] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel.
[0066] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, which includes a computing express link (CXL) or a Lingqu UB bus protocol.
[0067] In one possible implementation, the first back-end driver interception module determines that the first access request meets preset conditions, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0068] In one possible implementation, a second virtual machine is also provided on the second server, and a second front-end driver interception module and a second application are provided in the second virtual machine. The second front-end driver interception module is used to simulate the second part of the computing power of the first GPU to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is also used to obtain a second access request from the second application for the second VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the second access request from the second front-end driver interception module from the connection channel, and send the second access request to the first GPU driver module when it is determined that the second access request meets the preset conditions; the first GPU driver module is used to call the second part of the computing power of the first GPU to process the second access request.
[0069] In one possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The second front-end driver interception module of the second virtual machine is further used to simulate the second portion of the computing power of the second GPU to generate a fourth VGPU accessible to the second application, the second application is used to access the fourth VGPU, and the second front-end driver interception module is further used to obtain a fourth access request from the second application for the fourth VGPU; the second back-end driver interception module is used to obtain the fourth access request from the second front-end driver interception module, and send the fourth access request to the second GPU driver module if it is determined that the fourth access request meets preset conditions; the second GPU driver module is used to call the second portion of the computing power of the second GPU to process the fourth access request.
[0070] In one possible implementation, a second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
[0071] In one possible implementation, a fourth container is further provided in the second virtual machine, a fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application.
[0072] In a possible implementation, the second application program runs directly in the second virtual machine.
[0073] In one possible implementation, the first GPU driver module is further used to obtain the second processing result generated by the second part of the computing processing power of the first GPU to process the second access request, and send the second processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the second processing result from the first GPU driver module, and send the second processing result to the connection channel; the second front-end driver interception module is further used to obtain the second processing result from the first back-end driver interception module from the connection channel, and provide the second processing result to the second application.
[0074] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module and send the second access request to the connection channel.
[0075] In one possible implementation, the second back-end driver interception module determines that the second access request meets preset conditions, including: the second back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold; and / or, the second back-end driver interception module determines that the GPU computing power required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0076] In a sixth aspect, the present application provides a cloud management platform. The cloud management platform is used to manage infrastructure. The infrastructure includes multiple servers. The cloud management platform includes: an acquisition unit for acquiring a virtual machine creation request input by a tenant, the virtual machine creation request carrying specifications of a first VGPU of a first virtual machine to be created; a selection unit for selecting a second server from the multiple servers that can provide the specifications of the first VGPU; and a creation unit for creating the first virtual machine on the second server.
[0077] Among them, the multiple servers also include a first server, the first server is provided with a first back-end driver interception module, a first GPU driver module and a first GPU, a connection channel is established between the first server and the second server, and a first front-end driver interception module and a first application are provided in the first virtual machine. The first front-end driver interception module is used to simulate part of the computing power of the first GPU to generate a first VGPU accessible to the first application, and the first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request from the first application for the first VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions; the first GPU driver module is used to call part of the computing power of the first GPU to process the first access request.
[0078] In one possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further configured to simulate the first portion of the computing power of the second GPU to generate a third VGPU accessible to the first application, the first application being configured to access the third VGPU, and the first front-end driver interception module being configured to obtain a third access request from the first application for the third VGPU; the second back-end driver interception module is configured to obtain the third access request from the first front-end driver interception module, and upon determining that the third access request meets preset conditions, the third access request is sent to the second GPU driver module; and the second GPU driver module is configured to invoke the first portion of the computing power of the second GPU to process the third access request.
[0079] In a possible implementation, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container.
[0080] In one possible implementation, a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application.
[0081] In a possible implementation, the first application program runs directly in the first virtual machine.
[0082] In one possible implementation, the first GPU driver module is further used to obtain a first processing result generated by the first part of the computing processing power of the first GPU to process the first access request, and send the first processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the first processing result from the first GPU driver module, and send the first processing result to the connection channel; the first front-end driver interception module is further used to obtain the first processing result from the first back-end driver interception module from the connection channel, and provide the first processing result to the first application.
[0083] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel.
[0084] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, which includes a computing express link (CXL) or a Lingqu UB bus protocol.
[0085] In one possible implementation, the first back-end driver interception module determines that the first access request meets preset conditions, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0086] In one possible implementation, a second virtual machine is also provided on the second server, and a second front-end driver interception module and a second application are provided in the second virtual machine. The second front-end driver interception module is used to simulate the second part of the computing power of the first GPU to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is also used to obtain a second access request from the second application for the second VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the second access request from the second front-end driver interception module from the connection channel, and send the second access request to the first GPU driver module when it is determined that the second access request meets the preset conditions; the first GPU driver module is used to call the second part of the computing power of the first GPU to process the second access request.
[0087] In one possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The second front-end driver interception module of the second virtual machine is further used to simulate the second portion of the computing power of the second GPU to generate a fourth VGPU accessible to the second application, the second application is used to access the fourth VGPU, and the second front-end driver interception module is further used to obtain a fourth access request from the second application for the fourth VGPU; the second back-end driver interception module is used to obtain the fourth access request from the second front-end driver interception module, and send the fourth access request to the second GPU driver module if it is determined that the fourth access request meets preset conditions; the second GPU driver module is used to call the second portion of the computing power of the second GPU to process the fourth access request.
[0088] In one possible implementation, a second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
[0089] In one possible implementation, a fourth container is further provided in the second virtual machine, a fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application.
[0090] In a possible implementation, the second application program runs directly in the second virtual machine.
[0091] In one possible implementation, the first GPU driver module is further used to obtain the second processing result generated by the second part of the computing processing power of the first GPU to process the second access request, and send the second processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the second processing result from the first GPU driver module, and send the second processing result to the connection channel; the second front-end driver interception module is further used to obtain the second processing result from the first back-end driver interception module from the connection channel, and provide the second processing result to the second application.
[0092] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module and send the second access request to the connection channel.
[0093] In one possible implementation, the second back-end driver interception module determines that the second access request meets preset conditions, including: the second back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold; and / or, the second back-end driver interception module determines that the GPU computing power required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0094] In the seventh aspect, the present application provides a computing device comprising a memory and a processor, the memory storing program instructions, the processor executing the program instructions to implement the server provided in the first aspect of the present application and any possible implementation thereof, and to implement the server system provided in the first aspect of the present application and any possible implementation thereof.
[0095] In an eighth aspect, the present application provides a computing device cluster, comprising multiple computing devices, the multiple computing devices including multiple processors and multiple memories, the multiple memories storing program instructions, and the multiple processors running the program instructions, so that the computing device cluster implements the server provided in the first aspect of the present application and any possible implementation thereof, and implements the server system provided in the first aspect of the present application and any possible implementation thereof.
[0096] In the ninth aspect, the present application provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium, and the computer-readable storage medium includes program instructions. When the program instructions are executed on a computing device, the computing device implements the server provided in the first aspect of the present application and any possible implementation thereof, and implements the server system provided in the first aspect of the present application and any possible implementation thereof.
[0097] In the tenth aspect, the present application provides a computer program product comprising instructions, which, when the computer program product is run on a computer, enables the computer to implement the server provided in the first aspect of the present application and any possible implementation thereof, and implement the server system provided in the first aspect of the present application and any possible implementation thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] FIG1 is a schematic structural diagram of an implementation scenario provided by an embodiment of the present application;
[0099] FIG2 is a schematic diagram of a deployment of basic resources provided in an embodiment of the present application;
[0100] FIG3 is a schematic diagram of the structure of a server based on cloud computing technology provided in an embodiment of the present application;
[0101] FIG4 is a schematic diagram of the structure of another server based on cloud computing technology provided in an embodiment of the present application;
[0102] FIG5 is a schematic diagram of the structure of another server based on cloud computing technology provided in an embodiment of the present application;
[0103] FIG6 is a schematic diagram of the structure of another server based on cloud computing technology provided in an embodiment of the present application;
[0104] FIG7 is a schematic diagram of the structure of another server based on cloud computing technology provided in an embodiment of the present application;
[0105] FIG8 is a schematic diagram of the structure of another server based on cloud computing technology provided in an embodiment of the present application;
[0106] FIG9 is a schematic structural diagram of a server system based on cloud computing technology provided in an embodiment of the present application;
[0107] FIG10 is a schematic structural diagram of another server system based on cloud computing technology provided in an embodiment of the present application;
[0108] FIG11 is a schematic structural diagram of another server system based on cloud computing technology provided in an embodiment of the present application;
[0109] FIG12 is a schematic structural diagram of another server based on cloud computing technology provided in an embodiment of the present application;
[0110] 13 is a schematic diagram showing that the functions of a back-end driver interception module and a front-end driver interception module provided in an embodiment of the present application can be optionally implemented by multiple functional units;
[0111] FIG14 is a flowchart of a method for creating a virtual machine based on cloud computing technology provided in an embodiment of the present application;
[0112] FIG15 is a schematic diagram of a cloud management platform provided in an embodiment of the present application;
[0113] FIG16 is a flowchart of another method for creating a virtual machine based on cloud computing technology provided in an embodiment of the present application;
[0114] FIG17 is a schematic diagram of a cloud management platform provided in an embodiment of the present application;
[0115] FIG18 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0116] FIG19 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0117] FIG20 is a schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0118] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0119] To facilitate understanding, the technology and background involved in the embodiments of this application are explained below.
[0120] An internet data center (IDC) is an internet-based facility that provides operational maintenance and related service systems for equipment that centrally collects, stores, processes, and transmits data. Conceptually, it can be understood as a public, commercial internet "computer room." It's also a form of professional IT service and a crucial infrastructure for the IT industry. An IDC is not only a service concept, but also a network concept. It forms part of the network's infrastructure, much like backbone networks and access networks, providing high-end data delivery and high-speed access services. Generally speaking, a tenant's offline IDC can be understood as their own offline computer room. This is where the tenant leverages existing internet communication lines and bandwidth resources to establish a standardized, professional-grade telecommunications computer room environment, offering a full range of services such as server hosting, leasing, and related value-added services. A cloud data center is an internet data center deployed using the infrastructure resources owned by cloud vendors.
[0121] A resource pool is a collection of various hardware and software resources involved in a cloud data center. Generally, resources in a resource pool can be divided into computing resources, storage resources, and network resources by type.
[0122] Physical machine (PM): A physical resource used to host virtualization technology. A host is also called a physical machine. Typically, a physical server is a host used to deploy virtual instances. A physical machine has multiple physical devices. For example, a physical server has physical devices such as a processor and memory. Multiple virtual instances can be deployed on a host. Multiple virtual instances deployed on the same host share the host's physical resources. Depending on the usage scenario, multiple virtual instances deployed on a host can optionally belong to the same tenant or different tenants.
[0123] Virtualization is a resource management technology. It abstracts and transforms a host's physical resources, such as computing, networking, and storage, to create a more tangible representation. This breaks down the barriers between the host's physical structure and allows tenants to utilize these resources in a more efficient manner than their original configuration. The resources created through virtualization are called virtualized resources, and they are not restricted by the existing physical resource configuration, location, or physical configuration.
[0124] Virtualized resources are usually provided to tenants in the form of virtual instances. Virtual instances can use the hardware resources of the host and run on the host's operating system. Applications run in virtual instances and are used to implement the tenant's business. The hardware resources of the host can be used by one or more tenants at the granularity of virtual instances. Different virtual instances are isolated from each other, allowing tenants to use physical resources conveniently and flexibly under the premise of secure isolation, and greatly improving the utilization of physical resources. Typically, virtual instances can be virtual machines, containers, or independent processes (such as functions). Virtual instances can also be called cloud servers (elastic compute services, ECS) or elastic instances (different cloud service providers have different names).
[0125] Virtual machine (VM): refers to a complete computer system with complete hardware system functions and running in a completely isolated environment, which is simulated through virtualization technology. Some subsets of the virtual machine's instructions can be processed in the host machine, and other instructions can be executed in a simulated manner. A virtual machine is also called a virtual server. A virtual machine can be regarded as a collection of several virtual devices, which have complete hardware system functions and run in a completely isolated environment. Virtual devices are virtualized through virtualization technology on the basis of physical devices that can share resources. For example, a virtual processor virtualized on the basis of a processor based on virtualization technology is a virtual device. For another example, a training card virtualized on the basis of a field-programmable gate array (FPGA) based on virtualization technology is also a virtual device. For example, the virtual machine in this application can be a kernel-based virtual machine (KVM). All the work that can be done in a server can be achieved in a virtual machine. When creating a virtual machine on a server, a portion of the physical machine's hard disk and memory capacity is used as the virtual machine's hard disk and memory capacity. Each virtual machine has its own independent hard disk and operating system, and tenants can operate the virtual machine just like they would on a server. The operating environments (such as virtual machine applications, operating systems, and virtual hardware) within each virtual machine are completely isolated. Communication between virtual machines requires network packets to be forwarded by the virtualization manager.
[0126] Containers use the namespace and cgroup technologies supported by the Linux kernel to isolate the application APP process and its dependent packages (the runtime environment bins / libs, specifically all the files required to run the APP) in an independent runtime environment. Containers provide a lightweight virtual runtime environment. Containers can be obtained by packaging all the code, libraries, and dependencies of a tenant's application into an image. When the image is executed, the image runs in the virtual runtime environment. In this case, the container is a runtime instance of the image, similar to a lightweight sandbox, which can be started, started, stopped, and deleted. The container infrastructure can be the server hardware or a virtual machine on the cloud (that is, containers can also be deployed in virtual machines). The operating system uses the Linux kernel and supports namespaces and cgroups. Namespaces are used to isolate processes, and cgroups are used to allocate process resources, specifically the virtual processors and memory allocated to the process. The container engine is similar to a virtual machine manager, running in the operating system to manage containers. Compared with virtual machines that come with their own operating systems, containers do not have operating systems. Containers run as processes in the host's operating system, so containers start faster than virtual machines. They are particularly suitable for lightweight applications, and a host can run thousands of containers (processes) simultaneously.
[0127] Resource pooling integrates multiple computing and storage resources into a unified resource pool for dynamic allocation and management. Resource pooling enables high-level resource sharing, improves resource utilization, simplifies resource management, and provides users with flexible, on-demand services.
[0128] In view of this, the embodiments of the present application provide a server based on cloud computing technology, a server system, and a corresponding virtual machine creation method and cloud management platform based on cloud computing technology. In the present application, a virtual machine is running in the server. A front-end driver interception module and an application are provided in the virtual machine. The front-end driver interception module can provide the virtual machine with a virtual graphics processing unit (VGPU) that can be accessed by the application in the virtual machine based on part of the computing power of the graphics processing unit (GPU). After the application accesses the VGPU, the front-end driver interception module can obtain the application's access request for the VGPU and provide the access request to the back-end driver interception module. After the back-end driver interception module obtains the access request from the front-end driver interception module, it can send the access request to the GPU driver module if it determines that the access request meets the preset conditions. The GPU driver module can call on part of the computing power of the GPU to process the access request.
[0129] In this application, the front-end driver interception module can capture access requests for the VGPU from applications in the virtual machine and, through the back-end driver interception module, provide these access requests to the GPU driver module. This allows the GPU driver module to utilize some of the GPU's computing power to process the access requests. This allows applications in the virtual machine to access the GPU through software, reducing the reliance on GPU hardware information and other implementation details when invoking the GPU to process access requests, and increasing the flexibility of GPU access for applications in the virtual machine.
[0130] It should be noted that although this application uses the example of an application in a virtual machine accessing the GPU to illustrate, it does not exclude the implementation method of accessing the GPU in this application, and can also be applied to the application running in other types of virtual instances to access the GPU, and can also be applied to access other types of hardware. For example, the implementation method of accessing the GPU in this application can also be applied to the application running in the container to respond to the access request to the GPU. For another example, the implementation method of accessing the GPU in this application can also be applied to the application running in the virtual instance to access other hardware such as the hard disk, disk, graphics card, intelligent processing unit (IPU) and data stream processor unit (DPU). Moreover, when the implementation method of accessing hardware in this application is applied to other types of virtual instances and / or other types of hardware, please refer to the implementation method of accessing the GPU by the application running in the virtual machine, and this article will not repeat it.
[0131] This article provides a detailed introduction to the technical solution of this application from multiple perspectives, including implementation scenarios, method flow, hardware devices, and software devices.
[0132] The following first illustrates an implementation scenario of the embodiment of the present application with examples.
[0133] FIG1 is a schematic diagram of the structure of an implementation scenario provided by an embodiment of the present application. As shown in FIG1 , the implementation scenario includes: a data center 1 and a client 2. A communication connection can be established between the data center 1 and the client 2 via a network. Optionally, the network can be the Internet or other networks, which are not limited by the embodiment of the present application. Tenants can interact with the data center 1 through the client 2. For example, a tenant can send information such as a cloud service request to the data center 1 through the client 2. The data center 1 is used to respond based on the information sent by the client 2.
[0134] A large amount of infrastructure owned by the cloud service provider is deployed in the data center 1, such as computing resources, storage resources, and network resources. For example, computing resources can be computing devices (such as servers, etc.) that can provide computing capabilities. As shown in Figure 1, the data center 1 includes a cloud management platform and infrastructure (not shown in Figure 1). The cloud management platform and the infrastructure are connected through the data center's internal network. The cloud management platform is used to manage the infrastructure. The infrastructure is used to provide public cloud services. The infrastructure includes multiple servers. Cloud services can be optionally deployed in the server. Cloud services are implemented by running virtual instances, so they are also called virtual instances deployed in the server for implementing tenant services. Tenants can send cloud service requests and related information to the server through the client 2 they use. The server can process the cloud service requests and related information and provide cloud services to the tenant based on the processed cloud service requests and related information. For example, the server can receive a virtual instance creation request provided by the tenant and create a virtual instance for the tenant in the server based on the virtual instance creation request.
[0135] The cloud management platform can be logically divided into the following functional areas: the tenant console, compute management service, network management service, storage management service, authentication service, and image management service. The tenant console provides an interface or application program interface (API) for interacting with tenants. The compute management service manages servers running virtual instances and bare metal servers. The network management service manages network services (such as gateways and firewalls). The storage management service manages storage services (such as data bucket services). The authentication service manages tenant accounts and passwords. The image management service manages images for virtual instances.
[0136] In the implementation scenario shown in Figure 1, multiple servers are deployed in a data center. The servers include hardware and software layers. The hardware layer is the standard configuration of the server. Hardware devices such as processors, memory, network cards, disks, and buses are deployed in the hardware layer. The software layer includes the operating system installed and running on the server. The operating system of the virtual machine is called the host operating system. The host operating system runs a virtual machine manager (also called a hypervisor). The role of the virtual machine manager is to implement computing virtualization, network virtualization, and storage virtualization for the virtual machines, and is responsible for managing the virtual machines.
[0137] The cloud management platform client runs within the virtual machine manager. The cloud management platform client receives control plane commands from the cloud management platform, creates virtual instances on servers based on these commands, and manages the virtual instances throughout their lifecycle. For example, the cloud management platform client monitors the hardware resource usage of the server in real time and reports this information to the cloud management platform. When the cloud management platform confirms the creation of a virtual instance on a server, it sends a virtual instance creation command to the cloud management platform client on that server. Upon receiving this command, the cloud management platform client creates the virtual instance on that server. This allows tenants to create, manage, log in to, and operate virtual instances in the data center through the cloud management platform.
[0138] Servers can be used to run virtual machines of varying specifications. Virtual machine specifications are categorized as general-purpose computing, memory-optimized, and ultra-large memory, with each type further detailed. After a tenant selects a virtual machine specification, the cloud management platform selects a server in the data center that supports it, determines if it has sufficient available hardware resources, and then creates a virtual machine with that specification on that server. Server configuration through the cloud management platform enables analysis and planning of server hardware resources. Based on the server's hardware performance, computing products corresponding to the physical hardware can be planned, such as virtual machines of varying specifications, to meet the differentiated needs of different tenants. Furthermore, the performance differences between virtual machines of varying specifications enable differentiated pricing strategies. For example, virtual instances with high performance specifications can be sold at a higher price, while those with standard performance specifications can be sold at a lower price, allowing tenants to purchase virtual instances on demand.
[0139] In one implementation, as shown in Figure 2, the location of basic resources in a data center can be described using cloud resource deployment regions and availability zones (AZs). Tenants can optionally deploy cloud services based on resources in specific regions and AZs. Regions are defined based on geographic location and network latency. Within a region, the same resource pool is used, which can be understood as sharing public services such as elastic computing, block storage, object storage, virtual private cloud (VPC) networks, Elastic Internet Protocol (EIP) addresses, and images. Regions are categorized as general regions and dedicated regions. General regions provide general cloud services to public tenants. Dedicated regions are dedicated regions that carry the same type of business or provide services to specific tenants. A region typically includes multiple AZs. AZs within a region are connected by high-speed fiber optic cables to meet tenants' needs for building high-availability systems across AZs. An AZ is a collection of one or more data centers, as shown in Figure 2. Computing, networking, and storage resources within an AZ are logically divided into multiple clusters.
[0140] Tenants can send instructions to the cloud management platform through the client 2 they use to create, manage, log in and operate virtual instances in the server, and use the cloud services provided by the virtual instances. For example, the cloud management platform can provide an access interface. The access interface can be optionally provided in the form of an interface or an API. Tenants can operate the client to remotely access the access interface to register a cloud account and password on the cloud management platform, and use the cloud account and password to log in to the cloud management platform. The cloud management platform can also authenticate the cloud account and password. After successful authentication, the tenant can further select and pay to purchase a virtual instance of specific specifications (processor, memory, disk) on the cloud management platform. After the tenant successfully pays for the virtual instance, the cloud management platform provides the tenant with the remote login account and password of the purchased virtual instance. The tenant can use the remote login account and password to remotely log in to the virtual instance on the client, install and run the tenant's application in the virtual instance, and implement the tenant's business through the application.
[0141] Client 2 may be a computer, a personal computer, a laptop computer, a mobile phone, a smart phone, a tablet computer, a cloud host, a portable mobile terminal, a multimedia player, an e-book reader, a wearable device, a smart home appliance, an artificial intelligence device, a smart wearable device, a smart vehicle-mounted device or an Internet of Things device, etc.
[0142] In one implementation, the functions of the cloud computing-based server, server system, and corresponding cloud computing-based virtual machine creation method and cloud management platform provided in the embodiments of the present application can be implemented by running an executable program on a computing device in data center 1. Furthermore, the executable program that implements the function can optionally be presented in the form of an application installation package. After the server installs the application installation package, the server can implement the function by running the executable program therein.
[0143] It should be understood that the above content is an illustrative description of the implementation scenarios involved in the embodiments of the present application and does not constitute a limitation on the implementation scenarios involved in the embodiments of the present application. A person of ordinary skill in the art will know that as business needs change, the implementation scenarios can be adjusted according to application requirements, and the embodiments of the present application do not make specific limitations on them.
[0144] The following first introduces the implementation of the server based on cloud computing technology provided by this application.
[0145] FIG3 is a schematic diagram of the structure of a server based on cloud computing technology provided by an embodiment of the present application. As shown in FIG3 , a first virtual machine and a second virtual machine are running on the server. The server is also provided with a back-end driver interception module, a GPU driver module, and a GPU.
[0146] The first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is configured to perform device simulation on a first portion of the GPU's computing power to generate a first VGPU accessible to the first application. The first application is configured to access the first VGPU. The first front-end driver interception module is further configured to receive a first access request from the first application for the first VGPU. The first application is configured to implement a user service.
[0147] The second virtual machine is provided with a second front-end driver interception module and a second application. The second front-end driver interception module is configured to perform device simulation on the second portion of the GPU's computing power to generate a second VGPU accessible to the second application. The second application is configured to access the second VGPU. The second front-end driver interception module is further configured to receive a second access request from the second application for the second VGPU. The second application is configured to implement user services.
[0148] The back-end driver interception module is used to obtain a first access request from the first front-end driver interception module and / or a second access request from the second front-end driver interception module, and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions, and / or send the second access request to the GPU driver module when it is determined that the second access request meets the preset conditions.
[0149] The GPU driver module is used to call a first portion of the computing power of the GPU to process a first access request, and / or call a second portion of the computing power of the GPU to process a second access request.
[0150] It should be noted that the running of the first virtual machine and the second virtual machine on the server is only an example, and the number of virtual machines deployed in the server can be adjusted based on application requirements, and the embodiments of the present application do not specifically limit it. For example, fewer or more servers can be optionally deployed in the server according to application requirements, such as deploying at least one virtual machine in the server. When at least one virtual machine is deployed in the server, the access request of the application program in any virtual machine to access the virtual hardware obtained based on the hardware can be passed to the back-end driver interception module in accordance with the method provided in the present application, and passed to the driver of the hardware through the back-end driver module, so that the driver of the hardware calls the capabilities of the hardware to process the access request. Among them, Figure 4 is a schematic diagram of deploying a virtual machine in the server.
[0151] After the GPU driver module uses the GPU's computing power to process the access request, if it receives the processing result generated by the GPU processing the access request, the GPU driver module is also used to provide the processing result to the backend driver interception module, which then returns the processing result to the application that triggered the access request. The various modules in the server play the following roles in returning the processing result:
[0152] The GPU driver module is used to obtain the first processing result generated by the first part of the GPU's computing processing capability to process the first access request, and / or obtain the second processing result generated by the second part of the GPU's computing processing capability to process the second access request, and send the first processing result and / or the second processing result to the back-end driver interception module.
[0153] The back-end driver interception module is used to obtain the first processing result and / or the second processing result from the GPU driver module, and send the first processing result to the first front-end driver interception module, and / or send the second processing result to the second front-end driver interception module.
[0154] The first front-end driver interception module is used to obtain the first processing result from the back-end driver interception module and provide the first processing result to the first application.
[0155] The second front-end driver interception module is used to obtain the second processing result from the back-end driver interception module and provide the second processing result to the second application.
[0156] For example, when a first application in a first virtual machine needs to use the GPU for rendering, it accesses the first VGPU. This access operation triggers a first access request instructing access to the GPU. The first access request carries data related to the image to be rendered. The first front-end driver interception module can receive the first access request and forward it to the back-end driver interception module. After receiving the first access request, if the first access request meets preset conditions, the back-end driver interception module forwards the first access request to the GPU driver module. After receiving the first access request, the GPU driver module invokes the first portion of the GPU's computing power to process the first access request. After obtaining the GPU's processing result of the first access request, the GPU driver module forwards the result to the back-end driver interception module. The processing result carries the video memory address of the rendering result, which is the result of the GPU rendering based on the relevant data carried in the first access request. After receiving the processing result, the back-end driver interception module forwards the processing result to the first front-end driver interception module. After receiving the processing result, the first front-end driver interception module sends the processing result to the first application, so that the first application can retrieve the rendering result from the video memory address carried in the processing result.
[0157] As can be seen, the access request triggered by the application in the virtual machine accessing the VGPU is sequentially transmitted to the GPU driver module in the server via the front-end driver interception module in the virtual machine and the back-end driver interception module in the server. The processing result generated by the GPU processing the access request is sequentially transmitted to the application in the virtual machine via the GPU driver module in the server, the back-end driver interception module in the server, and the front-end driver interception module in the virtual machine. Because both transmission processes are implemented in software, the process of calling the GPU to process the access request is less dependent on GPU hardware information and other implementation details, and even eliminates the need to rely on GPU hardware information and other implementation details. For example, it does not rely on GPU manufacturer hardware information, achieving decoupling from the GPU and improving the flexibility of GPU access for applications in the virtual machine. Furthermore, when the virtual machine uses part of the GPU's computing power, it is equivalent to virtualizing and partitioning the GPU. The software implementation of transmitting access requests and processing results in this application allows the virtualization of GPU partitioning to be divided beyond the GPU hardware control, supporting multiple scheduling strategies for GPU scheduling, and improving GPU scheduling flexibility.
[0158] In this application, there are multiple ways to implement the server. The following three implementations are used as examples to illustrate them.
[0159] In the first implementation, the first application runs directly in the first virtual machine, and the second application runs directly in the second virtual machine. As shown in Figure 3, this implementation presents a scenario where the first application, running directly in a virtual machine deployed within a server, accesses the server's GPU. Because virtual machine operations are managed by the hypervisor, and the GPU, GPU driver module, and backend driver interception module are managed by the server's operating system (OS), this scenario can be referred to as GPU access through the hypervisor and OS.
[0160] In this first implementation method, the VGPU used by the virtual machine deployed in the server is obtained based on part of the computing power of the GPU in the server, which is equivalent to virtualizing the GPU of the server. The present application forwards access requests and processing results between the GPU driver module and the application through the back-end driver interception module and the front-end driver interception module, thereby reducing the dependence on implementation details such as the GPU hardware information, and even eliminating the need to rely on implementation details such as the GPU hardware information, thereby improving the flexibility of the application in the virtual machine to access the GPU, and the server of the present application supports multiple scheduling strategies for scheduling the GPU, thereby improving the flexibility of scheduling the GPU.
[0161] In a second implementation, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount a first VGPU for the first application to access the first VGPU in the first container. A second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount a second VGPU for the second application to access the second VGPU in the second container. Figure 5 is a schematic diagram of a server provided in an embodiment of the present application. As shown in Figure 5, the scenario presented by this implementation is: a virtual machine is deployed in the server, a container is deployed in the virtual machine, and the application in the container has an access requirement for the GPU of the VGPU provided for it. Since the execution operations of the virtual machine and the container are all managed by the hypervisor, and the GPU, GPU driver module and back-end driver interception module are managed by the server's OS, this scenario can be called a GPU access scenario implemented by the hypervisor, OS and container.
[0162] Optionally, the number of containers set in the virtual machine can be adjusted according to application requirements. For example, a virtual machine can be provided with one container, or a virtual machine can be provided with multiple containers. For example, as shown in Figure 6, the first virtual machine is also provided with a third container, a third application runs in the third container, and the third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container. The first front-end driver interception module is used to perform device simulation on the third portion of the GPU's computing power to generate a third VGPU accessible to the third application. For the implementation method of the third application accessing the third VGPU, please refer to the implementation method of the first application accessing the first VGPU and the second application accessing the second VGPU, which will not be repeated here. Similarly, the second virtual machine is also provided with a fourth container (not shown in Figure 6), a fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth portion of the GPU's computing power to generate a fourth VGPU accessible to the fourth application. For the implementation method of the fourth application accessing the fourth VGPU, please refer to the implementation method of the first application accessing the first VGPU and the implementation method of the second application accessing the second VGPU, which will not be repeated here.
[0163] This second implementation not only has the effect of the first implementation, but also further divides the virtual machine on the basis of virtualizing the GPU, thereby realizing secondary division and comprehensive scheduling of the virtual machine after GPU resource division, further improving the flexibility of accessing the GPU and scheduling the GPU. It should be noted that running a container in a virtual machine is only an example, and other types of virtual instances can also be run in the virtual machine. It is known to those skilled in the art that as business needs change, the type of virtual instance running in the virtual machine can be adjusted according to application requirements, and the embodiments of the present application do not specifically limit it. For example, the way of deploying virtual instances in a server can also be further evolved on this basis.
[0164] It should be noted that the above two implementation methods of the server can be used separately or in combination. For example, when there are multiple virtual machines running in the server, each of the multiple virtual machines can be deployed in accordance with the first implementation method mentioned above, that is, the application program runs directly in the virtual machine. Alternatively, each of the multiple virtual machines can be deployed in accordance with the second implementation method mentioned above, that is, a container is provided in the virtual machine, and an application program runs in the container. Alternatively, some of the multiple virtual machines are deployed in accordance with the first implementation method mentioned above, and some of the virtual machines are deployed in accordance with the second implementation method mentioned above. By way of example, Figure 7 is a schematic diagram of a server provided in an embodiment of the present application. As shown in Figure 7, the second application program runs directly in the second virtual machine in the server. The first application program runs in the first container in the first virtual machine in the server, and the third application program runs in the third container.
[0165] When the back-end driver interception module determines that the access request meets the preset conditions, the back-end driver interception module forwards the access request to the GPU driver module. When the back-end driver interception module determines that the access request does not meet the preset conditions, the access request is not forwarded to the GPU driver module. At this time, the back-end driver interception module may return an error message to the virtual machine to which the access request belongs to indicate that the virtual machine cannot access the GPU. In one possible implementation, the back-end driver interception module determines that the access request meets the preset conditions, including: the back-end driver interception module determines that the video memory of the GPU required by the access request is not greater than the video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the access request is not greater than the specified computing power or the first computing power threshold. For example, the back-end driver interception module determines that the first access request meets the preset conditions, including: the back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first part of the computing power or the first computing power threshold. The back-end driver interception module determines that the second access request meets the preset conditions, including: the back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold; and / or, the back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0166] In one possible implementation, the backend driver interception module receives an access request, obtains the virtual machine initiating the access request, and determines whether the virtual machine meets the preset conditions for the GPU. Meeting the preset conditions can be indicated in various ways. For example, when the tenant has not configured GPU capabilities for the virtual machine, the access request from the virtual machine does not meet the preset conditions. Alternatively, when the virtual machine's usage quota for the GPU exceeds the quota configured by the tenant for the virtual machine, the access request from the virtual machine does not meet the preset conditions. An access request from a virtual machine refers to an access request generated by an application running in the virtual machine to access a VGPU. For example, when the usage quota configured by the tenant for the virtual machine is the available video memory quota for the virtual machine, the allocated video memory quota can be considered the maximum amount of video memory available to the virtual machine. This video memory quota can then be used as a video memory threshold. If the video memory requested by the virtual machine is not greater than the video memory threshold, the access request from the virtual machine is determined to meet the preset conditions. For example, the backend driver interception module determines that a first access request meets the preset conditions by determining that the video memory requested by the first access request is not greater than the first video memory threshold. The back-end driver interception module determines that the second access request meets the preset conditions, including: the back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than the second video memory threshold. For another example, when the usage quota configured by the tenant for the virtual machine is the quota of computing power that the virtual machine can use, the quota of computing power configured by the tenant for the virtual machine can be regarded as the maximum usage of computing power that the virtual machine can use, and the quota of computing power can be used as the computing power threshold. When the computing power required by the virtual machine is not greater than the computing power threshold, it is determined that the access request from the virtual machine meets the preset conditions. For example, the back-end driver interception module determines that the first access request meets the preset conditions, including: the back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first computing power threshold. The back-end driver interception module determines that the second access request meets the preset conditions, including: the back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second computing power threshold. For another example, since the VGPU is generated by the front-end driver interception module in the virtual machine by simulating part of the GPU's computing power, when judging whether the access request meets the preset conditions, it can also be judged whether the GPU's computing power required by the access request is greater than this part of the computing power. When the GPU's computing power required by the access request is not greater than this part of the computing power, it is determined that the access request meets the preset conditions. For example, the back-end driver interception module determines that the first access request meets the preset conditions, including: the back-end driver interception module determines that the GPU's computing power required by the first access request is not greater than the first part of the computing power. The back-end driver interception module determines that the second access request meets the preset conditions, including: the back-end driver interception module determines that the GPU's computing power required by the second access request is not greater than the second part of the computing power.Among them, the server records the GPU configured by the tenant for the virtual machine and the GPU's computing power, computing power threshold, video memory threshold and other specifications. Therefore, the back-end driver interception module can obtain this information and determine whether the access request from the virtual machine meets the preset conditions based on this information.
[0167] Similarly, when a container is also running in the virtual machine, the front-end driver interception module set in the virtual machine may optionally determine whether the access request meets the preset conditions after receiving the access request, and then provide the access request to the back-end driver interception module if the access request meets the preset conditions. If the access request does not meet the preset conditions, the virtual machine may return an error message to the container to indicate that the container cannot access the GPU. In one possible implementation, the virtual machine may optionally determine, on a container basis, whether the GPU memory required by the access request triggered by the application running in the container is greater than the memory threshold; and / or whether the GPU computing power required by the access request is greater than the partial computing power or computing power threshold of the GPU available for use by the VGPU accessible by the application. Here, the memory threshold is the amount of video memory configured by the tenant for the container, and the computing power threshold is the amount of computing power configured by the tenant for the container.
[0168] The following first describes the working principle of the front-end driver interception module in the virtual machine obtaining an access request triggered by an application running in the virtual machine and providing the access request to the back-end driver interception module.
[0169] Typically, when an application needs to access a GPU, it calls the interface provided by the application middleware to access that GPU, triggering an access request for the GPU. In this application, the front-end driver interception module in the virtual machine simulates part of the GPU's computing power to generate a virtual GPU accessible to applications running in the virtual machine. When an application calls the interface provided by the application middleware to access the GPU, it is actually accessing the VGPU generated based on this GPU simulation. Therefore, when an application calls the interface provided by the application middleware to access the VGPU, an access request is triggered. The VGPU can be considered a virtual driver file for the GPU provided in the virtual machine. When an application calls the interface provided by the application middleware to access the VGPU, it is actually accessing the virtual driver file for the GPU. Therefore, when an application calls the interface provided by the application middleware to access the virtual driver file, an access request is triggered. In one possible implementation, the GPU's virtual driver file is generated by the front-end driver interception module using virtualization technology to simulate the GPU. When the virtual machine runs the application directly, the virtual driver file is actually stored in the virtual machine. When the application runs in a container, the virtual driver file is actually stored in the container.
[0170] To respond to virtual driver files, the virtual machine also provides a virtual access interface for the driver file system. This virtual access interface is used to respond to accesses to the virtual driver file. When an application accesses the virtual driver file, triggering an access request, the driver file system processes the access request according to file access logic. However, in this application, the front-end driver interception module is configured to intercept the virtual access interface for the driver file system. Therefore, when an application accesses the GPU's virtual driver file, the front-end driver interception module can intercept the triggered access request. In one possible implementation, the driver file system's virtual access interface is virtualized using virtualization technology. When the virtual machine runs the application directly, the virtual access interface is actually provided within the virtual machine. When the application runs in a container, the virtual access interface is actually provided within the container. For example, when an application needs to access the VGPU, it can call an interface provided by the application middleware. For example, for a certain manufacturer's GPU, the application can call a CUDA-related interface. By calling this interface, the application can open the corresponding GPU's virtual driver file and perform file system access operations on it, such as ioctl operations or mmap operations. When an application performs an ioctl operation on the virtual driver file of the GPU, an ioctl access request is triggered, and the virtual access interface that responds to the ioctl access request is the virtual ioctl interface. The front-end driver interception module can intercept the ioctl access request by intercepting the virtual ioctl interface. When an application performs an mmap operation on the virtual driver file of the GPU, an mmap access request is triggered, and the virtual access interface that responds to the mmap access request is the virtual mmap interface. The front-end driver interception module can intercept the mmap access request by intercepting the virtual mmap interface. Among them, the mmap access request is used to indicate that data is mapped to the GPU's video memory. The ioctl access request is used to send control and configuration commands to the GPU.
[0171] After the front-end driver interception module intercepts the access request sent by the application, it needs to provide the access request to the back-end driver interception module. In a possible implementation, information is transmitted between the back-end driver interception module and the front-end driver interception module through one or more of the following methods: interrupt call or network connection.
[0172] The transmitting end in the back-end driver interception module and the front-end driver interception module transmits information to the receiving end thereof via an interrupt call, which means that the transmitting end sends an interrupt to the receiving end, and the interrupt carries the information to be transmitted. Taking the front-end driver interception module as the transmitting end and the back-end driver interception module as the receiving end as an example, since the virtual machines in the server are managed by the virtual machine manager in the server, the implementation process of the front-end driver interception module sending an access request to the back-end driver interception module via an interrupt call is as follows: the front-end driver interception module sends an interrupt to the virtual machine manager, and the virtual machine manager sends an interrupt to the back-end driver interception module. When the virtual machine manager is a virtual machine monitor (VMM), the interrupt call is a hypercall of the VMM. It should be noted that when information is transmitted between the back-end driver interception module and the front-end driver interception module via a hypercall, a customized hypercall type needs to be added to support this transmission. When the back-end driver interception module is the transmitting end and the front-end driver interception module is the receiving end, the implementation method of transmitting information between the two via interrupt calls should refer to the implementation method when the front-end driver interception module is the transmitting end and the back-end driver interception module is the receiving end, and will not be repeated here.
[0173] Transmitting information between the back-end driver interception module and the front-end driver interception module via a network connection means that a network connection is established between the back-end driver interception module and the front-end driver interception module, and the two transmit information to be transmitted via the network. The type of network connection between the two can be determined based on application requirements and is not specifically limited in the present embodiment.
[0174] The following describes the working principle of the backend driver interception module receiving the access request and forwarding the access request to the GPU driver module.
[0175] After the back-end driver interception module obtains the access request provided by the front-end driver interception module, it needs to send the access request to the real driver of the GPU that provides the VGPU for the application, that is, the GPU driver module, in order to achieve access to the GPU. After the front-end driver interception module provides the access request to the back-end driver interception module through a specified transmission method, the back-end driver interception module can obtain the access request using the transmission method corresponding to the specified transmission method. For example, after the front-end driver interception module provides the access request to the back-end driver interception module through an interrupt call, the back-end driver interception module can obtain the access request by receiving the interrupt sent by the virtual management component. For another example, after the front-end driver interception module sends the access request to the back-end driver interception module through the network, the back-end driver interception module can receive the access request through the network. It should be noted that the information transmission method used between the back-end driver interception module and the front-end driver interception module can be optionally pre-configured or determined by negotiation between the two.
[0176] After receiving the access request provided by the front-end driver interception module, the back-end driver interception module may optionally process the access request first and then forward the access request to the GPU driver module if it determines that the second access request meets the preset conditions. For example, after receiving the access request provided by the front-end driver interception module, the back-end driver interception module schedules GPU resources for the virtual machine that initiated the access request based on the access request, and sends the access request to the GPU driver module according to the scheduled GPU resources. For example, when the virtual machine's access request indicates access to the GPU, after receiving the access request, the back-end driver interception module schedules the video memory blocks that need to be allocated to the virtual machine in the GPU's video memory, and then sends the access request carrying information indicating the video memory blocks to the GPU driver module. In another possible implementation, after receiving the access request, the back-end driver interception module can also schedule the access request from the virtual machine to determine the timing of providing the access request to the GPU driver module. For example, the back-end driver interception module can schedule the access request according to a scheduling method such as task preemption, queuing, or multi-type hybrid scheduling to determine the order in which to send the multiple access requests received to the GPU driver module, and then send the access requests to the GPU driver module in sequence according to the order. It should be noted that the back-end driver interception module can also perform other processing on the access request, which is not specifically limited in the embodiment of the present application.
[0177] There are multiple ways for the back-end driver interception module to communicate with the GPU driver module. For example, the back-end driver interception module and the GPU driver module can optionally communicate through interrupts. For example, after receiving an access request, the back-end driver interception module sends an interrupt to the GPU driver module indicated by the access request to notify the GPU driver module of the access requirement indicated by the access request. It should be noted that the back-end driver interception module and the GPU driver module can also communicate through other means. The communication between the two through interrupts is an example, and the embodiments of the present application do not specifically limit the communication method between the two.
[0178] After the back-end driver interception module forwards the access request to the GPU driver module, the GPU driver module can then utilize the GPU's computing power to process the access request. After the GPU driver module obtains the GPU's processing result for the access request, it can transmit the processing result to the virtual machine along a path reverse to the path by which the access request was transmitted from the virtual machine to the GPU driver module. For example, the GPU driver module uses the communication method between it and the back-end driver interception module to provide the processing result to the back-end driver interception module. After obtaining the processing result, the back-end driver interception module uses the communication method between the back-end driver interception module and the front-end driver interception module to provide the processing result to the front-end driver interception module. After obtaining the processing result, the front-end driver interception module provides the processing result to the application. Furthermore, as shown in Figures 3 to 7, when the server includes a virtual machine manager, if information is transmitted between the back-end driver interception module and the front-end driver interception module via the virtual machine manager, the process of the back-end driver interception module transmitting the processing result to the front-end driver interception module is as follows: the back-end driver interception module provides the processing result to the virtual machine manager, and the virtual machine manager provides the processing result to the front-end driver interception module. For example, in response to information being transmitted between the back-end driver interception module and the front-end driver interception module via an interrupt call, the back-end driver interception module needs to send an interrupt indicating a processing result to the virtual machine manager, and the virtual machine manager sends an interrupt indicating a processing result to the front-end driver interception module. Regarding the communication method between the GPU driver module and the back-end driver interception module, the communication method between the back-end driver interception module and the front-end driver interception module, the communication method between the back-end driver interception module and the virtual machine manager, and the communication method between the virtual machine manager and the front-end driver interception module, please refer to the relevant description in the previous content, which will not be repeated here.
[0179] When the processing result includes a storage address, the back-end driver interception module and the front-end driver interception module also need to perform address conversion on the storage address so that the application can recognize the storage address. For example, in response to the processing result including the host physical address (HPA) of the video memory block allocated to the virtual machine in the video memory, the back-end driver interception module is further used to convert the host physical address into a client physical address (GPA) and provide the client physical address to the front-end driver interception module. The front-end driver interception module is also used to convert the client physical address into a client virtual address (GVA) and provide the client virtual address to the application. For example, when the operation that triggers the access request is an mmap operation, the processing result is the host physical address of the video memory block allocated to the virtual machine. However, because the address space of the video memory is different from that of the virtual machine, the host physical address in the GPU cannot be directly mmap-mapped to the client virtual address in the virtual machine. Therefore, after obtaining the processing result, the back-end driver interception module needs to map the host physical address in the processing result to the host virtual address (HVA), then convert the host virtual address to the client physical address and provide the client physical address to the front-end driver interception module. After obtaining the client physical address, the front-end driver interception module also needs to convert the client physical address to the client virtual address to achieve shared access to the video memory of the virtual machine and the server. Among them, the virtual address is the address in the virtual address space used to load program data during program execution. In other words, the virtual address is the address allocated to the application process during the application execution. The virtual address can be mapped to a physical storage block. The data indicated by the virtual address is recorded in the physical storage block to which it is mapped. The physical address is the address of the physical storage block.
[0180] When the GPU allocates video memory blocks to virtual machines, the virtual machine manager also needs to provide virtual video memory to the virtual machines based on virtualization technology. The purpose of virtualization technology is to provide virtual machines with a continuous physical video memory space starting at address 0, effectively isolating and scheduling video memory resources between virtual machines. This virtualization technology primarily involves the conversion of client virtual addresses -> client physical addresses -> host virtual addresses -> host physical addresses. In virtualization technology, a physical host often runs multiple virtual machines, each of which assumes exclusive access to the host's video memory space. Therefore, a virtual machine uses client physical addresses to represent its own video memory space, which the virtual machine considers to be continuous (i.e., it can be understood as a virtual machine that owns a complete physical video memory bank). A client virtual address is the address generated by the virtual machine's operating system mapping the client physical address. The virtual machine's operating system provides the client virtual address to processes or applications running on the virtual machine's operating system. Based on the mapping, the virtual machine's operating system can determine the mapping relationship between the client virtual address and the client physical address. In one possible implementation, the conversion of the client virtual address to the client physical address is performed by the virtual machine's operating system's page table. The host physical address is the actual physical video memory address, and the host virtual address is the address formed by the host operating system mapping the host physical address. The host operating system provides the host virtual address to the process (such as a virtual machine) on the operating system for use. The host operating system can obtain the mapping relationship between the host virtual address and the host physical address based on the mapping situation. In one possible implementation, the conversion of the host virtual address to the host physical address is implemented by the page table of the host operating system. Before the virtual machine manager simulates and obtains virtual video memory for the virtual machine, it needs to first determine the video memory block allocated to the virtual machine based on the specifications of the virtual machine, and then perform device simulation based on the video memory block allocated to the virtual machine to obtain virtual video memory. Therefore, the virtual machine manager can obtain the client physical address of the virtual video memory and the host virtual address of the video memory block allocated to it, and based on this, obtain the mapping relationship between the host virtual address and the client physical address. In one possible implementation, the conversion of the host virtual address to the client physical address is implemented by the page table of the virtual machine's operating system. Therefore, the present application can achieve the above conversion between the host physical address and the client virtual address.
[0181] In Figure 8, virtual machine 1 and virtual machine 2 are set up in the same server (hereinafter referred to as the host). The host's virtual machine manager sets the address range of the client physical address of virtual machine 1 to 0-5GB, which corresponds to the address range of the host physical address on the physical video memory, 1.5GB-4.5GB and 6.5GB-8.5GB. In addition, the host's virtual machine manager sets the address range of the client physical address of virtual machine 2 to 0-4GB, which corresponds to the address range of the host physical address on the physical video memory, 9GB-11GB and 13GB-15GB. Therefore, virtual machine 1 exclusively uses the address range of the client physical address of 0-5GB, and virtual machine 2 exclusively uses the address range of the client physical address of 0-4GB. The address range of the client physical address of 0-5GB and the address range of the client physical address of 0-4GB can both correspond to different address ranges of the host physical address on the physical video memory, thereby achieving isolation of the virtual machine video memory.
[0182] The following is an introduction to the implementation of the server system based on cloud computing technology provided by this application.
[0183] Figure 9 is a structural diagram of a server system based on cloud computing technology provided in an embodiment of the present application. As shown in Figure 9, the server system includes a first server and a second server. The first server is provided with a first back-end driver interception module, a first GPU driver module and a first GPU. A first virtual machine and a second virtual machine are running on the second server. A connection channel is established between the first server and the second server. In one possible implementation, the connection channel is implemented by a high-speed interconnection protocol, and the high-speed interconnection protocol includes a compute express link (CXL) protocol or a Lingqu bus (also known as UB bus) protocol. It should be noted that different servers can also be optionally connected to each other through other implementation methods, and the embodiments of the present application do not specifically limit them.
[0184] The first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is configured to perform device simulation on a first portion of the computing power of the first GPU to generate a first VGPU accessible to the first application. The first application is configured to access the first VGPU. The first front-end driver interception module is further configured to receive a first access request from the first application for the first VGPU and send the first access request to a connection channel. The first application is configured to implement a user service.
[0185] The second virtual machine is provided with a second front-end driver interception module and a second application. The second front-end driver interception module is configured to perform device simulation on the second portion of the computing power of the first GPU to generate a second VGPU accessible to the second application. The second application is configured to access the second VGPU. The second front-end driver interception module is further configured to receive a second first access request from the second application for the second VGPU and send the first access request to the connection channel. The second application is configured to implement user services.
[0186] The first back-end driver interception module is used to obtain a first access request from the first front-end driver interception module and / or a second first access request from the second front-end driver interception module from the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions, and / or send the second first access request to the first GPU driver module when it is determined that the second first access request meets the preset conditions.
[0187] The first GPU driver module is configured to call a first portion of the computing power of the first GPU to process a first access request, and / or call a second portion of the computing power of the first GPU to process a second first access request.
[0188] It should be noted that the running of the first virtual machine and the second virtual machine on the second server is only an example, and the number of virtual machines deployed in the second server can be determined based on application requirements, and the embodiments of the present application do not specifically limit it. For example, fewer or more servers can be deployed in the second server, such as deploying at least one virtual machine in the second server. When at least one virtual machine is deployed in the second server, the access request of the application program in any virtual machine to the virtual hardware obtained based on the hardware can be passed to the first back-end driver interception module in accordance with the method provided in the present application, and passed to the hardware driver through the first back-end driver module, so that the hardware driver calls the hardware's capabilities to process the access request.
[0189] After the first GPU driver module uses the computing power of the first GPU to process the access request, if it receives a result generated by the first GPU processing the access request, the first GPU driver module is further configured to provide the result to the first backend driver interception module, which then returns the result to the application that triggered the access request. The roles of the various modules in the server in returning the result are as follows:
[0190] The first GPU driver module is used to obtain the first processing result generated by the first part of the computing processing power of the first GPU to process the first access request, and / or obtain the second processing result generated by the second part of the computing processing power of the first GPU to process the second first access request, and send the first processing result and / or the second processing result to the first back-end driver interception module.
[0191] The first back-end driver interception module is used to obtain the first processing result and / or the second processing result from the first GPU driver module, send the first processing result to the connection channel, and / or send the second processing result to the connection channel.
[0192] The first front-end driver interception module is used to obtain the first processing result from the first back-end driver interception module from the connection channel, and provide the first processing result to the first application.
[0193] The second front-end driver interception module is used to obtain the second processing result from the first back-end driver interception module through the connection channel, and provide the second processing result to the second application.
[0194] For example, when a first application in a first virtual machine needs to use the first GPU for rendering, the first application accesses the first VGPU. This access operation triggers a first access request instructing access to the first GPU. The first access request carries relevant data of the image to be rendered. The first front-end driver interception module can obtain the first access request and forward it to the first back-end driver interception module. After receiving the first access request, if the first access request meets preset conditions, the first back-end driver interception module forwards the first access request to the first GPU driver module. After receiving the first access request, the first GPU driver module calls on the first portion of the computing power of the first GPU to process the first access request. After obtaining the processing result of the first GPU on the first access request, the first GPU driver module forwards the processing result to the first back-end driver interception module. The processing result carries the video memory address of the rendering result, which is the result of the first GPU rendering based on the relevant data carried in the first access request. After receiving the processing result, the first back-end driver interception module forwards the processing result to the first front-end driver interception module. After receiving the processing result, the first front-end driver interception module sends the processing result to the first application, so that the first application obtains the rendering result from the video memory address based on the video memory address carried by the processing result.
[0195] As can be seen from this, an access request triggered by an application running in a virtual machine on the second server accessing the VGPU can be transmitted to the first GPU driver module in the first server, sequentially via the front-end driver interception module in the virtual machine and the first back-end driver interception module in the first server. The processing result generated by the first GPU in the first server processing the access request can be transmitted to the application in the virtual machine, sequentially via the first GPU driver module, the first back-end driver interception module, and the front-end driver interception module in the virtual machine. Because both transmission processes are implemented in software, the virtual machine in the second server can reduce its reliance on GPU hardware information and other implementation details when calling the first GPU in the first server to process the access request. This can even eliminate the need to rely on GPU hardware information and other implementation details, such as hardware information from the GPU manufacturer. This allows for decoupling from the GPU and improves the flexibility of GPU access for applications in the virtual machine. Furthermore, when a virtual machine uses part of the GPU's computing power, it is equivalent to virtualizing and partitioning the GPU. The software implementation of transmitting access requests and processing results in this application allows the virtualization of GPU partitioning to be sharded beyond the GPU's hardware control, supporting a variety of GPU scheduling strategies and improving GPU scheduling flexibility.
[0196] In this application, there are multiple ways to implement the server system. The following three implementations are used as examples to illustrate them.
[0197] In a first implementation, the second server is not provided with a back-end driver interception module, a GPU driver module, and a GPU. As shown in FIG9 , the scenario presented by this implementation is that a virtual machine deployed inside the second server accesses the GPU possessed by the first server. For example, the scenario of this implementation may be a scenario in which the first server provides hardware resources for the first server. For example, this scenario is a pooling scenario in which GPUs are pooled. In this pooling scenario, a plurality of servers are deployed in the resource pool, each server having a GPU, a GPU driver module, and a back-end driver interception module, and the plurality of servers include the first server. After the second server applies for resources from the resource pool, the first server in the resource pool provides the GPU to the second server. This implementation can break through the limitation that the virtual machines currently deployed in the server can only access the GPU possessed by the server itself, and further improves the flexibility of accessing the GPU and the flexibility of scheduling the GPU.
[0198] The first implementation method can be divided into multiple situations, and this application uses the following situations as examples to illustrate it.
[0199] In the first case, the first application runs directly in the first virtual machine, and the second application runs directly in the second virtual machine. As shown in Figure 9, this implementation method presents a scenario in which the first application runs directly in the virtual machine deployed inside the second server and accesses the first GPU of the first server. Because the virtual machine execution operations in the second server are managed by the hypervisor, and the first GPU, first GPU driver module, and first back-end driver interception module in the first server are managed by the operating system (OS) of the first server, this scenario can be referred to as a scenario in which GPU access is achieved through the cross-server hypervisor and OS.
[0200] In this first implementation method, the VGPU used by the virtual machine deployed in the second server is obtained based on part of the computing power of the first GPU in the first server, which is equivalent to virtualizing the first GPU of the first server. The present application forwards the first access request and processing result between the first GPU driver module and the application running in the virtual machine in the second server through the first back-end driver interception module and the front-end driver interception module in the second server, thereby reducing the dependence on implementation details such as GPU hardware information, and even eliminating the need to rely on implementation details such as GPU hardware information, thereby improving the flexibility of application programs in virtual machines to access the GPU, and the server of the present application supports multiple scheduling strategies for GPU scheduling, thereby improving the flexibility of GPU scheduling.
[0201] In the second case, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container. A second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container. Figure 10 is a schematic diagram of a server system provided in an embodiment of the present application. As shown in Figure 10, the scenario presented by this implementation method is: a virtual machine is deployed in the second server, a container is deployed in the virtual machine, and the application in the container has an access requirement for the first GPU in the first server that provides the VGPU for it. Since the execution operation of the virtual machine in the second server is managed by the hypervisor, and the first GPU, the first GPU driver module and the first back-end driver interception module in the first server are managed by the OS of the first server, this scenario can be called a GPU access scenario implemented by the hypervisor, OS and container across servers.
[0202] Optionally, the number of containers set in the virtual machine can be adjusted according to application requirements. For example, a virtual machine can be provided with one container, or a virtual machine can be provided with multiple containers. For example, as shown in Figure 11, the first virtual machine is also provided with a third container, a third application runs in the third container, and the third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third portion of the computing power of the first GPU to generate a third VGPU accessible to the third application. For the implementation method of the third application accessing the third VGPU, please refer to the implementation method of the first application accessing the first VGPU and the second application accessing the second VGPU, which will not be repeated here. Similarly, the second virtual machine is also provided with a fourth container, a fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth portion of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application. For the implementation method of the fourth application accessing the fourth VGPU, please refer to the implementation method of the first application accessing the first VGPU and the implementation method of the second application accessing the second VGPU, which will not be repeated here.
[0203] The second case not only has the effect of the first case, but also further divides the virtual machine in the second server on the basis of virtualizing and dividing the first GPU in the first server, thereby realizing secondary division and comprehensive scheduling of the virtual machine in the second server based on the division of the first GPU resources, further improving the flexibility of accessing the GPU and the flexibility of scheduling the GPU. It should be noted that running a container in a virtual machine is only an example, and other types of virtual instances can also be run in the virtual machine. It is known to those skilled in the art that as business needs change, the type of virtual instance running in the virtual machine can be adjusted according to application requirements, and the embodiments of the present application do not specifically limit it. For example, the way of deploying virtual instances in servers can also be further evolved on this basis.
[0204] In the second implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. At this time, the access of the application running on the virtual machine in the second server to the VGPU obtained based on the second GPU is actually the access of the application running on the virtual machine in the server to the VGPU obtained by the GPU on the server itself. For its implementation process, please refer to the implementation process in the aforementioned server. For example, the first front-end driver interception module of the first virtual machine is used to perform device simulation on the first part of the computing power of the second GPU to generate a third VGPU accessible to the first application. The first front-end driver interception module used by the first application to access the third VGPU is also used to obtain the third first access request of the first application to the third VGPU. The second front-end driver interception module of the second virtual machine is also used to perform device simulation on the second part of the computing power of the second GPU to generate a fourth VGPU accessible to the second application. The second application is used to access the fourth VGPU. The second front-end driver interception module is also used to obtain the fourth first access request of the second application to the fourth VGPU. The second back-end driver interception module is configured to obtain a third first access request from the first front-end driver interception module and / or a fourth first access request from the second front-end driver interception module, and if the third first access request is determined to meet a preset condition, send the third first access request to the second GPU driver module, and / or, if the fourth first access request is determined to meet the preset condition, send the fourth first access request to the second GPU driver module. The second GPU driver module is configured to utilize the first portion of the computing power of the second GPU to process the third first access request, and / or utilize the second portion of the computing power of the second GPU to process the fourth first access request.
[0205] In this second implementation, the deployment of virtual machines and applications in the second server can also be categorized into various scenarios. This application uses the following scenarios as examples to illustrate these scenarios. In the first scenario, the first application runs directly in the first virtual machine, and the second application runs directly in the second virtual machine. In the second scenario, the first virtual machine is configured with a first container, in which the first application runs. The first container is used to mount a first VGPU, allowing the first application to access the first VGPU within the first container. The second virtual machine is configured with a second container, in which the second application runs. The second container is used to mount a second VGPU, allowing the second application to access the second VGPU within the second container. Furthermore, the number of containers configured in the virtual machine can be adjusted based on application requirements. For example, the virtual machine can be configured with one container, or multiple containers. For example, the first virtual machine is further configured with a third container, in which a third application runs. The third container is used to mount a third VGPU, allowing the third application to access the third VGPU within the third container. The first front-end driver interception module is configured to perform device simulation on a third portion of the computing power of the first GPU to generate a third VGPU accessible to the third application. Similarly, a fourth container is also provided in the second virtual machine, and a fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container. Among them, the second front-end driver interception module is used to simulate the fourth part of the computing power of the second GPU to generate a fourth VGPU that can be accessed by the fourth application. For the implementation methods of various situations in this second implementation method, please refer to the previous relevant descriptions accordingly, which will not be repeated here. In addition, the first server can also be deployed with a virtual machine, and the application running in the virtual machine can also access the first GPU or the second GPU, and the deployment method of the virtual machine and application in the first server and the method of accessing the GPU can all refer to the deployment method of the virtual machine and application in the second server and the relevant description of the method of accessing the GPU, which will not be repeated here.
[0206] It should be noted that the above two implementations of the server can be used separately or in combination. For example, when the server system includes multiple servers, each of the multiple servers can be deployed according to the first implementation method described above, or each of the multiple servers can be deployed according to the second implementation method described above, or some of the multiple servers can be deployed according to the first implementation method described above, and some of the servers can be deployed according to the second implementation method described above. Similarly, when multiple virtual machines are running on any server, the virtual machines in the multiple virtual machines can be deployed according to any of the various scenarios in the above two implementation methods. For example, the virtual machines can directly run applications, while some of the virtual machines can be provided with containers, and the containers can all run applications. Figure 12 is a schematic diagram of a server system provided in an embodiment of the present application. As shown in Figure 12, the server system includes a first server and a second server having a connection channel. The first server is provided with a first back-end driver interception module, a first GPU driver module, and a first GPU. The second server is provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first virtual machine in the second server directly runs the first application, and the second virtual machine in the second server is provided with a second container, and the second application runs in the second container. The fifth virtual machine in the first server directly runs the fifth application program, and a fifth front-end driver interception module is provided in the fifth virtual machine.
[0207] In the server system, the back-end driver interception module forwards the access request to the GPU driver module, and the method of forwarding the access request to the GPU driver by the candidate driver interception module in the aforementioned server can also refer to the method of forwarding the access request to the GPU driver by the back-end driver interception module. For example, when the back-end driver interception module determines that the access request meets the preset conditions, the back-end driver interception module forwards the access request to the GPU driver module. When the back-end driver interception module determines that the access request does not meet the preset conditions, the access request is not forwarded to the GPU driver module. At this time, the back-end driver interception module may return an error message to the virtual machine to which the access request belongs to indicate that the virtual machine cannot access the GPU. In one possible implementation, the first back-end driver interception module determines that the first access request meets the preset conditions, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold. The second back-end driver interception module determines that the second first access request meets a preset condition, including: the second back-end driver interception module determining that the GPU video memory required by the second first access request is not greater than a second video memory threshold; and / or the second back-end driver interception module determining that the GPU computing power required by the second first access request is not greater than a second partial computing power or a second computing power threshold. For the implementation of this process, please refer to the relevant description of the server above and will not be repeated here.
[0208] Similarly, when a container is also running in the virtual machine, the front-end driver interception module set in the virtual machine may optionally determine whether the access request meets the preset conditions after receiving the access request, and then provide the access request to the back-end driver interception module if the access request meets the preset conditions. If the access request does not meet the preset conditions, the virtual machine may return an error message to the container to indicate that the container cannot access the GPU. In one possible implementation, the virtual machine may optionally determine, on a container basis, whether the GPU memory required by the access request triggered by the application running in the container is greater than the memory threshold; and / or whether the GPU computing power required by the access request is greater than the partial computing power or computing power threshold of the GPU available for use by the VGPU accessible by the application. Here, the memory threshold is the amount of video memory configured by the tenant for the container, and the computing power threshold is the amount of computing power configured by the tenant for the container.
[0209] In this server system, an application triggers an access request, a front-end driver interception module obtains the access request, a back-end driver interception module processes the received access request and forwards the access request to a GPU driver module, the GPU driver module calls upon the computing power of the GPU to process the access request, and the GPU driver module transmits the processing result to the virtual machine along a reverse path of the path through which the access request is transmitted from the virtual machine to the GPU driver module. When the processing result includes a storage address, the implementation process of address conversion of the storage address can refer to the relevant description in the previous server and will not be repeated here.
[0210] Unlike the previous server implementation, the first backend driver interception module and the frontend driver interception module in the second server transmit information via one or more of the following methods: shared memory or a network connection. The following example uses information transmission between the first backend driver interception module and the first frontend driver interception module in the second server as an example.
[0211] The transmitting end in the first back-end driver interception module and the first front-end driver interception module transmits information to the receiving end therein through a shared memory, which means that a shared memory for transmitting information is pre-configured for the transmitting end and the receiving end. When the transmitting end needs to transmit information to the receiving end, the information to be transmitted is first stored in the shared memory. When the receiving end determines that there is data update in the shared memory, the updated data is obtained from the shared memory to obtain the information to be transmitted. Among them, the shared memory can be configured on the server where either the first back-end driver interception module or the first front-end driver interception module is located, or on a remote server other than the server where the first back-end driver interception module and the first front-end driver interception module are located. It should be noted that there are many ways to implement shared memory, and the technical personnel of this application can choose the implementation method they need to use according to their needs, and they are not given examples one by one here.
[0212] Transmitting information between the first back-end driver interception module and the first front-end driver interception module via a network connection means that a network connection is established between the first back-end driver interception module and the first front-end driver interception module, and the two transmit information to be transmitted via the network. The type of network connection between the two can be determined based on application requirements and is not specifically limited in this embodiment of the application.
[0213] In addition, the second server also includes: a virtual machine manager. Then information can be transmitted between the first back-end driver interception module and the front-end driver interception module in the second server through the virtual machine manager. For example, the virtual machine manager is used to obtain a first access request from the first front-end driver interception module and send the first access request to the connection channel, and obtain a second access request from the second front-end driver interception module and send the second access request to the connection channel. The virtual machine manager is used to obtain a first processing result from the first back-end driver interception module from the connection channel, and send the first processing result to the first front-end driver interception module, so that the first front-end driver interception module provides the first processing result to the first application, and obtain a second processing result from the first back-end driver interception module from the connection channel, and send the second processing result to the second front-end driver interception module, so that the second front-end driver interception module provides the second processing result to the second application.
[0214] The following three examples illustrate the implementation process of application access to GPU.
[0215] As shown in FIG3 , a first virtual machine is deployed in a server, and the GPU to be accessed is the GPU in the server. The implementation process of the first application in the first virtual machine accessing the GPU includes the following steps:
[0216] S11. The first front-end driver interception module in the first virtual machine virtualizes the access interface of the driver file system in the first virtual machine to obtain a virtual access interface of the driver file system, performs device simulation on the first part of the computing power of the GPU to generate a virtual driver file (i.e., the first VGPU) accessible to the first application.
[0217] S12. The first application runs in the first virtual machine. The first application performs system access to the virtual driver file of the GPU by calling an interface provided by the application middleware, thereby triggering an access request indicating access to the GPU.
[0218] S13. The first front-end driver interception module intercepts the virtual access interface, obtains an access request, and sends the access request to the back-end driver interception module via a hypercall.
[0219] S14. After receiving the access request transmitted by the hypercall, the Hypervisor in the server sends the access request to the backend driver interception module.
[0220] S15. After receiving the access request, the backend driver interception module first performs GPU global resource monitoring and scheduling based on the access request, and then forwards the access request to the GPU driver module in the server.
[0221] S16. The GPU driver module in the server accesses the GPU based on the access request to utilize part of the GPU's computing power to process the access request. After obtaining a processing result generated by the processing of the access request by the GPU's computing power, the GPU driver module transmits the processing result to the first application program along the reverse path of the above-mentioned path.
[0222] As shown in Figure 5, a first virtual machine and a second virtual machine are deployed in a server. A first container is deployed in the first virtual machine, and a second container is deployed in the second virtual machine. The GPU being accessed is the GPU in the server. The implementation process of a first application in the first container accessing the GPU includes the following steps:
[0223] S21. A first front-end driver interception module in a first virtual machine virtualizes an access interface of a driver file system in the first virtual machine to obtain a virtual access interface of the driver file system.
[0224] S22. When creating the first container, the container runtime performs device simulation on the first part of the GPU's computing power to generate a virtual driver file (i.e., a first VGPU) of the GPU that can be accessed by the first application, and embeds (mounts) the virtual driver file into the first container.
[0225] S23. The first application runs in the first container. The first application performs system access to the virtual driver file of the GPU in the container by calling the interface provided by the application middleware, thereby triggering an access request indicating access to the GPU.
[0226] S24. The first front-end driver interception module intercepts the virtual access interface, obtains an access request, monitors and schedules resources in the GPU virtual machine, and then sends the access request to the back-end driver interception module via a hypercall.
[0227] S25. After receiving the access request transmitted by the hypercall, the Hypervisor in the server sends the access request to the backend driver interception module.
[0228] S26. After receiving the access request, the backend driver interception module first performs GPU global resource monitoring and scheduling based on the access request, and then forwards the access request to the GPU driver module in the server.
[0229] S27: The GPU driver module in the server accesses the GPU based on the access request to invoke a portion of the GPU's computing power to process the access request. After obtaining a processing result generated by the processing of the access request by the GPU's computing power, the GPU driver module transmits the processing result to the first application program along a reverse path of the aforementioned path.
[0230] As shown in Figure 9, the server system includes a first server and a second server. The first server and the second server are connected via a high-speed interconnect protocol such as CXL / UB. A first virtual machine is deployed in the second server, and the GPU being accessed is the first GPU in the first server. The process of implementing the first virtual machine accessing the first GPU includes the following steps:
[0231] S31. The first front-end driver interception module in the first virtual machine virtualizes the access interface of the driver file system in the first virtual machine to obtain a virtual access interface of the driver file system, performs device simulation on the first part of the computing power of the first GPU to generate a virtual driver file of the first GPU (i.e., the first VGPU) that can be accessed by the first application.
[0232] S32: The first application runs in the first virtual machine. The first application performs system access to the virtual driver file of the first GPU by calling an interface provided by the application middleware, thereby triggering an access request indicating access to the first GPU.
[0233] S33. The first front-end driver interception module intercepts the virtual access interface, obtains the access request, and uses the CXL bus connected to the second server to access the memory shared by the first server to the second server, and stores the access request in the memory.
[0234] S34. The first back-end driver interception module in the first server obtains an access request by accessing the memory. After the first back-end driver interception module performs GPU global resource monitoring and scheduling based on the access request, the access request is forwarded to the first GPU driver module in the first server.
[0235] S35: The first GPU driver module accesses the first GPU based on the access request to utilize a portion of the computing power of the first GPU to process the access request. After obtaining a processing result generated by the processing of the access request by the portion of the computing power of the first GPU, the first GPU driver module transmits the processing result to the first application program along a reverse path of the aforementioned path.
[0236] In one implementation, the functions of the back-end driver interception module and the front-end driver interception module can optionally be implemented by multiple functional units. For example, as shown in FIG13 , the functions of the front-end driver interception module are implemented by the device simulation and interception system access logic unit, the front-end data transmission logic unit, and the front-end video memory mapping logic unit. The device simulation and interception system access logic unit is used to utilize virtualization technology to obtain a virtual access interface for the driver file system and a virtual driver file for the GPU, intercept access requests, provide the access requests to the front-end data transmission logic unit, and provide the processing results to the application. The front-end data transmission logic unit is used to transmit access requests to the back-end data transmission logic unit and provide the processing results obtained from the back-end data transmission logic unit to the device simulation and interception system access logic unit. The front-end video memory mapping logic unit is used to convert client physical addresses into client virtual addresses and provide the client virtual addresses to the application. The functions of the back-end driver interception module are implemented by the back-end data transmission logic unit, the back-end video memory mapping logic unit, and the resource access monitoring and scheduling logic unit. The back-end data transmission logic unit is used to receive the access request transmitted by the front-end data transmission logic unit, and forward the access request to the resource access monitoring and scheduling logic unit, as well as receive the processing result provided by the resource access monitoring and scheduling logic unit, and provide the processing result to the front-end data transmission logic unit. The back-end video memory mapping logic unit is used to convert the host physical address into the client physical address, and provide the client physical address to the front-end driver interception module. The resource access monitoring and scheduling logic unit processes the access request, and then forwards the access request to the GPU driver module after the processing result indicates that the access request needs to be forwarded to the GPU driver module, and after receiving the processing result provided by the GPU driver module, provides the processing result to the back-end data transmission logic unit.
[0237] It should be noted that compared with the existing API interception methods in the industry, this application intercepts the file system access interface at the driver level, and the interception level is different. Moreover, the access request intercepted by this application is transmitted from the virtual machine to the accessed hardware, which does not occupy the tenant network and does not affect the network used by the tenant. In addition, this application can also be applied to access to other types of resources. For example, the object of access can be not only physical resources, but also virtual resources (such as VGPU). At the same time, the method of charging for resources can also be flexibly selected. For example, it can be charged according to the amount of divided resources, or it can be charged according to the actual usage of resources. The embodiments of this application do not make specific limitations on this.
[0238] As can be seen above, the improvements to the server system of this application primarily lie in the process by which the server and virtual machines in the server system access the GPU. Corresponding to the aforementioned server, embodiments of this application also provide a method for creating a virtual machine based on cloud computing technology. This method is applied to a cloud management platform, which is used to manage infrastructure, including multiple servers. The implementation process of this method for creating a virtual machine based on cloud computing technology is described below.
[0239] Figure 14 is a flow chart of a method for creating a virtual machine based on cloud computing technology provided by an embodiment of the present application. As shown in Figure 14, the method for creating a virtual machine based on cloud computing technology includes the following steps:
[0240] Step 1401: Obtain a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specifications of a first VGPU of a first virtual machine to be created.
[0241] When a tenant needs to create a virtual machine based on the infrastructure managed by the cloud management platform, they can perform specified operations on the client used by the tenant to trigger a virtual machine creation request, so that the cloud management platform can create a virtual machine for the tenant under the instruction of the virtual machine creation request. In one possible implementation, the cloud management platform can provide the tenant with a virtual machine creation interface, and the tenant can trigger the virtual machine creation request based on the virtual machine creation interface. The virtual machine creation request carries the specifications of the first virtual machine to be created. After the tenant triggers the virtual machine creation request, the cloud management platform can obtain the virtual machine creation request through the virtual machine creation interface and obtain the specifications of the first virtual machine to be created from the virtual machine creation request.
[0242] For example, the virtual machine creation interface is implemented through one or more of the following: an application programming interface (API), an interaction template, and a configuration interface. Among them, the interaction template is a template provided by the cloud management platform to tenants for implementing different functions. When a tenant needs to use a certain function, the tenant can download a template for implementing the function, add relevant information of the tenant to the template, and then feed back the template with the relevant information of the tenant to the cloud management platform. After receiving the template with the relevant information of the tenant, the cloud management platform can obtain the function that the template needs to implement and customize the function according to the information of the tenant. The configuration interface means that the tenant can operate in the configuration interface to indicate the function that the tenant needs to implement.
[0243] Step 1402: Select a target server from multiple servers that can provide the specifications of the first VGPU.
[0244] After the cloud management platform obtains the specifications of the first VGPU of the first virtual machine to be created, it can select a server in the infrastructure that can provide the specifications of the first VGPU to obtain a target server.
[0245] Step 1403: Create a first virtual machine on the target server.
[0246] After the cloud management platform selects a target server in the infrastructure that can provide the specifications of the first VGPU, it can create a virtual machine in the target server that meets the specifications of the first VGPU. In addition, to ensure that the application running in the first virtual machine can access the first VGPU using the method provided in this application, it is also necessary to set up a first front-end driver interception module and a first application in the first virtual machine. The first front-end driver interception module is used to virtualize the access interface of the driver file system in the first virtual machine to obtain a virtual access interface of the driver file system, and to perform device simulation on the first portion of the computing power of the target server's GPU to generate a virtual driver file (i.e., the first VGPU) accessible to the first application. The parameters of the first VGPU match the specifications of the first VGPU. For example, when the specifications of the first VGPU indicate multiple parameters (such as computing power, video memory size, video memory bit width, and video memory bandwidth), the multiple parameters of the first VGPU obtained by the device virtualization correspond one-to-one with the multiple parameters indicated by the specifications, and any of the multiple parameters of the first VGPU is equal to or slightly greater than the corresponding parameter indicated by the specifications. At the same time, if the target server is not equipped with a back-end driver interception module, it is also necessary to configure a back-end driver interception module in the target server. The first front-end driver interception module is further configured to obtain a first access request from the first application to the first VGPU. The back-end driver interception module is configured to obtain the first access request from the first front-end driver interception module and, if it determines that the first access request meets a preset condition, send the first access request to the GPU driver module. Accordingly, the GPU driver module is configured to utilize a first portion of the GPU's computing power to process the first access request.
[0247] In a possible implementation, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container.
[0248] In one possible implementation, a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the GPU's computing power to generate a third VGPU that can be accessed by the third application.
[0249] In a possible implementation, the first application program runs directly in the first virtual machine.
[0250] In one possible implementation, the GPU driver module is also used to obtain the first processing result generated by the first part of the GPU's computing processing power to process the first access request, and send the first processing result to the back-end driver interception module; the back-end driver interception module is also used to obtain the first processing result from the GPU driver module, and send the first processing result to the first front-end driver interception module; the first front-end driver interception module is also used to obtain the first processing result from the back-end driver interception module, and provide the first processing result to the first application.
[0251] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module, and send the first access request to the back-end driver interception module.
[0252] In one possible implementation, the back-end driver interception module determines that the first access request meets preset conditions, including: the back-end driver interception module determines that the GPU video memory required by the first access request is not greater than the first video memory threshold; and / or, the back-end driver interception module determines that the GPU computing power required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0253] In one possible implementation, a second virtual machine is also provided on the target server, and a second front-end driver interception module and a second application are provided in the second virtual machine. The second front-end driver interception module is used to simulate the second part of the GPU's computing power to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is also used to obtain a second access request from the second application for the second VGPU; the back-end driver interception module is used to obtain the second access request from the second front-end driver interception module, and send the second access request to the GPU driver module when it is determined that the second access request meets preset conditions; the GPU driver module is used to call the second part of the GPU's computing power to process the second access request.
[0254] In one possible implementation, a second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
[0255] In a possible implementation, the second application program runs directly in the second virtual machine.
[0256] In one possible implementation, the GPU driver module is also used to obtain the second processing result generated by the second part of the GPU's computing processing power to process the second access request, and send the second processing result to the back-end driver interception module; the back-end driver interception module is also used to obtain the second processing result from the GPU driver module, and send the second processing result to the second front-end driver interception module; the second front-end driver interception module is also used to obtain the second processing result from the back-end driver interception module, and provide the second processing result to the second application.
[0257] In a possible implementation, the target server further includes:
[0258] The virtual machine manager is configured to obtain a second access request from the second front-end driver interception module and send the second access request to the back-end driver interception module.
[0259] In one possible implementation, the backend driver interception module determines that the second access request meets a preset condition, including:
[0260] The back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold; and / or, the back-end driver interception module determines that the GPU computing power required by the second access request is not greater than the second partial computing power or the second computing power threshold.
[0261] The above describes the virtual machine creation method based on cloud computing technology in an embodiment of the present application. Corresponding to the above method, an embodiment of the present application also provides a cloud management platform. The cloud management platform is used to manage infrastructure. The infrastructure includes multiple servers. Figure 15 is a structural diagram of a cloud management platform provided by an embodiment of the present application. Based on the following multiple components shown in Figure 15, the cloud management platform shown in Figure 15 can perform all or part of the operations shown in Figure 14 above. It should be understood that the device may include more additional components than the components shown or omit some of the components shown therein, and the embodiment of the present application does not limit this. As shown in Figure 15, the cloud management platform 150 may include:
[0262] The acquiring unit 1501 is configured to acquire a virtual machine creation request input by a tenant, where the virtual machine creation request carries specifications of a first VGPU of a first virtual machine to be created.
[0263] The selection unit 1502 is configured to select a target server that can provide the specifications of the first VGPU from a plurality of servers.
[0264] The creating unit 1503 is configured to create a first virtual machine on the target server.
[0265] Among them, the target server is provided with a back-end driver interception module, a GPU driver module and a GPU, and the first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to simulate the first part of the GPU's computing power to generate a first VGPU accessible to the first application. The first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request from the first application for the first VGPU; the back-end driver interception module is used to obtain the first access request from the first front-end driver interception module, and send the first access request to the GPU driver module when it is determined that the first access request meets the preset conditions; the GPU driver module is used to call the first part of the GPU's computing power to process the first access request.
[0266] In a possible implementation, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container.
[0267] In one possible implementation, a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the GPU's computing power to generate a third VGPU that can be accessed by the third application.
[0268] In a possible implementation, the first application program runs directly in the first virtual machine.
[0269] In one possible implementation, the GPU driver module is also used to obtain the first processing result generated by the first part of the GPU's computing processing power to process the first access request, and send the first processing result to the back-end driver interception module; the back-end driver interception module is also used to obtain the first processing result from the GPU driver module, and send the first processing result to the first front-end driver interception module; the first front-end driver interception module is also used to obtain the first processing result from the back-end driver interception module, and provide the first processing result to the first application.
[0270] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module, and send the first access request to the back-end driver interception module.
[0271] In one possible implementation, the back-end driver interception module determines that the first access request meets preset conditions, including: the back-end driver interception module determines that the GPU video memory required by the first access request is not greater than the first video memory threshold; and / or, the back-end driver interception module determines that the GPU computing power required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0272] In one possible implementation, a second virtual machine is also provided on the target server, and a second front-end driver interception module and a second application are provided in the second virtual machine. The second front-end driver interception module is used to simulate the second part of the GPU's computing power to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is also used to obtain a second access request from the second application for the second VGPU; the back-end driver interception module is used to obtain the second access request from the second front-end driver interception module, and send the second access request to the GPU driver module when it is determined that the second access request meets preset conditions; the GPU driver module is used to call the second part of the GPU's computing power to process the second access request.
[0273] In one possible implementation, a second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
[0274] In a possible implementation, the second application program runs directly in the second virtual machine.
[0275] In one possible implementation, the GPU driver module is also used to obtain the second processing result generated by the second part of the GPU's computing processing power to process the second access request, and send the second processing result to the back-end driver interception module; the back-end driver interception module is also used to obtain the second processing result from the GPU driver module, and send the second processing result to the second front-end driver interception module; the second front-end driver interception module is also used to obtain the second processing result from the back-end driver interception module, and provide the second processing result to the second application.
[0276] In a possible implementation, the target server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module, and send the second access request to the back-end driver interception module.
[0277] In one possible implementation, the back-end driver interception module determines that the second access request meets preset conditions, including: the back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold; and / or the back-end driver interception module determines that the GPU computing power required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0278] For the detailed working procedures of acquisition unit 1501, selection unit 1502, and creation unit 1503, please refer to the description in the previous method embodiment. For example, acquisition unit 1501 uses the aforementioned step 1401 to obtain a virtual machine creation request input by a tenant. Selection unit 1502 uses the aforementioned step 1402 to select a target server from multiple servers that can provide the specifications of the first VGPU. Creation unit 1503 uses the aforementioned step 1403 to create the first virtual machine on the target server. This embodiment of the present application will not be repeated here.
[0279] Corresponding to the aforementioned server system, embodiments of the present application also provide a method for creating a virtual machine based on cloud computing technology. This method is applied to a cloud management platform, which is used to manage infrastructure, including multiple servers. The implementation process of this method for creating a virtual machine based on cloud computing technology is described below.
[0280] Figure 16 is a flow chart of a method for creating a virtual machine based on cloud computing technology provided by an embodiment of the present application. As shown in Figure 16, the method for creating a virtual machine based on cloud computing technology includes the following steps:
[0281] Step 1601: Obtain a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specifications of a first VGPU of a first virtual machine to be created.
[0282] For the implementation process of step 1601, please refer to the implementation process of the aforementioned step 1401, which will not be repeated here.
[0283] Step 1602: Select a second server from multiple servers that can provide the specifications of the first VGPU.
[0284] After the cloud management platform obtains the specifications of the first VGPU of the first virtual machine to be created, it can select a server in the infrastructure that can provide the specifications of the first VGPU to obtain the second server.
[0285] Step 1603: Create a first virtual machine on the second server.
[0286] After the cloud management platform selects a second server in the infrastructure that can provide the specifications of the first VGPU, it can create a virtual machine in the second server that meets the specifications of the first VGPU. Furthermore, to ensure that the application running in the first virtual machine can access the first VGPU using the method provided in this application, it is also necessary to set up a first front-end driver interception module and the first application in the first virtual machine, and use the first front-end driver interception module to virtualize the access interface of the driver file system in the first virtual machine to obtain a virtual access interface of the driver file system. Furthermore, the cloud management platform selects a first server, performs device simulation on the first portion of the computing power of the GPU of the first server, and generates a virtual driver file (i.e., the first VGPU) that can be accessed by the first application. The parameters of the first VGPU match the specifications of the first VGPU. For example, when the specifications of the first VGPU indicate multiple parameters (such as computing power, video memory size, video memory bit width, and video memory bandwidth), the multiple parameters of the first VGPU obtained by the device virtualization correspond one-to-one with the multiple parameters indicated by the specifications, and any one of the multiple parameters of the first VGPU is equal to or slightly greater than the corresponding parameter indicated by the specifications. Furthermore, if the first server is not equipped with a back-end driver interception module, it is also necessary to configure a first back-end driver interception module in the first server. The first front-end driver interception module is used to simulate a portion of the computing power of the first GPU to generate a first VGPU accessible to a first application. The first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request from the first application for the first VGPU and send the first access request to the connection channel. The first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel and send the first access request to the first GPU driver module if it is determined that the first access request meets preset conditions. Accordingly, the first GPU driver module is used to call upon a portion of the computing power of the first GPU to process the first access request.
[0287] In one possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further configured to simulate the first portion of the computing power of the second GPU to generate a third VGPU accessible to the first application, the first application being configured to access the third VGPU, and the first front-end driver interception module being configured to obtain a third access request from the first application for the third VGPU; the second back-end driver interception module is configured to obtain the third access request from the first front-end driver interception module, and upon determining that the third access request meets preset conditions, the third access request is sent to the second GPU driver module; and the second GPU driver module is configured to invoke the first portion of the computing power of the second GPU to process the third access request.
[0288] In a possible implementation, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container.
[0289] In one possible implementation, a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application.
[0290] In a possible implementation, the first application program runs directly in the first virtual machine.
[0291] In one possible implementation, the first GPU driver module is further used to obtain a first processing result generated by the first part of the computing processing power of the first GPU to process the first access request, and send the first processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the first processing result from the first GPU driver module, and send the first processing result to the connection channel; the first front-end driver interception module is further used to obtain the first processing result from the first back-end driver interception module from the connection channel, and provide the first processing result to the first application.
[0292] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel.
[0293] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, which includes a computing express link (CXL) or a Lingqu UB bus protocol.
[0294] In one possible implementation, the first back-end driver interception module determines that the first access request meets preset conditions, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0295] In one possible implementation, a second virtual machine is also provided on the second server, and a second front-end driver interception module and a second application are provided in the second virtual machine. The second front-end driver interception module is used to simulate the second part of the computing power of the first GPU to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is also used to obtain a second access request from the second application for the second VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the second access request from the second front-end driver interception module from the connection channel, and send the second access request to the first GPU driver module when it is determined that the second access request meets the preset conditions; the first GPU driver module is used to call the second part of the computing power of the first GPU to process the second access request.
[0296] In one possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The second front-end driver interception module of the second virtual machine is further used to simulate the second portion of the computing power of the second GPU to generate a fourth VGPU accessible to the second application, the second application is used to access the fourth VGPU, and the second front-end driver interception module is further used to obtain a fourth access request from the second application for the fourth VGPU; the second back-end driver interception module is used to obtain the fourth access request from the second front-end driver interception module, and send the fourth access request to the second GPU driver module if it is determined that the fourth access request meets preset conditions; the second GPU driver module is used to call the second portion of the computing power of the second GPU to process the fourth access request.
[0297] In one possible implementation, a second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
[0298] In one possible implementation, a fourth container is further provided in the second virtual machine, a fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application.
[0299] In a possible implementation, the second application program runs directly in the second virtual machine.
[0300] In one possible implementation, the first GPU driver module is further used to obtain the second processing result generated by the second part of the computing processing power of the first GPU to process the second access request, and send the second processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the second processing result from the first GPU driver module, and send the second processing result to the connection channel; the second front-end driver interception module is further used to obtain the second processing result from the first back-end driver interception module from the connection channel, and provide the second processing result to the second application.
[0301] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module and send the second access request to the connection channel.
[0302] In one possible implementation, the second back-end driver interception module determines that the second access request meets preset conditions, including: the second back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold; and / or, the second back-end driver interception module determines that the GPU computing power required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0303] The above describes the virtual machine creation method based on cloud computing technology in an embodiment of the present application. Corresponding to the above method, an embodiment of the present application also provides a cloud management platform. The cloud management platform is used to manage infrastructure. The infrastructure includes multiple servers. Figure 17 is a structural diagram of a cloud management platform provided by an embodiment of the present application. Based on the following multiple components shown in Figure 17, the cloud management platform shown in Figure 17 can perform all or part of the operations shown in Figure 16 above. It should be understood that the device may include more additional components than the components shown or omit some of the components shown therein, and the embodiment of the present application does not limit this. As shown in Figure 17, the cloud management platform 170 may include:
[0304] The acquiring unit 1701 is configured to acquire a virtual machine creation request input by a tenant, where the virtual machine creation request carries specifications of a first VGPU of a first virtual machine to be created.
[0305] The selection unit 1702 is configured to select a second server that can provide the specifications of the first VGPU from a plurality of servers.
[0306] The creating unit 1703 is configured to create a first virtual machine on the second server.
[0307] Among them, the multiple servers also include a first server, the first server is provided with a first back-end driver interception module, a first GPU driver module and a first GPU, a connection channel is established between the first server and the second server, and a first front-end driver interception module and a first application are provided in the first virtual machine. The first front-end driver interception module is used to simulate part of the computing power of the first GPU to generate a first VGPU accessible to the first application, and the first application is used to access the first VGPU. The first front-end driver interception module is also used to obtain a first access request from the first application for the first VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions; the first GPU driver module is used to call part of the computing power of the first GPU to process the first access request.
[0308] In one possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further configured to simulate the first portion of the computing power of the second GPU to generate a third VGPU accessible to the first application, the first application being configured to access the third VGPU, and the first front-end driver interception module being configured to obtain a third access request from the first application for the third VGPU; the second back-end driver interception module is configured to obtain the third access request from the first front-end driver interception module, and upon determining that the third access request meets preset conditions, the third access request is sent to the second GPU driver module; and the second GPU driver module is configured to invoke the first portion of the computing power of the second GPU to process the third access request.
[0309] In a possible implementation, a first container is provided in the first virtual machine, a first application runs in the first container, and the first container is used to mount the first VGPU so that the first application can access the first VGPU in the first container.
[0310] In one possible implementation, a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU so that the third application can access the third VGPU in the third container, wherein the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the first GPU to generate a third VGPU accessible to the third application.
[0311] In a possible implementation, the first application program runs directly in the first virtual machine.
[0312] In one possible implementation, the first GPU driver module is further used to obtain a first processing result generated by the first part of the computing processing power of the first GPU to process the first access request, and send the first processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the first processing result from the first GPU driver module, and send the first processing result to the connection channel; the first front-end driver interception module is further used to obtain the first processing result from the first back-end driver interception module from the connection channel, and provide the first processing result to the first application.
[0313] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel.
[0314] In a possible implementation, the connection channel is implemented through a high-speed interconnection protocol, which includes a computing express link (CXL) or a Lingqu UB bus protocol.
[0315] In one possible implementation, the first back-end driver interception module determines that the first access request meets preset conditions, including: the first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than the first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first partial computing power or the first computing power threshold.
[0316] In one possible implementation, a second virtual machine is also provided on the second server, and a second front-end driver interception module and a second application are provided in the second virtual machine. The second front-end driver interception module is used to simulate the second part of the computing power of the first GPU to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is also used to obtain a second access request from the second application for the second VGPU and send the first access request to the connection channel; the first back-end driver interception module is used to obtain the second access request from the second front-end driver interception module from the connection channel, and send the second access request to the first GPU driver module when it is determined that the second access request meets the preset conditions; the first GPU driver module is used to call the second part of the computing power of the first GPU to process the second access request.
[0317] In one possible implementation, the second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The second front-end driver interception module of the second virtual machine is further used to simulate the second portion of the computing power of the second GPU to generate a fourth VGPU accessible to the second application, the second application is used to access the fourth VGPU, and the second front-end driver interception module is further used to obtain a fourth access request from the second application for the fourth VGPU; the second back-end driver interception module is used to obtain the fourth access request from the second front-end driver interception module, and send the fourth access request to the second GPU driver module if it is determined that the fourth access request meets preset conditions; the second GPU driver module is used to call the second portion of the computing power of the second GPU to process the fourth access request.
[0318] In one possible implementation, a second container is provided in the second virtual machine, a second application runs in the second container, and the second container is used to mount the second VGPU so that the second application can access the second VGPU in the second container.
[0319] In one possible implementation, a fourth container is further provided in the second virtual machine, a fourth application runs in the fourth container, and the fourth container is used to mount a fourth VGPU so that the fourth application can access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the second GPU to generate a fourth VGPU accessible to the fourth application.
[0320] In a possible implementation, the second application program runs directly in the second virtual machine.
[0321] In one possible implementation, the first GPU driver module is further used to obtain the second processing result generated by the second part of the computing processing power of the first GPU to process the second access request, and send the second processing result to the first back-end driver interception module; the first back-end driver interception module is further used to obtain the second processing result from the first GPU driver module, and send the second processing result to the connection channel; the second front-end driver interception module is further used to obtain the second processing result from the first back-end driver interception module from the connection channel, and provide the second processing result to the second application.
[0322] In a possible implementation, the second server further includes: a virtual machine manager, configured to obtain the second access request from the second front-end driver interception module and send the second access request to the connection channel.
[0323] In one possible implementation, the second back-end driver interception module determines that the second access request meets preset conditions, including: the second back-end driver interception module determines that the GPU video memory required by the second access request is not greater than the second video memory threshold; and / or, the second back-end driver interception module determines that the GPU computing power required by the second access request is not greater than the second part of the computing power or the second computing power threshold.
[0324] For the detailed working procedures of acquisition unit 1701, selection unit 1702, and creation unit 1703, please refer to the description in the previous method embodiment. For example, acquisition unit 1701 uses the aforementioned step 1601 to obtain a virtual machine creation request input by a tenant. Selection unit 1702 uses the aforementioned step 1602 to select a second server from multiple servers that can provide the specifications of the first VGPU. Creation unit 1703 uses the aforementioned step 1603 to create the first virtual machine on the second server. This embodiment of the present application will not be repeated here.
[0325] Among them, the acquisition unit 1501, the selection unit 1502 and the creation unit 1503, the acquisition unit 1701, the selection unit 1702 and the creation unit 1703 can all be implemented by software, or can be implemented by a GPU. For example, the implementation of the acquisition unit 1501 is described below using the acquisition unit 1501 as an example. Similarly, the implementation of the selection unit 1502 and the creation unit 1503, the acquisition unit 1701, the selection unit 1702 and the creation unit 1703 can refer to the implementation of the acquisition unit 1501.
[0326] As an example of a software functional unit, the acquisition unit 1501 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the acquisition unit 1501 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one cloud data center or multiple geographically close cloud data centers. Typically, a region may include multiple AZs.
[0327] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0328] As an example of a hardware functional unit, acquisition unit 1501 may include at least one computing device, such as a server. Alternatively, acquisition unit 1501 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0329] The multiple computing devices included in acquisition unit 1501 can be distributed in the same region or in different regions. The multiple computing devices included in acquisition unit 1501 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in acquisition unit 1501 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0330] It should be noted that, in other embodiments, any one of the obtaining unit 1501, the selecting unit 1502, the creating unit 1503, or the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703 may be used to execute any step in the method for creating a virtual machine based on cloud computing technology. The steps that the obtaining unit 1501, the selecting unit 1502, the creating unit 1503, or the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703 are responsible for implementing may be specified as needed. The full functionality of the cloud management platform is achieved by having the obtaining unit 1501, the selecting unit 1502, the creating unit 1503, or the obtaining unit 1701, the selecting unit 1702, and the creating unit 1703 respectively implement different steps in the method for creating a virtual machine based on cloud computing technology.
[0331] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the various components described above can refer to the corresponding contents in the aforementioned method embodiments and will not be repeated here.
[0332] The following is an example of the basic hardware structure involved in the embodiments of the present application.
[0333] This application also provides a computing device 1800. As shown in Figure 18, computing device 1800 includes a bus 1802, a processor 1804, a memory 1806, and a communication interface 1808. Processor 1804, memory 1806, and communication interface 1808 communicate with each other via bus 1802. Computing device 1800 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1800.
[0334] Bus 1802 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG18 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1802 may include a path for transmitting information between various components of computing device 1800 (e.g., memory 1806, processor 1804, and communication interface 1808).
[0335] The processor 1804 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0336] Memory 1806 is used to store computer programs, including an operating system and executable code (i.e., program instructions). Memory 1806 may include volatile memory, such as random access memory (RAM). Processor 1804 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0337] Memory 1806 stores executable program code, and processor 1804 executes the executable program code to implement the functions of the aforementioned acquisition unit 1501, selection unit 1502, and creation unit 1503, as well as the acquisition unit 1701, selection unit 1702, and creation unit 1703, respectively, thereby implementing any cloud computing technology-based server in this application. In other words, memory 1806 stores instructions for implementing any cloud computing technology-based server in this application.
[0338] The communication interface 1808 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1800 and other devices or a communication network.
[0339] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0340] As shown in Figure 19, the computing device cluster includes at least one computing device 1800. The memory 1806 in one or more computing devices 1800 in the computing device cluster may store the same instructions for implementing any server based on cloud computing technology in this application.
[0341] In some possible implementations, the memory 1806 of one or more computing devices 1800 in the computing device cluster may also store partial instructions for implementing any cloud computing technology-based server in the present application. In other words, the combination of one or more computing devices 1800 can jointly execute instructions for implementing any cloud computing technology-based server in the present application.
[0342] It should be noted that the memory 1806 in different computing devices 1800 in the computing device cluster can store different instructions, each used to execute a portion of the functions of the cloud management platform. In other words, the instructions stored in the memory 1806 in different computing devices 1800 can implement the functions of one or more modules in the acquisition unit 1501, selection unit 1502, and creation unit 1503, or the acquisition unit 1701, selection unit 1702, and creation unit 1703.
[0343] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. The network can be a wide area network (WAN) or a local area network (LAN), among others. FIG. 20 illustrates one possible implementation. As shown in FIG. 20 , two computing devices 1800A and 1800B are connected via a network. Specifically, the connection to the network is achieved via a communication interface in each computing device.
[0344] It should be understood that the functionality of the computing device 1800A shown in FIG20 may also be implemented by multiple computing devices 1800. Similarly, the functionality of the computing device 1800B may also be implemented by multiple computing devices 1800.
[0345] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection methods of the computing device clusters in Figures 19 and 20. However, the memory 1806 in one or more computing devices 1800 in this computing device cluster can store the same instructions for implementing any of the cloud computing technology-based servers described in this application.
[0346] In some possible implementations, the memory 1806 of one or more computing devices 1800 in the computing device cluster may also store partial instructions for implementing any cloud computing technology-based server in this application. In other words, a combination of one or more computing devices 1800 can jointly implement the instructions of any cloud computing technology-based server in this application.
[0347] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be executed on a computing device or stored in any available medium. When the computer program product is executed on at least one computing device, the at least one computing device implements any of the cloud computing technology-based servers described in the present application.
[0348] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to implement any server based on cloud computing technology in the present application.
[0349] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0350] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the raw data and executable code involved in this application were obtained with full authorization.
[0351] In the embodiments of the present application, the terms "first," "second," and "third" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. The term "at least one" refers to one or more, and the term "plurality" refers to two or more, unless otherwise expressly limited.
[0352] In this application, the term "and / or" simply describes an association between related objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0353] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A server based on cloud computing technology, characterized in that, A first virtual machine and a second virtual machine are running on the server. The server is further provided with a backend driver interception module, a graphics processing unit (GPU) driver module, and a GPU. Among them, the first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on the first part of the computing power of the GPU to generate a first virtual graphics processing unit (VGPU) accessible to the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU; the second virtual machine is provided with a second front-end driver interception module and a second application. The second front-end driver interception module is used to perform device simulation on the second part of the computing power of the GPU to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is further used to obtain a second access request of the second application for the second VGPU; the backend driver interception module is used to obtain the first access request from the first front-end driver interception module and / or the second access request from the second front-end driver interception module. When it is determined that the first access request meets the preset conditions, the first access request is sent to the GPU driver module, and / or, when it is determined that the second access request meets the preset conditions, the second access request is sent to the GPU driver module; the GPU driver module is used to call the first part of the computing power of the GPU to process the first access request, and / or, call the second part of the computing power of the GPU to process the second access request.
2. The server according to claim 1, wherein a first container is provided in the first virtual machine, and the first application runs in the first container. The first container is used to mount the first VGPU for the first application to access the first VGPU in the first container; a second container is provided in the second virtual machine, and the second application runs in the second container. The second container is used to mount the second VGPU for the second application to access the second VGPU in the second container.
3. The server according to claim 2, wherein a third container is further provided in the first virtual machine, and a third application runs in the third container. The third container is used to mount a third VGPU for the third application to access the third VGPU in the third container. Among them, the first front-end driver interception module is used to perform device simulation on the third part of the computing power of the GPU to generate the third VGPU accessible to the third application; And / or, a fourth container is further provided in the second virtual machine, and a fourth application runs in the fourth container. The fourth container is used to mount a fourth VGPU for the fourth application to access the fourth VGPU in the fourth container. Wherein, the second front-end driver interception module is used to perform device simulation on the fourth part of the computing power of the GPU to generate the fourth VGPU accessible to the fourth application.
4. The server according to claim 1, wherein The first application runs directly in the first virtual machine, and the second application runs directly in the second virtual machine.
5. The server according to any one of claims 1 to 4, wherein The GPU driver module is further configured to obtain the first processing result generated by processing the first access request using the first part of the computing and processing power of the GPU, and / or obtain the second processing result generated by processing the second access request using the second part of the computing and processing power of the GPU, and send the first processing result and / or the second processing result to the back-end driver interception module; The back-end driver interception module is further configured to obtain the first processing result and / or the second processing result from the GPU driver module, send the first processing result to the first front-end driver interception module, and / or send the second processing result to the second front-end driver interception module; The first front-end driver interception module is further configured to obtain the first processing result from the back-end driver interception module and provide the first processing result to the first application; The second front-end driver interception module is further configured to obtain the second processing result from the back-end driver interception module and provide the second processing result to the second application.
6. The server according to any one of claims 1 to 4, characterized in that The server further includes: A virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the back-end driver interception module, obtain the second access request from the second front-end driver interception module, and send the second access request to the back-end driver interception module.
7. The server according to any one of claims 1 to 6, wherein The back-end driver interception module determines that the first access request meets the preset conditions, including: The back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or the back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first part of the computing power or a first computing power threshold; The back-end driver interception module determines that the second access request meets the preset conditions, including: The back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than a second video memory threshold; and / or the back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or a second computing power threshold.
8. A server system based on cloud computing technology, characterized in that, The server system includes a first server and a second server. The first server is provided with a first back-end driver interception module, a first GPU driver module, and a first GPU. A first virtual machine and a second virtual machine are running on the second server. A connection channel is established between the first server and the second server. Among them, The first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on a first part of the computing power of the first GPU to generate a first VGPU accessible to the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU and send the first access request to the connection channel; The second virtual machine is provided with a second front-end driver interception module and a second application. The second front-end driver interception module is used to perform device simulation on a second part of the computing power of the first GPU to generate a second VGPU accessible to the second application. The second application is used to access the second VGPU. The second front-end driver interception module is further used to obtain a second access request of the second application for the second VGPU and send the first access request to the connection channel; The first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module and / or the second access request from the second front-end driver interception module from the connection channel. When it is determined that the first access request meets the preset conditions, the first access request is sent to the first GPU driver module, and / or, when it is determined that the second access request meets the preset conditions, the second access request is sent to the first GPU driver module; The first GPU driver module is used to call a first part of the computing power of the first GPU to process the first access request, and / or, call a second part of the computing power of the first GPU to process the second access request.
9. The server system according to claim 8, wherein The second server is further provided with a second back-end driver interception module, a second GPU driver module, and a second GPU. The first front-end driver interception module of the first virtual machine is further used to perform device simulation on a first part of the computing power of the second GPU to generate a third VGPU accessible to the first application. The first application is used to access the third VGPU. The first front-end driver interception module is further used to obtain a third access request of the first application for the third VGPU; The second front-end driver interception module of the second virtual machine is further configured to perform device simulation on the second part of the computing power of the second GPU to generate a fourth VGPU accessible to the second application program, the second application program is used to access the fourth VGPU, and the second front-end driver interception module is further configured to obtain a fourth access request of the second application program for the fourth VGPU; The second back-end driver interception module is configured to obtain the third access request from the first front-end driver interception module and / or the fourth access request from the second front-end driver interception module, and send the third access request to the second GPU driver module when it is determined that the third access request meets the preset conditions, and / or send the fourth access request to the second GPU driver module when it is determined that the fourth access request meets the preset conditions; The second GPU driver module is configured to call the first part of the computing power of the second GPU to process the third access request, and / or call the second part of the computing power of the second GPU to process the fourth access request.
10. The server system according to claim 8 or 9, wherein A first container is provided in the first virtual machine, the first application program runs in the first container, and the first container is used to mount the first VGPU for the first application program to access the first VGPU in the first container; A second container is provided in the second virtual machine, the second application program runs in the second container, and the second container is used to mount the second VGPU for the second application program to access the second VGPU in the second container.
11. The server system according to claim 10, wherein A third container is further provided in the first virtual machine, a third application program runs in the third container, and the third container is used to mount a third VGPU for the third application program to access the third VGPU in the third container, wherein the first front-end driver interception module is configured to perform device simulation on the third part of the computing power of the first GPU to generate the third VGPU accessible to the third application program; and / or, a fourth container is further provided in the second virtual machine, a fourth application program runs in the fourth container, and the fourth container is used to mount a fourth VGPU for the fourth application program to access the fourth VGPU in the fourth container, wherein the second front-end driver interception module is configured to perform device simulation on the fourth part of the computing power of the second GPU to generate the fourth VGPU accessible to the fourth application program.
12. The server system according to claim 8 or 9, wherein The first application program runs directly in the first virtual machine, and the second application program runs directly in the second virtual machine.
13. The server system according to any one of claims 8 to 12, wherein The first GPU driver module is further configured to obtain a first processing result generated by processing the first access request using a first part of the computing and processing capabilities of the first GPU, and / or obtain a second processing result generated by processing the second access request using a second part of the computing and processing capabilities of the first GPU, and send the first processing result and / or the second processing result to the first back-end driver interception module; The first back-end driver interception module is further configured to obtain the first processing result and / or the second processing result from the first GPU driver module, and send the first processing result to the connection channel, and / or send the second processing result to the connection channel; The first front-end driver interception module is further configured to obtain the first processing result from the first back-end driver interception module from the connection channel, and provide the first processing result to the first application; The second front-end driver interception module is further configured to obtain the second processing result from the first back-end driver interception module from the connection channel, and provide the second processing result to the second application.
14. The server system according to any one of claims 8 to 13, characterized in that, The second server further includes: A virtual machine manager, configured to obtain the first access request from the first front-end driver interception module and send the first access request to the connection channel, obtain the second access request from the second front-end driver interception module, and send the second access request to the connection channel.
15. The server system according to any one of claims 8 to 14, characterized in that, The connection channel is implemented through a high-speed interconnection protocol, and the high-speed interconnection protocol includes Compute Express Link (CXL) or Lingqu Universal Bus (UB) protocol.
16. The server system according to any one of claims 8 to 15, wherein The first back-end driver interception module determines that the first access request meets a preset condition, including: The first back-end driver interception module determines that the video memory of the GPU required by the first access request is not greater than a first video memory threshold; and / or the first back-end driver interception module determines that the computing power of the GPU required by the first access request is not greater than the first part of the computing power or a first computing power threshold; The second back-end driver interception module determines that the second access request meets the preset condition, including: The second back-end driver interception module determines that the video memory of the GPU required by the second access request is not greater than a second video memory threshold; and / or the second back-end driver interception module determines that the computing power of the GPU required by the second access request is not greater than the second part of the computing power or a second computing power threshold.
17. A method for creating a virtual machine based on cloud computing technology, characterized in that, The method is applied to a cloud management platform for managing infrastructure, and the infrastructure includes multiple servers. The method includes: Obtaining a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of a first virtual GPU of a first virtual machine to be created; Selecting a target server from the multiple servers that can provide the specification of the first virtual GPU; Creating the first virtual machine on the target server; Among them, the target server is provided with a backend driver interception module, a GPU driver module, and a GPU. The first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on the first part of the computing power of the GPU to generate the first VGPU that can be accessed by the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU. The backend driver interception module is used to obtain the first access request from the first front-end driver interception module and send the first access request to the GPU driver module when it is determined that the first access request meets a preset condition. The GPU driver module is used to call the first part of the computing power of the GPU to process the first access request.
18. A cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure, and the infrastructure includes multiple servers. The cloud management platform includes: An acquisition unit, configured to acquire a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of the first VGPU of the first virtual machine to be created; A selection unit, configured to select a target server that can provide the specification of the first VGPU from the multiple servers; A creation unit, configured to create the first virtual machine on the target server; Among them, the target server is provided with a backend driver interception module, a GPU driver module, and a GPU. The first virtual machine is provided with a first front-end driver interception module and a first application. The first front-end driver interception module is used to perform device simulation on the first part of the computing power of the GPU to generate the first VGPU that can be accessed by the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU. The backend driver interception module is used to obtain the first access request from the first front-end driver interception module and send the first access request to the GPU driver module when it is determined that the first access request meets a preset condition. The GPU driver module is used to call the first part of the computing power of the GPU to process the first access request.
19. A method for creating a virtual machine based on cloud computing technology, characterized in that, The method is applied to a cloud management platform, and the cloud management platform is used to manage the infrastructure. The infrastructure includes multiple servers. The method includes: Acquiring a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of the first VGPU of the first virtual machine to be created; Selecting a second server that can provide the specification of the first VGPU from the multiple servers; Creating the first virtual machine on the second server; Among them, the multiple servers further include a first server, which is provided with a first back-end driver interception module, a first GPU driver module, and a first GPU. A connection channel is established between the first server and the second server. A first front-end driver interception module and a first application are set in the first virtual machine. The first front-end driver interception module is used to perform device simulation on part of the computing power of the first GPU to generate the first VGPU that can be accessed by the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU and send the first access request to the connection channel. The first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions. The first GPU driver module is used to call part of the computing power of the first GPU to process the first access request.
20. A cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure, and the infrastructure includes multiple servers. The cloud management platform includes: An acquisition unit, configured to acquire a virtual machine creation request input by a tenant, where the virtual machine creation request carries the specification of the first VGPU of the first virtual machine to be created; A selection unit, configured to select a second server that can provide the specification of the first VGPU from the multiple servers; A creation unit, configured to create the first virtual machine on the second server; Among them, the multiple servers further include a first server, which is provided with a first back-end driver interception module, a first GPU driver module, and a first GPU. A connection channel is established between the first server and the second server. A first front-end driver interception module and a first application are set in the first virtual machine. The first front-end driver interception module is used to perform device simulation on part of the computing power of the first GPU to generate the first VGPU that can be accessed by the first application. The first application is used to access the first VGPU. The first front-end driver interception module is further used to obtain a first access request of the first application for the first VGPU and send the first access request to the connection channel. The first back-end driver interception module is used to obtain the first access request from the first front-end driver interception module through the connection channel, and send the first access request to the first GPU driver module when it is determined that the first access request meets the preset conditions. The first GPU driver module is used to call part of the computing power of the first GPU to process the first access request.
21. A computing device, characterized in that, It includes a processor and a memory. Program instructions are stored in the memory, and the processor runs the program instructions to enable the computing device to implement the server according to any one of claims 1 to 16.
22. A computer-readable storage medium, characterized in that, including program instructions that, when run on a computing device, cause the computing device to implement the server according to any one of claims 1 to 16.
23. A computer program product comprising instructions, characterized in that, When the instructions are run on a computing device, cause the computing device to implement the server according to any one of claims 1 to 16.
Citation Information
Patent Citations
Image processing system and method
CN113240571A
Server system and virtual machine creating method and device
CN114691286A
Virtualization method and device of graphics processor, electronic equipment and medium
CN115861029A
Method and device for intercepting file access request of container based on virtual equipment
CN116414513A