Server system based on public cloud technology and access method therefor
By sinking the address translation function to external devices, the problem of degradation of address translation performance in cloud computing is solved, efficient memory management and improved external device operation performance, reducing memory costs, and enhancing the competitiveness of the public cloud.
Patent Information
- Application Number
- PCT/CN2025/071798
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-29
- Filing Date
- 2025-01-10
- Publication Date
- 2025-07-17
AI Technical Summary
In a cloud computing environment, as the number of external devices increases, the TLB of existing address translation devices such as IOMMU/SMMU cannot record sufficient GPA-HPA correspondence relationship, resulting in a degradation in address translation performance, affecting the operating performance and memory management efficiency of external devices.
Sink the address translation function to external devices, and work together through the virtual machine manager and address management module. The external devices directly manage memory and translate address, reducing the burden on IOMMU/SMMU in the server, leverage the address translation capabilities of external devices, reduce the number of corresponding relationships that need to be recorded in TLB, and improve address translation performance.
It improves the address translation performance of the server system, improves the operating performance of external devices such as DMA performance, and improves the memory management efficiency, reduces memory costs, and enhances the market competitiveness of the public cloud.
Smart Images

Figure CN2025071798_17072025_PF_FP_ABST
Abstract
Description
Server system based on public cloud technology and access method thereof
[0001] This application claims priority to Chinese patent application No. 202410050423.8 filed on January 12, 2024, with invention name “Method for managing virtual machines”, and priority to Chinese patent application No. 202410547732.6 filed on April 29, 2024, with invention name “Server system based on public cloud technology and access method thereof”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of cloud service technology, and in particular to a server system based on public cloud technology and an access method thereof. Background Art
[0003] As cloud computing becomes increasingly widespread, the number of external devices used on cloud servers is increasing. As the number of external devices increases, the number of GPA-HPA correspondences that the remapping device that performs address translation needs to record increases accordingly. However, since the number of GPA-HPA correspondences that the remapping device can record is fixed, some GPA-HPA correspondences may not be recorded. Therefore, when address translation needs to be performed based on these unrecorded GPA-HPA correspondences, address translation between the GPA and HPA cannot be performed, resulting in a decrease in the address translation performance of the remapping device, which in turn affects the operational performance of the external devices. Summary of the Invention
[0004] This application provides a server system based on public cloud technology and its access method. This application effectively improves the address translation performance of the server system, helps improve the operating performance of external devices (such as DMA performance) and the efficiency of memory management. The technical solutions provided by this application are as follows:
[0005] In a first aspect, the present application provides a server system based on public cloud technology. The server system includes a server and an external device plugged into the server. A connection channel is established between the server and the external device. A virtual machine and a virtual machine manager are provided in the server. The virtual machine manager is used to provide virtual hardware for the virtual machine. The virtual hardware includes virtual external devices and virtual memory. The virtual external device is obtained by device simulation based on the external device. The virtual memory is obtained by device simulation based on the memory configured for the external device. A virtual device driver for the virtual external device is provided in the virtual machine. The virtual device driver is used to send an access request to the external device when the virtual external device accesses the target client physical address of the virtual memory. The access request carries the target client physical address. The external device is used to query the first correspondence between the client physical address and the host physical address based on the access request, obtain the target host physical address corresponding to the target client physical address, obtain the target data recorded by the target host physical address, and send the target data to the virtual device driver.
[0006] As can be seen from the above, the server system is equivalent to sinking the address translation function into the external device, allowing the external device to directly manage the memory allocated for the virtual machine and share the task of address translation of the target client physical address in the access request. In this way, there is no need for the IOMMU / SMMU or other address translation devices in the server to perform the address translation task, and there is no need for the TLB of the IOMMU / SMMU or other address translation devices in the server to record the correspondence between the target client physical address accessed by the virtual external device of the external device, thus reducing the number of correspondences that need to be recorded in the TLB, alleviating the contradiction between the small number of page table entries in the TLB and the excessive number of external devices, reducing the probability of TLB page fault problems during the address translation process, effectively improving the address translation performance of the server system, and helping to improve the operating performance (such as DMA performance) of the external device and the efficiency of memory management.
[0007] In one possible implementation, the server is further provided with an address management module. The external device is further configured to send a message indicating a match failure to the address management module, the message carrying the target client physical address, when the first correspondence does not record the target host physical address corresponding to the target client physical address; the address management module is configured to forward the message to the virtual machine manager; the virtual machine manager is configured to allocate a memory block in memory for the target client physical address based on the message, establish a second correspondence between the host physical address of the memory block and the target client physical address, and send the second correspondence to the address management module; and the address management module is configured to update the first correspondence based on the second correspondence.
[0008] As can be seen from the above, the server system provided by the present application can perform relevant processing when the corresponding relationship based on the target client physical address query is not named, such as executing page fault exception processing to solve the problem of miss. In this way, the memory can be pin-free before the virtual machine is started, and the global memory lock can be released. In the case where the external device is directly connected to the virtual machine in the cloud scenario, the memory pin-free makes it possible to super-multiplex the virtual machine memory without the virtual machine user's perception (no special requirements for the virtual machine system), which can improve the overall utilization of the server memory and ensure that the server's performance does not decrease. This feature is particularly evident in scenarios with high memory pressure. At the same time, since the memory is pin-free, there is no need to configure a large amount of memory in the cloud scenario due to the virtual machine pre-occupying memory, which can reduce the cost of memory use in the cloud scenario, solve the problem of gradually increasing public cloud memory costs, and thus enhance the market competitiveness of the public cloud.
[0009] In one possible implementation, the server is further provided with an address management module. The virtual machine manager is configured to perform device emulation for the virtual machine based on the virtual machine's specifications to obtain virtual memory, configure a client physical address space for the virtual memory, establish a correspondence between the host physical address of the memory block configured for the external device in the memory and the client physical address in the client physical address space, and send the correspondence to the address management module.
[0010] In one possible implementation, the configuration information of the external device indicates that the external device supports translated requests, and the access request indicates that the address carried in the access request is the host physical address. In this way, if the server is compatible with the existing address translation function, it can ensure that the access request capable of the external device can be processed by the external device itself, rather than by the server's existing address translation function, thus bypassing the server's existing address translation function.
[0011] In a possible implementation, the memory configured for the external device includes one or more of the following: a memory inserted into the server or a shared memory of the server.
[0012] In a second aspect, the present application provides an access method for a server system based on public cloud technology. The server system includes a server and an external device inserted into the server, a connection channel is established between the server and the external device, a virtual machine and a virtual machine manager are provided in the server, the virtual machine manager is used to provide virtual hardware for the virtual machine, the virtual hardware includes a virtual external device and a virtual memory, the virtual external device is obtained by device simulation based on the external device, the virtual memory is obtained by device simulation based on the memory configured for the external device, and a virtual device driver for the virtual external device is provided in the virtual machine, the method comprising: when the virtual external device accesses the target client physical address of the virtual memory, the virtual device driver sends an access request to the external device, the access request carries the target client physical address; based on the access request, the external device queries the first correspondence between the client physical address and the host physical address, obtains the target host physical address corresponding to the target client physical address, obtains the target data recorded by the target host physical address, and sends the target data to the virtual device driver.
[0013] In one possible implementation, the server is further provided with an address management module. The method further includes: when the first correspondence does not record the target host physical address corresponding to the target client physical address, the external device sends a message to the address management module indicating a match failure, the message carrying the target client physical address; the address management module forwards the message to the virtual machine manager; based on the message, the virtual machine manager allocates a memory block in memory for the target client physical address, establishes a second correspondence between the host physical address of the memory block and the target client physical address, and sends the second correspondence to the address management module; and the address management module updates the first correspondence based on the second correspondence.
[0014] In one possible implementation, the server is further provided with an address management module. The method further includes: the virtual machine manager, based on the virtual machine's specifications, performing device simulation for the virtual machine to obtain virtual memory, configuring a client physical address space for the virtual memory, establishing a correspondence between the host physical address of a memory block configured for the external device in the memory and the client physical address in the client physical address space, and sending the correspondence to the address management module.
[0015] In a possible implementation, the configuration information of the external device indicates that the external device supports translated requests, and the access request indicates that the address carried in the access request is a host physical address.
[0016] In a possible implementation, the memory configured for the external device includes one or more of the following: a memory inserted into the server or a shared memory of the server.
[0017] In a third aspect, the present application provides a computing device comprising a memory and a processor, wherein the memory stores program instructions, and the processor runs the program instructions to execute the method provided in the second aspect of the present application and any possible implementation thereof.
[0018] In a fourth aspect, the present application provides a computing device cluster, comprising multiple computing devices, the multiple computing devices including multiple processors and multiple memories, the multiple memories storing program instructions, and the multiple processors running the program instructions, so that the computing device cluster executes the method provided in the second aspect of the present application and any possible implementation thereof.
[0019] In a fifth aspect, the present application provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium. The computer-readable storage medium includes program instructions. When the program instructions are executed on a computing device, the computing device executes the method provided in the second aspect of the present application and any possible implementation thereof.
[0020] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when run on a computer, enables the computer to execute the method provided in the second aspect of the present application and any possible implementation thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] FIG1 is a schematic diagram of a GPA and an HPA provided in an embodiment of the present application;
[0022] FIG2 is a schematic diagram of a structure of an implementation scenario involved in the present application provided by an embodiment of the present application;
[0023] FIG3 is a schematic diagram of a cloud resource deployment provided in an embodiment of the present application;
[0024] FIG4 is a schematic diagram of the structure of a server system provided in an embodiment of the present application;
[0025] FIG5 is a logical diagram of a server system provided in an embodiment of the present application;
[0026] FIG6 is a logical diagram of another server system provided in an embodiment of the present application;
[0027] FIG7 is a logical diagram of another server system provided in an embodiment of the present application;
[0028] FIG8 is a schematic diagram of a query correspondence relationship provided in an embodiment of the present application;
[0029] FIG9 is a schematic structural diagram of another server system provided in an embodiment of the present application;
[0030] FIG10 is a schematic diagram of a context entry provided in an embodiment of the present application;
[0031] FIG11 is a schematic diagram of a TLP header format provided in an embodiment of the present application;
[0032] FIG12 is a schematic diagram of interaction between a user state and a kernel state provided in an embodiment of the present application;
[0033] FIG13 is a schematic diagram of an ioctl command format provided in an embodiment of the present application;
[0034] FIG14 is a schematic diagram of an interaction between a kernel state and an external device provided in an embodiment of the present application;
[0035] FIG15 is a schematic diagram of a server system including multiple external devices provided in an embodiment of the present application;
[0036] FIG16 is a flowchart of a method for accessing a server system based on public cloud technology according to an embodiment of the present application;
[0037] FIG17 is a flowchart of another method for accessing a server system based on public cloud technology provided in an embodiment of the present application;
[0038] FIG18 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0039] FIG19 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0040] FIG20 is a schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0042] To facilitate understanding, the technology and background involved in the embodiments of this application are first introduced below.
[0043] A cloud data center is an internet-based network that provides operational maintenance facilities and related service systems for centralized data collection, storage, processing, and transmission. Conceptually, it can be understood as a public, commercial Internet "computer room." It also provides a professional information technology (IT) service and is a critical infrastructure for the IT industry. A cloud data center is not only a service concept but also a network concept. It forms part of the network infrastructure, much like backbone networks and access networks, providing high-end data delivery and high-speed access services.
[0044] Cloud computing: Cloud computing is a type of distributed computing that refers to a network that centrally manages and schedules large amounts of computing and storage resources to provide on-demand services to users. These computing and storage resources are provided by computing equipment located in data centers. Cloud computing can also provide users with a variety of service types, such as infrastructure as a service (IaaS), platform as a service (PaaS), and software as a service (SaaS). Infrastructure as a service provides virtual machines or other resources as a service to tenants. Platform as a service provides a development platform as a service to tenants. Software as a service provides applications (Apps) as a service to customers.
[0045] Physical machine (PM): A physical resource used to host virtualization technology. A host is also called a physical machine. Typically, a physical server is a host used to deploy virtual instances. A physical machine has multiple physical devices. For example, a physical server has physical devices such as a processor and memory. Multiple virtual instances can be deployed on a host. Multiple virtual instances deployed on the same host share the host's physical resources. Depending on the usage scenario, multiple virtual instances deployed on a host can optionally belong to the same tenant or different tenants.
[0046] Virtualization is a resource management technology. It abstracts and transforms a host's physical resources, such as computing, networking, and storage, to create a more tangible representation. This breaks down the barriers between the host's physical structure and allows tenants to utilize these resources in a more efficient manner than their original configuration. The resources created through virtualization are called virtualized resources, and they are not restricted by the existing physical resource configuration, location, or physical configuration.
[0047] Virtualized resources are usually provided to tenants in the form of virtual instances. The hardware resources of a physical server can be used by one or more tenants at the granularity of a virtual instance. Different virtual instances are isolated from each other, allowing tenants to use physical resources conveniently and flexibly under the premise of secure isolation, and can greatly improve the utilization of physical resources. In this application, a virtual instance can be a virtual machine, a container, or an independent process (such as a function), etc. A virtual instance can also be called a cloud server (elastic compute service, ECS) or an elastic instance (different cloud service providers have different names).
[0048] A virtual machine (VM) is a complete computer system with complete hardware system functions, simulated by software and running in a completely isolated environment. Any work that can be done on a server can also be performed in a VM. When creating a VM on a server, part of the physical machine's hard disk and memory capacity is used as the VM's hard disk and memory capacity. Each VM has its own independent hard disk and operating system, and VM tenants can operate the VM just like they would a server. The operating environments (such as VM applications, operating systems, and virtual hardware) in different VMs are completely isolated, and communication between VMs requires network packets to be forwarded by the virtualization manager.
[0049] Containers provide a lightweight virtual runtime environment. They can be created by packaging a tenant's application code, libraries, and dependencies into an image. When an image is executed, it runs within the virtual runtime environment. A container is a runtime instance of an image, similar to a lightweight sandbox, which can be started, started, stopped, and deleted. This image does not share the host's memory, processors (such as the central processing unit (CPU)), and disk resources with other images. This provides container isolation between the image and the host, and between the image and other images, ensuring that processes within the container cannot monitor any processes or resources outside the container. The container infrastructure can be server hardware or virtual machines in the cloud (i.e., containers can also be deployed within virtual machines). The operating system uses the Linux kernel and supports namespaces and control groups. Namespaces isolate processes, while control groups allocate resources to processes, specifically virtual processors and memory allocated to each process. A container engine, similar to a virtual machine manager, runs within the operating system and manages containers. Compared with virtual machines that come with their own operating systems, containers do not have operating systems. Containers run as processes in the host's operating system, so containers start faster than virtual machines. They are particularly suitable for lightweight applications, and a host can run thousands of containers (processes) simultaneously.
[0050] Direct memory access (DMA), also known as direct memory operation or group data transfer, refers to a data exchange mode in which an external device accesses data directly from a computer's memory, without going through the computer's central processing unit (CPU). When transferring data in DMA mode, the computer's CPU issues instructions to the DMA controller, instructing it to control data transfer. After completing the data transfer, the DMA controller sends a message to the CPU confirming the transfer is complete. This DMA data transfer process eliminates the need for the computer's CPU to perform transfer operations, eliminating the need for CPU operations such as instruction fetching, data fetching, and data sending. This reduces the CPU's resource usage and conserves system resources.
[0051] DMA can include remote direct memory access (RDMA) and local DMA. RDMA transfers data directly from one computer's memory to another over a network, without requiring the intervention of either computer's operating system. Local DMA transfers data without requiring a network connection.
[0052] Peripheral Component Interconnect Express (PCIe) bus: A high-speed serial computer expansion bus.
[0053] Memory: It can also be called internal memory or main memory. Its function is to temporarily store the calculation data in the CPU and the data exchanged with external memory such as hard disk.
[0054] External device, referred to as "peripheral". A general term for input and output devices (including external memory) in a computer system. It plays the role of transmitting, forwarding and storing data and information. It is an important component of a computer system. Peripheral devices refer to any device other than the host computer. Peripheral devices are auxiliary or auxiliary devices connected to the computer. Peripheral devices can expand the computer system. Examples of external devices are network cards, external memory (such as hard disks, disks and graphics cards), graphics processing units (GPUs), intelligent processing units (IPUs) and data processing units (DPUs). Among them, a network card can also be called a network interface controller (NIC), a network adapter, or a local area network receiver. It is a type of computer hardware designed to allow a host or computing device to communicate on a network.
[0055] The purpose of memory virtualization technology is to provide virtual machines with a continuous physical memory space starting at address 0, effectively isolating and scheduling memory resources between virtual machines. Memory virtualization technology primarily involves the translation of guest virtual address (GVA) -> guest physical address (GPA) -> host virtual address (HVA) -> host physical address (HPA).
[0056] In virtualization technology, multiple virtual machines often run on a physical host, and each virtual machine believes that it has exclusive access to the physical host's memory space. Therefore, the virtual machine uses GPA to represent the memory space owned by the virtual machine, where the memory space is considered continuous by the virtual machine (that is, it can be understood that the virtual machine believes that it owns a complete physical memory bar).
[0057] GVA is the address formed by the virtual machine's operating system mapping the GPA. The virtual machine's operating system provides the GVA to the process or application software installed on the virtual machine's operating system. The virtual machine's operating system records the mapping relationship between GVA and GPA. The conversion from GVA to GPA is implemented by the virtual machine's operating system's page table.
[0058] HPA is the actual physical memory address, and HVA is the address formed by the host operating system mapping HPA. The host operating system provides HVA to processes on the operating system (such as virtual machines) for use. The host operating system records the mapping relationship between HVA and HPA. The conversion from HVA to HPA is implemented by the page table of the host operating system.
[0059] In Figure 1, virtual machine 1 and virtual machine 2 are set in the same server (hereinafter referred to as the host machine). The virtual machine manager of the host machine (also called Hypervisor) sets the GPA address range of virtual machine 1 to 0-5GB, which corresponds to the HPA address range of 1.5GB-4.5GB and 6.5GB-8.5GB on the physical memory. In addition, the virtual machine manager of the host machine sets the GPA address range of virtual machine 2 to 0-4GB, which corresponds to the HPA address range of 9GB-11GB and 13GB-15GB on the physical memory. Therefore, virtual machine 1 exclusively uses the GPA address range of 0-5GB, and virtual machine 2 exclusively uses the GPA address range of 0-4GB. The GPA address range of 0-5GB and the GPA address range of 0-4GB can both correspond to different HPA address ranges on the physical memory, thereby achieving isolation of virtual machine memory.
[0060] The GPA address range is related to the VM specifications mentioned above. For example, a tenant can set VM 1 with a memory size of 5GB in the cloud management platform. At this time, the VM manager is notified by the cloud management platform to create VM 1 with a GPA of 0-5GB. A tenant can also set VM 2 with a memory size of 4GB in the cloud management platform. At this time, the VM manager is notified by the cloud management platform to create VM 2 with a GPA of 0-4GB.
[0061] Currently, in cloud scenarios, when external devices are directly connected to virtual machines, the hypervisor must allocate memory to the virtual machine all at once and lock (pin) the memory blocks allocated to the virtual machine. This requires the virtual machine to initialize the reserved memory before startup, ensuring that the memory blocks allocated to the virtual machine cannot be used by other memory users during the virtual machine's lifetime. At the same time, the hypervisor allocates GPA space for the virtual machine and establishes a correspondence between multiple GPAs in this GPA space and the HPA of the memory blocks allocated to the virtual machine. The correspondence between GPAs and HPA is maintained by the host operating system.
[0062] Among them, pass-through technology is a virtualization technology supported by the server's external devices (such as PCIe devices). Pass-through means skipping the hypervisor and providing the external device directly to the virtual machine for use. Pass-through technology supports the external device to set up multiple virtual functions (VFs). The external device sets storage resources or network resources in one or more VFs and provides one or more VFs to a tenant's virtual machine. The virtual machine can directly use the resources provided by one or more VFs. For example, the external device binds a 40G logical disk divided in a storage resource to a VF and provides the VF to virtual machine 1. After mounting the VF, virtual machine 1 can access the logical disk as if it were a local disk. Alternatively, the external device binds a network card interface eth0 in the network resource to a VF and provides the VF to virtual machine 1. After mounting the VF, virtual machine 1 can use eth0 as if it were a local network card.
[0063] The operating system of a virtual machine running on a virtual machine typically doesn't know the HPA of the memory block allocated to it by the hypervisor, but it does know the GPA space allocated to it. When a virtual machine accesses its allocated memory block through an external device directly connected to it, it sends a DMA request to the host operating system. The DMA request carries the GPA to be accessed. After receiving the DMA request, the host operating system queries the GPA-HPA mapping based on the GPA carried in the DMA request, obtains the HPA corresponding to the GPA requested by the virtual machine, reads the data in the memory block indicated by the HPA, and then provides this data to the virtual machine. The process by which the host operating system obtains the corresponding HPA based on the GPA is called address translation or address remapping. This process translates the GPA of the DMA request into a host physical address (HPA) that can be used by the external device, and is a memory management process. The device used to perform the address remapping process is called remapping hardware. Currently, remapping hardware typically includes the input / output memory management unit (IOMMU) and the system memory management unit (SMMU). Both the IOMMU and SMMU are set up in the host. The IOMMU and SMMU record the correspondence between the GPA and the HPA through the translation lookaside buffer (TLB). The number of page table entries (PTE) in the TLB is usually fixed, that is, the number of GPA and HPA correspondences that the IOMMU and SMMU can record is fixed. Among them, the PTE consists of a valid bit and an n-bit address field. The address field records the HPA corresponding to the GPA, that is, the starting position of the memory page corresponding to the GPA.
[0064] As the number of external devices increases, the number of GPA-HPA correspondences that the IOMMU and SMMU need to record increases accordingly. However, since the number of GPA-HPA correspondences that the IOMMU and SMMU can record is fixed, some GPA-HPA correspondences may not be recorded. When address translation needs to be performed based on the unrecorded GPA-HPA correspondence, address translation between the GPA and HPA cannot be achieved, resulting in a decrease in the address translation performance of the IOMMU and SMMU, which in turn affects the operating performance of external devices. For example, in the current cloud scenario, the host's external devices are constantly increasing, and the resources and DMA operations of external devices are constantly increasing, but the number of GPA-HPA correspondences that the IOMMU / SMMU can record is relatively small, resulting in a decrease in the address translation performance of the IOMMU / SMMU. Taking the IOMMU as an example, the current x86 IOMMU TLB only has 64 page table entries. In cloud scenarios with large virtual functions (VFs), the number of page table entries required to be recorded by the IOMMU TLB often exceeds 64. This results in some GPA-HPA mappings not being recorded. When the HPA corresponding to these GPAs needs to be queried, address translation performance is affected. For example, when the number of VFs exceeds 64, severe TLB page faults will occur. At 128 VFs, address translation performance drops by approximately 40%.
[0065] Currently, the address translation function does not support page fault handling, which requires external devices connected directly to the virtual machine to pre-occupy memory. Because the memory reserved by external devices is locked, it cannot be used by other memory users, resulting in low memory utilization. For example, in current cloud scenarios, when external devices are connected directly to virtual machines, the hypervisor must successfully allocate memory all at once and pin the allocated memory to the host. The allocated memory cannot be reused, resulting in low overall memory utilization. Furthermore, the virtual machine needs to initialize the pre-occupied memory before startup, which can also slow down the virtual machine startup.
[0066] Although a virtual IOMMU technology (coIOMMU) has been proposed, it requires modifying the virtual machine's operating system, making it difficult to implement in public clouds.
[0067] In view of this, the present application provides a server system based on public cloud technology and an access method thereof. The server system includes a server and an external device plugged into the server. A connection channel is established between the server and the external device. A virtual machine and a virtual machine manager are provided in the server. The virtual machine manager is used to provide virtual hardware for the virtual machine. The virtual hardware includes virtual external devices and virtual memory. The virtual external device is obtained by device simulation based on the external device. The virtual memory is obtained by device simulation based on the memory configured for the external device. A virtual device driver for the virtual external device is provided in the virtual machine. The virtual device driver is used to send an access request to the external device when the virtual external device accesses the target client physical address of the virtual memory. The access request carries the target client physical address. The external device is used to query the first correspondence between the client physical address and the host physical address based on the access request, obtain the target host physical address corresponding to the target client physical address, obtain the target data recorded by the target host physical address, and send the target data to the virtual device driver.
[0068] As can be seen from the above, the server system is equivalent to sinking the address translation function into the external device, allowing the external device to directly manage the memory allocated for the virtual machine and share the task of address translation of the target client physical address in the access request. In this way, there is no need for the IOMMU / SMMU or other address translation devices in the server to perform the address translation task, and there is no need for the TLB of the IOMMU / SMMU or other address translation devices in the server to record the correspondence between the target client physical address accessed by the virtual external device of the external device, thus reducing the number of correspondences that need to be recorded in the TLB, alleviating the contradiction between the small number of page table entries in the TLB and the excessive number of external devices, reducing the probability of TLB page fault problems during the address translation process, effectively improving the address translation performance of the server system, and helping to improve the operating performance (such as DMA performance) of the external device and the efficiency of memory management.
[0069] It should be noted that although this application uses a virtual machine using an external device as an example, this does not exclude the use of external devices by other types of virtual instances, and the embodiments of this application do not specifically limit this. Such other types of virtual instances may be, for example, containers. When other types of virtual instances use external devices, the implementation method for accessing memory through external devices should refer to the implementation method for virtual machines accessing memory through external devices in the embodiments of this application, and will not be further described herein.
[0070] This article provides a detailed introduction to the technical solution of this application from multiple perspectives, including implementation scenarios, method flow, hardware devices, and software devices.
[0071] The following first illustrates an implementation scenario of the embodiment of the present application with examples.
[0072] Figure 2 is a structural diagram of an implementation scenario involved in the present application provided in an embodiment of the present application. As shown in Figure 2, the implementation scenario includes: data center 1 and client 2. A communication connection can be established between data center 1 and client 2 through a network. Optionally, the network can be the Internet or other networks, which is not limited in the embodiment of the present application. Tenants can interact with data center 1 through client 2. For example, a tenant can send information such as a virtual instance creation request to data center 1 through client 2. Data center 1 is used to respond based on the information sent by client 2.
[0073] Data center 1 is deployed with a large amount of infrastructure owned by the cloud service provider, such as computing resources, storage resources, and network resources. For example, computing resources can be computing devices that can provide computing power, such as servers and external devices with computing power, such as GPUs, IPUs, and DPUs. Storage resources can be devices that can provide storage capabilities, such as external devices such as hard disks, disks, and graphics cards. Network resources can be devices that can provide network transmission capabilities, such as external devices such as network cards. As shown in Figure 2, data center 1 includes a cloud management platform and infrastructure (not shown in Figure 2). The cloud management platform and infrastructure are connected through the data center's internal network. The cloud management platform is used to manage the infrastructure. The infrastructure is used to provide public cloud services. The basic settings include multiple servers. Cloud services can be optionally deployed in the servers. Tenants can send cloud service requests to the server through the client 2 they use. The server can process the cloud service request and provide cloud services based on the processed cloud service request.
[0074] The cloud management platform can be logically divided into the following functional areas: the tenant console, compute management service, network management service, storage management service, authentication service, and image management service. The tenant console provides an interface or application program interface (API) for interacting with tenants. The compute management service manages servers running virtual instances and bare metal servers. The network management service manages network services (such as gateways and firewalls). The storage management service manages storage services (such as data bucket services). The authentication service manages tenant accounts and passwords. The image management service manages images for virtual instances.
[0075] In the implementation scenario shown in Figure 2, multiple servers are deployed in a data center. The server consists of a hardware layer and a software layer. The hardware layer is the standard configuration of the server. The hardware layer deploys hardware devices such as processors, memory, network cards, disks, and buses. The software layer includes the operating system installed and running on the server. The operating system of the virtual machine is called the host operating system. The host operating system runs a virtual machine manager (also called a hypervisor). The role of the virtual machine manager is to implement computing virtualization, network virtualization, and storage virtualization for the virtual machine, and is responsible for managing the virtual machine.
[0076] The cloud management platform client runs within the virtual machine manager. The cloud management platform client receives control plane commands from the cloud management platform, creates virtual instances on servers based on these commands, and manages the virtual instances throughout their lifecycle. For example, the cloud management platform client monitors the hardware resource usage of the server in real time and reports this information to the cloud management platform. When the cloud management platform confirms the creation of a virtual instance on a server, it sends a virtual instance creation command to the cloud management platform client on that server. Upon receiving this command, the cloud management platform client creates the virtual instance on that server. This allows tenants to create, manage, log in to, and operate virtual instances in the data center through the cloud management platform.
[0077] Servers can be used to run virtual machines of varying specifications. Virtual machine specifications are categorized as general-purpose computing, memory-optimized, and ultra-large memory, with each type further defined. After a tenant selects a virtual machine specification, the cloud management platform selects a server in the data center that supports that specification, determines if the server has sufficient available hardware resources, and then creates a virtual machine with that specification on that server. Configuring servers through the cloud management platform allows for analysis and planning of server hardware resources. Based on the server's hardware performance, computing products corresponding to the physical hardware can be planned, such as virtual machines of varying specifications, to meet the differentiated needs of different tenants. Furthermore, the performance differences between virtual machines of varying specifications can enable differentiated pricing strategies. For example, virtual instances with high performance specifications can be sold at a higher price, while those with standard performance specifications can be sold at a lower price, allowing tenants to purchase virtual instances on demand.
[0078] In one implementation, as shown in Figure 3, the location of basic resources in a data center can be described using cloud resource deployment regions and availability zones (AZs). Tenants can optionally deploy cloud services based on resources in specific regions and AZs. Regions are divided based on geographic location and network latency. Within a region, the same resource pool is used, which can be understood as sharing public services such as elastic computing, block storage, object storage, virtual private cloud (VPC) networks, Elastic Internet Protocol (EIP) addresses, and images. Regions are categorized as general regions and dedicated regions. General regions provide general cloud services to public tenants. Dedicated regions are dedicated regions that carry the same type of business or provide services to specific tenants. A region typically includes multiple AZs. AZs within a region are connected by high-speed fiber optic cables to meet tenants' needs for building high-availability systems across AZs. An AZ is a collection of one or more data centers, as shown in Figure 3. Computing, networking, and storage resources within an AZ are logically divided into multiple clusters.
[0079] Tenants can send instructions to the cloud management platform through the client 2 they use to create, manage, log in and operate virtual instances in the server, and use the cloud services provided by the virtual instances. For example, the cloud management platform can provide an access interface. The access interface can be optionally provided in the form of an interface or an API. Tenants can operate the client to remotely access the access interface to register a cloud account and password on the cloud management platform, and use the cloud account and password to log in to the cloud management platform. The cloud management platform can also authenticate the cloud account and password. After successful authentication, the tenant can further select and pay to purchase a virtual instance of specific specifications (processor, memory, disk) on the cloud management platform. After the tenant successfully pays for the virtual instance, the cloud management platform provides the tenant with the remote login account and password of the purchased virtual instance. The tenant can use the remote login account and password to remotely log in to the virtual instance on the client, install and run the tenant's application in the virtual instance, and implement the tenant's business through the application.
[0080] Client 2 may be a computer, a personal computer, a laptop computer, a mobile phone, a smart phone, a tablet computer, a cloud host, a portable mobile terminal, a multimedia player, an e-book reader, a wearable device, a smart home appliance, an artificial intelligence device, a smart wearable device, a smart vehicle-mounted device or an Internet of Things device, etc.
[0081] In one implementation, the method for accessing a server system based on public cloud technology provided in an embodiment of the present application can be implemented by running an executable program on a computing device in the data center 1. Optionally, the method for accessing a server system based on public cloud technology provided in an embodiment of the present application can be optionally applied to an access system for a server system based on public cloud technology. The access system for a server system based on public cloud technology is deployed in a server managed by a cloud management platform. The access system for a server system based on public cloud technology can implement the method for accessing a server system based on public cloud technology provided in an embodiment of the present application by running the executable program for the method for accessing a server system based on public cloud technology provided in an embodiment of the present application. Furthermore, the executable program for implementing the method for accessing a server system based on public cloud technology can be optionally presented in the form of an application installation package. After the server installs the application installation package, it can implement the method for accessing a server system based on public cloud technology provided in an embodiment of the present application by running the executable program therein.
[0082] It should be understood that the above content is an illustrative description of the implementation scenario of the access method for a server system based on public cloud technology provided in the embodiment of the present application, and does not constitute a limitation on the implementation scenario of the access method for a server system based on public cloud technology. A person of ordinary skill in the art will know that as business needs change, its implementation scenario can be adjusted according to application requirements, and the embodiment of the present application does not list them one by one.
[0083] The following describes an implementation of a server system based on public cloud technology provided in an embodiment of the present application. FIG4 is a schematic diagram of a server system provided in an embodiment of the present application. As shown in FIG4 , the server system 1 includes a server 11 and an external device 12 inserted into the server. A connection channel is established between the server and the external device. For example, a device slot is provided on the motherboard of the server, and the external device has an interface that matches the slot. The external device is inserted into the slot of the server through the interface, and the external device and the server can be connected by a bus or the like. Optionally, the external device in the present application may be a GPU, IPU, DPU, hard disk, disk, graphics card, network card, etc., and the external device in the present application may also be other types of devices other than those exemplified here, and the embodiments of the present application do not give examples one by one for them.
[0084] The server has hardware 111. A virtual machine manager 112 and a virtual machine 113 are provided within the server. The virtual machine manager provides virtual hardware 1121 for the virtual machine based on the server's hardware and its external devices. The virtual hardware includes virtual external devices 1121a and virtual memory 1121b. The virtual external devices are simulated by the virtual machine manager based on the external devices. Virtual memory is simulated by the virtual machine manager based on the memory configured for the external devices. It should be understood that the virtual machine manager can provide other types of virtual hardware for the virtual machine, such as virtual processors, to meet the performance requirements of the virtual machine. Examples of these are not provided here. The virtual machine also includes a virtual device driver for the virtual external device. This virtual device driver implements the access process for the virtual external device. When the virtual external device accesses a target client physical address in virtual memory, the virtual device driver sends an access request to the external device. The access request carries the target client physical address. For example, if an application is provided within the virtual machine and requires the capabilities of an external device during its execution, and the external device is directly passed to the virtual machine for use, the virtual external device will access the virtual memory to provide the application with the capabilities of the external device. The operation of the virtual external device accessing the virtual memory triggers an access request. The virtual device driver can obtain the access request and send the access request to the external device so that the external device can access the memory based on the access request, thereby achieving the purpose of the virtual external device accessing the virtual memory.
[0085] Among them, the memory in this application can be selected as the memory configured on the server (i.e., local memory), or the memory of other computing devices that have established a connection channel with the server, etc. The memory of other computing devices that have established a connection channel with the server, for example, is deployed on the other computing device, and the memory shared by the other computing device to the server is shared memory. The server can access the memory of the other computing device through remote access or a high-speed communication bus. For example, the server accesses the memory of the other computing device through remote direct memory access (RDMA) technology. The high-speed communication bus can be selected as a compute express link (CXL) or a Lingqu bus (also known as a UB bus).
[0086] After receiving the access request sent by the virtual device driver, the external device can query the first correspondence between the client physical address and the host physical address based on the access request to obtain the target host physical address corresponding to the target client physical address. The external device then obtains the target data recorded in the target host physical address from the memory and sends the target data to the virtual device driver. Among them, the virtual device driver can optionally send the access request to the external device through a work queue (WQ). Accordingly, the external device can obtain the access request from the work queue in the manner of obtaining a work queue element (WQE) from the work queue, and then obtain the input / output virtual address (IOVA) carried by the access request in the access request, that is, obtain the target client physical address.
[0087] The address translation function of an external device can be implemented in software or hardware. When implemented in software, the external device must be configured with an executable program that implements the function. When implemented in hardware, the external device must also be configured with the hardware. This hardware is called a device memory management unit (DMMU).
[0088] Figures 5, 6, and 7 are logical diagrams of several server systems presented from different perspectives in the embodiments of the present application. As shown in Figure 5, the server system can be logically divided into a server and external devices, and the server is further divided into user mode and kernel mode. A virtual machine is provided in the user mode, and the virtual machine runs an operating system for the virtual machine. The operating system manages virtual external devices (not shown in the figure), virtual device drivers, and virtual memory. The kernel mode manages memory and the DMMU device driver (DMMU driver), and a mapping table (map table) recording the first correspondence is stored in the kernel mode. The memory in Figure 5 is the local memory of the server. The hypervisor includes a user mode-oriented portion and a kernel mode-oriented portion. The external device is provided with a DMMU and a DMA enablement component (combo DMA). When a virtual external device triggers an access request, the virtual device driver will obtain the access request and send it to the DMMU. After receiving the access request, the DMMU queries the mapping table based on the GPA carried in the access request. After finding the HPA corresponding to the GPA, it provides the HPA to the DMA enablement component. Based on the HPA, the DMA-enabled component reads the data recorded in the memory block indicated by the HPA and provides this data to the virtual machine. Figure 6 differs from Figure 5 in that memory can be implemented in multiple ways. Depending on its function, memory can include multiple components. For example, as shown in Figure 6, memory includes real allocated memory, page on demand memory (pod mem), copy on write memory (cow mem), and swapped memory (swapped mem). When the memory is double data rate synchronous dynamic random access memory (DDR SDRAM), real allocated memory is also called real DDR. Figure 7 differs from Figure 5 in that the memory in Figure 7 is shared memory deployed on another server (such as the remote node in Figure 7). After obtaining the HPA, if the memory block indicated by the HPA is a shared memory block, the DMA-enabled component uses RDMA or a high-speed communication bus to access the shared memory to read the data recorded in the memory block indicated by the HPA. After obtaining the HPA, if the memory block indicated by the HPA is a shared memory block, the DMA-enabled component accesses the shared memory via the high-speed communication bus to read the data stored in the memory block indicated by the HPA. Furthermore, in Figure 7, the server and external device establish a connection channel via CXL / UB, and the external device and remote node achieve mutual access to shared memory via RDMA / UB.In addition, Figure 7 shows the process of providing GPA to DMMU after the application set in the virtual machine initiates the access. The DMMU is provided with a page table cache (cache) and an address mapping lookup component. The address mapping lookup component is used to query the corresponding relationship based on GPA. The page table cache is used to store the corresponding relationships that are more frequently used in the EPT. When performing address translation, the address mapping lookup component can first search for the corresponding relationship in the page table cache. When it cannot be found in the page table cache, it will then search in the address management module or EPT. When the target HPA corresponding to the target GPA can be queried in the page table cache, the address translation can be further accelerated. Figure 7 also shows a large memory management, which is used to manage the memory configured for external devices.
[0089] In one possible implementation, when a virtual machine manager provides a virtual external device for a virtual machine, it may optionally provide resources to the virtual machine in the form of a VF. Different VFs are distinguished by their VF identifiers (i.e., VFIDs). Furthermore, the memory allocated to the external device is also distinguished by VFs, so the first correspondence relationship may be optionally managed at the VF granularity. For example, when the first correspondence relationship is stored in the form of a page table, each page table is used to record the correspondence between the client physical address and the host physical address belonging to a VF. The correspondence between the client physical address and the host physical address belonging to a VF is used to record the correspondence between the client physical address in the client physical address space allocated to the VF and the host physical address in the physical memory space allocated to the VF. When the virtual machine manager provides resources to the virtual machine in the form of a VF, if the virtual machine accesses the resources provided in the form of a VF, the access request received by the external device will carry the VFID. After receiving the access request, the patching device can obtain the VFID based on the access request. Then, as shown in FIG8 , the external device can query the memory request context (MR context) of the VF indicated by the VFID from the external device's configuration space based on the VFID. From the MR context, the external device can then find the base address of the page table that records the correspondence between the client physical address and the host physical address belonging to the VF, i.e., mr_ctx->mtt_base_addr in FIG8 . The external device then queries the page table indicated by the base address based on the target client physical address carried in the access request to obtain the target host physical address corresponding to the target client physical address. The process of the external device querying the page table based on the target client physical address is a level-by-level search process.
[0090] As shown in FIG8 , the IOVA carried in the access request includes multiple address segments, each of which is represented by an address field represented by a binary number of a specified number of bits. The target client physical address shown in FIG8 includes four address segments: address segment 0 represented by binary bits 0 to 11, address segment 1 represented by binary bits 12 to 20, address segment 2 represented by binary bits 21 to 29, and address segment 3 represented by binary bits 30 to 38. After obtaining the base address of the page table, the external device can query the second segment page table (such as seg 2table in FIG8 ) based on this base address. Then, based on the address indicated by the third address segment, the base address of the first segment page table (such as seg 1table in FIG8 ) is found in the second segment page table. Then, based on the address indicated by the second address segment, the base address of the 0th segment page table (such as seg 0table in FIG8 ) is obtained by querying the first segment page table indicated by the base address of the first segment page table. Then, in the page table of segment 0 indicated by the base address of the page table of segment 0, a query is performed based on the address indicated by the first address segment to obtain the offset of the page table entry in the page table of segment 0 that records the correspondence between the target client physical address and the target host physical address. Then, based on the address indicated by the 0th address segment, the target host physical address (i.e., PA in FIG8 ) is obtained by querying the page table entry. It should be noted that the information such as the number of bits in the address field in FIG8 is only an example and can be changed according to application requirements.
[0091] Among them, the process of the external device obtaining the target host physical address corresponding to the target client physical address based on the access request can be optionally implemented by a microcode set in the external device, or by a hardware engine set in the external device, and the embodiments of the present application do not make specific limitations on it.
[0092] In one possible implementation, as shown in FIG9 , the server is further provided with an address management module 1111. The correspondence between the host physical address of a memory block in memory and the client physical address in the client physical address space can be stored in the address management module. That is, the mapping tables shown in FIG5 , FIG6 , and FIG7 are stored in the address management module. Accordingly, the address management module can optionally be deployed in the kernel state of the server. FIG9 is a schematic diagram of the address management module being implemented in the hardware 111.
[0093] At this time, the virtual machine manager is used to provide the address management module with the correspondence between the host physical address and the client physical address. Among them, the virtual machine manager can allocate memory blocks for the virtual machine from the memory based on the specifications of the virtual machine, and then perform device simulation for the virtual machine based on the memory blocks allocated to the virtual machine to obtain virtual memory, and configure the client physical address space of the virtual memory, and then establish a correspondence between the host physical address of the memory block in the memory and the client physical address in the client physical address space, and provide the correspondence to the address management module. For example, as shown in Figures 5, 6 and 7, the hypervisor can provide MR information to the address management module, which carries the correspondence between the host physical address of the memory block in the memory and the client physical address in the client physical address space. In addition, the correspondence also carries the identifier (VMID) of the virtual machine to which the correspondence belongs, to identify the virtual machine to which the correspondence belongs, so that the corresponding virtual machine can be found based on the identifier when a page is missing. Among them, before the virtual machine manager simulates and obtains virtual memory for the virtual machine, it is necessary to first determine the memory blocks allocated to the virtual machine based on the specifications of the virtual machine, and then perform device simulation based on the memory blocks allocated to the virtual machine to obtain virtual memory. Therefore, the virtual machine manager can obtain the host physical address of the memory block corresponding to the client physical address in the client physical address space and establish a correspondence between the host physical address and the client physical address based on the host physical address. Optionally, a DMMU driver is provided in the server kernel. In this case, the virtual machine manager can send MR information to the address management module through the DMMU driver.
[0094] In the present application, the server is compatible with the original address translation function. For example, as shown in Figures 5 and 7, the server is also provided with an MMU and an IOMMU. The IOMMU is used to receive access requests sent by the virtual machine, and based on the GPA carried by the access request, obtain the HPA corresponding to the GPA, read data from the HPA, and provide the read data to the virtual machine. The correspondence used by the IOMMU is also provided by the Hypervisor. When the server is compatible with the original address translation function, some changes need to be made in the server so that access requests using the capabilities of the external device can be processed by the external device itself, not by the original address translation function of the server, that is, bypassing the original address translation function of the server. In one possible implementation method, this function can be implemented by setting the configuration space of the external device and / or setting the access request. For example, the configuration information of the external device can be set to indicate that the access request is sent directly to the external device, and the access request indicates that the address carried by the access request is the host physical address obtained by translation. In this way, when the virtual external device accesses the target client physical address of the virtual memory, the virtual device driver can send an access request to the external device instead of sending the access request to the server's original remapping device, thereby bypassing the server's original remapping device. The following uses the server's original remapping device as an example to explain the implementation method for setting configuration information and access requests. When the remapping device is another type of remapping device, please refer to the corresponding implementation method for setting configuration information and access requests, and will not be repeated here.
[0095] In one possible implementation, the server's original remapping device is configured to treat the access request as a translated request. In this way, the server's original remapping device will no longer translate the target client physical address carried in the access request, so that the access request is sent to the external device as a translated request so that the external device can process it. A translated request means that the address carried by the request is a host physical address. To achieve this goal, on the one hand, the external device needs to support translated requests, that is, it needs to be able to translate the target client physical address carried by the request that the server's original remapping device believes is translated. On the other hand, it is necessary to ensure that the type of the access request is a translated request.
[0096] When the current state of the external device supports translated requests, no configuration of the external device is required. When the current state of the external device does not support translated requests, but the external device has the capability to support translated requests, configuration is required to enable the current state of the external device to support translated requests. In one possible implementation, the context entry in the external device's configuration space typically contains the device's IOMMU page table information. This IOMMU page table information can indicate the external device's ability to support translated, untranslated, and non-translation-required requests. As shown in Figure 10, the context entry uses 128 bits to represent multiple fields, each of which indicates different information. The TT field, represented by bits 3 and 2, indicates the external device's ability to support translated, untranslated, and non-translation-required requests. Table 1 shows the external device's support for translated, untranslated, and non-translation-required requests when the TT field is assigned different values. Table 1 shows that when the value of the TT field is "01," the external device's current state supports translated requests. Therefore, if the server's existing remapping device needs to be bypassed, the TT field value can be changed to "01" so that the external device's current state supports translated requests. An untranslated request refers to a request that does not require translation. For example, when an access request carries a host physical address, it is considered an untranslated request. Therefore, there is no need to determine whether the address is translated or untranslated. A translated request refers to a request that carries a translated address. A translation request indicates that the address carried in the request requires address translation. If the external device does not support translated requests, there are at least two solutions to circumvent the kernel's limitations. In the first implementation, the kernel is directly modified to set the TT field to "01b," causing the IOMMU to assume that the external device supports translated requests. In the second implementation, the external device reports to the server that it is capable of supporting translated requests, even if it does not actually support translated requests. At the same time, since the device does not initiate a Translation request, no exceptions caused by illegal access will be triggered.
[0097] Table 1
[0098] As can be seen from the above, in order to bypass the original remapping device of the server, it is necessary to ensure that the type of access request is a request that has been translated. Therefore, in order to cooperate with the implementation of this function, this application also needs to set the type of access request. The type of access request can be set by setting the field in the transaction layer packet (TLP) header of the access request. The TLP header format of the PCIe request is shown in Figure 11, where the AT field is used to specify the address type, that is, to indicate whether the address in the access request needs to be translated. Table 2 shows whether the address in the access request needs to be translated when the AT field is assigned different values. According to Table 2, this application can set the AT field to 10 to indicate that the address carried by the access request is an address that does not need to be translated, thereby bypassing the remapping device. In the eyes of the IOMMU, this situation is equivalent to the access request sent by the virtual device driver being a request that has been translated, and the GPA carried by the access request is an address that has been translated. The access request can be regarded as a request that does not require translation.
[0099] Table 2
[0100] When the external device queries the first corresponding relationship based on the target client physical address, it may be able to find the target host physical address corresponding to the target client physical address (i.e., a hit), or it may not be able to find the target host physical address corresponding to the target client physical address (i.e., a miss). For example, when the address management module does not record the corresponding relationship between the target client physical address and its corresponding target host physical address, when the external device queries the first corresponding relationship based on the target client physical address, it is unable to find the target host physical address corresponding to the target client physical address. At this time, some processing needs to be done for this situation so that the corresponding target host physical address can be queried based on the target client physical address. When the corresponding relationship is recorded through the page table, the external device cannot find the target host physical address corresponding to the target client physical address in the page table, which can be called a page fault exception. The above-mentioned processing done for it at this time is also called page fault exception processing.
[0101] In one possible implementation, the external device is further configured to send a message indicating a match failure to the address management module when the first correspondence does not record the target host physical address corresponding to the target client physical address. The message carries the target client physical address. The address management module is configured to forward the message to the virtual machine manager after receiving the message. The virtual machine manager is configured to reallocate a memory block for the target client physical address in the memory based on the message, establish a second correspondence between the host physical address of the reallocated memory block and the target client physical address, and send the second correspondence to the address management module. The address management module is configured to update the first correspondence based on the second correspondence after receiving the second correspondence. In the process of obtaining the second correspondence, the virtual machine manager may need to perform page filling, such as filling the correspondence between the missed client physical address and its corresponding host physical address in the page table. For the implementation process of determining the correspondence between the missed client physical address and its corresponding host physical address, please refer to the implementation process of obtaining the first correspondence, which will not be repeated here. When an extended page table (EPT) is provided in the kernel of the server, when processing a miss, the processing result may be optionally synchronized with the EPT.
[0102] As shown in Figure 5, the address management module is deployed in the kernel of the server, and the EPT is deployed in the kernel of the server. A DMMU is deployed in the external device, and the DMMU queries the correspondence based on the physical address of the target client. When the query based on the physical address of the target client does not hit, the server system performs page fault exception processing based on the page fault exception and synchronizes the processing result to the EPT. During the page fault exception processing process, the DMMU sends a notification message indicating that the match failed to the address management module. The message carries the physical address of the target client. After receiving the message, the address management module forwards the message to the Hypervisor. Based on the message, the Hypervisor reallocates a memory block for the target client physical address in the memory, establishes a second correspondence between the host physical address of the reallocated memory block and the physical address of the target client, and sends the second correspondence to the address management module. The address management module updates the first correspondence based on the second correspondence. Among them, the server kernel is provided with a DMMU driver, and the DMMU sends a message indicating a match failure to the address management module through the DMMU driver. The hypervisor sends a second correspondence to the address management module through the DMMU driver, and the address management module forwards the message to the hypervisor through the DMMU driver. The implementation method shown in Figure 6 is different from the implementation method shown in Figure 5 in that, in the implementation method shown in Figure 6, when a miss occurs, the DMMU sends a message indicating a match failure to the address management module through the DMA enable component and the DMMU driver, and synchronizes the processing results to the EPT, and the EPT executes the swap-in and swap-out process according to the page fault situation. The process of executing the page fault exception handling in Figure 7 is basically the same as the process in Figure 5, and will not be repeated here.
[0103] It's important to note that when updating the second correspondence based on the first correspondence, it's necessary to consider the evacuation of tasks using the first correspondence. That is, if a task using the first correspondence is already executing, the update to the second correspondence based on the first correspondence will not be executed. Once the task using the first correspondence is completed, the update to the second correspondence based on the first correspondence will be executed. This ensures that tasks using the first correspondence are correctly executed, improving the reliability of the update process.
[0104] As can be seen from the above, the server system provided by the present application can perform relevant processing when the corresponding relationship based on the target client physical address query is not named, such as executing page fault exception processing to solve the problem of miss. In this way, the memory can be pin-free before the virtual machine is started, and the global memory lock can be released. In the case where the external device is directly connected to the virtual machine in the cloud scenario, the memory pin-free makes it possible to super-multiplex the virtual machine memory without the virtual machine user's perception (no special requirements for the virtual machine system), which can improve the overall utilization of the server memory and ensure that the server's performance does not decrease. This feature is particularly evident in scenarios with high memory pressure. At the same time, since the memory is pin-free, there is no need to configure a large amount of memory in the cloud scenario due to the virtual machine pre-occupying memory, which can reduce the cost of memory use in the cloud scenario, solve the problem of gradually increasing public cloud memory costs, and thus enhance the market competitiveness of the public cloud.
[0105] In this application, user mode and kernel mode can optionally interact through commands. Kernel mode and external devices can optionally interact through commands. As shown in Figure 12, user mode and kernel mode interact through ioctl commands, and kernel mode and external devices interact through command queues (cmdq). For example, user mode and kernel mode use ioctl commands to achieve HPA address awareness. If events such as page faults or table changes occur during business operations, cmdq needs to be used to interact between kernel mode and external devices. Among them, ioctl commands mainly include mapping commands, unmapping commands, and query commands. As shown in Figures 12 and 13, the DMMU mapping command is DMMU_CMD_MAP (abbreviated as MAP), which is used to map the entire address space of a virtual machine as a memory region. When the page table is managed at the granularity of the virtual machine, DMMU_CMD_MAP carries a vmid field. The vmid field is used to indicate the virtual machine to which the client physical address recorded in the page table belongs. It should be noted that page tables can also be managed at other granularities. When managed at other granularities, the DMMU_CMD_MAP contains a field indicating the granularity to which the guest physical address recorded in the page table belongs. The page_shift field in the DMMU_CMD_MAP indicates the page offset, representing the size of the page used to record the correspondence between guest physical addresses and host physical addresses. The vaddr field in the DMMU_CMD_MAP indicates the virtual address of the mapped memory region. The iova field in the DMMU_CMD_MAP indicates the starting address for device access after mapping, which in this scenario is the starting address of the GPA and is typically 0. The length field in the DMMU_CMD_MAP indicates the length of the memory region in bytes. The cnt field in the DMMU_CMD_MAP indicates the total number of correspondences in the page table. The bdf_array field in the DMMU_CMD_MAP indicates the device performing the address translation. The DMMU unmap command, DMMU_CMD_UNMAP (abbreviated as UNMAP), is used to disassociate the guest physical address from the host physical address in the virtual machine. The DMMU query command, DMMU_CMD_QUERY (abbreviated as QUERY), instructs the client physical address carried in the command to be translated and the translation result returned to the user-mode program. The paddr field in DMMU_CMD_QUERY indicates the mapped physical address. The rsvd field in DMMU_CMD_MAP is a reserved bit.For the contents of the fields in DMMU_CMD_UNMAP and DMMU_CMD_QUERY that are identical to those in DMMU_CMD_MAP, see the descriptions of the corresponding fields in DMMU_CMD_MAP. In Figure 13, the u32, u64, and u16 above the fields indicate the length of the fields below them. For example, u16 indicates that the field below it occupies 16 bits.
[0106] In this application, the stream concurrency method can be optionally used to improve the interaction efficiency between the kernel state and the external device. As shown in Figure 14, the physical function (PF) device presented by the external device in this application can be initialized by the DMMU driver on the server side into a cmdq interaction channel. For example, the PF4 device presented by the external device can be initialized by the DMMU driver dmmu.ko on the server side into a cmdq interaction channel to realize the interaction between the kernel state and the external device. Similarly, the user state and the external device can also realize the interaction channel through the PF device presented by the external device, such as PF0, PF1, PF2, and PF3 in Figure 14.
[0107] In this application, one or more external devices can be plugged into the same server, meaning that multiple external devices can coexist. It should be noted that multiple external devices need to be isolated from each other. Figure 15 is a logical diagram of multiple external devices plugged into a server, each of which is configured with a DMMU. As shown in Figure 15, the page table resources of multiple DMMUs are stored in the DDR on the host side. The host kernel configures a DMMU device driver for each DMMU, which manages each virtual machine. Each DMMU is also configured with a DMMU memory image to support memory access. Furthermore, each virtual machine acts as a domain, and the multiple virtual functions (VFs) connected to it share the same page table resources. Therefore, a virtual machine-to-VF mapping table needs to be maintained on the DMMU device driver side. The external device side stores the external device's own basic context information and page table entry cache. The context information is independent for each VF, and different external devices are isolated from each other. In addition, Figure 15 is a schematic diagram of multiple external devices being network cards, each of which uses a dedicated PF for configuration. In this scenario, the kernel driver probes each network card and exports a dmmu device file. When starting a virtual machine, QEMU finds the corresponding device file based on the bus device function (BDF) information and uses it to perform operations such as deploying a MR. The BDF information uniquely identifies the virtual function (VF).
[0108] In addition, the granularity of the page table used in this application can be configured based on application requirements. Furthermore, this application can use multi-level page table mapping and hash mapping for page table lookups. For example, for memory pages with a granularity of 4 kilobytes (KB) or 8KB, a multi-level page table lookup can be used to save memory space, while for large pages with larger granularity, a hash mapping can be used for lookups to reduce lookup latency.
[0109] This application is equivalent to the system layer collaborating with the virtual function input / output (VFIO) device driver and the memory management unit (such as DMMU) on the external device to jointly solve the current address translation problems. When the hypervisor swaps in / out the virtual machine memory, it uses the system API to operate the VFIO driver, and then feeds back to the device's DMMU to ensure the correctness of the external device DMA, and refreshes the EPT table to ensure the correctness of the virtual machine address. The external device uses the address translation capability to sink the memory address management of the external device to the specific device, which can alleviate the pressure of IOMMU / SMMU translation and improve the performance of virtual address management.
[0110] From the above, it can be seen that the improvements of this application mainly have the following two values:
[0111] 1. As the number of VFs in the host device gradually increases, the memory management function of this application can prevent the operating performance of external devices (such as DMA performance) from decreasing. Compared with the address translation function of IOMMU / SMMU, this application enables external devices to directly manage memory, which can alleviate the performance bottleneck of IOMMU / SMMU, improve the efficiency of device address translation, and accelerate address translation.
[0112] 2. This application can improve public cloud memory utilization, reduce cloud memory costs, solve the key technical bottleneck of memory over-allocation, and enhance the market competitiveness of public clouds.
[0113] This application can initially be applied to external devices of the public cloud (such as network cards), and can later be gradually pushed to cloud core, storage and other product lines. After the standard is formed, it can be gradually pushed to the community to form the core standard of IPU / DPU.
[0114] The following describes the implementation process of the method for accessing a server system based on public cloud technology provided in an embodiment of the present application. FIG16 is a flow chart of the method for accessing a server system based on public cloud technology provided in an embodiment of the present application. As shown in FIG16 , the method for accessing a server system based on public cloud technology includes the following steps:
[0115] Step 1601: The cloud management platform obtains a virtual machine creation request input by the tenant. The virtual machine creation request carries the target specifications of the virtual machine to be created. A server that can provide the target specifications is selected from multiple servers managed by the cloud management platform, and the virtual machine is created on the selected server.
[0116] A cloud management platform is used to manage infrastructure, which includes multiple servers. When a tenant needs to create a virtual machine based on the infrastructure managed by the cloud management platform, they can perform a specified operation on the tenant's client to trigger a virtual machine creation request, which in turn causes the cloud management platform to create the virtual machine for the tenant in accordance with the virtual machine creation request. After the tenant triggers the virtual machine creation request, the cloud management platform receives the virtual machine creation request and obtains the target specifications from the request. After obtaining the target specifications set by the tenant, the cloud management platform selects a server in the infrastructure that can provide the target specifications and creates a virtual machine that meets the target specifications on the selected server. During virtual machine creation, the server's virtual machine manager must provision virtual hardware for the virtual machine based on the target specifications. For example, if the target specifications indicate that the virtual machine requires access to the server's external device capabilities and the server's memory capabilities, the virtual machine manager must perform device emulation based on the external device to obtain a virtual external device and its virtual device driver, enabling the virtual machine to utilize the external device capabilities. Furthermore, the virtual machine manager must perform device emulation based on the server's memory to obtain virtual memory, enabling the virtual machine to utilize the memory capabilities. The server's "memory" includes one or more of the following: internal memory or shared memory. When the virtual machine manager obtains virtual memory, it needs to configure the client physical address space of the virtual memory and establish a correspondence between the host physical addresses of the memory blocks configured for external devices and the client physical addresses in the client physical address space. Simultaneously, the virtual machine manager needs to send this correspondence to the address management module in the server kernel state.
[0117] Step 1602: When the virtual external device accesses the target client physical address of the virtual memory, the virtual device driver sends an access request to the external device, where the access request carries the target client physical address.
[0118] For the implementation process of step 1602, please refer to the relevant description in the previous server system. It should be noted that when the server is compatible with the original address translation function, some changes need to be made in the server so that the access request using the external device capability can be processed by the external device itself, and not by the original address translation function of the server, that is, bypassing the original address translation function of the server. In one possible implementation method, this function can be implemented by setting the configuration space of the external device, and / or setting the access request. For example, the configuration information of the external device can be set to indicate that the access request is sent directly to the external device, and the access request indicates that the address carried by the access request is the host physical address obtained through translation. The configuration information of the example external device indicates that the external device supports the request that has been translated, and the access request indicates that the address carried by the access request is the host physical address. For the implementation principle, please refer to the relevant description in the previous server system, and no further details will be given here.
[0119] Step 1603: The external device queries the first correspondence between the client physical address and the host physical address based on the access request.
[0120] For the implementation process of step 1603, please refer to the relevant description in the previous server system, which will not be repeated here.
[0121] Step 1604: After the external device finds the target host physical address corresponding to the target client physical address based on the access request, it obtains the target data recorded in the target host physical address and sends the target data to the virtual device driver.
[0122] For the implementation process of step 1604, please refer to the relevant description in the previous server system, which will not be repeated here.
[0123] When the external device queries the first correspondence based on the target client physical address, it may be able to find the target host physical address corresponding to the target client physical address (i.e., a hit), or it may not be able to find the target host physical address corresponding to the target client physical address (i.e., a miss). For example, when the address management module does not record the correspondence between the target client physical address and its corresponding target host physical address, when the external device queries the first correspondence based on the target client physical address, it is not able to find the target host physical address corresponding to the target client physical address. At this time, some processing needs to be done for this situation so that the corresponding target host physical address can be queried based on the target client physical address. Therefore, as shown in Figure 17, the access method optionally further includes:
[0124] Step 1605: When the first correspondence does not record the target host physical address corresponding to the target client physical address, the external device sends a message indicating a matching failure to the address management module, and the message carries the target client physical address.
[0125] For the implementation process of step 1605, please refer to the relevant description in the previous server system, which will not be repeated here.
[0126] Step 1606: The address management module forwards the message to the virtual machine manager.
[0127] For the implementation process of step 1606, please refer to the relevant description in the previous server system, which will not be repeated here.
[0128] Step 1607: Based on the message, the virtual machine manager allocates a memory block for the target client physical address in the memory, establishes a second correspondence between the host physical address of the memory block and the target client physical address, and sends the second correspondence to the address management module.
[0129] For the implementation process of step 1607, please refer to the relevant description in the previous server system, which will not be repeated here.
[0130] Step 1608: The address management module updates the first corresponding relationship based on the second corresponding relationship.
[0131] For the implementation process of step 1608, please refer to the relevant description in the previous server system, which will not be repeated here.
[0132] As can be seen from the above, the server system is equivalent to sinking the address translation function into the external device, allowing the external device to directly manage the memory allocated for the virtual machine and share the task of address translation of the target client physical address in the access request. In this way, there is no need for the IOMMU / SMMU or other address translation devices in the server to perform the address translation task, and there is no need for the TLB of the IOMMU / SMMU or other address translation devices in the server to record the correspondence between the target client physical address accessed by the virtual external device of the external device, thus reducing the number of correspondences that need to be recorded in the TLB, alleviating the contradiction between the small number of page table entries in the TLB and the excessive number of external devices, reducing the probability of TLB page fault problems during the address translation process, effectively improving the address translation performance of the server system, and helping to improve the operating performance (such as DMA performance) of the external device and the efficiency of memory management.
[0133] It should be noted that the order of the steps in the method for accessing a server system based on public cloud technology provided in the embodiments of the present application can be adjusted appropriately, and the number of steps can be increased or decreased accordingly. Any person skilled in the art who can easily conceive of a variation within the technical scope disclosed in this application should be included in the scope of protection of this application, and therefore will not be described in detail.
[0134] The following is an example of the basic hardware structure involved in the embodiments of the present application.
[0135] This application also provides a computing device. As shown in Figure 18 , computing device 1800 includes a bus 1802, a processor 1804, a memory 1806, and a communication interface 1808. Processor 1804, memory 1806, and communication interface 1808 communicate with each other via bus 1802. Computing device 1800 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1800.
[0136] Bus 1802 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG18 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1802 may include a path for transmitting information between various components of computing device 1800 (e.g., memory 1806, processor 1804, and communication interface 1808).
[0137] The processor 1804 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0138] Memory 1806 is used to store computer programs, including an operating system and executable code (i.e., program instructions). Memory 1806 may include volatile memory, such as random access memory (RAM). Processor 1804 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0139] The memory 1806 stores executable program code, and the processor 1804 executes the executable program code to implement the method for accessing a server system based on public cloud technology provided in the embodiment of the present application. In other words, the memory 1806 stores instructions for executing the method for accessing a server system based on public cloud technology.
[0140] The communication interface 1808 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1800 and other devices or a communication network.
[0141] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0142] As shown in Figure 19, the computing device cluster includes at least one computing device 1800. The memory 1806 in one or more computing devices 1800 in the computing device cluster may store the same instructions for executing the access method of the server system based on the public cloud technology.
[0143] In some possible implementations, the memory 1806 of one or more computing devices 1800 in the computing device cluster may also store partial instructions for executing the method for accessing a server system based on public cloud technology. In other words, the combination of one or more computing devices 1800 can jointly execute the instructions for executing the method for accessing a server system based on public cloud technology.
[0144] It should be noted that the memory 1806 in different computing devices 1800 in the computing device cluster can store different instructions, each for executing a portion of the functions of the method for accessing a server system based on public cloud technology. In other words, the instructions stored in the memory 1806 in different computing devices 1800 can implement a portion of the functions of the method for accessing a server system based on public cloud technology.
[0145] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. The network can be a wide area network (WAN) or a local area network (LAN), among others. FIG. 20 illustrates one possible implementation. As shown in FIG. 20 , two computing devices 1800A and 1800B are connected via a network. Specifically, the connection to the network is achieved via a communication interface in each computing device.
[0146] It should be understood that the functionality of the computing device 1800A shown in FIG20 may also be implemented by multiple computing devices 1800. Similarly, the functionality of the computing device 1800B may also be implemented by multiple computing devices 1800.
[0147] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection methods of the computing device clusters shown in Figures 19 and 20. However, the memory 1806 in one or more computing devices 1800 in this computing device cluster can store the same instructions for executing a method for accessing a server system based on public cloud technology.
[0148] In some possible implementations, the memory 1806 of one or more computing devices 1800 in the computing device cluster may also store partial instructions for executing the memory management method based on public cloud technology. In other words, the combination of one or more computing devices 1800 can jointly execute instructions for executing the method for accessing a server system based on public cloud technology.
[0149] The present application also provides a server. The server is used to implement the functions of the server in the server system provided in the present application. For the working principle and implementation of the server, please refer to the relevant description of the server system above.
[0150] This embodiment of the present application also provides a method for accessing a server based on public cloud technology, executed by the server provided herein. By executing this method, the server can implement the server functions described in the method for accessing a server system based on public cloud technology provided herein. For details on its operating principles and implementation, please refer to the previous description of the server system.
[0151] The present application also provides an external device. This external device is used to implement the functions of the external device in the server system provided in this application. For the working principle and implementation of this external device, please refer to the relevant description of the server system above.
[0152] This embodiment of the present application also provides a method for accessing a server based on public cloud technology, performed by an external device provided herein. By executing this method, the external device can implement the functions of the external device in the method for accessing a server system based on public cloud technology provided herein. For details on its operating principles and implementation, please refer to the relevant description of the server system above.
[0153] Embodiments of the present application also provide a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute a method for accessing a server system based on public cloud technology.
[0154] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute an access method for a server system based on public cloud technology, or instruct the computing device to execute an access method for a server system based on public cloud technology.
[0155] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0156] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the raw data and executable code involved in this application were obtained with full authorization.
[0157] In the embodiments of the present application, the terms "first," "second," and "third" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. The term "at least one" refers to one or more, and the term "plurality" refers to two or more, unless otherwise expressly limited.
[0158] In this application, the term "and / or" simply describes an association between related objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A server system based on public cloud technology, characterized in that, The server system includes a server and an external device inserted into the server, and a connection channel is established between the server and the external device; A virtual machine and a virtual machine manager are provided in the server. The virtual machine manager is used to provide virtual hardware for the virtual machine. The virtual hardware includes a virtual external device and virtual memory. The virtual external device is obtained by device simulation based on the external device, and the virtual memory is obtained by device simulation based on the memory configured for the external device. A virtual device driver of the virtual external device is provided in the virtual machine. The virtual device driver is used to send an access request to the external device when the virtual external device accesses a target guest physical address of the virtual memory, and the access request carries the target guest physical address; The external device is used to query a first correspondence between a guest physical address and a host physical address based on the access request, obtain a target host physical address corresponding to the target guest physical address, acquire target data recorded by the target host physical address, and send the target data to the virtual device driver.
2. The server system according to claim 1, wherein An address management module is further provided in the server; The external device is further used to send a message indicating a matching failure to the address management module when the first correspondence does not record a target host physical address corresponding to the target guest physical address, and the message carries the target guest physical address; The address management module is used to forward the message to the virtual machine manager; The virtual machine manager is used to allocate a memory block for the target guest physical address in the memory based on the message, establish a second correspondence between the host physical address of the memory block and the target guest physical address, and send the second correspondence to the address management module; The address management module is used to update the first correspondence based on the second correspondence.
3. The server system according to claim 1 or 2, characterized in that, An address management module is further provided in the server; The virtual machine manager is used to perform device simulation for the virtual machine based on the specification of the virtual machine to obtain the virtual memory, configure the guest physical address space of the virtual memory, establish a correspondence between the host physical address of the memory block configured for the external device in the memory and the guest physical address in the guest physical address space, and send the correspondence to the address management module.
4. The server system according to any one of claims 1 to 3, characterized in that The configuration information of the external device indicates that the external device supports a translated request, and the access request indicates that the address carried by the access request is a host physical address.
5. The server system according to any one of claims 1 to 4, characterized in that The memory configured for the external device includes one or more of the following: the memory inserted into the server or the shared memory of the server.
6. A method for accessing a server system based on public cloud technology, characterized in that, The server system includes a server and an external device inserted into the server. A connection channel is established between the server and the external device. A virtual machine and a virtual machine manager are set in the server. The virtual machine manager is used to provide virtual hardware for the virtual machine. The virtual hardware includes virtual external devices and virtual memory. The virtual external devices are obtained by device simulation based on the external device. The virtual memory is obtained by device simulation based on the memory configured for the external device. A virtual device driver of the virtual external device is set in the virtual machine. The method includes: When the virtual device driver accesses the target guest physical address of the virtual memory by the virtual external device, the virtual device driver sends an access request to the external device. The access request carries the target guest physical address. Based on the access request, the external device queries the first correspondence between the guest physical address and the host physical address, obtains the target host physical address corresponding to the target guest physical address, acquires the target data recorded by the target host physical address, and sends the target data to the virtual device driver.
7. The method according to claim 6, wherein An address management module is further set in the server. The method further includes: When the target host physical address corresponding to the target guest physical address is not recorded in the first correspondence, the external device sends a message indicating a matching failure to the address management module. The message carries the target guest physical address. The address management module forwards the message to the virtual machine manager. Based on the message, the virtual machine manager allocates a memory block for the target guest physical address in the memory, establishes a second correspondence between the host physical address of the memory block and the target guest physical address, and sends the second correspondence to the address management module. The address management module updates the first correspondence based on the second correspondence.
8. The method according to claim 6 or 7, characterized in that, An address management module is further set in the server. The method further includes: Based on the specification of the virtual machine, the virtual machine manager performs device simulation for the virtual machine to obtain the virtual memory, configures the guest physical address space of the virtual memory, establishes a correspondence between the host physical address of the memory block configured for the external device in the memory and the guest physical address in the guest physical address space, and sends the correspondence to the address management module.
9. The method according to any one of claims 6 to 8, characterized in that The configuration information of the external device indicates that the external device supports a translated request, and the access request indicates that the address carried by the access request is the host physical address.
10. The method according to any one of claims 6 to 9, characterized in that, The memory configured for the external device includes one or more of the following: the memory inserted into the server or the shared memory of the server.
11. A computing device, characterized in that, It includes a processor and a memory. Program instructions are stored in the memory. The processor runs the program instructions to enable the computing device to execute the method according to any one of claims 6 to 10.
12. A computer-readable storage medium, characterized in that, It includes program instructions. When the program instructions run on a computing device, the computing device is enabled to execute the method according to any one of claims 6 to 10.
13. A computer program product comprising instructions, characterized in that, When the instruction is run by a computing device, the computing device is caused to execute the method according to any one of claims 6 to 10.
Citation Information
Patent Citations
Memory access management unit and system of input and output device and address conversion method
CN112612574A
Computing system, memory missing page processing method and storage medium
CN114661414A
Memory access handling for peripheral component interconnect devices
US20220358049A1