Bridging across address spaces
By implementing address space bridging technology in peripheral devices, the problem of peripheral devices being unable to access the address space of other system components is solved, and the collaboration efficiency between system components and the overall performance of the computing system are improved.
Patent Information
- Application Number
- CN202210145282.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-02
- Filing Date
- 2022-02-17
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-02-17
AI Technical Summary
In existing computing systems, peripheral devices cannot efficiently access the address space of another system component while serving it, resulting in inefficient collaboration between system components.
By implementing address space bridging technology in a peripheral device, the peripheral device can access its address space while serving another system component, utilize virtualization software to define the association of address translation, and communicate through a peripheral bus.
It improves the efficiency of collaboration between system components, reduces the overhead of address translation and data transmission, and enhances the overall performance of the computing system.
Smart Images

Figure CN114996185B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates generally to computing systems, and more particularly to methods and systems for bridging memory address spaces and address translation in computing system components and peripherals. Background Art
[0002] Various types of computing systems include peripheral devices that service various system components via a peripheral bus (e.g., a Peripheral Bus Interconnect Express (PCIe) bus). Examples of such systems include network adapters that connect multiple processors to a network, or storage devices that store data for multiple processors. Such computing systems also typically include memory in which system components store data. As part of servicing the system components, the peripheral devices can access the memory to read or write data. Summary of the Invention
[0003] Embodiments of the invention described herein provide a computing system comprising at least one peripheral bus, a peripheral device connected to the at least one peripheral bus, at least one memory, and a first system component and a second system component. The first system component is (i) associated with a first address space in the at least one memory and (ii) connected to the peripheral device via the at least one peripheral bus. The second system component is (i) associated with a second address space in the at least one memory and (ii) connected to the peripheral device via the at least one peripheral bus. The first system component is arranged to enable the peripheral device to access the second address space associated with the second system component.
[0004] In one embodiment, one or both of the first and second system components are physical processors. In another embodiment, one or both of the first and second system components are virtual processors. In some embodiments, the first and second system components run on different servers. In other embodiments, the first and second system components run on the same server.
[0005] In some embodiments, the first system component is a first virtual processor arranged to use a first address translation to access a first address space, the second system component is a second virtual processor arranged to use a second address translation to access a second address space, and the system further includes virtualization software configured to define an association between at least a portion of the first address translation and at least a portion of the second address translation and enable the peripheral device to access the second address space using the association.
[0006] In an example embodiment, the peripheral device includes a network adapter, and virtualization software is configured to define an association between a first remote direct memory access (RDMA) memory key (MKEY) and a second remote direct memory access memory key corresponding to a first address translation and a second address translation, respectively. In a disclosed embodiment, the first system component is configured to cause the peripheral device to transfer data via remote direct memory access (RDMA) to the second address space using a queue pair (QP) controlled by the first system component. In one embodiment, the QP and the transfer of data via RDMA are associated with a network association of the second system component.
[0007] In another embodiment, the first virtual processor is arranged to subject the peripheral device to a security policy governing access to the address space to access the second address space. In yet another embodiment, the peripheral device includes a network adapter configured to save a local copy of an association between at least a portion of the first address translation and at least a portion of the second address translation, and to access the second address space using the local copy of the association.
[0008] In a disclosed embodiment, the second system component is a virtual processor, and the first system component is configured to run virtualization software that (i) allocates physical resources to the virtual processor and (ii) accesses the second address space of the virtual processor. In another embodiment, the second system component is configured to specify to the peripheral device whether the first system component is permitted to access the second address space, and the peripheral device is configured to access the second address space only upon verifying that the first system component is permitted to access the second address space.
[0009] In yet another embodiment, the at least one peripheral bus includes a first peripheral bus connected to a first system component and a second peripheral bus connected to a second system component; and the at least one memory includes a first memory having the first address space and associated with the first system component, and a second memory having the second address space and associated with the second system component. In an embodiment, the second system component runs in a server integrated into a network adapter.
[0010] In another embodiment, the peripheral device comprises a network adapter arranged to transmit a data packet to a network, the first system component and the second system component have respective first and second network addresses, and the second system component is arranged to: (i) construct a data packet comprising obtaining portions of the data packet from the first address space of the first system component; and (ii) transmit the data packet to the network via the network adapter using the second network association. In yet another embodiment, the peripheral device comprises a storage device.
[0011] In some embodiments, the peripheral device includes a processor that emulates a storage device by exposing a non-volatile memory express (NVMe) device on a peripheral bus. In an exemplary embodiment, the peripheral device that emulates the storage device is a network adapter. In an embodiment, the processor in the peripheral device is configured to complete input / output (I / O) transactions of the first system component by accessing a second address space. In an exemplary embodiment, the processor in the peripheral device is configured to directly access the second address space using remote direct memory access (RDMA). In disclosed embodiments, the processor in the peripheral device is configured to complete the I / O transactions using a network association of the second system component.
[0012] In some embodiments, the peripheral device includes a network adapter, and the first and second system components execute on different physical hosts serviced by the network adapter.
[0013] In some embodiments, the peripheral device is a network adapter configured to emulate a graphics processing unit (GPU) by exposing GPU functionality over a peripheral bus. In an embodiment, the second system component is configured to perform network operations on behalf of the emulated GPU.
[0014] According to an embodiment of the present invention, there is also provided a computing method, the computing method comprising: communicating between a peripheral device and a first system component via at least one peripheral bus, wherein the first system component is (i) associated with a first address space in at least one memory and (ii) connected to the peripheral device via the at least one peripheral bus; and communicating between the peripheral device and a second system component via the at least one peripheral bus, wherein the second system component is (i) associated with a second address space in the at least one memory and (ii) connected to the peripheral device via the at least one peripheral bus. Using the first system component, the peripheral device accesses a second address space associated with the second system component. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The present invention will be more fully understood from the following detailed description of embodiments with reference to the accompanying drawings, in which:
[0016] Figure 1 is a block diagram schematically illustrating a computing system according to one embodiment of the present invention;
[0017] Figure 2 is a block diagram schematically illustrating a virtualized computing and communication system according to an embodiment of the present invention;
[0018] Figure 3 is a flowchart schematically illustrating a method for address space bridging according to an embodiment of the present invention; and
[0019] Figure 4 is a block diagram schematically illustrating a computing and communication system using a Smart-NIC according to one embodiment of the present invention. DETAILED DESCRIPTION
[0020] Overview
[0021] Various types of computing systems include peripheral devices that serve multiple system components via at least one peripheral bus. Examples of such systems include network adapters that connect multiple processors to a network, or storage devices that store data for multiple processors. The peripheral bus may include, for example, a Peripheral Bus Interconnect Express (PCIe) bus.
[0022] Any system component and / or peripheral device can be physical or virtual. In a virtualization system, for example, a physical computer hosts one or more virtual machines (VMs). The physical computer typically runs virtualization software ("hypervisor") that allocates physical resources to the VMs. Specifically, the hypervisor allocates the resources of the peripheral devices to the VMs ("virtualized peripherals"). For example, each VM can be assigned a virtual network interface controller (VNIC) in a physical network adapter and / or a virtual disk (VDISK) in a physical disk. Typically, each VM and the hypervisor has a corresponding network association (also called a network identity, an example of which is an IP address). The hypervisor can use its own network association to provide some services, for example in storage services provided to the VMs. Other services (e.g., VM to VM communication) will use the network association of the VM.
[0023] Such a (physical or virtualized) computing system typically also includes at least one memory, wherein system components are assigned corresponding address spaces. The address space assigned to a certain system component is typically associated with corresponding address translations between, for example, virtual addresses used by the system component and physical addresses of the memory.
[0024] Systems such as those described above typically enable peripheral devices to directly access memory while servicing system components. Example use cases include network devices that scatter and / or gather data packets, storage devices that service requests and scatter and / or gather data, or remote direct memory access (RDMA) network interface controllers (NICs) that perform large memory transactions.
[0025] RDMA is a network protocol for transferring data from one memory to another. RDMA endpoints typically use queue pairs (QPs) for communication and have a network association (e.g., an IP address). Remote storage protocols such as NVMe-over-fabrics can be implemented via RDMA to access remote storage. In such protocols, data transfers are initiated by a storage target on the QP of an RDMAREAD / WRITE to transfer data. A network device that emulates a local NVMe device typically carries a network association (e.g., an IP address) that has a QP to service NVMe commands to NVMe-over-fabrics on the QP. The disclosed technology enables RDMA transactions arriving from a storage server to be passed to the PCIe bus via the PCIe association of the emulated NVMe device, where the network association is a network function and network operations are controlled by the software entity of the network.
[0026] To implement direct memory access, peripheral devices are typically made aware of the address spaces and corresponding address translations used by system components. For example, a VNIC serving a VM may save a local copy of the address translations used by the VM and use this local copy to directly access the address space of the VM in memory. In this implementation, each address translation may be identified by an identifier or handle called an "MKEY". The NIC may save a translation and protection table (TPT), accessed, for example, via {VM identifier, MKEY}, which stores a local copy of each address translation used by each VM. The MKEY is primarily used to convert virtual addresses to physical addresses, and to abstract complex scatter and gather operations.
[0027] Conventionally, a given system component (e.g., a physical CPU or VM) can only access its own address space in memory and is completely unaware of and unable to access the address spaces of other system components. Peripherals can also conventionally only access the address space of the system component they are assigned to service.
[0028] However, in various practical use cases, it would be very beneficial if a peripheral device could access the address space of one system component while servicing another. Such a capability would enable system components to collaborate in performing complex tasks while efficiently accessing memory with minimal address translation and data transfer. This capability would decouple the control of functions and their associated functions from the actual data transfer.
[0029] Embodiments of the invention described herein provide improved methods and systems in which peripheral devices are given the ability and permission to access the address space of one system component while servicing another system component. In various embodiments, the peripheral devices may include, for example, a network adapter or a storage device. The system components serviced by the peripheral devices may include, for example, a physical processor and / or a virtual machine (VM). Communication between system components and peripheral devices, as well as access to memory, is performed via at least one bus (e.g., a PCIe bus).
[0030] The disclosed technology is useful in a variety of system configurations. Several illustrative examples are described herein. Some system configurations are completely physical ("bare metal") configurations, while other configurations use virtualization.
[0031] For example, in one embodiment, the system includes multiple physical processors (eg, CPUs or GPUs) that share a peripheral device (eg, a network adapter). The disclosed technology allows a network adapter to access the address space of one processor while servicing another processor.
[0032] In another embodiment, the system includes multiple VMs that are assigned corresponding virtual NICs (VNICs) in a physical network adapter. The VNICs are configured to use RDMA to access the address space of their corresponding VMs in memory based on the corresponding MKEYs. Virtualization and resource allocation are typically coordinated by a hypervisor in this system. In such a system, the disclosed technology enables a VNIC serving a particular VM to access the address space of another VM. To this end, the hypervisor typically defines an association between MKEYs (between the MKEY of the accessing VM and the MKEY of the VM being accessed). The hypervisor typically records the association between the MKEYs in the NIC, for example, in a translation and protection table (TPT). In this way, the memory access operations of the VNIC (including accessing another VM on behalf of one VM) can be performed by the VNIC without involving the VM.
[0033] In some virtualization environments, each virtualized system component is assigned a unique PCIe bus-device-function (BDF), and virtualized peripherals use the BDF to access memory on behalf of the system component. When used in such an environment, the disclosed technology enables peripherals to access the address space associated with one BDF while servicing another BDF ("on behalf of another BDF"). Therefore, the disclosed technology is sometimes referred to as "cross-function bridging." In this context, the hypervisor is also considered a system component that has its own BDF but can use the VM's BDF to access the address space of a hosted VM.
[0034] Typically, a computing system defines a security policy that governs cross-functional bridging operations (or otherwise authorizes or denies system components access to the address spaces of other system components). In some embodiments, the security policy is also stored in the TPT along with the association between MKEYs. The security policy can be defined as, for example, a pair of {access MKEY, MKEY being accessed}.
[0035] It should be noted that in some use cases, the "accessed system component" (the system component whose address space is being accessed) and its memory reside in one physical machine, and the "accessing system component" (the system component accessing the address space of the accessed system component) resides in a different physical machine. For example, the disclosed technology can be deployed in a multi-host system where a peripheral device serves multiple system components (e.g., CPUs or VMs) that reside in two or more different physical computers (e.g., servers). Another example involves a "smart-NIC," i.e., a NIC with an on-NIC server. The disclosed technology can be used to enable an on-NIC server to access the address space of a system component (e.g., a CPU or VM) served by the smart-NIC.
[0036] The ability of a peripheral device to access the address space of one system component while servicing ("on behalf of") another system component is a powerful tool that can be used to enhance the efficiency of many computing tasks. Examples of peripheral devices that can use this tool are, for example, different virtualized peripherals that interact with memory, such as network adapters, CPUs, GPUs, solid-state drives (SSDs), and others. These examples are by no means limiting and are provided purely by way of example.
[0037] Cross-memory bridging in "bare metal" systems
[0038] Figure 1 is a block diagram schematically illustrating a computing system 20 according to one embodiment of the present invention. System 20 includes two central processing units (CPUs) 24, represented as CPU1 and CPU2. CPU1 and CPU2 are coupled to respective memories 28, represented by MEM1 and MEM2, which in this example are random access memory (RAM) devices. CPU1 and CPU2 communicate with MEM1 and MEM2 via respective peripheral bus interconnect (PCIe) buses 36, represented as PCIe1 and PCIe2. PCIe peripheral devices 32 are connected to the two PCIe buses 36 and are shared by PCU1 and CPU2. Peripheral devices 32 may include, for example, a network adapter (e.g., a network interface controller - NIC) or a disk. In this document, CPU1 and CPU2 are examples of system components served by peripheral devices.
[0039] In one example, system 20 is a multi-host system in which two servers are connected to the network via a single NIC. One server includes CPU1, MEM1, and PCIe1, and the other server includes CPU2, MEM2, and PCIe2. However, Figure 1 The system configuration shown in is merely an example chosen purely for conceptual clarity. Any other suitable configuration may be used in alternative embodiments.
[0040] For example, the disclosed technology is not limited to CPUs and may be used with any other suitable type of processor or other system component, such as a graphics processing unit (GPU). As another example, instead of two separate PCIe buses 36, the system 20 may include a single PCIe bus 36 connecting the CPU 24, the memory 28, and the peripheral devices 32. In addition or alternatively, instead of two separate memory devices 28, MEM1 and MEM2 may include two separate memory areas assigned to CPU1 and CPU2 in the same memory device. Such memory areas do not have to be contiguous. As another example, the system may include more than one single peripheral device, more than two CPUs (or other processors), and / or more than two memories. However, for clarity, the following description will refer to the description in Figure 1 The exemplary configuration described in .
[0041] In some embodiments, CPU1 stores and retrieves data from MEM1 by accessing the address space designated ADDRSPACE1. CPU2 stores and retrieves data from MEM2 by accessing the address space designated ADDRSPACE2. Peripheral device 32 typically knows ADDRSPACE1 when servicing CPU1 and knows ADDRSPACE2 when servicing CPU2. For example, peripheral device 32 can be configured to read data directly from MEM1 (or write data directly to MEM1) based on ADDRSPACE1 on behalf of CPU1. Similarly, peripheral device 32 can be configured to read data directly from MEM2 (or write data directly to MEM2) based on ADDRSPACE2 on behalf of CPU2.
[0042] Conventionally, the two address spaces ADDRSPACE1 and ADDRSPACE2 are separate and independent, and each CPU is unaware of the address space of the other CPU. However, in some practical scenarios, it is efficient for peripheral device 32 to access MEM1 while executing the tasks of CPU2 (or vice versa).
[0043] For example, consider an embodiment in which peripheral device 32 is a NIC that connects CPU1 and CPU2 to a packet communication network. In this example, CPU1 and CPU2 collaborate to transmit data packets to the network, with the work divided between them. CPU1 prepares the data (payload) of the data packet and stores the payload in MEM1. CPU2 is responsible for preparing the header of the data packet, storing the header in MEM2, constructing the data packet from the respective header and payload, and transmitting the data packet to the network. The data packet is transmitted to a network that has a network association (e.g., an IP address) with CPU2. CPU1 does not need to be involved or aware of this process, for example, CPU2 can obtain data from CPU1 and construct the data packet, while CPU1 can be requested to read data from the local disk.
[0044] In an embodiment, system 20 constructs and transmits data packets efficiently by enabling the NIC to access the packet payload (which is stored in MEM1) while serving CPU2 (which runs the process of constructing the packet from the header and payload).
[0045] In an exemplary implementation, CPU1 configures a peripheral device (a NIC in this example) to access ADDRSPACE1 in MEM1 when servicing CPU2. Configuring the NIC by CPU1 typically involves granting the NIC security permissions to access ADDRSPACE1 in MEM1 when servicing CPU2.
[0046] When the NIC is configured as described above, CPU2 and the NIC effectively execute the process of constructing and transmitting a packet. In an embodiment, when servicing CPU2, the NIC alternately accesses MEM2 (based on ADDRSPACE2) and MEM1 (based on ADDRSPACE1) (retrieving a header, retrieving the corresponding payload, retrieving the next header, retrieving the corresponding payload, and so on). This process is executed under the control of CPU2 without involving CPU1, even though the packet data resides in MEM1.
[0047] Cross-function bridging in virtualized systems
[0048] Figure 2is a block diagram schematically illustrating a virtualized computing and communication system 40 according to another embodiment of the present invention. System 40 may include, for example, a virtualization server running on one or more physical host computers. System 40 includes multiple virtual machines (VMs) 44. For clarity, the figure shows only two VMs, designated VM1 and VM2. VMs 44 store data in physical memory 52 (e.g., RAM) and communicate over a packet network 60 via a physical NIC 56. In this context, VMs 44 are examples of system components, and NIC 56 is an example of a peripheral device that services them. Typically, although not necessarily, system 40 communicates over network 60 using remote direct memory access (RDMA).
[0049] System 40 includes a hypervisor (HV) 48 (also referred to herein as "virtualization software") that allocates the physical resources of the hosted computer (e.g., computing, memory, disk, and network resources) to VMs 44. Specifically, HV 48 manages the storage of data by the VMs in memory 52 and the communication of data packets between the VMs and NIC 56.
[0050] For clarity, some physical elements of the system are omitted from the figure, such as the processor or processors hosting the VMs and HVs, and the PCIe bus connecting the processors, memory 52, and NIC 56. NIC 56 includes a NIC processor 84 (also referred to as the NIC's "firmware" (FW)).
[0051] In this example, VM1 runs a guest operating system (OS) 68, designated OS1. OS1 runs a virtual CPU (not shown), which runs one or more software processes 64. Similarly, VM2 runs a guest operating system 68, designated OS2, which in turn runs a virtual CPU, which runs one or more software processes 64. Each process 64 can store and retrieve data in memory 52 and / or send and receive data packets via NIC 56. As part of the virtualization scheme, FW 84 of NIC 56 runs two virtual NICs (VNICs) 80, designated VNIC1 and VNIC2. VNIC1 is assigned to VM1 (via HV 48), and VNIC2 is assigned to VM2 (via HV 48).
[0052] Typically, data flow within VM 44 and between VM 44, memory 52, and NIC 56 involves multiple address translations between multiple address spaces. In this example, process 64 in VM1 writes data to and reads data from memory based on a space of guest virtual addresses (GVAs) designated GVA1. Guest OS 68 of VM1 (OS1) translates between GVA1 and a space of guest physical addresses (GPAs) designated GPA1. Similarly, process 64 in VM2 writes data to and reads data from memory based on a space of GVAs designated GVA2. Guest OS 68 of VM2 (OS2) translates between GVA2 and a space of GPAs designated GPA2.
[0053] Among other components, HV 48 runs an input-output memory management unit (IOMMU) 72, which manages the storage of data in physical memory 52 and, in particular, converts GVAs into physical addresses (PAs). For read and write commands associated with VM1, the IOMMU converts between GPA1 and the space of the PA, designated PA1. For read and write commands associated with VM2, the IOMMU converts between GPA2 and the space of the PA, designated PA2. Data in memory 52 is stored and retrieved based on the PA.
[0054] In some embodiments, when sending and receiving packets, VNIC 1 is configured to directly read data from and / or write data to memory 52 on behalf of VM 1. Similarly, VNIC 2 is configured to directly read data from and / or write data to memory 52 on behalf of VM 2. Process 64 in VM 44 may issue an RDMA "scatter" or "gather" command that instructs VNIC 80 to write data to / read data from a list of address ranges in memory 52.
[0055] RDMA scatter and gather commands are typically expressed in terms of GVAs. To execute a scatter or gather command in memory 52, the VNIC must convert the GVA into the corresponding GPA (which the IOMMU in turn converts to PA). To simplify this process, in some embodiments, the NIC 56 includes a translation and protection table (TPT) 88 that stores the addresses of different VMs. Typically, each VM's guest OS is used The converted copy is used to configure the TPT 88. Using the TPT 88, the VNIC 80 can execute the The address is translated without going back to VM 44.
[0056] With above Figure 1As in the "bare metal" example of FIG, in a virtualized system such as system 40, it is highly advantageous to enable one processor to access the address space of another processor. In the context of this disclosure and in the claims, a VM 44 running a virtual processor (such as a virtual CPU) is also considered herein to be a type of processor. A hypervisor 48 is also considered herein to be a processor.
[0057] For example, consider again the example of two processors collaborating in transmitting a packet to a network. However, in this embodiment, Figure 2 VM1 and VM2 in virtualized system 40 perform packet transmission. VM1 prepares the data (payload) of the packet and stores the payload in memory 52. VM2 prepares the packet header, stores the header in memory 52, constructs the packet from the header and payload, and transmits the packet to VNIC2 in NIC 56 for transmission to network 60. The packet is transmitted using VM2's network association (e.g., IP address).
[0058] As mentioned above, VM1's task of storing the payload in memory 52 involves a number of address translations:
[0059] ■ The providing process 64 in VM1 writes the payload to memory according to the GVA in the address space GVA1.
[0060] ■ OS1 of VM1 converts GVA into GPA in address space GPA1
[0061] ■The IOMMU 72 in the HV 48 converts the GPA into the PA in the address space PA1
[0062] ■ HV 48 stores the payload in memory 52 according to PA1.
[0063] VM2's task of saving the header in memory 52 also involves multiple address translations, but with a different address space, using the same token:
[0064] ■ The header providing process 64 in VM2 writes the header to memory according to the GVA in the address space GVA2.
[0065] ■ OS2 of VM2 converts GVA into GPA in address space GPA2
[0066] ■The IOMMU 72 in the HV 48 converts the GPA into the PA in the address space PA2
[0067] ■ HV 48 stores these headers in memory 52 according to PA2.
[0068] In an embodiment, system 40 efficiently constructs and sends packets by enabling VNIC 2 (which serves VM 2) to access both packet headers (which are stored in memory 52 according to PA 2) and packet payloads (which are stored in memory 52 according to PA 1). In the virtualized environment of system 40, this example illustrates a technique for configuring a peripheral device (in this example, VNIC 2 in NIC 56) to access the address space of one processor (in this example, GVA 1 / GPA 1 of VM 1) while serving another processor (in this example, VM 2).
[0069] Configuring VNIC2 in this manner typically involves granting VNIC2 security permissions to access VM1's address space (GVA1 / GPA1) while servicing VM2. In some embodiments, HV 48 maintains associations 76 between the address spaces of different VM2s, including appropriate security permissions.
[0070] When VNIC 2 is configured as described above, the process of constructing and transmitting a packet can be efficiently performed. Typically, VNIC 2 alternately accesses the header (using VM 2's address space—GVA2 / GPA2 / PA2) and the payload (using VM 1's address space—GVA1 / GPA1 / PA1) in memory 52 (retrieves a header, retrieves the corresponding payload, retrieves the next header, retrieves the corresponding payload, and so on). This process is performed for VM 2 without involving VM 1, even though the packet data is stored by VM 1 using VM 1's address space.
[0071] In some embodiments, the virtualization scheme of system 40 assigns a unique PCIe Bus-Device-Function (BDF) to each VM. Specifically, VM1 and VM2 are assigned different BDFs, denoted as BDF1 and BDF2, respectively. Using the disclosed techniques, hypervisor 48 instructs VNIC2 to access VM1's address space using BDF1 on VM1's behalf.
[0072] In some embodiments, the disclosed techniques are used to enable the hypervisor 48 itself to access the address space of the VM 44 it hosts on behalf of the VM. In an embodiment, the hypervisor knows the BDF of the VM and uses the BDF to access the address space of the VM.
[0073] Figure 3 FIG2 is a flow chart schematically illustrating a method for address space bridging according to an embodiment of the present invention. The address space bridging technology is explained using an example of packet transmission, but is by no means limited to this use case.
[0074] Figure 3 The method begins at an access authorization step 90 where the NIC processor 84 ("NIC firmware") of the NIC 56 grants VM2 security permission to access data belonging to VM1.
[0075] In indication step 94, VM2 sends a collect command to VNIC2 that specifies the construction of the data packet to be transmitted. In this example, the collect command includes a list of entries that alternate between the address spaces of VM2 and VM1:
[0076] ■ Retrieve the header of Packet 1 from the GVA in GVA2 (of VM2).
[0077] ■ Retrieve the payload of packet 1 from the GVA in GVA1 (of VM1).
[0078] ■ Retrieve the header of packet 2 from the GVA in GVA2 (of VM2).
[0079] ■ Retrieve the payload of packet 2 from the GVA in GVA1 (of VM1).
[0080] ■ Retrieve the header of packet 3 from the GVA in GVA2 (of VM2).
[0081] ■…
[0082] In execution step 98 , VNIC 2 executes the collection commands in memory 52 , thereby constructing data packets and transmitting them over network 60 .
[0083] As can be appreciated, packet reception and separation into headers and payloads can be implemented in a similar manner. In such an embodiment, VNIC 2 receives packets, separates each packet into a header and payload, stores the header in memory 52 according to VM 2's address space, and stores the payload in memory 52 according to VM 1's address space. To perform this task, VM 2 sends a scatter command to VNIC 2, which includes a list of entries that alternate between VM 2's and VM 1's address spaces:
[0084] ■ Save the header of packet 1 in the GVA in GVA2 (of VM2).
[0085] ■ Save the payload of packet 1 in the GVA in GVA1 (of VM1).
[0086] ■ Save the header of packet 2 in the GVA in GVA2 (of VM2).
[0087] ■ Save the payload of packet 2 in the GVA in GVA1 (of VM1).
[0088] ■ Save the header of packet 3 in the GVA in GVA2 (of VM2).
[0089] ■…
[0090] Another example use case for address space bridging involves RDMA scatter and gather commands. In some embodiments, VM1 can perform an RDMA READ from a remote server and direct the remote read to VM2's memory space. The fact that this access is performed due to the networking protocol is completely abstracted from VM2.
[0091] Another example use case involves the Non-Volatile Memory Express (NVMe) storage protocol. In some embodiments, VM1 can issue an NVMe request to an NVMe disk (which collects data from VM2) and write the data to an address on the disk specified by VM1. Control of this operation is performed by VM1.
[0092] Another possible use case involves NVMe emulation technology. NVMe emulation is addressed, for example, in U.S. Patent 9,696,942, entitled "Accessing remote storage devices using a local bus protocol," the disclosure of which is incorporated herein by reference. In some embodiments, VM2 is exposed to an emulated NVMe function whose commands are terminated by VM1. VM1 performs the appropriate networking tasks to service these NVMe commands to the remote storage server and instructs the NIC to directly distribute data retrieved from the network on behalf of the NVMe function to VM2's buffer in the NVMe command.
[0093] The above use cases are described by way of example only. In alternative embodiments, the disclosed technology can be used in a variety of other use cases and scenarios, and can be used with a variety of other types of peripheral devices.
[0094] Using MKEY and TPT bridge
[0095] In some embodiments, each of the system 40 An address translation is identified by a corresponding handle (also referred to as a "translation identifier" or "translation pointer"). In the following description, this handle is referred to as a memory key (MKEY). Typically, the MKEY is assigned by the guest OS 68 of the VM 44 and is therefore unique only within the scope of the VM 44.
[0096] In some embodiments, when a VM defines a new When converting and assigning an MKEY to it, the VM also A copy of the translation is stored in the TPT 88 of the NIC 56. The TPT 88 thus includes the translations used by different processes 64 on different VMs 44. Since MKEY is not unique outside the scope of the VM, each The conversion is resolved by {MKEY, VM identifier}.
[0097] By using the TPT, VNIC 80 in NIC 56 can access (read and write) data in memory 62, for example, for executing scatter and gather commands, without involving the VM for address translation. The use of TPT 88 enables the VNIC to "zero-copy" data into the VM's process buffer using the GVA, i.e., copy the data without involving the VM's guest operating system.
[0098] In some embodiments of the invention, NIC 56 uses TPT 88 to enable a VNIC 80 assigned to serve one VM 44 to access the address space of another VM 44 .
[0099] In one embodiment, referring to the packet generation example above, VM2 may have two MKEYs, represented as MKEY_A and MKEY_B. MKEY_A points to VM2's own address space, while MKEY_B points to VM1's MKEY. Both MKEY_A and MKEY_B have copies stored in TPT 88. In order for VM2 to access VM1's address space, VM2 instructs VNIC2 (in NIC 56) to use MKEY_B to access memory.
[0100] So, for example, to construct a sequence of packets, VM2 may send a collect command to VNIC2 that alternates between VM2's address space (holding the header) and VM1's address space (holding the payload):
[0101] ■ Retrieve the header of packet 1 from the GVA in GVA2 (of VM2) using MKEY_A.
[0102] ■ Retrieve the payload of packet 1 from the GVA in GVA1 (of VM1) using MKEY_B.
[0103] ■ Retrieve the header of packet 2 from the GVA in GVA2 (of VM2) using MKEY_A.
[0104] ■ Retrieve the payload of packet 2 from the GVA in GVA1 (of VM1) using MKEY_B.
[0105] ■ Retrieve the header of packet 3 from the GVA in GVA2 (of VM2) using MKEY_A.
[0106] ■…
[0107] VNIC2 typically executes the gather command without involving VM2 by using a copy of the address translation stored in TPT 88 (and the corresponding MKEY).
[0108] For example, a scatter command for disassembling a packet can also be implemented by alternating between MKEYs pointing to the address spaces of different VMs.
[0109] In some embodiments, TPT 88 also stores security permissions that specify which VMs are permitted to access which MKEYs of other VMs. In the above example, for example, VM1 initially specifies in TPT 88 that VM2 is permitted to access VM1's address space (holding a payload) using MKEY_B. When NIC processor 84 (NIC "FW") receives a request from VNIC2 to access VM1's address space using MKEY_B, it will grant the request only if TPT 88 indicates that such access is permitted. In the above security mechanism, there is no centralized entity that defines security permissions. Instead, each VM specifies which VMs (if any) are permitted to access its memory space and which MKEYs to use.
[0110] Smart-NIC, emulated NVME storage devices, and decoupling the data plane from the control plane
[0111] In paravirtualization, as is known in the art, a hypervisor emulates network or storage (e.g., SSD) functionality. At a lower level, the hypervisor services requests by translating addresses into its own address space, adding headers or executing some logic, and transmitting to the network or storage using the hypervisor's own network associations (sometimes referred to as the underlying network). In some disclosed embodiments, such functionality can be performed by the peripheral device (e.g., by a hypervisor running on a different server or within a network adapter). Instead of the hypervisor performing the address translation, the hypervisor can use the disclosed techniques to request that the peripheral device perform DMA operations on its behalf, in which case the address translation will be performed by the IOMMU.
[0112] Figure 4This is a block diagram schematically illustrating a computing and communication system 100 according to another embodiment of the present invention. This example demonstrates several aspects of the disclosed technology: (i) implementation using a "smart-NIC"; (ii) implementation emulating an NVMe storage device; and (iii) decoupling between the control plane and the data plane. This example also illustrates how, using a queue pair (QP) controlled by the first system component, a first system component can cause a peripheral device to transfer data to a second address space (associated with a second system component) via RDMA. The data carries the network association (e.g., IP address) of the second system component.
[0113] System 100 includes a NIC 104 and a processor 108. NIC 104 is used to connect processor 108 to a network 112. Processor 108 can include a physical processor (e.g., a CPU or GPU) or a virtual processor (e.g., a virtual CPU of a VM running on a hosted physical computer). NIC 104 is referred to as a "smart-NIC" because, in addition to a NIC processor ("FW") 120, it includes a server on the NIC 116. Server on the NIC 116 includes memory 122 and a CPU 124. CPU 124 is configured to run an operating system and general-purpose software.
[0114] In this example, processor 108 includes memory 144 in which data is stored. Processor 108 stores data according to address translation with an MKEY, designated MKEY_A. CPU 124 of server 116 on the NIC is assigned another MKEY, designated MKEY_B. To efficiently transmit data from memory 144 to network 112 and / or efficiently receive data from network 112 to memory 144, processor 108 configures CPU 124 of server 116 on the NIC to access data in memory 144 on its behalf. CPU 124 may be unaware of memory accesses performed by server 116 on the NIC.
[0115] In an embodiment, the NIC processor 120 of the NIC 104 includes a TPT 132, which, as described above, stores a copy of the address translations, including MKEY_A and MKEY_B, used by the processor 108 and the CPU 124. The TPT 132 also indicates permission for the CPU 124 to access data in the memory 144 on behalf of the processor 108. This permission can be indicated as, for example, an association between MKEY_A and MKEY_B.
[0116] A diagonal line 136 in the diagram marks the boundary between the portion of system 100 associated with MKEY_A and the portion associated with MKEY_B. The portion associated with MKEY_B is shown in a shaded pattern (to the left of line 136), while the portion associated with MKEY_A is shown in a clear pattern (above line 136). The diagram illustrates the boundary created between the two VNICs (the VNIC for processor 108 and the VNIC for server 116 on the NIC) and shows that the two memory keys in each domain are tightly coupled.
[0117] As shown, the control path for transmitting and receiving data between the network 112 and the memory 144 includes: (i) a queue 140 in the processor 108; (ii) a queue 126 in the server 116 on the NIC; and a queue pair (QP) 128 in the NIC processor 120. On the other hand, the data path for data transmission is directly between the memory 144 and the network 112, via the NIC processor 120. In particular, the data path does not traverse the server 116 on the NIC.
[0118] In an example implementation of a virtualized system, a Smart-NIC 104 serves one or more servers, each of which hosts multiple VMs 108. In this example, a server 116 on the NIC runs a hypervisor that serves the various VMs. The disclosed technology enables the hypervisor (running on the CPU 124) to access the memory and address space of the various VMs 108.
[0119] Use Cases for Multiple Hosts
[0120] In some cases, peripheral devices are configured to connect to multiple hosts. In this architecture (referred to as "multi-host"), a given virtualized device ("logical device") serves multiple hosts that have corresponding physical address spaces and virtual address spaces in memory. This type of device can include, for example, a network adapter or a dual-port storage device (e.g., a solid-state drive - SSD) that supports multiple hosts. For example, multi-host operation is proposed in U.S. Patent 7,245,627, entitled "Sharing a network interface card among multiple hosts," the disclosure of which is incorporated herein by reference. In some embodiments, using the disclosed bridging technology, such peripheral devices can access the address space of one host on behalf of another host.
[0121] Use cases with peripherals connected to FPGAs or GPUs
[0122] In some embodiments, the peripheral device includes a network adapter (e.g., a NIC) connected to another device that serves a different processor (e.g., a physical or virtual host and / or VM). However, unlike the use case of the smart-NIC, the other device in this use case is not a server. The other device may include, for example, a field programmable gate array (FPGA) or a graphics processing unit (GPU). In one example use case, the other device (e.g., an FPGA or GPU) is accessible to the entire PCIe topology connected to the NIC and through its slices.
[0123] For example, such devices may support single root input / output virtualization (SRIOV) and implement the disclosed techniques similar to those described above for network adapters and storage devices. In some embodiments, one or more FPGAs and / or GPUs may be emulated by a PCIe endpoint and abstracted to a VM, where a VM uses the disclosed techniques to serve different emulated GPUs / FPGAs.
[0124] Use case with peripherals emulated as NVME devices
[0125] In some embodiments, a peripheral device (e.g., a Smart-NIC) exposes itself on the PCIe bus as a Non-Volatile Memory Express (NVMe) compatible storage device or any other suitable storage interface on PCIe. However, internally, the peripheral device includes a processor (e.g., a microcontroller) that virtualizes the actual physical storage device. The physical storage device can be located locally or remotely. Using the disclosed techniques, the processor can access the address space of a different processor (physical or virtual) on its behalf (e.g., access the address space of a VM on behalf of the VM) to complete IO transactions.
[0126] Completing the transaction may include using RDMA operations to remote storage on the QP / Ethernet NIC associated with the service VM, which in turn will perform the network transaction (RDMA or non-RDMA) directly towards the buffer specified in the NVMe request in the VM that sees the NVMe emulation. The VM is unaware of the network protocol in the emulated device and is unaware of RDMA or MKEYs. The VM simply issues an NVMe request with scatter-gathering being serviced by another VM that performs all networking tasks.
[0127] Figure 1 、 2 The configurations of the various systems, hosts, peripherals, and other elements shown in FIG4 and their respective components are example configurations depicted purely for conceptual clarity. Any other suitable configuration may be used in alternative embodiments.
[0128] In various embodiments, Figure 1 、 2Elements of the various systems, hosts, peripherals, and other elements shown in and 4 may be implemented using suitable hardware (such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs)), using software, or using a combination of hardware and software elements.
[0129] In some embodiments, some or all of the functionality of certain system elements described herein may be implemented using one or more general-purpose processors that are programmed in software to perform the functions described herein. The software may be downloaded to the processor in electronic form, for example, over a network, or it may alternatively or additionally be provided and / or stored on non-transitory tangible media, such as magnetic, optical, or electronic memory.
[0130] It should be understood that the above embodiments are cited by way of example, and the present invention is not limited to what is specifically shown and described above. On the contrary, the scope of the present invention includes combinations and sub-combinations of the various features described above, as well as variations and modifications thereof that occur to those skilled in the art upon reading the above description and that are not disclosed in the prior art. The documents incorporated by reference in this patent application are considered to be an integral part of this application, except to the extent that any term is defined in these incorporated documents in a manner that conflicts with the explicit or implicit definition in this specification, only the definition in this specification should be considered.
Claims
1. A computing system for bridging memory address spaces and address translation, comprising: at least one peripheral bus; peripheral devices connected to the at least one peripheral bus; at least one memory; a first system component (i) associated with a first address space in the at least one memory and arranged to access the first address space using a first address translation, and (ii) connected to the peripheral device via the at least one peripheral bus; as well as a second system component that (i) is associated with a second address space in the at least one memory and is arranged to access the second address space using a second address translation, and (ii) is connected to the peripheral device via the at least one peripheral bus, Wherein, the first system component is arranged to cause the peripheral device to access the second address space associated with the second system component based on an association between at least a portion of the first address translation and at least a portion of the second address translation. 2 . The system of claim 1 , wherein one or both of the first system component and the second system component are physical processors. 3 . The system of claim 1 , wherein one or both of the first system component and the second system component are virtual processors.
4. The system of claim 1, wherein the first system component and the second system component run on different servers.
5. The system of claim 1, wherein the first system component and the second system component run on the same server.
6. The system according to claim 1, wherein the first system component is a first virtual processor, the first virtual processor being arranged to access the first address space using the first address translation; wherein the second system component is a second virtual processor, the second virtual processor being arranged to access the second address space using the second address translation; and includes virtualization software configured to define an association between at least a portion of the first address translation and at least a portion of the second address translation, and to enable the peripheral device to access the second address space using the association.
7. The system of claim 6, wherein the peripheral device comprises a network adapter, and wherein: The virtualization software is configured to define an association between a first remote direct memory access, RDMA, memory key, MKEY and a second remote direct memory access, RDMA, memory key, MKEY corresponding to the first and second address translations, respectively. 8 . The system of claim 6 , wherein the first system component is configured to enable the peripheral device to transfer data to the second address space through remote direct memory access (RDMA) using a queue pair (QP) controlled by the first system component.
9. The system of claim 8, wherein the QP and the transfer of data via RDMA are associated with a network association of the second system component.
10. The system of claim 6, wherein the first virtual processor is arranged to subject the peripheral device to accessing the second address space subject to a security policy governing access to address spaces.
11. The system of claim 6, wherein the peripheral device comprises a network adapter configured to save a local copy of the association between the at least a portion of the first address translation and the at least a portion of the second address translation and to access the second address space using the local copy of the association.
12. The system of claim 1 , wherein the second system component is a virtual processor, and wherein the first system component is configured to run virtualization software that (i) allocates physical resources to the virtual processor and (ii) accesses the second address space of the virtual processor.
13. The system of claim 1 , wherein the second system component is configured to specify to the peripheral device whether the first system component is permitted to access the second address space, and wherein The peripheral device is configured to access the second address space only upon verifying that the first system component is permitted to access the second address space.
14. The system of claim 1 , wherein: The at least one peripheral bus includes a first peripheral bus connected to the first system component, and a second peripheral bus connected to the second system component; and The at least one memory includes a first memory having the first address space and associated with the first system component, and a second memory having the second address space and associated with the second system component.
15. The system of claim 1, wherein the second system component runs in a server integrated into a network adapter.
16. The system of claim 1 , wherein the peripheral device comprises a network adapter arranged to transmit data packets to a network, wherein the first system component and the second system component have respective first and second network addresses, and wherein, The second system component is arranged to (i) construct a data packet, including obtaining portions of the data packet from the first address space of the first system component, and (ii) transmit the data packet to the network via the network adapter using the network association of the second system component.
17. The system of claim 1, wherein the peripheral device comprises a storage device.
18. The system of claim 1, wherein the peripheral device comprises a processor that emulates a storage device by exposing a Non-Volatile Memory Express (NVMe) device on the peripheral bus.
19. The system of claim 18, wherein the peripheral device that emulates the storage device is a network adapter.
20. The system of claim 18, wherein the processor in the peripheral device is configured to complete input / output (I / O) transactions of the first system component by accessing the second address space.
21. The system of claim 20, wherein the processor in the peripheral device is configured to directly access the second address space using Remote Direct Memory Access (RDMA).
22. The system of claim 20, wherein the processor in the peripheral device is configured to use a network association of the second system component to complete the I / O transaction.
23. The system of claim 1, wherein the peripheral device comprises a network adapter, and wherein: The first system component and the second system component run on different physical hosts serviced by the network adapter.
24. The system of claim 1, wherein the peripheral device is a network adapter configured to emulate a graphics processing unit (GPU) by exposing GPU functionality over the peripheral bus.
25. The system of claim 24, wherein the second system component is configured to perform network operations on behalf of the emulated GPU.
26. A computational method for bridging memory address spaces and address translation, comprising: communicating between a peripheral device and a first system component via at least one peripheral bus, wherein the first system component (i) is associated with a first address space in at least one memory and is arranged to access the first address space using a first address translation, and (ii) is connected to the peripheral device via the at least one peripheral bus; communicating between the peripheral device and a second system component via the at least one peripheral bus, wherein the second system component (i) is associated with a second address space in the at least one memory and is arranged to access the second address space using a second address translation, and (ii) is connected to the peripheral device via the at least one peripheral bus; as well as Using the first system component, the peripheral device is caused to access the second address space associated with the second system component based on an association between at least a portion of the first address translation and at least a portion of the second address translation.
27. The method of claim 26, wherein one or both of the first system component and the second system component are physical processors.
28. The method of claim 26, wherein one or both of the first system component and the second system component are virtual processors.
29. The method of claim 26, wherein the first system component and the second system component run on different servers.
30. The method of claim 26, wherein the first system component and the second system component run on the same server.
31. The method according to claim 26, wherein the first system component is a first virtual processor, the first virtual processor being arranged to access the first address space using the first address translation; wherein the second system component is a second virtual processor, the second virtual processor being arranged to access the second address space using the second address translation; and includes: using virtualization software; defining an association between at least a portion of the first address translation and at least a portion of the second address translation; and enabling the peripheral device to access the second address space using the association.
32. The method of claim 31 , wherein the peripheral device comprises a network adapter, and wherein: Defining the association includes defining an association between a first remote direct memory access, RDMA, memory key, MKEY and a second remote direct memory access, RDMA, memory key, MKEY corresponding to the first address translation and the second address translation, respectively.
33. The method of claim 31 , wherein enabling the peripheral device to access the second address space comprises: The peripheral device is enabled to use a queue pair QP controlled by the first system component to transmit data to the second address space through remote direct memory access RDMA.
34. The method of claim 33, wherein the QP and transfer of data via RDMA are associated with a network association of the second system component.
35. The method of claim 31 , wherein causing the peripheral device to access the second address space comprises: Access to the second address space by the peripheral device is subject to a security policy governing access to address spaces.
36. The method of claim 31 , wherein the peripheral device comprises a network adapter, and wherein: Enabling the peripheral device to access the second address space includes: saving a local copy of the association between the at least a portion of the first address translation and the at least a portion of the second address translation in the network adapter; and accessing the second address space using the local copy of the association.
37. The method of claim 26, wherein the second system component is a virtual processor, and wherein: Enabling the peripheral device to access the second address space includes running virtualization software in the first system component, the virtualization software (i) allocating physical resources to the virtual processor and (ii) accessing the second address space of the virtual processor.
38. The method of claim 26, wherein enabling the peripheral device to access the second address space comprises: specifying, by the second system component, to the peripheral device whether to permit the first system component to access the second address space; And the second address space is accessed by the peripheral device only upon verification that the first system component is permitted to access the second address space.
39. The method of claim 26, wherein: The at least one peripheral bus includes a first peripheral bus connected to the first system component, and a second peripheral bus connected to the second system component; and The at least one memory includes a first memory having the first address space and associated with the first system component, and a second memory having the second address space and associated with the second system component.
40. The method of claim 26, wherein the second system component runs in a server integrated into a network adapter.
41. A method according to claim 26, wherein the peripheral device includes a network adapter, the network adapter is arranged to transmit a data packet to a network, wherein the first system component and the second system component have corresponding first network addresses and second network addresses, and the method includes performing, by the second system component: (i) constructing a data packet, including obtaining multiple parts of the data packet from the first address space of the first system component; and (ii) transmitting the data packet to the network via the network adapter using the network association of the second system component.
42. The method of claim 26, wherein the peripheral device comprises a storage device.
43. The method of claim 26, wherein the peripheral device comprises a processor that emulates a storage device by exposing a Non-Volatile Memory Express (NVMe) device on the peripheral bus.
44. The method of claim 43, wherein the peripheral device that emulates the storage device is a network adapter.
45. The method of claim 43, wherein causing the peripheral device to access the second address space comprises: Using the processor in the peripheral device, an input / output (I / O) transaction of the first system component is completed by accessing the second address space.
46. The method of claim 45, wherein accessing the second address space comprises: The second address space is directly accessed by the processor in the peripheral device using Remote Direct Memory Access (RDMA).
47. The method of claim 45, wherein completing the I / O transaction is performed by the processor in the peripheral device using the network association of the second system component.
48. The method of claim 26, wherein the peripheral device comprises a network adapter, and wherein: The first system component and the second system component run on different physical hosts serviced by the network adapter.
49. The method of claim 26, wherein the peripheral device is a network adapter, and the method comprises: By exposing a graphics processing unit (GPU) functionality on the peripheral bus, a GPU is emulated by the network adapter.
50. The method of claim 49, comprising performing, by the second system component, network operations on behalf of the emulated GPU.
Citation Information
Patent Citations
Sharing a network interface card among multiple hosts
US7245627B2
Accessing remote storage devices using a local bus protocol
US9696942B2
Network interface controller with flexible memory handling
US20130067193A1
Sharing address translation between CPU and peripheral devices
US20140122828A1
Techniques for routing packets between virtual machines
US20170054659A1