A data transmission method and apparatus
By deploying daemons and hardware accelerators in NFV systems, using the predefined data structure of virtual input/output rings and bidirectional zero-copy technology, the problem of low packet processing efficiency after NFV hardware is solved, and efficient packet transmission is achieved.
Patent Information
- Application Number
- CN202111358155.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2016-06-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2036-06-15
AI Technical Summary
After NFV hardware is generalized, the processing capacity of general hardware devices is insufficient, resulting in a longer processing time of data packets and insufficient throughput, and hardware acceleration devices are required to accelerate the processing of data packets.
By deploying daemons in the host, using hardware accelerators and virtual accelerators, the information required to perform acceleration operations is transmitted in the virtual input/output rings using predefined data structures, realizing bidirectional zero-copy technology and improving the transmission efficiency of data packets.
It improves the transmission efficiency of data packets, enhances the throughput capability of virtual input/output rings, reduces the number of context switching between kernel state and user state, and reduces CPU consumption.
Smart Images

Figure CN114217902B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a data transmission method and device. Background Art
[0002] Network Function Virtualization (NFV) achieves hardware generalization in the form of a combination of "software + general hardware", so that network device functions no longer rely on dedicated hardware. Hardware resources can be fully and flexibly shared to achieve rapid development and deployment of new services, and perform automatic deployment, elastic scaling, fault isolation and self-healing based on actual business needs.
[0003] However, after the NFV hardware is universalized, the processing capacity of universal hardware devices is insufficient, resulting in extended processing time of data packets and insufficient throughput. Therefore, it is necessary to introduce hardware acceleration devices to accelerate data packet processing. It can be seen that how to enhance the throughput capacity of data packets and improve the transmission efficiency of data packets is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The embodiments of the present invention provide a data transmission method and device, which can improve the transmission efficiency of data packets.
[0005] A first aspect of an embodiment of the present invention discloses a data transmission method, which is applied to a daemon process in a host machine, wherein a virtual machine is deployed on the host machine, and a hardware accelerator and at least one virtual accelerator configured for the virtual machine are also deployed in the host machine, wherein the method comprises:
[0006] Acquire information required for performing an acceleration operation in a virtual input / output ring of a target virtual accelerator, wherein the information required for performing an acceleration operation adopts a predefined data structure, and the data structure occupies a unit item of the virtual input / output ring of the target virtual accelerator;
[0007] Determining information that can be recognized by the hardware accelerator according to the information required for performing the acceleration operation;
[0008] The information that can be recognized by the hardware accelerator is sent to the hardware accelerator, and the hardware accelerator is used to obtain the data to be accelerated according to the information that can be recognized by the hardware accelerator, and perform an acceleration operation on the data to be accelerated.
[0009] The information required for performing the acceleration operation includes the virtual machine physical address of the data to be accelerated, the length of the data to be accelerated, the virtual machine physical address storing the acceleration result, and the acceleration type parameter.
[0010] Among them, the hardware accelerator can be a multi-queue hardware accelerator (i.e., supporting multiple virtual functions VFs, each VF is equivalent to a hardware accelerator), or the hardware accelerator can be a single-queue hardware accelerator.
[0011] It can be seen that the information required for performing the acceleration operation adopts a predefined data structure, and this data structure occupies a unit item of the virtual input / output ring of the target virtual accelerator, which can greatly enhance the throughput capacity of the virtual input / output ring. More data packets can be transmitted per unit time, thereby improving the transmission efficiency of the data packets.
[0012] In addition, the daemon process in the host runs in the vhost-user user mode and directly accesses the hardware accelerator in the user mode without passing through the kernel protocol stack, so as to minimize the number of context switches between the kernel mode and the user mode and reduce the switching overhead.
[0013] In a possible implementation manner, the determining the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation includes:
[0014] Determining the host physical address of the data to be accelerated according to the virtual machine physical address of the data to be accelerated in the information required for performing the acceleration operation and the preset mapping relationship between the virtual machine physical address and the host physical address;
[0015] The sending the information recognizable by the hardware accelerator to the hardware accelerator includes:
[0016] Sending the host physical address of the data to be accelerated to the hardware accelerator;
[0017] Among them, the hardware accelerator is used to obtain the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
[0018] In this possible implementation manner, since the mapping relationship between the virtual machine physical address and the host physical address is established in advance and the hardware accelerator can recognize the host physical address, the virtual machine does not need to copy the data to be accelerated in the virtual memory to the memory buffer, thus realizing zero-copy in the packet sending direction. Zero-copy in the packet sending direction means that the virtual machine does not need to copy the data to be accelerated in the virtual machine memory to the memory buffer during the process of sending the information required for performing the acceleration operation to the hardware accelerator.
[0019] In a possible implementation manner, the hardware accelerator supports multiple virtual functions VFs. After determining the host physical address of the data to be accelerated, the method further includes:
[0020] Query the target VF bound to the target virtual accelerator from the preset binding relationship between the virtual accelerator and the VF;
[0021] The sending the host physical address of the data to be accelerated to the hardware accelerator includes:
[0022] Send the host physical address of the data to be accelerated to the target VF, and the target VF is used to obtain the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
[0023] Among them, a binding relationship between the virtual accelerator and the VF is established in advance, so that each virtual accelerator does not interfere with each other and has the optimal performance.
[0024] In a possible implementation manner, the method further includes:
[0025] After the data to be accelerated is accelerated and processed, add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator in the binding relationship.
[0026] Among them, for the case where the hardware accelerator supports multiple VFs, the identifier of the target virtual accelerator can be obtained from the binding relationship.
[0027] In a possible implementation manner, before obtaining the information required for the acceleration operation in the virtual input / output ring of the target virtual accelerator, the method further includes:
[0028] If the hardware accelerator supports multiple virtual functions VFs, select an unused target VF from the multiple VFs, and establish a binding relationship between the target virtual accelerator and the target VF.
[0029] It can be seen that for the case of a multi-queue hardware accelerator, the above-mentioned transceiver two-way zero-copy scheme (that is, zero-copy in the packet sending direction and zero-copy in the packet receiving direction) is adopted, and the entire acceleration data stream has zero-copy in the full path, with almost no additional CPU consumption. The solution of the present invention retains the advantages of the para-virtualization technology, such as the migratability of the VM and the portability of the VNFC code (that is, hardware insensitivity).
[0030] In a possible implementation manner, before sending the host physical address of the data to be accelerated to the hardware accelerator, the method further includes:
[0031] Record the identifier of the target virtual accelerator.
[0032] Among them, the identifier of the target virtual accelerator can be recorded at a specified position in the memory buffer.
[0033] In a possible implementation, the method further includes:
[0034] After the data to be accelerated is accelerated and processed, according to the identifier of the target virtual accelerator recorded, add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator.
[0035] It can be seen that for the case of a single-queue hardware accelerator, the above-mentioned two-way zero-copy scheme for sending and receiving (i.e., zero-copy in the packet sending direction and zero-copy in the packet receiving direction) is adopted, and the entire accelerated data stream has zero-copy throughout the path, with almost no additional CPU consumption. This makes the solution of the present invention retain the advantages of the para-virtualization technology, such as the migratability of VMs and the portability of VNFC code (i.e., hardware unawareness).
[0036] In a possible implementation, the determining the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation includes:
[0037] According to the physical machine address of the virtual machine of the data to be accelerated in the information required for performing the acceleration operation, and the preset mapping relationship between the physical machine address of the virtual machine and the virtual address of the host, determine the virtual address of the host of the data to be accelerated in the virtual machine memory;
[0038] According to the virtual address of the host of the data to be accelerated in the virtual machine memory, copy the data to be accelerated into the memory buffer;
[0039] According to the virtual address of the host of the data to be accelerated in the memory buffer and the preset mapping relationship between the virtual address of the host and the physical address of the host, determine the physical address of the host of the data to be accelerated in the memory buffer;
[0040] The sending the information recognizable by the hardware accelerator to the hardware accelerator includes:
[0041] Send the physical address of the host of the data to be accelerated to the hardware accelerator.
[0042] Wherein, the hardware accelerator is used to obtain the data to be accelerated from the memory buffer according to the physical address of the host of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
[0043] Among them, in this possible implementation, since only the mapping relationship between the virtual machine physical address and the host virtual address is established in advance, and the mapping relationship between the virtual machine physical address and the host physical address is not established, and the hardware accelerator can only recognize the host physical address, it is necessary to convert the host virtual address of the data to be accelerated into the host physical address of the data to be accelerated that the hardware accelerator can recognize. In this way, the virtual machine needs to copy the data to be accelerated in the virtual memory into the memory buffer.
[0044] In a possible implementation, the method further includes:
[0045] After the data to be accelerated is accelerated, copy the generated acceleration result into the virtual machine memory;
[0046] According to the identifier of the target virtual accelerator, add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator.
[0047] Among them, after the data to be accelerated is accelerated, the daemon process in the host also needs to copy the generated acceleration result from the memory buffer to the virtual machine memory, so that the virtual machine can obtain the acceleration result from the virtual machine memory subsequently.
[0048] In a possible implementation, after adding the identifier of the unit item to the virtual input / output ring of the target virtual accelerator, the method further includes:
[0049] Send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator, so as to trigger the virtual machine to query the identifier of the unit item in the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is accelerated.
[0050] Specifically, the daemon process in the host can use the eventfd mechanism of the Linux kernel to send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator.
[0051] A second aspect of the embodiments of the present invention discloses a data transmission method, which is applied to a daemon process in a host. The host deploys a virtual machine, and the host also deploys a hardware accelerator and at least one virtual accelerator configured for the virtual machine. The method includes:
[0052] After the data to be accelerated is accelerated, obtain the identifier of the target virtual accelerator;
[0053] According to the identifier of the target virtual accelerator, add the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, so that the virtual machine can obtain the acceleration result generated after the data to be accelerated is processed by acceleration according to the identifiers of at least one of the unit items, where the unit item stores information required for performing an acceleration operation on the data to be accelerated.
[0054] In a possible implementation, the hardware accelerator supports multiple virtual functions VFs. Before obtaining the identifier of the target virtual accelerator after the data to be accelerated is processed by acceleration, the method further includes:
[0055] Select an unused target VF from the multiple VFs, and establish a binding relationship between the target virtual accelerator and the target VF;
[0056] Wherein, the identifier of the target virtual accelerator is obtained from the binding relationship.
[0057] In a possible implementation, before obtaining the identifier of the target virtual accelerator after the data to be accelerated is processed by acceleration, the method further includes:
[0058] Record the identifier of the target virtual accelerator.
[0059] In a possible implementation, the manner of adding the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, so that the virtual machine can obtain the acceleration result generated after the data to be accelerated is processed by acceleration according to the identifiers of at least one of the unit items is specifically as follows:
[0060] According to the identifier of the target virtual accelerator, add the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, and send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator, so that the virtual machine can respond to the interrupt request, query the identifiers of at least one of the unit items in the virtual input / output ring of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is processed by acceleration; or,
[0061] According to the identifier of the target virtual accelerator, add the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, so that the virtual machine can automatically monitor the identifiers of at least one of the unit items in the virtual input / output ring of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is processed by acceleration according to the identifiers of at least one of the unit items.
[0062] In the third aspect of the embodiments of the present invention, a data transmission device is disclosed, which is applied to a daemon process in a host computer. A virtual machine is deployed on the host computer, and a hardware accelerator and at least one virtual accelerator configured for the virtual machine are also deployed in the host computer. The data transmission device includes functional units for performing some or all of the steps of any method in the first aspect of the embodiments of the present invention. Among them, when the data transmission device performs some or all of the steps of any method in the first aspect, the transmission efficiency of data packets can be improved.
[0063] In the fourth aspect of the embodiments of the present invention, a data transmission device is disclosed, which is applied to a daemon process in a host computer. A virtual machine is deployed on the host computer, and a hardware accelerator and at least one virtual accelerator configured for the virtual machine are also deployed in the host computer. The data transmission device includes functional units for performing some or all of the steps of any method in the second aspect of the embodiments of the present invention. Among them, when the data transmission device performs some or all of the steps of any method in the second aspect, the consumption of CPU resources can be reduced.
[0064] In the fifth aspect of the embodiments of the present invention, a host computer is disclosed. The host computer includes: a processor, a hardware accelerator, and a memory. The memory is configured to store instructions, and the processor is configured to run the instructions. The processor runs the instructions to perform some or all of the steps of any method in the first aspect of the embodiments of the present invention. Among them, when the host computer performs some or all of the steps of any method in the first aspect, the transmission efficiency of data packets can be improved.
[0065] In the sixth aspect of the embodiments of the present invention, a host computer is disclosed. The host computer includes: a processor, a hardware accelerator, and a memory. The memory is configured to store instructions, and the processor is configured to run the instructions. The processor runs the instructions to perform some or all of the steps of any method in the second aspect of the embodiments of the present invention. Among them, when the host computer performs some or all of the steps of any method in the second aspect, the transmission efficiency of data packets can be improved.
[0066] In the seventh aspect of the embodiments of the present invention, a computer storage medium is disclosed. The computer storage medium stores a program, and the program specifically includes instructions for performing some or all of the steps of any method in the first aspect of the embodiments of the present invention.
[0067] In the eighth aspect of the embodiments of the present invention, a computer storage medium is disclosed. The computer storage medium stores a program, and the program specifically includes instructions for performing some or all of the steps of any method in the second aspect of the embodiments of the present invention.
[0068] In some feasible embodiments, after adding the identifier of the unit item to the virtual input / output ring of the target virtual accelerator, the virtual machine can also monitor the virtual input / output ring of the target virtual accelerator to obtain the acceleration result generated after the data to be accelerated is processed. Specifically, the virtual machine can adopt a polling method to monitor the vring_used table of the virtual input / output ring of the target virtual accelerator in real time. If the vring_used table is updated, the identifier of the unit item in the vring_used table is obtained, where the unit item has been accelerated. Further, the virtual machine can obtain the virtual machine physical address storing the acceleration result from the vring_desc table of the virtual input / output ring according to the identifier of the unit item. Furthermore, the virtual machine can obtain the acceleration result from the virtual machine memory corresponding to the virtual machine physical address storing the acceleration result.
[0069] This method can completely avoid generating interrupts, prevent the virtual machine from executing interrupt handling operations and interrupting the service flow processing operations, and can improve the performance of the VNF.
[0070] In some feasible embodiments, the data structure used for the virtual machine to issue the information required for the acceleration operation occupies at least one unit item of the virtual input / output ring of the target virtual accelerator, and the zero-copy scheme is adopted for both the sending and receiving directions. In this way, although it does not enhance the throughput capacity of the virtual input / output ring, no memory copy is required for both the sending and receiving directions, and the consumption of CPU resources can be reduced compared with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0072] Figure 1.1 It is a schematic diagram of the network architecture of an NFV system disclosed in an embodiment of the present invention;
[0073] Figure 1.2 It is a schematic diagram of the internal acceleration system architecture of a VM disclosed in an embodiment of the present invention;
[0074] Figure 1.3 It is a schematic diagram of the internal acceleration system architecture of a Hypervisor disclosed in an embodiment of the present invention;
[0075] Figure 1.4 It is a schematic diagram of the network architecture of another NFV system disclosed in an embodiment of the present invention;
[0076] Figure 2 It is a schematic flowchart of a data transmission method disclosed in an embodiment of the present invention;
[0077] Figure 2.1 It is the storage form of information required for performing an acceleration operation in the Vring in an embodiment of the present invention;
[0078] Figure 2.2 It is another storage form of information required for performing an acceleration operation in the Vring in an embodiment of the present invention;
[0079] Figure 3 It is a schematic flowchart of another data transmission method disclosed in an embodiment of the present invention;
[0080] Figure 4 It is a schematic flowchart of another data transmission method disclosed in an embodiment of the present invention;
[0081] Figure 5 It is a schematic structural diagram of a data transmission device disclosed in an embodiment of the present invention;
[0082] Figure 6 It is a schematic structural diagram of another data transmission device disclosed in an embodiment of the present invention;
[0083] Figure 7 It is a schematic structural diagram of a data transmission device disclosed in an embodiment of the present invention;
[0084] Figure 8 It is a schematic structural diagram of a host computer in an embodiment of the present invention;
[0085] Figure 9 It is a schematic structural diagram of another host computer disclosed in an embodiment of the present invention. Detailed implementation manners
[0086] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0087] In the description, claims and the above-mentioned drawings of the present invention, the terms "first", "second", "third", etc. are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units is not limited to the listed steps or units, but optionally further comprises steps or units not listed, or optionally further comprises other steps or units inherent to these processes, methods, products or devices.
[0088] Reference to "an embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0089] Embodiments of the present invention disclose a data transmission method and apparatus, which can improve the transmission efficiency of data packets. The following will be described in detail respectively.
[0090] To better understand a data transmission method disclosed in embodiments of the present invention, the network architecture applicable to embodiments of the present invention will be described below.
[0091] Please refer to Figure 1.1 , Figure 1.1 which is a schematic diagram of the network architecture of an NFV system disclosed in embodiments of the present invention, where NFV (Network Function Virtualization). As Figure 1.1 shown, the NFV system includes: VNF (Virtualized Network Function), hypervisor, and general hardware.
[0092] Among them, the VNF may include at least one VM (Virtual Machine). Please refer to Figure 1.2 , Figure 1.2 which is a schematic diagram of the internal acceleration system architecture of a VM disclosed in embodiments of the present invention. As Figure 1.2As shown in the figure, the VM internal acceleration system includes a virtual network function application layer, an accelerator abstraction layer, an accelerator core layer, and a software input / output interface layer. Among them, the virtual network function application layer (Virtualization Network Function Applications) is for VMs that need to use virtual accelerators. The accelerator abstraction layer (Abstract Acceleration Layer, AAL) is mainly used to provide an interface layer for different virtual accelerators with a common API (Application Programming Interface). The accelerator core layer (Acceleration Core, AC), such as: DPDK (Data Plane Development Kit), ODP (Open Data Plane), or other high-performance acceleration software implementation frameworks. AC includes the accelerator core layer specified interfaces (AC specific APIs, s-API) and the front-end drivers configured for each virtual accelerator. Among them, the virtual accelerator is a virtualized form of the physical hardware accelerator, which can be understood as a virtual hardware accelerator. The virtual machine can only see the virtualized virtual accelerator. The virtual accelerator is similar to a virtual network card and a virtual disk. For example, the virtual accelerator includes: VirtIO-crypto, VirtIO-ipsec, VirtIO-compression. The software input / output interface layer (software I / O interface, sio) includes multiple virtual accelerator data packet headers, such as: VirtIO-crypto header, VirtIO-ipsec header. Among them, sio is the interface for data transmission between the front-end drivers of various virtual accelerators and the back-end device (Hypervisor), such as: VirtIO (virtual I / O) interface.
[0093] Among them, the Hypervisor can simulate at least one virtual accelerator for each virtual machine. Please also refer to Figure 1.3 , Figure 1.3 which is a schematic diagram of the internal acceleration system architecture of a Hypervisor disclosed in an embodiment of the present invention. As Figure 1.3As shown in the figure, the Hypervisor internal acceleration system includes a VirtIO backend accelerator, a VirtIO-based user space interface, an accelerator core layer, and a physical accelerator layer. Among them, the VirtIO backend accelerator (backend device) is a virtual accelerator simulated by a virtual machine emulator (Qemu). The VirtIO-based user space interface (User space based VirtIO interface, vHost-user) is a technology that adopts the vhost-user user space solution, realizes packet forwarding based on the VirtIO interface through the high-speed data transmission channel between the virtual machine and the host. The accelerator core layer includes the specified interface of the accelerator core layer, hardware resource management, packet management, and a general driver layer. Among them, the hardware resource management (hardware mgnt) supports ordinary hardware accelerators (referred to as single-queue hardware accelerators in the present invention) and hardware accelerators that support the SR-IOV function (referred to as multi-queue hardware accelerators in the present invention). For the multi-queue hardware accelerator, a mechanism for queue allocation and recycling is provided. The hardware accelerator that supports the SR-IOV function can virtualize multiple VFs (Virtual Function, virtual functions), and each VF corresponds to a hardware accelerator. The packet management (packet mgnt) relies on the memory management mechanism of the AC and provides an identification function for accelerated data packets for the single-queue accelerator. The physical accelerator layer (Acceleration) includes various accelerator software, hardware, and CPU instruction acceleration resources, etc.
[0094] Among them, general hardware may include, but is not limited to, a CPU (Central Processing Unit, central processing unit) that can provide special instructions, an SoC chip (System-on-a-Chip), and other hardware devices that can provide acceleration functions, such as: a GPU (Graphics Processing Unit, graphics processing unit), an FPGA (Field-Programmable Gate Array, field programmable gate array), etc. Among them, acceleration means offloading some functions in the program to the hardware for execution to achieve the effect of shortening the program execution time.
[0095] Please refer to Figure 1.4 , Figure 1.4 is a schematic diagram of the network architecture of an NFV system disclosed in an embodiment of the present invention. Among them, Figure 1.4 is Figure 1.1 a specific form of Figure 1.4As shown, each VM is configured with at least one virtual accelerator, and each virtual accelerator has a corresponding virtual accelerator front-end driver and a Vring (Virtual I / O ring). The virtual accelerator front-end driver is located on the virtual machine side, and the virtual accelerator is located on the Hypervisor side. In Figure 1.4 In the NFV system shown, the information required for the data to be accelerated sent by the VNF of the VM passes through the virtual accelerator front-end driver. In the virtual accelerator front-end driver, the information required for the data to be accelerated is assembled into the information required for performing the acceleration operation. The information required for performing the acceleration operation passes through the Vring of the virtual accelerator and, after being parsed and processed by the daemon process, is finally transmitted to the hardware accelerator. The hardware accelerator performs acceleration processing on the data to be accelerated. Among them, the data to be accelerated may include, but is not limited to, network packets, storage packets, files to be encrypted, etc. The data to be accelerated is stored in memory (host memory or virtual machine memory). The daemon process polls and monitors the destination packets generated after the data to be accelerated is accelerated. After the acceleration processing is completed, the daemon process puts the identifier of the destination packet into the usage table of the Vring of the virtual accelerator and, after being parsed and processed by the virtual accelerator front-end driver, is finally transmitted to the VNF of the VM. It should be noted that the information required for performing the acceleration operation does not include the data to be accelerated that needs to be accelerated, but only includes the information required for performing the acceleration operation, such as packet length, encryption / decryption algorithm type, key length, length of the data to be accelerated, and related GPA (Guest Physical Address, virtual machine physical address, also known as the client physical address), such as the GPA of the data to be accelerated, the GPA for storing the acceleration result, etc.
[0096] Optionally, if the hardware accelerator supports multiple VFs, when creating a virtual accelerator for the VM, immediately bind an unused VF to the virtual accelerator on the Hypervisor side. In this way, it is possible to notify the virtual machine corresponding to the virtual accelerator of the data packet received from the specified VF without the need for memory copying. Optionally, if the hardware accelerator is an ordinary single-queue hardware accelerator, when sending data packets on the Hypervisor side, use a specified position (such as a space offset by 4 bytes) of the mbuf (Memory Buffer) structure to record the identifier of the virtual accelerator of the virtual machine. After the hardware accelerator completes the acceleration processing of the data to be accelerated, the Hypervisor side retrieves the identifier of the virtual accelerator recorded in the mbuf structure, and can notify the virtual machine corresponding to the virtual accelerator without the need for memory copying.
[0097] Please refer to Figure 2 , Figure 2It is a schematic flowchart of a data transmission method disclosed in an embodiment of the present invention. Among them, this method is written from multiple aspects such as virtual machines, daemon processes in the host, and hardware accelerators, and this method adopts the bidirectional zero-copy technology. As Figure 2 shown, this method may include the following steps.
[0098] 201. If the hardware accelerator supports multiple virtual functions VF, the daemon process in the host selects an unused target VF from multiple VFs and establishes a binding relationship between the target virtual accelerator and the target VF.
[0099] In the embodiment of the present invention, virtual machines, hardware accelerators, and at least one virtual accelerator configured for the virtual machines are deployed on the host. Among them, the hardware accelerator can be a single-queue hardware accelerator or a multi-queue hardware accelerator (that is, the hardware accelerator supports multiple virtual functions VF, and each VF corresponds to a hardware accelerator).
[0100] If the hardware accelerator supports multiple VFs, the daemon process in the host can select an unused target VF from multiple VFs during the initialization phase and establish a binding relationship between the target virtual accelerator and the target VF. Among them, the target VF can be any one of the unused VFs.
[0101] It should be noted that in the embodiment of the present invention, multiple binding relationships can be established. After establishing multiple binding relationships, each virtual accelerator does not interfere with each other and has the optimal performance. The target virtual accelerator
[0102] As an optional implementation manner, before step 201, this method further includes the following steps:
[0103] 11) Start the daemon process and the full polling service on the host.
[0104] 12) The daemon process on the host creates a memory buffer pool.
[0105] 13) Configure at least one virtual accelerator for the virtual machine.
[0106] 14) The virtual machine simulator (Qemu) creates a virtual machine.
[0107] 15) Qemu sends the physical memory address layout of the virtual machine to the daemon process on the host through the vhost-user protocol.
[0108] 16) The daemon process in the host establishes a mapping relationship between the GPA and the host physical address HPA, and establishes a mapping relationship between the GPA and the host virtual address HVA according to the physical memory address layout of the virtual machine.
[0109] In this alternative embodiment, the daemon process in the host machine also serves as a service of the host machine and starts automatically after the host machine boots up. After the daemon process in the host machine starts, it immediately scans the hardware accelerator, initializes the hardware accelerator management structure, creates a memory buffer pool, and registers the vhost-user server listening port. The administrator configures at least one virtual accelerator (which can be called the target virtual accelerator) for the virtual machine and starts the QEMU process to create the virtual machine. The QEMU process sends a Unix message to the daemon process through the vhost-user protocol; after receiving the message sent by the QEMU process, the daemon process in the host machine creates a server-side virtual accelerator. The QEMU process sends the physical memory address layout of the virtual machine to the daemon process in the host machine through the vhost-user protocol. After receiving the message, the daemon process in the host machine establishes the mapping relationship between GPA and the host physical address HPA and the mapping relationship between GPA and the host virtual address HVA according to the physical memory address layout of the virtual machine.
[0110] 202. The virtual machine adds the information required for performing the acceleration operation to a predefined data structure.
[0111] Among them, the information required for performing the acceleration operation may include, but is not limited to, the virtual machine physical address of the data to be accelerated, the length of the data to be accelerated, the virtual machine physical address for storing the acceleration result, and the acceleration type parameter (such as: algorithm type, operation type), etc.
[0112] Common acceleration functions include encryption / decryption, compression / decompression, audio / video encoding / decoding, packet processing, etc. Taking encryption / decryption as an example, if a symmetric encryption operation of associated data is to be performed, the information to be passed to the hardware accelerator (i.e., the information required for performing the acceleration operation) includes: initialization vector (iv addr), associated data (auth_data addr), source data to be encrypted (src_data addr), memory for storing the encryption result (dst_data addr), memory for storing the digest result (digest_result addr), and memory for storing the success flag of encryption / decryption and verification (inhdr addr). The data to be accelerated may include, but is not limited to, network packets, storage packets, files to be encrypted, etc., and the data to be accelerated are all stored in memory (host memory or VM memory).
[0113] In the prior art, such an encryption / decryption data packet requires 6 entry items in the virtual input / output ring (Vring) of the virtual accelerator. Please refer to Figure 2.1 , Figure 2.1 which is the storage form of the information required for performing the acceleration operation in the Vring disclosed in the embodiments of the present invention. AsFigure 2.1 As shown, taking the information required to perform the acceleration operation, i.e., the encrypted and decrypted data packets, as an example, 6 data parameter information needs to be passed to the hardware accelerator. Among them, each data parameter information occupies one entry item of the Vring. Assuming the size of the Vring created for the encrypted and decrypted virtual accelerator is 4096, then the maximum number of encrypted and decrypted packets that can be stored at one time is 682 (4096 / 6).
[0114] In the embodiments of the present invention, in order to improve the throughput of data packets, a data structure is defined in advance for each piece of information required to perform the acceleration operation. This data structure includes the information required for the data to be accelerated, such as: the virtual machine physical address (Guest Physical Address, GPA) address and length information of the data to be accelerated, where GPA can also be referred to as the client physical address.
[0115] Taking the encrypted and decrypted virtual accelerator as an example, the following data structure is defined for symmetric encryption and decryption operations:
[0116]
[0117] The application program of the virtualized network function (VNF) only needs to pass the parameter information recorded in the struct virtio_crypto_sym_op_data to the front-end driver of the encrypted and decrypted virtual accelerator according to the above data structure. The front-end driver of the encrypted and decrypted virtual accelerator directly assigns the struct virtio_crypto_sym_op_data data structure, and then adds this data structure as a unit item (i.e., an entry item) to the Vring of the virtual accelerator. Therefore, for the encrypted and decrypted virtual accelerator, each symmetric encryption and decryption algorithm only needs to occupy one item in the Vring. Among them:
[0118] vring_desc.addr = the GPA address of the struct virtio_crypto_sym_op_data
[0119] vring_desc.len = sizeof(struct virtio_crypto_sym_op_data)
[0120] vring_desc.flag = ~NEXT
[0121] Please also participate Figure 2.2 , Figure 2.2 is another storage form of the information required to perform the acceleration operation in the Vring disclosed in the embodiments of the present invention. From Figure 2.2It can be seen that this data structure occupies one unit item (i.e., one entry) of the virtual input / output ring of the virtual accelerator; for a Vring with a size of 4096, 4096 encryption and decryption data packets can be stored at one time. Figure 2.2 The storage form of the source data in the Vring ring relative to Figure 2.1 the traditional scheme shown in
[0122] has a 5-fold performance improvement, effectively enhancing the throughput capacity of the data packets in the virtual input / output ring.
[0123] In the embodiments of the present invention, the target virtual accelerator is any one of the multiple virtual accelerators configured for the virtual machine, and for the convenience of subsequent description, it is referred to as the target virtual accelerator.
[0124] Specifically, the VNF of the virtual machine calls the virtual accelerator front-end driver API, and passes the information required for performing the acceleration operation (including the GPA of the data to be accelerated, the GPA for storing the acceleration result, etc.) to the front-end driver. The front-end driver applies for a memory buffer in the memory buffer pool and constructs the information required for performing the acceleration operation using this information. Finally, the front-end driver calls the VirtIO API and puts the GPA and length of the information required for performing the acceleration operation into the vring_desc table of the Vring.
[0125] In the embodiments of the present invention, there are a total of three tables in the Vring shared area: the vring_desc table, which is used to store the addresses of the IO requests generated by the virtual machine; the vring_avail table, which is used to indicate which items in the vring_desc are available; and the vring_used table, which is used to indicate which items in the vring_desc have been submitted to the hardware.
[0126] For the vring_desc table, it stores the address (such as the GPA address) of the IO request generated by the virtual machine in the virtual machine memory. Each row in this table contains four fields: Addr, len, flags, and next. As shown below:
[0127] struct vring_desc{
[0128] / *Address(guest-physical).* /
[0129] __virtio64 addr;
[0130] / *Length.* /
[0131] __virtio32 len;
[0132] / *The flags as indicated above.* /
[0133] __virtio16 flags;
[0134] / *We chain unused descriptors via this,too* /
[0135] __virtio16 next;
[0136] };
[0137] Among them, Addr stores the memory address of the IO request in the virtual machine, usually a GPA value; len represents the length of this IO request in memory; flags indicates whether the data in this row is readable, writable, and whether it is the last item of an IO request; each IO request may contain multiple rows in the vring_desc table, and the next field indicates where the next item of this IO request is located in which row. Through next, multiple rows stored by an IO request in vring_desc can be connected into a linked list. When flag =~ VRING_DESC_F_NEXT, it means the end of this linked list.
[0138] For the vring_avail table, it stores the head position of the linked list formed by each IO request in vring_desc. The data structure is as follows:
[0139] struct vring_avail{
[0140] __virtio16 flags;
[0141] __virtio16 idx;
[0142] __virtio16 ring[];
[0143] };
[0144] Among them, in the vring_desc table, the position of the head of the linked list connected by the next field in the vring_desc table is stored in the ring[] array; idx points to the next available free position in the ring array; flags is a flag field.
[0145] For the vring_used table, after the data to be accelerated is accelerated and processed, the daemon process on the host can update this data structure:
[0146] / *u32 is used here for ids for padding reasons.* /
[0147] struct vring_used_elem{
[0148] / *Index of start of used descriptor chain.* /
[0149] __virtio32 id;
[0150] / *Total length of the descriptor chain which was used(written to)* /
[0151] __virtio32 len;
[0152] };
[0153] struct vring_used{
[0154] __virtio16 flags;
[0155] __virtio16 idx;
[0156] struct vring_used_elem ring[];
[0157] };
[0158] Among them, the ring[] array in vring_uesd has two members: id and len. id represents the position of the head node of the linked list formed by the processed IO requests in the vring_desc table; len represents the length of the linked list. idx points to the next available position in the ring array; flags is a flag bit.
[0159] 204. The daemon process in the host obtains the information required for the execution acceleration operation in the virtual input / output ring of the target virtual accelerator.
[0160] In the embodiment of the present invention, the daemon process (Daemon process) in the host runs in the vhost-user user mode, without passing through the kernel protocol stack, and directly accesses the hardware accelerator in the user mode, minimizing the number of context switches between the kernel mode and the user mode and reducing the switching overhead. Among them, the daemon process in the host is located in the Hypervisor.
[0161] Optionally, the specific implementation of the daemon process in the host to obtain the information required for the execution acceleration operation in the virtual input / output ring of the target virtual accelerator may be as follows:
[0162] The daemon process in the host uses a full polling method to continuously monitor the virtual input / output ring of each virtual accelerator to obtain the information required for the execution acceleration operation in the virtual input / output ring of the target virtual accelerator.
[0163] Specifically, in this optional implementation manner, the daemon process in the host may use a full polling method to continuously monitor the vring_desc table in the virtual input / output ring of each virtual accelerator to obtain the information required for the execution acceleration operation in the vring_desc table in the virtual input / output ring. This method reduces the consumption of CPU resources.
[0164] Optionally, the specific implementation of the daemon process in the host to obtain the information required for the execution acceleration operation in the virtual input / output ring of the target virtual accelerator may be as follows:
[0165] The virtual machine sends an information acquisition notification carrying the identifier of the target virtual accelerator to the daemon process in the host, and the daemon process in the host obtains the information required for the execution acceleration operation from the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator.
[0166] Specifically, in this optional implementation manner, the virtual machine sends an information acquisition notification carrying the identifier of the target virtual accelerator to the daemon process in the host. After receiving the information acquisition notification, the daemon process in the host obtains the information required for the execution acceleration operation from the vring_desc table in the virtual input / output ring of the target virtual accelerator.
[0167] 205. The daemon process in the host determines the host physical address of the data to be accelerated according to the virtual machine physical address of the data to be accelerated in the information required for the execution acceleration operation and the preset mapping relationship between the virtual machine physical address and the host physical address.
[0168] In the embodiments of the present invention, the hardware accelerator can only recognize the host physical address (HPA), and the memory address passed from the virtual machine to the daemon process in the host is the virtual machine physical address (GPA) of the data to be accelerated, which cannot be directly passed to the hardware accelerator for use.
[0169] In the initialization phase, the daemon process in the host has pre-established the mapping relationship between GPA and HPA. After having this mapping relationship, when processing the data to be accelerated subsequently, memory copying (i.e., zero-copy) can be avoided, and only simple address conversion is required. For example, the data to be accelerated is initially stored in the VM memory corresponding to GPA. After the daemon process in the host has established the mapping relationship between GPA and HPA, the VM does not need to copy the data to be accelerated in the VM memory to the memory buffer corresponding to the Host Virtual Address (HVA). Instead, only by converting GPA to HPA, the data to be accelerated can be obtained from the memory buffer corresponding to HPA.
[0170] 206. The daemon process in the host queries the target VF bound to the target virtual accelerator from the preset binding relationship between the virtual accelerator and the VF.
[0171] In the embodiments of the present invention, a hardware accelerator supporting the SR-IOV function can virtualize multiple VFs (Virtual Functions), and each VF corresponds to a hardware accelerator. Among them, the SR-IOV (Single-Root I / O Virtualization) technology is a hardware-based virtualization solution that can improve performance and scalability. The SR-IOV standard allows efficient sharing of PCIe (Peripheral Component Interconnect Express) devices between virtual machines, and it is implemented in hardware, enabling I / O performance comparable to that of the host.
[0172] The two new function types in SR-IOV are: Physical Function (PF) is a PCI function used to support the SR-IOV function, as defined in the SR-IOV specification. The PF contains the SR-IOV function structure for managing the SR-IOV function. The PF is a full-function PCIe function and can be discovered, managed, and processed like any other PCIe device. The PF has fully configured resources that can be used to configure or control PCIe devices. A VF is a function associated with the physical function. A VF is a lightweight PCIe function that can share one or more physical resources with the physical function and other VFs associated with the same physical function. The VF is only allowed to have the configuration resources for its own behavior.
[0173] Each SR-IOV device can have one PF, and each PF can have up to 64,000 VFs associated with it. The PF can create VFs through registers, and these registers are designed with attributes dedicated for this purpose.
[0174] In the embodiment of the present invention, the binding relationship between the virtual accelerator and the VF has been established in the initialization stage. At this time, the daemon process in the host can query the target VF bound to the target virtual accelerator from the preset binding relationship between the virtual accelerator and the VF.
[0175] 207. The daemon process in the host sends the host physical address of the data to be accelerated to the target VF.
[0176] Specifically, the daemon process in the host can call the (Application Programming Interface, API) of the hardware accelerator to send the host physical address of the data to be accelerated to the target VF.
[0177] 208. The target VF obtains the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated, and performs an acceleration operation on the data to be accelerated.
[0178] In the embodiment of the present invention, after the target VF receives the host physical address of the data to be accelerated, the hardware accelerator driver will obtain the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address HPA of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
[0179] 209. When the data to be accelerated is accelerated and processed, the daemon process in the host adds the identifier of the unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator in the binding relationship.
[0180] In the embodiment of the present invention, since the zero-copy technology is adopted, when the VM issues the information required for the acceleration operation, it is not necessary to copy the data to be accelerated in the virtual machine memory corresponding to the GPA to the memory buffer corresponding to the HVA. Therefore, the data to be accelerated and the acceleration result generated after the data to be accelerated is accelerated and processed are both stored in the virtual machine memory corresponding to the GPA. The daemon process in the host can poll and monitor whether the target VF has completed the acceleration processing operation on the data to be accelerated. If it is monitored that the operation has been completed, the daemon process in the host can add the identifier of the unit item to the vring_used table of the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator in the binding relationship.
[0181] 210. The virtual machine monitors the virtual input / output ring of the target virtual accelerator to obtain the acceleration result generated after the data to be accelerated is accelerated and processed.
[0182] Specifically, in this optional implementation, the virtual machine can use a polling method to continuously monitor the vring_used table of the virtual input / output ring of the target virtual accelerator. If the vring_used table is updated, the identifier of the unit item in the vring_used table is obtained. Here, the unit item has been accelerated. Further, the virtual machine can obtain the virtual machine physical address storing the acceleration result from the vring_desc table of the virtual input / output ring according to the identifier of the unit item. Furthermore, the virtual machine can obtain the acceleration result from the virtual machine memory corresponding to the virtual machine physical address storing the acceleration result. Optionally, the virtual machine can also perform further processing on the acceleration result, such as sending the acceleration result to the network.
[0183] Among them, this method of the virtual machine automatically monitoring the virtual input / output ring of the target virtual accelerator is beneficial to reducing the interrupt overhead and saving resources.
[0184] As another optional implementation, step 210 can be not executed and replaced by the following steps:
[0185] 21) The daemon process in the host sends an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator;
[0186] 22) The virtual machine queries the identifier of the unit item in the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator;
[0187] 23) The virtual machine obtains the virtual machine physical address storing the acceleration result according to the identifier of the unit item, and obtains the acceleration result from the virtual machine memory corresponding to the virtual machine physical address storing the acceleration result.
[0188] In this optional implementation, after the daemon process in the host adds the identifier of the accelerated unit item to the virtual input / output ring of the target virtual accelerator, the daemon process in the host can use the eventfd mechanism of the Linux kernel to send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator. Among them, when the virtual machine is created, the Qemu process creates an eventfd (event handle) for the Vring ring of the virtual accelerator.
[0189] After receiving the interrupt request, the virtual machine can stop executing the current program, and instead, according to the identifier of the target virtual accelerator, query the identifier of the unit item in the vring_used table of the virtual input / output ring of the target virtual accelerator, and according to the identifier of the unit item, obtain the physical address of the virtual machine where the acceleration result is stored from the vring_desc table of the virtual input / output ring of the target virtual accelerator. Further, the virtual machine can obtain the acceleration result from the virtual machine memory corresponding to the physical address of the virtual machine where the acceleration result is stored.
[0190] It can be seen that in the embodiments of the present invention, a two-way zero-copy scheme for transceiver (i.e., zero-copy in the packet sending direction and zero-copy in the packet receiving direction) is adopted, and zero-copy is performed for the entire acceleration data flow path, with almost no additional CPU consumption. The solution of the present invention retains the advantages of the para-virtualization technology, such as the migratability of the VM and the portability of the VNFC code (i.e., hardware insensitivity). Among them, zero-copy in the packet sending direction means that when the virtual machine sends the information required for the acceleration operation to the hardware accelerator, it is not necessary to copy the data to be accelerated in the virtual machine memory to the memory buffer. Zero-copy in the packet receiving direction means that after the hardware accelerator finishes accelerating the data to be accelerated, the generated acceleration result does not need to be copied from the memory buffer to the virtual machine memory either. During the whole process, the data to be accelerated and the acceleration result are always in the virtual machine memory.
[0191] It should be noted that the above-mentioned daemon process in the host is applicable to various types of virtual machines.
[0192] In Figure 2 the described method flow, the information required for the acceleration operation sent by the virtual machine adopts a predefined data structure, and this data structure occupies a unit item of the virtual input / output ring of the target virtual accelerator, which can greatly enhance the throughput capacity of the virtual input / output ring. In unit time, more data packets are transmitted, thereby improving the transmission efficiency of the data packets.
[0193] Please refer to Figure 3 , Figure 3 which is a schematic flow diagram of another data transmission method disclosed in the embodiments of the present invention. Among them, this method is written from multiple sides such as the virtual machine, the daemon process in the host, and the hardware accelerator, and this method adopts a two-way copy technology. As Figure 3 shown, this method may include the following steps.
[0194] 301. The virtual machine adds the information required for the acceleration operation to a predefined data structure.
[0195] 302. The virtual machine puts the information required for the acceleration operation into the virtual input / output ring of the target virtual accelerator.
[0196] 303. The daemon process in the host obtains the information required for performing the acceleration operation in the virtual input / output ring of the target virtual accelerator.
[0197] 304. The daemon process in the host determines the host physical address of the data to be accelerated according to the virtual machine physical address of the data to be accelerated in the information required for performing the acceleration operation and the preset mapping relationship between the virtual machine physical address and the host physical address.
[0198] Among them, the above steps 301-304 can specifically refer to Figure 2 the relevant descriptions in. It should be noted that in the embodiments of the present invention, the hardware accelerator is a single-queue hardware accelerator.
[0199] 305. The daemon process in the host records the identifier of the target virtual accelerator.
[0200] Specifically, the daemon process in the host can record the identifier of the target virtual accelerator in the space with an offset of 4 bytes from the head of the structure in the memory buffer. Among them, the memory buffer is any one in the memory buffer pool applied for in the initialization stage.
[0201] 306. The daemon process in the host sends the host physical address of the data to be accelerated to the hardware accelerator.
[0202] 307. The hardware accelerator obtains the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated, and performs an acceleration operation on the data to be accelerated.
[0203] 308. After the data to be accelerated is accelerated, the daemon process in the host adds the identifier of the unit item to the virtual input / output ring of the target virtual accelerator according to the recorded identifier of the target virtual accelerator.
[0204] Specifically, after the data to be accelerated is accelerated by the hardware accelerator, the daemon process in the host can query the identifier of the target virtual accelerator recorded in the space with an offset of 4 bytes from the head of the structure in the memory buffer, and add the identifier of the accelerated unit item to the vring_used table of the virtual input / output ring of the target virtual accelerator.
[0205] 309. The daemon process in the host sends an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator.
[0206] Specifically, the daemon process in the host can use the eventfd mechanism of the Linux kernel to send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator. Among them, when the virtual machine is created, the Qemu process creates an eventfd (event handle) for the Vring ring of the virtual accelerator.
[0207] 310. The virtual machine queries the identifier of the unit item in the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator.
[0208] Among them, after the virtual machine receives the interrupt request, it can stop executing the current program and instead query the identifier of the unit item in the vring_used table of the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator.
[0209] 311. The virtual machine obtains the physical address of the virtual machine where the acceleration result is stored according to the identifier of the unit item, and obtains the acceleration result from the virtual machine memory corresponding to the physical address of the virtual machine where the acceleration result is stored.
[0210] Specifically, the virtual machine can obtain the physical address of the virtual machine where the acceleration result is stored from the vring_desc table of the virtual input / output ring of the target virtual accelerator according to the identifier of the unit item, and obtain the acceleration result from the virtual machine memory corresponding to the physical address of the virtual machine where the acceleration result is stored. Optionally, the virtual machine can also perform further processing on the acceleration result, such as sending the acceleration result to the network.
[0211] As another optional implementation manner, steps 309-311 may not be executed, but the following steps may be executed:
[0212] The virtual machine monitors the virtual input / output ring of the target virtual accelerator to obtain the acceleration result generated after the data to be accelerated is processed by acceleration.
[0213] Specifically, in this optional implementation manner, the virtual machine can use the polling method to monitor the vring_used table of the virtual input / output ring of the target virtual accelerator in real time. If the vring_used table is updated, the identifier of the unit item in the vring_used table is obtained. Among them, this unit item has been accelerated. Further, the virtual machine can obtain the physical address of the virtual machine where the acceleration result is stored from the vring_desc table of the virtual input / output ring according to the identifier of the unit item. Furthermore, the virtual machine can obtain the acceleration result from the virtual machine memory corresponding to the physical address of the virtual machine where the acceleration result is stored. Optionally, the virtual machine can also perform further processing on the acceleration result, such as sending the acceleration result to the network.
[0214] Among them, this method of automatically monitoring the virtual input / output ring of the target virtual accelerator by the virtual machine is beneficial to reducing the interruption overhead and saving resources.
[0215] In Figure 3 the described method flow, the information required for the virtual machine to issue an execution acceleration operation adopts a predefined data structure, and this data structure occupies a unit item of the virtual input / output ring of the target virtual accelerator, which can greatly enhance the throughput capacity of the virtual input / output ring. In unit time, more data packets are transmitted, thereby improving the transmission efficiency of the data packets. At the same time, the hardware accelerator is a single-queue hardware accelerator, and by adopting the method of recording the identifier of the target virtual accelerator, the entire acceleration data stream has zero copy along the full path, with almost no additional CPU consumption, and the solution of the present invention also retains the advantages of the para-virtualization technology, such as the migratability of the VM and the portability of the VNFC code (i.e., hardware insensitivity).
[0216] Please refer to Figure 4 , Figure 4 which is a schematic flow diagram of another data transmission method disclosed in an embodiment of the present invention. Among them, this method is described from multiple sides such as the virtual machine, the daemon process in the host, and the hardware accelerator, and this method adopts the two-way copy technology. As Figure 4 shown, this method may include the following steps.
[0217] 401. If the hardware accelerator supports multiple virtual functions VF, the daemon process in the host selects an unused target VF from the multiple VFs and establishes a binding relationship between the target virtual accelerator and the target VF.
[0218] 402. The virtual machine adds the information required for the execution acceleration operation to the predefined data structure.
[0219] 403. The virtual machine stores the data structure in the virtual input / output ring of the target virtual accelerator.
[0220] 404. The daemon process in the host obtains the information required for the execution acceleration operation in the virtual input / output ring of the target virtual accelerator.
[0221] 405. The daemon process in the host determines the host virtual address of the data to be accelerated in the virtual machine memory according to the physical address of the virtual machine of the data to be accelerated in the information required for the execution acceleration operation and the preset mapping relationship between the physical address of the virtual machine and the host virtual address.
[0222] In the embodiment of the present invention, the daemon process in the host can only recognize the HVA and cannot recognize the virtual machine physical address GPA. Therefore, in the initialization stage, a mapping relationship between the virtual machine physical address GPA and the host virtual address HVA is established in advance. After the daemon process in the host obtains the virtual machine physical address GPA of the data to be accelerated, it can determine the host virtual address of the data to be accelerated according to the set mapping relationship between the virtual machine physical address and the host virtual address.
[0223] It should be noted that in the embodiment of the present invention, in the initialization stage, no mapping relationship is established between the virtual machine physical address GPA and the host physical address HPA, and only the mapping relationship between GPA and HVA is established.
[0224] 406. The daemon process in the host copies the data to be accelerated into the memory buffer according to the host virtual address of the data to be accelerated in the virtual machine memory.
[0225] In the embodiment of the present invention, since no mapping relationship is established between the virtual machine physical address GPA and the host physical address HPA, the daemon process in the host needs to copy the data to be accelerated into the memory buffer according to the host virtual address of the data to be accelerated in the virtual machine memory, so as to facilitate subsequent address conversion.
[0226] 407. The daemon process in the host determines the host physical address of the data to be accelerated in the memory buffer according to the host virtual address of the data to be accelerated in the memory buffer and the preset mapping relationship between the host virtual address and the host physical address.
[0227] In the embodiment of the present invention, in the initialization stage, no mapping relationship is established between GPA and HPA, and the hardware accelerator can only recognize the host physical address HPA of the data to be accelerated. The daemon process in the host needs to convert the host virtual address HVA of the data to be accelerated in the memory buffer into the host physical address HPA of the data to be accelerated that the hardware accelerator can recognize.
[0228] Specifically, the daemon process in the host can query the HPA corresponding to the HVA of the data to be accelerated in the memory buffer according to the / proc / pid / pagemap file in the Linux system.
[0229] 408. The daemon process in the host sends the host physical address of the data to be accelerated to the target VF.
[0230] 409. The target VF obtains the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated, and performs an acceleration operation on the data to be accelerated.
[0231] 410. After the data to be accelerated is processed by acceleration, the daemon process in the host copies the generated acceleration result to the virtual machine memory.
[0232] Specifically, after the data to be accelerated is processed by acceleration, the daemon process in the host also needs to copy the generated acceleration result from the memory buffer to the virtual machine memory, so that the virtual machine can obtain the acceleration result from the virtual machine memory subsequently.
[0233] 411. The daemon process in the host adds the identifier of the unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator.
[0234] 412. The virtual machine monitors the virtual input / output ring of the target virtual accelerator to obtain the acceleration result generated after the data to be accelerated is processed by acceleration.
[0235] In Figure 4 the described method flow, the information required for performing the acceleration operation sent by the virtual machine adopts a predefined data structure, and this data structure occupies a unit item of the virtual input / output ring of the target virtual accelerator, which can greatly enhance the throughput capacity of the virtual input / output ring. More data packets are transmitted per unit time, thereby improving the transmission efficiency of the data packets.
[0236] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a data transmission device disclosed in an embodiment of the present invention. Among them, the data transmission device 500 can be used to execute Figure 2 all or part of the steps in the data transmission method disclosed in Figure 2 . Specifically, please refer to Figure 5 described, and details are not repeated here. As
[0237] shown, the data transmission device 500 may include:
[0238] An obtaining unit 501, configured to obtain the information required for performing the acceleration operation in the virtual input / output ring of the target virtual accelerator, where the information required for performing the acceleration operation adopts a predefined data structure, and the data structure occupies a unit item of the virtual input / output ring of the target virtual accelerator;
[0239] A determining unit 502, configured to determine the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation;
[0240] A sending unit 503, configured to send information recognizable by the hardware accelerator to the hardware accelerator, where the hardware accelerator is configured to obtain data to be accelerated according to the information recognizable by the hardware accelerator, and perform an acceleration operation on the data to be accelerated.
[0241] As an optional implementation manner, the manner in which the determining unit 502 determines the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation is specifically as follows:
[0242] According to the virtual machine physical address of the data to be accelerated in the information required for performing the acceleration operation, and a preset mapping relationship between the virtual machine physical address and the host physical address, determine the host physical address of the data to be accelerated;
[0243] The manner in which the sending unit 503 sends the information recognizable by the hardware accelerator to the hardware accelerator is specifically as follows:
[0244] Send the host physical address of the data to be accelerated to the hardware accelerator;
[0245] Wherein, the hardware accelerator is configured to obtain the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
[0246] As an optional implementation manner, the hardware accelerator supports multiple virtual functions VFs. Figure 5 The data transmission device 500 further includes:
[0247] A query unit 504, configured to query the target VF bound to the target virtual accelerator from a preset binding relationship between the virtual accelerator and the VF after the determining unit 502 determines the host physical address of the data to be accelerated;
[0248] The manner in which the sending unit 503 sends the host physical address of the data to be accelerated to the hardware accelerator is specifically as follows:
[0249] Send the host physical address of the data to be accelerated to the target VF, and the target VF is configured to obtain the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
[0250] As another optional implementation manner, Figure 5 The data transmission device 500 further includes:
[0251] A first adding unit 505, configured to, after the data to be accelerated is accelerated and processed, add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator in the binding relationship.
[0252] As another optional implementation manner, Figure 5 the data transmission device 500 shown further includes:
[0253] A establishing unit 506, configured to, before the obtaining unit 501 obtains the information required for performing an acceleration operation in the virtual input / output ring of the target virtual accelerator, if the hardware accelerator supports multiple virtual functions VFs, select an unused target VF from the multiple VFs, and establish a binding relationship between the target virtual accelerator and the target VF.
[0254] As an optional implementation manner, the sending unit 503 is further configured to send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator, so as to trigger the virtual machine to query the identifier of the unit item in the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is accelerated and processed.
[0255] Please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of another data transmission device disclosed in an embodiment of the present invention. Among them, the data transmission device 600 can be used to execute Figure 3 all or part of the steps in the disclosed data transmission method. Specifically, please refer to Figure 3 described, which will not be elaborated here. As shown in Figure 6 , the data transmission device 600 may include:
[0256] An obtaining unit 601, configured to obtain the information required for performing an acceleration operation in the virtual input / output ring of the target virtual accelerator. The information required for performing the acceleration operation adopts a predefined data structure, and the data structure occupies one unit item of the virtual input / output ring of the target virtual accelerator;
[0257] Among them, the information required for performing the acceleration operation includes the virtual machine physical address of the data to be accelerated, the length of the data to be accelerated, the virtual machine physical address for storing the acceleration result, and the acceleration type parameter.
[0258] A determining unit 602, configured to determine the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation;
[0259] A sending unit 603, configured to send information recognizable by the hardware accelerator to the hardware accelerator, where the hardware accelerator is configured to obtain data to be accelerated according to the information recognizable by the hardware accelerator, and perform an acceleration operation on the data to be accelerated.
[0260] As an optional implementation manner, Figure 6 the data transmission device 600 shown further includes:
[0261] A recording unit, configured to record an identifier of the target virtual accelerator before the sending unit 603 sends a host physical address of the data to be accelerated to the hardware accelerator.
[0262] As an optional implementation manner, Figure 6 the data transmission device 600 shown further includes:
[0263] A second adding unit 605, configured to, after the data to be accelerated is accelerated, add an identifier of the unit item to a virtual input / output ring of the target virtual accelerator according to the recorded identifier of the target virtual accelerator.
[0264] As an optional implementation manner, the sending unit 603 is further configured to send an interrupt request carrying the identifier of the target virtual accelerator to a virtual machine corresponding to the target virtual accelerator, so as to trigger the virtual machine to query the identifier of the unit item in the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, and obtain an acceleration result generated after the data to be accelerated is accelerated.
[0265] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of another data transmission device disclosed in an embodiment of the present invention. Among them, the data transmission device 700 can be used to execute Figure 2 or Figure 4 partial steps in the disclosed data transmission method. Specifically, please refer to Figure 2 or Figure 4 for description, and details are not described herein again. As Figure 7 shown, the data transmission device 700 may include:
[0266] An obtaining unit 701, configured to obtain information required for performing an acceleration operation in a virtual input / output ring of a target virtual accelerator, where the information required for performing the acceleration operation adopts a predefined data structure, and the data structure occupies a unit item of the virtual input / output ring of the target virtual accelerator;
[0267] Among them, the information required for performing the acceleration operation includes the virtual machine physical address of the data to be accelerated, the length of the data to be accelerated, the virtual machine physical address for storing the acceleration result, and the acceleration type parameter.
[0268] A determination unit 702, configured to determine the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation;
[0269] A sending unit 703, configured to send the information recognizable by the hardware accelerator to the hardware accelerator, where the hardware accelerator is configured to obtain the data to be accelerated according to the information recognizable by the hardware accelerator, and perform an acceleration operation on the data to be accelerated.
[0270] As an optional implementation manner, the determination unit 702 includes:
[0271] A determination subunit 7021, configured to determine the host virtual address of the data to be accelerated in the virtual machine memory according to the virtual machine physical address of the data to be accelerated in the information required for performing the acceleration operation and a preset mapping relationship between the virtual machine physical address and the host virtual address;
[0272] A copying subunit 7022, configured to copy the data to be accelerated to the memory buffer according to the host virtual address of the data to be accelerated in the virtual machine memory;
[0273] The determination subunit 7021 is further configured to determine the host physical address of the data to be accelerated in the memory buffer according to the host virtual address of the data to be accelerated in the memory buffer and a preset mapping relationship between the host virtual address and the host physical address;
[0274] The sending unit 703 is specifically configured to send the host physical address of the data to be accelerated to the hardware accelerator.
[0275] Among them, the hardware accelerator is configured to obtain the data to be accelerated from the memory buffer according to the host physical address of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
[0276] As an optional implementation manner, Figure 7 The data transmission device 700 shown further includes:
[0277] A copying unit 704, configured to copy the generated acceleration result to the virtual machine memory after the data to be accelerated is accelerated;
[0278] A third adding unit 705, configured to add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator.
[0279] As an alternative implementation, the sending unit 703 is further configured to send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator, so as to trigger the virtual machine to query the identifier of the unit item in the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is accelerated and processed.
[0280] In Figures 5 - 7 the described data transmission device, the information required for the virtual machine to issue an acceleration operation still uses a predefined data structure, which occupies a unit item of the virtual input / output ring of the target virtual accelerator. This can greatly enhance the throughput capacity of the virtual input / output ring. More data packets can be transmitted per unit time, thereby improving the transmission efficiency of data packets.
[0281] Please refer to Figure 8 , Figure 8 FIG. is a schematic structural diagram of a host computer disclosed in an embodiment of the present invention. Among them, the host computer 800 can be used to execute Figures 2 - 4 all or part of the steps in the disclosed data transmission method. For details, please refer to Figures 2 - 4 described, which will not be elaborated here. As Figure 8 shown, the host computer 800 may include: at least one processor 801, such as a CPU (Central Processing Unit), a hardware accelerator 802, and a memory 803. Among them, the processor 801, the hardware accelerator 802, and the memory 803 are respectively connected to a communication bus. The memory 803 can be a high-speed RAM memory or a non-volatile memory. Those skilled in the art can understand that Figure 8 the structure of the host computer 800 shown in Figure 8 does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than
[0282] shown, or combine some components, or different component arrangements.
[0283] Obtain the information required for performing an acceleration operation in the virtual input / output ring of the target virtual accelerator. The information required for performing the acceleration operation adopts a predefined data structure, and the data structure occupies a unit item of the virtual input / output ring of the target virtual accelerator;
[0284] Determine the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation;
[0285] Send the information recognizable by the hardware accelerator to the hardware accelerator, and the hardware accelerator is used to obtain the data to be accelerated according to the information recognizable by the hardware accelerator and perform an acceleration operation on the data to be accelerated.
[0286] Among them, the information required for performing the acceleration operation includes the virtual machine physical address of the data to be accelerated, the length of the data to be accelerated, the virtual machine physical address for storing the acceleration result, and the acceleration type parameter.
[0287] Optionally, the way for the processor 801 to determine the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation is specifically as follows:
[0288] Determine the host physical address of the data to be accelerated according to the virtual machine physical address of the data to be accelerated in the information required for performing the acceleration operation and the preset mapping relationship between the virtual machine physical address and the host physical address;
[0289] The way for the processor 801 to send the information recognizable by the hardware accelerator to the hardware accelerator is specifically as follows::
[0290] Send the host physical address of the data to be accelerated to the hardware accelerator;
[0291] Among them, the hardware accelerator is used to obtain the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated and perform an acceleration operation on the data to be accelerated.
[0292] Optionally, the hardware accelerator supports multiple virtual functions VFs. After the processor 801 determines the host physical address of the data to be accelerated, it can also call the program code stored in the memory 803 to perform the following operations:
[0293] Query the target VF bound to the target virtual accelerator from the preset binding relationship between the virtual accelerator and the VF;
[0294] The processor 801 sending the host physical address of the data to be accelerated to the hardware accelerator includes:
[0295] Send the host physical address of the data to be accelerated to the target VF. The target VF is used to obtain the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
[0296] Optionally, the processor 801 may also call the program code stored in the memory 803 to perform the following operations:
[0297] After the data to be accelerated is accelerated, according to the identifier of the target virtual accelerator in the binding relationship, add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator.
[0298] Optionally, before the processor 801 obtains the information required for the acceleration operation in the virtual input / output ring of the target virtual accelerator, it may also call the program code stored in the memory 803 to perform the following operations:
[0299] If the hardware accelerator supports multiple virtual functions VF, select an unused target VF from the multiple VFs, and establish a binding relationship between the target virtual accelerator and the target VF.
[0300] Optionally, before the processor 801 sends the host physical address of the data to be accelerated to the hardware accelerator, it may also call the program code stored in the memory 803 to perform the following operations:
[0301] Record the identifier of the target virtual accelerator.
[0302] Optionally, the processor 801 may also call the program code stored in the memory 803 to perform the following operations:
[0303] After the data to be accelerated is accelerated, according to the recorded identifier of the target virtual accelerator, add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator.
[0304] Optionally, the specific way for the processor 801 to determine the information that the hardware accelerator can recognize according to the information required for the acceleration operation is:
[0305] According to the virtual machine physical address of the data to be accelerated in the information required for the acceleration operation, and the preset mapping relationship between the virtual machine physical address and the host virtual address, determine the host virtual address of the data to be accelerated in the virtual machine memory;
[0306] Copy the data to be accelerated to the memory buffer according to the host virtual address of the data to be accelerated in the virtual machine memory;
[0307] Determine the host physical address of the data to be accelerated in the memory buffer according to the host virtual address of the data to be accelerated in the memory buffer and the preset mapping relationship between the host virtual address and the host physical address;
[0308] The processor 801 sending the information recognizable by the hardware accelerator to the hardware accelerator includes:
[0309] Send the host physical address of the data to be accelerated to the hardware accelerator.
[0310] Wherein, the hardware accelerator is used to obtain the data to be accelerated from the memory buffer according to the host physical address of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
[0311] Optionally, the processor 801 can also call the program code stored in the memory 803 to perform the following operations:
[0312] After the data to be accelerated is accelerated, copy the generated acceleration result to the virtual machine memory;
[0313] Add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator.
[0314] Optionally, after the processor 801 adds the identifier of the unit item to the virtual input / output ring of the target virtual accelerator, it can also call the program code stored in the memory 803 to perform the following operations:
[0315] Send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator, so as to trigger the virtual machine to query the identifier of the unit item in the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is accelerated.
[0316] In Figure 8 In the described host computer, the information required for the virtual machine to issue an acceleration operation still uses a predefined data structure, which occupies a unit item of the virtual input / output ring of the target virtual accelerator. This can greatly enhance the throughput capacity of the virtual input / output ring, and more data packets can be transmitted per unit time, thereby improving the transmission efficiency of the data packets.
[0317] Please refer toFigure 9 , Figure 9 is a schematic structural diagram of another host computer disclosed in an embodiment of the present invention. Among them, the host computer 900 can be used to execute Figures 2 - 4 all or part of the steps in the disclosed data transmission method. For details, please refer to Figures 2 - 4 as described, which will not be elaborated here. As Figure 9 shown, the host computer 900 may include: at least one processor 901, such as a CPU (Central Processing Unit), a hardware accelerator 902, and a memory 903. Among them, the processor 901, the hardware accelerator 902, and the memory 903 are respectively connected to a communication bus. The memory 903 can be a high-speed RAM memory or a non-volatile memory. Those skilled in the art can understand that Figure 9 the structure of the host computer 900 shown in Figure 9 does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than Figure 9 shown, or combine some components, or have different component arrangements.
[0318] Among them, the processor 901 is the control center of the host computer 900 and can be a Central Processing Unit (CPU). The processor 901 uses various interfaces and lines to connect all parts of the entire host computer 900. By running or executing software programs and / or modules stored in the memory 903, and calling program codes stored in the memory 903, it is used to perform the following operations:
[0319] After the data to be accelerated is accelerated and processed, obtain the identifier of the target virtual accelerator;
[0320] According to the identifier of the target virtual accelerator, add the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, so that the virtual machine can obtain the acceleration result generated after the data to be accelerated is accelerated and processed according to the identifiers of at least one of the unit items, where the unit item stores information required for performing an acceleration operation on the data to be accelerated.
[0321] Optionally, the hardware accelerator supports multiple virtual functions VF. Before the processor 901 obtains the identifier of the target virtual accelerator after the data to be accelerated is accelerated and processed, it can also call the program codes stored in the memory 903 to perform the following operations:
[0322] Select an unused target VF from the multiple VFs and establish a binding relationship between the target virtual accelerator and the target VF;
[0323] Among them, the identifier of the target virtual accelerator is obtained from the binding relationship.
[0324] Optionally, after the data to be accelerated is accelerated and processed, before the processor 901 obtains the identifier of the target virtual accelerator, the program code stored in the memory 903 can also be called to perform the following operations:
[0325] Record the identifier of the target virtual accelerator.
[0326] Optionally, according to the identifier of the target virtual accelerator, the processor 801 adds the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, so that the virtual machine can obtain the acceleration result generated after the data to be accelerated is accelerated and processed according to the identifiers of at least one unit item. The specific method is as follows:
[0327] According to the identifier of the target virtual accelerator, add the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, and send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator, so that the virtual machine can respond to the interrupt request, query the identifiers of at least one unit item in the virtual input / output ring of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is accelerated and processed; or,
[0328] According to the identifier of the target virtual accelerator, add the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, so that the virtual machine can automatically monitor the identifiers of at least one unit item in the virtual input / output ring of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is accelerated and processed according to the identifiers of at least one unit item.
[0329] In Figure 9 In the described host machine, the data to be accelerated is initially stored in the virtual machine memory, and the acceleration result generated after the data to be accelerated is accelerated and processed is also stored in the virtual machine memory. The entire process does not perform copying, reducing the occupation of CPU resources.
[0330] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and units involved are not necessarily essential to this application.
[0331] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0332] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above various methods. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0333] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A data transmission method, characterized in that, a virtual machine is deployed on a host computer, and a hardware accelerator and at least one virtual accelerator configured for the virtual machine are also deployed in the host computer. The method includes: obtaining information required for performing an acceleration operation in a virtual input / output ring of a target virtual accelerator, where the information required for performing the acceleration operation adopts a predefined data structure, and the data structure occupies a unit item of the virtual input / output ring of the target virtual accelerator; the information required for performing the acceleration operation includes the virtual machine physical address of data to be accelerated and the length of the data to be accelerated; determining information recognizable by the hardware accelerator according to the information required for performing the acceleration operation; sending the information recognizable by the hardware accelerator to the hardware accelerator, where the hardware accelerator is used to obtain the data to be accelerated according to the information recognizable by the hardware accelerator and perform an acceleration operation on the data to be accelerated.
2. The method according to claim 1, characterized in that, the information required for performing the acceleration operation further includes the virtual machine physical address for storing an acceleration result and an acceleration type parameter.
3. The method according to claim 1 or 2, characterized in that, the determining the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation includes: determining the host physical address of the data to be accelerated according to the virtual machine physical address of the data to be accelerated in the information required for performing the acceleration operation and a preset mapping relationship between the virtual machine physical address and the host physical address; the sending the information recognizable by the hardware accelerator to the hardware accelerator includes: sending the host physical address of the data to be accelerated to the hardware accelerator; wherein, the hardware accelerator is used to obtain the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated and perform an acceleration operation on the data to be accelerated.
4. The method according to claim 3, characterized in that, the hardware accelerator supports multiple virtual functions VF. After determining the host physical address of the data to be accelerated, the method further includes: querying a target VF bound to the target virtual accelerator from a preset binding relationship between the virtual accelerator and VF; the sending the host physical address of the data to be accelerated to the hardware accelerator includes: sending the host physical address of the data to be accelerated to the target VF, where the target VF is used to obtain the data to be accelerated from the virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated and perform an acceleration operation on the data to be accelerated.
5. The method according to claim 4, characterized in that, the method further includes: after the data to be accelerated is accelerated, adding an identifier of the unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator in the binding relationship.
6. The method according to claim 4 or 5, characterized in that, Before obtaining the information required for performing an acceleration operation in the virtual input / output ring of the target virtual accelerator, the method further includes: If the hardware accelerator supports multiple virtual functions (VFs), select an unused target VF from the multiple VFs, and establish a binding relationship between the target virtual accelerator and the target VF.
7. The method according to claim 3, wherein, Before sending the host physical address of the data to be accelerated to the hardware accelerator, the method further includes: Record the identifier of the target virtual accelerator.
8. The method according to claim 7, wherein, The method further includes: After the data to be accelerated is accelerated, according to the recorded identifier of the target virtual accelerator, add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator.
9. The method according to claim 1 or 2, wherein, Determining the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation includes: According to the virtual machine physical address of the data to be accelerated in the information required for performing the acceleration operation, and a preset mapping relationship between the virtual machine physical address and the host virtual address, determine the host virtual address of the data to be accelerated in the virtual machine memory; According to the host virtual address of the data to be accelerated in the virtual machine memory, copy the data to be accelerated to a memory buffer; According to the host virtual address of the data to be accelerated in the memory buffer and a preset mapping relationship between the host virtual address and the host physical address, determine the host physical address of the data to be accelerated in the memory buffer; Sending the information recognizable by the hardware accelerator to the hardware accelerator includes: Sending the host physical address of the data to be accelerated to the hardware accelerator; wherein, the hardware accelerator is configured to obtain the data to be accelerated from the memory buffer according to the host physical address of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
10. The method according to claim 9, wherein, The method further includes: After the data to be accelerated is accelerated, copy the generated acceleration result to the virtual machine memory; According to the identifier of the target virtual accelerator, add the identifier of the unit item to the virtual input / output ring of the target virtual accelerator.
11. The method according to claim 5, 8 or 10, wherein, After adding the identifier of the unit item to the virtual input / output ring of the target virtual accelerator, the method further includes: Send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator, so as to trigger the virtual machine to query the identifier of the unit item in the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is accelerated.
12. A data transmission method, wherein, A virtual machine is deployed on a host machine. A hardware accelerator and at least one virtual accelerator configured for the virtual machine are also deployed in the host machine. The method includes: After the data to be accelerated is accelerated and processed, obtain the identifier of the target virtual accelerator; According to the identifier of the target virtual accelerator, add the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, so that the virtual machine can obtain the acceleration result generated after the data to be accelerated is accelerated and processed according to the identifiers of at least one of the unit items. Wherein, the unit item stores information required for performing an acceleration operation on the data to be accelerated; the information required for performing the acceleration operation includes the virtual machine physical address of the data to be accelerated and the length of the data to be accelerated.
13. The method according to claim 12, wherein, the hardware accelerator supports multiple virtual functions VF. Before obtaining the identifier of the target virtual accelerator after the data to be accelerated is accelerated and processed, the method further includes: Select an unused target VF from the multiple VFs and establish a binding relationship between the target virtual accelerator and the target VF; wherein, the identifier of the target virtual accelerator is obtained from the binding relationship.
14. The method according to claim 12, wherein, before obtaining the identifier of the target virtual accelerator after the data to be accelerated is accelerated and processed, the method further includes: Record the identifier of the target virtual accelerator.
15. The method according to any one of claims 12 to 14, wherein, the manner of adding the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, so that the virtual machine can obtain the acceleration result generated after the data to be accelerated is accelerated and processed according to the identifiers of at least one of the unit items is specifically: According to the identifier of the target virtual accelerator, add the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, and send an interrupt request carrying the identifier of the target virtual accelerator to the virtual machine corresponding to the target virtual accelerator, so that the virtual machine can respond to the interrupt request, query the identifiers of at least one of the unit items in the virtual input / output ring of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is accelerated and processed; or, According to the identifier of the target virtual accelerator, add the identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator, so that the virtual machine can automatically monitor the identifiers of at least one of the unit items in the virtual input / output ring of the target virtual accelerator, and obtain the acceleration result generated after the data to be accelerated is accelerated and processed according to the identifiers of at least one of the unit items.
16. A data transmission device, wherein, includes: An acquisition unit, configured to acquire information required for performing an acceleration operation in a virtual input / output ring of a target virtual accelerator, where the information required for performing the acceleration operation adopts a predefined data structure, and the data structure occupies one cell item of the virtual input / output ring of the target virtual accelerator; the information required for performing the acceleration operation includes a virtual machine physical address of data to be accelerated and a length of the data to be accelerated. A determination unit, configured to determine information recognizable by a hardware accelerator according to the information required for performing the acceleration operation. A sending unit, configured to send the information recognizable by the hardware accelerator to the hardware accelerator, where the hardware accelerator is configured to acquire the data to be accelerated according to the information recognizable by the hardware accelerator and perform an acceleration operation on the data to be accelerated.
17. The apparatus according to claim 16, wherein, the information required for performing the acceleration operation further includes a virtual machine physical address for storing an acceleration result and an acceleration type parameter.
18. The apparatus according to claim 16 or 17, wherein, the manner in which the determination unit determines the information recognizable by the hardware accelerator according to the information required for performing the acceleration operation is specifically: determining a host physical address of the data to be accelerated according to the virtual machine physical address of the data to be accelerated in the information required for performing the acceleration operation and a preset mapping relationship between the virtual machine physical address and the host physical address; the manner in which the sending unit sends the information recognizable by the hardware accelerator to the hardware accelerator is specifically: sending the host physical address of the data to be accelerated to the hardware accelerator; wherein, the hardware accelerator is configured to acquire the data to be accelerated from a virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated and perform an acceleration operation on the data to be accelerated.
19. The apparatus according to claim 18, wherein, the hardware accelerator supports multiple virtual functions VFs, and the apparatus further includes: A query unit, configured to query a target VF bound to the target virtual accelerator from a preset binding relationship between the virtual accelerator and the VF after the determination unit determines the host physical address of the data to be accelerated; the manner in which the sending unit sends the host physical address of the data to be accelerated to the hardware accelerator is specifically: sending the host physical address of the data to be accelerated to the target VF, and the target VF is configured to acquire the data to be accelerated from a virtual machine memory corresponding to the virtual machine physical address of the data to be accelerated according to the host physical address of the data to be accelerated and perform an acceleration operation on the data to be accelerated.
20. The apparatus according to claim 19, wherein, the apparatus further includes: A first adding unit, configured to add an identifier of the cell item to a virtual input / output ring of the target virtual accelerator according to an identifier of the target virtual accelerator in the binding relationship after the data to be accelerated is accelerated.
21. The apparatus according to claim 19 or 20, wherein, The device further includes: A creation unit, configured to, before the acquisition unit acquires information required for performing an acceleration operation in a virtual input / output ring of a target virtual accelerator, if the hardware accelerator supports multiple virtual functions (VFs), select an unused target VF from the multiple VFs, and establish a binding relationship between the target virtual accelerator and the target VF.
22. The device according to claim 18, wherein, The device further includes: A recording unit, configured to record an identifier of the target virtual accelerator before the sending unit sends a host physical address of the data to be accelerated to the hardware accelerator.
23. The device according to claim 22, wherein, The device further includes: A second addition unit, configured to, after the data to be accelerated is accelerated, add an identifier of the unit item to a virtual input / output ring of the target virtual accelerator according to the recorded identifier of the target virtual accelerator.
24. The device according to claim 16 or 17, wherein, The determination unit includes: A determination subunit, configured to determine a host virtual address of the data to be accelerated in the virtual machine memory according to a virtual machine physical address of the data to be accelerated in the information required for performing the acceleration operation and a preset mapping relationship between the virtual machine physical address and the host virtual address; A copy subunit, configured to copy the data to be accelerated to a memory buffer according to the host virtual address of the data to be accelerated in the virtual machine memory; The determination subunit is further configured to determine a host physical address of the data to be accelerated in the memory buffer according to the host virtual address of the data to be accelerated in the memory buffer and a preset mapping relationship between the host virtual address and the host physical address; The sending unit is specifically configured to send the host physical address of the data to be accelerated to the hardware accelerator; wherein, the hardware accelerator is configured to acquire the data to be accelerated from the memory buffer according to the host physical address of the data to be accelerated, and perform an acceleration operation on the data to be accelerated.
25. The device according to claim 24, wherein, The device further includes: A copy unit, configured to copy a generated acceleration result to the virtual machine memory after the data to be accelerated is accelerated; A third addition unit, configured to add an identifier of the unit item to a virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator.
26. The device according to claim 20, 23, or 25, wherein, The sending unit is further configured to send an interrupt request carrying the identifier of the target virtual accelerator to a virtual machine corresponding to the target virtual accelerator, so as to trigger the virtual machine to query the identifier of the unit item in the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, and acquire the acceleration result generated after the data to be accelerated is accelerated.
27. A data transmission device, wherein, It includes: An obtaining unit, configured to obtain an identifier of a target virtual accelerator after the data to be accelerated is accelerated and processed; An adding unit, configured to add identifiers of at least one unit item to a virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, so that a virtual machine can obtain an acceleration result generated after the data to be accelerated is accelerated and processed according to the identifiers of at least one of the unit items, where the unit item stores information required for performing an acceleration operation on the data to be accelerated; the information required for performing the acceleration operation includes a virtual machine physical address of the data to be accelerated and a length of the data to be accelerated.
28. The apparatus according to claim 27, wherein, the hardware accelerator supports multiple virtual functions VF, and the apparatus further includes: A establishing unit, configured to select an unused target VF from the multiple VFs and establish a binding relationship between the target virtual accelerator and the target VF before the obtaining unit obtains the identifier of the target virtual accelerator after the data to be accelerated is accelerated and processed; wherein, the identifier of the target virtual accelerator is obtained from the binding relationship.
29. The apparatus according to claim 27, wherein, the apparatus further includes: A recording unit, configured to record an identifier of a target virtual accelerator before the obtaining unit obtains the identifier of the target virtual accelerator after the data to be accelerated is accelerated and processed.
30. The apparatus according to any one of claims 27 to 29, wherein, the manner in which the adding unit adds identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, so that a virtual machine can obtain an acceleration result generated after the data to be accelerated is accelerated and processed according to the identifiers of at least one of the unit items is specifically as follows: adding identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, and sending an interrupt request carrying the identifier of the target virtual accelerator to a virtual machine corresponding to the target virtual accelerator, so that the virtual machine responds to the interrupt request, queries the identifiers of at least one of the unit items in the virtual input / output ring of the target virtual accelerator, and obtains an acceleration result generated after the data to be accelerated is accelerated and processed; or, adding identifiers of at least one unit item to the virtual input / output ring of the target virtual accelerator according to the identifier of the target virtual accelerator, so that the virtual machine automatically monitors the identifiers of at least one of the unit items in the virtual input / output ring of the target virtual accelerator, and obtains an acceleration result generated after the data to be accelerated is accelerated and processed according to the identifiers of at least one of the unit items.
Citation Information
Patent Citations
Xen-based FPGA accelerator virtualization platform and application
CN105389199A