Data transmission system, method, device, storage medium and computer program product
By introducing virtual RDMA devices between the virtual machine and the host, the problems of RDMA connection scalability and resource sharing in the virtualized environment are solved, and efficient data transmission and system stability are achieved.
Patent Information
- Application Number
- CN202410132919.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-01
AI Technical Summary
The scalability and resource sharing problems of RDMA connections in a virtualized environment in the prior art lead to low network transmission efficiency and large host resource consumption.
Introduce virtual RDMA devices between virtual machines and hosts to achieve decoupling between front-end and back-end, manage memory and connection mapping relationships through virtual RDMA devices, reduce host overhead, and support connection and memory sharing.
It improves the connection scalability and memory scalability between the virtual machine and the host, reduces the processing burden of the host, and improves the flexibility and stability of the data transmission system.
Smart Images

Figure CN120407497A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and particularly to a data transmission system, method, device, storage medium, and computer program product. Background Art
[0002] As the scale of cloud computing scenarios becomes larger and larger, the communication load in the data center grows exponentially, and thus it is necessary to expand the network communication capacity to cope with increasingly high data transmission requests. Remote Direct Memory Access (RDMA), as a network transmission technology with high performance, is widely used in data centers. In order to maintain isolation, portability, etc. in cloud computing scenarios, RDMA virtualization is required, that is, applying the RDMA technology in virtual machines. Therefore, there is an urgent need to provide a data transmission system to implement RDMA virtualization and improve data transmission performance. Summary of the Invention
[0003] This application provides a data transmission system, method, device, storage medium, and computer program product for implementing RDMA data transmission.
[0004] In a first aspect, a data transmission system is provided. The data transmission system includes a virtual machine, a virtual RDMA device, and a host, and the virtual RDMA device runs between the virtual machine and the host.
[0005] Among them, the virtual machine is used to run applications; the virtual RDMA device is used to register the virtual memory and virtual RDMA connection corresponding to the virtual machine, configure the memory mapping relationship between the virtual memory and the physical memory, and the connection mapping relationship between the virtual RDMA connection and the physical RDMA connection, and send the memory mapping relationship and the connection mapping relationship to the host; the host is used to transmit the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship.
[0006] By adding a layer of virtual RDMA device between the virtual machine at the front end and the host at the back end, this data transmission system realizes the decoupling of the front-end virtual machine and the back-end host, enabling the virtual machine to interact with the host without a dedicated connection, improving the scalability of the connection, and making the data transmission system more modular and easy to manage. Moreover, the mapping overhead between the virtual and the physical is transferred from the host to this virtual RDMA device, reducing the host overhead.
[0007] In a possible implementation, the virtual machine is further configured to send an RDMA write request to the virtual RDMA device based on the running application. The RDMA write request includes a first source virtual memory, a first destination virtual memory, and a first destination communication address. The RDMA write request is used to write a first data payload in the first source physical memory corresponding to the first source virtual memory to the first destination physical memory corresponding to the first destination virtual memory. The virtual RDMA device is further configured to add the RDMA write request to a send queue. The host is further configured to access the RDMA write request in the send queue, determine a first virtual RDMA connection according to the first destination communication address, and determine a first physical RDMA connection corresponding to the first virtual RDMA connection according to a connection mapping relationship. The host is further configured to determine the first source physical memory corresponding to the first source virtual memory according to a memory mapping relationship, read the first data payload in the first source physical memory, and send a first data stream to a destination host corresponding to the first destination communication address based on the first physical RDMA connection. The first data stream includes the first destination virtual memory and the first data payload, and the first data stream is used for the destination host to write the first data payload to the first destination physical memory corresponding to the first destination virtual memory.
[0008] The send queue is maintained by the virtual RDMA device, and the host processes the RDMA write request of the virtual machine by accessing the send queue maintained by the virtual RDMA device, reducing the overhead of the host for maintaining the send queue.
[0009] In a possible implementation, the host includes a central processing unit (CPU) and a physical network card. The CPU is configured to obtain a first data stream based on the first data payload in the first source physical memory and send the first data stream to the physical network card. The physical network card is configured to send the first data stream to the destination host based on the first physical RDMA connection. By invoking the CPU, the first data payload can be encapsulated into the first data stream transmitted by the physical network card.
[0010] In a possible implementation, the host includes a field-programmable gate array (FPGA) and a physical network card. The FPGA is configured to obtain a first data stream based on the first data payload in the first source physical memory and send the first data stream to the physical network card. The physical network card is configured to send the first data stream to the destination host based on the first physical RDMA connection. By invoking the FPGA, the first data payload can be encapsulated into the first data stream transmitted by the physical network card without involving the CPU, reducing the burden on the CPU and improving the response speed and processing capacity of the data transmission system.
[0011] In a possible implementation, the host is further configured to switch the first data stream transmitted by the first physical RDMA connection to be transmitted by the second physical RDMA connection in the case of a failure or an anomaly of the first physical RDMA connection, where the second physical RDMA connection is a standby connection of the first physical RDMA connection. This method can ensure the continuity and reliability of services in the case of failures or anomalies.
[0012] In a possible implementation, the virtual machine is further configured to send an RDMA read request to the virtual RDMA device based on the running application, where the RDMA read request includes a second source virtual memory, a second destination virtual memory, and a second destination communication address, and the RDMA read request is used to write a second data payload in the second destination physical memory corresponding to the second destination virtual memory to the second source physical memory corresponding to the second source virtual memory; the virtual RDMA device is further configured to add the RDMA read request to a receiving queue; the host is further configured to access the RDMA read request in the receiving queue, determine a second virtual RDMA connection according to the second destination communication address, and determine a third physical RDMA connection corresponding to the second virtual RDMA connection according to a connection mapping relationship; read the second data payload based on the third physical RDMA connection; and determine the second source physical memory corresponding to the second source virtual memory according to a memory mapping relationship, and write the second data payload to the second source physical memory.
[0013] The virtual RDMA device maintains the receiving queue, and the host processes the RDMA read request of the virtual machine by accessing the receiving queue maintained by the virtual RDMA device, reducing the overhead of the host for maintaining the receiving queue.
[0014] In a possible implementation, the transmission of RDMA data is data transmission based on an unreliable data (UD) transmission protocol, the virtual RDMA connection is a queue pair (QP) connection, and one QP state is used to serve multiple QP connections; the host is configured to implement the transmission of RDMA data according to a memory mapping relationship, a connection mapping relationship, a QP state, and the UD transmission protocol. Using the UD transmission protocol can implement stateless and message-oriented UD transmission, and one-to-many transmission can be achieved by multiplexing the QP state, simplifying the QP states that need to be maintained inside the host and greatly improving the scalability.
[0015] In a possible implementation, the virtual RDMA connections in the connection mapping relationship include priorities; the host is further configured to determine the processing order of multiple virtual RDMA connections according to the priorities of the multiple virtual RDMA connections, and implement the transmission of RDMA data according to the processing order. Adopting a scheduling strategy based on priorities can ensure that requests with higher priorities are processed in a timely manner.
[0016] In a possible implementation, multiple virtual RDMA connections in the connection mapping relationship correspond to a shared physical RDMA connection. Thereby, connection sharing between virtual machines is achieved, improving the scalability of connections.
[0017] In a possible implementation, multiple virtual memories in the memory mapping relationship correspond to a shared physical memory. Thereby, connection sharing between virtual machines is achieved. Memory sharing between virtual machines is realized, improving the scalability of memory.
[0018] In a second aspect, a data transmission method is provided. The data transmission method is applied to a data transmission system. The data transmission system includes a virtual machine, a virtual RDMA device, and a host. The virtual RDMA device runs between the virtual machine and the host. The method includes: the virtual machine runs an application; the virtual RDMA device registers the virtual memory and virtual RDMA connection corresponding to the virtual machine, configures the memory mapping relationship between the virtual memory and the physical memory, and the connection mapping relationship between the virtual RDMA connection and the physical RDMA connection, and sends the memory mapping relationship and the connection mapping relationship to the host; the host transmits the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship.
[0019] In a possible implementation, the virtual machine sends an RDMA write request to the virtual RDMA device based on the running application. The RDMA write request includes a first source virtual memory, a first destination virtual memory, and a first destination communication address. The RDMA write request is used to write a first data payload in the first source physical memory corresponding to the first source virtual memory to the first destination physical memory corresponding to the first destination virtual memory; the virtual RDMA device adds the RDMA write request to the send queue. In this case, the manner in which the host transmits the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship may include that the host accesses the RDMA write request in the send queue, determines a first virtual RDMA connection according to the first destination communication address, and determines a first physical RDMA connection corresponding to the first virtual RDMA connection according to the connection mapping relationship; determines the first source physical memory corresponding to the first source virtual memory according to the memory mapping relationship, and reads the first data payload in the first source physical memory; sends a first data stream to the destination host corresponding to the first destination communication address based on the first physical RDMA connection. The first data stream includes the first destination virtual memory and the first data payload, and the first data stream is used for the destination host to write the first data payload to the first destination physical memory corresponding to the first destination virtual memory.
[0020] In a possible implementation, the host includes a CPU and a physical network card. The method for sending a first data stream to a destination host corresponding to a first destination communication address based on a first physical RDMA connection may include: the CPU obtaining the first data stream based on a first data payload in a first source physical memory and sending the first data stream to the physical network card; and the physical network card sending the first data stream to the destination host based on the first physical RDMA connection.
[0021] In a possible implementation, the host includes an FPGA and a physical network card. The method for sending a first data stream to a destination host corresponding to a first destination communication address based on a first physical RDMA connection may include: the FPGA obtaining the first data stream based on a first data payload in a first source physical memory and sending the first data stream to the physical network card; and the physical network card sending the first data stream to the destination host based on the first physical RDMA connection.
[0022] In a possible implementation, when a first physical RDMA connection fails or malfunctions, the host switches the first data stream transmitted by the first physical RDMA connection to be transmitted by a second physical RDMA connection, where the second physical RDMA connection is a standby connection of the first physical RDMA connection.
[0023] In a possible implementation, a virtual machine sends an RDMA read request to a virtual RDMA device based on a running application. The RDMA read request includes a second source virtual memory, a second destination virtual memory, and a second destination communication address, and is used to write a second data payload in a second destination physical memory corresponding to the second destination virtual memory to a second source physical memory corresponding to the second source virtual memory. The virtual RDMA device is further configured to add the RDMA read request to a receive queue. In this case, the method for the host to transmit the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship may include: accessing the RDMA read request in the receive queue, determining a second virtual RDMA connection according to the second destination communication address, and determining a third physical RDMA connection corresponding to the second virtual RDMA connection according to the connection mapping relationship; reading the second data payload based on the third physical RDMA connection; and determining a second source physical memory corresponding to the second source virtual memory according to the memory mapping relationship and writing the second data payload to the second source physical memory.
[0024] In a possible implementation, the transmission of RDMA data is data transmission based on the UD transmission protocol, the virtual RDMA connection is a QP connection, and one QP state is used to serve multiple QP connections. The method for the host to transmit the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship may include: the host implementing the transmission of RDMA data according to the memory mapping relationship, the connection mapping relationship, the QP state, and the UD transmission protocol.
[0025] In a possible implementation, the virtual RDMA connections in the connection mapping relationship include priorities; the host also determines the processing order of the multiple virtual RDMA connections according to the priorities of the multiple virtual RDMA connections, and implements the transmission of RDMA data according to the processing order.
[0026] In a possible implementation, multiple virtual RDMA connections in the connection mapping relationship correspond to a shared physical RDMA connection.
[0027] In a possible implementation, multiple virtual memories in the memory mapping relationship correspond to a shared physical memory.
[0028] In a third aspect, a network device is provided. The network device includes: a processor, the processor is coupled to a memory, and at least one program instruction or code is stored in the memory. The at least one program instruction or code is loaded and executed by the processor so that the network device implements the data transmission method described in the second aspect or any one of the second aspects above.
[0029] Optionally, the processor is one or more, and the memory is one or more.
[0030] Optionally, the memory may be integrated with the processor, or the memory is separately arranged from the processor.
[0031] In a specific implementation process, the memory may be a non-transitory memory, such as a read only memory (ROM). It may be integrated with the processor on the same chip, or may be separately arranged on different chips. The present application does not limit the type of the memory and the setting manner of the memory and the processor.
[0032] In a fourth aspect, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium. The instruction is loaded and executed by a processor so that a computer implements the method in the second aspect or any one of the possible implementations of the second aspect above.
[0033] In a fifth aspect, a computer program (product) is provided. The computer program (product) includes: computer program code. When the computer program code is run by a computer, the computer is caused to execute the method in the second aspect or any one of the possible implementations of the second aspect above.
[0034] In a sixth aspect, a chip is provided, including a processor for calling and running an instruction stored in a memory, so that a communication device installed with the chip executes the method in the second aspect or any one of the possible implementations of the second aspect above.
[0035] In a seventh aspect, another chip is provided, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute the code in the memory. When the code is executed, the processor is configured to execute the method in the second aspect or any possible implementation manner of the second aspect described above.
[0036] It should be understood that for the beneficial effects achieved by the technical solutions of the second aspect to the seventh aspect of this application and the corresponding possible implementation manners, reference can be made to the technical effects of the first aspect and its corresponding possible implementation manners described above, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of a virtualization architecture provided in the related art;
[0038] Figure 2 A schematic structural diagram of a data transmission system provided in an embodiment of this application;
[0039] Figure 3 A schematic structural diagram of another data transmission system provided in an embodiment of this application;
[0040] Figure 4 A flowchart of a data transmission method provided in an embodiment of this application;
[0041] Figure 5 A schematic structural diagram of a network device provided in an embodiment of this application;
[0042] Figure 6 A schematic structural diagram of another network device provided in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0044] With the development of communication technologies, the rise of cloud computing has led to the continuous evolution of communication technologies and network architectures. For example, the cloud-native architecture, in a service-intensive deployment manner, covers the hierarchical structure from the host, virtual machine to container. In the cloud-native architecture, the full-mesh connection method also results in a large number of end-to-end network connections. Cloud computing has been practically applied in fields such as artificial intelligence training, high-performance computing, and data storage. These application scenarios have put forward more stringent requirements for network infrastructure. Since the bottleneck of system performance has gradually shifted from the central part to the edge side, it is necessary to start from optimizing virtualization to improve data forwarding efficiency and network utilization.
[0045] In the related art, taking virtual input output (virtIO)-RDMA in the host para-virtualization solution as an example, virtIO-RDMA is based on the virtIO protocol and allows virtualization of RDMA connections. The virtual RDMA connections can be used by virtual machines (VMs). The virtIO protocol is a virtual device interface standard. Exemplarily, referring to Figure 1 the virtualization architecture in the related art shown in the figure, the guest operating system (OS) includes multiple virtual machines, and the host OS includes a network card supporting RDMA. Communication occurs between the virtIO-frontend (FE) of the guest OS and the backend (BE) of the host OS through virtual RDMA connections, where the number of virtual RDMA connections can reach more than 160. To implement virtualization of RDMA connections, memory mapping is required between the guest virtual memory and the host physical memory for the network card to perform direct memory access (DMA).
[0046] However, although virtio-RDMA can achieve virtualization at the RDMA connection level and the number of connections can reach more than 160, there are still scalability problems, especially in ultra-large-scale virtualization environments. And virtualization operations usually consume some host resources, thus introducing additional latency and overhead to the host system. In addition, virtio-RDMA does not support resource sharing. Each virtual machine has its own dedicated RDMA connection, which requires establishing and maintaining an independent connection state on the network card for each connection. As a result, a large number of connection states need to be cached inside the network card, which not only increases the complexity of hardware and software but also limits the concurrent transmission ability.
[0047] Embodiments of the present application provide a data transmission system, as Figure 2As shown in the figure, the data transmission system includes a virtual machine 201, a virtual RDMA device 202, and a host 203. The virtual RDMA device 202 runs between the virtual machine 201 and the host 203. Among them, the virtual machine 201 is used to run applications; the virtual RDMA device 202 is used to register the virtual memory and virtual RDMA connection corresponding to the virtual machine 201, configure the memory mapping relationship between the virtual memory and the physical memory, and the connection mapping relationship between the virtual RDMA connection and the physical RDMA connection, and send the memory mapping relationship and the connection mapping relationship to the host 203; the host 203 is used to transmit the RDMA data of the application running on the virtual machine 201 according to the memory mapping relationship and the connection mapping relationship.
[0048] A virtual RDMA device 202 is added between the virtual machine 201 at the front end and the host 203 at the back end of the data transmission system, realizing the decoupling of the front-end virtual machine 201 and the back-end host 203. There are the same virtual RDMA connections between different virtual machines 201, so that the virtual machine 201 does not need to interact with the host 203 through an exclusive connection, improving the scalability of the connection and making the data transmission system more modular and easy to manage. Moreover, the mapping overhead between the virtual and the physical is transferred from the host 203 to the virtual RDMA device 202, reducing the overhead of the host 203.
[0049] Therefore, through the decoupling of the front end and the back end, the virtual machine 201 and the host 203 can be independently optimized without affecting the whole body by pulling one hair, and the architecture implementation is relatively flexible. That is, it makes the data transmission system easier to expand and upgrade. For example, the front-end application can be independently upgraded, or the back-end transmission layer technology can be independently upgraded, without affecting the stable operation of the entire data transmission system. Moreover, since the virtual RDMA device 202 provides a unified and stable interface for the application, if the back-end transmission layer technology changes, the front-end application does not need to make any adjustments, realizing the complete insensitivity of the application to the transmission layer and improving the stability of the application operation.
[0050] In an RDMA-based data transmission scenario, an application typically submits a work request (WR) to a work queue (WQ). The work queue includes a send queue (SQ) and a receive queue (RQ). Each element of the work queue is called a work queue element (WQE). Since SQ and RQ are usually created in pairs, they are called a queue pair (QP). Hardware with an RDMA engine continuously fetches WRs from the WQ for execution. After execution, it places a work completion (WC) into the completion queue (CQ). Each element in the completion queue is called a completion queue element (CQE). The application fetches WCs from the CQ.
[0051] Taking the work request as an RDMA write request as an example. The virtual machine 201 is also used to send an RDMA write request to the virtual RDMA device 202 based on the running application. The RDMA write request includes a first source virtual memory, a first destination virtual memory, and a first destination communication address. The RDMA write request is used to write the first data payload in the first source physical memory corresponding to the first source virtual memory to the first destination physical memory corresponding to the first destination virtual memory. The virtual RDMA device 202 is also used to add the RDMA write request to the send queue. The host 203 is also used to access the RDMA write request in the send queue, determine a first virtual RDMA connection according to the first destination communication address, and determine the first physical RDMA connection corresponding to the first virtual RDMA connection according to the connection mapping relationship. Determine the first source physical memory corresponding to the first source virtual memory according to the memory mapping relationship, and read the first data payload in the first source physical memory. Send a first data stream to the destination host corresponding to the first destination communication address based on the first physical RDMA connection. The first data stream includes the first destination virtual memory and the first data payload, and the first data stream is used for the destination host to write the first data payload to the first destination physical memory corresponding to the first destination virtual memory.
[0052] The virtual RDMA device 202 maintains the send queue. The host 203 processes the RDMA write request of the virtual machine 201 by accessing the send queue maintained by the virtual RDMA device 202, reducing the overhead of the host 203 for maintaining the send queue.
[0053] In a possible implementation, the host 203 includes a CPU and a physical network card; wherein, the CPU is configured to obtain a first data stream based on a first data payload in a first source physical memory and send the first data stream to the physical network card; the physical network card is configured to send the first data stream to a destination host based on a first physical RDMA connection. By invoking the CPU, the first data payload can be encapsulated into the first data stream transmitted by the physical network card. Wherein, the physical network card may refer to any network card that supports RDMA.
[0054] Optionally, obtaining the first data stream based on the first data payload in the first source physical memory may include performing packet processing on the first data payload in the first source physical memory according to the RDMA packet format, and determining the first data payload after packet processing as the first data stream. Exemplarily, the RDMA packet format includes an Ethernet (Eth) header field, an internet protocol (IP) header field, a user datagram protocol (UDP) header field, an infiniband (IB) basic transport header (BTH) field, and an IB payload field, and the IB payload field is used to carry the first data payload.
[0055] In another possible implementation, the host 203 includes an FPGA and a physical network card; wherein, the FPGA is configured to obtain a first data stream based on a first data payload in a first source physical memory and send the first data stream to the physical network card; the physical network card is configured to send the first data stream to a destination host based on a first physical RDMA connection. By invoking the FPGA, the first data payload can be encapsulated into the first data stream transmitted by the physical network card without involving the CPU, reducing the burden on the CPU and improving the response speed and processing capacity of the data transmission system.
[0056] Optionally, the manner in which the FPGA obtains the first data stream is the same as the manner in which the above-mentioned CPU obtains the first data stream, and will not be elaborated here. Wherein, the physical RDMA connection is a connection pre-created by the hosts on both sides of the communication. After mapping the first virtual RDMA connection to the first physical RDMA connection, the two communication parties of the first physical RDMA connection are the local host and the destination host. Furthermore, the physical network card can send the first data stream to the destination host through the first physical RDMA connection.
[0057] In the embodiment of the present application, the host 203 is further configured to switch the first data stream transmitted by the first physical RDMA connection to the second physical RDMA connection for transmission when the first physical RDMA connection fails or is abnormal. The second physical RDMA connection is a standby connection of the first physical RDMA connection. Thereby, the continuity and reliability of the service can be ensured in the event of a failure or an abnormal situation.
[0058] The embodiment of the present application does not limit the manner in which the host 203 senses that the first physical RDMA connection fails or is abnormal. Optionally, the connection status of the first physical RDMA connection is sensed through a physical network card to determine whether the first physical RDMA connection fails or is abnormal. Alternatively, whether the first physical RDMA connection fails or is abnormal is determined by analyzing the latency or success rate of the data transmitted based on the first physical RDMA connection. Alternatively, whether the first physical RDMA connection fails or is abnormal is actively determined by sending a probe packet. For example, a probe packet is sent based on the first physical RDMA connection. If an acknowledgment packet feedback based on the probe packet is received, it is determined that the first physical RDMA connection does not fail or is abnormal; if an acknowledgment packet feedback based on the probe packet is not received, it is determined that the first physical RDMA connection fails or is abnormal.
[0059] Exemplarily, taking the host 203 including a CPU and a physical network card as an example, the physical network card is configured to sense that the first physical RDMA connection fails or is abnormal and send failure notification information of the first physical RDMA connection to the CPU; the CPU switches the first data stream transmitted by the first physical RDMA connection to the second physical RDMA connection for transmission based on the failure notification information.
[0060] In the embodiment of the present application, the virtual RDMA device 202 is further configured to configure the primary-standby mapping relationship between physical RDMA connections and send the primary-standby mapping relationship to the host 203. The host 203 is configured to determine that the standby connection of the first physical RDMA connection is the second physical RDMA connection according to the primary-standby mapping relationship.
[0061] Optionally, the host 203 is further configured to report an indication message indicating that the first physical RDMA connection fails or is abnormal to the virtual RDMA device 202 when the first physical RDMA connection fails or is abnormal; the virtual RDMA device 202 reconfigures the connection mapping relationship based on the indication information.
[0062] Taking the work request as an RDMA read request as an example, the virtual machine 201 is further configured to send an RDMA read request to the virtual RDMA device 202 based on the running application. The RDMA read request includes a second source virtual memory, a second destination virtual memory, and a second destination communication address. The RDMA read request is used to write the second data payload in the second destination physical memory corresponding to the second destination virtual memory into the second source physical memory corresponding to the second source virtual memory. The virtual RDMA device 202 is further configured to add the RDMA read request to the receive queue. The host 203 is further configured to access the RDMA read request in the receive queue, determine a second virtual RDMA connection according to the second destination communication address, and determine a third physical RDMA connection corresponding to the second virtual RDMA connection according to the connection mapping relationship. Read the second data payload based on the third physical RDMA connection. Determine the second source physical memory corresponding to the second source virtual memory according to the memory mapping relationship, and write the second data payload into the second source physical memory.
[0063] The receive queue is maintained by the virtual RDMA device 202. The host 203 processes the RDMA read request of the virtual machine 201 by accessing the receive queue maintained by the virtual RDMA device 202, reducing the overhead of the host 203 for maintaining the receive queue.
[0064] In a possible implementation manner, the transmission of RDMA data is data transmission based on the UD transmission protocol. The virtual RDMA connection is a QP connection, and one QP state is used to serve multiple QP connections. The host 203 is configured to implement the transmission of RDMA data according to the memory mapping relationship, the connection mapping relationship, the QP state, and the UD transmission protocol. The UD transmission protocol can implement stateless, message-oriented UD transmission. By multiplexing the QP state, one-to-many transmission can be achieved, simplifying the QP state that needs to be maintained inside the host 203 and greatly improving the scalability.
[0065] Thus, the data transmission system can provide a stateless design, such that regardless of the traffic pattern of the front-end data stream, the virtual RDMA device 202 performs transmission based on the stateless UD transport layer. For example, the traffic pattern of the data stream provided by the front-end virtual machine 201 is the RDMA semantics, but through the middle layer of the virtual RDMA device 202, finally at the transport layer of the host 203, it will be a completely stateless, message-oriented UD transmission.
[0066] Among them, although the data stream sent based on the UD transport protocol is stateless, some QP states still need to be maintained on the physical network card of host 203. The difference is that the QP states maintained on the physical network card can be reused. That is to say, a single QP state can serve multiple QP connections, realizing one-to-many communication. Therefore, this data transmission system realizes the combination of message-oriented and connection-oriented, that is, it takes into account both the message-oriented advantages of the UD transport protocol and the stability and reliability of the QP connection-oriented.
[0067] In a possible implementation manner, the virtual RDMA connections in the connection mapping relationship include priorities; host 203 is further configured to determine the processing order of multiple virtual RDMA connections according to the priorities of the multiple virtual RDMA connections, and implement the transmission of RDMA data according to the processing order. By adopting a scheduling strategy based on priorities, it can ensure that requests with higher priorities are processed in a timely manner, and further ensure that requests with higher priorities obtain more resources.
[0068] In the embodiments of the present application, multiple virtual RDMA connections in the connection mapping relationship correspond to a shared physical RDMA connection. The connection sharing between virtual machines 201 is realized, and the scalability of the connection is improved. Multiple virtual memories in the memory mapping relationship correspond to a shared physical memory. The memory sharing between virtual machines 201 is realized, and the scalability of the memory is improved.
[0069] Next, taking the number of virtual machines 201 as multiple as an example, in combination with Figure 3 the data transmission system provided by the embodiments of the present application will be illustrated by examples. As Figure 3 shown, the guest OS includes multiple virtual machines 201 and a virtual RDMA device 202, and the host OS includes a host 203. The virtual RDMA device 202 is located between the multiple virtual machines 201 and the host 203. Among them, the virtual RDMA device 202 includes a registration queue and a scheduling strategy, and the host 203 includes request processing, policy execution, failover, and a network card supporting RDMA.
[0070] This data transmission system implements a registration queue and a scheduling strategy on the control path. Among them, the registration queue includes, but is not limited to, registering the virtual memory and virtual RDMA connections corresponding to the virtual machine 201, and configuring the memory mapping relationship. For example, the registration queue is implemented by establishing a shared memory queue and application memory registration. Memory registration includes registering QP or CQ, doorbell, and memory area. When the QP needs to send data, the driver generates a WQE, and then notifies the network card to obtain the WQE for RDMA data transmission. This process of notifying the network card is called a doorbell.
[0071] The scheduling policies include, but are not limited to, the configuration of connection mapping relationships, the selection and configuration of transport layer technologies, the priority configuration of virtual RDMA connections, and the configuration of fault backup groups. Among them, the configuration of connection mapping relationships is used to map the virtual RDMA connections at the front end to physical RDMA connections, thereby determining the transmission path. The selection and configuration of transport layer technologies are used to decide whether to use hardware offloading or other transport layer technologies according to requirements. For example, the transport layer technology can adopt the UD transport layer technology, and the hardware offloading can be to offload the CPU overhead to the FPGA. The priority configuration of virtual RDMA connections is used to ensure that requests with higher priorities obtain more resources. For example, the priority can be to classify the quality of service (QoS) of virtual RDMA connections. The configuration of fault backup groups is used to set up a backup scheme to ensure the continuity of services in case of failures.
[0072] Request processing, policy execution, and failover are implemented on the data path. Request processing is used to process work requests in the work queue of the virtual RDMA device 202 according to the memory mapping relationship and the connection mapping relationship. Policy execution is used to execute the above scheduling policies. For example, scheduling is performed according to the priority of virtual RDMA connections to ensure that high-priority requests are processed in a timely manner. Failover is used to immediately switch to the backup scheme if any failures or anomalies are detected during the transmission process. The network card is used for data transmission at the transport layer. In the embodiment of the present application, when the virtual machine 201 interacts with the network card, the exchanged data is directly exchanged through DMA. For example, when submitting a WQE, the data transmission system sends the payload of the WQE to the CPU through DMA, and the CPU re-packages the payload and performs other processing, and finally sends the processed payload out through the network card.
[0073] Optionally, in the design of the transport layer, the user-space implementation can be optimized. By using FPGA technology, the CPU processing is replaced with an FPGA-based RDMA to implement a stateless data transmission method, ensuring that the data path does not involve the CPU and is exchanged only through a single peripheral component interconnect express (PCIE) path, thereby greatly improving the efficiency. In addition, the data transmission system also supports multi-path transmission, scattered transmission of data packets, and optimization of transmission reliability. By combining the characteristics of the network and services, through the hardware offloading function, a direct one-to-one mapping relationship between network packets and service operations is ensured.
[0074] Thus, by adding a layer of virtual RDMA device 202 between the front end and the back end, a software-hardware combination solution for decoupling the front-end virtual machine 201 and the back-end host 203 is implemented, aiming to build a smooth transmission mechanism between the front end and the back end to ensure efficient data transmission and interaction, while simplifying the development and maintenance work. In this data transmission system, the front end only needs to consider what kind of interfaces to provide for the application and which device types to virtualize, while the back end is completely focused on the optimization of the transport layer. The optimization of the transport layer can be carried out from the perspectives of scalability, load balancing, packet out-of-order, congestion control, support for more semantics, offloading primitives, providing corresponding interfaces, hardware offloading, and other transport layer technologies.
[0075] Exemplarily, in the data transmission system provided by the embodiments of the present application, the hardware components may include a network card and an FPGA, etc. The network card may be a network interface card (NIC) enabled with RDMA (RDMA-enabled). The RDMA-enabled NIC can support virtual RDMA technology, has a hardware acceleration function, and is suitable for optimization for the cloud computing environment. The FPGA is used to provide hardware acceleration for the transmission of RDMA data, ensure that the data only passes through the PCIE path once, and implement direct packetization.
[0076] The software components may include virtual RDMA management software, a transport layer optimization toolkit, and a network configuration management tool, etc. The virtual RDMA management software provides a virtualization interface for the front end and manages the communication with the back-end transport layer. For example, it can handle requests from the guest and perform tasks such as priority scheduling. The transport layer optimization toolkit can provide functions such as load balancing, congestion control, and out-of-order processing. The network configuration management tool provides configuration, management, and failover functions for hardware offloading and multi-path optimization.
[0077] This data transmission system provides a more efficient data processing speed through a software-hardware combined RDMA large-scale architecture. Combining the flexibility of software and the high-speed processing ability of hardware, it overcomes the limitations of relying solely on software or hardware. Through the function of hardware offloading, some tasks can be directly processed on the hardware, reducing the burden on the CPU and improving the response speed and processing ability of the overall system. Through the FPGA acceleration technology, it is ensured that the data directly passes through the PCIE path once, reducing the data processing and transmission delay in the intermediate links.
[0078] Optionally, the embodiments of the present application do not limit the application scenarios of the data transmission system. For example, the application scenarios of the data transmission system include, but are not limited to, cloud-native high-performance computing architectures and distributed machine learning platforms, etc. Among them, the cloud-native high-performance computing architecture means that high-performance applications under large-scale cloud-native require highly optimized networks to achieve high-speed communication between nodes. High-performance applications are usually carried out in large-scale computing clusters and interact data through thousands of parallel tasks. By using the data transmission system provided by the embodiments of the present application, multi-path load balancing, hardware offloading, and congestion control and flow control can be provided for this cloud-native high-performance computing architecture. Multi-path load balancing enables effective utilization of network resources, ensures uniform data transmission on multiple paths, and increases network utilization. Hardware offloading enables packet processing on the FPGA, significantly reducing the burden on the central processing unit, so that the central processing unit can focus more on computing tasks. Congestion control and flow control ensure smooth data communication in high-concurrency data scenarios.
[0079] The distributed machine learning platform refers to machine learning models, such as deep learning networks. Machine learning models need to process huge datasets and a large number of parameters. The training of machine learning models requires multiple server nodes and high-speed inter-node communication to synchronize data. By using the data transmission system provided by the embodiments of the present application, ultra-low latency, high bandwidth, and load balancing can be provided for this distributed machine learning platform. Ultra-low latency can ensure rapid synchronization of model parameters between nodes, thus accelerating the training process. High bandwidth can meet the transmission requirements of large-scale datasets and model parameters. Load balancing can dynamically allocate resources according to the workload of each server node during model training.
[0080] The embodiments of the present application provide a data transmission method. Refer to Figure 4 , Figure 4 which is a flowchart of a data transmission method provided by the embodiments of the present application. This data transmission method can be applied to Figure 2 or Figure 3 the data transmission system shown. Taking the data transmission system including a virtual machine, a virtual RDMA device, and a host as an example, the virtual RDMA device runs between the virtual machine and the host. The method includes, but is not limited to, the following steps 401-step 403.
[0081] Step 401, the virtual machine runs an application.
[0082] Step 402, the virtual RDMA device registers the virtual memory and virtual RDMA connection corresponding to the virtual machine, configures the memory mapping relationship between the virtual memory and the physical memory, and the connection mapping relationship between the virtual RDMA connection and the physical RDMA connection, and sends the memory mapping relationship and the connection mapping relationship to the host.
[0083] Step 403: The host transfers the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship.
[0084] In a possible implementation, the virtual machine sends an RDMA write request to the virtual RDMA device based on the running application. The RDMA write request includes a first source virtual memory, a first destination virtual memory, and a first destination communication address. The RDMA write request is used to write the first data payload in the first source physical memory corresponding to the first source virtual memory to the first destination physical memory corresponding to the first destination virtual memory. The virtual RDMA device adds the RDMA write request to the send queue.
[0085] In this case, the way for the host to transfer the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship may include: The host accesses the RDMA write request in the send queue, determines the first virtual RDMA connection according to the first destination communication address, and determines the first physical RDMA connection corresponding to the first virtual RDMA connection according to the connection mapping relationship; determines the first source physical memory corresponding to the first source virtual memory according to the memory mapping relationship, and reads the first data payload in the first source physical memory; sends a first data stream to the destination host corresponding to the first destination communication address based on the first physical RDMA connection. The first data stream includes the first destination virtual memory and the first data payload, and the first data stream is used for the destination host to write the first data payload to the first destination physical memory corresponding to the first destination virtual memory.
[0086] In a possible implementation, the host includes a CPU and a physical network card. The way to send the first data stream to the destination host corresponding to the first destination communication address based on the first physical RDMA connection may include: The CPU obtains the first data stream based on the first data payload in the first source physical memory and sends the first data stream to the physical network card; the physical network card sends the first data stream to the destination host based on the first physical RDMA connection.
[0087] In a possible implementation, the host includes an FPGA and a physical network card. The way to send the first data stream to the destination host corresponding to the first destination communication address based on the first physical RDMA connection may include: The FPGA obtains the first data stream based on the first data payload in the first source physical memory and sends the first data stream to the physical network card; the physical network card sends the first data stream to the destination host based on the first physical RDMA connection.
[0088] In a possible implementation, when the first physical RDMA connection fails or is abnormal, the host switches the first data stream transmitted by the first physical RDMA connection to be transmitted by a second physical RDMA connection, and the second physical RDMA connection is a standby connection of the first physical RDMA connection.
[0089] In a possible implementation, the virtual machine sends an RDMA read request to the virtual RDMA device based on the running application. The RDMA read request includes a second source virtual memory, a second destination virtual memory, and a second destination communication address. The RDMA read request is used to write the second data payload in the second destination physical memory corresponding to the second destination virtual memory to the second source physical memory corresponding to the second source virtual memory. The virtual RDMA device is further configured to add the RDMA read request to the receive queue. In this case, the manner in which the host transfers the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship may include accessing the RDMA read request in the receive queue, determining the second virtual RDMA connection according to the second destination communication address, and determining the third physical RDMA connection corresponding to the second virtual RDMA connection according to the connection mapping relationship; reading the second data payload based on the third physical RDMA connection; and determining the second source physical memory corresponding to the second source virtual memory according to the memory mapping relationship, and writing the second data payload to the second source physical memory.
[0090] In a possible implementation, the transmission of the RDMA data is data transmission based on the UD transmission protocol, the virtual RDMA connection is a QP connection, and one QP state is used to serve multiple QP connections. The manner in which the host transfers the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship may include that the host implements the transmission of the RDMA data according to the memory mapping relationship, the connection mapping relationship, the QP state, and the UD transmission protocol.
[0091] In a possible implementation, the virtual RDMA connections in the connection mapping relationship include priorities. The host further determines the processing order of the multiple virtual RDMA connections according to the priorities of the multiple virtual RDMA connections, and implements the transmission of the RDMA data according to the processing order.
[0092] In a possible implementation, multiple virtual RDMA connections in the connection mapping relationship correspond to a shared physical RDMA connection.
[0093] In a possible implementation, multiple virtual memories in the memory mapping relationship correspond to a shared physical memory.
[0094] For the implementation of this data transmission method, reference can be made to Figure 2 or Figure 3 the implementation of the data transmission system shown, which will not be elaborated here.
[0095] Reference is made to Figure 5 , Figure 5 which shows a schematic structural diagram of a network device 2000 provided by an exemplary embodiment of the present application. Figure 5The network device 2000 shown is used to perform the operations involved in the above-mentioned Figure 4 data transmission method. The network device 2000 is, for example, a switch, a router, etc., and the network device 2000 can be implemented by a general bus architecture.
[0096] As Figure 5 shown, the network device 2000 includes at least one processor 2001, a memory 2003, and at least one communication interface 2004.
[0097] The processor 2001 is, for example, a general-purpose central processing unit (CPU), a digital signal processor (DSP), a network processor (NP), a graphics processing unit (GPU), a neural-network processing unit (NPU), a data processing unit (DPU), a microprocessor, or one or more integrated circuits for implementing the solution of this application. For example, the processor 2001 includes an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The PLD is, for example, a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It can implement or execute various logic blocks, modules, and circuits described in connection with the disclosure of the embodiments of the present invention. The processor can also be a combination for implementing computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on.
[0098] Optionally, the network device 2000 further includes a bus. The bus is used to transfer information between the components of the network device 2000. The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 only one line is used to represent it in Figure 5 , but it does not mean that there is only one bus or one type of bus.
[0099] The memory 2003 is, for example, a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, or a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2003 exists independently, for example, and is connected to the processor 2001 through the bus. The memory 2003 can also be integrated with the processor 2001.
[0100] The communication interface 2004 uses any transceiver-like device for communicating with other devices or communication networks, which can be Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc. The communication interface 2004 can include a wired communication interface and can also include a wireless communication interface. Specifically, the communication interface 2004 can be an Ethernet interface, a Fast Ethernet (FE) interface, a Gigabit Ethernet (GE) interface, an Asynchronous Transfer Mode (ATM) interface, a wireless local area networks (WLAN) interface, a cellular network communication interface, or a combination thereof. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. In the embodiments of this application, the communication interface 2004 can be used for the network device 2000 to communicate with other devices.
[0101] In a specific implementation, as an embodiment, the processor 2001 can include one or more CPUs, such as Figure 5 CPU0 and CPU1 shown in. Each of these processors can be a single-core CPU processor or a multi-core CPU processor. The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0102] In a specific implementation, as an embodiment, the network device 2000 can include multiple processors, such as Figure 5 the processor 2001 and the processor 2005 shown in. Each of these processors can be a single-core CPU processor or a multi-core CPU processor. The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0103] In a specific implementation, as an embodiment, the network device 2000 may further include an output device and an input device. The output device communicates with the processor 2001 and can display information in various ways. For example, the output device can be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device communicates with the processor 2001 and can receive user input in various ways. For example, the input device can be a mouse, a keyboard, a touch screen device, or a sensing device, etc.
[0104] In some embodiments, the memory 2003 is used to store the program code 2010 for executing the solution of this application, and the processor 2001 can execute the program code 2010 stored in the memory 2003. That is to say, the network device 2000 can implement the data transmission method provided by the method embodiment through the processor 2001 and the program code 2010 in the memory 2003. The program code 2010 may include one or more software modules. Optionally, the processor 2001 itself can also store the program code or instructions for executing the solution of this application.
[0105] Figure 4 Each step of the data transmission method shown is completed by the integrated logic circuit of the hardware in the processor of the network device 2000 or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of this application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0106] See Figure 6 , Figure 6 shows a schematic structural diagram of a network device 2100 provided by another exemplary embodiment of this application. Figure 6 The network device 2100 shown is used to execute all or part of the operations involved in the data transmission method shown above. Figure 4 The network device 2100 is, for example, a switch, a router, etc., and the network device 2100 can be implemented by a general bus architecture.
[0107] As Figure 6 shown, the network device 2100 includes: a main control board 2110 and an interface board 2130.
[0108] The main control board is also known as the main processing unit (MPU) or the route processor card. The main control board 2110 is used for controlling and managing each component in the network device 2100, including routing calculation, device management, device maintenance, and protocol processing functions. The main control board 2110 includes: a central processing unit 2111 and a memory 2112.
[0109] The interface board 2130 is also known as the line processing unit (LPU), line card, or service board. The interface board 2130 is used to provide various service interfaces and implement packet forwarding. The service interfaces include, but are not limited to, Ethernet interfaces, POS (Packet over SONET / SDH) interfaces, etc. The Ethernet interface is, for example, a Flexible Ethernet Clients (FlexE Clients) interface. The interface board 2130 includes: a central processing unit 2131, a network processor 2132, a forwarding table entry memory 2134, and a physical interface card (PIC) 2133.
[0110] The central processing unit 2131 on the interface board 2130 is used to control and manage the interface board 2130 and communicate with the central processing unit 2111 on the main control board 2110.
[0111] The network processor 2132 is used to implement the forwarding processing of packets. The form of the network processor 2132 can be a forwarding chip. The forwarding chip can be a network processor (NP). In some embodiments, the forwarding chip can be implemented by an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA). Specifically, the network processor 2132 is used to forward the received packets based on the forwarding table entries stored in the forwarding table entry memory 2134. If the destination address of the packet is the address of the network device 2100, the packet is sent to the CPU (such as the central processor 2131) for processing; if the destination address of the packet is not the address of the network device 2100, the next hop and output interface corresponding to the destination address are found from the forwarding table according to the destination address, and the packet is forwarded to the output interface corresponding to the destination address. Among them, the processing of the upstream packets can include: the processing of the packet input interface and the forwarding table lookup; the processing of the downstream packets can include: the forwarding table lookup, etc. In some embodiments, the central processor can also perform the functions of the forwarding chip, such as implementing software forwarding based on a general-purpose CPU, so that the forwarding chip is not required in the interface board.
[0112] The physical interface card 2133 is used to implement the docking function at the physical layer. The original traffic enters the interface board 2130 from here, and the processed packets are sent out from the physical interface card 2133. The physical interface card 2133 is also called a daughter card and can be installed on the interface board 2130. It is responsible for converting optical and electrical signals into packets, performing a legality check on the packets, and then forwarding them to the network processor 2132 for processing. In some embodiments, the central processor 2131 can also perform the functions of the network processor 2132, such as implementing software forwarding based on a general-purpose CPU, so that the network processor 2132 is not required in the physical interface card 2133.
[0113] Optionally, the network device 2100 includes multiple interface boards. For example, the network device 2100 further includes an interface board 2140, and the interface board 2140 includes: a central processor 2141, a network processor 2142, a forwarding table entry memory 2144, and a physical interface card 2143. The functions and implementation methods of the components in the interface board 2140 are the same as or similar to those in the interface board 2130, and will not be elaborated here.
[0114] Optionally, the network device 2100 further includes a switch fabric board 2120. The switch fabric board 2120 can also be referred to as a switch fabric unit (SFU). When the network device 2100 has multiple interface boards, the switch fabric board 2120 is used to complete data exchange between the interface boards. For example, the interface board 2130 and the interface board 2140 can communicate through the switch fabric board 2120.
[0115] The main control board 2110 is coupled to the interface board. For example, the main control board 2110, the interface board 2130, the interface board 2140, and the switch fabric board 2120 are interconnected through a system bus and a system backplane. In a possible implementation, an inter-process communication (IPC) channel is established between the main control board 2110 and the interface board 2130 and the interface board 2140, and the main control board 2110 communicates with the interface board 2130 and the interface board 2140 through the IPC channel.
[0116] Logically, the network device 2100 includes a control plane and a forwarding plane. The control plane includes the main control board 2110 and the central processing unit 2111, and the forwarding plane includes various components that perform forwarding, such as the forwarding table entry memory 2134, the physical interface card 2133, and the network processor 2132. The control plane performs functions such as acting as a router, generating a forwarding table, processing signaling and protocol messages, configuring and maintaining the status of the network device, etc. The control plane distributes the generated forwarding table to the forwarding plane. In the forwarding plane, the network processor 2132 looks up the table and forwards the packets received by the physical interface card 2133 based on the forwarding table distributed by the control plane. The forwarding table distributed by the control plane can be stored in the forwarding table entry memory 2134. In some embodiments, the control plane and the forwarding plane can be completely separated and not on the same network device.
[0117] It should be noted that there may be one or more main control boards. When there are multiple main control boards, it may include an active main control board and a standby main control board. There may be one or more interface boards. The stronger the data processing capacity of the network device, the more interface boards are provided. There may also be one or more physical interface cards on the interface board. There may be no switching fabric board, or there may be one or more switching fabric boards. When there are multiple switching fabric boards, they can jointly achieve load sharing and redundant backup. In a centralized forwarding architecture, the network device may not require a switching fabric board, and the interface board undertakes the processing function of the service data of the entire system. In a distributed forwarding architecture, the network device may have at least one switching fabric board, and data exchange between multiple interface boards is achieved through the switching fabric board, providing a large-capacity data exchange and processing capacity. Therefore, the data access and processing capacity of the network device in the distributed architecture is greater than that of the network device in the centralized architecture. Optionally, the form of the network device may also be a single board card, that is, there is no switching fabric board, and the functions of the interface board and the main control board are integrated on this single board card. At this time, the central processing unit on the interface board and the central processing unit on the main control board can be combined into one central processing unit on this single board card to execute the functions after their superposition. The data exchange and processing capacity of this form of network device is relatively low (for example, network devices such as low-end switches or routers). Which architecture is specifically adopted depends on the specific networking deployment scenario and is not limited here.
[0118] It should be understood that the above-mentioned processor may be a CPU, or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It should be noted that the processor may be a processor supporting the advanced RISC machines (ARM) architecture.
[0119] Furthermore, in an optional embodiment, the above-mentioned memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. The memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0120] The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0121] An embodiment of the present application further provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by a processor to enable a computer to implement any one of the above data transmission methods.
[0122] An embodiment of the present application further provides a computer program (product), when the computer program is executed by a computer, it can enable the processor or the computer to execute the corresponding steps and / or processes in the above method embodiments.
[0123] An embodiment of the present application further provides a chip, including a processor, configured to call and run an instruction stored in a memory, so that a communication device equipped with the chip executes any one of the above data transmission methods.
[0124] An embodiment of the present application further provides another chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute the code in the memory, and when the code is executed, the processor is configured to execute any one of the above data transmission methods.
[0125] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk), etc.
[0126] Those of ordinary skill in the art can realize that, in combination with the method steps and modules described in the embodiments disclosed herein, they can be implemented by software, hardware, firmware, or any combination thereof. To clearly illustrate the interchangeability of hardware and software, the steps and components of each embodiment have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0127] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or by a program instructing relevant hardware to complete. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0128] When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer program instructions. As an example, the method according to an embodiment of the present application may be described in the context of machine-executable instructions, such as program modules executed in devices included in a target real or virtual processor. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., which perform specific tasks or implement specific abstract data structures. In various embodiments, the functions of program modules may be combined or split among the described program modules. The machine-executable instructions for program modules may be executed within local or distributed devices. In a distributed device, program modules may be located in both local and remote storage media.
[0129] The computer program code for implementing the method according to an embodiment of the present application may be written in one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0130] In the context of an embodiment of the present application, the computer program code or related data may be carried by any suitable carrier so that a device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like.
[0131] Examples of signals may include electrical, optical, radio, acoustic, or other forms of propagating signals, such as carrier waves, infrared signals, etc.
[0132] A machine-readable medium may be any tangible medium that contains or stores a program for or related to an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination thereof. More specific examples of machine-readable storage media include electrical connections with one or more wires, portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0133] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0134] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or modules, and can also be in electrical, mechanical, or other forms of connection.
[0135] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present application.
[0136] Furthermore, in each embodiment of the present application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0137] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc., which can store program codes.
[0138] In this application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first image can be referred to as the second image, and similarly, the second image can be referred to as the first image. Both the first image and the second image can be images, and in some cases, they can be separate and different images.
[0139] It should also be understood that in various embodiments of this application, the magnitude of the serial numbers of each process does not imply the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not impose any limitation on the implementation process of the embodiments of this application.
[0140] In this application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "a plurality of" refers to two or more. For example, a plurality of second messages refers to two or more second messages. In this document, the terms "system" and "network" are often used interchangeably.
[0141] It should be understood that the terms used in the description of various examples herein are only for describing specific examples and are not intended to be limiting. As used in the description of various examples and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0142] It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" describes the associative relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this application generally represents an "or" relationship between the associated objects before and after.
[0143] It should also be understood that the term "comprising" (also known as "includes", "including", "comprises", and / or "comprising") when used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups.
[0144] It should also be understood that the terms "if" and "when" can be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrases "if it is determined that..." or "if [the stated condition or event] is detected" can be interpreted to mean "when it is determined that..." or "in response to determining..." or "when [the stated condition or event] is detected" or "in response to detecting [the stated condition or event]".
[0145] It should be understood that determining B based on A does not mean determining B solely based on A. B can also be determined based on A and / or other information.
[0146] It should also be understood that the "one embodiment", "an embodiment", and "a possible implementation" mentioned throughout the specification mean that the specific features, structures, or characteristics related to the embodiment or implementation are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment", "a possible implementation" throughout the specification do not necessarily refer to the same embodiment. Additionally, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.
[0147] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included within the protection scope of the present application.
Claims
1. A data transmission system, characterized in that, The data transmission system includes a virtual machine, a virtual Remote Direct Memory Access (RDMA) device, and a host. The virtual RDMA device runs between the virtual machine and the host; The virtual machine is used to run applications; The virtual RDMA device is used to register the virtual memory and virtual RDMA connection corresponding to the virtual machine, configure the memory mapping relationship between the virtual memory and the physical memory, and the connection mapping relationship between the virtual RDMA connection and the physical RDMA connection, and send the memory mapping relationship and the connection mapping relationship to the host; The host is used to transmit the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship.
2. The system according to claim 1, wherein The virtual machine is further used to send an RDMA write request to the virtual RDMA device based on the running application. The RDMA write request includes a first source virtual memory, a first destination virtual memory, and a first destination communication address. The RDMA write request is used to write a first data payload in the first source physical memory corresponding to the first source virtual memory to the first destination physical memory corresponding to the first destination virtual memory; The virtual RDMA device is further used to add the RDMA write request to the send queue; The host is further used to access the RDMA write request in the send queue, determine a first virtual RDMA connection according to the first destination communication address, and determine a first physical RDMA connection corresponding to the first virtual RDMA connection according to the connection mapping relationship; Determine the first source physical memory corresponding to the first source virtual memory according to the memory mapping relationship, read the first data payload in the first source physical memory; send a first data stream to the destination host corresponding to the first destination communication address based on the first physical RDMA connection. The first data stream includes the first destination virtual memory and the first data payload, and the first data stream is used for the destination host to write the first data payload to the first destination physical memory corresponding to the first destination virtual memory.
3. The system according to claim 2, wherein The host includes a Central Processing Unit (CPU) and a physical network card; The CPU is used to obtain the first data stream based on the first data payload in the first source physical memory and send the first data stream to the physical network card; The physical network card is used to send the first data stream to the destination host based on the first physical RDMA connection.
4. The system according to claim 2, wherein The host includes a Field Programmable Gate Array (FPGA) and a physical network card; The FPGA is used to obtain the first data stream based on the first data payload in the first source physical memory and send the first data stream to the physical network card; The physical network card is used to send the first data stream to the destination host based on the first physical RDMA connection.
5. The system according to any one of claims 2-4, characterized in that, The host is further configured to switch the first data stream transmitted by the first physical RDMA connection to be transmitted by a second physical RDMA connection in the case of a failure or anomaly of the first physical RDMA connection, where the second physical RDMA connection is a backup connection of the first physical RDMA connection.
6. The system according to any one of claims 1-5, characterized in that, The virtual machine is further configured to send an RDMA read request to the virtual RDMA device based on a running application, where the RDMA read request includes a second source virtual memory, a second destination virtual memory, and a second destination communication address, and the RDMA read request is used to write a second data payload in the second destination physical memory corresponding to the second destination virtual memory to the second source physical memory corresponding to the second source virtual memory. The virtual RDMA device is further configured to add the RDMA read request to a receive queue. The host is further configured to access the RDMA read request in the receive queue, determine a second virtual RDMA connection according to the second destination communication address, and determine a third physical RDMA connection corresponding to the second virtual RDMA connection according to the connection mapping relationship. Read the second data payload based on the third physical RDMA connection; determine the second source physical memory corresponding to the second source virtual memory according to the memory mapping relationship, and write the second data payload to the second source physical memory.
7. The system according to any one of claims 1-6, characterized in that, The transmission of the RDMA data is data transmission based on an unreliable data UD transmission protocol, and the virtual RDMA connection is a queue pair QP connection, and one QP state is used to serve multiple QP connections; the host is configured to implement the transmission of the RDMA data according to the memory mapping relationship, the connection mapping relationship, the QP state, and the UD transmission protocol.
8. The system according to any one of claims 1-7, characterized in that, The virtual RDMA connections in the connection mapping relationship include priorities; the host is further configured to determine a processing order of the multiple virtual RDMA connections according to the priorities of the multiple virtual RDMA connections, and implement the transmission of the RDMA data according to the processing order.
9. The system according to any one of claims 1-8, characterized in that, Multiple virtual RDMA connections in the connection mapping relationship correspond to a shared physical RDMA connection.
10. The system according to any one of claims 1-9, characterized in that, Multiple virtual memories in the memory mapping relationship correspond to a shared physical memory.
11. A data transmission method, characterized in that, The data transmission method is applied to a data transmission system, where the data transmission system includes a virtual machine, a virtual remote direct memory access RDMA device, and a host, and the virtual RDMA device runs between the virtual machine and the host; the method includes: The virtual machine runs an application. The virtual RDMA device registers the virtual memory and the virtual RDMA connection corresponding to the virtual machine, configures a memory mapping relationship between the virtual memory and the physical memory, and a connection mapping relationship between the virtual RDMA connection and the physical RDMA connection, and sends the memory mapping relationship and the connection mapping relationship to the host. The host transmits the RDMA data of the application run by the virtual machine according to the memory mapping relationship and the connection mapping relationship.
12. The method according to claim 11, wherein The method further includes: The virtual machine sends an RDMA write request to the virtual RDMA device based on the running application. The RDMA write request includes a first source virtual memory, a first destination virtual memory, and a first destination communication address. The RDMA write request is used to write a first data payload in the first source physical memory corresponding to the first source virtual memory to the first destination physical memory corresponding to the first destination virtual memory. The virtual RDMA device adds the RDMA write request to the send queue. The host transmits the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship, including: The host accesses the RDMA write request in the send queue, determines a first virtual RDMA connection according to the first destination communication address, and determines a first physical RDMA connection corresponding to the first virtual RDMA connection according to the connection mapping relationship; determines the first source physical memory corresponding to the first source virtual memory according to the memory mapping relationship, and reads the first data payload in the first source physical memory; sends a first data stream to the destination host corresponding to the first destination communication address based on the first physical RDMA connection. The first data stream includes the first destination virtual memory and the first data payload, and the first data stream is used for the destination host to write the first data payload to the first destination physical memory corresponding to the first destination virtual memory.
13. The method according to claim 12, wherein The host includes a central processing unit (CPU) and a physical network card. The sending of the first data stream to the destination host corresponding to the first destination communication address based on the first physical RDMA connection includes: The CPU obtains the first data stream based on the first data payload in the first source physical memory and sends the first data stream to the physical network card. The physical network card sends the first data stream to the destination host based on the first physical RDMA connection.
14. The method according to claim 12, wherein The host includes a field programmable gate array (FPGA) and a physical network card. The sending of the first data stream to the destination host corresponding to the first destination communication address based on the first physical RDMA connection includes: The FPGA obtains the first data stream based on the first data payload in the first source physical memory and sends the first data stream to the physical network card. The physical network card sends the first data stream to the destination host based on the first physical RDMA connection.
15. The method according to any one of claims 12 - 14, characterized in that, The method further includes: When the first physical RDMA connection fails or is abnormal, the host switches the first data stream transmitted by the first physical RDMA connection to be transmitted by a second physical RDMA connection. The second physical RDMA connection is a standby connection of the first physical RDMA connection.
16. The method according to any one of claims 11-15, characterized in that The method further includes: The virtual machine sends an RDMA read request to the virtual RDMA device based on the running application. The RDMA read request includes a second source virtual memory, a second destination virtual memory, and a second destination communication address. The RDMA read request is used to write the second data payload in the second destination physical memory corresponding to the second destination virtual memory to the second source physical memory corresponding to the second source virtual memory. The virtual RDMA device adds the RDMA read request to the receive queue. The host transmits the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship, including: The host accesses the RDMA read request in the receive queue, determines a second virtual RDMA connection according to the second destination communication address, and determines a third physical RDMA connection corresponding to the second virtual RDMA connection according to the connection mapping relationship; reads the second data payload based on the third physical RDMA connection; determines the second source physical memory corresponding to the second source virtual memory according to the memory mapping relationship, and writes the second data payload to the second source physical memory.
17. The method according to any one of claims 11-16, characterized in that, The transmission of the RDMA data is data transmission based on the unreliable data UD transmission protocol. The virtual RDMA connection is a queue pair QP connection, and one QP state is used to serve multiple QP connections. The host transmits the RDMA data of the application running on the virtual machine according to the memory mapping relationship and the connection mapping relationship, including: The host implements the transmission of the RDMA data according to the memory mapping relationship, the connection mapping relationship, the QP state, and the UD transmission protocol.
18. The method according to any one of claims 11-17, characterized in that The virtual RDMA connections in the connection mapping relationship include priorities, and the method further includes: The host determines the processing order of the multiple virtual RDMA connections according to the priorities of the multiple virtual RDMA connections, and implements the transmission of the RDMA data according to the processing order.
19. The method according to any one of claims 11-18, characterized in that, Multiple virtual RDMA connections in the connection mapping relationship correspond to a shared physical RDMA connection.
20. The method according to any one of claims 11-19, characterized in that Multiple virtual memories in the memory mapping relationship correspond to a shared physical memory.
21. A network device, characterized in that, The network device includes: a processor, the processor is coupled with a memory, and at least one program instruction or code is stored in the memory. The at least one program instruction or code is loaded and executed by the processor to enable the network device to implement the data transmission method according to any one of claims 11-20.
22. A computer-readable storage medium, characterized in that, At least one instruction is stored in the computer storage medium. The at least one instruction is loaded and executed by a processor to enable a computer to implement the data transmission method according to any one of claims 11-20.
23. A computer program product, characterized in that, The computer program product includes: computer program code. The computer program code is loaded and executed by a computer to enable the computer to implement the data transmission method according to any one of claims 11-20.
Citation Information
Cited By
Control system and control method for controller virtualization
CN120848354A
Data transmission method and computer equipment
CN121217797A