Network Transmission Method Based on Lock-Free Queue
By adopting a lock-free queue network transmission method in RDMA network, the problem of multi-producer data competition is solved, and efficient network transmission performance is achieved, which is suitable for parallel processing of multi-core processors.
Patent Information
- Application Number
- CN202411011996.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-07-26
AI Technical Summary
In the existing RDMA network implementation method, in order to solve the data competition problem of multiple producers, mutex locks are usually used to use the entire network task dispatch process as a critical area, resulting in performance losses and the inability to fully utilize the parallel processing capabilities of multi-core processors.
Using a network transmission method based on lock-free queues, the completion queue and lock-free work queue are initialized, the ring buffer of the queue is allocated, and the work queue and completion queue are associated in QP. When the user program issues network tasks, the network card driver performs WQE enqueue operation of the lock-free work queue. The network card hardware processes the tasks in sequence according to the WQE enqueue order, and reports the task completion status to the completion queue.
It realizes parallel enqueue and poll dequeuing of multiple RDMA network tasks, improves network transmission performance, adapts to the semantics of RDMA network communication, and reduces memory copy operations.
Smart Images

Figure CN119052344B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network technology, and particularly relates to a network transmission method based on a lock-free queue. Background Art
[0002] An asynchronous queue is an important part of the RDMA network I / O model. During the communication process of RDMA, the user program issues network tasks through a work queue (WQ). After the network card hardware completes the network tasks in sequence according to the enqueue order of the work queue elements (WQE), it reports the completion status of each task through the completion queue (CQ). Compared with the TCP / IP network, the RDMA user program does not need to block and wait for the data buffer to be available, nor is there a process of synchronizing data copying, thus realizing high-bandwidth and low-latency network transmission based on the asynchronous queue.
[0003] The enqueue and dequeue operations of the queue are typical producer-consumer problems. In the existing open-source RDMA network implementation methods, to solve the data competition problem of multiple producers, a mutex is often used to regard the entire network task issuing process as a critical section, so as to ensure mutually exclusive access to the shared resources of the queue. However, the mutex itself has certain performance losses, and the method using the mutex cannot fully utilize the parallel processing ability of the multi-core processor during the enqueue operation of the WQE. Summary of the Invention
[0004] The present invention provides a network transmission method based on a lock-free queue to solve the above-mentioned technical problems, and specifically adopts the following technical solutions:
[0005] A network transmission method based on a lock-free queue includes:
[0006] Initialize the completion queue and the lock-free work queue, and allocate the circular buffer of the queue;
[0007] Associate the work queue and the completion queue in the QP (Queue Pair, communication pair) and perform unified management;
[0008] When the user program issues a network task, the network card driver performs the enqueue operation of the WQE of the lock-free work queue;
[0009] The network card hardware processes the tasks in sequence according to the enqueue order of the WQE, and reports the task completion status to the completion queue after completion;
[0010] The network card driver polls the completion queue to obtain available CQE and parses the completion status of the network task;
[0011] The network card driver dequeues the WQE corresponding to the CQE from the lock-free work queue and reports the completion status to the user program.
[0012] Furthermore, the completion queue and the lock-free work queue are located in the host memory, and the occupied memory space is allocated by the network card driver according to the queue size specified by the user program.
[0013] Furthermore, the circular buffer of the queue is an array with continuous addresses and the same element size in the memory. The circular buffer of the completion queue is used to store the task completion element CQE. The lock-free work queue contains two circular buffers, one for saving the context information of the send task and the other for storing the SWQE, and the SWQE contains all the information required for the network card hardware to fill the network packet.
[0014] Furthermore, when creating the lock-free work queue, configure the type of the queue. When only a single thread adds WQEs to the work queue, configure the queue as a single-producer type, and no atomic instructions are required to protect the shared resources of the queue during the enqueue operation. When multiple threads concurrently add WQEs to the work queue, configure the queue as a multi-producer type, and atomic instructions are used to protect the shared resources of the queue during the enqueue operation.
[0015] Furthermore, when creating the lock-free work queue, configure the size of the circular buffer of the queue. The work queues in the QP are divided into a send work queue and a receive work queue, and the user program specifies the work queue depth as needed, that is, the maximum number of WQEs that the send queue and the receive queue can accommodate.
[0016] Furthermore, when the network card driver performs the enqueue operation of the lock-free work queue, the network card driver performs a two-stage enqueue operation. The first stage is responsible for allocating the corresponding queue element position for each thread. After successfully obtaining the right to use the queue element position, the user thread saves the context information of the send task. The second stage is responsible for filling the registers in sequence according to the enqueue order of the threads to notify the network card hardware that the SWQE at the corresponding position has arrived, and finally restores the producer tail pointer of the queue.
[0017] Furthermore, when the network card driver issues a network task, it notifies the hardware that a new WQE has arrived by writing to the network card register.
[0018] Furthermore, when the network card driver performs the enqueue and dequeue operations of the WQE, it uses atomic variables and atomic instructions to ensure the consistency of the queue operation semantics among multiple threads.
[0019] Furthermore, after the network card hardware completes the network task, it reports the task completion event by setting a flag in the Completion Queue Element (CQE) to 1.
[0020] Furthermore, the network card driver uses a proxy thread to poll the completion queue. The network card driver obtains available CQEs according to the set flag bits, and the CQEs contain the success or failure completion status of a transmission task.
[0021] The beneficial effect of the present invention lies in the network transmission method based on a lock-free queue, which allows multiple RDMA network tasks to be enqueued in parallel, and dequeues the network tasks sequentially when polling the CQEs. When multiple threads share the work queue of the same QP, the performance of network transmission can be effectively improved.
[0022] The beneficial effect of the present invention also lies in the network transmission method based on a lock-free queue. When enqueuing or dequeuing the WQEs, no memory copy is required, and it is more adaptable to the semantics of RDMA network communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0024] Figure 1 It is a schematic structural diagram of an embodiment of a network transmission based on a lock-free queue provided by the present invention;
[0025] Figure 2 It is a flowchart of a network transmission method based on a lock-free queue provided by the present invention;
[0026] Figure 3 It is a schematic diagram of an embodiment of multiple threads performing lock-free queue enqueue operations provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following details the embodiments of the present application. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application, but should not be construed as a limitation of the present application.
[0028] The following illustrates a network transmission method based on a lock-free queue provided by the present invention through an embodiment. As Figure 1 shown, the purpose of this embodiment is to allow multiple threads to issue network tasks through the same lock-free send queue. Utilizing the enqueue characteristic of the lock-free queue, the user threads issuing tasks can write SWQEs into the circular buffer in parallel, thereby improving the data transmission performance of the RDMA network.
[0029] As shown in Figure 2 Figure [0000067], a network transmission method based on a lock-free queue according to the present application specifically includes:
[0030] Step S1: Initialize the completion queue and the lock-free work queue, and allocate resources such as the circular buffer of the queue.
[0031] In the embodiment of the present application, the completion queue and the lock-free work queue are located in the host memory, and the occupied memory space is allocated by the network card driver according to the queue size specified by the user program. A part of the memory resources of the queue is used to store the control information of the queue, such as the size of the queue, the producer head and tail pointers prod_head and prod_tial, the consumer head and tail pointers cons_head and cons_tail, the pointer of the RDMA callback function that needs to be called when the work queue performs an enqueue operation, etc. Another part of the memory resources serves as the circular buffer of the queue.
[0032] Specifically, the circular buffer of the queue is an array with continuous addresses and the same element size in memory. The circular buffer of the completion queue is used to store the task completion elements CQE. The lock-free work queue contains two circular buffers. As shown in Figure 1 Figure [0000074], one of them is used to save the context information of the send task, such as the work_id of the network task, the pointer to the corresponding SWQE, etc. The other is used to store the SWQE, and the SWQE contains all the information required for the network card hardware to fill the network packet, such as the source address of the network data, the destination address of the network data, the length, etc. The two circular buffer elements of the lock-free work queue correspond one by one. When initializing the work queue, it is necessary to point the SWQE pointer in the send task element to the address where the SWQE is located at the corresponding position, as shown in Figure 1 Figure [0000075].
[0033] In addition, when creating the lock-free work queue, the type of the queue can be configured. As shown in Figure 1 Figure [0000078], in the embodiment of the present application, when there are multiple threads concurrently adding SWQEs to the work queue, the queue needs to be configured as a multi-producer type, and atomic instructions are required to protect the shared resources of the queue during the enqueue operation. Since this implementation uses a single proxy thread to poll the completion queue and mutex locks are used to protect the completion queue in general scenarios, this means that only one consumer will perform the dequeue operation of the work queue. Therefore, the lock-free work queue can be configured as a single-consumer type, and atomic instructions are not required to protect the queue resources during the dequeue operation of the SWQE, further improving the dequeue performance of the lock-free work queue.
[0034] In an embodiment of the present application, when creating a lock-free work queue, the size of the circular buffer of the queue is configured. The work queues in the QP (Queue Pair) are divided into a send work queue and a receive work queue. The user program specifies the work queue depth as needed, that is, the maximum number of WQEs that the send queue and the receive queue can accommodate.
[0035] Step S2: Associate the work queue and the completion queue in the QP and perform unified management.
[0036] In RDMA programming, the QP is an interface for the local user program to establish a connection with the remote user program and issue network tasks. After the network card driver completes the initialization of the completion queue, it needs to save the association information of the work queue and the completion queue in the QP context.
[0037] Step S3: When the user program issues a network task, the network card driver performs an enqueue operation on the lock-free work queue.
[0038] In an embodiment of the present application, the user program creates multiple threads from a thread pool, and each thread issues network tasks to the same send work queue. If the queue resources are not atomically protected, multiple threads may fill the SWQE at the same position, resulting in a data competition problem. As a preferred embodiment, this embodiment implements a two-stage enqueue operation method to solve the above problem. As Figure 3 shown, the first stage is responsible for allocating a corresponding queue element position for each thread. After successfully obtaining the usage right of the queue element position, the user thread saves the context information of the send task, including the work_id of the task, and obtains the address of the SWQE initialized in step S1. Finally, it calls the RDMA callback function registered in step S1 to fill the SWQE. The second stage is responsible for filling the register in sequence according to the enqueue order of the threads to notify the network card hardware that the SWQE at the corresponding position has arrived, and finally restoring the producer tail pointer of the queue.
[0039] Specifically, in the first stage, two threads simultaneously request to add send tasks to the work queue. At this time, the global prod_head and prod_tail pointers both point to the next available position in the queue, that is, the position after send task 2. First, thread 1 and thread 2 copy the global prod_head to the local variables prod1_head and prod2_head. Secondly, according to the number of send tasks requested to be enqueued, the local variables prod1_tail and prod2_tail are updated respectively. Then thread 1 and thread 2 execute the Compare And Swap (CAS) atomic instruction. This instruction will assign the local prod1_tail or prod2_tail to the global prod_head and return success only when the global prod_head is equal to the local prod1_head or prod2_head, otherwise it returns failure. Since the CAS instruction is an atomic operation, only one thread will successfully execute this operation. In the embodiment of the present application, thread 1 successfully executes the CAS instruction, and thread 2 needs to re-execute the above steps. Finally, thread 1 successfully obtains the right to use the queue element where send task 3 is located, and thread 2 successfully obtains the right to use the queue element where send task 4 is located. At this time, the states of the global producer head and tail pointers and the local producer head and tail pointers are as Figure 3 shown in Stage 1.
[0040] After successfully obtaining the right to use the queue element, multiple threads concurrently call the RDMA callback function to fill in the relevant information of the send task into the SWQE at the corresponding position. Since the network card hardware fetches the SWQE in sequence, fills the network packet according to the content of the SWQE and sends it to the remote node, so the write register operation cannot be performed in the callback function, but wait until Stage 2 to write the register in sequence according to the order in which the threads obtain the queue element.
[0041] Specifically, as Figure 3 shown in Stage 2, each thread needs to update the global prod_tail to the local prod1_tail or prod2_tail position. First, thread 1 and thread 2 use the atomic instruction to compare the global prod_tail and the local variables prod1_head or prod2_head. Only when the global prod_tail is equal to the local variables prod1_head or prod2_head, will it exit the infinite loop. In the embodiment of the present application, after thread 1 successfully exits the loop, it writes a register to notify the network card hardware that send task 3 has arrived, and uses the atomic instruction to assign prod1_tail to prod_tail. Then thread 2 can successfully exit the loop and write a register to notify the network card hardware that send task 4 has arrived. Finally, the global producer head and tail pointers point to the next available position, that is, the position after send task 4.
[0042] At this point, multiple user threads have completed the enqueue operation of the sending tasks.
[0043] Step S4: The network card hardware processes the tasks in the order of WQE enqueue, and reports the task completion status to the completion queue after completion.
[0044] As a preferred implementation, after the network card hardware completes the network sending task, the network card hardware sets the flag bit corresponding to the CQE in the completion queue to 1 to indicate that the completion event has arrived.
[0045] Step S5: The network card driver polls the completion queue to obtain available CQEs and parses the completion status of the network tasks.
[0046] As a preferred implementation, this embodiment uses a proxy thread to poll the completion queue. The network card driver obtains available CQEs according to the flag bits set in step S4. The CQE contains the success or failure completion status of a sending task.
[0047] Step S6: The network card driver dequeues the WQE corresponding to the CQE from the lock-free work queue and reports the completion status to the user program.
[0048] Specifically, the network card driver dequeues the sending task corresponding to the CQE obtained in step S5. Since there is only one consumer using the work queue at this time, the dequeue operation of the queue can be implemented using a single-threaded implementation method without using mutex locks or atomic instructions. Finally, the network card driver fills the fields in the CQE and the corresponding sending task into the completion event context, which includes the work_id filled in step S3, and returns it to the proxy thread. The proxy thread can distribute the completion status of the sending task to each thread according to different work_ids.
[0049] In the implementation of this application, when the network card driver issues a network task, it notifies the hardware of the arrival of a new WQE by writing to the network card register. Moreover, when the network card driver performs the enqueue and dequeue operations of the WQE, it uses atomic variables and atomic instructions to ensure the consistency of the queue operation semantics among multiple threads.
[0050] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the above embodiments do not limit the present invention in any form. Any technical solutions obtained by using equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A network transmission method based on a lock-free queue, characterized in that: include: Initialize the completion queue and lock-free work queue, and allocate the queue's ring buffer; Associate the work queue and the completion queue in the communication pair QP and manage them uniformly; When the user program issues a network task, the network card driver executes the WQE enqueue operation of the lock-free work queue; The NIC hardware processes tasks in the order they are queued by WQE, and reports the task completion status to the completion queue after completion; The network card driver polls the completion queue to obtain the available completion queue element CQE and parses the completion status of the network task; The network card driver dequeues the WQE corresponding to the CQE from the lock-free work queue and reports the completion status to the user program; When creating a lock-free work queue, configure the queue type. When only a single thread adds WQE to the work queue, configure the queue as a single-producer type, and there is no need to use atomic instructions to protect queue shared resources during enqueue operations. When multiple threads add WQE to the work queue concurrently, configure the queue as a multi-producer type, and use atomic instructions to protect queue shared resources during enqueue operations.
2. According to the lock-free queue-based network transmission method of claim 1, it is characterized in that: The completion queue and the lock-free work queue are located in the host memory, and the memory space they occupy is allocated by the network card driver according to the queue size specified by the user program.
3. According to the lock-free queue-based network transmission method of claim 1, it is characterized in that: The queue's circular buffer is an array with consecutive addresses and elements of the same size in memory. The completion queue's circular buffer is used to store the task completion element CQE. The lock-free work queue contains two circular buffers, one of which is used to save the context information of the sending task, and the other is used to store the sending work queue element SWQE. SWQE contains all the information required by the network card hardware to fill the network message.
4. A network transmission method based on a lock-free queue according to claim 1, characterized in that: When creating a lock-free work queue, configure the queue's ring buffer size. The work queues in QP are divided into send work queues and receive work queues. The user program specifies the work queue depth as needed, that is, the maximum number of WQEs that the send queue and receive queue can accommodate.
5. A network transmission method based on a lock-free queue according to claim 1, characterized in that: When the network card driver performs the enqueue operation of the lock-free work queue, the network card driver performs a two-stage enqueue operation. The first stage is responsible for allocating the corresponding queue element position for each thread. After successfully obtaining the right to use the queue element position, the user thread saves the context information of the sending task. The second stage is responsible for filling in the registers in the order of thread enqueue to notify the network card hardware that the SWQE at the corresponding position has arrived, and finally restores the producer tail pointer of the queue.
6. A network transmission method based on a lock-free queue according to claim 5, characterized in that: When the network card driver issues a network task, it notifies the hardware of the arrival of a new WQE by writing to the network card register.
7. A network transmission method based on a lock-free queue according to claim 6, characterized in that: When the network card driver performs WQE enqueue and dequeue operations, it uses atomic variables and atomic instructions to ensure the consistency of queue operation semantics among multiple threads.
8. A network transmission method based on a lock-free queue according to claim 7, characterized in that: After the network card hardware completes the network task, it reports the task completion event by setting a flag in CQE to 1.
9. A network transmission method based on a lock-free queue according to claim 8, characterized in that: The network card driver uses a proxy thread to poll the completion queue. The network card driver obtains the available CQE according to the flag bit set to 1. The CQE contains the success or failure completion status of a sending task.
Citation Information
Patent Citations
Multi-thread data lock-free processing method and device and electronic equipment
CN113051057A
Practical contention-free distributed weighted fair-share scheduler
US20100211954A1