Data transmission synchronization processing methods and network cards
Patent Information
- Application Number
- CN202610954295.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-30
AI Technical Summary
[0004]上述处理方式需要主机CPU参与控制,增加了主机开销与软硬件交互延迟
(1)实现数据同步:网卡在读操作完成后才生成CQE,保证主机收到CQE时数据已真正写入目标设备,避免了数据不同步问题;
Smart Images

Figure CN122470398B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data communication technology, and in particular to a data transmission synchronization processing method and a network interface card (NIC). Background Technology
[0002] like Figure 1 As shown, in data reception and transmission scenarios, after the network interface card (NIC) receives a write message and writes the data to the target device, it typically reports a completion queue entry (CQE) to the host to notify that the transmission is complete. In systems incorporating switching equipment (such as PCIe switches, CXL switches, etc.), data writing and CQE reporting may follow different routes of the switching equipment, resulting in different queuing and forwarding delays on the bus, and even potentially causing out-of-order packets. Therefore, the NIC may generate a CQE report before the data is actually written to the target device, causing the host to receive the CQE without the data actually being written, leading to data asynchrony or system logic errors.
[0003] Existing technologies employ a loopback RDMA read flush scheme to ensure reliable data write confirmation: After the network interface card (NIC) writes data to the target device via Direct Memory Access (DMA), it directly generates a CQE and writes it to the host memory. Upon receiving the CQE, the host application layer sends a loopback RDMA read operation request to the NIC. The NIC reads data from the target device at the specified address via DMA, encapsulates it into a standard RDMA read response message, and sends it to its own receive queue through the internal loopback path. Upon receiving the message, the NIC writes the data back to the target device and generates an acknowledgment (ACK). Upon receiving the ACK, the NIC generates the corresponding read operation completion queue entry (Read CQE) and writes it to the host memory, thus confirming that the data has been written.
[0004] The above processing method requires host CPU involvement, increasing host overhead and hardware / software interaction latency. Furthermore, loopback RDMA Read involves the complete RDMA protocol stack processing flow, requiring the network interface card's own queue pair (QP) resources, resulting in a heavy overall workload. Therefore, in high-concurrency scenarios, this approach significantly reduces system processing efficiency and impacts data consistency. Summary of the Invention
[0005] This application provides a data transmission synchronization processing method and a network card, which adopts a lightweight DMA read operation to realize the synchronous confirmation of data persistence status within the network card hardware, ensuring that CQE generation is later than data writing completion, and eliminating the application layer loopback RDMA Read process, thereby reducing latency, increasing throughput, and saving network card hardware resources.
[0006] In a first aspect, one embodiment of this application provides a data transmission synchronization processing method, the method being autonomously executed by hardware logic in a network interface card (NIC), the method comprising: The network card receives a write message from the sender and performs a write operation via Direct Memory Access (DMA) to write the data in the write message to the target device; the storage space of the target device does not belong to the host memory. After initiating the write operation, if the network interface card (NIC) receives a packet carrying an immediate value, it will proactively initiate a read operation via DMA to read data from the specified address of the target device. The read operation and the write operation target the same device, and based on the bus transaction ordering rules, the read operation and the write operation are executed sequentially in the order they are initiated. For write packets without an immediate value, the NIC only initiates the write operation, does not initiate the read operation, and does not generate a completion queue entry (CQE). The network interface card (NIC) generates a CQE and writes the CQE to the host memory only after the read operation is completed; before the read operation is completed, the NIC does not perform the CQE generation and writing to the host memory operation.
[0007] Optionally, the message carrying an immediate value is a write message carrying an immediate value, and the write message carries a destination address; the network interface card initiates the read operation using the destination address carried in the write message as the specified address.
[0008] Optionally, the network interface card (NIC) has received at least one write message without an immediate value before receiving a write message carrying an immediate value.
[0009] Optionally, the write message does not carry an immediate value; the message carrying an immediate value is a notification message that carries an immediate value but does not carry a destination address or data; the method further includes: The network interface card records the target address carried in the write message; In response to receiving the notification message, after all write operations prior to the notification message have been initiated, the network interface card initiates the read operation using the recorded target address as the specified address.
[0010] Optionally, the method further includes: the network card records the target address of each write packet into the same address register, and the target address of the later received write packet overwrites the target address of the earlier received write packet.
[0011] Optionally, the specified address is a pre-configured fixed address.
[0012] Secondly, one embodiment of this application provides a network interface card (NIC), comprising: The receiving module is used to receive write messages from the sending end; The DMA engine is used to perform write operations via Direct Memory Access (DMA) to write data in the write message to the target device; the storage space of the target device is not part of the host memory. The operation control logic is used to, after initiating the write operation, if a message carrying an immediate value is received, control the DMA engine to actively initiate a read operation to read data from the specified address in the target device; wherein, the read operation and the write operation are for the same target device, and the read operation and the write operation are executed sequentially according to the initiation order based on the bus transaction sorting rules; for write messages without immediate values, only the DMA engine is controlled to initiate the write operation, without initiating the read operation or generating a CQE; The queue management logic is used to generate a completion queue entry (CQE) and write the CQE to the host memory only after the read operation is completed; the CQE generation and writing to the host memory operations are not performed before the read operation is completed.
[0013] Compared with the prior art, the data transmission synchronization processing method and network card provided in this application have the following advantages: (1) Achieve data synchronization: The network card generates CQE only after the read operation is completed, ensuring that the data has been truly written to the target device when the host receives the CQE, thus avoiding the problem of data asynchrony; (2) No host CPU involvement required: The read operation is automatically triggered by the network card hardware based on the packet carrying the immediate value, without the host issuing instructions, thus releasing host resources; (3) Lightweight implementation: Read operations are initiated directly through DMA, without the need to construct RDMA messages or occupy QP resources, thus reducing latency and increasing throughput. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. Obviously, the drawings introduced below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a schematic diagram illustrating an application scenario of the data transmission synchronization processing method provided in an embodiment of this application. Figure 2 This is a flowchart illustrating a data transmission synchronization processing method according to an embodiment of this application. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of this application provide method operation steps as shown in the following embodiments or drawings, the method may include more or fewer operation steps based on conventional or non-inventive effort. For steps that do not logically have a necessary causal relationship, the execution order of these steps is not limited to the execution order provided in the embodiments of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other.
[0017] For ease of understanding, the terms used in the embodiments of this application are explained below: Sender: refers to the network device or host that sends the write message. It can be a device that actively sends data or a device that returns data in response to a request. In this application, the sender is... Figure 1 The network card shown is connected to a peer network device via a network, rather than to a local device connected to the network card.
[0018] Network Interface Card (NIC): Refers to a Network Interface Controller used to connect a computer to a network and receive and send data packets. In this application, the NIC has Direct Memory Access (DMA) capability and can actively initiate read operations.
[0019] Direct Memory Access (DMA): This refers to a data transfer method that allows devices (such as network interface cards) to directly read and write data in memory without going through a central processing unit (CPU). In this application, both write and read operations are performed using DMA.
[0020] Write message: A write request message sent by the sender, carrying data and the destination address. Write messages may or may not carry immediate values.
[0021] Target address: Indicates the storage location in the target device where data is written.
[0022] Target device: refers to the destination of data writing, including but not limited to graphics processing unit (GPU), field programmable gate array (FPGA), neural network processor (NPU) or other storage devices coupled to the network card via a bus, whose storage space does not belong to the host memory.
[0023] Read operation: refers to the operation initiated by the network card to read data from the target device. This read operation is executed via DMA and is executed sequentially with the write operation in the order of initiation. The bus transaction ordering rules ensure that the read operation is not executed before the write operation.
[0024] Bus transaction ordering rules: These refer to the transaction ordering rules defined by a bus (such as a PCIe bus) to ensure data consistency. In default or strict ordering mode, read operations originating from the same bus interface are forced to be executed after all previously initiated write operations. Based on this rule, read and write operations targeting the same device and initiated by the same network interface card are executed sequentially in the order they were initiated.
[0025] Specified address: Specifies the memory address in the target device to be read by the read operation. This specified address can be the target address carried in the write packet, the write packet address recorded by the network card, or a pre-configured fixed address.
[0026] Immediate value: This refers to user-defined data embedded in a message, used to indicate that the message needs to generate a completion notification. The network interface card (NIC) determines whether to trigger a read operation based on whether the message carries an immediate value.
[0027] Notification message: This refers to an RDMA write message that carries an immediate value but not the destination address and data. It is used to indicate that a completion notification needs to be generated. After receiving the notification message, the network card triggers a read operation to confirm that the previous data write has been completed.
[0028] Completion Queue Entry (CQE): This refers to an entry in the Completion Queue (CQ), generated by the network interface card (NIC) hardware and written to the host memory. It is used to notify the application layer on the host CPU that an operation has been completed. In this application, the CQE is generated only after a read operation is completed.
[0029] Data transmission synchronization processing: In this application, the read operation verification mechanism automatically executed by the hardware ensures that the moment the host receives the completion queue entry (CQE) is consistent with the moment the data is actually written to the target device, thereby achieving synchronization between the host notification status and the device data status.
[0030] Data landing: refers to the data being actually written to the physical storage medium of the target device (such as GPU memory), rather than just arriving at the network card or being in transit.
[0031] refer to Figure 1 The solution proposed in this application is implemented in a computer system that includes the following components: (1) Network card: It is a smart network card with DMA engine and hardware logic control capabilities. It can automatically initiate DMA read operation after initiating write operation and generate CQE after the read operation is completed.
[0032] (2) Target device: Storage or computing device coupled to the network card via a high-speed bus, including but not limited to GPU (graphics processing unit), FPGA (field programmable gate array), NPU (neural network processor) or SSD (solid-state drive).
[0033] (3) Host: Includes CPU (Central Processing Unit) and host memory, running operating system and applications.
[0034] (4) Switching devices are used to connect network interface cards (NICs), hosts, and target devices to enable high-speed data exchange between multiple devices. NICs and target devices can be connected via switching devices (such as PCIe switches, CXL switches, etc.) to support point-to-point (P2P) DMA transmission, or directly connected via a bus root node, or coupled through other high-speed bus interfaces. During bus transaction forwarding, switching devices may select different forwarding paths for different types of transactions (such as data write transactions and completion notification transactions), resulting in different queuing and forwarding delays between transactions. NICs are coupled to hosts and target devices via high-speed buses, including but not limited to PCIe, CXL, or NVLink, which ensure that bus transactions are executed in the order they are initiated.
[0035] The proposed solution can be applied to scenarios such as distributed AI training, high-performance computing, data center storage, and edge computing, and is especially suitable for environments that require highly reliable data transmission and time-series deterministic guarantees.
[0036] refer to Figure 2 This application provides a data transmission synchronization processing method, applied to a network interface card (NIC), specifically including the following steps: S101. The network card receives the write message and performs the write operation through direct memory access (DMA) to write the data in the write message to the target device.
[0037] Specifically, the network interface card (NIC) receives a write message from the sender, which carries data and a destination address. The NIC parses the write message, extracts the data and destination address, and then initiates a DMA write operation to directly write the data to the storage location corresponding to the destination address in the target device. This DMA write operation does not require host intervention and is completed autonomously by the NIC hardware.
[0038] S102. After initiating a write operation, if a message carrying an immediate value is received, the network card will actively initiate a read operation via DMA to read data from the specified address in the target device; the read operation and the write operation are executed sequentially in the order of initiation.
[0039] Specifically, after initiating the DMA write operation in step S101, the network card parses the received packet to determine whether it carries an immediate value; if it does, it actively initiates a DMA read operation. The network card actively initiates a DMA read operation. This read operation targets a specified address and reads data from the target device at that specified address.
[0040] In this embodiment, the high-speed bus employs strict transaction ordering rules to ensure that all operations (including read and write operations) are executed sequentially in the order they are initiated. This transaction ordering rule can be implemented through relevant configurations. Taking the PCIe bus as an example, setting the Relaxed Ordering configuration bit to 0 enables the PCIe bus to adopt a strict ordering mode, meaning that later initiated operations are not allowed to exceed previously queued operations. For details, please refer to the PCIe bus specification.
[0041] Since the read operation in this embodiment is initiated after the write operation, the transaction ordering rules of the high-speed bus ensure that the read operation is executed after all previous write operations. Therefore, when the read operation is completed, it means that all previous write operations have been completed. The network card does not need to wait or confirm; it can generate the corresponding CQE based on the completion status of the read operation and write the CQE to the host memory, thereby ensuring that the data has been actually written to the target device when the host receives the CQE.
[0042] S103. The network card generates a completion queue entry (CQE) only after the read operation is completed, and writes the CQE to the host memory.
[0043] In this embodiment, the network interface card (NIC) generates a CQE and writes it to the host memory only after the read operation is completed. Before the read operation is completed, the NIC does not perform the CQE generation operation, nor does it perform the CQE writing operation to the host memory.
[0044] Specifically, after the DMA read operation in step S102 is completed, the network card generates a corresponding CQE, which may contain information such as operation status and number of bytes. The network card writes the CQE into the completion queue (CQ) in the host memory via DMA to notify the host application layer that the data corresponding to the write operation has been successfully stored in the target address of the target device. This ensures that the time when the host receives the CQE is synchronized with the time when the data is actually stored.
[0045] The data transmission synchronization method in this embodiment initiates a read operation after a write operation and utilizes the transaction ordering mechanism of the high-speed bus to ensure that the read operation is executed after the write operation. Therefore, the completion of the read operation indicates that all previous write operations have been completed. The network card only generates a CQE after the read operation is completed, ensuring that when the host receives the CQE, the data corresponding to the write operation has actually been written to the target device, thereby avoiding the data synchronization problem in the prior art where the host receives the CQE but the data has not yet been actually written.
[0046] To address the out-of-order data writing and CQE reporting issues caused by routing in switching devices, existing technologies employ a loopback RDMA Read flush scheme. This involves the host application layer actively issuing an RDMA Read command upon receiving a CQE request to verify data completion. This approach suffers from high host CPU overhead and high hardware / software interaction latency. In contrast, in this embodiment, the read operation is automatically initiated by the network interface card (NIC) hardware, requiring no host CPU intervention and effectively freeing up host CPU resources. Furthermore, the NIC directly initiates the read operation via DMA, eliminating the need to construct RDMA packets and consume NIC QP resources, resulting in a more lightweight approach. Because it eliminates the need for host CPU involvement, RDMA protocol stack processing, and QP resource consumption, the proposed solution significantly reduces end-to-end data transmission latency. This advantage is particularly pronounced in high-concurrency scenarios, effectively improving system throughput.
[0047] As described in the previous embodiments, the network interface card (NIC) initiates a read operation after initiating a write operation. This embodiment further specifies that the read operation is initiated by the NIC after receiving a packet carrying an immediate value.
[0048] Specifically, the network interface card (NIC) receives packets from the sender. These packets can be write packets carrying immediate values, write packets without immediate values, or notification packets carrying immediate values but without the destination address and data. The immediate value in the packet indicates that a completion notification needs to be generated; that is, after completing data writing and read verification, the NIC needs to generate a CQE and report it to the host. Packets carrying immediate values trigger read operations, while packets without immediate values do not trigger read operations.
[0049] The workflow of this embodiment is as follows: Step 1: The network card receives the packet, parses the packet header, and determines whether it carries an immediate value. If it carries an immediate value, proceed to Step 2; otherwise, proceed to Step 3. Step 2: The network card further determines whether the packet carries data; if the packet carries data (i.e., a write packet), proceed to step 2-1; if the packet does not carry data (i.e., a notification packet), proceed to step 2-2. Step 2-1: The network card initiates a DMA write operation to write the data in the packet to the target device; after initiating the write operation, it actively initiates a DMA read operation to read the data at the specified address from the target device; after the read operation is completed, a CQE is generated and written to the host memory. Step 2-2: The network card does not perform a DMA write operation; it directly triggers a read operation based on the immediate value; after the read operation is completed, a CQE is generated and written to the host memory. Step 3: The network card performs a DMA write operation to write the data in the packet to the target device without triggering a read operation or generating a CQE.
[0050] In this embodiment, the network card automatically triggers the read operation based on the packet carrying an immediate value, without the need for the host CPU to issue instructions, thus freeing up host CPU resources. The speed at which the network card hardware parses the packet and triggers the read operation is much faster than the CPU issuing instructions, reducing response latency. Whether it is a write packet carrying data or a notification packet without data, the read operation can be triggered uniformly by an immediate value, which is convenient for hardware implementation.
[0051] According to the triggering conditions in the foregoing embodiments, packets carrying immediate values trigger read operations, while packets without immediate values do not trigger read operations. Based on the different types of packets that trigger read operations, this application provides the following three application scenarios: Scenario 1: Each write message sent by the sender carries an immediate value.
[0052] The network card uses the target address carried in the write message as the specified address for the read operation. That is, when the network card initiates a DMA read operation, it directly reads the data at the target address.
[0053] The specific workflow is as follows: The network card receives a write message carrying an immediate value and extracts the target address and data from the message; the network card initiates a DMA write operation to write the data to the storage location corresponding to the target address in the target device; after initiating the write operation, the network card actively initiates a DMA read operation using the target address as the specified address to read the data at the target address from the target device; after the read operation is completed, the network card generates a CQE and writes it to the host memory.
[0054] In the embodiment of Scenario 1, the specified address for the read operation comes directly from the target address carried in the write message, requiring no additional recording or configuration, thus simplifying implementation. Since the address read by the read operation is the same as the address written by the write operation, according to the transaction sorting rules of the high-speed bus, read operations for the same address are forced to be executed after write operations, ensuring that read operations are executed only after write operations are completed.
[0055] Scenario 1 is applicable to single-message scenarios. When a write message independently completes data transmission and acknowledgment, this embodiment provides the most direct and efficient processing method.
[0056] Scenario 2: After sending one or more write messages without immediate values, the sender then sends a notification message.
[0057] In Scenario 2, the write packets sent by the sending end do not carry immediate values, but only the destination address and data. The network interface card (NIC) records the destination address carried in the received write packets. In response to receiving a notification packet, the NIC waits until all write operations initiated before receiving the notification packet are completed, and then initiates a read operation using the recorded destination address as the specified address.
[0058] The specific workflow is as follows: After receiving a write packet, the network card performs the following operations: writes the data in the write packet to the target address in the target device and records the target address; after receiving a notification packet, the network card waits for all write operations before the notification packet to be initiated, and actively initiates a DMA read operation with the recorded target address as the specified address; after the read operation is completed, the network card generates a CQE and writes it to the host memory.
[0059] Since the read operation is initiated after all write operations preceding the notification message, and the bus transaction ordering rules guarantee that the read operation executes after the write operation, the completion of this read operation indicates that the data in all previous write messages without immediate values has been successfully written to the target device. The generated CQE is used to notify the host that the above data has been reliably stored.
[0060] In scenario two, the network card can record the destination address of the write packet in any of the following ways: (1) Record the last address: The network card records the destination address of each write packet into the same address register, and the destination address of the later received write packet overwrites the destination address of the earlier received write packet. The final recorded address is the address of the last write packet. This method is simple to implement and does not require checking whether it is the first packet.
[0061] (2) Record the first address: The network card only records the target address of the first received write message, and the record is not updated for subsequent write messages. This method records the starting address of the message.
[0062] The implementation of Scenario 2 triggers read operations via notification messages and uses the recorded target address as the specified address for the read operation, which better supports multi-message scenarios. When a large message is split into multiple write messages for transmission, it is not necessary to trigger a read operation for each message; only the last notification message needs to trigger a read operation to confirm that all data has been written, reducing the number of read operations and lowering bus overhead. This solution is also applicable to out-of-order scenarios, ensuring data integrity when messages arrive out of order through a target address recording mechanism.
[0063] Scenario 3: After the sender sends one or more write messages without immediate values, it then sends a write message with an immediate value.
[0064] The network interface card (NIC) has received at least one write message without an immediate value before receiving a write message carrying an immediate value. In response to receiving a write message carrying an immediate value, the NIC initiates a read operation using the destination address carried in that write message as the specified address.
[0065] The specific workflow is as follows: The network card receives packets, parses the packet header, and determines whether it carries an immediate value. For write packets that do not carry immediate values, the network card performs a DMA write operation, writing the data in the write packet to the storage location corresponding to the target address in the target device, without triggering a read operation; For a write packet carrying an immediate value, the network card first initiates a DMA write operation to write the data in the write packet to the storage location corresponding to the target address in the target device; then, after initiating the above write operation, the network card actively initiates a DMA read operation using the target address carried in the write packet as the specified address; after the read operation is completed, the network card generates a CQE and writes it to the host memory.
[0066] Since the read operation is initiated after the write operation, and the transaction ordering rules of the high-speed bus guarantee that the read operation is executed after the write operation, the completion of this read operation not only indicates that the data of the current write message has been successfully written to the target device, but also that the data of all previous write messages without immediate values has been successfully written to the target device. The generated CQE is used to notify the host that the above data has been reliably stored.
[0067] In Scenario 3, the write message carrying an immediate value serves a dual role of "data transmission" and "triggering notification." Therefore, there's no need to send a separate notification message; the last write message completes both data transmission and trigger confirmation simultaneously, reducing the number of messages. Furthermore, there's no need to wait for a separate notification message; read verification is triggered immediately after data writing, further reducing end-to-end latency. This provides a flexible processing method between Scenario 1 and Scenario 2, selectable according to application requirements.
[0068] The data transmission synchronization processing method provided in this application can also be applied to scenario four, namely the data request-response scenario (Dump WQE).
[0069] In existing data request-response scenarios (Dump WQE), the following method is typically used to confirm data persistence: After receiving the data packet (i.e., write packet) returned by the peer, the network card (NIC) acting as the data requester writes the data to the target device via DMA; the NIC generates a write operation CQE1 and writes it to the host memory; after receiving the CQE1, the host CPU issues a DMADump WQE command to the NIC; the NIC performs a DMA read operation, reads the data from the target device at the specified address and generates a CQE2 corresponding to the Dump WQE, and writes the CQE2 to the host memory; after the host CPU checks the CQE2 of the Dump WQE, it confirms that the data has been stored in the target device.
[0070] The above processing method requires the host-side CPU to actively issue a Dump WQE instruction to trigger a read operation for verification after receiving the CQE1 of the write operation, which increases CPU overhead and software-hardware interaction latency.
[0071] To address the aforementioned issues, this application provides an improved data request-response scenario processing method. In Scenario 4, the network interface card (hereinafter referred to as "NIC A") provided in this application acts as the data requester, sending a data request to the peer NIC (hereinafter referred to as "NIC B") and receiving a data packet (i.e., a write packet) containing the requested data returned by the peer NIC.
[0072] The specific workflow is as follows: Step 4-1: Network card A sends a data request message to network card B.
[0073] Step 4-2: Network card B receives the data request message, prepares the requested data, and returns a write message to network card A, which carries the data and the destination address.
[0074] Step 4-3: Network card A receives write messages and extracts the target address and data from the messages.
[0075] Step 4-4: Network card A initiates a DMA write operation to write data to the storage location corresponding to the target address in the target device.
[0076] Steps 4-5: After initiating a DMA write operation, network card A immediately initiates a DMA read operation to read data from the target device at the specified address.
[0077] In Scenario 4, after each DMA write operation is initiated, the network card hardware automatically triggers a DMA read operation without requiring the host CPU to issue any verification instructions.
[0078] The specified address in steps 4-5 can be: the target address carried in the write message, or a pre-configured fixed address (such as the Flush Memory address).
[0079] Steps 4-6: After the DMA read operation is completed, NIC A generates a CQE and writes it to the host memory to notify the host-side application layer that the requested data has been successfully written to the target device.
[0080] Since the read operation is initiated after the write operation, and the transaction ordering rules of the high-speed bus ensure that the read operation is executed after the write operation, the completion of this read operation indicates that network card A has successfully written the received data to the target device.
[0081] In Scenario 4, after receiving the write message returned by the peer, NIC A actively initiates a DMA read operation through the NIC hardware to confirm that the data has been written to the target device. This eliminates the need for the host CPU to issue an additional Dump WQE instruction, reducing CPU overhead and response latency, and improving the processing efficiency of data requests.
[0082] Based on any of the above embodiments, this application also provides an implementation method in which the specified address for the read operation is a pre-configured fixed address (such as a Flush Memory address). This fixed address is different from the target address for writing data in the write packet; it is a dedicated memory area pre-allocated in the target device, the content of which has no actual business meaning and is only used to trigger read verification. When the network card initiates a read operation, it directly reads the data at this fixed address and uses the bus transaction sorting rules to confirm whether the previous write operation has been completed.
[0083] The workflow of this embodiment is as follows: The network card receives a write message and extracts the target address and data from the message; the network card initiates a DMA write operation to write the data to the storage location corresponding to the target address in the target device; after initiating the write operation, the network card actively initiates a DMA read operation using a pre-configured fixed address as the specified address to read the data at the fixed address in the target device; after the read operation is completed, the network card generates a CQE and writes it to the host memory.
[0084] Since read operations are initiated after write operations, and the high-speed bus employs a strict ordering mode to ensure that read operations execute after write operations, the completion of this read operation indicates that all previous write operations have been completed.
[0085] The fixed address scheme provided in this embodiment can be used as a general solution to replace the specified address source in any of the aforementioned embodiments. That is, in scenarios one, two, three, and four, the network card can use a pre-configured fixed address as the specified address for read operations, without needing to extract or dynamically record the target address from the packet. This method simplifies hardware design and is suitable for various packet types and transmission modes.
[0086] This embodiment uses a pre-configured fixed address as the designated address for read operations, eliminating the need to obtain or dynamically record the address from the message, thus simplifying hardware implementation. Regardless of the address carried by the write message, the read operation always accesses the same fixed address, making it suitable for various message types and highly versatile. Furthermore, the fixed address can be pre-configured by the system software and can be flexibly adjusted according to the hardware platform.
[0087] The data transmission synchronization processing method in this embodiment is executed by hardware logic in the network interface card (NIC). This hardware logic includes, but is not limited to, the following functional modules: DMA engine: Used to perform DMA write and DMA read operations. It can directly write received data to the target device and read data from a specified address on the target device.
[0088] Operation control logic: Used to control the timing of write and read operations, ensuring that a read operation is initiated after a write operation is initiated, and that read and write operations are executed in the order they are initiated.
[0089] The completion queue management logic is used to generate a CQE after a read operation is completed and write the CQE into the completion queue in the host memory.
[0090] Taking the PCIe bus as an example, the workflow of the network card hardware logic is as follows: After receiving a write packet, the DMA engine automatically initiates a DMA write operation to write the data to the target device; the operation control logic automatically initiates a DMA read operation after the write operation is initiated (without waiting for the host to issue instructions); after the read operation is completed, the queue management logic automatically generates a CQE and writes it to the host memory (without the host CPU polling). The entire process is completed autonomously by the network card hardware logic without the need for host-side participation, which is in stark contrast to the existing technology that relies on the host CPU to issue instructions.
[0091] In addition, the network interface card also integrates the following optional logic: Address recording logic: Used to record the target address carried in the write message in scenario 2, and use the recorded address as the specified address for the read operation after receiving the notification message.
[0092] Address overwrite logic: Used to implement address overwrite recording in scenario 2, recording the target address of each write message in the same address register, and the target address of the later received write message overwrites the target address of the earlier received write message.
[0093] By integrating data transmission synchronization processing methods into the network interface card (NIC) hardware logic, the entire data synchronization confirmation process is completed autonomously by the NIC hardware, without host intervention, thus freeing up host resources. Simultaneously, hardware logic executes much faster than software and eliminates the need for operating system scheduling and RDMA protocol stack processing, significantly reducing end-to-end latency. Furthermore, the execution timing of hardware logic is deterministic, eliminating the need for software scheduling and avoiding the impact of uncertainties such as operating system interrupts and task scheduling. This makes it suitable for scenarios with high real-time requirements, thereby achieving efficient synchronization between host notification status and device data status.
[0094] Based on the same inventive concept, this application also provides a network card. Since the principle of the network card in solving the problem is similar to that of a data transmission synchronization processing method, the implementation of the network card can refer to the implementation of the method, and the repeated parts will not be described again.
[0095] In some implementations, the network interface card provided in this application includes: The receiving module is used to receive write messages from the sending end; The DMA engine is used to perform write operations via Direct Memory Access (DMA) to write data in the write message to the target device; the storage space of the target device is not part of the host memory. The operation control logic is used to, after initiating the write operation, if a message carrying an immediate value is received, control the DMA engine to actively initiate a read operation to read data from the specified address in the target device; wherein, the read operation and the write operation are for the same target device, and the read operation and the write operation are executed sequentially according to the initiation order based on the bus transaction sorting rules; for write messages without immediate values, only the DMA engine is controlled to initiate the write operation, without initiating the read operation or generating a CQE; The queue management logic is used to generate a completion queue entry (CQE) and write the CQE to the host memory only after the read operation is completed; the CQE generation and writing to the host memory operations are not performed before the read operation is completed.
[0096] In some optional embodiments, the message carrying an immediate value is a write message carrying an immediate value, and the write message carries a target address; the operation control logic is used to control the DMA engine to initiate the read operation with the target address carried by the write message as the specified address.
[0097] In some alternative embodiments, the network interface card (NIC) has received at least one write message without an immediate value before receiving a write message carrying an immediate value.
[0098] In some optional embodiments, the write message does not carry an immediate value; the message carrying an immediate value is a notification message that carries an immediate value but does not carry a destination address or data; the network interface card further includes: Address recording logic is used to record the target address carried in the write message; The operation control logic is further configured to: in response to receiving the notification message, after all write operations prior to the notification message have been initiated, control the DMA engine to initiate the read operation using the recorded target address as the specified address.
[0099] In some optional embodiments, the address recording logic is specifically used to: record the target address of each write message into the same address register, and the target address of the later received write message overwrites the target address of the earlier received write message.
[0100] In some alternative embodiments, the designated address is a pre-configured fixed address.
[0101] In some alternative embodiments, the receiving module, the DMA engine, the operation control logic, and the completion queue management logic are implemented by hardware logic in the network interface card.
[0102] The network card provided in this application embodiment adopts the same inventive concept as the above-mentioned data transmission synchronization processing method and can achieve the same beneficial effects, so it will not be described again here.
[0103] The above embodiments are only used to provide a detailed description of the technical solutions of this application. However, the description of the above embodiments is only for the purpose of helping to understand the methods of the embodiments of this application and should not be construed as a limitation on the embodiments of this application. Any changes or substitutions that can be easily conceived by those skilled in the art should be covered within the protection scope of the embodiments of this application.
Claims
1. A data transmission synchronization processing method, characterized in that, The method is executed autonomously by the hardware logic in the network interface card, and the method includes: The network card receives a write message from the sender and performs a write operation via Direct Memory Access (DMA) to write the data in the write message to the target device; the storage space of the target device does not belong to the host memory. After initiating the write operation, if the network interface card (NIC) receives a packet carrying an immediate value, it will proactively initiate a read operation via DMA to read data from the specified address of the target device. The read operation and the write operation target the same device, and based on the bus transaction ordering rules, the read operation and the write operation are executed sequentially in the order they are initiated. For write packets without an immediate value, the NIC only initiates the write operation, does not initiate the read operation, and does not generate a completion queue entry (CQE). The network interface card (NIC) generates a CQE and writes the CQE to the host memory only after the read operation is completed; before the read operation is completed, the NIC does not perform the CQE generation and writing to the host memory operation.
2. The method according to claim 1, characterized in that, The message carrying an immediate value is a write message carrying an immediate value, and the write message carries a destination address; the network interface card initiates the read operation using the destination address carried in the write message as the designated address.
3. The method according to claim 2, characterized in that, The network interface card (NIC) has received at least one write message without an immediate value before receiving a write message carrying an immediate value.
4. The method according to claim 1, characterized in that, The write message does not carry an immediate value; the message carrying an immediate value is a notification message that carries an immediate value but does not carry a destination address or data; the method further includes: The network interface card records the target address carried in the write message; In response to receiving the notification message, after all write operations prior to the notification message have been initiated, the network interface card initiates the read operation using the recorded target address as the specified address.
5. The method according to claim 4, characterized in that, The method further includes: the network card records the target address of each write packet into the same address register, and the target address of the later received write packet overwrites the target address of the earlier received write packet.
6. The method according to claim 1, characterized in that, The specified address is a pre-configured fixed address.
7. A network interface card (NIC), characterized in that, include: The receiving module is used to receive write messages from the sending end; A DMA engine is used to perform write operations via direct memory access (DMA) to write data in the write message to the target device. The storage space of the target device does not belong to the host memory; The operation control logic is used to, after initiating the write operation, if a message carrying an immediate value is received, control the DMA engine to actively initiate a read operation to read data from the specified address in the target device; wherein, the read operation and the write operation are for the same target device, and the read operation and the write operation are executed sequentially according to the initiation order based on the bus transaction sorting rules; for write messages without immediate values, only the DMA engine is controlled to initiate the write operation, without initiating the read operation or generating a CQE; The queue management logic is used to generate a completion queue entry (CQE) and write the CQE to the host memory only after the read operation is completed; the CQE generation and writing to the host memory operations are not performed before the read operation is completed.
8. The network interface card according to claim 7, characterized in that, The message carrying an immediate value is a write message carrying an immediate value, and the write message carries a target address; the operation control logic is used to control the DMA engine to initiate the read operation with the target address carried in the write message as the specified address.
9. The network interface card according to claim 8, characterized in that, The network interface card (NIC) has received at least one write message without an immediate value before receiving a write message carrying an immediate value.
10. The network interface card according to claim 7, characterized in that, The write message does not carry an immediate value; the message carrying an immediate value is a notification message that carries an immediate value but does not carry a destination address or data; the network interface card also includes: Address recording logic is used to record the target address carried in the write message; The operation control logic is further configured to: in response to receiving the notification message, after all write operations prior to the notification message have been initiated, control the DMA engine to initiate the read operation using the recorded target address as the specified address.
Citation Information
Patent Citations
Data processing method and equipment
CN113688072A
Read-write operation optimization method and device of SLC NAND flash memory controller
CN118778885A