Processing method, device and system for sending data packet, storage medium and chip product

By improving the send queue submission ring entry structure of the Jnet protocol, adopting a 32B queue submission ring entry, and using the enable bit signal to indicate multiple data packets, the number of PCIE read operations is reduced, and the problem of limited small packet performance in the Jnet protocol is solved. The performance of pps and bps is significantly improved, and the efficiency of the driver software is improved.

CN120723686APending Publication Date: 2025-09-30BEIJING JAGUAR MICROSYSTEMS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510741156.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

The existing Jnet protocol has limited performance when sending data packets in small packets, with low pps and bps performance indicators, low driver software processing efficiency, and frequent PCIE read and write operations, resulting in the inability to further improve performance.

Method used

By improving the submission ring entry structure of the sending queue, adopting a 32B queue submission ring entry structure, using the enable bit signal to indicate multiple data packets to be sent, and obtaining the cache data through a single PCIE read operation, the number of read operations is reduced and the packet processing efficiency is improved.

Benefits of technology

It significantly improves the PPS and BPS performance of small packets, improves the processing efficiency of driver software, releases CPU resources, and improves PCIE bandwidth utilization, with performance improvements of 2.4 times and more than 5 times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723686A_ABST
    Figure CN120723686A_ABST
Patent Text Reader

Abstract

The invention provides a processing method for sending a data packet, which comprises the following steps of: receiving a notification sent by a host, and analyzing the notification to obtain an address of a sending queue entry; reading and analyzing the sending queue entry according to the address of the sending queue entry, and obtaining an enable bit signal of the sending queue entry, data initial address information of a to-be-sent data packet and data total length information of the to-be-sent data packet; acquiring data corresponding to the plurality of to-be-sent data packets in a cache region through a single PCIE read operation; and obtaining effective data of each to-be-sent data packet according to the data initial address information and the data total length information, and sending the effective data of each to-be-sent data packet. The invention further discloses a corresponding device and system, a storage medium and a chip product. By implementing the method and the device, the PCIE reading operation times required when the DPU reads the items and the cache data in the host can be reduced, and the path delay is reduced, so that the communication performance is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of communication between a host and a data processing unit (DPU), and in particular to a method, device, system, storage medium, and chip product for processing sent data packets. Background Art

[0002] With the rapid development of the cloud computing industry, virtualization services have been fully developed. However, with the development of high-computing services such as AI (Artificial Intelligence), equipment has put forward higher requirements for performance, latency, etc. There is an urgent need for a more efficient solution to improve the processing performance of network cards. Therefore, Yunbao Intelligent Company proposed the Jnet Protocol. The Jnet protocol is adopted on the host driver side of a series of smart network interface cards (SmartNetwork Interface Card, SmartNIC) and data processing unit (Data Processing Unit, DPU) products. The Jnet protocol can provide higher performance and lower latency.

[0003] Smart NICs and DPUs are back-end devices based on the Jnet protocol, interacting with front-end drivers through Jnet queues. NICs play a crucial role in the entire data path, and a high-performance Smart NIC or DPU can significantly impact a product's market success.

[0004] The performance of the network card is an indicator that needs to be paid attention to. It mainly includes two indicators: the number of packets processed per second (Packet Per Second, pps) and the port rate (Bits Per Second, bps). Among them: pps is one of the important indicators to measure the processing capacity of the network card and the performance of the network. It indicates the number of data packets that the network card can process or forward within the time range of one second. The data packet is the basic unit in network communication and contains data transmission information, such as source address, destination address, protocol type, etc. The pps capability of the network card directly affects the throughput and response speed of the network. A higher pps value means that the server can process data packets faster and provide higher network performance. Therefore, pps is an important indicator when designing and selecting a server, especially in scenarios where a large number of data packets need to be processed, such as network routers, firewalls, load balancers, etc.

[0005] The existing Jnet tx (send) direction performance has the following main shortcomings:

[0006] The Jnet transmission performance limit for small packets (64 bytes or less) is 250Mpps (Million Packets Per Second). Because the Jnet protocol's I / O data transmission for small packets, like large packets, requires multiple high-speed Peripheral Component Interconnect Express (PCIE) read and write operations, and PCIE bandwidth is a significant factor affecting performance, achieving higher performance is limited.

[0007] Due to the limitations of the Jnet protocol, the maximum pps is 250M. Therefore, the bps of the data packets to be sent in the form of small packets can only reach a maximum of 16Gbps (Gigabits Per Second). Compared with the 200Gbps large packet limit performance, it is already very low, affecting the bps performance index of the network card.

[0008] The Jnet driver software processes small packets of data to be sent in the same way as large packets. When preparing each packet to be sent, the driver needs to maintain a complete set of complex processes, but can only transmit data less than 64 bytes. This occupies a lot of central processing unit (CPU) resources but can only process a small amount of data. Therefore, the driver software is inefficient when processing small packets of data to be sent.

[0009] In summary, performance metrics are crucial for network cards. However, the current performance of small data packets is limited by the JNET protocol standard, reaching a maximum of only 250 Mpps. This impacts the bit-per-second (bps) and page-per-second (pps) performance of small data packets. Furthermore, the driver software's efficiency in processing small data packets is relatively low. Summary of the Invention

[0010] The technical problem to be solved by this application is to provide a processing method, device, system, storage medium and chip product for sending data packets, which can reduce the number of PCIE read operations when the DPU reads entries and cached data in the host, reduce path delays, and thus effectively improve communication performance.

[0011] To solve the above technical problems, as one aspect of the present application, a method for processing a sent data packet is provided, which is applied to a data processing unit and includes the following steps:

[0012] Receiving a notification sent by the host, parsing the notification to obtain the address of a sending queue entry;

[0013] Reading and parsing the transmit queue entry according to the address of the transmit queue entry to obtain an enable bit signal of the transmit queue entry, data starting address information of a data packet to be transmitted, and total data length information of the data packet to be transmitted, wherein the enable bit signal is used to indicate that the transmit queue entry includes information of a plurality of data packets to be transmitted;

[0014] Acquire data corresponding to the plurality of data packets to be sent in the buffer area through a single PCIE read operation;

[0015] The valid data of each data packet to be sent is obtained according to the data starting address information and the data total length information, and the valid data of each data packet to be sent is sent.

[0016] When the enable bit signal indicates that the sending queue entry includes information of a plurality of data packets to be sent, the sending queue entry further includes valid data length information corresponding to each data packet to be sent;

[0017] The step of splitting the valid data of each data packet to be sent according to the data starting address information and the data length information specifically includes:

[0018] The valid data of each data packet to be sent is obtained according to the cache line alignment principle and the data start address information, the total data length information, and the valid data length information corresponding to the data packet to be sent.

[0019] When the enable bit signal indicates that the sending queue entry includes only one packet to be sent, data corresponding to the packet to be sent in the buffer area is acquired through a PCIE read operation, and the data corresponding to the packet to be sent is sent.

[0020] Among them, the enable bit is located in a reserved bit in the front byte area. When the bit is 1, the enable bit signal indicates that the sending queue entry includes multiple data packet information to be sent, and the rear byte area is parsed as the length field of the data packet to be sent; when the bit is 0, the enable bit signal indicates that the sending queue entry includes only one data packet information to be sent.

[0021] As another aspect of the present application, a method for processing a data packet is further provided, which comprises the following steps:

[0022] The host places the data of multiple packets to be sent into the cache according to the cache alignment principle;

[0023] The host generates a sending queue entry according to the starting addresses of the data of the multiple data packets to be sent, the total length information of the data packets to be sent, the data length information of each data packet to be sent, and the enable bit information, and stores the sending queue entry in the sending queue cache unit, wherein the enable bit information is used to indicate that the sending queue entry includes information of multiple data packets to be sent;

[0024] The host generates a notification according to the address of the send queue entry in the send queue cache unit, and sends the notification to the data processing unit;

[0025] The data processing unit receives a notification sent by the host, and parses the notification to obtain an address of a sending queue entry;

[0026] The data processing unit reads and parses the sending queue entry according to the address of the sending queue entry to obtain the enable bit information of the sending queue entry, the data starting address information of the data packet to be sent, and the total length information of the data packet to be sent;

[0027] When the enable bit information indicates that the sending queue entry includes information of multiple data packets to be sent, the data processing unit obtains data corresponding to the multiple data packets to be sent in the buffer area through a single PCIE read operation;

[0028] The data processing unit obtains the valid data of each data packet to be sent according to the data starting address information and the data total length information, and sends the valid data of each data packet to be sent.

[0029] The host generates a sending queue entry according to the starting addresses of the data of the plurality of data packets to be sent, the total length information of the data packets to be sent, the data length information of each data packet to be sent, and the enable bit information, and stores the sending queue entry in the sending queue cache unit, further comprising:

[0030] The host generates a transmit queue entry, the entry comprising a front byte area and a back byte area; wherein the front byte area comprises a base address field, a total length field, a buffer identifier, and an enable bit; and the back byte area comprises a plurality of fixed-length length fields for lengths of packets to be transmitted;

[0031] The host stores multiple data packets to be sent in a continuous cache-aligned address space on the host side, fills the starting cache address of the first data packet to be sent into the base address field, fills the total data length into the total length field, fills the cache identifier field with the cache identifier corresponding to the cache-aligned address space, and sets the enable position to 1; and fills the length information of each data packet to be sent into each data packet length field to be sent in the rear byte area in sequence.

[0032] Among them, the enable bit is located in a reserved bit in the front byte area. When the bit is 1, the enable bit signal indicates that the sending queue entry includes multiple data packet information to be sent, and the rear byte area is parsed as the length field of the data packet to be sent; when the bit is 0, the enable bit signal indicates that the sending queue entry includes only one data packet information to be sent.

[0033] The host places the data of multiple data packets to be sent into the cache area according to the cache alignment principle, including:

[0034] When allocating the buffer area, the driver software adopts the cache line alignment principle to store multiple data packets to be sent in the continuous physical memory corresponding to the single buffer area identifier, and ensures that all data packets to be sent occupy a continuous buffer address segment;

[0035] The data processing unit obtains data corresponding to the plurality of data packets to be sent in the cache area through a single PCIE read operation, including: the data amount of the single PCIE read operation is the entire block of data after the total length of all data packets to be sent and the cache line are padded.

[0036] The data processing unit obtains valid data of each data packet to be sent according to the data starting address information and the data total length information, and sends the valid data of each data packet to be sent, further comprising:

[0037] The entire block of data read is calculated according to the length of each data packet to be sent and the cache line alignment principle, the starting address of each data packet to be sent is split into multiple data packets to be sent, and the data packets are sent to the corresponding target ports in sequence;

[0038] After the transmission is completed, the information of updating the completion queue is sent to the host and an interrupt is triggered to notify the host to release resources.

[0039] As another aspect of the present invention, there is also provided a processing device for sending data packets to be sent, which is applied to a data processing unit and includes:

[0040] A receiving module, configured to receive a notification sent by a host, and parse the notification to obtain an address of a sending queue entry;

[0041] a parsing module, configured to read and parse the send queue entry according to the address of the send queue entry, and obtain an enable bit signal of the send queue entry, data starting address information of a data packet to be sent, and total data length information of the data packet to be sent, wherein the enable bit signal is used to indicate that the send queue entry includes information of multiple data packets to be sent;

[0042] A data reading module is used to obtain data corresponding to the multiple data packets to be sent in the buffer area through a single PCIE read operation;

[0043] The splitting and sending module is used to obtain the valid data of each data packet to be sent according to the data starting address information and the data total length information, and send the valid data of each data packet to be sent.

[0044] Wherein, the parsing module further includes:

[0045] A length information parsing unit, used to identify the enable bit in the entry and the length field of the data packet to be sent after 16 bytes;

[0046] An address calculation unit, configured to generate a PCIE read request for a continuous data block according to a base address, the length of each data packet to be sent, and a cache line alignment rule, and obtain data corresponding to the multiple data packets to be sent in the cache area through a single PCIE read operation;

[0047] The splitting and sending module further includes:

[0048] The data splitting unit is used to obtain the valid data of each data packet to be sent according to the cache line alignment principle and the data starting address information, the total data length information, and the valid data length information corresponding to the data packet to be sent, and send it to the target port in sequence;

[0049] The sending completion trigger unit is used to send information about the update completion queue to the host after the sending is completed and trigger an interrupt to notify the host to release resources.

[0050] Among them, the enable bit is located in a reserved bit in the front byte area. When the bit is 1, the enable bit signal indicates that the sending queue entry includes multiple data packet information to be sent, and the rear byte area is parsed as the length field of the data packet to be sent; when the bit is 0, the enable bit signal indicates that the sending queue entry includes only one data packet information to be sent.

[0051] As another aspect of the present application, a processing system for sending a data packet is further provided, comprising:

[0052] A host, configured to place data of a plurality of to-be-sent data packets into a cache area according to a cache alignment principle; generate a send queue entry based on the starting addresses of the data of the plurality of to-be-sent data packets, the total length information of the data of the to-be-sent data packets, the data length information of each to-be-sent data packet, and enable bit information; and store the send queue entry in a send queue cache unit, wherein the enable bit information is used to indicate that the send queue entry includes information of a plurality of to-be-sent data packets; and generate a notification based on the address of the send queue entry in the send queue cache unit, and send the notification to the data processing unit;

[0053] A data processing unit is used to receive a notification sent by a host, obtain a send queue entry and detect an enable bit, and based on the content of the send queue entry, obtain all data packets to be sent in the cache area through a single PCIE read operation; and split the read data to obtain the data of each data packet to be sent and send it, wherein a processing device for sending data packets as described in any one of claims 10-12 is provided.

[0054] Wherein, the host further includes:

[0055] a sending queue entry generating unit, configured to generate a sending queue entry, wherein the entry includes a front byte area and a back byte area;

[0056] The front byte area includes a base address field, a total length field, a buffer identifier, and an enable bit; the back byte area includes a plurality of to-be-sent data packet length fields with fixed length bits, each length field storing the length of a to-be-sent data packet within a continuous address space;

[0057] a filling processing unit, configured to store data of a plurality of data packets to be sent in a continuous cache-aligned address space on the host side, fill the starting cache address of the first data packet to be sent into the base address field, fill the total data length into the total length field, fill the cache identifier field with the cache identifier corresponding to the cache-aligned address space, and set the enable bit to 1; and fill the length information of the data of each data packet to be sent into each length field of the data packet to be sent in the back-segment byte area in sequence;

[0058] a notification sending unit, configured to send a notification to the data processing unit, wherein the notification includes a queue identifier and an entry index;

[0059] The sending completion processing module is used to receive feedback from the data processing unit, update the entry in the completion ring of the queue and release the entry resources in the submission ring of the queue.

[0060] Accordingly, as another aspect of the present application, a computer-readable storage medium is provided, storing a computer program, which implements the steps of the aforementioned method when executed by a processor.

[0061] Accordingly, as another aspect of the present application, a chip product is also provided, in which the aforementioned system is deployed.

[0062] Implementing this embodiment has the following beneficial effects:

[0063] The present application provides a method, device, system, storage medium, and chip product for processing data packets. By improving the entry structure of the submission ring of the sending queue, the number of PCIE read operations required by the DPU when reading entries and cache data can be effectively reduced, reducing path latency, thereby significantly improving performance. This improvement has doubled the pps (packets per second) performance and bps (bits per second) performance of data packets to be sent in small packets, while significantly improving the efficiency of the driver software in processing data packets to be sent, freeing up more CPU resources to handle other tasks.

[0064] Furthermore, the solution provided by this application does not require changes to the chip's main architecture and maintains backward compatibility. It not only significantly reduces the number of read entries but also increases the amount of valid data returned by a single read operation, thereby improving PCIE bandwidth utilization. Furthermore, this solution eliminates the need for software to maintain multiple cache identifiers and fragmented address spaces. Instead, it uses a single, continuous block of space to transmit multiple data packets to be sent, further improving software efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, without paying any creative labor, obtaining other drawings based on these drawings still falls within the scope of the present application.

[0066] Figure 1 A schematic diagram of the data communication flow between the host and the data processing unit (DPU) involved in the operating environment of this application;

[0067] Figure 2 for Figure 1 Schematic diagram of the principle of the two ring structures used by the transmission queue;

[0068] Figure 3 Schematic diagram of the delay model involved in the transmission process;

[0069] Figure 4 Schematic diagram of the data structure of the queue submission ring using 16 bytes;

[0070] Figure 5 A schematic diagram of the main flow of an embodiment of a method for processing data packets sent provided by the present application;

[0071] Figure 6 for Figure 5 Schematic diagram of the meaning of each field in the first 16 bytes of the 32-byte queue submission ring;

[0072] Figure 7 for Figure 5 Schematic diagram of the meaning of each field in the last 16-byte area of ​​the 32-byte queue submission ring;

[0073] Figure 8 for Figure 5 Schematic diagram of the relationship between the data structure of the first 16-byte area of ​​the 32-byte queue submission ring and the buffer area;

[0074] Figure 9 for Figure 5 Schematic diagram of the relationship between the data structure of the 16-byte area in the back section of the 32-byte queue submission ring and the buffer area;

[0075] Figure 10 A schematic structural diagram of an embodiment of a processing system for sending data packets provided by the present application;

[0076] Figure 11 for Figure 10 Schematic diagram of the structure of the host;

[0077] Figure 12 for Figure 10 Schematic diagram of the structure of the data processing unit (DPU);

[0078] Figure 13 for Figure 12 Schematic diagram of the structure of the analysis module;

[0079] Figure 14 for Figure 12 Schematic diagram of the structure of the split sending module. DETAILED DESCRIPTION

[0080] In order to make the objectives, technical solutions and advantages of this application clearer, this application will be described in further detail below with reference to the accompanying drawings.

[0081] Before describing the detailed method of this application, first Figures 1 to 4 Describe the implementation environment involved in this application method.

[0082] like Figure 1As shown, a flow chart of data communication between the host (Host) and the data processing unit (DPU) involved in the operating environment of this application is shown; the data communication protocol adopted in this application is the Jnet protocol, and Jnet uses paired rings (Ring) to implement Jnet queues (Queue). Each Queue is a queue that carries a large amount of data. The specific number of queues used depends on business needs. In a specific example, the Jnet network driver generally uses multiple Tx Queues (multiple pairs of Tx-Rings) to implement network data transmission.

[0083] Data is sent between the host and the DPU through the Jnet-tx Sq-Ring and the Jnet-tx Cq-Ring of the Jnet send queue. The main steps are as follows:

[0084] Step 1: The host driver fills the data buffer information into the submission ring (tx Sq-Ring) of the transmit queue and updates the index corresponding to this request;

[0085] Step 2: The host sends a notification to the DPU, which includes the corresponding virtual function identifier (vf_id), queue identifier (queue_id), and index.

[0086] Step 3: The DPU parses the notification, maps it to the locally stored queue configuration, and initiates a request to read the entry in the transmit queue's submission ring (tx Sq-Ring). It obtains the address and length of the buffer corresponding to the data packet data to be sent in the form of a small packet, as well as some network packet configuration information.

[0087] Step 4: The DPU initiates a read data request, obtains the data from the host locally, and then sends it to the destination port;

[0088] Step 5: After the data is successfully sent, the DPU updates the completion ring (Tx Cq Ring) entry of the host-side transmit queue, including the buffer identifier (buffer ID), length (length), timestamp, and other information;

[0089] Step 6: The DPU sends an MSI-X interrupt to the Host to inform the Host that the data transmission is completed;

[0090] Step 7: The host updates its locally maintained completion ring index (Cqring index) and wrap counter (Wrap cnt) based on the uploaded entry information, parses the entry information of the send queue's commit ring (tx Sq-Ring), and releases the entry and the buffer information stored in it for next use.

[0091] like Figure 2 Figure 1 shows the structural relationship between the two rings of the current Jnet transmit queue. Each entry in the send queue's send queue ring (Sq Ring) stores information such as the buffer address (Addr8B), buffer length (Len4B), type (Type), and info for the data to be sent. The Sq Ring has a total of queue_size (Queue_Size) entries. After the buffer is filled, this information is added to the Sq Ring, the driver's index (DriverSqIndex) is updated, and the kick function is called to send a Notify message to the backend device. Upon receiving the request, the backend device initiates a DMA operation, reads the data packet to be sent to the hardware, and transmits it through the network port. It then updates the queue's completion queue ring (Cq Ring) with the processed Sq Ring entry information and notifies the frontend driver via an interrupt. Upon receiving the interrupt, the frontend driver updates its local wrap counter (Wrapcnt) and cq index (index) based on the Cq Ring information sent, freeing the corresponding buffer and entry (Sq Ring entry) to the free queue for future allocation. At this point, the data packet sending process is completed.

[0092] By reviewing the structure of the tx-sq-ring above, we can find that whether sending data packets in the form of large or small packets, they must go through the complete Jnet protocol process: updating the provider index, sending the doorbell message, obtaining the submission ring entry (sq entry) data, parsing the submission ring entry (sq entry) data, obtaining the buffer data, sending the completion ring entry (cq entry) data, sending the interrupt data, and parsing the completion ring entry (cq entry) data.

[0093] Looking at the entire transmission process, the TX direction delay model is as follows Figure 3 As shown:

[0094] T1: Host notification delay, assuming the typical PCIE latency is 1000ns;

[0095] T2: DPU single read submission ring entry descriptor delay, the delay value is 1000ns;

[0096] T3: The delay of the DPU reading a single TX packet to be sent. The delay value is 1000ns.

[0097] T4: The delay between the DPU processing the TX data and sending it to the MAC port. The delay value is the delay inside the DPU.

[0098] t0: The scheduling delay of the DPU internal queue, which is an uncertain delay;

[0099] t1: The uncertain delay between T2 and T3. This delay varies depending on the structure of the commit ring entries, such as whether the entry header type is 16 bytes, 32 bytes, or 48 bytes, as well as the I / O chain length. For example, for a chain-structured entry descriptor, the PSUB must cache all commit ring entries before initiating a read request.

[0100] From the above analysis, we can see that the total time for processing an I / O data entry is: T1 + T2 + T3 + T4 + t0 + t1. The idea behind the technical solution of this application is to reduce the total data processing time, thereby improving data transmission performance. T1 can be achieved by software reducing the number of notify messages sent. T4, t0, and t1 are the internal processing delays of the DPU. High-performance DPUs can already achieve very low delays, so improvements are limited. T2 and T3 are the time it takes to read entries and actual data in the cache. They account for a large proportion of the time in the I / O processing process, especially when processing multiple data entries. If T2 and T3 can be effectively reduced, the benefits will be enormous. Reducing T2 delay mainly reduces the number of times the DPU reads entries. That is, when the driver software prepares a single entry, it needs to carry information about multiple I / O entries. When the DPU reads such an entry, it will read more cache data, so that each read request returns more valid data, which can significantly reduce the number of PCIE reads. Therefore, T3 time can be effectively reduced.

[0101] From the above analysis, we can know that for the processing of small packets to be sent, the number of PCIE reads of the DPU can be reduced by increasing the number of buffers carried by a single entry, thereby reducing the processing delay of multiple I / Os and improving performance. By analyzing the 16-Byte entry of Jnet, it is found that the reserved fields are only as follows Figure 4As shown in the gray area, there are only 6 bits. A 64-Byte length requires at least 7 bits, which is obviously not enough. Therefore, this application will subsequently adopt a 32B queue submission ring entry structure to realize the transmission of multiple message information carried by a single entry. The first 16Bytes reuse the original field structure, and the last 16Bytes adopt a brand new structure.

[0102] like Figure 5 FIG. 1 is a schematic diagram showing a main flow of an embodiment of a method for processing a data packet sent according to the present application; and FIG. Figures 6 to 8 As shown, in this embodiment, the method includes at least the following steps:

[0103] In step S10, the host places the data of multiple data packets to be sent into the cache according to the principle of cache alignment, generates a send queue entry, fills in parameter information according to the packet data to be sent, and sends a notification to the data processing unit (DPU); wherein the send queue entry includes at least a front byte area and a back byte area; the front byte area is filled with the base address, total length, cache identifier, and enable bit information, and the back byte area is sequentially filled with the length information of each packet data to be sent;

[0104] In a specific example, the step S10 further includes:

[0105] In step S100, the host stores multiple small packet data to be sent in a continuous cache alignment address space on the host side according to the cache alignment principle. When allocating the cache area, the driver software stores multiple small packet data in a continuous physical memory corresponding to a single cache area identifier, and ensures that all small packet data occupy a continuous buffer address segment.

[0106] Specifically, the driver software needs to ensure that the transmitted data packets are in a continuous address space and are cache line aligned. Only in this way, the starting address of the next buffer can be calculated based on the length and cache line alignment principle, just by knowing the address of the previous buffer. Figure 8 The lower half of the .

[0107] Step S101: The host generates a transmit queue entry, wherein the entry includes a front byte area and a back byte area, both of which are 16 bytes; wherein the front byte area includes a base address field, a total length field, a buffer identifier, and an enable bit; and the back byte area includes multiple fixed-length packet length fields.

[0108] Among them, the enable bit is located in a reserved bit in the front byte area, which is used to indicate whether the small packet performance enhancement feature is turned on; when the bit is 1, the enable bit signal indicates that the sending queue entry includes multiple data packet information to be sent, and the rear byte area is parsed as the length field of the data packet to be sent; when the bit is 0, the enable bit signal indicates that the sending queue entry includes only one data packet information to be sent, and the original Jnet protocol format is used.

[0109] Specifically, if Figure 6 As shown in the figure, considering protocol compatibility, the improved 32-byte format uses bit 7 of the first 16 bytes of the descriptor as the feature enable bit, named tx_small_packet_perform_enhance_enable (abbreviated as tx_sp_perf_en). If this bit is 1, the feature is enabled, and the next 16 bytes adopt the new meaning. If this bit is 0, the feature is disabled, and the next 16 bytes retain the original meaning. This feature is an enable bit negotiated between software and hardware, and can be accurately determined for each I / O transaction, making it very flexible.

[0110] The format of the last 16-byte entry in the improved 32-byte format is as follows: Figure 7 As shown, it is used to indicate the length of each buffer to be transmitted. Each 6 bits represents the length of a small packet buffer, completing the one-time transmission of up to 21 small packets of data. The arrangement of its fields Figure 7 shown.

[0111] The meaning of each field is shown in Table 1 below.

[0112] Table 1

[0113]

[0114] Step S102, and fill the starting cache address of the first packet into the base address field, fill the total data length into the total length field, fill the cache identifier field with the cache area identifier corresponding to the cache alignment address space, and set the enable position to 1; fill the length information of each packet data into each packet length field in the latter byte area in sequence.

[0115] For details, please refer to Figure 8 as well as Figure 9 The first 16-byte structure of the improved 32-byte entry structure in this application is as shown in Figure 8This application requires the coordination of hardware and software. The driver software needs to ensure that the transmitted data packets are in a continuous address space and are cache_line aligned. Only in this way, the starting address of the next buffer can be calculated based on the length and cache_line alignment principle by knowing the address of the previous buffer. Put the starting address addr0 of the first packet as the starting address (addr_base) into Figure 6 Calculate the total length of all messages len_total=len0+len1+…+lenn, and fill it into Figure 6 The buf_len field indicates the total length of all transmitted packets and is mainly used to calculate the total number of transmitted messages. The driver software needs to ensure that all transmitted packets occupy an entire buffer address segment, so only one buffer identifier (buffer_id) is used to identify a unique buffer space, which is convenient for software management and resource allocation and recycling. When allocating resources, the buffer_id will be filled in. Figure 6 The buf_id field. Figure 6 In the tx_sp_perf_en bit 7, it is set to 1, and the rest of the fields are the same as the original settings.

[0116] The last 16 bytes of the improved 32-byte entry structure in this application are as follows: Figure 9 The gray part is shown. It is mainly used to identify the buffer length of each small packet. This design aims to improve the transmission performance of small packets. Small packets generally refer to messages with a length of less than 64 bytes. Therefore, each small packet only needs 6 bits to represent the length of a buffer. 0 represents a packet length of 1 byte, 1 represents a packet length of 2 bytes, and so on. 63 represents a packet length of 64 bytes. A length of 16 bytes (i.e., 128 bits = 16 bytes * 8 bits / byte) can represent up to 21 6-bits. In other words, a 16-byte data structure can carry the length information of up to 21 small packets. The total number of messages can be calculated by subtracting the length of each buffer in the last 16 bytes from the total length of the buffer in the first 16 bytes.

[0117] Step S103: The host sends a notification to the DPU, where the notification includes a queue identifier and an entry index.

[0118] Step S11: The DPU receives and parses the notification, obtains a send queue entry, detects the enable bit, obtains buffer address information based on the send queue entry content, and obtains the packet data to be sent in the buffer through a single PCIE read operation.

[0119] Specifically, the step S11 further includes:

[0120] In step S110, the DPU receives and parses the notification, obtains a send queue entry, and identifies the enable bit and the packet length field in the last 16 bytes of the entry. Specifically, the DPU reads and parses the send queue entry according to the address of the send queue entry to obtain the enable bit signal of the send queue entry, the data starting address information of the data packet to be sent, and the total data length information of the data packet to be sent. The enable bit signal is used to indicate that the send queue entry includes information about multiple data packets to be sent.

[0121] Step S112: Generate a PCIE read request for a continuous data block based on the base address, the length of each packet, and the cache line alignment rule, and obtain all packet data in the continuous address space through a single PCIE read operation;

[0122] The data volume of a single PCIE read operation is the total length of all small packets and the entire block of data after the cache line is padded.

[0123] It can be understood that in the solution of the present invention, when the enable bit signal is 0, indicating that the sending queue entry only includes information of one data packet to be sent, the data corresponding to the data packet to be sent in the cache area is obtained through a PCIE read operation, and the data corresponding to the data packet to be sent is sent.

[0124] In step S12, the DPU obtains the valid data of each data packet to be sent based on the data starting address information and the total data length information, and sends the valid data of each data packet to be sent. Specifically, the DPU splits the read data according to the content of the sending queue entry, obtains the small packet data to be sent, and sends them.

[0125] Specifically, the step S12 further includes the following steps:

[0126] Step S120 , the read entire block of data is calculated based on the length of each packet and the cache line alignment principle, the starting address of each packet is split into independent packets, the valid data in each packet is obtained, and the data is sent to the target port in sequence;

[0127] It is understandable that Figure 9 This shows how multiple message lengths are mapped into the last 16-byte entry. Note that cache_line alignment must be ensured. For example, len0 only indicates the length of buffer0 and does not include the cache_line padding.

[0128] Therefore, when the DPU calculates the starting address of the next buffer1, it needs to add the length of cache_line. Generally, cache_line alignment is based on 64-byte alignment, so the calculated starting address of buffer1 is {addr_hi, addr_lo}+len0+cache_line. When processing this entry data, in order to improve PCIE bandwidth utilization, the DPU will obtain buffer data of multiple small packets through a PCIE read. The DPU will then split the data into smaller packets based on the length of each packet and parse them out through the Ethernet interface. It can be found that after the improvement, the data that originally required n reads can be read and sent through a single PCIE read, which can effectively reduce the T3 time and improve performance.

[0129] In step S1210, after the DPU completes splitting and sending multiple small packets, it updates the completion queue and triggers an interrupt to notify the host to release resources.

[0130] It is understandable that this application effectively solves the performance issues of the current Jnet protocol in sending small packets. By sorting out the delays of the current Jnet tx in each stage of small packet transmission, it proposes an entry structure that improves the performance of Jnet-tx small packets. This solution can effectively reduce the number of PCIE read operations when the DPU reads entry and buffer data, reducing path delays, thereby effectively improving performance. Compared with projects that do not use this method: the pps performance of small packets can reach a maximum of 600Mpps, with a performance improvement of 2.4 times; the bps has also been greatly improved, reaching over 100G, with a performance improvement of more than 5 times; at the same time, it can effectively improve the efficiency of the driver software in processing small packets, freeing up more CPU resources to handle other tasks.

[0131] This application addresses the low performance of small packets in the tx direction in the current Jnet protocol. By analyzing the inefficiencies in the entire transmission process, we found that frequent PCIE read operations are required when processing small packets, just to process data smaller than 64 bytes, resulting in very low software and hardware efficiency. This application can reduce the number of PCIE reads when processing small packets in the tx direction and increase the amount of data returned by a single PCIE read.

[0132] This solution requires no changes to the chip's underlying architecture and is backward compatible. It significantly reduces the number of entry reads while increasing the amount of valid data returned by a single read, improving PCIE bandwidth utilization. This method has been proven to exponentially improve the performance of small packets in the send direction. Furthermore, it eliminates the need for software to maintain multiple cache identifiers and fragmented address spaces. Multiple small packets can be transmitted using a single, continuous block of space, making it software-friendly and improving efficiency.

[0133] like Figure 10 FIG. 1 is a block diagram showing an embodiment of a processing system for sending data packets provided by the present application. Figures 11 to 14 As shown, in this embodiment, the system at least includes:

[0134] Host 1 is configured to place data of multiple to-be-sent data packets into a cache area according to a cache alignment principle; generate a send queue entry based on the starting addresses of the data of the multiple to-be-sent data packets, the total length information of the data of the to-be-sent data packets, the data length information of each to-be-sent data packet, and enable bit information; and store the send queue entry in a send queue cache unit, wherein the enable bit information is used to indicate that the send queue entry includes information of multiple to-be-sent data packets; generate a notification based on the address of the send queue entry in the send queue cache unit, and send the notification to the data processing unit (DPU);

[0135] Data processing unit (DPU) 2 is used to receive notifications sent by the host, obtain the send queue entry and detect the enable bit, and based on the content of the send queue entry, obtain the small packet data to be sent in the cache area through a single PCIE read operation; and split the read small packet data in sequence and send them to the target port.

[0136] More specifically, if Figure 11 As shown, in a specific example, the host 1 further includes:

[0137] The sending queue entry generating unit 10 is used to generate a sending queue entry, wherein the front byte area and the back byte area included in the entry are both 16 bytes;

[0138] The front byte area includes a base address field, a total length field, a buffer identifier, and an enable bit; the back byte area includes multiple packet length fields with fixed length bits, each of which stores the length of packet data within a continuous address space; the enable bit is located in a reserved bit in the front byte area and is used to indicate whether the packet performance enhancement feature is enabled; when this bit is 1, the back byte area is interpreted as a packet length field; otherwise, the original Jnet protocol format is used.

[0139] The filling processing unit 11 is configured to store the data of multiple packets to be sent in a continuous cache-aligned address space on the host side, fill the starting cache address of the first packet into the base address field, fill the total data length into the total length field, fill the cache identifier corresponding to the cache-aligned address space into the cache identifier field, set the enable bit to 1; and fill the length information of each packet data into each packet length field in the back byte area in sequence;

[0140] a notification sending unit 12, configured to send a notification to the DPU, wherein the notification includes a queue identifier and an entry index;

[0141] The sending completion processing module 13 is configured to receive feedback from the DPU, update entries in the completion ring of the queue, and release entry resources in the submission ring of the queue.

[0142] It is understandable that the cache area maintained by the host driver software is a continuous physical memory area, and multiple small packet data is managed through a single cache area identifier.

[0143] Specifically, if Figure 12 As shown, in a specific example, the data processing unit 2 further includes:

[0144] The receiving module 20 is configured to receive a notification sent by the host and parse the notification to obtain an address of a sending queue entry;

[0145] a parsing module 21 configured to read and parse the send queue entry according to the address of the send queue entry, and obtain an enable bit signal of the send queue entry, data starting address information of the data packet to be sent, and total data length information of the data packet to be sent; wherein the enable bit signal is used to indicate that the send queue entry includes multiple data packets to be sent; when the enable bit is 1, the enable bit signal indicates that the send queue entry includes multiple data packets to be sent, and the subsequent byte area is parsed as a length field of the data packet to be sent; when the enable bit is 0, the enable bit signal indicates that the send queue entry includes only one data packet to be sent;

[0146] The data reading module 22 is configured to obtain data corresponding to the plurality of data packets to be sent in the buffer area through a single PCIE read operation;

[0147] The splitting and sending module 23 is used to split the read data according to the data starting address information and the data total length information, obtain the valid data of each data packet (packet) to be sent, and send each data packet to be sent.

[0148] More specifically, if Figure 13As shown, in a specific example, the parsing module 21 further includes:

[0149] The length information parsing unit 210 is used to identify the enable bit and the packet length field of the last 16 bytes in the entry;

[0150] The address calculation unit 211 is used to generate a PCIE read request for a continuous data block based on the base address, the length of each packet and the cache line alignment rule, and obtain all the packet data in the continuous address space through a single PCIE read operation; the data volume of the single PCIE read operation is the total length of all packets and the entire block of data after the cache line is padded.

[0151] More specifically, if Figure 14 As shown, in a specific example, the splitting and sending module 23 further includes:

[0152] The data splitting unit 230 is used to split the read data into independent packets according to the length of each packet and the cache line alignment principle, and send them to the target port in sequence;

[0153] The sending completion triggering unit 231 is used to send information about the update completion queue to the host after the sending is completed and trigger an interrupt to notify the host to release resources.

[0154] For more details, please refer to and combine the above Figures 1 to 9 The description is not repeated here.

[0155] Accordingly, as another aspect of the present application, a computer-readable storage medium is provided, which stores a computer program, and when the program is executed by a processor, the following is achieved: Figures 1 to 9 For more details, please refer to and combine the above Figures 1 to 9 The description is not repeated here.

[0156] Accordingly, as another aspect of the present application, a chip product is provided, wherein the Figures 10 to 14 For more details, please refer to and combine the above Figures 10 to 14 The description is not repeated here.

[0157] Accordingly, as another aspect of the present application, a chip product is provided, wherein at least Figures 12 to 14 For more details of the data processing unit described, please refer to and combine the above Figures 12 to 14 The implementation of this embodiment has the following beneficial effects:

[0158] The present application provides a method, device, system, storage medium, and chip product for processing data packets. By improving the entry structure of the submission ring of the sending queue, the number of PCIE read operations required by the DPU when reading entries and cache data can be effectively reduced, reducing path latency, thereby significantly improving performance. This improvement has doubled the pps (packets per second) performance and bps (bits per second) performance of data packets to be sent in small packets, while significantly improving the efficiency of the driver software in processing data packets to be sent, freeing up more CPU resources to handle other tasks.

[0159] Furthermore, the solution provided by this application does not require changes to the main chip architecture and maintains backward compatibility. It not only significantly reduces the number of read entries, but also increases the amount of valid data returned by a single read operation, thereby improving PCIE bandwidth utilization. Furthermore, this solution eliminates the need for software to maintain multiple cache identifiers and fragmented address spaces. Instead, it uses a single continuous block of space to transmit multiple small packets of data to be sent, further improving software efficiency.

[0160] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0161] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0162] The above disclosure is only a preferred embodiment of the present application and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A method for processing a data packet, applied to a data processing unit, characterized in that: The following steps are involved: Receiving a notification sent by the host, parsing the notification to obtain the address of a sending queue entry; Reading and parsing the transmit queue entry according to the address of the transmit queue entry to obtain an enable bit signal of the transmit queue entry, data starting address information of a data packet to be transmitted, and total data length information of the data packet to be transmitted, wherein the enable bit signal is used to indicate that the transmit queue entry includes information of a plurality of data packets to be transmitted; Acquire data corresponding to the plurality of data packets to be sent in the buffer area through a single PCIE read operation; The valid data of each data packet to be sent is obtained according to the data starting address information and the data total length information, and the valid data of each data packet to be sent is sent.

2. The processing method according to claim 1, characterized in that When the enable bit signal indicates that the sending queue entry includes information of a plurality of data packets to be sent, the sending queue entry further includes valid data length information corresponding to each data packet to be sent; The step of splitting the valid data of each data packet to be sent according to the data starting address information and the data length information specifically includes: The valid data of each data packet to be sent is obtained according to the cache line alignment principle and the data start address information, the total data length information, and the valid data length information corresponding to the data packet to be sent.

3. The method according to claim 2, wherein: When the enable bit signal indicates that the sending queue entry includes only one packet of data to be sent, data corresponding to the packet to be sent in the buffer area is acquired through a PCIE read operation, and the data corresponding to the packet to be sent is sent.

4. The method according to claim 3, wherein The enable bit is located in a reserved bit in the front byte area. When the bit is 1, the enable bit signal indicates that the send queue entry includes multiple data packet information to be sent, and the rear byte area is parsed as the length field of the data packet to be sent; when the bit is 0, the enable bit signal indicates that the send queue entry includes only one data packet information to be sent.

5. A method for sending a data packet, characterized in that: The following steps are involved: The host places the data of multiple packets to be sent into the cache according to the cache alignment principle; The host generates a sending queue entry according to the starting addresses of the data of the multiple data packets to be sent, the total length information of the data packets to be sent, the data length information of each data packet to be sent, and the enable bit information, and stores the sending queue entry in the sending queue cache unit, wherein the enable bit information is used to indicate that the sending queue entry includes information of multiple data packets to be sent; The host generates a notification according to the address of the send queue entry in the send queue cache unit, and sends the notification to the data processing unit; The data processing unit receives a notification sent by the host, and parses the notification to obtain an address of a sending queue entry; The data processing unit reads and parses the sending queue entry according to the address of the sending queue entry to obtain the enable bit information of the sending queue entry, the data starting address information of the data packet to be sent, and the total length information of the data packet to be sent; When the enable bit information indicates that the sending queue entry includes information of multiple data packets to be sent, the data processing unit obtains data corresponding to the multiple data packets to be sent in the buffer area through a single PCIE read operation; The data processing unit obtains the valid data of each data packet to be sent according to the data starting address information and the data total length information, and sends the valid data of each data packet to be sent.

6. The method according to claim 5, wherein The host generates a sending queue entry according to the starting addresses of the data of the plurality of data packets to be sent, the total length information of the data packets to be sent, the data length information of each data packet to be sent, and the enable bit information, and stores the sending queue entry in the sending queue cache unit, further comprising: The host generates a transmit queue entry, the entry comprising a front byte area and a back byte area; wherein the front byte area comprises a base address field, a total length field, a buffer identifier, and an enable bit; and the back byte area comprises a plurality of fixed-length length fields for lengths of packets to be transmitted; The host fills the starting cache address of the first data packet to be sent into the base address field, fills the total data length into the total length field, fills the cache identifier field with the cache identifier corresponding to the cache alignment address space, and sets the enable position to 1; and fills the length information of each data packet to be sent into each data packet length field to be sent in the latter byte area in sequence.

7. The method according to claim 6, wherein The enable bit is located in a reserved bit in the front byte area. When the bit is 1, the enable bit signal indicates that the send queue entry includes multiple data packet information to be sent, and the rear byte area is parsed as the length field of the data packet to be sent; when the bit is 0, the enable bit signal indicates that the send queue entry includes only one data packet information to be sent.

8. The method according to claim 7, wherein The host places the data of multiple packets to be sent into the cache according to the cache alignment principle, including: When allocating the cache area, the driver software adopts the cache line alignment principle to store multiple data packets to be sent in the continuous physical memory corresponding to a single cache area identifier, and ensures that all data packets to be sent occupy a continuous buffer address segment.

9. The method according to claim 8, wherein The data processing unit obtains valid data of each data packet to be sent according to the data starting address information and the data total length information, and sends the valid data of each data packet to be sent, further comprising: The entire block of data read is calculated according to the length of each data packet to be sent and the cache line alignment principle, the starting address of each data packet to be sent is split into multiple data packets to be sent, and the data packets are sent to the corresponding target ports in sequence; After the transmission is completed, the information of updating the completion queue is sent to the host and an interrupt is triggered to notify the host to release resources.

10. A processing device for sending data packets, used in a data processing unit, characterized in that: include: A receiving module, configured to receive a notification sent by a host, and parse the notification to obtain an address of a sending queue entry; a parsing module, configured to read and parse the send queue entry according to the address of the send queue entry, and obtain an enable bit signal of the send queue entry, data starting address information of a data packet to be sent, and total data length information of the data packet to be sent, wherein the enable bit signal is used to indicate that the send queue entry includes information of multiple data packets to be sent; A data reading module is used to obtain data corresponding to the multiple data packets to be sent in the buffer area through a single PCIE read operation; The splitting and sending module is used to obtain the valid data of each data packet to be sent according to the data starting address information and the data total length information, and send the valid data of each data packet to be sent.

11. The device according to claim 10, wherein in: The parsing module further comprises: A length information parsing unit, used to identify the enable bit in the entry and the length field of the data packet to be sent after 16 bytes; An address calculation unit, configured to generate a PCIE read request for a continuous data block according to a base address, the length of each data packet to be sent, and a cache line alignment rule, and obtain data corresponding to the multiple data packets to be sent in the cache area through a single PCIE read operation; The splitting and sending module further includes: The data splitting unit is used to obtain the valid data of each data packet to be sent according to the cache line alignment principle and the data starting address information, the total data length information, and the valid data length information corresponding to the data packet to be sent, and send it to the target port in sequence; The sending completion trigger unit is used to send information about the update completion queue to the host after the sending is completed and trigger an interrupt to notify the host to release resources.

12. The device according to claim 11, wherein The enable bit is located in a reserved bit in the front byte area. When the bit is 1, the enable bit signal indicates that the send queue entry includes multiple data packet information to be sent, and the rear byte area is parsed as the length field of the data packet to be sent; when the bit is 0, the enable bit signal indicates that the send queue entry includes only one data packet information to be sent.

13. A processing system for sending data packets, characterized in that: include: The host is used to place data of multiple data packets to be sent into a buffer area according to the principle of buffer alignment; generating a send queue entry based on the starting addresses of the data of the plurality of data packets to be sent, the total length information of the data packets to be sent, the data length information of each data packet to be sent, and the enable bit information, and storing the send queue entry in a send queue cache unit; generating a notification based on the address of the send queue entry in the send queue cache unit, and sending the notification to the data processing unit; wherein the enable bit information is used to indicate that the send queue entry includes information of a plurality of data packets to be sent; A data processing unit is configured to receive a notification from a host, obtain a send queue entry and detect an enable bit, and based on the contents of the send queue entry, obtain all data packets to be sent in a cache area through a single PCIE read operation; and split the read data to obtain and send the data packets of each to be sent, wherein a data packet processing device for sending a data packet as described in any one of claims 10 to 12 is provided.

14. The system according to claim 13, wherein: The host further comprises: a sending queue entry generating unit, configured to generate a sending queue entry, wherein the entry includes a front byte area and a back byte area; The front byte area includes a base address field, a total length field, a buffer identifier, and an enable bit; the back byte area includes a plurality of to-be-sent data packet length fields with fixed length bits, each length field storing the length of a to-be-sent data packet within a continuous address space; a filling processing unit, configured to store data of a plurality of data packets to be sent in a continuous cache-aligned address space on the host side, fill the starting cache address of the first data packet to be sent into the base address field, fill the total data length into the total length field, fill the cache identifier field with the cache identifier corresponding to the cache-aligned address space, and set the enable bit to 1; and fill the length information of the data of each data packet to be sent into each length field of the data packet to be sent in the back-segment byte area in sequence; a notification sending unit, configured to send a notification to the data processing unit, wherein the notification includes a queue identifier and an entry index; The sending completion processing module is used to receive feedback from the data processing unit, update the entry in the completion ring of the queue and release the entry resources in the submission ring of the queue.

15. A computer-readable storage medium storing a computer program, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

16. A chip product, characterized in that: At least the device according to any one of claims 10 to 12 is deployed therein.