A virtualized network packet receiving method, chip, and network interface card based on virtio packed ring.
Patent Information
- Application Number
- CN202610895457.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-06-22
AI Technical Summary
然而,现有packed ring的接收路径中,硬件仍需依赖描述符中的next标志位遍历描述符链以确定每个报文的边界,导致硬件需要多次读取描述符并解析链表结构,增加了PCIe带宽消耗和硬件内部的缓存开销
本发明的基于virtio packed ring的虚拟化网络报文接收方法,与现有技术相比,首先,硬件通过预先探测每个描述符链的固定长度,在后续报文接收过程中无需读取和解析描述符的next标志位,仅需根据报文长度确定实际使用的描述符数量,并以描述符链长度为固定步长直接跳转到下一个报文起始地址,从而大幅减少了硬件对主机内存的描述符读取次数,节省了PCIe带宽;其次,硬件无需缓存整个描述符链的内容,仅需维护起始指针和计数器,显著降低了芯片内部的缓存资源占用;再次,硬件自动跳过当前链中未被实际使用的描述符,避免了软件干预,提升了接收处理的一致性;最后,整个方法兼容现有virtio packed ring格式,无需修改驱动与硬件的标准接口,具有良好的实用性和可扩展性。
Smart Images

Figure CN122420259B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network interface card (NIC) technology, and in particular to a virtualized network packet receiving method, chip, and NIC based on virtio packed ring. Background Technology
[0002] In virtualization scenarios, Virtio, as a standard I / O paravirtualization framework, improves cache utilization and PCIe transmission efficiency through its packed ring format by merging the descriptor table, available rings, and used rings into a single descriptor ring. However, in the existing packed ring receive path, the hardware still needs to rely on the next flag in the descriptor to traverse the descriptor chain to determine the boundary of each packet. This results in the hardware needing to read descriptors and parse the linked list structure multiple times, increasing PCIe bandwidth consumption and internal hardware cache overhead. Furthermore, when the number of descriptors actually used by a packet is less than the total length of the descriptor chain, the hardware cannot automatically skip unused descriptors, requiring software intervention or additional processing, further reducing receive efficiency. Therefore, there is an urgent need for a packet receive method that can simplify hardware logic, reduce PCIe interactions, and automatically identify packet boundaries. Summary of the Invention
[0003] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows: According to a first aspect of this application, a method for receiving virtualized network packets based on a virtio packed ring is provided, the method comprising the following steps: S100, the driver allocates a ring of descriptors with contiguous physical addresses in the host memory, each descriptor pointing to a data buffer of a preset size, the preset size being equal to the system's page size; S200: The hardware automatically detects the number of descriptors contained in a descriptor chain in each virtual queue, determines the detected number of descriptors as the descriptor chain length, and sends the descriptor chain length to the driver. S300, the driver determines the allocation of descriptor rings according to the descriptor chain length, such that each descriptor chain consists of a plurality of consecutive descriptors of a number equal to the length of the descriptor chain; S400 When the hardware receives an Ethernet packet, it starts from the beginning address of the currently available descriptor chain and sequentially uses the descriptors in the currently available descriptor chain for direct memory access writing until the end of the packet; wherein, the hardware determines the actual number of descriptors used according to the packet length, and the actual number of descriptors used is less than or equal to the length of the descriptor chain. S500: After the hardware completes the direct memory access write of the current message, it directly sets the starting descriptor address of the next message to the current starting address plus the product of the length of the descriptor chain and the fixed number of bytes occupied by each descriptor, thereby skipping the descriptors in the current chain that are not actually used. S600, the hardware traverses the descriptor ring with a fixed step size of the descriptor chain length to determine the starting position of each message in the descriptor ring.
[0004] According to another aspect of this application, a chip is also provided, including hardware logic circuitry for performing the method described in the first aspect; the hardware logic circuitry includes: The detection module is used to automatically read the descriptor ring and count the length of the descriptor chain, packed_desc_num. The direct memory access engine is used to read descriptors from host memory and write packet data to the corresponding data buffer according to a fixed step size of packed_desc_num×16. An address generator is used to calculate the starting address of the next descriptor chain: next_addr = current_addr + packed_desc_num × 16, and supports modulo operations on the ring size; The interrupt controller is used to notify the host driver when the probe is complete or the message reception is complete. The configuration register stores the cfg_pref_num and vq_init_done flags written by the driver, as well as the packed_desc_num result written by the hardware.
[0005] According to another aspect of this application, a network interface card is also provided, including the chip described in the second aspect, as well as an Ethernet physical layer interface and a PCIe interface.
[0006] The present invention has at least the following beneficial effects: The virtualized network packet receiving method based on virtio packed ring of this invention, compared with the prior art, firstly, by pre-detecting the fixed length of each descriptor chain, the hardware does not need to read and parse the next flag bit of the descriptor during subsequent packet reception. It only needs to determine the number of descriptors actually used based on the packet length and directly jump to the next packet start address with the descriptor chain length as a fixed step size, thereby significantly reducing the number of descriptor reads from the host memory and saving PCIe bandwidth. Secondly, the hardware does not need to cache the contents of the entire descriptor chain, but only needs to maintain the start pointer and counter, significantly reducing the cache resource occupation inside the chip. Thirdly, the hardware automatically skips descriptors that are not actually used in the current chain, avoiding software intervention and improving the consistency of reception processing. Finally, the entire method is compatible with the existing virtio packed ring format, does not require modification of the standard interface between the driver and the hardware, and has good practicality and scalability. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 A flowchart illustrating a virtualized network packet receiving method based on a virtio packed ring, provided in an embodiment of the present invention. Detailed Implementation
[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0010] It should be noted that, based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Furthermore, this device and / or practice the method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.
[0011] The following will refer to Figure 1The flowchart shown is a method for receiving virtualized network packets based on virtio packed rings, which introduces a method for receiving virtualized network packets based on virtio packed rings.
[0012] The virtualized network packet receiving method based on virtio packed ring may include the following steps: S100, the driver allocates a ring of physical address descriptors in the host memory, each descriptor pointing to a data buffer of a preset size, the preset size being equal to the system page size.
[0013] Furthermore, the system has a page size of 4KB, a maximum transmission unit (MTU) of 9600 bytes for the Ethernet packet, and a 12-byte network header information is appended to each packet, occupying a separate descriptor.
[0014] In this embodiment, the driver allocates a contiguous memory region in the host memory as a descriptor ring. Each descriptor occupies a fixed 16 bytes, and the total number of descriptors in the ring (i.e., the queue depth) is typically set to an integer multiple of the descriptor chain length, such as 256. Each descriptor contains an address field and a length field. The address field points to a separately allocated data buffer, the size of which is equal to the system's page size. In x86 architecture Linux systems, the page size is typically 4KB, so each data buffer is 4KB. The driver fills the physical addresses of these buffers into the addr field of each descriptor, fills the 4KB field into the len field, and sets the availability flag in the flags field, indicating that these descriptors are available for hardware use. The base address of the descriptor ring, desc_ring_pa, and the queue depth are written into the hardware's configuration register.
[0015] For example, assuming the descriptor chain length is 16 and the queue depth is 256 chains, the descriptor ring has a total of 256 × 16 = 4096 descriptors, occupying 4096 × 16 = 64KB of contiguous physical memory. The driver also allocates 4096 4KB data buffers, for a total memory of 16MB, with the address of each buffer filled with the corresponding descriptor.
[0016] The driver allocates a ring of descriptors with contiguous physical addresses, and each descriptor points to a data buffer of fixed page size, resulting in a linear arrangement of the descriptor ring in memory. Subsequent hardware can directly access any descriptor using a base address plus an offset, without address translation or linked list traversal. Simultaneously, the preset page size (e.g., 4KB) standardizes buffer granularity, simplifies hardware address calculation logic, and reduces hardware design complexity.
[0017] S200: The hardware automatically detects the number of descriptors contained in a descriptor chain in each virtual queue, determines the detected number of descriptors as the descriptor chain length, and sends the descriptor chain length to the driver.
[0018] Furthermore, step S200 includes the following steps: S210, each virtual queue maintains an initialization completion flag vq_init_done, with an initial value of 0.
[0019] In this embodiment, each virtual queue has a corresponding control register in the hardware, which contains a 1-bit vq_init_done flag. Upon system power-on or hardware reset, the vq_init_done flag for all queues is automatically cleared. The driver can also write this flag to 0 at runtime via PCIe configuration space or memory-mapped I / O to force the hardware to re-probe.
[0020] The hardware internally has a state machine specifically for handling queue initialization. This state machine checks the current queue's vq_init_done value every clock cycle. If it is 0, the probe process begins; if it is 1, the probe is skipped, and the system directly enters the normal message receiving state.
[0021] Using a separate `vq_init_done` flag, the hardware can distinguish the initialization state of each queue, supporting independent probing and resetting of different queues. The driver can dynamically adjust the probing behavior at runtime without restarting the entire device. This flag implementation requires only one register bit, resulting in minimal hardware overhead.
[0022] S220, when vq_init_done is 0, the hardware starts from the descriptor ring base address desc_ring_pa and reads descriptor data of length cfg_pref_num × 16 bytes; cfg_pref_num is the number of descriptors to be read in advance.
[0023] When the hardware detects that `vq_init_done=0`, it starts reading 16 bytes of data from the descriptor ring base address `desc_ring_pa` via the DMA engine, using the `cfg_pref_num` × 16 bytes of data. Here, `cfg_pref_num` is a value pre-configured by the driver, typically written to another configuration register in the hardware. During initialization, the driver sets a sufficiently large value (e.g., 32 or 64) based on the expected chain length to ensure that the read data contains at least one complete descriptor chain, and may even contain two complete chains for verification.
[0024] The hardware initiates a DMA read operation, reading data into an internal temporary buffer (such as a 512-byte FIFO or RAM). After the read operation is complete, the hardware notifies the parsing module that the data is ready.
[0025] It should be noted that, to avoid excessive data reading causing an oversized internal cache, the driver can set cfg_pref_num appropriately, for example, equal to the driver's known maximum possible chain length plus 4. For most scenarios, 32 is sufficient to cover commonly used chain lengths (16, 32, 64, etc.).
[0026] By pre-reading multiple descriptors at once, the hardware reduces the number of PCIe transactions and improves probing efficiency. The driver can flexibly configure the number of pre-reads, striking a balance between cache size and probing speed. The pre-read data is temporarily stored internally in the hardware, eliminating the need for subsequent access to host memory during parsing, thus saving bandwidth.
[0027] S230, the hardware parses the next flag bit in the descriptor read, and counts the number of descriptors contained in a complete descriptor chain, packed_desc_num.
[0028] In this embodiment, the hardware parsing module checks the descriptors in the internal buffer one by one. The least significant bit (bit 0) of the 14th byte of each descriptor (i.e., the low-order byte of the flags field) is the next flag bit. It is usually denoted as VRING_DESC_F_NEXT in the specification.
[0029] The parsing process is as follows: 1. Initialize the counters count=0 and chain_found=0.
[0030] 2. Starting with the first descriptor, increment count.
[0031] 3. Check the next bit of the current descriptor: if it is 1, it means that the current descriptor is not the end of the chain, continue to the next descriptor (count++); if it is 0, it means that the end of the chain has been found, record the current count as the length of the first chain, chain_found=1.
[0032] 4. Then, starting from the next descriptor, repeat the above process to count the length of the second chain.
[0033] 5. If the lengths of the two chains are equal, then the length is confirmed to be packed_desc_num; if they are not equal, the third chain can be counted, or the length of the majority of the chains can be taken, or an error interruption can be triggered.
[0034] Furthermore, to ensure correctness, the hardware can also check whether other bits of the flags of the in-chain descriptor (such as the available / used flags) are reasonable, but the most basic is the next bit.
[0035] For example: Suppose the contents of the 32 descriptors read are as follows (only the next bit is listed): Descriptor 0: next=1; Descriptor 1: next=1; ... Descriptor 14: next=1; Descriptor 15: next=0 → End of the first chain, count=16; Descriptor 16: next=1; Descriptor 17: next=1; ... Descriptor 30: next=1; Descriptor 31: next=0 → End of the second chain, count=16; Both statistics show a value of 16, and the hardware confirms that packed_desc_num=16.
[0036] In this step, the hardware can determine the chain length by parsing the existing next flag in the descriptor, without relying on any external configuration. The method of counting two chains and comparing their lengths improves the reliability of the probe and avoids false positives caused by incomplete descriptor ring initialization. The entire process is completed internally in the hardware, is fast, and does not consume CPU resources.
[0037] S240 hardware saves packed_desc_num, sets vq_init_done=1, and triggers the interrupt notification driver.
[0038] The hardware writes the confirmed `packed_desc_num` to a read-only register for the driver to read. Then, the hardware sets the current queue's `vq_init_done` to 1, indicating that the probe is complete. Finally, the hardware sends an interrupt request to the interrupt controller; the interrupt vector typically corresponds to either the "probe complete" event or a queue-specific event.
[0039] After the driver's interrupt handler is triggered, it reads the packed_desc_num register of the queue to obtain the chain length. The driver can adjust the management method of subsequent descriptor rings according to this length, as described in step S300.
[0040] If the driver finds that the length is unreasonable (e.g., beyond the expected range), the driver can write vq_init_done to 0 again for the queue, and may modify cfg_pref_num or reinitialize the descriptor ring, and then wait for a new interrupt.
[0041] For example: the hardware detects that packed_desc_num = 16 and stores it in the register REG_PACKED_DESC_NUM (address offset 0x100). Then it sets vq_init_done = 1 and triggers interrupt number 42. After receiving the interrupt, the driver reads REG_PACKED_DESC_NUM and finds 16, confirming that it is valid. Subsequent reception will then be managed according to 16 descriptors per chain.
[0042] In this step, the hardware actively notifies the driver by saving the probe results and triggering an interrupt, eliminating the need for polling. This mechanism ensures synchronization between the hardware and software, allowing the driver to promptly obtain the chain length and perform corresponding configurations. The interrupt-driven approach is more efficient than polling, especially when multiple queues complete probes simultaneously; interrupt merging reduces processing overhead. Furthermore, the driver has the ability to reject probe results and trigger reprobing, increasing system robustness.
[0043] Furthermore, S200 also includes the following steps: S250, if the driver determines that the packed_desc_num obtained by the hardware detection is unreasonable, the driver will set the vq_init_done flag of the corresponding virtual queue to 0.
[0044] Upon receiving a hardware interrupt, the driver reads the value of `packed_desc_num` from the register. Internally, the driver has an expected reasonable range, such as between 1 and 64, or the driver has already negotiated the chain length used by the peer (e.g., 16) through virtio features. The driver then compares the read value with the expected value.
[0045] If any of the following conditions are found, the driver can consider the detection results to be unreasonable: packed_desc_num is 0 (obviously incorrect); packed_desc_num exceeds the maximum value supported by the hardware or software (e.g., greater than 64). The length is inconsistent with the expected length obtained by the driver through other means (for example, the driver knows from the virtual machine configuration that the descriptor ring is organized in 16-length, but the hardware probes a different value). The results of multiple consecutive probes are inconsistent (the driver can record previous results); When the determination is unreasonable, the driver writes 0 to the control register corresponding to the virtual queue, clearing the vq_init_done flag. This write operation is usually completed through PCIe MMIO (memory-mapped I / O), and the driver only needs to execute a register write instruction.
[0046] For example, the driver expects the chain length to be 16, but reads the register and finds packed_desc_num = 32. The driver considers 32 to be inconsistent with expectations, so it writes 0 to the control register at base address BAR0+0x08. Bit 0 of this register is the vq_init_done flag. After the write, the hardware immediately sees this flag become 0.
[0047] It should be noted that before writing 0, the driver may need to reconfigure cfg_pref_num (e.g., increase the prefetch count) or reinitialize the contents of the descriptor ring (e.g., refill the next bit of the descriptor) to ensure that the next probe will obtain the correct result. These preparatory actions are driver-driven and not hardware-mandated.
[0048] The driver has the final say on the probe results and can make flexible decisions based on upper-layer configurations or feature negotiations. If a probe anomaly is detected, the driver can proactively trigger a hardware re-probe without restarting the entire device or resetting PCIe functionality. This hardware-software collaborative fault-tolerance mechanism improves system reliability and adaptability, especially in virtualized environments where different virtual machines may use different configured virtio drivers; the hardware can automatically adapt, and the driver can intervene to correct the issue.
[0049] S260: After the hardware detects that vq_init_done has become 0, S210 is re-executed.
[0050] The control logic for each virtual queue within the hardware continuously monitors the value of the vq_init_done register. This register is typically a hardware register; once written by the driver, the hardware combinational or timing logic can detect the change immediately (or within a few clock cycles).
[0051] The hardware is typically designed as a finite state machine (FSM) to manage queue initialization. When vq_init_done changes from 1 to 0, the state machine transitions from the "initialized" state (e.g., IDLE state) to the "start probing" state, re-executing S210 and subsequent S220~S240. Specifically: the hardware clears the internally stored packed_desc_num value; resets the internal descriptor buffer pointer; restarts reading cfg_pref_num × 16 bytes of descriptor data from the descriptor ring base address desc_ring_pa; re-parses the next flag and counts the chain length. The result is then saved, vq_init_done is set to 1, and the interrupt is triggered again.
[0052] The hardware will not change the probe parameters due to a previous probe failure (e.g., cfg_pref_num remains unchanged unless the driver modifies this parameter simultaneously). If the driver wants to use a different prefetch count, it needs to update the cfg_pref_num register before or after setting it to 0.
[0053] For example, if the driver previously detected packed_desc_num = 32, which it considered unreasonable, it wrote 0 to vq_init_done. Upon hardware detection, it ignores the previously saved 32 and rereads the descriptor from desc_ring_pa. If the driver has now corrected the next flag in the descriptor ring (e.g., changing the next flag of the 16th descriptor in each chain to 0), the hardware will parse it again and get 16, thus returning the correct result.
[0054] It should be noted that, to prevent infinite retries, the driver typically gives up and reports an error after a certain number of retries. The hardware itself does not limit the number of retries; it will continue to try as long as vq_init_done is 0.
[0055] The hardware can respond instantly to driver reset commands, automatically returning to the beginning of the probing process without requiring a power-on reset. This design allows the hardware to attempt probing multiple times until the driver is satisfied. The hardware does not need to store complex error states internally; it only needs to check flag bits, resulting in simple and reliable logic. Simultaneously, the driver can dynamically adjust the descriptor ring layout (e.g., changing the chain length) without power interruption, allowing the hardware to re-probe and adapt to the new layout, enabling virtualization hot migration or dynamic configuration.
[0056] Furthermore, the actual number of descriptors required to receive a message is: 1 + ceil(MTU / page size); ceil() is the floor function; the length of the descriptor chain obtained by the hardware detection, packed_desc_num, is not less than 4.
[0057] Understandably, virtio-net requires that each received packet be prepended with a 12-byte network header (net_hdr) to store virtio-net control information (such as hash values, fragmentation status, etc.). This net_hdr occupies a separate descriptor and cannot be merged with the data into the same descriptor. Therefore, the total number of descriptors required to receive a complete packet = the number of descriptors occupied by the net_hdr + the number of descriptors needed to store the packet data.
[0058] In computer systems, buffer allocation must be in units of integer pages. Even if a message is only 1 byte, it will still occupy a full 4KB page (although in practice, virtio allows the len field of the descriptor to indicate the actual length used, but the buffer size of the descriptor is still PAGE_SIZE). Rounding up ensures that the number of allocated pages is sufficient to cover the entire message length and prevents overflow.
[0059] `packed_desc_num` is the total length of each descriptor chain detected by the hardware (i.e., how many fixed-size descriptors a chain contains). This length must be able to hold at least one complete maximum message; otherwise, the hardware cannot receive frames of MTU size.
[0060] A maximum packet (MTU=9600) actually requires 4 descriptors (1 net_hdr + 3 data). Therefore, the chain length packed_desc_num should be at least 4; otherwise, the chain cannot even hold a maximum packet.
[0061] In practical systems, packed_desc_num is often greater than 4 (e.g., 16 or 32) to support LRO (Large Packet Reassembly) or reduce software processing overhead. However, a value of not less than 4 is the minimum requirement, a necessary but not sufficient condition.
[0062] This step uses a mathematical formula to precisely express the minimum number of descriptors required for a standard maximum received message, and establishes a constraint relationship between this requirement and the chain length `packed_desc_num` obtained by hardware probing: `packed_desc_num` must not be less than this required value. This constraint ensures the basic functional integrity of the hardware receive path and is also one of the bases for judging whether the probing results are reasonable.
[0063] Furthermore, when the hardware supports the receiver merging LRO function, the descriptor chain length packed_desc_num is configured to be no less than 16, so that the total buffer size of each descriptor chain is no less than 64KB.
[0064] Furthermore, each descriptor occupies a fixed number of 16 bytes; when the hardware traverses the descriptor ring with a fixed step size of descriptor chain length, the step size is calculated as packed_desc_num × 16 bytes. The hardware directly generates the starting address of the next descriptor chain through the address generator, without needing to read the flags.next field of any descriptor in the current chain.
[0065] When the hardware enables LRO (Large Receive Offload) functionality, it merges multiple small packets into a large aggregate packet and submits it to the driver all at once. The length of the merged packet may far exceed the normal MTU (for example, multiple 1500-byte small packets can be merged into 64KB or even larger). Therefore, the hardware needs a larger descriptor chain to accommodate this aggregated packet.
[0066] The descriptor chain length packed_desc_num is configured to be no less than 16.
[0067] Each descriptor points to a 4KB data buffer, and the total buffer size of the chain = packed_desc_num × 4KB ≥ 16 × 4KB = 64KB. 64KB is a common upper limit for LRO aggregation (for example, the default maximum aggregation length of LRO in the Linux kernel is 64KB).
[0068] 16 descriptors × 4KB = 64KB, which perfectly meets the maximum aggregation length of most LRO implementations. If the upper limit of the aggregation length is higher, packed_desc_num can be increased accordingly, for example, 32 corresponds to 128KB.
[0069] This step ensures that the hardware can handle the giant frames generated by LRO, preventing packet loss or fragmentation due to insufficient descriptor chain capacity. The driver increases packed_desc_num through negotiation or configuration, and the hardware probe will also obtain a correspondingly large value, thus ensuring that each chain is long enough to leverage the performance advantages of LRO.
[0070] S300, the driver determines the allocation of descriptor rings based on the descriptor chain length, such that each descriptor chain consists of a plurality of consecutive descriptors of a number equal to the length of the descriptor chain.
[0071] In this embodiment, after the driver reads packed_desc_num from the hardware register, it needs to re-examine the previously allocated descriptor ring to ensure that the entire ring can be neatly divided into several consecutive chains, each containing packed_desc_num descriptors.
[0072] During initialization (S100), the driver allocates a contiguous ring of physical address descriptors, and the total number of descriptors is usually a multiple of a certain number (e.g., 4096). Now, after obtaining `packed_desc_num`, the driver first checks if the total number of descriptors is divisible by `packed_desc_num`. If it is not divisible, it means the initially allocated size does not match the current chain length, and the driver needs to reallocate the descriptor ring to make the total number of descriptors an integer multiple of `packed_desc_num`.
[0073] In this embodiment, during S100, the driver first allocates a relatively large descriptor ring according to a default value (e.g., 256), or allocates a sufficiently large ring (e.g., 4096 descriptors). After obtaining `packed_desc_num`, if 4096 is divisible by `packed_desc_num` (16, 32, 64, etc. are all divisible by 4096), no reallocation is needed; if it is not divisible (e.g., `packed_desc_num=12, 4096÷12 is not an integer), the driver needs to release the original ring and reallocate a new ring with a total number of `k×packed_desc_num`, where `k` is the queue depth (e.g., 256). This is done to ensure that each chain is arranged continuously within the ring and does not cross the ring boundary.
[0074] Internally, the driver divides the descriptor ring into groups based on `packed_desc_num`. For example, assuming `packed_desc_num` = 16, descriptors 0 to 15 belong to chain 0, descriptors 16 to 31 belong to chain 1, and so on. The driver maintains an array that records the starting descriptor index and offset within each chain.
[0075] Simultaneously, the driver needs to ensure that the descriptors within each chain are contiguous in physical address (since the ring itself is a contiguous array, this is naturally satisfied). The flags.next of the first descriptor in each chain should be set to 1 (unless the chain length is equal to 1), next=1 for the middle descriptors, and next=0 for the last descriptor. When initializing the descriptor ring, the driver needs to set the next flag bit of each descriptor according to this rule so that the hardware can correctly read the chain length during probing. The probing process itself will read these flag bits, and the driver only needs to fill them in according to the standard virtio specification.
[0076] The driver writes `packed_desc_num` to the hardware's step register. The hardware uses this value to calculate the starting address of the next chain when receiving subsequent packets. Some hardware implementations may also require the driver to write the total length of the descriptor ring (in terms of the number of descriptors) so that the hardware can perform a wraparound modulo operation.
[0077] For example: Suppose the driver initially allocates a ring of 4096 descriptors. The hardware detects that packed_desc_num = 16. 4096 ÷ 16 = 256, which is an integer. Internally, the driver creates 256 chains, each containing 16 consecutive descriptors. The driver sets the next flag of the first descriptor of each chain (0, 16, 32, etc.) to 1 (because the chain length is greater than 1), and sets the next flag of the last descriptor of each chain (15, 31, 47, etc.) to 0. Then, the driver writes the value 16 into the hardware's step register.
[0078] If the detected value is packed_desc_num=12, then 4096÷12≈341.33, which is not an integer. The driver needs to reallocate the descriptor ring, for example, allocating a total of 256×12=3072 descriptors (3072×16=49152 bytes). Then, it reorganizes the descriptors into chains of 12, fills in the next flag, and writes 12 into the step register.
[0079] Furthermore, if `packed_desc_num=1`, it means there is only one descriptor per chain. In this case, each descriptor's `next` flag should be set to 0, and the hardware will only use one descriptor to receive each packet. This approach is suitable for scenarios where the packet length does not exceed one page minus `net_hdr` (e.g., MTU ≤ 4084 bytes). The driver needs to manage this using a single descriptor chain.
[0080] In this step, the driver reconfirms the descriptor ring division based on the chain length obtained from hardware detection, ensuring that the software and hardware agree on "how many consecutive descriptors a chain contains." This forms the basis for subsequent fixed-step reception. By ensuring that the total number of ring descriptors is an integer multiple of the chain length, each packet is strictly aligned to the start of the chain, preventing packets from crossing chain boundaries. The driver's ability to actively adjust the descriptor ring size allows it to adapt to any valid length detected by the hardware, eliminating the need for a pre-set fixed chain length and enhancing versatility. Simultaneously, the driver fills in the next flag in the descriptor, maintaining compatibility with the standard virtio protocol, ensuring correct reading during hardware detection. This step translates the hardware detection results into the actual memory layout and organization structure, achieving parameter synchronization between the software and hardware.
[0081] S400: When the hardware receives an Ethernet packet, it starts from the beginning address of the currently available descriptor chain and sequentially uses the descriptors in the currently available descriptor chain for direct memory access writing until the end of the packet; wherein, the hardware determines the actual number of descriptors used based on the packet length, and the actual number of descriptors used is less than or equal to the length of the descriptor chain.
[0082] In this embodiment, when a complete packet is received on the Ethernet port, the hardware receive engine obtains the packet length information (usually provided by the MAC layer, in bytes). Internally, the hardware maintains two key states for each virtual queue: `chain_start_ptr` (the starting address of the currently available descriptor chain, pointing to the first descriptor in the chain) and `desc_index` (the number of descriptors currently in use within the chain, counting from 0). Before receiving a new packet, `desc_index` is 0.
[0083] The hardware obtains the total length of the packet from the MAC layer (including the Ethernet header, payload, FCS, etc., although the FCS is usually stripped). For virtio-net, the hardware needs to distinguish between two parts: a 12-byte virtio_net_hdr (placed at the beginning of the packet), and the actual network data. Therefore, from the hardware's perspective, the total length of data to be written is: 12 bytes (net_hdr) + packet payload length (i.e., MTU or actual frame length minus Ethernet header, etc.). For simplicity, the hardware can assume that each received packet first writes a 12-byte net_hdr (the value can be all 0s or filled in by the hardware), and then writes the remaining payload.
[0084] In practical applications, the hardware does not need to know the specific contents of net_hdr; it only needs to write it sequentially. The driver will parse it based on the fields in net_hdr.
[0085] Calculate the actual number of descriptors required: Assume the total data length to be written for a message is total_len bytes, including 12 bytes of net_hdr and data_len bytes of payload (total_len = 12 + data_len). Each descriptor points to a 4KB (4096 bytes) data buffer.
[0086] It's important to note that while the first descriptor stores the net_hdr, its buffer size remains 4KB. After the hardware writes 12 bytes to the first descriptor, the remaining space (4096-12=4084 bytes) is typically not used by the current packet in standard virtio designs (because net_hdr occupies a separate descriptor, which is the premise of the original scheme). Therefore, a strict distinction must be made: the first descriptor is used only for writing net_hdr, and subsequent descriptors are used only for writing payload data.
[0087] The number of descriptors required is: net_hdr descriptor: 1 Number of payload data descriptors = ceil(data_len / 4096); The total number of descriptors actually used is used_num = 1 + ceil(data_len / 4096).
[0088] The hardware knows `data_len` (i.e., the packet payload length), which can be obtained through an internal hardware counter or from the MAC layer. The hardware has an internal arithmetic logic unit for division or rounding up, or it can be implemented using a comparator: for example, checking if `data_len` exceeds 4096, 8192, etc., to obtain the required number of descriptors. Since `used_num` is not very large (usually less than or equal to 64), the hardware can use a simple comparator implementation.
[0089] Write data descriptor by descriptor: The hardware first reads the first descriptor pointed to by chain_start_ptr (address = chain start address) and obtains the addr field (buffer physical address). The hardware writes 12 bytes of net_hdr to this address. After writing, the hardware updates desc_index to 1.
[0090] The hardware then begins processing the payload data. It maintains a remaining data length, `remaining = data_len`. For the i-th data descriptor (i starts from 1, corresponding to the 2nd descriptor in the chain, with the actual index `desc_index`), the hardware calculates the address of the current descriptor: `desc_addr = chain_start_ptr + desc_index × 16`, and reads this descriptor to obtain the data buffer address. The hardware writes a maximum of 4096 bytes of data to this buffer (if the remaining data is greater than or equal to 4096 bytes, it fills the buffer; otherwise, it writes the remaining data). After writing, `remaining` is subtracted from the actual number of bytes written, and `desc_index` is incremented by 1. This process is repeated until `remaining` becomes 0.
[0091] Once all data has been written, the hardware no longer uses the descriptors that follow in the chain (even if `desc_index` might be less than `packed_desc_num`). The hardware records the actual number of descriptors used for the message, `used_num` (which is equal to `desc_index`). Then, the hardware updates the used ring or directly modifies the `used` bit in the descriptor's flags (according to the virtio packed ring specification) and updates the producer pointer. Finally, the hardware triggers a receive completion interrupt to notify the driver that a new message has arrived.
[0092] It's important to note that during the entire write process, the hardware never checks the next flag of the descriptor, nor does it rely on the next bit to jump to the next descriptor. Instead, it uses consecutive addresses (because the descriptor ring is a contiguous array, the address of the next descriptor is the current address + 16). Therefore, even if the next bit of a descriptor is 0, the hardware will still continue using the next descriptor because it uses the chain range given by `packed_desc_num`, not the next bit. This is how descriptor removal works.
[0093] For example: Assuming packed_desc_num=16, page size=4KB, and a jumbo frame of 9600 bytes is received (payload 9600 bytes, no VLAN, etc.), then total_len=12+9600=9612 bytes. The payload data requires ceil(9600 / 4096)=3 descriptors. Therefore, used_num=4.
[0094] The hardware's current chain start address is 0x10000000 (assumed). The hardware operates sequentially as follows: Read descriptor 0 (address 0x10000000), obtain buffer address buf0=0x20000000, and write 12 bytes of net_hdr.
[0095] Read descriptor 1 (address 0x10000010), obtain buf1=0x20001000, and write 4096 bytes of data (bytes 0-4095 of payload).
[0096] Read descriptor 2 (address 0x10000020), obtain buf2=0x20002000, and write 4096 bytes of data (bytes 4096-8191 of payload).
[0097] Read descriptor 3 (address 0x10000030), obtain buf3=0x20003000, and write the remaining 9600-8192=1408 bytes of data.
[0098] At this point, desc_index=4, and the message ends. The hardware no longer uses descriptors 4 to 15 (addresses 0x10000040 to 0x100000F0), nor does it read them.
[0099] If the received message is small, such as a 64-byte ARP packet with a payload of only 64 bytes, then used_num = 1 + ceil(64 / 4096) = 1 + 1 = 2. The hardware only uses the first two descriptors: the first writes 12 bytes of net_hdr, and the second writes 64 bytes of data (the remaining space in the buffer is unused). Then the remaining 14 descriptors are skipped.
[0100] If the message length is 0 (an abnormal situation), the hardware should only write net_hdr (12 bytes) and use a single descriptor.
[0101] In this embodiment, the hardware can operate in this way only if the driver has allocated descriptor rings according to `packed_desc_num` and each chain has enough descriptors. `used_num` is always less than or equal to `packed_desc_num`. If a received packet requires more descriptors than `packed_desc_num` (e.g., a packet that is too large after LRO merging, but not a sufficiently large chain is configured), the hardware should discard the packet and report an error. This is why `packed_desc_num` needs to be set relatively large (e.g., 16 corresponds to 64KB) when supporting LRO.
[0102] In this step, the hardware dynamically calculates the actual number of descriptors needed based on the message length and only uses the first few descriptors in the chain for DMA writing, without occupying the entire chain. This approach allows the hardware to efficiently utilize the data buffer, preventing small messages from wasting too many descriptors and cache space. Simultaneously, the hardware relies entirely on fixed-step address calculations to obtain the next descriptor, eliminating the need to parse the `next` flag in the descriptor, thus avoiding multiple PCIe read operations and internal state machine overhead caused by traversing the linked list. Since the hardware only reads the descriptors actually used in the current chain (typically only the first descriptor is needed to obtain the buffer address, and subsequent descriptor addresses are obtained through addition), the number of interactions between the hardware and host memory is significantly reduced, saving PCIe bandwidth. Furthermore, this method is insensitive to changes in message length, has a constant processing time, and is beneficial for achieving line-rate reception.
[0103] S500: After the hardware completes the direct memory access write of the current message, it directly sets the starting descriptor address of the next message to the current starting address plus the product of the length of the descriptor chain and the fixed number of bytes occupied by each descriptor, thereby skipping descriptors that are not actually used in the current chain.
[0104] After the current message is received, the hardware needs to prepare to receive the next message. In the traditional way, the hardware needs to traverse the next flag of the descriptor chain to find the end of the chain, but this invention uses a fixed step size jump.
[0105] The hardware internally maintains a `chain_start_ptr` pointer that currently points to the address of the first descriptor of the chain being used. The hardware directly increments this pointer by `packed_desc_num × 16` bytes to obtain the address of the starting descriptor of the next chain.
[0106] Furthermore, step S500 skips descriptors that are not actually used in the current chain, including the following steps: The S510 hardware internally maintains a chain start pointer chain_start_ptr and an intra-chain descriptor index desc_index for each virtual queue.
[0107] Internally, the hardware allocates a separate set of registers for each virtual queue, containing at least two key variables: chain_start_ptr: The starting address of the currently used descriptor chain (points to the first descriptor in the chain).
[0108] desc_index: A counter (starting from 0) for how many descriptors have been used in the current chain.
[0109] These registers are cleared upon power-on or queue reset. Once the hardware has completed probing (vq_init_done=1) and begins receiving packets normally, chain_start_ptr is initialized to the base address of the descriptor ring, and desc_index is 0.
[0110] For example: Suppose the descriptor ring base address is 0x10000000, and packed_desc_num=16. Initially, chain_start_ptr=0x10000000, desc_index=0. Each time the hardware receives a new packet, it starts using the descriptor from this position.
[0111] By maintaining independent start pointers and index counters for each queue, the hardware can precisely track the current packet's position in the descriptor ring without relying on any state in host memory. These two registers occupy very little space (typically 64 bits and 16 bits respectively), but provide the hardware foundation for subsequent fixed-step skipping mechanisms.
[0112] S520, when a new message is received, reset desc_index=0.
[0113] When the hardware detects the start of a new packet on the Ethernet port (usually indicated by the MAC layer's start-of-frame delimiter or receive-valid signal), and the previous packet has been processed, the hardware prepares to receive new packets. At this time, the hardware clears the desc_index register to zero. Meanwhile, chain_start_ptr remains unchanged, still pointing to the first descriptor of the current chain.
[0114] It's important to note that this step occurs before the hardware actually reads the first descriptor. If the hardware currently has no available descriptor chains (e.g., all chains are occupied by the driver), it may pause receiving or drop packets. In normal operation, the driver has already provided available descriptor chains, so `chain_start_ptr` is valid.
[0115] For example, after the previous message is processed, chain_start_ptr points to the starting address of chain 5, 0x10000100, and desc_index is reset to 0 after the previous message ends. Now, when a new message is received, the hardware sees that desc_index is already 0, and without any additional action, it directly starts using the first descriptor.
[0116] Resetting `desc_index` ensures that each new packet is written starting from the first descriptor of the current chain, preventing confusion with residual index values from previous packets. This operation is simple and has minimal hardware overhead, but effectively avoids data misalignment between packets.
[0117] For the S530, desc_index is incremented by 1 after each direct memory access write to a descriptor is completed.
[0118] Each time the hardware completes a DMA write to the data buffer pointed to by a descriptor, it immediately increments desc_index by 1. Writing to a descriptor means that the hardware has read the addr and len of that descriptor (or obtained directly through address calculation), written a portion of the packet data into the corresponding buffer, and is ready to process the next descriptor.
[0119] The value of `desc_index` indicates how many descriptors have been used in the current packet. For example, after writing the first descriptor (net_hdr), `desc_index` = 1; after writing the second descriptor (the first data page), `desc_index` = 2, and so on. Internally, the hardware can calculate the address of the next descriptor based on the value of `desc_index`: `next_desc_addr = chain_start_ptr + desc_index × 16`. Since the descriptor ring is a contiguous array, this calculation is very simple, requiring only an adder and a multiplier (or a shifter, since ×16 can be implemented by left shifting by 4 bits).
[0120] By using simple index increment and address calculation, the hardware can traverse the descriptors in the chain in fixed steps, without parsing the next bit or reading the flags field of each descriptor to determine the location of the next descriptor. This not only saves PCIe read bandwidth (because no additional descriptor reading is required), but also simplifies hardware design, allowing the receive engine to process data in a pipelined manner every clock cycle.
[0121] In S540, when the message ends, the hardware directly updates chain_start_ptr to chain_start_ptr + packed_desc_num × 16, takes the modulo of the length of the descriptor ring, and resets desc_index; it does not perform any read or write operations on the remaining descriptors whose index is greater than or equal to the number of descriptors actually used.
[0122] When the hardware has finished writing all the data in the message (i.e., the remaining data length is 0, and the last descriptor has been written), the hardware enters the message end processing flow. At this time: Update the chain start pointer: chain_start_ptr is assigned the value chain_start_ptr + packed_desc_num × 16. That is, skip the entire current chain (including all descriptors that are actually used and not used) and point to the first descriptor of the next chain.
[0123] Modulo operation on ring length: The descriptor ring is a circular buffer with a total number of bytes: ring_total_bytes = ring_desc_count × 16. The hardware needs to compare the updated chain_start_ptr with the ring boundary; if it exceeds the boundary, it subtracts ring_total_bytes (or uses an AND gate to take the modulo). Typically, the hardware maintains a ring mask, enabling fast modulo operation.
[0124] Reset the descriptor index: desc_index=0, in preparation for receiving the next message.
[0125] No operations are performed on the remaining descriptors: For descriptors in the current chain that are not actually used (indexes from used_num to packed_desc_num-1), the hardware neither reads nor modifies them. Their contents remain as is, and the driver will handle them when the entire chain is subsequently reclaimed. The hardware completely ignores their existence.
[0126] For example: Assuming `packed_desc_num=16`, the current chain start address is 0x10000000, and 4 descriptors are actually used. After the message ends, the hardware updates `chain_start_ptr` to 0x10000000 + 16 × 16 = 0x10000100. Assuming the total number of ring descriptors is 4096, the total number of ring bytes is 4096 × 16 = 65536. 0x10000100 is still within the range (base address 0x10000000 to 0x1000FFFF), so no modulo operation is needed. Then the hardware sets `desc_index=0`. The address range of descriptors 4 to 15, 0x10000040 to 0x100000F0, is no longer accessed by the hardware.
[0127] If the starting address of the next chain exceeds the end of the ring, for example, after the update it is 0x10010000, the ring ends at 0x1000FFFF, and the excess part needs to wrap around: subtract 65536 to get 0x10010000-0x10000=0x10000000, and return to the beginning of the ring.
[0128] This step implements the core idea of "descriptor removal": the hardware doesn't care which descriptors in the chain are actually used, nor does it care about the next flag of the descriptors; it directly jumps to the next packet start position using the complete chain length as the step size. The advantages of this are: regardless of the actual number of descriptors used, the hardware processing logic remains consistent (fixed step size), avoiding conditional branching; simultaneously, the hardware doesn't need to read the contents of unused descriptors, saving PCIe read transactions. This design enables the receive engine to process mixed-size packet streams with extremely high efficiency and is easy to implement in hardware pipelines.
[0129] Step S500 achieves automatic packet boundary alignment in the receive path by maintaining the chain start pointer and descriptor index internally in the hardware and employing a mechanism of "fixed-step jump, ignoring unused descriptors". Throughout the process, the hardware does not need to parse the next flag of any descriptor or read unused descriptors, thus significantly reducing the number of accesses to host memory (at least a reduction of packed_desc_num - used_num descriptor read operations per packet). Simultaneously, the hardware internal state requires only two registers (start pointer and index), resulting in minimal cache usage. This concise hardware model is not only easy to implement but also has constant processing time, contributing to line-rate forwarding performance. Furthermore, since the hardware completely skips unused descriptors, the driver can still recognize the integrity of the entire chain when recycling it, eliminating the need for additional synchronization overhead between the hardware and software.
[0130] S600, the hardware traverses the descriptor ring with a fixed step size of the descriptor chain length to determine the starting position of each message in the descriptor ring.
[0131] During subsequent message reception, the hardware consistently moves `chain_start_ptr` in fixed steps using `packed_desc_num`. This means that the starting descriptor position of each message within the descriptor ring is strictly determined: the starting address of the k-th message = descriptor ring base address + (k × packed_desc_num × 16) modulo the total number of bytes in the ring. The hardware does not need to read the `next` flag of any descriptor, nor does it need to parse the linked list structure of the descriptors, to know where each message starts and where the next message starts.
[0132] As far as the driver is concerned, it also knows packed_desc_num. Therefore, when the driver finishes processing the queue, it can traverse the descriptor ring with the same fixed step size and quickly find the descriptor chain corresponding to each packet without traversing the next chain.
[0133] This application pre-detects the fixed length of each descriptor chain in hardware and jumps directly to it in steps of this length during message reception, eliminating the need to read and parse the next flag of the descriptor. This significantly reduces descriptor read operations on host memory and saves PCIe bandwidth. Internally, the hardware does not need to cache the entire descriptor chain; it only needs to maintain a start pointer and a counter, reducing chip cache resource consumption. For messages whose actual usage length is less than the chain length, the hardware automatically skips unused descriptors without software intervention, improving the consistency of reception processing. The entire method is compatible with the existing virtio packedring format, and the driver and hardware interfaces do not require modification of the standard protocol, demonstrating good practicality and scalability. Real-world testing shows that in a scenario with an MTU of 9600 bytes, a page size of 4KB, and a chain length of 16, each message can save 15 descriptor read operations (a total of 240 bytes of PCIe reads), and the hardware cache requirement is reduced from 256 bytes to approximately 32 bytes.
[0134] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0135] Embodiments of the present invention also provide a chip, including hardware logic circuitry for performing the methods described in the above embodiments; the hardware logic circuitry includes: The detection module is used to automatically read the descriptor ring and count the length of the descriptor chain, packed_desc_num.
[0136] The direct memory access engine is used to read descriptors from host memory and write packet data to the corresponding data buffers according to a fixed step size of packed_desc_num×16.
[0137] An address generator is used to calculate the starting address of the next descriptor chain: next_addr = current_addr + packed_desc_num × 16, and supports modulo operations on the ring size.
[0138] The interrupt controller is used to notify the host driver when the probe is complete or the message reception is complete.
[0139] The configuration register stores the cfg_pref_num and vq_init_done flags written by the driver, as well as the packed_desc_num result written by the hardware.
[0140] An embodiment of the present invention also provides a network interface card (NIC), characterized in that it includes the chip in the above embodiments, as well as an Ethernet physical layer interface and a PCIe interface.
[0141] While specific embodiments of the invention have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention.
Claims
1. A method for receiving virtualized network packets based on a virtio packed ring, characterized in that, Includes the following steps: S100, the driver allocates a ring of descriptors with contiguous physical addresses in the host memory, each descriptor pointing to a data buffer of a preset size, the preset size being equal to the system's page size; S200: The hardware automatically detects the number of descriptors contained in a descriptor chain in each virtual queue, determines the detected number of descriptors as the descriptor chain length, and sends the descriptor chain length to the driver. S300, the driver determines the allocation of descriptor rings according to the descriptor chain length, such that each descriptor chain consists of a plurality of consecutive descriptors of a number equal to the length of the descriptor chain; S400 When the hardware receives an Ethernet packet, it starts from the beginning address of the currently available descriptor chain and sequentially uses the descriptors in the currently available descriptor chain for direct memory access writing until the end of the packet; wherein, the hardware determines the actual number of descriptors used according to the packet length, and the actual number of descriptors used is less than or equal to the length of the descriptor chain. S500: After the hardware completes the direct memory access write of the current message, it directly sets the starting descriptor address of the next message to the current starting address plus the product of the length of the descriptor chain and the fixed number of bytes occupied by each descriptor, thereby skipping the descriptors in the current chain that are not actually used. S600, the hardware traverses the descriptor ring with a fixed step size of the descriptor chain length to determine the starting position of each message in the descriptor ring.
2. The virtualized network packet receiving method based on virtio packed ring according to claim 1, characterized in that, S200 includes the following steps: S210, each virtual queue maintains an initialization completion flag vq_init_done, with an initial value of 0; S220, when vq_init_done is 0, the hardware starts from the descriptor ring base address desc_ring_pa and reads descriptor data of length cfg_pref_num × 16 bytes; cfg_pref_num is the number of descriptors to be read in advance; S230, the hardware parses the next flag bit in the descriptor read, and counts the number of descriptors contained in a complete descriptor chain: packed_desc_num; S240 hardware saves packed_desc_num, sets vq_init_done=1, and triggers the interrupt notification driver.
3. The virtualized network packet receiving method based on virtio packed ring according to claim 2, characterized in that, S200 also includes the following steps: S250, if the driver determines that the packed_desc_num obtained by the hardware detection is unreasonable, the driver will set the vq_init_done flag of the corresponding virtual queue to 0; S260: After the hardware detects that vq_init_done has become 0, S210 is re-executed.
4. The virtualized network packet receiving method based on virtio packed ring according to claim 1, characterized in that, Step S500 skips descriptors that are not actually used in the current chain, including the following steps: The S510 hardware internally maintains a chain start pointer chain_start_ptr and an in-chain descriptor index desc_index for each virtual queue. S520: When a new message is started being received, desc_index is reset to 0; For S530, desc_index is incremented by 1 after each direct memory access write to a descriptor is completed; In S540, when the message ends, the hardware directly updates chain_start_ptr to chain_start_ptr + packed_desc_num × 16, takes the modulo of the length of the descriptor ring, and resets desc_index; it does not perform any read or write operations on the remaining descriptors whose index is greater than or equal to the number of descriptors actually used.
5. The virtualized network packet receiving method based on virtio packed ring according to claim 1, characterized in that, The system has a page size of 4KB, the Ethernet packet's maximum transmission unit (MTU) is 9600 bytes, and each packet is preceded by a 12-byte network header that occupies a separate descriptor. The actual number of descriptors required to receive a message is: 1 + ceil(MTU / page size); ceil() is the floor function; the length of the descriptor chain obtained by the hardware detection, packed_desc_num, is not less than 4.
6. The virtualized network packet receiving method based on virtio packed ring according to claim 5, characterized in that, When the hardware supports the receiver merging LRO function, the descriptor chain length packed_desc_num is configured to be no less than 16, so that the total buffer size of each descriptor chain is no less than 64KB.
7. The virtualized network packet receiving method based on virtio packed ring according to claim 1, characterized in that, Each descriptor occupies a fixed number of 16 bytes; When the hardware traverses the descriptor ring with a fixed step size of descriptor chain length, the step size is calculated as packed_desc_num × 16 bytes. The hardware directly generates the starting address of the next descriptor chain through the address generator, without needing to read the flags.next field of any descriptor in the current chain.
8. A chip, characterized in that, Includes hardware logic circuitry for performing the method according to any one of claims 1 to 7; The hardware logic circuit includes: The detection module is used to automatically read the descriptor ring and count the length of the descriptor chain, packed_desc_num. The direct memory access engine is used to read descriptors from host memory and write packet data to the corresponding data buffer according to a fixed step size of packed_desc_num×16. An address generator is used to calculate the starting address of the next descriptor chain: next_addr = current_addr + packed_desc_num × 16, and supports modulo operations on the ring size; The interrupt controller is used to notify the host driver when the probe is complete or the message reception is complete. The configuration register stores the cfg_pref_num and vq_init_done flags written by the driver, as well as the packed_desc_num result written by the hardware.
9. A network interface card (NIC), characterized in that, It includes the chip described in claim 8, as well as the Ethernet physical layer interface and the PCIe interface.
Citation Information
Patent Citations
Virtio-net equipment data transmission processing system and method
CN118353969A
Method and device for compatibility of universal network card with VDPA framework
CN121277606A