A multi-core based PCIe network card and a working method thereof

CN115794721BActive Publication Date: 2026-09-08TIH MICROELECTRONIC TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211542689.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-02
Publication Date
2026-09-08
Estimated Expiration
2042-12-02

AI Technical Summary

Technical Problem

[0004]现有的计算机网卡一般采用专用的ASIC(Application Specific IntegratedCircuit,专用集成电路)实现,该实现方法不仅需要专业的IC(Integrated Circuit,集成电路)设计人员开发并且研发周期长、灵活度低,后期产品升级困难,很难应对日新月异的网络需求

Benefits of technology

[0017] The multi-core PCIe network card and its operating method of this invention employ a dual-core architecture design, which can effectively balance the network card load and ensure smooth network traffic. It can also effectively simplify software design complexity, shorten product development cycles, and improve product competitiveness. Therefore, this invention does not require professional IC designers for development, has low cost, short R&D cycle, high flexibility, low software design complexity, is easy to upgrade and maintain, facilitates product iteration, and offers fast data transmission speeds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115794721B_ABST
    Figure CN115794721B_ABST
Patent Text Reader

Abstract

The application discloses a kind of PCIe network cards based on multi-core and working method thereof, belong to network card technical field, the PCIe network cards based on multi-core includes main CPU core, slave CPU core, PCIe module, hardware acceleration module HWA, address conversion AT controller, flash memory Flash, random access memory RAM, media access controller MAC and port physical layer PHY, wherein: main CPU core connects slave CPU core, HWA, Flash and MAC, slave CPU core connects HWA and RAM, AT controller connects PCIe module and RAM, MAC connects PHY;Main CPU core and slave CPU core are shared by using lockless ring Data.The application does not need professional IC designer to develop, low in cost, short in research and development cycle, high in flexibility, low in software design complexity, easy to upgrade maintenance, conducive to product iteration, and data transmission speed is fast.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network interface card (NIC) technology, and in particular to a multi-core PCIe NIC and its operating method. Background Technology

[0002] With the development and progress of network technology, the transmission and acquisition of information through computer networks has become an indispensable part of our daily lives. As network traffic increases day by day, how to efficiently transmit network data has become an urgent problem to be solved.

[0003] A network interface card (NIC), also known as a network adapter, is a piece of computer hardware used to connect to an Ethernet network. The NIC acts as a bridge for communication between the computer and the Ethernet network. Figure 1 As shown, when a computer needs to send data, the network interface card (NIC) obtains data from the user device and sends it to the network; when it needs to receive network data, the NIC obtains data from the network and presents it to the user through the computer. The NIC mainly implements the data link layer and physical layer of the OSI (Open System Interconnection) 7-layer model and interacts with the computer's operating system to complete the sending and receiving of data.

[0004] Existing computer network interface cards (NICs) are generally implemented using dedicated ASICs (Application Specific Integrated Circuits). This approach requires specialized IC (Integrated Circuit) designers, involves long development cycles, lacks flexibility, and makes future product upgrades difficult, making it hard to meet ever-changing network demands. Some NICs are implemented using a single-core CPU (central processing unit). This approach results in interleaved and mixed data transmission and reception processes, making performance optimization a bottleneck and increasing software design complexity, which is detrimental to product iteration. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide a multi-core PCIe network card with fast data transmission speed, low cost, and easy upgrade and maintenance, as well as its working method.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] On the one hand, a multi-core PCIe network interface card (NIC) is provided, including a main CPU core, a slave CPU core, a PCIe module, a hardware acceleration module (HWA), an address translation (AT) controller, flash memory (Flash), random access memory (RAM), a media access controller (MAC), and a port physical layer (PHY), wherein:

[0008] The main CPU core is connected to the slave CPU core, HWA, Flash, and MAC; the slave CPU core is connected to the HWA and RAM; the AT controller is connected to the PCIe module and RAM; and the MAC is connected to the PHY.

[0009] The master CPU core and slave CPU core share data by using a lock-free Ring.

[0010] On the other hand, the above-mentioned operating method for a multi-core PCIe network card is provided, the operating method including a data transmission method, the data transmission method including:

[0011] Step 11: Poll the virtual registers from the CPU core. If the transmit data flag of the virtual register has been modified, immediately start HWA to retrieve the management information of the data frame to be transmitted to the internal shadow descriptor; if the transmit data flag of the virtual register has not been modified, repeat the polling action.

[0012] Step 12: Wait for the CPU core to complete the data retrieval and notify the main CPU core to send the data via inter-process communication;

[0013] Step 13: The main CPU core parses the shadow descriptor and initiates MAC to send data based on the parsing result;

[0014] Step 14: The MAC transmits the data to the Ethernet via the PHY;

[0015] Step 15: After the MAC transmission is completed, notify the main CPU core that the data transmission is complete via an interrupt.

[0016] The present invention has the following beneficial effects:

[0017] The multi-core PCIe network card and its operating method of this invention employ a dual-core architecture design, which can effectively balance the network card load and ensure smooth network traffic. It can also effectively simplify software design complexity, shorten product development cycles, and improve product competitiveness. Therefore, this invention does not require professional IC designers for development, has low cost, short R&D cycle, high flexibility, low software design complexity, is easy to upgrade and maintain, facilitates product iteration, and offers fast data transmission speeds. Attached Figure Description

[0018] Figure 1This is a schematic diagram illustrating the operation of a network interface card (NIC) in the existing technology.

[0019] Figure 2 This is a schematic diagram of the structure of the multi-core PCIe network card of the present invention. Detailed Implementation

[0020] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0021] Traditional network interface cards (NICs) are typically implemented using dedicated ASIC chips. This approach is not only costly and time-consuming to develop, but also hinders upgrades and maintenance. Therefore, this invention proposes a multi-core-based PCIe NIC deployment method. Leveraging a multi-core architecture and underlying high-speed hardware modules, it enables the rapid development of an ecosystem of products. The multi-core design effectively reduces software complexity, shortens product development cycles, and enhances product competitiveness.

[0022] On the one hand, this invention provides a multi-core PCIe network card, such as... Figure 2 As shown, it includes the main CPU core (Mcore), the slave CPU core (Ccore), the PCIe module (PCIe), the hardware acceleration module (HWA), the address translator (AT) controller (AT), flash memory, random access memory (RAM), the media access controller (MAC), and the port physical layer (PHY), where:

[0023] The main CPU core connects to the CPU core, HWA, Flash, and MAC; the CPU core connects to the HWA and RAM; the AT controller connects to the PCIe module and RAM; and the MAC connects to the PHY.

[0024] The master CPU core and slave CPU cores share data by using a lock-free Ring (circular queue).

[0025] Preferably, the PCIe module may include a base address register (BAR) and a direct memory access (DMA) module (shown as DMA in the figure); the main CPU core may also be connected to RAM for convenient related control.

[0026] The following is a brief description of the structure of the multi-core PCIe network card of the present invention:

[0027] 1. PHY: Used to connect a MAC to a physical medium, usually optical fiber or twisted pair, to perform data encoding and decoding functions;

[0028] 2. MAC: Used as a network data frame transceiver to receive and send network frames;

[0029] 3. Mcore (Main CPU Core): It is the commander of the network card operation, responsible for coordinating and directing network data transmission and reception, data frame filtering, network card configuration information retrieval and other functions. New functions can be flexibly added according to business needs in the future.

[0030] 4. Ccore (from CPU core): It is the executor of data transfer. It mainly interacts with the computer through the PCIe module to send and receive data frames according to the business needs of the main CPU core.

[0031] 5. RAM: Program execution and data caching;

[0032] 6. Flash: Stores firmware code and network card configuration parameters;

[0033] 7. HWA: Primarily responsible for prefetching and re-flashing operations of the internal shadow desc in the firmware. New acceleration modules can be added in a timely manner according to actual business needs.

[0034] 8. AT Controller: Primarily responsible for expanding the internal registers of the firmware. This controller maps the RAM within the firmware to the computer's memory space via a PCIe module, thus simulating a register space (virtual register). This allows the computer to directly manipulate the network card's RAM as if it were a local register. When the computer reads or writes to the virtual register, an AT interrupt is generated and sent to the CPU. The CPU responds to the interrupt and queries the AT table to determine which virtual register the host is operating on, then executes the corresponding operation.

[0035] 9. PCIe module: The bridge connecting the computer and the network card, mainly responsible for data exchange, network card configuration, and transmission of network card upgrade firmware.

[0036] On the other hand, the present invention provides a method for operating the above-mentioned multi-core PCIe network card, the method including a data transmission method, the data transmission method including:

[0037] Step 11: Poll the virtual registers from the CPU core (Ccore). If the transmit data flag bit of the virtual register (e.g., FREG_TDTx) has been modified, immediately start HWA to retrieve the management information of the data frame to be transmitted to the internal shadow descriptor (shadow desc); if the transmit data flag bit of the virtual register has not been modified, repeat the polling action.

[0038] Step 12: The CPU core Ccore waits for the data retrieval to complete and notifies the main CPU core Mcore to send the data via inter-process communication (IPC).

[0039] Step 13: The main CPU core Mcore parses the shadow descriptor and starts the MAC to send data based on the parsing result;

[0040] Step 14: The MAC transmits the data to the Ethernet via the PHY;

[0041] Step 15: After the MAC transmission is completed, notify the main CPU core that the data transmission is complete via an interrupt.

[0042] In this embodiment of the invention, the following may be included before step 11:

[0043] Step 101: The computer initializes the network card hardware and writes the descriptor base address into the corresponding virtual register (such as FREG_TDBAx); here, the virtual register is obtained by mapping RAM to the computer memory space through the PCIe module;

[0044] Step 102: The computer prepares the data frame to be sent and triggers the network card to send the data by modifying the transmit data flag bit (such as FREG_TDTx) in the virtual register.

[0045] In a further embodiment, step 15 may be followed by:

[0046] Step 16: The main CPU core Mcore notifies the computer that data transmission is complete and clears the relevant status variable values.

[0047] Furthermore, the working method may also include a data receiving method, the data receiving method comprising:

[0048] Step 21: Retrieve the receive configuration information from the CPU core (Ccore) to the network card's internal shadow descriptor via HWA;

[0049] Step 22: After receiving the data, the MAC notifies the main CPU core Mcore of the data received via an interrupt;

[0050] Step 23: The main CPU core Mcore parses the MAC descriptor information and fills the information into the corresponding shadow descriptor, while incrementing the tail index value (tail value) of the internal Ring;

[0051] Step 24: Poll the CPU core (Ccore) for changes to the internal ring. If the tail index value of the internal ring increases, start the DMA module of the PCIe module to send the data to the computer; if the tail index value of the internal ring does not increase, continue polling.

[0052] Step 25: The CPU core (Ccore) notifies the computer of the data reception status based on the write-back result.

[0053] In this embodiment of the invention, the following may be included before step 21:

[0054] Step 201: The computer initializes the network card hardware and writes the descriptor base address and descriptor parameters into the corresponding virtual registers (for example, the descriptor base address can be written to FREG_RDBAx, and the descriptor parameters can include the receive data flag bit, which is written to FREG_RDTx).

[0055] Step 202: The computer starts the data receiving thread.

[0056] In a further embodiment, step 25 may be followed by:

[0057] Step 26: The computer acquires the received data;

[0058] And / or, step 27: release internal resources from CPU core Ccore.

[0059] This invention employs an embedded processor, which typically has a high operating frequency, allowing it to fully utilize its I / O capabilities, enhance data processing, improve data throughput, and ensure the network card's transmission rate. The main CPU core, slave CPU core, PCIe module, HWA, and other modules can be connected via an AXI bus. During power-on initialization, the main CPU core performs basic configuration of the HWA via the AXI bus, including address setting, operating mode selection, Function ID setting, and prefetching settings. The main CPU core and slave CPU core communicate via a lock-free ring. The main CPU core operates on the ring's hdr to inform the slave CPU core that resources are sufficient for data operations, while the slave CPU core operates on the tail (tail index) to inform the main CPU core that data processing is complete and resources can be reclaimed for further processing. This lock-free ring approach fully utilizes the CPU's processing power and reduces the performance overhead caused by hardware synchronization instructions such as atomic locks.

[0060] In summary, the multi-core PCIe network card and its working method of the present invention have the following advantages:

[0061] 1. It can effectively simplify software design complexity, shorten product development cycle, and improve product competitiveness;

[0062] 2. The dual-core architecture design can effectively balance the network card load and ensure smooth network traffic.

[0063] 3. By adopting a general component design approach, business modules can be flexibly added or removed according to business needs;

[0064] 4. A network card can be used to assist the computer in preprocessing network data packets, reducing the computer's CPU load and improving I / O (input / output) utilization;

[0065] 5. It can provide accurate performance references for hardware iteration, accelerating the localization process of PCIe network card products;

[0066] 6. Provides flexible address translation functionality to adapt to the needs of different hosts, increasing product compatibility and stability.

[0067] This invention implements network interface card (NIC) functionality using multiple CPU cores and high-speed hardware. During application, it adopts a top-down approach, focusing on communication speed, to evaluate multi-core performance, CPU core count, and acceleration hardware. After initial evaluation, the iotrace testing method can be used to trace the latency and time consumption of each stage of the NIC's sending and receiving process in detail. This method accurately analyzes the complexity of each task, providing a solid basis for multi-core task allocation. When analyzing test data, data analysis software based on a production scheduling Gantt chart can be created using Python. This software can quickly analyze hardware and software bottlenecks, providing support for hardware additions / removals and software task allocation. Data sharing among multiple CPUs can be achieved using shared memory, and resource sharing among multiple CPUs is implemented using a Ring-based lock-free approach. These methods maximize CPU resource utilization while ensuring timely task processing. During the continuous adjustment of hardware and software deployment strategies, Tcpreplay and RAW sockets can be used to test the performance of TX (transmit) and RX (receive), gradually adjusting to the optimal state through continuous testing.

[0068] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multi-core PCIe network interface card, characterized in that, This includes the main CPU core, slave CPU cores, PCIe module, hardware acceleration module (HWA), address translation (AT) controller, flash memory, random access memory (RAM), media access controller (MAC), and port physical layer (PHY), among which: The main CPU core is connected to the slave CPU core, HWA, Flash, and MAC; the slave CPU core is connected to the HWA and RAM; the AT controller is connected to the PCIe module and RAM; and the MAC is connected to the PHY. The master CPU core and slave CPU core achieve data sharing by using a lock-free Ring; The working method includes a data transmission method, which includes: Step 11: Poll the virtual registers from the CPU core. If the transmit data flag of the virtual register has been modified, immediately start HWA to retrieve the management information of the data frame to be transmitted to the internal shadow descriptor; if the transmit data flag of the virtual register has not been modified, repeat the polling action. Step 12: Wait for the CPU core to complete the data retrieval and notify the main CPU core to send the data via inter-process communication; Step 13: The main CPU core parses the shadow descriptor and initiates MAC to send data based on the parsing result; Step 14: The MAC transmits the data to the Ethernet via the PHY; Step 15: After the MAC transmission is completed, the main CPU core is notified of the completion of data transmission via an interrupt; The working method includes a data receiving method, which includes: Step 21: Retrieve the receive configuration information from the CPU core to the network card's internal shadow descriptor via HWA; Step 22: After receiving the data, the MAC notifies the main CPU core that data has been received via an interrupt; Step 23: The main CPU core parses the MAC descriptor information and fills the information into the corresponding shadow descriptor, while simultaneously incrementing the tail index value of the internal circular queue Ring; Step 24: Poll the CPU core for changes to the internal ring. If the tail index value of the internal ring increases, start the DMA module of the PCIe module to send the data to the computer; if the tail index value of the internal ring does not increase, continue polling. Step 25: The CPU core notifies the computer of the data reception status based on the write-back result.

2. The multi-core PCIe network card according to claim 1, characterized in that, The PCIe module includes a base address register (BAR) and a direct memory access (DMA) module. And / or, the main CPU core is also connected to the RAM.

3. The working method according to claim 2, characterized in that, The preceding steps, prior to step 11, include: Step 101: The computer initializes the network card hardware and writes the descriptor base address into the corresponding virtual registers; Step 102: The computer prepares the data frame to be sent and triggers the network card to send the data by modifying the transmit data flag bit of the virtual register.

4. The working method according to claim 3, characterized in that, Step 15 is followed by: Step 16: The main CPU core notifies the computer that data transmission is complete and clears the relevant status variable values.

5. The working method according to claim 4, characterized in that, The preceding steps, prior to step 21, include: Step 201: The computer initializes the network card hardware and writes the descriptor base address and descriptor parameters into the corresponding virtual registers; Step 202: The computer starts the data receiving thread.

6. The working method according to claim 5, characterized in that, Step 25 is followed by: Step 26: The computer acquires the received data; And / or, step 27: release internal resources from the CPU core.