FPGA-based high-speed low-delay 10-gigabit switching method and system

By designing data cache modules and three-level pipeline architecture in FPGAs and dynamically adjusting the transmission rate, the problem of high delay of ASIC chip switches is solved, low-latency and high-throughput data forwarding is achieved, and high-speed interconnection is suitable for edge intelligent computing.

CN120342951APending Publication Date: 2025-07-18CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510484424.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, 10 Gigabit Ethernet switches based on ASIC chips have problems of high switching delays and low processing efficiency, and it is difficult to meet the needs of edge intelligent computing for low latency and high bandwidth.

Method used

FPGA is used to realize the high-speed, low-latency 10,000 Gigabit switching method, store Ethernet data frames through the data cache module, dynamically update the forwarding address table, adopt a three-level pipeline architecture for address search, and dynamically adjust the sending rate to optimize VLAN division and queue scheduling, and reduce the residence time of data packets within the switching unit.

Benefits of technology

It realizes low-latency and high-throughput data forwarding, meets the network requirements of distributed inference on the side of the artificial intelligence model, and improves the operation efficiency and flexibility of side applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342951A_ABST
    Figure CN120342951A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of Ethernet high-speed interconnection, and provides a high-speed low-delay 10-gigabit switching method and system based on an FPGA, and the method comprises the steps: receiving Ethernet data frames, storing effective data frames in a data caching module, and generating a caching write-in address and a control signal; reading a destination address, a source address and a type field of the analyzed data frame from the data caching module, and updating a forwarding address table through a dynamic learning mode or a preset mode; based on the updated forwarding address table, a three-level assembly line framework is adopted to complete address searching; reading the data frame from the cache according to the forwarding decision, repackaging the Ethernet header information and the check code, and controlling the sending standard rate; and meanwhile, the sending rate is dynamically adjusted, and VLAN division, port rate and queue scheduling parameters are optimized in combination with an external configuration command. According to the invention, the onboard processor is utilized to cooperate with the FGPA to realize the fast search of the forwarding table, the residence time of the data packet in the switching unit is reduced, and the operation efficiency of the side application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of high-speed Ethernet interconnection, and particularly relates to a high-speed and low-latency 10 Gigabit switching method and system based on FPGA. Background Art

[0002] With the rapid development of artificial intelligence-related technologies, the demand for computing resources, data elements, etc. shows an exponential growth trend. This makes the traditional centralized processing mode face huge bandwidth pressure and energy consumption, and unable to meet requirements such as bandwidth, real-time performance, and reliability. Therefore, edge computing emerges as the times require and gradually develops towards edge intelligence. The edge intelligence solution can instantaneously process data at the edge and quickly feedback results. By sinking the storage, computing, and intelligent resources of the cloud computing center to the network edge, this solution promotes the migration of artificial intelligence applications from the cloud to the edge, thereby meeting the key requirements of related industries in aspects such as real-time response, intelligent application, and agile perception.

[0003] Industrial Ethernet, as a current mature technology, is widely used in the high-speed interconnection of edge intelligent computing. Currently, the relatively mature technologies include 100 Mbps, 1 Gbps, and 10 Gbps Ethernet technologies. Among them, the single-hop store-and-forward delay of a 100 Mbps Ethernet switch is usually about 10 us, the 1 Gbps Ethernet delay is usually about 3 - 4 us, and the 10 Gbps forwarding processing delay is usually about 1 us. The implementation method usually uses an ASIC (Application Specific Integrated Circuit) chip. The processing process usually includes the reception of data packets, collision and flow control, maintenance and lookup of the forwarding address table, implementation of layer 2 functions, and transmission of data packets. The entire process, especially the reception, parsing, table lookup, and implementation of layer 2 functions of data frames, consumes a relatively large amount of delay. Although the table lookup algorithm implemented using a custom ASIC switching chip can perform table lookup operations quickly, due to its custom characteristics, the logic and storage resources required for table lookup are fixed and unchangeable, which will make the optimization and upgrade of the system extremely complex and costly, and this will be disadvantageous to the scalability and innovation of network research. Summary of the Invention

[0004] The purpose of the present invention is to provide a high-speed and low-latency 10 Gigabit switching method and system based on FPGA to solve the problems of high internal switching delay and low processing efficiency in the prior art.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a high-speed and low-latency 10 Gigabit switching method based on FPGA, including: After the received Ethernet data frame is processed by the RS sublayer, it is sent to the data reception module, and the valid data frame is stored in the data cache module, generating a cache write address and a control signal; Read the destination address, source address, and type field of the parsed data frame from the data cache module, and update the forwarding address table through the dynamic learning mode or the preset mode; Based on the updated forwarding address table, complete the address lookup using a three-stage pipeline architecture; Read the data frame from the cache according to the forwarding decision, re-encapsulate the Ethernet header information and checksum, and control the sending standard rate; at the same time, dynamically adjust the sending rate, and optimize the VLAN division, port rate, and queue scheduling parameters in combination with external configuration commands to achieve network congestion avoidance and low-latency transmission.

[0006] Further, after the received Ethernet data frame is processed by the RS sublayer, it is sent to the data reception module, and the valid data frame is stored in the data cache module, generating a cache write address and control signals, including: Perform a preliminary analysis on the received 10 Gigabit Ethernet data frame. The analysis process includes extracting the preamble, frame header information, and verifying the integrity of the frame. According to the analysis results, store the valid data frame in the data cache module and generate the corresponding cache write address and control signals; the Ethernet data frame processed by the RS sublayer will be sent to the data reception module for frame parsing processing, and finally sent to the packet header extraction module and the DMA part respectively according to its type.

[0007] Further, the data cache module is designed with a high-speed data cache inside the Field Programmable Gate Array (FPGA) to temporarily store the received data frames; the cache adopts a multi-port RAM and FIFO structure.

[0008] Further, the process of reading the destination address, source address, and type field of the parsed data frame from the data cache module and updating the forwarding address table through the dynamic learning mode or the preset mode includes: In the dynamic learning mode, the FPGA automatically learns the source and destination addresses of different ports and updates the address table. In the preset mode, the static MAC address table is configured through an external CPU to form a port forwarding mapping relationship.

[0009] Further, when forwarding the data frame, it specifically includes: Receive the data frame and extract the source MAC address and destination MAC address from it; According to the initial address table maintenance mode, generate a MAC address forwarding table according to the preset mode and MAC table format; Lookup in the forwarding address table. First, predict whether the destination MAC address is a broadcast or multicast address. If so, perform a broadcast / multicast mark and wait for further forwarding; if it is a unicast address, look up the destination MAC address in the MAC address table. Based on the lookup algorithm, if a matching item is found, obtain the corresponding port number. If no matching item is found, it is marked as an "unknown unicast frame"; According to the lookup result, the switching device executes the following forwarding logic: Known unicast frame: If the destination MAC address is in the table, and the corresponding port is different from the receiving port, and the VLAN tables are the same, then forward the data frame from this port; if the port corresponding to the destination MAC address is the same as the receiving port, then discard the data frame; Unknown unicast frame: If the destination MAC address is not in the table, determine whether there is a VLAN tag. If not, flood the data frame to all ports except the receiving port. If there is a VLAN tag, flood the data frame to all ports within a specific VLAN except the receiving port; Broadcast / multicast frame: Determine whether there is a VLAN tag. If not, directly flood the data frame to all ports except the receiving port. If there is, flood the data frame to all ports within a specific VLAN except the receiving port.

[0010] Furthermore, based on the updated forwarding address table, a three-level pipeline architecture is adopted to complete address lookup, including: The first level uses a fast prediction strategy to determine whether the current MAC address is in the possible current MAC address table. If not, it is directly determined as an unknown unicast. If so, the second level uses a hash table to accurately locate the position of the bucket. The third level is a parallel storage unit, which uses the parallel storage unit of the FPGA to store multiple values to avoid hash values being overwritten.

[0011] Furthermore, read the data frame from the cache according to the forwarding decision, re-encapsulate the Ethernet header information and checksum, and control the sending standard rate; at the same time, dynamically adjust the sending rate, and optimize the VLAN division, port rate, and queue scheduling parameters in combination with external configuration commands to achieve network congestion avoidance and low-latency transmission, including: Reconstruct and encapsulate the data frame, send it to the physical layer, and synchronously update the aging time. Each time a MAC address is learned or looked up, update the aging time of the corresponding table entry; if the aging time of the table entry times out, delete the table entry from the MAC address table; Receive the configuration commands of the external CPU through the SPI interface, set the PHY working mode, forwarding rules, and queue scheduling parameters, provide configuration management software, and support users to configure device parameters through a graphical interface; Based on the credit-based flow control algorithm, dynamically adjust the data sending rate to prevent network congestion, and optimize the data transmission efficiency of the sending end according to the feedback information from the receiving end; obtain the device running status in real time through the management interface, and support VLAN division and port rate configuration.

[0012] In a second aspect, the present invention provides a high-speed and low-latency 10 Gigabit Ethernet switching system based on FPGA, including: A data receiving module, which is used to send the received Ethernet data frame to the data receiving module after being processed by the RS sublayer, store the valid data frame in the data buffer module, and generate a buffer write address and a control signal; An address updating module, which is used to read the destination address, source address and type field of the parsed data frame from the data buffer module, and update the forwarding address table through a dynamic learning mode or a preset mode; An address lookup module, which is used to complete address lookup based on the updated forwarding address table by adopting a three-stage pipeline architecture; A data sending module, which is used to read the data frame from the buffer according to the forwarding decision, re-encapsulate the Ethernet header information and the check code, and control the sending standard rate; at the same time, dynamically adjust the sending rate, and optimize the VLAN division, port rate and queue scheduling parameters in combination with external configuration commands to achieve network congestion avoidance and low-latency transmission.

[0013] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the high-speed low-latency 10Gigabit switching method based on FPGA are implemented.

[0014] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the high-speed low-latency 10Gigabit switching method based on FPGA are implemented.

[0015] Compared with the prior art, the present invention has the following technical effects: By streamlining the layer-2 redundancy function, optimizing the switching processing flow, and using the on-board processor to cooperate with the FGPA to achieve fast lookup of the forwarding table, the present invention reduces the residence time of data packets inside the switching unit, and controls the single-node forwarding latency within hundreds of nanoseconds to meet the operation requirements of high-bandwidth and low-latency networks for edge-side distributed inference of artificial intelligence models, and improves the operation efficiency of edge-side applications.

[0016] Through the hardware acceleration technology of FPGA, the present invention optimizes the switching processing flow, and uses the on-board processor to cooperate with the FGPA to achieve fast lookup of the forwarding table, reduces the residence time of data packets inside the switching unit, realizes low-latency and high-throughput data forwarding, and controls the single-node forwarding latency within hundreds of nanoseconds to meet the operation requirements of high-bandwidth and low-latency networks for edge-side distributed inference of artificial intelligence models, and improves the operation efficiency of edge-side applications. The device has the advantages of high throughput, low latency, high reliability and flexible configuration, and is suitable for scenarios with extremely high requirements for network performance. Description of the Drawings

[0017] Figure 1It is a block diagram of the logic for an ultra-low latency 10 Gigabit Ethernet switch.

[0018] Figure 2 It is the overall logic flow chart of the 10 Gigabit switching function based on FPGA.

[0019] Figure 3 It is the data frame forwarding processing flow chart.

[0020] Figure 4 It is the flow chart of the present invention. Detailed implementation manners

[0021] The present invention will be further described below with reference to the accompanying drawings: Example 1. Please refer to Figure 4 The present invention provides a high-speed and low-latency 10 Gigabit switching method based on FPGA, including: After the received Ethernet data frame is processed by the RS sublayer, it is sent to the data receiving module, and the valid data frame is stored in the data cache module, generating a cache write address and a control signal; Read the destination address, source address and type field of the parsed data frame from the data cache module, and update the forwarding address table through the dynamic learning mode or the preset mode; Based on the updated forwarding address table, complete the address lookup using a three-stage pipeline architecture; Read the data frame from the cache according to the forwarding decision, re-encapsulate the Ethernet header information and the check code, and control the sending standard rate; at the same time, dynamically adjust the sending rate, and optimize the VLAN division, port rate and queue scheduling parameters in combination with external configuration commands to achieve network congestion avoidance and low-latency transmission.

[0022] The present invention reduces the residence time of data packets inside the switching unit by streamlining the layer-2 redundant functions, optimizing the switching processing flow, and using the on-board processor to cooperate with the FGPA to quickly find the forwarding table, controlling the single-node forwarding latency within hundreds of nanoseconds to meet the operation requirements of high bandwidth and low latency of the artificial intelligence model for edge-side distributed inference, and improving the operation efficiency of edge-side applications.

[0023] Example 2. The present invention provides a high-speed and low-latency 10 Gigabit switching method based on FPGA, which specifically includes the following four parts: data packet reception and transmission, data packet parsing and encapsulation, data exchange control, and configuration and management.

[0024] (1) Data packet reception and transmission 1) Design of the 10 Gigabit Ethernet physical layer interface: Design a physical layer interface circuit compatible with the 10 Gigabit Ethernet standard, including an optical module interface or an electrical port interface. Implement signal conditioning, encoding / decoding functions at the physical layer, such as 64B / 66B encoding and decoding, to ensure reliable data transmission over the physical medium.

[0025] 2) Data frame reception processing: Perform a preliminary parsing of the received 10 Gigabit Ethernet data frame. The parsing process includes extracting the preamble, frame header information, and verifying the integrity of the frame. According to the parsing results, store the valid data frame in the data cache module and generate the corresponding cache write address and control signals.

[0026] The Ethernet MAC frame processed by the RS sublayer will be sent to the data reception module for frame parsing processing and finally sent to the packet header extraction module and the DMA part for further processing according to its type.

[0027] 3) Data cache module design: Design a high-speed data cache inside the FPGA to temporarily store the received data frames. The cache adopts a multi-port RAM and FIFO structure. Determine the size and depth of the cache according to the actual application requirements to meet the cache requirements under different data traffic conditions and avoid data overflow.

[0028] 4) Data frame transmission processing: According to the forwarding decision result, send the data frame out from the corresponding output port. Before sending, repackage the data frame and add the necessary frame header and check information. Control the transmission rate to ensure that the data transmission on the output port meets the standard rate requirements of 10 Gigabit Ethernet and avoid transmission buffer overflow.

[0029] 5) Error frame detection and marking: After the switching device starts forwarding the frame, continuously calculate the CRC. If an error is detected, immediately stop forwarding the frame data and use the invalid symbol encoding rule of the physical layer 8B / 10B encoding to immediately send the physical layer termination symbol (such as 4 consecutive / K28.5 / ). After the receiving end detects the termination symbol, discard the received frame fragment. Record the error counter for network management analysis.

[0030] (2) Packet parsing and encapsulation Design the packet parsing logic to parse the received Ethernet packet and extract fields such as the destination address, source address, and type of the packet. According to the parsed information, determine the forwarding path of the packet and perform corresponding processing on the packet, such as adding or modifying certain fields.

[0031] Before the data packet is forwarded to the output port, encapsulate the data packet, add necessary Ethernet header information to ensure that the data packet can be correctly transmitted to the destination device, and at the same time add a preamble to achieve external encapsulation of the data frame.

[0032] (3)Data exchange control 1)Maintenance of the forwarding address table Design the MAC address forwarding table format, including MAC address, port number, aging time, VLAN number, etc. The forwarding address table can be pre-statically configured in the on-chip memory of the FPGA or updated in real time through a dynamic learning algorithm. Therefore, there are two modes for maintaining the forwarding address table, namely the preset mode and the dynamic learning mode. In the preset mode, the external CPU reads / writes and deletes the MAC address forwarding table to the FPGA to implement the mapping management between the internal ports of the switching device; in the dynamic learning mode, the internal MAC address table management module of the FPGA automatically learns the source and destination addresses from different ports and dynamically updates the mapping relationships of different ports.

[0033] 2)Data frame forwarding decision For the specific process of data frame forwarding decision, please refer to Figure 3 Data frame forwarding processing flow: ①: Receive the data frame and extract the source MAC address and destination MAC address from it.

[0034] ②: According to the initial address table maintenance mode, generate the MAC address forwarding table according to the preset mode and MAC table format.

[0035] ③: Forwarding address table lookup. First, pre-judge whether the destination MAC address is a broadcast or multicast address. If it is, perform broadcast / multicast marking and wait for further forwarding. If it is a unicast address, look up the destination MAC address in the MAC address table. The three-stage pipeline architecture address table forwarding lookup algorithm is used to improve the lookup speed and accuracy, which is divided into three sub-modules. Specifically, ① Bloom Filter pre-filtering: Use 3 different hash functions (improved based on CRC32) to reduce the misjudgment probability. Each table entry is stored with 2 bits (instead of the traditional 1 bit), further reducing misjudgment. It occupies about 1.5 Mb of storage space and can support the pre-judgment of 8000 addresses. ② Hierarchical hash table: The first layer (main table) maps the MAC address to 4096 storage buckets through a hash function, and each bucket directly stores a common address. The second layer (overflow handling), each bucket has an additional 8 reserved positions to solve the hash conflict problem (similar to adding partitions when the drawers are not enough). Each table entry only needs 84 bits, including the MAC address (48 bits), port mapping (4 bits), and timestamp (32 bits). Use the high-speed storage unit (Block RAM) of FPGA to achieve parallel access. ③ Target location: 8 entries in each storage bucket are matched and compared simultaneously, and the result can be obtained within 1 clock cycle. Utilize the hardware parallel characteristics of FPGA to significantly shorten the lookup time. Based on the above lookup algorithm, if a matching item is found, obtain the corresponding port number. If no matching item is found, it is marked as an "unknown unicast frame".

[0036] ④: Forwarding decision. According to the lookup result, the switching device executes the following forwarding logic: Known unicast frame: If the destination MAC address is in the table, and the corresponding port is different from the receiving port, and the VLAN table is the same, forward the data frame from this port. If the port corresponding to the destination MAC address is the same as the receiving port, discard the data frame (to avoid loops).

[0037] Unknown unicast frame: If the destination MAC address is not in the table, judge whether there is a VLAN tag. If not, flood the data frame to all ports except the receiving port. If there is a VLAN tag, flood the data frame only to all ports within a specific VLAN except the receiving port.

[0038] Broadcast / multicast frame: Judge whether there is a VLAN tag. If not, directly flood the data frame to all ports except the receiving port. If there is, flood the data frame only to all ports within a specific VLAN except the receiving port.

[0039] ⑤: Data queue scheduling: Perform traffic prediction and dynamic priority adjustment at the input port, and schedule each port queue based on microsecond-level time slots and dynamic credit mechanisms to avoid buffer overflow and collision loss of packets.

[0040] ⑥: Reconstruct and encapsulate the data frame and send it to the physical layer.

[0041] ⑦: Synchronously update the aging time. Each time a MAC address is learned or found, update the aging time of the corresponding entry. If the aging time of an entry times out, delete the entry from the MAC address table.

[0042] 3) Data queue scheduling Design a dynamic time-sharing cross-scheduling algorithm, which combines the dynamic priority assignment of the input queue, the crossbar scheduling based on time slices, and the credit flow control, and realizes collision-free switching with low latency and high throughput through FPGA hardware.

[0043] ① Construct the dynamic priority sub-queues of the input ports and the output port queues.

[0044] ② Conduct traffic prediction and dynamic priority adjustment. Analyze the traffic patterns of each queue in real time (such as bandwidth requirements, burstiness), and dynamically adjust the priority weights. Use sliding window statistics (such as the average packet length and rate in the recent N time slices) to predict the congestion risk of the output port. If it is detected that a certain output port is about to be overloaded, automatically increase the weight of the low-priority traffic sent to this port (such as P1→P0) to avoid tail drop.

[0045] ③ Design a time slot cross-scheduling algorithm. Divide time into microsecond-level time slots, split the data packet into cells of a fixed size (64 bytes), and dynamically allocate crossbar connections within each time slot to support parallel transmission of fragmented data units from multiple input ports to the same output port. The scheduler selects the input port with the highest weight to establish a connection within each time slot according to the queue priority and credit value.

[0046] ④ Credit-based flow control. Introduce a dynamic credit counter for each link from the input port to the output port, set the credit value as the remaining capacity of the output port buffer, consume the credit when sending a cell, and return a credit increment through in-band signaling (such as a dedicated control bit) every time the receiving end processes a cell. The FPGA hardware automatically blocks the input port with exhausted credit to avoid buffer overflow.

[0047] ⑤ Data packet recombination and out-of-order correction. Design a hardware-level cell recombination engine at the output port, maintain a recombination buffer based on cell sequence numbers, and support the fast recombination of out-of-order cells.

[0048] (IV) Configuration and management 1) Basic configuration Implement a communication interface with external devices. For example, receive configuration commands and parameters through the SPI (Serial Peripheral Interface) interface, and configure and manage the switching device, such as the working mode of the PHY, the forwarding rules of the switching matrix, the parameters of the queue scheduling algorithm, etc.

[0049] Write configuration and management software to provide a friendly user interface for users to configure and manage devices conveniently.

[0050] Implement a logging function to record the running status, error messages, etc. of the system for system debugging and maintenance.

[0051] 2) Flow control mechanism Implement a flow control mechanism, such as credit-based flow control or window-based flow control. The flow control algorithm is used to adjust the sending and receiving rates of data to prevent network congestion.

[0052] Dynamically adjust the data sending rate of the sending end according to the feedback information from the receiving end to ensure reliable data transmission.

[0053] 3) Layer 2 management Implement device management functions, including device status monitoring, fault diagnosis, logging, etc. Through the management interface, the running status of the device can be obtained in real time to detect and solve faults in a timely manner.

[0054] Write a configuration management program to implement parameter configuration and status monitoring of the switching device. Through the configuration management program, parameters such as port rate, working mode, VLAN (Virtual Local Area Network) division, etc. can be set.

[0055] Design a control module to be responsible for managing the operation of the entire switching device. The control module includes functions such as port status monitoring, configuration management, and flow control.

[0056] Example 3. This example aims to design a high-performance switching device based on a Field Programmable Gate Array (FPGA) for the core switching node in the edge aggregation point network. Through the hardware acceleration technology of FPGA, low-latency and high-throughput data forwarding is achieved to meet the strict requirements of the edge network for real-time performance and high performance.

[0057] As Figure 1 shown, the switching device mainly consists of the following parts: 10G Ethernet optical interface modules 1 - 4: Responsible for receiving / sending data packets on the network, implementing the conversion, encoding / decoding of optical and electrical signals, and interconnecting and exchanging data with the internal circuit through the SERDES interface.

[0058] Power input module: Responsible for providing power for the whole device, with an input voltage of 24V and a current not less than 3A.

[0059] Power distribution module: Responsible for distributing the input voltage as required into the voltages required by each module on the board: 3.3V, 2.5V, 1V, etc., and connecting to each module through the power plane.

[0060] Configuration Management Module: Responsible for device configuration, status monitoring, and exception handling, including initializing the configuration management forwarding address table learning mode, presetting forwarding address table values, starting / stopping settings for advanced functions such as VLAN, etc.

[0061] Clock Module: Provides various frequency clocks required at the board level, including 125MHz for gigabit networks, 156.25MHz and 200MHz for 10-gigabit networks, 25MHz and 32.768KHz for the configuration management module, etc.

[0062] Reset Module: Provides a starting signal for the entire board to facilitate the synchronous startup of each module, ensuring the reliability, stability, and predictability of the system.

[0063] Gigabit SERDES Interconnection Interface Module: Used to achieve interconnection and interoperability with low-speed networks.

[0064] Debug Interface Module: Used for device monitoring, debugging, and configuration, etc.

[0065] Ultra-Low Latency Switching Module: Used to implement functions such as receiving, parsing, processing, routing decision-making, forwarding, encapsulating, and sending data frames, which is the core function of the device.

[0066] FPGA Functional Logic Implementation The FPGA serves as the core processing unit and realizes the high-speed processing of data packets through hardware acceleration technology. The parallel processing ability of the FPGA enables it to process multiple data streams simultaneously, significantly improving the data processing efficiency. At the same time, the low-latency characteristic of the FPGA makes it particularly suitable for real-time data processing scenarios in the edge network. The overall logic of the 10-gigabit switching function based on the FPGA is as Figure 2 shown.

[0067] Adopt a high-performance FPGA (such as the Xilinx Kintex UltraScale+ series, model: xcku5p-ffvb676-2-i) as the core processing unit to realize data packet reception, parsing, switching control, and sending functions. The 10-gigabit Ethernet interface supports optical module interfaces and electrical interfaces with a rate of 10Gbps and is compatible with the IEEE 802.3ae standard. The cache module adopts a multi-port RAM and FIFO structure, with a designed cache depth of 16KB, supporting the efficient processing of burst traffic. The external configuration interface uses an SPI interface to communicate with an external CPU and supports dynamic configuration and management. The specific implementation plan is as follows: 1) Data Packet Reception and Transmission Physical Layer: Connect the fiber optic link of the edge network using an optical module interface, supporting a transmission rate of 10Gbps. Implement 64B / 66B encoding and decoding functions to ensure the reliable transmission of data on the physical link. Optimize the signal quality through a signal conditioning circuit to reduce the bit error rate.

[0068] Data frame reception processing: After the received Ethernet data frame is processed by the RS sublayer, it is sent to the data reception module. The frame preamble and frame header information are extracted, and the frame integrity is verified. The valid data frame is stored in the data cache module, and the cache write address and control signal are generated.

[0069] Data cache module: A multi-port RAM is designed inside the FPGA as the data cache, supporting simultaneous read and write operations. The cache depth is 16KB, which can handle burst traffic and avoid data overflow.

[0070] Data frame transmission processing: According to the forwarding decision result, the data frame is read from the cache and sent through the specified port. Before sending, the data frame is re-encapsulated, and Ethernet header information and checksum are added. The transmission rate is controlled to ensure compliance with the 10Gbps standard rate. When an error frame is detected, the physical layer error frame processing strategy is promptly used to notify the receiving end to discard the frame data and request again.

[0071] 2) Packet parsing and encapsulation The received Ethernet packet is parsed to extract the destination address, source address, and type fields. The forwarding address table is searched according to the destination address to determine the forwarding path of the packet. For specific details, refer to 3) Data exchange control. Before forwarding, the packet is re-encapsulated, and necessary Ethernet header information and preamble are added.

[0072] 3) Data exchange control There are two modes for maintaining the forwarding address table, namely the dynamic learning mode and the preset mode. In the dynamic learning mode, the FPGA automatically learns the source and destination addresses from different ports and updates the forwarding address table. In the preset mode, the static MAC address table is configured through an external CPU.

[0073] After the MAC forwarding address table is formed, it has the port forwarding mapping relationship. The data needs to be mapped to the corresponding port through the table lookup forwarding module. In this implementation case, a three-stage pipeline architecture address table forwarding lookup algorithm is used to improve the table lookup efficiency and achieve the extreme-speed forwarding of data frames. The first stage uses a fast prediction strategy to determine whether the current MAC address is likely in the current MAC address table. If not, it is directly determined as an unknown unicast; if it is likely to be in, the second stage uses a hash table to accurately locate the position of the bucket where it is located; the third stage is a parallel storage unit. Since hash values may conflict, the parallel storage unit of the FPGA is used to store multiple values to avoid hash values being overwritten. Since it is parallel storage, the lookup speed is independent of the depth and only consumes one clock cycle. This algorithm can determine the forwarding direction of the data frame in the fastest one clock cycle and complete the lookup of the destination port in the MAC address table within three clock cycles.

[0074] 4) Configuration and Management Basic Configuration: Receive configuration commands from an external CPU through the SPI interface to set the PHY working mode, forwarding rules, and queue scheduling parameters. Provide configuration management software to support users in configuring device parameters through a graphical interface.

[0075] Flow Control Mechanism: Implement a credit-based flow control algorithm to dynamically adjust the data transmission rate and prevent network congestion. Optimize the data transmission efficiency of the sender based on the feedback information from the receiver.

[0076] Layer 2 Management: Implement device status monitoring, fault diagnosis, and logging functions. Obtain the device operating status in real time through the management interface, and support VLAN division and port rate configuration.

[0077] 5) Performance Metrics High Throughput: Support full-duplex data transmission at 10 Gbps to meet the high-bandwidth requirements of edge aggregation nodes.

[0078] Low Latency: The end-to-end transmission latency is less than 800 nanoseconds to meet low-latency application scenarios.

[0079] High Reliability: The bit error rate is less than 10^-12 to ensure the reliability of data transmission.

[0080] Flexibility: Support dynamic configuration and management to adapt to the requirements of different network environments.

[0081] In the edge center network, this switching device can efficiently process data exchange between edge-side intelligent computing clusters, significantly reducing network latency. Through preset address table policies and an efficient forwarding address table lookup algorithm, it realizes the extremely fast forwarding of network data packets, and a dynamic traffic analysis and credit adjustment mechanism to optimize the forwarding performance and avoid network congestion. At the same time, it provides flexible configuration and management functions to facilitate network administrators to monitor and maintain the device in real time.

[0082] In another embodiment of the present invention, a high-speed low-latency 10 Gigabit Ethernet switching system based on FPGA is provided, which can be used to implement the above-mentioned high-speed low-latency 10 Gigabit switching method based on FPGA. Specifically, this system includes: A data receiving module, which is used to send the received Ethernet data frame to the data receiving module after being processed by the RS sublayer, store the valid data frame in the data cache module, and generate a cache write address and a control signal; An address update module, which is used to read and parse the destination address, source address, and type field of the data frame from the data cache module, and update the forwarding address table through a dynamic learning mode or a preset mode; An address lookup module, which is used to complete address lookup based on the updated forwarding address table using a three-stage pipeline architecture; A data sending module, which is used to read data frames from a cache according to a forwarding decision, re-encapsulate Ethernet header information and check codes, and control the sending standard rate; at the same time, dynamically adjust the sending rate, and optimize VLAN division, port rate, and queue scheduling parameters in combination with external configuration commands to achieve network congestion avoidance and low-latency transmission.

[0083] The division of modules in the embodiments of the present invention is illustrative, only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present invention, each functional module may be integrated in a processor, may also exist independently physically, or two or more modules may be integrated in one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0084] In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiments of the present invention may be used for the operation of the high-speed low-latency 10Gigabit switching method based on FPGA.

[0085] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. Moreover, one or more instructions suitable for being loaded and executed by the processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the high-speed low-latency 10 Gigabit switching method based on FPGA in the above embodiments.

[0086] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0087] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0088] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in Figure 1 one flow or multiple flows and / or blocksFigure 1 The functions specified in one or more boxes.

[0089] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for realizing the steps of the functions specified in one or more processes and / or boxes. Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A high-speed and low-latency 10 Gigabit switching method based on FPGA, characterized in that include: After the received Ethernet data frame is processed by the RS sublayer, it is sent to the data receiving module, which stores the valid data frame in the data cache module and generates the cache write address and control signal; Read the destination address, source address and type field of the parsed data frame from the data cache module, and update the forwarding address table through the dynamic learning mode or the preset mode; Based on the updated forwarding address table, a three-stage pipeline architecture is used to complete the address lookup; Reads data frames from the cache based on forwarding decisions, re-encapsulates Ethernet header information and checksums, and controls the standard rate of transmission; At the same time, the sending rate is dynamically adjusted, and the VLAN division, port rate and queue scheduling parameters are optimized in combination with external configuration commands to achieve network congestion avoidance and low-latency transmission.

2. The method for high-speed and low-latency 10 Gigabit Ethernet switching based on FPGA according to claim 1, wherein The received Ethernet data frame is processed by the RS sublayer and sent to the data receiving module, and the valid data frame is stored in the data cache module, and the cache write address and control signal are generated, including: The received 10 Gigabit Ethernet data frames are initially parsed, and the parsing process includes extracting the frame preamble, frame header information, and checking the integrity of the frame. According to the parsing results, the valid data frames are stored in the data cache module, and the corresponding cache write address and control signal are generated; the Ethernet data frames processed by the RS sublayer will be sent to the data receiving module for frame parsing, and finally sent to the packet header extraction module and DMA part according to their types.

3. The high-speed and low-latency 10 Gigabit switching method based on FPGA according to claim 2, characterized in that, The data cache module is a high-speed data cache designed inside the field programmable gate array FPGA, which is used to temporarily store received data frames; the cache adopts a multi-port RAM and FIFO structure.

4. The high-speed and low-latency 10 Gigabit switching method based on FPGA according to claim 1, characterized in that, The method of reading the destination address, source address and type field of the parsed data frame from the data cache module and updating the forwarding address table through a dynamic learning mode or a preset mode includes: In dynamic learning mode, the FPGA automatically learns the source and destination addresses of different ports and updates the address table. In preset mode, the static MAC address table is configured through the external CPU to form a port forwarding mapping relationship.

5. The high-speed and low-latency 10Gigabit switching method based on FPGA according to claim 4, characterized in that When data frames are forwarded, it specifically includes: Receive data frames and extract source MAC address and destination MAC address from them; According to the initial address table maintenance mode, a MAC address forwarding table is generated according to the preset mode and MAC table format; The forwarding address table is searched to first determine whether the destination MAC address is a broadcast or multicast address. If so, it is marked as a broadcast / multicast address for further forwarding. If it is a unicast address, the destination MAC address is searched in the MAC address table. Based on the search algorithm, if a match is found, the corresponding port number is obtained. If no match is found, it is judged and marked as an "unknown unicast frame". Based on the search results, the switch executes the following forwarding logic: Known unicast frame: If the destination MAC address is in the table, and the corresponding port is different from the receiving port, but the VLAN table is the same, the data frame is forwarded from this port; if the port corresponding to the destination MAC address is the same as the receiving port, the data frame is discarded; Unknown unicast frame: If the destination MAC address is not in the table, determine whether there is a VLAN tag. If not, flood the data frame to all ports except the receiving port. If there is a VLAN tag, flood the data frame to all ports within the specific VLAN except the receiving port. Broadcast / multicast frame: Determine whether there is a VLAN tag. If not, directly flood the data frame to all ports except the receiving port. If there is, flood the data frame to all ports within the specific VLAN except the receiving port.

6. The high-speed and low-latency 10 Gigabit switching method based on FPGA according to claim 4, characterized in that Based on the updated forwarding address table, the address lookup is completed using a three-stage pipeline architecture, including: In the first stage, through a fast prediction strategy, determine whether the current MAC address is in the possible current MAC address table. If not, directly determine it as an unknown unicast. If so, use the second stage to accurately locate the position of the bucket using a hash table. The third stage is a parallel storage unit that uses the parallel storage unit of the FPGA to store multiple values to avoid hash value overwriting.

7. The high-speed and low-latency 10 Gigabit switching method based on FPGA according to claim 1, characterized in that, Read the data frame from the cache according to the forwarding decision, re-encapsulate the Ethernet header information and checksum, and control the standard transmission rate. At the same time, dynamically adjust the transmission rate, optimize the VLAN division, port rate, and queue scheduling parameters in combination with external configuration commands to achieve network congestion avoidance and low-latency transmission, including: Reconstruct and encapsulate the data frame and send it to the physical layer, synchronously update the aging time. Each time a MAC address is learned or found, update the aging time of the corresponding table entry. If the aging time of the table entry times out, delete the table entry from the MAC address table. Receive configuration commands from the external CPU through the SPI interface, set the PHY working mode, forwarding rules, and queue scheduling parameters, provide a configuration management software, and support users to configure device parameters through a graphical interface. Based on the credit-based flow control algorithm, dynamically adjust the data transmission rate to prevent network congestion, and optimize the data transmission efficiency of the sending end according to the feedback information from the receiving end. Obtain the device operating status in real time through the management interface, and support VLAN division and port rate configuration.

8. A 10 Gigabit switching system with high speed and low latency based on FPGA, characterized in that, Including: Data receiving module, which is used to send the received Ethernet data frame to the data receiving module after RS sublayer processing, store the valid data frame in the data cache module, and generate a cache write address and control signal. Address update module, which is used to read and parse the destination address, source address, and type field of the data frame from the data cache module, and update the forwarding address table through the dynamic learning mode or preset mode. Address lookup module, which is used to complete the address lookup using a three-stage pipeline architecture based on the updated forwarding address table. Data sending module, which is used to read the data frame from the cache according to the forwarding decision, re-encapsulate the Ethernet header information and checksum, and control the standard transmission rate. At the same time, dynamically adjust the transmission rate, optimize the VLAN division, port rate, and queue scheduling parameters in combination with external configuration commands to achieve network congestion avoidance and low-latency transmission.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the high-speed and low-latency 10 Gigabit switching method based on FPGA according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the FPGA-based high-speed and low-latency 10 Gigabit Ethernet switching method according to any one of claims 1 to 7.

Citation Information

Cited By

  • High-speed optical fiber data flow parallel processing and cache optimization method and system based on FPGA

    CN120825461A

  • Network system fast addressing method and device and storage medium

    CN121603475A

  • Data transmission method and device of computing system and electronic equipment

    CN122285326A