A multi-stage flow control system for lossless transmission of cross-chip traffic in switches

Through the load balancing and rate control strategies of the multi-level flow control system, cross-chip traffic distribution is dynamically adjusted, solving the data congestion and packet loss problems of multi-FPGA switch systems under high load, achieving lossless transmission, and improving network performance and stability.

CN119254724BActive Publication Date: 2025-09-30XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411334902.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-09-30
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

Multi-FPGA switch systems are prone to data congestion and packet loss under high load conditions, affecting network performance.

Method used

A multi-level flow control system is adopted, through the crossbar architecture, the series connection of the first unit and the second unit, combined with load balancing and rate control strategies, to dynamically adjust the cross-slice traffic distribution and achieve lossless transmission of cross-slice data frames.

Benefits of technology

It reduces the probability of data congestion and packet loss, improves network performance and stability, and meets the high bandwidth requirements of modern networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119254724B_ABST
    Figure CN119254724B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-stage flow control system for lossless transmission of cross-chip traffic of a switch, each flow control system comprising: a crossbar architecture, a first unit and a second unit; the crossbar architecture is used to send a current-chip cross-chip data frame to the first unit via a cross-chip bus; the first unit is used to update the address table entry of the current-chip cross-chip data frame, and under the joint action of a load balancing strategy and a rate control strategy, select a corresponding sending channel and a receiving rate matching the second unit of an adjacent flow control system, and send the current-chip cross-chip data frame to the adjacent flow control system based on the result of the address table entry update; the second unit is used to obtain adjacent cross-chip data frames, and under the joint action of a load balancing strategy and a rate control strategy, determine the destination port number according to the destination MAC address of the adjacent cross-chip data frame; the crossbar architecture is also used to forward data for the adjacent cross-chip data frame according to the destination port number.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data communication and network switching, and in particular to a multi-stage flow control system for lossless transmission of cross-chip traffic of a switch. Background Art

[0002] With the rapid development of information technology, data communications have become essential infrastructure in modern society. Switches play a crucial role in building efficient data transmission networks. Through internal logic processing, switches forward data packets from one port to another, enabling data exchange between different network nodes. However, with the surge in data traffic, switches face increasing challenges in handling large-scale data transmission.

[0003] In today's data communications landscape, with the explosive growth of network traffic, traditional single-chip FPGA switches are gradually reaching a processing bottleneck. To address this challenge, modern switch designs are trending towards a multi-chip FPGA parallel processing architecture. By integrating multiple FPGA units, this architecture exponentially increases switching capacity and effectively expands the switch's data processing capabilities. However, for multi-FPGA systems to work together efficiently, cross-chip traffic management becomes crucial. This not only affects the throughput of the entire switching system but also directly impacts the timeliness and stability of data transmission. Therefore, developing advanced cross-chip traffic control technologies has become a core technical issue for improving the performance of multi-FPGA switches.

[0004] In existing multi-FPGA switch systems, multiple FPGA channels are independent of each other, and traffic from different ports passes through fixed channels to reach another FPGA. This may lead to data congestion and packet loss under high load conditions, affecting network performance. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a multi-stage flow control system for lossless transmission of cross-chip traffic of a switch.

[0006] The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0007] The present invention provides a multi-stage flow control system for lossless transmission of cross-chip traffic of a switch, wherein a plurality of flow control systems are connected in series, and each flow control system comprises: a crossbar architecture, a first unit and a second unit;

[0008] The crossbar architecture is used to send the current slice cross-slice data frame to the first unit via the cross-slice bus;

[0009] The first unit is configured to update the address table entry for the current inter-slice data frame, and under the combined effects of the load balancing strategy and the rate control strategy, select a corresponding sending channel and a receiving rate that matches the second unit of the adjacent flow control system. Based on the result of the address table update, the current inter-slice data frame is sent to the adjacent flow control system.

[0010] The second unit is used to obtain adjacent cross-slice data frames and, under the combined effect of the load balancing strategy and the rate control strategy, search the address table entry according to the destination MAC address of the adjacent cross-slice data frames to determine the destination port number;

[0011] The crossbar architecture is also used to forward data frames across adjacent slices according to the destination port number.

[0012] Optionally, the first unit includes: an output arbitration module, a transmission load balancing module, a transmission unit, an LVDS address self-learning module, and an Aurora cross-chip transmission module; the output arbitration module, the transmission load balancing module, the transmission unit, and the Aurora cross-chip transmission module are connected in sequence; a first input end of the LVDS address self-learning module is connected to the output arbitration module, and a first output end of the LVDS address self-learning module is connected to the Aurora cross-chip transmission module;

[0013] The second unit includes: Aurora cross-chip receiving module, receiving unit, receiving load balancing module, input arbitration module and LVDS address self-learning module;

[0014] The Aurora cross-chip receiving module, receiving unit, receiving load balancing module, and input arbitration module are connected in sequence; the second input end of the LVDS address self-learning module is connected to the Aurora cross-chip receiving module, and the second output end of the LVDS address self-learning module is connected to the input arbitration module;

[0015] The input end of the output arbitration module is connected to the output end of the crossbar architecture; the output end of the input arbitration module is connected to the input end of the crossbar architecture.

[0016] Optionally, the output arbitration module is used to parse the current inter-chip data frame to obtain parsed data information, and send the parsed data information to the LVDS address self-learning module, and send the current inter-chip data frame to the transmission load balancing module;

[0017] The LVDS address self-learning module is used to update the parsed data information as an address table entry to the lookup table, obtain an updated lookup table, and send the parsed data information to the LVDS address self-learning module of the adjacent flow control system;

[0018] The transmit load balancing module selects a corresponding transmission path based on channel congestion. It then uses the transmit unit and Aurora cross-slice transmit module under the corresponding transmission path to send the current cross-slice data frame to the adjacent flow control system at the matching transmission rate. The matching transmission rate is determined based on the data congestion of the second unit in the adjacent flow control system.

[0019] The Aurora cross-chip receiving module is used to receive adjacent cross-chip data frames and adjacent parsed data information, send the adjacent cross-chip data frames to the receiving unit, and send the adjacent parsed data information to the LVDS address self-learning module of the current chip;

[0020] The receiving unit is used to convert the formats of adjacent cross-slice data frames to obtain converted data frames, and send the converted data frames to the receiving load balancing module;

[0021] The receiving load balancing module is used to cache the converted data frames and determine whether to send the converted data frames to the input arbitration module through the back pressure signal of the crossbar architecture;

[0022] When the converted data frame is sent to the input arbitration module, the input arbitration module is used to parse the destination MAC address of the converted data frame and search the updated lookup table to determine the destination port number.

[0023] Optionally, the sending load balancing module includes: an output channel selection module, multiple sending buffer modules and an NFC flow control module connected in sequence;

[0024] The output channel selection module is used to store the current slice inter-slice data frame into the destination sending buffer module with the lowest congestion level among multiple sending buffer modules according to the congestion status of the sending buffer module;

[0025] The destination sending buffer module is used to obtain a token from the NFC flow control module, and based on the token and the sending unit and Aurora cross-chip sending module under the corresponding sending path, sends the current cross-chip data frame to the adjacent flow control system at the matching sending rate.

[0026] Optionally, the NFC flow control module further includes: an NFC token bucket unit and an NFC flow control unit;

[0027] The input end of the NFC token bucket unit is connected to the output end of the multiple sending buffer modules; the output end of the NFC token bucket unit is connected to the input end of the NFC flow control unit;

[0028] The output terminal of the NFC flow control unit is connected to the input terminal of the sending unit;

[0029] The NFC flow control unit is used to monitor the load watermark of the receiving load balancing module in the adjacent flow control system and control the token generation rate of the NFC token bucket unit according to the load watermark;

[0030] The NFC token bucket unit is used to generate tokens according to the token generation rate and use the tokens to control the traffic rate of the sending buffer module.

[0031] Optionally, the LVDS address self-learning module includes: a lookup table module, a hash module, a learning table module and an LVDS cross-chip transmission module;

[0032] The output end of the hash module is connected to the input end of the lookup table module and the learning table module; the input end of the hash module is connected to the output arbitration module;

[0033] The input of the learning table module is connected to the Aurora cross-chip receiving module, and the output of the learning table module is connected to the input of the lookup table module and the LVDS cross-chip transmission module;

[0034] The output end of the lookup table module is connected to the input arbitration module; the LVDS cross-chip transmission module is connected to the Aurora cross-chip transmission module.

[0035] Optionally, the hash module is used to convert the parsed data information into converted data information that meets the query format of the lookup table module and the learning table module; wherein the data table entries within the lookup table module and the learning table module are synchronized in real time;

[0036] The lookup table module is used to perform table lookup processing using the converted data information to obtain the destination port number;

[0037] The learning table module is used to learn the address table entries of the current slice and the adjacent slices, update the learning results of the address table entries to the lookup table module, and transmit the data table entries of the current slice to the learning table module of the adjacent flow control system using the LVDS cross-slice transmission module; wherein, the current slice is the current flow control system, and the adjacent slice is the adjacent flow control system;

[0038] The LVDS cross-chip transmission module is used to convert the data table entries of the current chip into the format required by the Aurora cross-chip transmission module, and send it to the LVDS address self-learning module of the adjacent flow control system based on the Aurora cross-chip transmission module.

[0039] Optionally, the sending unit is provided with multiple sub-sending units; the Aurora cross-slice sending module is provided with multiple sub-Aurora cross-slice sending modules; the sub-sending units are connected to the sub-Aurora cross-slice sending modules in a one-to-one correspondence, and one sub-sending unit and one sub-Aurora cross-slice sending module constitute a sending path;

[0040] The sub-sending unit is configured to convert the current slice cross-slice data frame into a data format that satisfies the sub-Aurora cross-slice sending module, thereby obtaining the current slice cross-slice converted data frame.

[0041] The sub-Aurora cross-slice sending module is used to send the cross-slice converted data frame of the current slice to the Aurora cross-slice receiving module of the adjacent flow control system.

[0042] The present invention provides a multi-stage flow control system for lossless transmission of cross-chip traffic of a switch, wherein multiple flow control systems are connected in series, and each flow control system comprises: a crossbar architecture, a first unit, and a second unit; the crossbar architecture is used to send a current-chip cross-chip data frame to the first unit via a cross-chip bus; the first unit is used to update the address table entry of the current-chip cross-chip data frame, and under the joint action of a load balancing strategy and a rate control strategy, select a corresponding sending channel and a receiving rate matching the second unit of an adjacent flow control system, and send the current-chip cross-chip data frame to the adjacent flow control system based on the result of the address table entry update; the second unit is used to obtain adjacent cross-chip data frames, and under the joint action of a load balancing strategy and a rate control strategy, search the address table entry according to the destination MAC address of the adjacent cross-chip data frame to determine the destination port number; the crossbar architecture is also used to forward data for the adjacent cross-chip data frame according to the destination port number. In the present invention, since the first unit dynamically determines the corresponding sending rate based on the receiving rate of the second unit, and uses different sending channels to send data under the action of the load balancing strategy, dynamic traffic adjustment and distribution is achieved, and the distribution efficiency of cross-chip channel loads is improved. Since the traffic distribution in the embodiment of the present invention is dynamically executed, the probability of data congestion and packet loss problems is greatly reduced, and network performance is improved.

[0043] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic diagram of the structure of a multi-stage flow control system for lossless transmission of cross-chip traffic in a switch provided by an embodiment of the present invention;

[0045] Figure 2 A schematic diagram of the structure of a transmission load balancing module provided in an embodiment of the present invention;

[0046] Figure 3 A schematic diagram of the selection process of the output channel selection module provided in an embodiment of the present invention;

[0047] Figure 4 A schematic diagram of a process for flow control based on a second unit of an adjacent flow control system according to an embodiment of the present invention;

[0048] Figure 5 This is a structural diagram of the LVDS address self-learning module provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.

[0050] In order to reduce the probability of data congestion and packet loss problems and improve network performance, an embodiment of the present invention provides a multi-stage flow control system for lossless transmission of cross-chip traffic of a switch. Figure 1 The present invention provides a schematic diagram of a multi-stage flow control system for lossless transmission of cross-chip traffic in a switch. Figure 1 As shown, a plurality of fluidic control systems are connected in series, each fluidic control system comprising: a crossbar architecture, a first unit, and a second unit;

[0051] The crossbar architecture is used to send the current slice cross-slice data frame to the first unit via the cross-slice bus;

[0052] The first unit is configured to update the address table entry for the current inter-slice data frame, and under the combined effects of the load balancing strategy and the rate control strategy, select a corresponding sending channel and a receiving rate that matches the second unit of the adjacent flow control system. Based on the result of the address table update, the current inter-slice data frame is sent to the adjacent flow control system.

[0053] The second unit is used to obtain adjacent cross-slice data frames and, under the combined effect of the load balancing strategy and the rate control strategy, search the address table entry according to the destination MAC address of the adjacent cross-slice data frames to determine the destination port number;

[0054] The crossbar architecture is also used to forward data frames across adjacent slices according to the destination port number.

[0055] It should be noted that, in the embodiment of the present invention, the current slice specifically refers to the current flow control system, and the adjacent slice specifically refers to the adjacent flow control system.

[0056] An embodiment of the present invention provides a multi-stage flow control system for lossless transmission of cross-chip traffic in a switch. Since the first unit dynamically determines the corresponding sending rate based on the receiving rate of the second unit, and uses different sending channels to send data under the action of the load balancing strategy, dynamic traffic adjustment and distribution are achieved, thereby improving the efficiency of cross-chip channel load distribution. Since the traffic distribution in the embodiment of the present invention is performed dynamically, the probability of data congestion and packet loss problems is greatly reduced, thereby improving network performance.

[0057] Optionally, the first unit includes: an output arbitration module, a transmission load balancing module, a transmission unit, an LVDS address self-learning module, and an Aurora cross-chip transmission module; the output arbitration module, the transmission load balancing module, the transmission unit, and the Aurora cross-chip transmission module are connected in sequence; a first input end of the LVDS address self-learning module is connected to the output arbitration module, and a first output end of the LVDS address self-learning module is connected to the Aurora cross-chip transmission module;

[0058] The second unit includes: Aurora cross-chip receiving module, receiving unit, receiving load balancing module, input arbitration module and LVDS address self-learning module;

[0059] The Aurora cross-chip receiving module, receiving unit, receiving load balancing module, and input arbitration module are connected in sequence; the second input end of the LVDS address self-learning module is connected to the Aurora cross-chip receiving module, and the second output end of the LVDS address self-learning module is connected to the input arbitration module;

[0060] The input end of the output arbitration module is connected to the output end of the crossbar architecture; the output end of the input arbitration module is connected to the input end of the crossbar architecture.

[0061] It should be noted that the output arbitration module and the input arbitration module have the same structure, the transmit load balancing module and the receive load balancing module have the same structure, the transmit unit and the receive unit have the same structure, and the Aurora cross-slice transmit module and the Aurora cross-slice receive module have the same structure. Because these corresponding modules have the same structure and only process different data, the subsequent embodiments will use the transmit path formed by the first unit as an example to describe the data transmission process.

[0062] Optionally, the output arbitration module is used to parse the current inter-chip data frame to obtain parsed data information, and send the parsed data information to the LVDS address self-learning module, and send the current inter-chip data frame to the transmission load balancing module;

[0063] The LVDS address self-learning module is used to update the parsed data information as an address table entry to the lookup table, obtain an updated lookup table, and send the parsed data information to the LVDS address self-learning module of the adjacent flow control system;

[0064] The transmit load balancing module selects a corresponding transmission path based on channel congestion. It then uses the transmit unit and Aurora cross-slice transmit module under the corresponding transmission path to send the current cross-slice data frame to the adjacent flow control system at the matching transmission rate. The matching transmission rate is determined based on the data congestion of the second unit in the adjacent flow control system.

[0065] The Aurora cross-chip receiving module is used to receive adjacent cross-chip data frames and adjacent parsed data information, send the adjacent cross-chip data frames to the receiving unit, and send the adjacent parsed data information to the LVDS address self-learning module of the current chip;

[0066] The receiving unit is used to convert the formats of adjacent cross-slice data frames to obtain converted data frames, and send the converted data frames to the receiving load balancing module;

[0067] The receiving load balancing module is used to cache the converted data frames and determine whether to send the converted data frames to the input arbitration module through the back pressure signal of the crossbar architecture;

[0068] When the converted data frame is sent to the input arbitration module, the input arbitration module is used to parse the destination MAC address of the converted data frame and search the updated lookup table to determine the destination port number.

[0069] Optionally, the sending load balancing module includes: an output channel selection module, multiple sending buffer modules and an NFC flow control module connected in sequence;

[0070] The output channel selection module is used to store the current slice inter-slice data frame into the destination sending buffer module with the lowest congestion level among multiple sending buffer modules according to the congestion status of the sending buffer module;

[0071] The destination sending buffer module is used to obtain a token from the NFC flow control module, and based on the token and the sending unit and Aurora cross-chip sending module under the corresponding sending path, sends the current cross-chip data frame to the adjacent flow control system at the matching sending rate.

[0072] It should be noted that multiple sending buffer modules are connected in series. For ease of description, the present invention uses two sending buffer modules as an example. The two sending buffer modules are: bus 1 sending buffer module (bus 1) and bus 2 sending buffer module (bus 2). Correspondingly, Figure 2 A structural diagram of a sending load balancing module provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the selection process of the output channel selection module provided by an embodiment of the present invention.

[0073] like Figure 3 The figure shows the process of selecting multiple transmit buffer modules for enqueuing and dequeuing a current inter-slice data frame (enqueuing: data frame entering a buffer module, dequeuing: data frame exiting a buffer module). Before selecting a corresponding transmit buffer module for the current inter-slice data frame, the multiple transmit buffer modules are first checked to see whether a data frame for the port corresponding to the current inter-slice data frame has already been cached.

[0074] Case 1: If a send buffer module has already cached data frames for this port, the system determines whether the send buffer module has reached its maximum watermark (no data can be stored if it has). If so, the system applies back pressure to the crossbar switch fabric (the parent module), preventing the crossbar switch fabric from outputting new data frames. If not, the send buffer module is used as the destination send buffer module, and the current cross-slice data frame is queued to the destination send buffer module. Dequeuing data frames from the destination send buffer module requires determining whether there are tokens in the token bucket. If there are tokens, a token is taken out, and the data frames from the destination send buffer module are dequeued.

[0075] It can be understood that the above-mentioned step setting of situation 1 can prevent the data frames from being out of order.

[0076] Case 2: If multiple send buffers have not cached the data frame for the port corresponding to the current cross-slice data frame, the system checks whether all send buffers have reached their maximum watermark. If so, the crossbar switching fabric is backpressured, preventing it from outputting new data frames. If none have reached their maximum watermark, the crossbar data frame is queued to the send buffer with the lowest watermark. The dequeuing process for the send buffer is the same as in Case 1.

[0077] like Figure 2 As shown, the output channel selection module compares the watermarks of the send cache modules on Bus 1 and Bus 2, prioritizing the channel with the lower watermark for cross-slice data frame transmission. Furthermore, to prevent cross-slice data frames from being out of order due to load balancing, this module first determines whether a frame with the same MAC address as the current cross-slice data frame is cached in either cache channel before selecting a channel. If so, the channel with the same source MAC address is selected regardless of congestion. If no frame with the same MAC address as the current cross-slice data frame is cached, the channel with the lowest congestion, i.e., the channel with the lowest watermark, is selected for cross-slice transmission.

[0078] The transmit buffer module uses the FIFO (First In, First Out) principle to cache the current inter-slice data frames. In this embodiment, the FIFO can be set to a data width of 384 bits and a depth of 1024. The watermark can be set to five levels, with the corresponding FIFO depths being 400 for level 1, 700 for level 2, 800 for level 3, 900 for level 4, and 950 for level 5. A depth of at least the longest Ethernet frame is reserved.

[0079] Optionally, the NFC flow control module further includes: an NFC token bucket unit and an NFC flow control unit;

[0080] The input end of the NFC token bucket unit is connected to the output end of the multiple sending buffer modules; the output end of the NFC token bucket unit is connected to the input end of the NFC flow control unit;

[0081] The output terminal of the NFC flow control unit is connected to the input terminal of the sending unit;

[0082] The NFC flow control unit is used to monitor the load watermark of the receiving load balancing module in the adjacent flow control system and control the token generation rate of the NFC token bucket unit according to the load watermark;

[0083] The NFC token bucket unit is used to generate tokens according to the token generation rate and use the tokens to control the traffic rate of the sending buffer module.

[0084] The NFC traffic control unit is specifically used to receive the load waterline of the load balancing module and adjust the token generation speed in the NFC token bucket unit according to the load waterline to achieve the purpose of traffic speed limit or shutdown.

[0085] It should be noted that in the embodiment of the present invention, the current cross-slice data frame needs to pass through the NFC flow control module before it can cross the slice. Each time a frame passes through, a token in the NFC token bucket unit needs to be taken. When the tokens in the NFC token bucket unit are consumed, the current cross-slice data frame cannot pass through and is temporarily stored in the sending cache module, thereby achieving the purpose of traffic speed limiting. The token generation speed in the NFC token bucket unit can meet the maximum transmission speed of 25Gb / s data frames. According to the 5-level waterline of the receiving cache module of the second unit, the first unit can divide the traffic into 25Gb / s, 20Gb / s, 10Gb / s, 5Gb / s and 0Gb / s.

[0086] It is understandable that the above settings can avoid data congestion at the sending end, greatly avoiding the occurrence of data packet loss problems.

[0087] Further, Figure 4 Schematic diagram of the process of flow control according to the second unit of the adjacent flow control system provided by the embodiment of the present invention. Specifically, the NFC flow control unit of the current flow control system monitors the load waterline ( Figure 4 In the example of the 5-level waterline, the token generation rate of the NFC token bucket unit is controlled according to the load waterline to obtain an NFC flow control frame (NFC frame); the NFC token bucket unit generates tokens according to the NFC flow control frame, and uses the tokens to control the data frame sending port to send data at the corresponding flow rate.

[0088] Figure 5 Schematic diagram of the structure of the LVDS address self-learning module provided by the embodiment of the present invention. Figure 5As shown, the LVDS address self-learning module includes: a lookup table module, a hash module, a learning table module and an LVDS cross-chip transmission module;

[0089] The output end of the hash module is connected to the input end of the lookup table module and the learning table module; the input end of the hash module is connected to the output arbitration module;

[0090] The input of the learning table module is connected to the Aurora cross-chip receiving module, and the output of the learning table module is connected to the input of the lookup table module and the LVDS cross-chip transmission module;

[0091] The output end of the lookup table module is connected to the input arbitration module; the LVDS cross-chip transmission module is connected to the Aurora cross-chip transmission module.

[0092] It should be noted that, in the embodiment of the present invention, the Aurora cross-chip receiving module is used to convert the data frame from the axi format to the local link format, and the Aurora cross-chip sending module is used to convert the data frame from the local link format to the axi format.

[0093] Optionally, the hash module is used to convert the parsed data information into converted data information that meets the query format of the lookup table module and the learning table module; wherein the data table entries within the lookup table module and the learning table module are synchronized in real time;

[0094] The lookup table module is used to perform table lookup processing using the converted data information to obtain the destination port number;

[0095] The learning table module is used to learn the address table entries of the current slice and the adjacent slices, update the learning results of the address table entries to the lookup table module, and transmit the data table entries of the current slice to the learning table module of the adjacent flow control system using the LVDS cross-slice transmission module; wherein, the current slice is the current flow control system, and the adjacent slice is the adjacent flow control system;

[0096] The LVDS cross-chip transmission module is used to convert the data table entries of the current chip into the format required by the Aurora cross-chip transmission module, and send it to the LVDS address self-learning module of the adjacent flow control system based on the Aurora cross-chip transmission module.

[0097] In addition, the hash module is also used to compress the 48-bit MAC address hash into 10 bits to save space in the lookup table module.

[0098] For each new data frame entering the switch, the learning table module analyzes its source MAC address and creates a new address table entry based on its port number. The learning table module compares the newly learned address entry with the already learned address entry. If the learned address entry has not been found, a new entry is created in the address table. If the learned address entry has been found, the duplicate entry is discarded.

[0099] For the lookup table module, every time a new data frame enters the switch, its destination MAC address will be parsed, and the corresponding MAC address table entry and port number will be searched in the address table to obtain the destination port number.

[0100] Optionally, the sending unit is provided with multiple sub-sending units; the Aurora cross-slice sending module is provided with multiple sub-Aurora cross-slice sending modules; the sub-sending units are connected to the sub-Aurora cross-slice sending modules in a one-to-one correspondence, and one sub-sending unit and one sub-Aurora cross-slice sending module constitute a sending path;

[0101] The sub-sending unit is configured to convert the current slice cross-slice data frame into a data format that satisfies the sub-Aurora cross-slice sending module, thereby obtaining the current slice cross-slice converted data frame.

[0102] The sub-Aurora cross-slice sending module is used to send the cross-slice converted data frame of the current slice to the Aurora cross-slice receiving module of the adjacent flow control system.

[0103] In summary, the present invention provides a multi-level flow control system for lossless transmission of cross-chip traffic of a switch. Through the sending load balancing unit, the system can monitor the congestion status of the cross-chip channel in real time, and dynamically adjust the traffic distribution, giving priority to channels with lower congestion for data transmission, thereby improving the efficiency of cross-chip channel load distribution. In addition, the built-in GTY high-speed interface in the Aurora cross-chip sending module can achieve a data transmission rate of up to 100Gb / s, meeting the high bandwidth requirements of modern networks. The NFC traffic control unit in the sending load balancing unit sends a control signal in time to adjust the sending port status according to the congestion status of the receiving channel, effectively preventing the loss of data frames due to insufficient processing power of the receiving end, and ensuring the integrity of data transmission. The LVDS address self-learning cross-chip unit learns the correspondence between the switch input port and the MAC address, and transmits it to another flow control system using the LVDS interface, thereby realizing cross-chip synchronization of address table entries, ensuring that data frames can be looked up and forwarded according to the correct address information.

[0104] The system design of this invention solves the problem of out-of-order data frame transmission caused by load balancing, ensures the sequential nature of data frames, and improves data transmission reliability. Through hierarchical token bucket management, the system can flexibly adjust flow control strategies based on varying traffic demands and network conditions, achieving more refined traffic management. The collaborative working mechanism of the multi-level flow control system improves the stability and reliability of the switch in complex network environments and high load conditions.

[0105] In the description of this specification, the reference terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" mean that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.

[0106] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the above-mentioned disclosed embodiments by viewing the drawings and the disclosed content. In the description of the present invention, the word "comprising" does not exclude other components or steps, "one" or "an" does not exclude multiple situations, and the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0107] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention cannot be considered to be limited to these descriptions. For those skilled in the art of the present invention, several simple deductions or substitutions can be made without departing from the concept of the present invention, and all of these should be considered to fall within the scope of protection of the present invention.

Claims

1. A multi-stage flow control system for lossless transmission of cross-chip traffic in a switch, characterized in that: A plurality of fluidic control systems are connected in series, each of the fluidic control systems comprising: a crossbar architecture, a first unit, and a second unit; The crossbar architecture is used to send the current slice cross-chip data frame to the first unit through the cross-chip bus; The first unit is configured to update an address table entry for the current inter-slice data frame, select a corresponding sending channel and a receiving rate that matches the second unit of the adjacent flow control system under the combined effect of a load balancing strategy and a rate control strategy, and send the current inter-slice data frame to the adjacent flow control system based on a result of the address table entry update; The second unit is configured to obtain adjacent cross-slice data frames, and under the combined effect of the load balancing strategy and the rate control strategy, search the address table entry according to the destination MAC address of the adjacent cross-slice data frames to determine the destination port number; The crossbar architecture is further configured to forward the adjacent cross-slice data frames according to the destination port number.

2. The multi-stage flow control system for lossless transmission of cross-chip traffic of a switch according to claim 1, characterized in that: The first unit includes: an output arbitration module, a transmission load balancing module, a transmission unit, an LVDS address self-learning module, and an Aurora cross-chip transmission module; the output arbitration module, the transmission load balancing module, the transmission unit, and the Aurora cross-chip transmission module are connected in sequence; a first input end of the LVDS address self-learning module is connected to the output arbitration module, and a first output end of the LVDS address self-learning module is connected to the Aurora cross-chip transmission module; The second unit includes: an Aurora cross-chip receiving module, a receiving unit, a receiving load balancing module, an input arbitration module, and an LVDS address self-learning module; The Aurora cross-chip receiving module, the receiving unit, the receiving load balancing module, and the input arbitration module are connected in sequence; the second input end of the LVDS address self-learning module is connected to the Aurora cross-chip receiving module, and the second output end of the LVDS address self-learning module is connected to the input arbitration module; The input end of the output arbitration module is connected to the output end of the crossbar architecture; the output end of the input arbitration module is connected to the input end of the crossbar architecture.

3. The multi-stage flow control system for lossless transmission of cross-chip traffic of a switch according to claim 2, characterized in that: The output arbitration module is used to parse the current inter-chip data frame to obtain parsed data information, send the parsed data information to the LVDS address self-learning module, and send the current inter-chip data frame to the transmission load balancing module; The LVDS address self-learning module is used to update the parsed data information as an address table entry to a lookup table to obtain an updated lookup table, and send the parsed data information to the LVDS address self-learning module of an adjacent flow control system; The transmission load balancing module is configured to select a corresponding transmission path based on the channel congestion situation, and send the current slice cross-slice data frame to the adjacent flow control system at a matched transmission rate through the transmission unit and the Aurora cross-slice transmission module under the corresponding transmission path. The matched transmission rate is obtained based on the data congestion situation of the second unit in the adjacent flow control system. The Aurora cross-chip receiving module is used to receive adjacent cross-chip data frames and adjacent parsed data information, send the adjacent cross-chip data frames to the receiving unit, and send the adjacent parsed data information to the LVDS address self-learning module of the current chip; The receiving unit is configured to perform format conversion on the adjacent cross-slice data frames to obtain converted data frames, and send the converted data frames to the receiving load balancing module; The receiving load balancing module is used to cache the converted data frame and determine whether to send the converted data frame to the input arbitration module through the back pressure signal of the crossbar architecture; When the converted data frame is sent to the input arbitration module, the input arbitration module is configured to parse the destination MAC address of the converted data frame and search the updated lookup table to determine the destination port number.

4. The multi-stage flow control system for lossless transmission of cross-chip traffic of a switch according to claim 3, characterized in that: The transmission load balancing module includes: an output channel selection module, multiple transmission buffer modules and an NFC flow control module connected in sequence; The output channel selection module is configured to store the current slice inter-slice data frame into a destination sending buffer module with the lowest congestion level among the multiple sending buffer modules according to the congestion level of the sending buffer module; The destination sending buffer module is used to obtain a token from the NFC flow control module, and based on the token and the sending unit and the Aurora cross-chip sending module under the corresponding sending path, send the current cross-chip data frame to the adjacent flow control system at a matched sending rate.

5. The multi-stage flow control system for lossless transmission of cross-chip traffic of a switch according to claim 4, characterized in that: The NFC flow control module further includes: an NFC token bucket unit and an NFC flow control unit; The input end of the NFC token bucket unit is connected to the output end of the multiple sending buffer modules; the output end of the NFC token bucket unit is connected to the input end of the NFC flow control unit; The output end of the NFC flow control unit is connected to the input end of the sending unit; The NFC flow control unit is used to monitor the load waterline of the receiving load balancing module in the adjacent flow control system, and control the token generation rate of the NFC token bucket unit according to the load waterline; The NFC token bucket unit is used to generate tokens according to the token generation rate, and use the tokens to control the flow rate of the sending buffer module.

6. The multi-stage flow control system for lossless transmission of cross-chip traffic of a switch according to claim 3, characterized in that: The LVDS address self-learning module includes: a lookup table module, a hash module, a learning table module and an LVDS cross-chip transmission module; The output end of the hash module is connected to the input end of the lookup table module and the learning table module; the input end of the hash module is connected to the output arbitration module; The input end of the learning table module is connected to the Aurora cross-chip receiving module, and the output end of the learning table module is connected to the input end of the lookup table module and the LVDS cross-chip transmission module; The output end of the lookup table module is connected to the input arbitration module; the LVDS cross-chip transmission module is connected to the Aurora cross-chip transmission module.

7. The multi-stage flow control system for lossless transmission of cross-chip traffic of a switch according to claim 6, characterized in that: The hash module is used to convert the parsed data information into converted data information that meets the query format of the lookup table module and the learning table module; wherein the internal data table entries of the lookup table module and the learning table module are synchronized in real time; The table lookup module is used to perform table lookup processing using the converted data information to obtain the destination port number; The learning table module is used to learn the address table entries of the current slice and the adjacent slices, update the learning results of the address table entries to the lookup table module, and transmit the data table entries of the current slice to the learning table module of the adjacent flow control system using the LVDS cross-chip transmission module; wherein the current slice is the current flow control system and the adjacent slice is the adjacent flow control system; The LVDS cross-chip transmission module is used to convert the data table entries of the current slice into a format that meets the requirements of the Aurora cross-chip sending module, and send the data to the LVDS address self-learning module of the adjacent flow control system based on the Aurora cross-chip sending module.

8. The multi-stage flow control system for lossless transmission of cross-chip traffic of a switch according to claim 2, characterized in that: The sending unit is provided with multiple sub-sending units; the Aurora cross-slice sending module is provided with multiple sub-Aurora cross-slice sending modules; the sub-sending units are connected to the sub-Aurora cross-slice sending modules in a one-to-one correspondence, and one sub-sending unit and one sub-Aurora cross-slice sending module form a sending path; The sub-slice sending unit is configured to convert the current slice cross-slice data frame into a data format that meets the data format of the sub-Aurora cross-slice sending module to obtain a current slice cross-slice converted data frame; The sub-Aurora cross-slice sending module is used to send the current slice cross-slice converted data frame to the Aurora cross-slice receiving module of the adjacent flow control system.