Crossbar switching network system based on shared cross nodes
By completing arbitration requests and responses within the FPGA, cross-chip transmission latency is reduced, solving the problem of reduced scheduling efficiency in two-FPGA systems and enabling more efficient expansion of switching capacity and port count.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIDIAN UNIV
- Filing Date
- 2023-04-12
- Publication Date
- 2026-05-05
AI Technical Summary
When using two FPGAs to implement a switching system, there is a time delay in cross-chip data or signal interaction, which leads to a decrease in scheduling efficiency.
A crossbar switching network system based on shared cross nodes is adopted, including an input queue management module, a normal cross node buffer module, a shared cross node buffer module, an RR column arbitration module, a high-speed Aurora interface module, and a WRR independent column arbitration module. By completing arbitration requests and responses within the same FPGA, cross-chip transmission latency is reduced.
It improves the scheduling efficiency of the switching system, reduces cross-chip transmission latency, expands the switching capacity and number of ports, solves the problems of scheduling efficiency and arbitration fairness, and improves link transmission efficiency.
Smart Images

Figure CN116455839B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communication technology, specifically relating to a crossbar switching network system based on a shared cross node. Background Technology
[0002] Currently, the mainstream single-stage switching structures are mainly divided into three types: shared bus, shared buffer, and crossbar. In a shared bus structure, all data from all input ports is transmitted on the bus in a time-division multiplexed manner. Therefore, a shared bus structure requires the bus speed to be greater than the sum of the speeds of all ports to ensure no blocking occurs. Because the switching capacity of a shared bus structure is limited by the bus speed and its scalability is low, this structure is generally not used. A shared buffer structure achieves data frame reception and transmission by reading and writing to the same buffer area. Compared to a shared bus structure, it is easier to achieve line-speed data processing. However, the switching capacity of a single shared buffer is limited by the buffer's write and read rates, and it also suffers from the problem of not being able to freely expand. A crossbar switching structure can effectively solve the problem of limited switching capacity in shared bus and shared buffer structures. A crossbar switching structure uses a high-speed crossbar switch matrix circuit to achieve multi-input to multi-output channel switching. Switching any input to output channel does not affect other connected channels, achieving strictly non-blocking. Crossbar switching networks are categorized by queuing strategy into Input Queued (IQ), Output Queued (OQ), Combined Input and Output Queued (CIOQ), and Combined Input and Crosspoint Queued (CICQ). Among these, Combined Input and Crosspoint Queued (CICQ) effectively isolates input and output terminals and facilitates capacity expansion, making it a widely adopted switching network structure. In CICQ, the shared buffers of the input queue management module and the crossbar buffers consume significant storage resources. When a single FPGA is insufficient to support the storage requirements of the current switching capacity, using multiple FPGAs can effectively address this issue. If the number of physical ports to be handled by the switch exceeds the capacity of a single FPGA, using two FPGAs to jointly perform the switching function is a viable solution. However, implementing a switching system using two FPGAs inherently presents a significant latency issue in data or signal interaction between the two FPGAs. Taking horizontal partitioning as an example, horizontal partitioning retains the logic from the input bus to the cross-node in the same row. Therefore, the interaction between the input queue management module and the cross-node buffer in the same row is completed within the same FPGA. However, when arbitrating the output of the cross-node buffer in the same column, the transmitted data and interaction signals need to be transmitted across the FPGA. How to mitigate the reduction in scheduling efficiency caused by transmitting data or signals across the FPGA has become a problem to be solved. Summary of the Invention
[0003] To address the aforementioned problems in the existing technology, this invention provides a horizontally splitting Crossbar switching network system based on shared cross nodes. The technical problem to be solved by this invention is achieved through the following technical solution:
[0004] This invention provides a crossbar switching network system based on a shared crossbar node, comprising: several input queue management modules, a regular crossbar node buffer module, a shared crossbar node buffer module, an RR column arbitration module, a high-speed Aurora interface module, a WRR independent column arbitration module, and several configuration interfaces, all configured on each FPGA.
[0005] The input queue management modules are used to receive data frames from physical ports and add a tag header to the data frame header according to the destination port of the data frame;
[0006] The ordinary cross-node buffer module is used to receive data frames with added tag headers and temporarily store the data frames with added tag headers into different ordinary cross-node buffers according to the destination port;
[0007] The shared cross-node cache module is used to receive and store data frames from another FPGA in the inter-chip high-speed Aurora interface module, and participate in the arbitration of the WRR independent column arbitration module together with the data frames in the ordinary cross-node cache in the same column on the local FPGA.
[0008] The RR column arbitration module is used to provide an arbitration result based on the transmission request of the ordinary cross node buffer on the same column, and according to the arbitration result, the data frame with added tag header in this FPGA is moved from the corresponding ordinary cross node buffer to the transmission part of the inter-chip high-speed Aurora interface module to form a data frame to be transmitted.
[0009] The inter-chip high-speed Aurora interface module is used to send the data frame to be sent to the receiving part of the inter-chip high-speed Aurora interface module of another FPGA, and to receive the data frame output by the sending part of the inter-chip high-speed Aurora interface module of another FPGA, and transmit it to the shared cross node buffer module of this FPGA.
[0010] The WRR independent column arbitration module is used to receive request signals sent by the cross-node caches of the same column in the ordinary cross-node cache module and the shared cross-node cache module, and to perform weighted round-robin scheduling, and to move the data frames in the cross-node caches that are scheduled in the round-robin to the output port;
[0011] The configuration interfaces are used to configure the weights required for scheduling by the WRR independent column arbitration module and the maximum and minimum thresholds of the queues of the input queue management modules.
[0012] In one embodiment of the present invention, the ordinary cross-node cache module includes a plurality of ordinary cross-node caches, the shared cross-node cache module includes a plurality of shared cross-node caches, the RR column arbitration module includes a plurality of RR column arbitration sub-modules, the inter-chip high-speed Aurora interface module includes a plurality of first inter-chip high-speed Aurora interfaces and a plurality of second inter-chip high-speed Aurora interfaces, and the WRR independent column arbitration module includes a plurality of WRR independent column arbitration sub-modules, wherein...
[0013] The ordinary cross node caches are arranged in an array, and each row of ordinary cross node caches is connected to the input queue management module.
[0014] Each of the shared cross node caches is connected to the ordinary cross node caches in the same column, each of the RR column arbitration submodules is connected to the ordinary cross node caches in the same column, and the sum of the number of the shared cross node caches and the number of the RR column arbitration submodules is equal to the number of columns of the ordinary cross node cache;
[0015] The first inter-chip high-speed Aurora interface connects the RR column arbitration submodule of this FPGA to the receiving part of the inter-chip high-speed Aurora interface module of another FPGA, and the second inter-chip high-speed Aurora interface connects the shared cross node buffer of this FPGA to the transmitting part of the inter-chip high-speed Aurora interface module of another FPGA.
[0016] Each of the WRR independent column arbitration submodules is connected to the ordinary cross node cache in the same column and is located in the same column as the shared cross node cache.
[0017] In one embodiment of the present invention, the first inter-chip high-speed Aurora interface includes a first Aurora IP core, a first cross-clock module, and a locallink-to-AXI module; the second inter-chip high-speed Aurora interface includes a second Aurora IP core, a second cross-clock module, and an AXI-to-locallink module.
[0018] The locallink-to-AXI module is used to convert the format of the data frame to be sent into the AXI data format to obtain a format-converted data frame; the first cross-clock module is used to cross the format-converted data frame from the system master clock domain to the user-side clock domain of the Aurora IP core to obtain a first cross-clock data frame; the first Aurora IP core is used to send the first cross-clock data frame to the receiving part of the high-speed Aurora interface module between another FPGA chip.
[0019] The second Aurora IP core is used to receive data frames output from the transmitting part of the high-speed Aurora interface module between another FPGA chip, and obtain a received data frame; the second cross-clock module is used to cross the received data frame from the user-side clock domain of the Aurora IP core to the system master clock domain, and obtain a second cross-clock data frame; the AXI to locallink module is used to convert the format of the second cross-clock data frame into the locallink data format used internally by the system.
[0020] In one embodiment of the present invention, both the first Aurora IP core and the second Aurora IP core adopt a 4-way bonding configuration, with a line rate of up to 40Gbps.
[0021] In one embodiment of the present invention, the number of input queue management modules on each FPGA is n / 2, where n is the number of buses; the number of ordinary cross-node buffers is n. 2 / 2; the number of shared cross node caches is n / 2; the number of RR column arbitration submodules is n / 2; the number of the first inter-chip high-speed Aurora interfaces is n / 2; the number of the second inter-chip high-speed Aurora interfaces is n / 2; and the number of WRR independent column arbitration modules is n / 2.
[0022] In one embodiment of the present invention, both the ordinary cross-node cache and the shared cross-node cache include a first storage area and a second storage area, wherein,
[0023] The first storage area is used for multicast or broadcast data, and the second storage area is used for storing unicast data.
[0024] In one embodiment of the present invention, the tag header includes a unicast / multicast identifier, a queue number, a frame length, and a destination port number.
[0025] In one embodiment of the present invention, the data frame transmission path with added tag headers is divided into a local transmission path and a cross-segment transmission path according to the destination port, wherein,
[0026] The transmission path of this chip is as follows: the input queue management module of this FPGA, the ordinary cross node buffer module of this FPGA, the WRR independent column arbitration module of this FPGA, and the output port corresponding to the destination port;
[0027] The cross-chip transmission path is as follows: the input queue management module of this FPGA, the ordinary cross node buffer module of this FPGA, the RR column arbitration module of this FPGA, the high-speed Aurora interface module of this FPGA, the high-speed Aurora interface module of the other FPGA, the shared cross node buffer module of the other FPGA, the WRR independent column arbitration module of the other FPGA, and the output port corresponding to the destination port.
[0028] In one embodiment of the present invention, the scheduling period of a data frame in the local transmission path includes a retrieval interval and a data transmission delay, and the scheduling period of a data frame in the cross-slice transmission path includes a retrieval interval and a data transmission delay.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. The Crossbar switching network system based on shared cross-nodes of this invention horizontally splits the Crossbar switching network. It adopts a shared cross-node approach, including a shared cross-node buffer module, an RR column arbitration module, and a high-speed Aurora interface module. As long as the shared cross-node can store the next longest frame (i.e., the almost signal is not pulled high), it directly receives data from another FPGA and then performs the arbitration request operation. In this way, the arbitration request and arbitration response are completed within the chip. The link transmission efficiency is the same as the data scheduling efficiency within the same chip. The only difference is that the data frame comes from another chip, so the switching latency will be greater than when implementing Crossbar with a single FPGA. Using a shared cross-node buffer avoids the problem of reduced scheduling efficiency caused by the latency of inter-chip arbitration. Theoretically, as long as the shared cross-node buffer is large enough, the WRR independent column arbitration scheduling of horizontal splitting has the same efficiency as the scheduling without splitting, improving the scheduling efficiency of the entire switching system and thus improving the link transmission efficiency.
[0031] 2. This invention adds a tag header to the data frame header based on the destination port of the data frame, which can effectively solve the problem of reduced scheduling efficiency caused by the asynchronous arrival of key information and data frames on another FPGA. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the structure of a crossbar switching network system based on a shared cross node, provided in an embodiment of the present invention.
[0033] Figure 2 This is a schematic diagram of the tag header structure provided in an embodiment of the present invention;
[0034] Figure 3This is a schematic diagram of the structure of the first inter-chip high-speed Aurora interface and the second inter-chip high-speed Aurora interface provided in an embodiment of the present invention;
[0035] Figure 4 This is a schematic diagram of two typical data paths provided in embodiments of the present invention;
[0036] Figure 5 This is a diagram illustrating the scheduling cycle of the local packet in an embodiment of the present invention.
[0037] Figure 6 This is a cross-slice group scheduling cycle analysis diagram provided in an embodiment of the present invention;
[0038] Figure 7 A diagram illustrating the scheduling cycle of shared cross-nodes used across different chips, provided in an embodiment of the present invention.
[0039] Figure 8 This is a simulation diagram of the cross-slice transmission path of the design and system of the horizontally split Crossbar switching network based on shared cross nodes provided in an embodiment of the present invention. Detailed Implementation
[0040] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0041] Example 1
[0042] To address the limited switching capacity achievable by a single FPGA and the cross-chip data and signal transmission latency issues when multiple FPGAs are used to construct a switching system, the Crossbar switching network can be split horizontally or vertically. Vertical splitting preserves the cross-node buffers and column arbitration logic for a single column of the Crossbar on the same FPGA, but data and control signals for the cross-node buffers and input queue management module in the same row still need to be transmitted across chips. Therefore, regardless of whether it's horizontal or vertical splitting, cross-chip transmission latency for data and control signals exists. Vertical splitting without introducing shared cross-node buffers leads to significant latency, resulting in decreased scheduling speed and issues with scheduling fairness between ports. Introducing shared cross-node buffers into the vertical split structure introduces head-of-line blocking. Therefore, a horizontally split Crossbar switching network based on shared cross-nodes is chosen here.
[0043] This embodiment provides a horizontally split Crossbar switching network system based on a shared cross node. This system includes an input queue management module 10, a horizontally split Crossbar switching network module, and a configuration interface 70. The input queue management module receives data frames from stream classification and packet processing, stores them in corresponding virtual output queues according to their destination port number and priority, and moves the data frames from the virtual output queues to the internal buffer of the Crossbar switching network module when it is ready. The horizontally split Crossbar switching network module performs data exchange between the input and output ports of the two FPGAs based on the destination port number provided by the input queue management module 10. The configuration interface 70 configures the required parameters within the input queue management module and the polling weights of each node in the WRR independent column arbitration module.
[0044] Please see Figure 1 , Figure 1 This is a schematic diagram of a horizontally split Crossbar switching network system based on a shared cross node, provided by an embodiment of the present invention. Specifically, the Crossbar switching network is horizontally split across two FPGAs. Each FPGA is equipped with several input queue management modules 10, ordinary cross node buffer modules 20, shared cross node buffer modules 30, RR column arbitration modules 40, high-speed Aurora interface modules 50, WRR independent column arbitration modules 60, and several configuration interfaces 70. The ordinary cross node buffer modules 20, shared cross node buffer modules 30, RR column arbitration modules 40, high-speed Aurora interface modules 50, and WRR independent column arbitration modules 60 form the horizontally split Crossbar switching network modules. The two FPGAs communicate using the inter-chip high-speed Aurora interface 50.
[0045] Specifically, each bus is equipped with one input queue management module 10. Each input queue management module 10 receives data frames from physical ports and adds a tag header to the data frame header according to the destination port of the data frame. The ordinary cross-node buffer module 20 receives the data frames with added tag headers and temporarily stores the data frames with added tag headers into different ordinary cross-node buffers 201 according to the destination port. The shared cross-node buffer module 30 receives and stores data frames from another FPGA in the inter-chip high-speed Aurora interface module 50, and participates in the arbitration of the WRR independent column arbitration module 60 with the data frames in the same column of the ordinary cross-node buffer 201 on the local FPGA. That is, the shared cross-node buffer module 30 is used to store the data frames output by the ordinary cross-node buffer module in the same column from another FPGA after RR column arbitration. The RR column arbitration module 40 is used to provide an arbitration result based on the transmission request of the ordinary cross-node buffer 201 in the same column, and according to the arbitration result, moves the data frame with added tag header in the local FPGA from the corresponding ordinary cross-node buffer 201 to the transmission part of the inter-chip high-speed Aurora interface module 50 to form a data frame to be transmitted. That is, the RR column arbitration module uses a fair round-robin method to aggregate the data in the same column that needs to be output across chips to another FPGA into one output. The inter-chip high-speed Aurora interface module 50 is used to send the data frame to be transmitted to the receiving part of the inter-chip high-speed Aurora interface module 50 of another FPGA, and receive the data frame output by the transmission part of the inter-chip high-speed Aurora interface module 50 of another FPGA, and transmit it to the shared cross-node buffer module 30 of the local FPGA. The WRR independent column arbitration module 60 is used to receive the request signals sent by the cross-node buffers in the same column in the ordinary cross-node buffer module 20 and the shared cross-node buffer module 30 and perform weighted round-robin scheduling, moving the data frame in the cross-node buffer that is scheduled to be transmitted to the output port. Several configuration interfaces 70 are used to configure the weights required for scheduling by the WRR independent column arbitration module 60 and parameters such as the maximum and minimum thresholds of the queues of several input queue management modules 10.
[0046] This embodiment employs WRR independent column arbitration scheduling to alleviate the scheduling fairness problem caused by the convergence of multiple data streams into a shared cross node.
[0047] Please see Figure 2 , Figure 2 This is a schematic diagram of the tag header structure provided in an embodiment of the present invention. Specifically, the tag header includes key information such as unicast / multicast identifier, queue number, frame length, and destination port number. Furthermore, data is transmitted between the same FPGA and across FPGAs using the tag header plus the original data, and the WRR column arbitration module 60 parses and deletes the tag header.
[0048] It is understood that the system in the above embodiment sets up a horizontally split Crossbar switching network on each FPGA, which can be used for data transmission between two FPGAs or for data transmission between multiple FPGAs. The two or more FPGAs transmit data through the high-speed Aurora interface module 50.
[0049] In one specific embodiment, the ordinary cross-node cache module 20 includes a plurality of ordinary cross-node caches 201, the shared cross-node cache module 30 includes a plurality of shared cross-node caches 301, the RR column arbitration module 40 includes a plurality of RR column arbitration sub-modules 401, and the inter-chip high-speed Aurora interface module 50 includes a plurality of first inter-chip high-speed Aurora interfaces 501 and a plurality of second inter-chip high-speed Aurora interfaces 502.
[0050] The ordinary cross-node caches 201 are arranged in an array, and each row of ordinary cross-node caches 201 is connected to the input queue management module 10. Each shared cross-node cache 301 is connected to the ordinary cross-node caches 201 in the same column, and each RR column arbitration submodule 401 is connected to the ordinary cross-node caches 201 in the same column. The sum of the number of shared cross-node caches 301 and RR column arbitration submodules 401 is equal to the number of columns of ordinary cross-node caches 201. The first inter-chip high-speed Aurora interface 501 connects the RR column arbitration submodule 401 of the FPGA on this chip to the receiving part of the inter-chip high-speed Aurora interface module 50 of another FPGA. The second inter-chip high-speed Aurora interface 502 connects the shared cross-node cache 301 of the FPGA on this chip to the transmitting part of the inter-chip high-speed Aurora interface module 50 of another FPGA. Each WRR independent column arbitration submodule 601 is connected to the ordinary cross-node cache 201 in the same column and is located in the same column as the shared cross-node cache 301.
[0051] It is understandable that the sum of the number of shared cross-node caches 301 and RR column arbitration submodules 401 is equal to the number of columns formed by several ordinary cross-node caches 201, the sum of the number of first inter-slice high-speed Aurora interfaces 501 and several second inter-slice high-speed Aurora interfaces 502 is equal to the number of columns formed by several ordinary cross-node caches 201, and the number of WRR independent column arbitration modules 60 is equal to the number of shared cross-node caches 301.
[0052] Specifically, when this system is used for data transmission between two FPGAs, the number of input queue management modules 10 on each FPGA is n / 2, where n is the number of buses; the number of ordinary cross-node buffers 201 is n. 2 / 2; the number of shared cross node caches 301 is n / 2; the number of RR column arbitration submodules 401 is n / 2; the number of first inter-chip high-speed Aurora interfaces 501 is n / 2; the number of second inter-chip high-speed Aurora interfaces 502 is n / 2; and the number of WRR independent column arbitration modules 60 is n / 2.
[0053] It should be noted that the system in the above embodiment is applicable to data transmission between two FPGAs, wherein the number of buses on the FPGA can be 4×4, 6×6, 8×8, etc.
[0054] When the number of buses on the two FPGAs is 4×4, the number of buses on each FPGA is 4, the number of input queue management modules 10 is 2, the number of ordinary cross node buffers 201 is 8, distributed in 2×4, the number of shared cross node buffers 301 is 2, the number of RR column arbitration submodules 401 is 2, the number of high-speed Aurora interfaces 501 between the first FPGAs is 2, the number of high-speed Aurora interfaces 502 between the second FPGAs is 2, and the number of WRR independent column arbitration modules 60 is 2.
[0055] When the number of buses on the two FPGAs is 6×6, the number of buses on each FPGA is 6, the number of input queue management modules 10 is 3, the number of ordinary cross-node buffers 201 is 18, distributed in 3×6, the number of shared cross-node buffers 301 is 3, the number of RR column arbitration submodules 401 is 3, the number of high-speed Aurora interfaces 501 between the first FPGAs is 3, the number of high-speed Aurora interfaces 502 between the second FPGAs is 3, and the number of WRR independent column arbitration modules 60 is 3.
[0056] In one specific embodiment, both the ordinary cross-node cache 201 and the shared cross-node cache 301 include a first storage area and a second storage area, wherein the first storage area is used for multicast or broadcast data, and the second storage area is used for storing unicast data.
[0057] This embodiment stores multicast or broadcast data and unicast data in the ordinary cross-node cache and the shared cross-node cache partitions, which can give multicast or broadcast data a higher priority when it is scheduled, so as to ensure the service quality of the system.
[0058] Please see Figure 3 , Figure 3This is a schematic diagram of the structure of the first inter-chip high-speed Aurora interface and the second inter-chip high-speed Aurora interface provided in an embodiment of the present invention. The first inter-chip high-speed Aurora interface 501 includes a first Aurora IP core, a first cross-clock module crossbar2Aurora, and a locallink to AXI module locallink2aix. The second inter-chip high-speed Aurora interface 502 includes a second Aurora IP core, a second cross-clock module Aurora2crossbar, and an AXI to locallink module axi2loacllink.
[0059] The locallink-to-AXI module locallink2aix is used to convert the format of the data frame to be sent into the AXI data format to obtain the format-converted data frame; the first cross-clock module crossbar2Aurora is used to cross the format-converted data frame from the system master clock domain to the user-side clock domain of the Aurora IP core to obtain the first cross-clock data frame; the first Aurora IP core is used to send the first cross-clock data frame to the receiving part of the high-speed Aurora interface module 50 between another FPGA chip.
[0060] The second Aurora IP core is used to receive data frames output from the transmitting part of another FPGA inter-chip high-speed Aurora interface module (50) to obtain received data frames; the second cross-clock module Aurora2crossbar is used to cross the received data frames from the user-side clock domain of the Aurora IP core to the system master clock domain to obtain the second cross-clock data frame; the AXI to locallink module axi2loacllink is used to convert the format of the second cross-clock data frame into the locallink data format used internally by the system.
[0061] Specifically, both the first and second Aurora IP cores use a 4-way bonding configuration, with a line rate of up to 40Gbps.
[0062] Please see Figure 4 , Figure 4 This diagram illustrates two typical data paths provided in embodiments of the present invention. Specifically, the data frame transmission path with an added tag header is divided into a local transmission path and a cross-chip transmission path based on the destination port.
[0063] The transmission path for this chip does not require cross-chip transmission. The path is as follows: the input queue management module 10 of this FPGA, the general cross-node buffer module 20 of this FPGA, the WRR independent column arbitration module 60 of this FPGA, and the output port corresponding to the destination port. Specifically, within the input queue management module 10, the data frame is enqueued based on key information such as the destination port number, the data frame priority, and the frame length, and is temporarily stored in the shared buffer of the input queue management module 10. When the general cross-node buffer 201 in the same row is ready to receive data frames and can store the next longest frame, it pulls the ready signal to the input queue management module 10 high. The input queue management module 10 internally polls each virtual output queue, and assembles the key information of the data frame (unicast / multicast identifier, frame length, destination port number, etc.) into a tag header, which is placed at the beginning of the data frame and moved together with the data frame from the shared buffer of the input queue management module 10 to the corresponding general cross-node buffer 201 in the same row. If the destination port of the data frame is on the FPGA chip, the data frame will directly participate in the WRR polling and scheduling on the chip. When the WRR independent column arbitration module 60 grants the request response to the ordinary cross node buffer 201, the data frame will be output from the ordinary cross node buffer 201 to the WRR independent column arbitration module 60. The WRR independent column arbitration module 60 will parse the key information in the tag, such as the destination port number, and use the key information to output the data frame to the output port with the corresponding destination port number.
[0064] The cross-chip transmission path is as follows: the input queue management module 10 of this FPGA, the ordinary cross-node buffer module 20 of this FPGA, the RR column arbitration module 40 of this FPGA, the high-speed Aurora interface module 50 of this FPGA, the high-speed Aurora interface module 50 of the other FPGA, the shared cross-node buffer module 30 of the other FPGA, the WRR independent column arbitration module 60 of the other FPGA, and the output port corresponding to the destination port. Specifically, within the input queue management module 10, the data frame is enqueued based on key information such as the destination port number, the data frame priority, and the frame length, and the data frame is temporarily stored in the shared buffer of the input queue management module 10. When the ordinary cross-node buffer 201 in the same row is ready to receive data frames, it pulls the ready signal to the input queue management module 10 high. The input queue management module 10 internally polls each virtual output queue and places a tag header composed of key information of the data frame (unicast / multicast identifier, frame length, destination port number, etc.) at the beginning of the data frame. This tag header is then moved from the shared buffer of the input queue management module 10 to the corresponding ordinary cross-node buffer 201 in the same row along with the data frame. If the destination port of the data frame is on another FPGA, the data frame participates in the RR column arbitration on the same column while it is temporarily stored in the ordinary cross-node buffer 201 of this FPGA. When the RR column arbitration module 40 issues a request response grant to the ordinary cross-node buffer 201, the data frame is output from the ordinary cross-node buffer 201 to the RR column arbitration module 40. The RR column arbitration module 40 directly passes the data through to the inter-chip Aurora high-speed interface module 50. After the inter-chip Aurora high-speed interface module 50 performs format conversion and cross-clock domain operation on the data frame, it sends the data frame to the other FPGA through the Aurora IP core. After receiving the data frame, the other FPGA also performs cross-clock domain and format conversion operations on the data frame. Next, when the shared cross-node buffer 301 is ready to receive the data frame (when the almost signal is not high), the data frame is moved from the receive buffer of the inter-chip Aurora high-speed interface module 50 to the shared cross-node buffer module 30. Next, the shared cross-node buffer 301, along with the other two ordinary cross-node buffers 201 in the same column, participates in the WRR polling scheduling for this column. When the WRR independent column arbitration module 60 grants a request response to the shared cross-node buffer module 30, the data frame is output from the shared cross-node buffer module 30 to the WRR independent column arbitration module 60. The WRR independent column arbitration module 60 parses the key information within the tag, such as the destination port number, and uses this key information to output the data frame to the output port corresponding to the destination port number.
[0065] Please see Figure 5 , Figure 5 This is a diagram illustrating the scheduling cycle of the local packet, provided as an embodiment of the present invention. (See diagram below.) Figure 5 As shown, in an unsplit Crossbar switching network, the scheduling cycle of a data frame of this FPGA consists of two parts. The first part is that the ordinary shared cross-node module generates a sending request and the column arbitration module gives the scheduling result. The other part is the data frame transmission delay, which includes the retrieval interval and the data transmission delay.
[0066] Please see Figure 6 , Figure 6 This is a cross-slice group scheduling cycle analysis diagram provided in an embodiment of the present invention. (See diagram below.) Figure 6 As shown, in a crossbar switching network with horizontal splitting without introducing shared cross nodes, the scheduling cycle of a packet can be divided into four parts: scheduling interval, scheduling result being sent across slices to ordinary cross node modules, latency of packet being moved from cross node to inter-slice interface, and data cross-slice latency.
[0067] Please see Figure 7 , Figure 7 This is a diagram illustrating the scheduling cycle analysis of shared cross-nodes used across different chips, provided as an embodiment of the present invention. Figure 7 As shown, after introducing a shared cross-connect node, when the shared cross-connect node can always accept data frames, the scheduling cycle of a packet is the same as that of the unsplit Crossbar switching network, consisting of two parts: the retrieval interval and the data transmission delay. After introducing a shared cross-connect node, the data frame is first transmitted across fragments to the shared cross-connect node, and then participates in the WRR independent column arbitration, so its scheduling cycle is the same as that without splitting.
[0068] Please see Figure 8 , Figure 8 This is a simulation diagram of the cross-slice transmission path of the design and system of the crossbar switching network based on a shared cross node provided in an embodiment of the present invention. The following conclusions can be drawn from the simulation diagram: the cross-slice data in the crossbar switching network using a shared cross node buffer only has a larger latency, but the scheduling efficiency is close to that of the non-crossbar switching network.
[0069] In this embodiment, without using tags and shared cross-node buffers, the control signals that need to be transmitted across chips include key information such as the unicast / multicast identifier, frame length, and destination port number of the data frame. The transmission of this information across chips may be asynchronous with the transmission of the data frame. Therefore, each data frame scheduling requires waiting for all the key information to be used to be transmitted to the other FPGA, which will increase the scheduling latency of the exchange, affect the data frame scheduling efficiency, and affect the scheduling fairness between ports. However, including key information such as the unicast / multicast identifier, frame length, and destination port number in the tag header can effectively solve the problem of reduced scheduling efficiency caused by the asynchronous arrival of key information and data frames on another FPGA.
[0070] This embodiment proposes a design and system for a horizontally split Crossbar switching network based on a shared cross-node. By horizontally splitting the Crossbar switching network, the WRR independent column arbitration module needs to exchange two signals to schedule data frames from the cross-node buffer to the output port: a request signal from the ordinary cross-node buffer to the WRR independent column arbitration module, and a request-response signal from the WRR independent column arbitration module to the ordinary cross-node buffer. If horizontal splitting is directly adopted without using a shared cross-node buffer, two additional GPIO signals are needed to complete the interaction of the request and request-response signals. The latency of these two signals during inter-chip transmission will ultimately affect the link transmission efficiency. If a shared cross-connect node is used, as long as the shared cross-connect node can store the next longest frame (i.e., the almost signal is not high), it can directly receive data from another chip and then perform the arbitration request operation. In this way, the arbitration request and arbitration response are completed within the chip. The link transmission efficiency is the same as the efficiency of intra-chip data scheduling. The only difference is that the data frame comes from another chip, so the switching latency will be greater than when implementing Crossbar with a single FPAG. Using a shared cross-connect node buffer avoids the problem of reduced scheduling efficiency caused by the latency of inter-chip arbitration. Theoretically, as long as the shared cross-connect node buffer is large enough, the horizontally split WRR independent column arbitration scheduling has the same efficiency as the scheduling without splitting, improving the scheduling efficiency of the entire switching system and thus improving the link transmission efficiency.
[0071] In summary, this embodiment provides a feasible technical solution for the requirements of large switching capacity and a large number of ports. The use of shared cross-connect nodes effectively avoids the problem of latency in cross-chip data and control signal transmission. It not only expands the switching capacity and the number of ports, but also alleviates the arbitration fairness problem and scheduling efficiency reduction problem caused by cross-chip data and control signal transmission, thereby improving the scheduling efficiency of the entire switching system, thereby improving the link transmission efficiency and achieving a larger switching capacity.
[0072] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A horizontally splitting crossbar switching network system based on shared cross nodes, characterized in that, include: Each FPGA is equipped with several input queue management modules (10), general cross-node buffer modules (20), shared cross-node buffer modules (30), RR column arbitration modules (40), inter-chip high-speed Aurora interface modules (50), WRR independent column arbitration modules (60), and several configuration interfaces (70). The input queue management modules (10) are used to receive data frames from physical ports and add a tag header to the data frame header according to the destination port of the data frame; The ordinary cross node cache module (20) is used to receive data frames with added tag headers and temporarily store the data frames with added tag headers into different ordinary cross node caches (201) according to the destination port; The shared cross-node cache module (30) is used to receive and store data frames from another FPGA in the inter-chip high-speed Aurora interface module (50), and participate in the arbitration of the WRR independent column arbitration module (60) with the data frames in the ordinary cross-node cache (201) in the same column on the local PFGA; The RR column arbitration module (40) is used to give an arbitration result based on the transmission request of the ordinary cross node buffer (201) on the same column, and according to the arbitration result, the data frame with added tag header in this FPGA is moved from the corresponding ordinary cross node buffer (201) to the transmission part of the inter-chip high-speed Aurora interface module (50) to form a data frame to be transmitted. The inter-chip high-speed Aurora interface module (50) is used to send the data frame to be sent to the receiving part of the inter-chip high-speed Aurora interface module (50) of another FPGA, and to receive the data frame output by the sending part of the inter-chip high-speed Aurora interface module (50) of the other FPGA, and transmit it to the shared cross node buffer module (30) of the local FPGA. The WRR independent column arbitration module (60) is used to receive request signals sent by the cross-node caches of the same column in the ordinary cross-node cache module (20) and the shared cross-node cache module (30) and perform weighted round-robin scheduling, and move the data frames in the cross-node caches that are scheduled to be round-robin to the output port; The configuration interfaces (70) are used to configure the weights required for scheduling by the WRR independent column arbitration module (60) and the maximum and minimum thresholds of the queues of the input queue management modules (10).
2. The lateral splitting Crossbar switching network system based on shared cross nodes according to claim 1, characterized in that, The ordinary cross-node cache module (20) includes several ordinary cross-node caches (201), the shared cross-node cache module (30) includes several shared cross-node caches (301), the RR column arbitration module (40) includes several RR column arbitration sub-modules (401), the inter-chip high-speed Aurora interface module (50) includes several first inter-chip high-speed Aurora interfaces (501) and several second inter-chip high-speed Aurora interfaces (502), and the WRR independent column arbitration module (60) includes several WRR independent column arbitration sub-modules (601), wherein, The plurality of ordinary cross node caches (201) are arranged in an array, and each row of ordinary cross node caches (201) is connected to the input queue management module (10); Each of the shared cross node caches (301) is connected to the ordinary cross node caches (201) in the same column, each of the RR column arbitration submodules (401) is connected to the ordinary cross node caches (201) in the same column, and the sum of the number of the shared cross node caches (301) and the RR column arbitration submodules (401) is equal to the number of columns of the ordinary cross node caches (201); The first inter-chip high-speed Aurora interface (501) connects the RR column arbitration submodule (401) of the FPGA and the receiving part of the inter-chip high-speed Aurora interface module (50) of another FPGA. The second inter-chip high-speed Aurora interface (502) connects the shared cross node buffer (301) of the FPGA and the transmitting part of the inter-chip high-speed Aurora interface module (50) of another FPGA. Each of the WRR independent column arbitration submodules (601) is connected to the common cross node cache (201) in the same column and is located in the same column as the shared cross node cache (301).
3. The lateral splitting Crossbar switching network system based on shared cross nodes according to claim 2, characterized in that, The first inter-chip high-speed Aurora interface (501) includes a first Aurora IP core, a first cross-clock module, and a locallink to AXI module; the second inter-chip high-speed Aurora interface (502) includes a second Aurora IP core, a second cross-clock module, and an AXI to locallink module. The locallink to AXI module is used to convert the format of the data frame to be sent into the AXI data format to obtain a format-converted data frame; the first cross-clock module is used to cross the format-converted data frame from the system master clock domain to the user-side clock domain of the Aurora IP core to obtain a first cross-clock data frame; the first Aurora IP core is used to send the first cross-clock data frame to the receiving part of the high-speed Aurora interface module (50) between another FPGA chip. The second Aurora IP core is used to receive data frames output from the transmitting part of another FPGA inter-chip high-speed Aurora interface module (50) to obtain a received data frame; the second cross-clock module is used to cross the received data frame from the user-side clock domain of the Aurora IP core to the system master clock domain to obtain a second cross-clock data frame; the AXI to locallink module is used to convert the format of the second cross-clock data frame into the locallink data format used internally by the system.
4. The lateral splitting Crossbar switching network system based on shared cross nodes according to claim 3, characterized in that, Both the first and second Aurora IP cores are configured with 4-way bonding, and the line rate can reach 40Gbps.
5. The horizontally splitting Crossbar switching network system based on shared cross nodes according to claim 2, characterized in that, On each FPGA, the number of input queue management modules (10) is n / 2, where n is the number of buses; the number of ordinary cross-node buffers (201) is n. 2 / 2; the number of shared cross node caches (301) is n / 2; the number of RR column arbitration submodules (401) is n / 2; the number of the first inter-chip high-speed Aurora interfaces (501) is n / 2; the number of the second inter-chip high-speed Aurora interfaces (502) is n / 2; and the number of WRR independent column arbitration modules (60) is n / 2.
6. The lateral splitting Crossbar switching network system based on shared cross nodes according to claim 2, characterized in that, Both the ordinary cross-node cache (201) and the shared cross-node cache (301) include a first storage area and a second storage area, wherein, The first storage area is used for multicast or broadcast data, and the second storage area is used for storing unicast data.
7. The lateral splitting Crossbar switching network system based on shared cross nodes according to claim 1, characterized in that, The tag header includes a unicast / multicast identifier, a queue number, a frame length, and a destination port number.
8. The horizontally splitting Crossbar switching network system based on a shared cross node according to claim 1, characterized in that, Based on the destination port, the data frame transmission path with the added tag header is divided into a local transmission path and a cross-chip transmission path, wherein, The transmission path of this chip is as follows: the input queue management module (10) of this FPGA, the ordinary cross node buffer module (20) of this FPGA, the WRR independent column arbitration module (60) of this FPGA, and the output port corresponding to the destination port; The cross-chip transmission path is as follows: the input queue management module (10) of this FPGA, the ordinary cross node buffer module (20) of this FPGA, the RR column arbitration module (40) of this FPGA, the high-speed Aurora interface module (50) of this FPGA, the high-speed Aurora interface module (50) of another FPGA, the shared cross node buffer module (30) of another FPGA, the WRR independent column arbitration module (60) of another FPGA, and the output port corresponding to the destination port.
9. The lateral splitting Crossbar switching network system based on shared cross nodes according to claim 8, characterized in that, The scheduling period of a data frame in the local transmission path includes the retrieval interval and the data transmission delay, and the scheduling period of a data frame in the cross-slice transmission path includes the retrieval interval and the data transmission delay.
Citation Information
Patent Citations
Multi-die higher-order photonic switching structure based on high-density memory
CN108111930A
Method for designing Crossbar switching units interconnected among FPGA chips
CN110290074A