System and method for generating internal traffic in a switch

By configuring replication lists and generating seed packets in multi-node switches, and using hardware replication units for internal transmission, the real-time and efficiency problems of queue monitoring in large-scale switches are solved, real-time monitoring and congestion detection for each queue are realized.

CN115695085BActive Publication Date: 2025-05-02HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111259939.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-28
Filing Date
2021-10-28
Publication Date
2025-05-02
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

In multi-node switches, monitoring the congestion state and connectivity of a large number of queues is a difficult task, especially in large-scale switches. Line cards need to monitor a large number of queues in real time, resulting in excessive CPU resource consumption and software solutions cannot meet real-time requirements.

Method used

Granular monitoring of each queue is achieved by configuring a replication list in the switch, generating and copying seed packets, and using existing hardware replication units to transmit and monitor these packets internally in the switch.

Benefits of technology

It effectively reduces the consumption of CPU resources, realizes real-time monitoring of each queue in the switch, can promptly detect congestion and connectivity problems, and ensures good operating conditions of the switch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115695085B_ABST
    Figure CN115695085B_ABST
Patent Text Reader

Abstract

The present application relates to systems and methods for generating internal traffic in a switch. One aspect of the present application provides a system and method for generating internal traffic for a switch. During operation, the system configures a replication list including a plurality of replication entries, wherein the respective replication entries correspond to destination ports on the switch. The system generates seed packets to be replicated for each replication entry in the replication list, wherein the destination address of the respective replicated packets corresponds to the replication entry. All replicated packets are associated with a virtual local area network (VLAN) reserved for internal traffic. The system then forwards the replicated packets together with external packets received by the switch to the corresponding destination ports on the switch.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to determining congestion status and connectivity in a multi-node switch system. More specifically, the present disclosure relates to a system and method for generating internal traffic in a switch to determine the status of queues and connectivity between nodes in the switch. Background Art

[0002] In a multi-node switch that implements a virtual output queue (VOQ), the physical buffer of each input port maintains a separate virtual queue for each output port, so that congestion on the output port only blocks the virtual queue for that specific output port. The queuing algorithm or scheduling of the packet requires that the queue state of the destination be propagated to the source node. For large-scale switches, the number of queues that the line card needs to monitor can be huge. For example, for a switch cabinet with ten line cards, each of which handles up to 48 ports, and each of which has up to eight queues, the line card may need to monitor 3840 queues at any given time. Each of these queues may suffer from structural connectivity problems, delay problems, or congestion. In order to ensure good operating conditions for the cabinet, a monitoring mechanism that can take into account the granularity of the queue is required. Summary of the invention

[0003] The present disclosure provides a computer-implemented method for generating internal traffic for a switch.

[0004] According to one aspect, a computer-implemented method for generating internal traffic for a switch is provided, comprising: configuring a replication list including a plurality of replication entries, wherein a respective replication entry corresponds to a destination port on the switch; generating a seed packet to be replicated for each replication entry in the replication list, wherein a destination address of the respective replicated packet corresponds to the replication entry, and wherein all replicated packets are associated with a virtual local area network (VLAN) reserved for internal traffic; and forwarding the replicated packets together with an external packet received by the switch to the corresponding destination port on the switch.

[0005] The present invention can achieve beneficial technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 An exemplary multi-node switch system according to one aspect of the present application is illustrated.

[0007] Figure 2 A block diagram of an exemplary node in a multi-node switch system according to one aspect of the present application is illustrated.

[0008] Figure 3An exemplary configuration of nodes in a multi-node switch system according to one aspect of the present application is illustrated.

[0009] Figure 4 A flow chart illustrating an exemplary process for configuring a switch system to facilitate generation of internal traffic according to one aspect of the present application is provided.

[0010] Figure 5 A flow chart illustrating an exemplary process for monitoring queue status according to one aspect of the present application is provided.

[0011] Figure 6 An exemplary computer system that facilitates congestion monitoring on a per-queue level according to one aspect of the present application is illustrated.

[0012] In the various drawings, like reference numerals refer to the same drawing elements. DETAILED DESCRIPTION

[0013] The following description is intended to enable any person skilled in the art to utilize the examples and is provided in the context of a specific application and its requirements. Various modifications to the disclosed examples will be apparent to those skilled in the art, and the general principles defined herein may be applied to other examples and applications without departing from the spirit and scope of the present disclosure. Therefore, the scope of the present disclosure is not limited to the examples shown but is to be given the widest scope consistent with the principles and features disclosed herein.

[0014] In a multi-node switch that implements a virtual output queue (VOQ), the physical buffer of each input port maintains a separate virtual queue for each output port, so that congestion on the output port only blocks the virtual queue for that specific output port. The queuing algorithm or scheduling of the packet requires that the queue state of the destination be propagated to the source node. For large-scale switches, the number of queues that the line card needs to monitor can be huge. For example, for a switch cabinet with ten line cards, each of which handles up to 48 ports, and each of which has up to eight queues, the line card may need to monitor 3840 queues at any given time. Each of these queues may suffer from structural connectivity problems, delay problems, or congestion. In order to ensure good operating conditions for the cabinet, a monitoring mechanism that can take into account the granularity of the queue is required.

[0015] One approach is to have the CPU on each line card generate packets for transmission within the switch and monitor the delivery of these internal packets. However, given the size of the queues in large switches, such an approach is inefficient. Using a previous switch with ten line cards as an example, each line card CPU needs to generate packets and send packets to 3840 destinations, which will consume a lot of CPU resources, leaving less CPU cycles for other tasks. In addition, software-based solutions are often too slow to meet the almost real-time requirements in hardware to effectively detect congestion problems. Hardware solutions are required to effectively monitor delays and congestion within the switch.

[0016] In one example, an existing hardware replication unit in a switch can be used to generate packets sent internally in the switch. In conventional switches, replication units are mainly used for the purpose of IP multicast and layer 2 (L2) replication. In both cases, the replication unit replicates packets received by the switch from an external device and the replicated packets are sent to an external destination. Here, the replication unit can be modified and configured to operate in a traffic generation mode. More specifically, the replication unit can maintain a replication list including multiple replication entries, each of which corresponds to a destination on the switch (i.e., a specific port on a specific node). For each replication entry, the replication unit can continuously replicate a single seed packet for each queue associated with the replication entry. For example, if a replication entry (e.g., a specific port at a specific node) has multiple queues (e.g., eight priority queues) active, the replication unit can receive multiple seed packets (one for each queue) and replicate each seed packet for the replication entry. If the replication list has 100 entries, each seed packet will be replicated 100 times; and if each replication entry has eight queues, a total of 800 packets will be generated. The destination of each packet can be controlled and defined. According to one aspect of the present application, the destination of the generated packet may be of the form {node, port, queue}, effectively targeting all possible destinations for a given source node.

[0017] Figure 1 An exemplary multi-node switch system according to one aspect of the present application is illustrated. The multi-node switch system 100 may include a number of nodes (e.g., nodes 102-108) interconnected via a switch node 110. In one example, the multi-node switch system 100 may include a switch cabinet, the nodes 102-108 may include line cards, each of which has a plurality of network ports; and the switch node 110 may include a switch card. The network port of each line card may be coupled to an external device (e.g., a computer, a wireless access point, other switches, etc.). For example, node 102 is coupled to external devices 112, 114, and 116.

[0018] Figure 1 It is also shown that each node maintains a replication list that can be used to replicate packets. For example, node 102 maintains a replication list 120, which includes multiple entries. More specifically, the replication unit on each node maintains a replication list. Each replication entry in the replication list corresponds to a significant packet destination for the node, and the packet destination is a specific port on a specific node. When the replication unit is configured to operate in a traffic generation mode, it can continuously replicate a seed packet (which can be generated by the CPU of the node) for each replication entry in the replication list. In one example, the replication unit can be configured to go through the replication list and replicate the seed packet for each entry in the list, one at a time at a predetermined rate, and continuously repeat the process until it is interrupted by, for example, a user command.

[0019] It should be noted that although a certain amount of bandwidth will be consumed by the internal traffic, the internal traffic generated by the replication unit when the replication unit is operating in traffic generation mode does not interfere with normal traffic. Internally generated packets (also referred to as internal packets) can be inserted into the packet processing pipeline and will be processed in the same manner as normal packets received from external devices, except that these internally replicated packets will not leave the switch system. In other words, like normal traffic, these internal packets will be forwarded to their corresponding destinations on the switch (e.g., a specific queue of a specific port on a specific node), as determined by Figure 1 As shown by the double-headed arrow in . After the status of the delivery of these internal packets (e.g., multiple packets actually arrive at their corresponding destinations) is determined (e.g., by using a number of counters), these internal packets will be discarded at the destination port. Note that the generation rate of internal packets can be configurable. In one example, the generation rate of internal packets can be adjusted based on the current load of the switch system to avoid excessive consumption of bandwidth by internal traffic. For example, when the switch system experiences lighter traffic, the generation rate of internal packets can be faster (e.g., one packet is generated every millisecond); on the other hand, when the switch system experiences heavier traffic, the generation rate of internal packets can be slower (e.g., one packet is generated every ten milliseconds). In one example, the bandwidth consumed by internal traffic can be a predetermined fraction of the available bandwidth (e.g., between 0.1% and 1%).

[0020] The congestion state of the switch system can be determined based on the delivery state of the internal packets (e.g., which packets reach the destination and which packets are dropped because the destination queue is full). Because the internal packets are targeted at each active queue, the congestion state of the switch system can be determined with granularity on a per-queue level. In addition, the delivery state of the internal packets can also be used to determine the connectivity between all nodes in the switch system. For example, packets that are always dropped can indicate a loss of connectivity.

[0021] Figure 2 A block diagram of an exemplary node in a multi-node switch system according to one aspect of the present application is illustrated. Figure 2 , node 200 may be part of a multi-node switch system. For example, node 200 may be a line card inserted into a switch cabinet including multiple line cards. Node 200 may include multiple switch ports (e.g., ports 202-208), a CPU 210, a packet replication unit 212, a replication list 214, a tunneling logic block 216, an internal recirculation port 218, a packet processing logic 220, and multiple per-queue counters 222.

[0022] Node 200 can receive packets from and send packets to external devices (eg, computers, wireless access points, other switches, etc.) via switch ports. Switch ports on node 200 can be coupled to each other and to switch ports of different nodes.

[0023] CPU 210 may sometimes be referred to as a line card CPU and is often responsible for processing control plane traffic. According to one aspect of the present application, CPU 210 may be responsible for generating seed packets to be copied and forwarded to various switch ports in a multi-node switch system. In one example, the multi-node switch system may implement a VOQ architecture, and CPU 210 may generate different seed packets for different types of output queues. For example, the queues of a switch port may be organized based on priority, and CPU 210 may generate different seed packets for each priority queue. In another example, each port may support up to eight priority queues, and CPU 210 may generate eight different seed packets accordingly, one for each priority queue.

[0024] According to one aspect of the present application, in order to reduce the bandwidth consumed by internal traffic, the size of each seed packet can be kept very small. For example, the minimum size of the packet payload can be one byte. When generating a seed packet, CPU 210 can add an internal header to the packet, wherein the internal header defines many attributes associated with the packet. In one example, the internal header can define a replication group, a priority, and a specific virtual local area network (VLAN). The replication group refers to a packet replication list 214, which can include multiple replication entries. Priority refers to the type of queue associated with the seed packet. A specific VLAN is a VLAN reserved for internal traffic. All internal packets are associated with this specific VLAN.

[0025] The packet replication unit 212 may be responsible for replicating the seed packets generated by the CPU 210. The packet replication unit 212 may maintain a replication list 214, which may be implemented using hardware logic. The replication list 214 may be similar to Figure 1 , and may include many replication entries, each of which specifies a {node, port} combination. The number of replication entries in the replication list 214 may be the product of the number of nodes and the number of ports on each node. For example, if there are 10 line cards in the switch system and each line card has 48 ports, then the replication list 214 will include up to 10×48=480 entries. Because not all ports in the switch system are configured or activated, not all replication entries in the replication list 214 are enabled. In one example, only entries corresponding to configured ports are enabled. In another example, a system administrator may manually enable or disable entries in the replication list 214. As previously discussed, the packet replication unit 212 may replicate seed packets for each enabled replication entry in the replication list 214.

[0026] After receiving the seed packet from the CPU 210, the packet replication unit 212 can remove the internal header of the seed packet but the priority and VLAN information indicated by the internal header will be retained and used to replicate the packet. For example, the priority tag can be used to associate the replicated packet with a specific type of priority queue. VLAN is a specific VLAN reserved for internal traffic and can be represented as VLAN_R in the present disclosure. The priority tag combined with the destination MAC address can map the replicated packet to a specific queue of a specific port on a specific node. In this way, internal packets can be generated for all possible active destinations on the switch. Note that the destination port is not a member of a specific VLAN, which means that the internal packet will be discarded at the destination port without leaving the switch system.

[0027] The tunneling logic block 216 may be responsible for creating a number of tunnel entries, one for each duplicate entry. The number of tunnel entries may be the same as the number of entries in the replication list 214. All tunnels may use the same specific VLAN (i.e., VLAN_R) for their encapsulation, where the destination media access control (MAC) address for each tunnel entry corresponds to a {node, port} combination. More specifically, the destination MAC address of a tunnel instance corresponds to a specific duplicate entry in the replication list 214. In one example, the destination MAC address may have the following style: X:X:X:X:NODE:PORT, where NODE and PORT correspond to a node identifier and a port identifier, respectively. In another example, the destination MAC address may be: 08:00:09:00:NODE:PORT. Other formats are also possible, as long as the {node, port} combination can be uniquely identified. The scope of the present disclosure is not limited by the format of the destination MAC address of the replicated packet. According to one aspect, all tunnels point to the same destination, which may be an internal recycle port 218.

[0028] Internal recycle port 218 is an internal port on node 200. In other words, this port is not visible to external devices. Internal recycle port 218 can be responsible for inserting internal packets into the packet processing pipeline (i.e., packet processing logic 220). In this way, internal packets can be processed in a manner similar to external packets (which are packets received by node 200).

[0029] Each queue counter 222 can be responsible for counting the number of internal packets received for each queue on the node 200. For example, if there are 48 ports on the node 200 and each port supports eight queues, there will be 384 counters, one for each queue. It should be noted that because all internal packets are associated with a specific VLAN (VLAN_R), each queue counter 222 can be configured to count only packets associated with a specific VLAN. According to one aspect, each queue counter 222 can be monitored by a control application running in the CPU 210, which uses this information to check switch card connectivity and possible congestion from one node to another. More specifically, each queue counter 222 can provide congestion information for each queue. Such information can be used to determine whether a queue (e.g., VOQ) is dropping packets because it is full or saturated. Congestion information can also be used to determine whether the utilization of the queue is at a desired level. In one aspect, if utilization of one or more queues is not at a desired level, control and management logic in the switch system may take action, such as adjusting the transmission rate of certain ports, rebalancing traffic, or reconfiguring replication lists to target specific ports.

[0030] Figure 3 An exemplary configuration of nodes in a multi-node switch system according to one aspect of the present application is illustrated. In this example, it is assumed that there are a maximum of 10 nodes and each node may have up to 52 ports. Figure 3 The configuration of node 0 is shown.

[0031] exist Figure 3 , a replication list 302 maintained by a replication unit on a node (e.g., a line card) may include a plurality of entries (e.g., 52×9=468 entries). The replication entries, when enabled, facilitate the replication unit to replicate seed packets. In one example, the replication entries may be enabled if the corresponding destination port is configured. For example, when node 1, port 0 is configured, replication entry 1 will be enabled. In addition, if at any point a particular flow is desired to be stopped (e.g., a particular port no longer needs to be monitored), the corresponding replication entry in the replication list 302 may be disabled, thereby stopping packet replication for that particular port.

[0032] Each duplicate entry points to a tunnel entry in tunnel table 304, which may be maintained by tunneling logic. The number of tunnel entries in tunnel table 304 is the same as the number of duplicate entries in duplicate list 302. Each tunnel entry may set a MAC address for each {node, port} combination. For example, duplicate entry 1 points to tunnel 1 which sets a MAC address for node 1, port 0. All tunnels use a specific VLAN (VLAN_R) for encapsulation and they all point to recycle port 306. Figure 3 As shown in , the difference between different tunnel entries is the destination MAC address used to identify each destination (e.g., each unique {node, port} combination). Figure 3 In the example shown in , the destination MAC address may have the following format: X:X:X:X:NODE:PORT.

[0033] The recirculation port 306 tags all recirculated packets with the VLAN ID of a specific VLAN so that the recirculated packets can retain the specific VLAN. According to one aspect, when the recirculation port 306 is configured, the control software running in the line card CPU can create a set of L2 entries for the L2 table 308. The number of L2 entries created can be the same as the number of duplicate entries in the replication list 302. Figure 3, the replication list 302 has 468 entries, and the L2 table 308 also has 468 entries. Each L2 entry matches a specific VLAN (i.e., VLAN_R) and a destination MAC address (e.g., in the style of 08:00:09:00:NODE:PORT) to a corresponding packet destination specified by a {node, port} combination. In addition, because each packet is also tagged with a priority based on which seed packet is used for replication, the packet will actually target a significant queue (i.e., each unique {node, port, queue} combination). The packet is sent to the corresponding fabric port 310 based on the {node, port} combination (such as {node 1, port 0} or {node 9, port 51}).

[0034] Figure 4 A flowchart illustrating an exemplary process for configuring a switch system to facilitate the generation of internal traffic according to one aspect of the present application is provided. The switch system may include a plurality of line cards, each of which has a plurality of ports. During operation, the system reserves a specific VLAN for internal traffic (operation 402). All internal packets will be sent on the specific VLAN (which is referred to as VLAN_R in this disclosure). The system then configures a replication list for each line card (operation 404). The entries in the replication list correspond to destination ports in the switch system, wherein each entry corresponds to a unique {node, port} or {line card, port} combination. In one example, the number of entries in the replication list is equal to the total number of ports with which the line card is communicating. For example, if a particular line card is communicating with nine line cards in a switch, each of which has 52 ports, then the number of entries in the replication list is 52×9=468. Configuring the replication list may include enabling certain entries and / or disabling certain entries in the list. For example, an entry may be enabled when the corresponding port is configured and ready for sending and receiving packets. Similarly, when it is no longer necessary to monitor a particular port, its corresponding entry in the replication list may be disabled. Once the replication list is configured, the replication unit on the line card may replicate received seed packets for each enabled entry in the list.

[0035] In addition to configuring the replication list, the system can also configure a tunnel table including a number of tunnel entries (operation 406). The number of entries in the tunnel table can be equal to the number of entries in the replication list. More specifically, each entry in the replication list can be responsible for sending a packet (i.e., a packet generated for the entry) to a unique tunnel. All tunnels use a specific VLAN (i.e., VLAN_R) for their encapsulation, and all tunnels point to an internal recirculation port on a line card.

[0036] The system may also configure a recirculation port (operation 408). In one embodiment, the recirculation port may be configured to tag the internal packets with a specific VLAN so that the recirculated packets retain the VLAN information. It should be noted that the destination port of the recirculated packets is not a member of the specific VLAN so that the recirculated packets will be discarded by their destination port without leaving the switch. A properly configured recirculation port may insert the internal packets into the packet processing pipeline to allow the internal packets to be processed in a manner similar to packets received from external devices.

[0037] The system also configures the L2 forwarding table (operation 410). The number of entries in the L2 forwarding table can also be equal to the number of entries in the replication list. Each entry in the L2 forwarding table can map a packet belonging to a specific VLAN (which indicates that the packet is an internal packet) and having a destination MAC that complies with the pattern X:X:X:X:NODE:PORT to a destination specified by a unique {node, port} combination. In other words, a packet with a VLAN ID that matches VLAN_R and a destination MAC address that matches X:X:X:X:NODE:PORT will be sent to a designated port. Depending on the priority tag, a packet can also be sent to a specific priority queue of the port. Each node in the switch system needs to be correctly configured in order to facilitate the successful generation and forwarding of internal traffic.

[0038] Figure 5 A flow chart illustrating an exemplary process for monitoring queue states according to one aspect of the present application is provided. During operation, the system determines that a trigger condition has been met (operation 502). The trigger condition may be a command entered by a system administrator for congestion or connectivity monitoring. In one example, internal traffic generation may be part of an initialization process of a switch (e.g., as part of a built-in self-test (BIST)), and the trigger condition may be the completion of some other initialization operation. In another example, the trigger condition may be a set of predetermined rules. An exemplary rule may be that the trigger condition is met when the overall congestion level of the network reaches a predetermined level. For example, if the network is experiencing packet loss, the per-queue congestion monitoring mechanism may be triggered.

[0039] Once the trigger condition is met, the line card CPU can generate a number of seed packets and send the seed packets to the replication unit on the line card (operation 504). More specifically, the number of seed packets depends on the number of queues supported by each port. For example, if each port supports eight priority queues, the line card CPU can generate eight seed packets, one for each queue. When generating seed packets, the CPU can add an internal header to each packet, where the internal header includes information such as VLAN, priority, and replication group. It should be noted that VLAN is a specific VLAN reserved for internal traffic, and the replication group identifier is used to replicate the replication list of the seed packet. The payload of each seed packet can be kept to a minimum (e.g., one byte).

[0040] After receiving the seed packet, the replication unit can go through the entire replication list and replicate each seed packet for each entry in the replication list (operation 506). When replicating the seed packet, the replication unit can remove and process the internal header to obtain information (e.g., VLAN and priority) included in the internal header. The packet replicated according to a specific replication entry can have its destination MAC address (which can be part of the packet header) set based on the specific replication entry. In one example, each replication entry defines a unique {node, port} combination, and the destination MAC address for the corresponding replicated packet can adopt a style similar to X:X:X:X:NODE:PORT, where the initial four octets are undefined and the last two octets indicate the node and port. In another example, the destination MAC address can be 08:00:09:00:NODE:PORT. In addition to setting the destination MAC address, the replication unit can also add a priority tag to the replicated packet based on the internal header of the seed packet.

[0041] In one example, the replication unit can replicate each seed packet for each entry by replicating the list, one entry at a time. The rate of replication (i.e., the rate at which internal packets are generated) can be configured by a control algorithm running in the line card CPU or by a system administrator. In order to prevent internal traffic from consuming too much bandwidth, according to one aspect, the replication rate can be adjusted based on the traffic load on the switch. In one example, the bandwidth occupied by internal traffic can be a predetermined fraction of the available bandwidth (e.g., between 0.1% and 1%). In another example, when the bandwidth consumed by internal traffic is greater than a first threshold (e.g., 1%) or less than a second threshold (e.g., 0.1%), the replication rate will be reduced or increased, respectively.

[0042] The duplication unit sends the duplicated packets to the recirculation port on the line card (operation 508). In one example, the duplication unit may send the duplicated packets to the recirculation port via multiple tunnels, each tunnel corresponding to each destination MAC address. All tunnels use a specific VLAN for encapsulation. The recirculation port then inserts the duplicated packets into the packet processing pipeline (operation 510). While inserting the packets, the recirculation port may also mark the packets with the ID of a specific VLAN to retain the VLAN information. Inserting or recirculating the packets into the packet processing pipeline can ensure that these internal packets are forwarded in the same manner as normal external packets. This also ensures that the generation and forwarding of internal traffic does not interfere with normal traffic. Unlike some solutions that require a packet processing application-specific integrated circuit (ASIC) to stop normal forwarding operations to operate in a traffic generation mode, this solution allows the packet processing ASIC to operate normally by treating internal and external packets in the same manner.

[0043] The packet processing ASIC then forwards the packets to their corresponding destinations (operation 512). To do so, the packet processing ASIC processes information included in the packet header, such as VLAN ID, priority, and destination MAC address. The packet can be forwarded to a specific priority queue of a specific port on a specific line card based on the information included in the packet header. When the packet arrives at its destination (e.g., a specific port on a specific line card), the corresponding per-queue counter increments its value and the packet is discarded (operation 514). The packet is discarded because the destination port is not a member of a specific VLAN. In this way, the internal packet will not leave the switch. In one example, each line card maintains multiple per-queue counters, one for each queue. The counter value is sent to the line card CPU (operation 516), thereby allowing the CPU to monitor and analyze the status of each queue. For example, based on the packet replication rate and the value of the per-queue counter, the CPU can determine the packet loss rate for each queue. The CPU can also determine whether the packet loss is due to the queue being full or saturated. In addition, the packet loss rate of a specific queue or port can also indicate connectivity issues between ports to the CPU.

[0044] Internal traffic may be continuously generated and forwarded (e.g., since initialization of the switch) to allow for continuous monitoring of congestion and / or connectivity. However, according to one aspect, the system may also determine whether an interruption condition is met (operation 518). If so, the process ends and there is no further replication of the seed packet. If not, the replication unit continues to replicate the seed packet (operation 506). The interruption condition may be a command from a system administrator to stop the internal traffic or the expiration of a predetermined timer.

[0045] Figure 6An exemplary computer system that facilitates congestion monitoring at a per-queue level according to one aspect of the present application is illustrated. Computer system 600 includes a processor 602, a memory 604, and a storage device 606. In addition, computer system 600 can be coupled to an external network input / output (I / O) user device 610, such as a display device 612, a keyboard 614, and a pointing device 616. Storage device 606 can store an operating system 618, a congestion monitoring system 620, and data 650.

[0046] The congestion monitoring system 620 may include instructions that, when executed by the computer system 600, may cause the computer system 600 or the processor 602 to perform the methods and / or processes described in the present disclosure. Specifically, the congestion monitoring system 620 may include instructions for determining whether a trigger condition is met (trigger condition determination instructions 622), instructions for generating a seed packet (seed packet generation instructions 624), instructions for configuring a replication list (replication list configuration instructions 626), instructions for configuring a packet replication rate (replication rate configuration instructions 628), instructions for reserving a VLAN (VLAN reservation instructions 630), instructions for configuring a tunnel table (tunnel table configuration instructions 632), instructions for configuring a recirculation port (recirculation port configuration instructions 634), instructions for configuring an L2 forwarding table (forwarding table configuration instructions 636), instructions for monitoring per-queue counters (counter monitoring instructions 638), and instructions for analyzing the status of individual queues (queue status analysis instructions 640).

[0047] In general, the present disclosure provides a solution to the problem of generating internal traffic in a switch, which can be used to monitor the congestion state of the switch at a per-queue level and determine connectivity between ports. Instead of using software to generate all internal packets, the disclosed solution uses software to generate seed packets that can be copied by hardware logic (e.g., a packet copying unit) in each line card. Such a packet copying unit has been conventionally used for multicast or L2 replication purposes. In order to achieve per-queue level granularity, multiple seed packets can be generated (one packet for each type of queue). The copying unit maintains a copy list, wherein the entries of the list correspond to the destination port on the switch. For each entry in the copy list, the copying unit can continue to copy the seed packet, which can be forwarded to the destination port afterwards. Different copies of different seed packets target different queues of the destination port. Each line card can also implement multiple per-queue counters to count the number of internal packets received at each individual queue. These per-queue counters can be monitored by control algorithms running in the line card CPU, which can collect and use this information to check the fabric card connectivity and possible congestion in the queue. In addition to monitoring congestion during normal operation of the switch, internal traffic can also be generated and monitored as an internal BIST during switch initialization to verify connectivity between nodes in the switch. The solution can also be used to measure the utilization of the VOQ of each destination port. Internal packets can be used to determine whether the VOQ is dropping traffic because it is full or saturated. Determining the utilization of the VOQ can allow the system administrator to take remedial measures (e.g., rate adjustment or traffic rebalancing) if the utilization of one or more VOQs is not as expected. Switch internal traffic can also be generated for purposes other than monitoring congestion and connectivity in a multi-node switch system.

[0048] One aspect of the present application provides a system and method for generating internal traffic for a switch. During operation, the system configures a replication list including a plurality of replication entries, wherein the respective replication entries correspond to destination ports on the switch. The system generates a seed packet to be replicated for each replication entry in the replication list, wherein the destination address of the respective replicated packet corresponds to the replication entry. All replicated packets are associated with a virtual local area network (VLAN) reserved for internal traffic. The system then forwards the replicated packets together with external packets received by the switch to the corresponding destination ports on the switch.

[0049] In a variation on this aspect, the corresponding destination port supports multiple queues, and the system generates multiple seed packets, one seed packet for each queue.

[0050] In another variation, the plurality of queues are priority queues; the seed packet includes an internal header indicating a type of priority queue targeted by packets copied based on the seed packet; and copying the seed packet includes removing the internal header and priority marking the copied packet.

[0051] In another variation, the system further receives from a per-queue counter a counter value counting the number of packets received at a particular port for a particular queue, and determines a state of the particular queue based on the counter value.

[0052] In a variation on this aspect, the switch includes a plurality of interconnected nodes, and wherein the respective replicated entries specify a unique {node, port} combination.

[0053] In a variation on this aspect, forwarding the copied packet includes configuring an internal recirculation port to insert the copied packet into a packet processing pipeline to allow the copied packet to be processed similarly to the external packet.

[0054] In another variation, the system configuration includes a tunnel table comprising a plurality of tunnel entries, wherein respective tunnel entries correspond to duplicate entries.All tunnel entries point to the recirculation port, thus facilitating tunneling of the duplicated packets to the internal recirculation port.

[0055] In a variation on this aspect, configuring the replication list includes adjusting a rate at which the seed packets are replicated based on a traffic load on the switch.

[0056] In a variation on this aspect, configuring the replication list includes disabling a replication entry in the replication list to stop replication of the seed packet for the disabled replication entry.

[0057] In a variation on this aspect, the system continuously replicates the seed packet until an interruption condition is met.

[0058] The methods and processes described in the specific implementation section can be embodied as code and / or data, which can be stored in a computer-readable storage medium as described above. When a computer system reads and runs the code and / or data stored on the computer-readable storage medium, the computer system executes the methods and processes embodied as data structures and codes and stored in the computer-readable storage medium.

[0059] In addition, the methods and processes described above may be included in hardware modules or devices. Hardware modules or devices may include, but are not limited to, ASIC chips, field programmable gate arrays (FPGAs), dedicated or shared processors that run specific software modules or code fragments at specific times, and other programmable logic devices now known or later developed. When the hardware modules or devices are activated, they execute the methods and processes included in them.

[0060] The foregoing description of the embodiments has been presented for the purpose of illustration and description only. It is not intended to be exhaustive or to limit the scope of the present disclosure to the disclosed forms. Therefore, many modifications and variations will be apparent to those skilled in the art.

Claims

1. A computer-implemented method for generating internal traffic for a switch, the method comprising: configuring a replication list including a plurality of replication entries, wherein respective replication entries correspond to destination ports on the switch; generating a seed packet to be replicated for each replication entry in the replication list, wherein a destination address of a respective replicated packet corresponds to the replication entry, and wherein all replicated packets are associated with a virtual local area network (VLAN) reserved for the internal traffic; forwarding the copied packet to a corresponding destination port on the switch along with an external packet received by the switch, wherein forwarding the copied packet includes configuring an internal recirculation port to insert the copied packet into a packet processing pipeline to allow the copied packet to be processed similarly to the external packet; as well as A tunnel table including a plurality of tunnel entries is configured, wherein respective tunnel entries correspond to duplicate entries, and wherein all tunnel entries point to the internal recirculation port, thereby facilitating tunneling of the duplicated packets to the internal recirculation port.

2. The method of claim 1, wherein the corresponding destination port supports a plurality of queues, and wherein the method further comprises generating a plurality of seed packets, one seed packet for each queue.

3. A method according to claim 2, wherein the multiple queues are priority queues, wherein the seed packet includes an internal header, the internal header indicates the type of priority queue targeted by the packet copied based on the seed packet, and wherein copying the seed packet includes removing the internal header and marking the priority of the copied packet.

4. The method according to claim 2, further comprising: receiving a counter value from a per-queue counter, the counter value counting a number of packets received at a particular port for a particular queue; as well as A state of the particular queue is determined based on the counter value.

5. The method of claim 1, wherein the switch comprises a plurality of interconnected nodes, and wherein the corresponding replication entry specifies a unique {node, port} combination.

6. The method of claim 1, wherein configuring the replication list comprises adjusting a rate at which the seed packets are replicated based on a traffic load on the switch. 7 . The method of claim 1 , wherein configuring the replication list comprises disabling replication entries in the replication list to stop replication of the seed packet for the disabled replication entries.

8. The method of claim 1, further comprising continuously replicating the seed packet until an interruption condition is met.

9. A computer system comprising: processor; as well as a memory coupled to the processor and storing instructions that, when executed by the processor, cause the processor to perform a method for generating internal traffic, the method comprising: configuring a replication list including a plurality of replication entries, wherein respective replication entries correspond to destination ports on the switch; generating a seed packet to be replicated for each replication entry in the replication list, wherein a destination address of a respective replicated packet corresponds to the replication entry, and wherein all replicated packets are associated with a virtual local area network (VLAN) reserved for the internal traffic; forwarding the copied packet along with an external packet received by the switch to a corresponding destination port on the switch, wherein forwarding the copied packet includes configuring an internal recirculation port to insert the copied packet into a packet processing pipeline to allow the copied packet to be processed similarly to the external packet; and A tunnel table including a plurality of tunnel entries is configured, wherein respective tunnel entries correspond to duplicate entries, and wherein all tunnel entries point to the internal recirculation port, thereby facilitating tunneling of the duplicated packets to the internal recirculation port.

10. The computer system of claim 9, wherein the corresponding destination port supports a plurality of queues, and wherein the method further comprises generating a plurality of seed packets, one seed packet for each queue.

11. A computer system according to claim 10, wherein the plurality of queues are priority queues, wherein the seed packet includes an internal header indicating a type of priority queue targeted by packets copied based on the seed packet, and wherein copying the seed packet includes removing the internal header and marking the priority of the copied packet.

12. The computer system of claim 10, wherein the method further comprises: receiving a counter value from a per-queue counter, the counter value counting a number of packets received at a particular port for a particular queue; as well as A state of the particular queue is determined based on the counter value.

13. The computer system of claim 9, wherein the switch comprises a plurality of interconnected nodes, and wherein the corresponding replication entry specifies a unique {node, port} combination.

14. The computer system of claim 9, wherein configuring the replication list comprises adjusting a rate at which the seed packets are replicated based on a traffic load on the switch.

15. The computer system of claim 9, wherein configuring the replication list comprises disabling replication entries in the replication list to stop replication of the seed packet for the disabled replication entries.

16. The computer system of claim 9, wherein the method further comprises continuously replicating the seed packet until an interrupt condition is met.

Citation Information

Patent Citations

  • System and method for using subnet prefix values in global route header (GRH) for linear forwarding table (LFT) lookup in a high performance computing environment

    CN108028813A

  • Apparatus and method to trace packets in a packet processing pipeline of a software defined networking switch

    CN112262553A