Adaptive port route notification

By monitoring the port bandwidth in the computing system and generating adaptive routing notification packets, the problem of high power consumption in management of multipath routing and low traffic periods is solved, and automatic deactivation of switches or ports and power savings are achieved.

CN120075113APending Publication Date: 2025-05-30MELLANOX TECHNOLOGIES LTD(IL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411727727.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Managing multipath routing and ensuring optimal path selection presents challenges in the prior art and power consumption may be unnecessary high during low traffic periods.

Method used

By monitoring the port bandwidth in a computing system (such as a switch), if the bandwidth is below the threshold, an adaptive routing notification packet is generated, stop receiving traffic and may deactivate the port to reduce power consumption.

Benefits of technology

Automatically deactivate underutilized switches or ports during low traffic periods, saving power and optimizing network performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075113A_ABST
    Figure CN120075113A_ABST
Patent Text Reader

Abstract

The invention discloses adaptive port route notification. A device, a communication system and a method are provided. In one example, a system for routing traffic is described that includes a plurality of ports to facilitate communications over a network. The system also includes a controller to selectively activate or deactivate ports of the system according to the queue depth and additional information to improve power efficiency of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to networking, and more particularly to networked devices, switches, and methods of operating the same. Background Art

[0002] Switches and similar network devices represent core components of many communication, security, and computing networks. Switches are typically used to connect multiple devices, device types, networks, and network types.

[0003] Devices (including but not limited to personal computers, servers, or other types of computing devices) can be interconnected using network devices such as switches. These interconnected entities form a network that enables data communication and resource sharing between nodes. Typically, there may be multiple potential paths for data flow between any pair of devices. This feature, commonly referred to as multipath routing, allows data (usually encapsulated in packets) to traverse different routes from a source device to a destination device. This network design enhances the robustness and flexibility of data communication as it provides alternatives in the event of path failures, congestion, or other adverse conditions. Additionally, it helps with load balancing across the network, thus optimizing overall network performance and efficiency. However, managing multipath routing and ensuring optimal path selection can pose significant challenges, requiring advanced mechanisms and algorithms for network control and data routing, and power consumption can be unnecessarily high, especially during low-traffic periods. Summary of the Invention

[0004] According to one or more embodiments described herein, a computing system (such as a switch) can enable various systems (such as other switches, servers, personal computers, and other computing devices) to communicate across a network. Ports of the computing system can serve as communication endpoints, allowing the computing system to manage multiple simultaneous network connections with one or more nodes.

[0005] Each port of the computing system can be regarded as a passageway and has an output queue for packets / data waiting to be sent through the port. In effect, each port can be used as an independent channel for data communication with the computing system. The ports allow concurrent network communication, enabling the computing system to perform multiple data exchanges with different network nodes simultaneously.

[0006] The traffic received and / or sent from a port can be measured in terms of bandwidth or data rate over time. For example, the bandwidth of a port can be measured by tracking the amount of data sent or received by the port and determining the rate. As an example, bandwidth can be measured in bits per second or other terms.

[0007] As described herein, bandwidth can be monitored for the entire system (considering all ingress and / or egress ports of the system) or individually for different ports of the system. If the bandwidth is less than a specific amount, the system can generate a notification and send the notification to the traffic source. This notification may cause the source to be less likely to continue sending traffic to the system. Thus, system ports that previously received traffic from the source may stop receiving traffic from the source. In response to determining that no traffic is being received (e.g., stopped), the port and / or the entire system can be deactivated.

[0008] This disclosure describes systems and methods for enabling a switch or other computing system to potentially deactivate one or more ports in response to receiving a relatively small amount of traffic over a given period of time. Using such a system or method, unutilized switches in a switch network can be turned off until needed, thereby saving power. In some embodiments, one or more ports of the switch can be individually deactivated, which also results in power savings. Embodiments of this disclosure are intended to address the above disadvantages and other problems by implementing improved routing methods. The routing methods shown and described herein can be applied to switches, routers, or any other suitable type of network device, known or yet to be developed.

[0009] In an illustrative example, a system for providing adaptive routing is disclosed, the system including circuitry for: receiving data from a network via a first port; determining bandwidth; determining that the bandwidth is below a threshold; generating an instruction packet in response to determining that the bandwidth is below the threshold; and sending the instruction packet via the first port.

[0010] In another example, a system for providing adaptive routing is disclosed, the system including one or more circuits for: receiving data from a network; determining the total bandwidth associated with the received data; determining that the total bandwidth associated with the received data is below a threshold; generating an instruction packet in response to determining that the total bandwidth is below the threshold; and sending the instruction packet via a plurality of ports.

[0011] In another example, a system for providing adaptive routing is disclosed, the system including one or more circuits for: receiving data via a first port of a plurality of ports; determining the bandwidth associated with the received data; determining that the bandwidth associated with the received data is below a threshold; generating an instruction packet in response to determining that the bandwidth is below the threshold; and sending the instruction packet via the first port.

[0012] Any of the above example aspects includes where the bandwidth includes the bandwidth of the first port, and where the one or more circuits are further for determining a second bandwidth associated with a second port.

[0013] Any of the above example aspects includes where the bandwidth includes the total bandwidth of a plurality of ports including the first port, and determining the bandwidth includes determining the total bandwidth of each port in the plurality of ports.

[0014] Any of the above example aspects includes where the bandwidth includes the egress bandwidth and the first port includes an egress port.

[0015] Any of the above example aspects includes where one or more circuits are further configured to enter a sleep mode after a link has been idle for a period of time.

[0016] Any of the above example aspects includes where entering the sleep mode includes disabling one or more serializer / deserializer (SerDes) circuits associated with the link.

[0017] Any of the above example aspects includes where after transmitting an instruction packet, traffic at the first port is stopped.

[0018] Any of the above example aspects includes where the total network bandwidth is low relative to the total network capacity.

[0019] Any of the above example aspects includes where the first port includes a port of a backbone switch and where data is received from a top-of-rack (TOR) switch.

[0020] Any of the above example aspects includes where the first port includes a port of an L2 switch and where data is received from an L3 switch and where the data is forwarded to an L1 switch.

[0021] Any of the above example aspects includes where the total bandwidth includes the total egress bandwidth.

[0022] Any of the above example aspects includes where transmitting the instruction packet includes: transmitting data packets to a plurality of destinations via a plurality of ports.

[0023] Any of the above example aspects includes where one or more circuits are further configured to enter a sleep mode after a link associated with a plurality of ports has been idle for a period of time.

[0024] Any of the above example aspects includes where entering the sleep mode includes disabling one or more serializer / deserializer (SerDes) circuits associated with the link.

[0025] Any of the above example aspects includes where after transmitting an instruction packet, traffic at a plurality of ports is stopped.

[0026] Any of the above example aspects includes where before determining the bandwidth associated with the first port, one or more circuits select the first port using a selection process.

[0027] Any of the above example aspects includes that the selection process includes round-robin scheduling.

[0028] Any of the above example aspects includes that one or more circuits are further configured to select a second port using the selection process after sending an instruction packet.

[0029] Any of the above example aspects includes that one or more circuits are further configured to update a routing table based on an updated queue depth.

[0030] Other features and advantages are described herein and will be apparent from the following description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The present disclosure is described in conjunction with the accompanying drawings, which are not necessarily drawn to scale:

[0032] Figure 1 is a block diagram showing an illustrative configuration of a computing system according to at least some embodiments of the present disclosure;

[0033] Figure 2 shows a network of a computing system and nodes according to at least some embodiments of the present disclosure;

[0034] Figure 3 shows a network of a computing system and nodes according to at least some embodiments of the present disclosure;

[0035] Figure 4 shows a network of a computing system and nodes according to at least some embodiments of the present disclosure;

[0036] Figure 5 is a flowchart showing a method according to at least some embodiments of the present disclosure;

[0037] Figure 6 is a flowchart showing a method according to at least some embodiments of the present disclosure; and

[0038] Figure 7 is a flowchart showing a method according to at least some embodiments of the present disclosure. DETAILED DESCRIPTION

[0039] The following description provides only embodiments and is not intended to limit the scope, applicability, or configuration of the claims. Instead, the following description will provide a viable description of implementing the embodiments for those skilled in the art. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope of the appended claims.

[0040] It can be understood from the following description that, for reasons of computational efficiency, the components of the system can be arranged at any appropriate location in a distributed component network without affecting the operation of the system.

[0041] In addition, it should be understood that the various links of the connecting elements can be wired, trace, or wireless links, or any suitable combination thereof, or any other suitable known or later-developed element capable of providing data to and / or transmitting data from the connected elements. For example, the transmission medium used as a link can be any suitable carrier of electrical signals, including coaxial cables, copper wires, and optical fibers, electrical traces on a printed circuit board (PCB), etc.

[0042] As used herein, the phrases "at least one", "one or more", "or", and "and / or" are open-ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions "at least one of A, B, and C", "at least one of A, B, or C", "one or more of A, B, and C", "one or more of A, B, or C", "A, B, and / or C", and "A, B, or C" represents A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.

[0043] The term "automatically" and its variants as used herein refer to any suitable process or operation that can be completed without substantial human input during its execution or operation. However, if an input is received before the execution or operation, the process or operation can be automatic even if the execution of the process or operation uses substantial or non-substantial human input. If the human input affects the manner in which the process or operation is executed, then the human input is considered substantial. A human input that consents to the execution of the process or operation should not be considered "substantial".

[0044] The terms "determine", "calculate", and "compute" and their variants as used herein are used interchangeably and include any suitable type of method, process, operation, or technique.

[0045] Aspects of the present disclosure will be described herein with reference to the accompanying drawings, which are schematic diagrams of idealized configurations.

[0046] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It should also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with the meaning in the context of the relevant art and this disclosure.

[0047] As used herein, unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" are also intended to include the plural forms. It should also be understood that when the terms "comprises," "comprising," and / or "including" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The term "and / or" includes any and all combinations of one or more of the associated listed items.

[0048] Now referring to Figures 1 to 7 , various systems and methods for providing adaptive routing of data packets between communication nodes will be described. The concepts of data packet routing shown and described herein can be applied to routing information from one computing device to another. The term "data packet" as used herein should be construed to mean any suitable discrete quantity of digitized information. The information being routed can be in the form of a single data packet or multiple data packets without departing from the scope of the present disclosure. Additionally, some embodiments will be described in connection with systems configured to make centralized routing decisions, while other embodiments will be described in connection with systems configured to make distributed and possibly uncoordinated routing decisions. It should be understood that the features and functionality of a centralized architecture can be applied to a distributed architecture or vice versa.

[0049] According to one or more embodiments described herein, as Figure 1 shown, the computing system 103 can enable various systems (such as switches, servers, personal computers, and other computing devices) to communicate across a network. Such a computing system 103 described herein can be a switch or any computing device, including a plurality of ports 106 for connecting to nodes on the network. In some embodiments, the computing system 103 described herein can be a switch that connects multiple top-of-rack (TOR) switches. The switch can, for example, be used as a backbone switch. In some embodiments, the computing system 103 can be part of a topology (such as a three-level fat tree topology). For example, a sending switch can be regarded as an L3 switch, a receiving switch can be regarded as an L2 switch, and a TOR switch can be regarded as an L1 switch. The computing system 103 can operate as any one of a backbone, TOR, L1, L2, or L3 switch. Nodes that can send data to the computing system 103 described herein can, in some embodiments, include TOR switches, backbone switches, or other computing devices capable of receiving and transmitting data.

[0050] The port 106 of the computing system 103 can be used as a communication endpoint, allowing the computing system 103 to manage multiple simultaneous network connections with one or more nodes. As described above, the computing system 103 can be a backbone switch or a TOR switch, or can be one of the L1, L2, or L3 switches in a three-tier fat-tree topology. The data received at the port 106 of the computing system 103 can be associated with the target device directly or indirectly connected to the computing system 103. After receiving such data, the computing system 103 can be configured to process the data, determine the destination, and make a routing decision. Making a routing decision can include determining from which port to forward the data. The destination can be directly connected to the computing system 103 via a single cable, or can be indirectly connected to the computing system via one or more intermediate switches. Each port 106 can be used to transmit data associated with one or more flows. Each of the ports 106a - 106c can be associated with a corresponding queue 121a - 121c, enabling the ports 106a - 106c to handle data packets of incoming and / or outgoing data associated with the flows.

[0051] Each port 106 of the computing system can be regarded as a path and is associated with an egress queue 121 for packets / data waiting to be sent through the port. In fact, each port 106 can be used as an independent channel for data communication with the computing system 103. The port 106 can allow concurrent network communication, enabling the computing system 103 to perform multiple data exchanges with different network nodes simultaneously.

[0052] Each of the ports 106a - 106c can be associated with an egress queue 121a - 121c, which can store data (such as data packets) waiting to be transmitted from the corresponding ports 106a - 106d. When a data packet or other form of data is ready to be sent from the computing system 103, the data packet can be assigned to the port 106 from which the data packet will be sent, and the data packet can be stored in the queue 121 associated with that port.

[0053] In some embodiments, each of the queues 121a - 121c can store data packets received by the corresponding ports 106a - 106c. For example, each of the ports 106a - 106c can be used as an ingress port and can receive data from one or more sources. After receiving a data packet, the receiving port 106 can store the data packet data in the corresponding queue 121.

[0054] The port 106 of the computing system 103 can be a physical connection point, allowing a network cable (such as an Ethernet cable) to connect the computing system 103 to one or more network nodes. Each port 106 can be of the same or different types, including, for example, 100 Mbps, 1000 Mbps, or 10 Gigabit Ethernet ports, each providing various levels of bandwidth.

[0055] As described herein, port 106 can be deactivated based on a variety of factors. Deactivating port 106 can include reducing or turning off the power of the switching hardware 109 associated with the corresponding port 106 (described in more detail below). In some embodiments, the switching hardware 109 can be turned off on a port-by-port basis, or the switching hardware 109 for all ports can be turned off together. As described below, port 106 may be effectively disabled due to lack of received traffic over a particular period of time. As a result of disabling the port, computing system 103 can reduce the amount of power consumption.

[0056] The switching hardware 109 of a computing system can include a fabric or path within the computing system 103 through which data is transmitted between two ports 106. In some embodiments, the switching hardware 109 can include one or more network interface cards (NICs). In some embodiments, each of ports 106a - 106d can be associated with a different NIC. The NIC or NICs can include hardware and / or circuitry that can be used to transmit data between ports 106.

[0057] The switching hardware 109 can also or alternatively include one or more application - specific integrated circuits (ASICs) to perform tasks such as determining to which port a received data packet should be sent. The switching hardware 109 can include various components, including for example, port controllers that manage the operation of individual ports, NICs that facilitate data transmission, and internal data paths that direct the flow of data within the computing system 103. The switching hardware 109 can also include memory elements for temporarily storing data and management software for controlling the operation of the hardware. Such a configuration can enable the switching hardware 109 to accurately track port usage and provide data to the processor 115 upon request.

[0058] Data packets received by the computing system 103 can be placed in the buffer 112 until they are placed in the queue 121 and then transmitted by the corresponding port 106. The buffer 112 can effectively be an ingress queue where data packets received can be temporarily stored. As described herein, the port 106 through which a given data packet is to be sent can be determined based on many factors.

[0059] In some embodiments, the buffer 112 can store packet data, while the queue 121 can store packet identification data. The queue 121 as described herein can be a data structure for managing data to be forwarded by the computing system 103. Each port 106 of the computing system 103 can have an associated queue. When the computing system 103 receives a data packet, the data packet can first be written to the buffer 112 until the data packet is assigned to the queue associated with the selected port for forwarding the data packet.

[0060] The switching hardware 109 may include serializer / deserializer (SerDes) circuitry 130. Each port 106 may be associated with dedicated SerDes circuitry or may share the SerDes circuitry 130. Disabling a port 106 may include disabling the SerDes circuitry 130 associated with the port 106. Disabling the SerDes circuitry 130 may include turning off or reducing the electrical power applied to the SerDes circuitry 130. Thus, disabling a port 106 may reduce the power consumption of the computing system 103 because less electrical power is consumed by the SerDes circuitry 130 when the port 106 is disabled.

[0061] As Figure 1 shown, the computing system 103 may also include processing circuitry, referred to herein as the processor 115, and may be a processor, CPU, microprocessor, or any circuitry or device capable of reading instructions from the memory 118 and performing operations. The processor 115 may, for example, execute software instructions to control the operation of the computing system 103. The processor 115 may also write data to the memory 118.

[0062] The processor 115 may serve as the central processing unit of the computing system 103 and may perform functions that implement the operational capabilities of the computing system 103. The processor 115 may be enabled to communicate with other components of the computing system 103 to manage and perform computing operations in accordance with the systems and methods described herein.

[0063] The processor 115 may be programmed to perform various computing tasks. The functions of the processor 115 may include executing program instructions, managing data within the computing system 103, and controlling the operation of other hardware components of the computing system 103 (such as the switching hardware 109). The processor 115 may be a single-core or multi-core processor and may include one or more processing units, depending on the specific design and requirements of the computing system 103.

[0064] The computing system 103 may also include one or more memory 118 components. The memory 118 may be configured to communicate with the processor 115 of the computing system 103. Communication between the memory 118 and the processor 115 may enable various operations, including but not limited to data exchange, command execution, and memory management. In accordance with the embodiments described herein, the memory 118 may be used to store data related to the use of the ports 106a - 106c of the computing system 103 (such as threshold data 124 and bandwidth data 127).

[0065] Memory 118 can be composed of a variety of physical components, depending on the specific type and design. For example, memory 118 can include one or more storage units capable of storing data in the form of binary information. These storage units can be composed of transistors, capacitors, or other suitable electronic components, depending on the memory type, such as DRAM, SRAM, or flash memory. To enable data transfer and communication with other parts of computing system 103, memory 118 can also include one or more data lines or buses, address lines, and control lines or be in contact with one or more data lines or buses, address lines, and control lines.

[0066] Threshold data 124 can be stored in memory 118 and can be a user-configurable variable that can be used to determine or detect when the bandwidth of one or more ports is low. The thresholds stored as threshold data 124 can be data in the form of a percentage, bit rate, number of bits, or other formats. The percentage threshold can be a percentage of the maximum bandwidth capacity of one or more ports, or it can be a percentage of a number indicating the bandwidth less than the maximum bandwidth capacity of one or more ports.

[0067] Bandwidth data 127 that can be stored in memory 118 can be a record of port usage. For example, bandwidth data 127 can store the moving average bandwidth of each port 106a - 106c and / or the aggregated total moving average bandwidth of each port 106a - 106c or a subset of ports. In some embodiments, historical data reflecting past bandwidth can be recorded as bandwidth data 127. For example, processor 115 can record a log of the bandwidth received through each port 106a - 106c with a timestamp so that trends related to the usage of ports 106a - 106c can be identified or determined, and the future usage of ports 106a - 106c can be predicted based on the identified trends.

[0068] In one or more embodiments of the present disclosure, processor 115 of computing system 103 (such as a switch) can perform a polling operation to retrieve data related to the activities of ports 106a - 106d, such as by polling switch hardware 109. As used herein, polling can involve processor 115 periodically querying or requesting data from switch hardware 109. The polling process can include processor 115 sending a request to switch hardware 109 to retrieve the required data. After receiving the request, switch hardware 109 can compile the requested port usage data and send it back to processor 115.

[0069] The bandwidth data 127 may include various metrics, such as the amount of data or the number of data packets in each queue 121a - 121c, the amount of data or the number of data packets in the buffer 112, and / or other information, such as the data transfer rate, the error rate, and the status of each port. After receiving this data, the processor 115 may perform further operations based on the information obtained, such as determining a moving average of the bandwidth, aggregating the bandwidths of the individual ports 106a - 106c, comparing the bandwidth with a threshold, and other functions described herein.

[0070] In one or more embodiments of the present disclosure, the computing system 103 (e.g., a switch) may communicate with a plurality of network nodes 200a - 200c, as Figure 2 shown, to form a network. Each network node 200a - 200c may be a computing system capable of sending and receiving data. Each node 200a - 200c may be any of a variety of devices, including but not limited to switches, personal computers, servers, or any other device capable of sending and receiving data in the form of data packets. For example, the nodes 200a - 200c may be the computing system 103 as Figure 1 shown. Each node 200a - 200c may communicate with a plurality of computing systems 103, such as the computing systems 103a and 103b as Figure 2 shown. The nodes 200a - 200c and the computing systems 103a - 103b may form a leaf - spine network, where the nodes 200a - 200c act as leaf switches, and the computing systems 103a - 103b act as spine switches. Figure 2 The limited number of nodes 200a - 200c and computing systems 103a - 103b shown in

[0071] should be considered for illustrative purposes only and should not be considered limiting in any way. For example, the network may include additional computing systems 103, more or fewer nodes 200a - 200c, and other levels of other computing devices, such as computing systems 103 that communicate with the nodes 200a - 200c through the computing systems 103a - 103b, or computing systems 103 that communicate with the computing systems 103a - 103b through the nodes 200a - 200c. Additionally, although each node 200a - 200c is shown as directly communicating with each computing system 103, it should be understood that other topologies and connection arrangements may be used.

[0071] Each computing system 103a - 103c can establish a communication channel with the network nodes 200a - 200c through a port. As Figure 2As shown, the computing system 103a communicates with the node 200a via the channel 203a connected to the port 106a, communicates with the node 200b via the channel 203b connected to the port 106b, and communicates with the node 200c via the channel 203c connected to the port 106c. Similarly, the computing system 103b communicates with the node 200a via the channel 203d connected to the port 106d, communicates with the node 200b via the channel 203e connected to the port 106e, and communicates with the node 200c via the channel 203f connected to the port 106f. Such channels can support data transmission in the form of data packet flows, following a predetermined protocol for control packet format, size, transmission method, and other aspects.

[0072] Each network node 200a - 200c can interact with each computing system 103a - 103b in various ways. The node 200 can send data packets to the computing system 103 for processing, transmission, or other operations, or forward them to another node 200 or another computing system 103. Conversely, each node 200 can receive data from the computing system 103, which may originate from the computing system 103 itself or other network nodes 200a - 200c or other computing systems 103 via the computing system 103. In this way, the computing systems 103a - 103b and the nodes 200a - 200c together form a network, thus facilitating data exchange, resource sharing, and other collaborative operations.

[0073] As Figure 3 and Figure 4 shown, based on routing decisions made by the nodes 200a - 200c and / or the computing systems 103a - 103b, the communication via the channels 203a - f can be active or inactive. When the ports 106a - 106f of the computing systems 103a - 103b do not receive (or send) data for a period of time, the ports 106a - 106f can be selectively disabled. In Figure 3 the example shown, the communication channels 203a - 203c are inactive (represented by dashed lines). Therefore, each of the ports 106a - 106c can be disabled by the computing system 103a. In Figure 4 the example shown, the communication channels 203a and 203e are inactive. Therefore, the port 106a can be disabled by the computing system 103a, and the port 106e can be disabled by the computing system 103b.

[0074] When the port 106 is disabled, the electronic device associated with that port (such as the SerDes circuitry 130) may be powered off. Therefore, when the port 106 is disabled, the computing system 103 may consume less or no electrical energy. When each port 106 of the computing system 103 (such as Figure 3When the ports 106a - 106c of the computing system 103a shown in [figure] are all disabled, the computing system 103 may enter a sleep mode, shut down, or otherwise be enabled to consume less or no electrical power. Using the systems or methods described herein, when it is determined that the ports 106 are not being fully utilized, the computing system 103 (e.g., a switch) can be enabled to reduce power consumption by sending an adaptive routing notification packet to the node 200. In response to such a notification packet, the node 200 can be configured to reroute packets if possible. Thus, traffic across multiple ports 106 and the computing system 103 can be consolidated, and under - utilized ports can be avoided, so that the ports 106 can be disabled to reduce power.

[0075] Figure 3 and Figure 4 The scenario shown in [figure] may be the result of an adaptive routing notification packet generated by the computing system 103a and sent to each of the nodes 200a - 200c. Because it is determined that the traffic received at each of the ports 106a - 106c is below a threshold, the adaptive routing notification packet can be generated by the computing system 103a. The adaptive routing notification packet generated by the computing system 103a can effectively convey to the nodes 200 that receive the adaptive routing notification packet the information that the computing system 103a is experiencing congestion, while the computing system 103a is actually experiencing the opposite of congestion. As a result of receiving such an adaptive routing notification packet, the nodes 200 can (if possible) reroute traffic to avoid the computing system 103a. The nodes 200 can be programmed to avoid or resolve congestion at the computing system 103 by rerouting traffic in response to the adaptive routing notification packet; however, by using the systems or methods described herein, the nodes 200 will help the computing system 103 reduce power consumption by rerouting traffic from the less - used ports 106. This technological advancement can be achieved by performing one or more of the methods 500, 600, 700 as shown in [figure] and described below. Figures 5 - 7 shown and described below.

[0076] Figure 3 An example implementation is shown where the computing system 103a determines that the aggregate total bandwidth of the traffic bandwidth received at each of the ports 106a - 106c is below an aggregate bandwidth threshold. In response to this determination, the computing system 103a has sent an adaptive routing notification packet to each of the nodes 200a - 200c and has stopped receiving traffic from each of the nodes 200a - 200c. After not receiving traffic from the nodes 200a - 200c for a period of time, the computing system 103a has disabled the ports 106a - 106c, entered a sleep mode, or otherwise reduced the amount of power consumption. This implementation can be as in Figure 5 method 500 of Figure 7as shown in method 700 and described below.

[0077] As Figure 5 shown, and according to the computing system 103 as shown in Figure 1 and described herein, method 500 can be executed to cause the computing system 103 to indicate to a node sending data to the computing system 103 to reroute traffic away from the computing system 103 or be more likely to reroute traffic away from the computing system 103 in the case of receiving relatively low traffic on multiple ports. Accordingly, the computing system 103 can reduce power consumption by turning off or reducing the power of switching hardware and / or other power-consuming components associated with the ports. Although the description of method 500 provided herein describes the steps of method 500 as being performed by the processor 115 of the computing system 103, the steps of method 500 can be performed by one or more of the processors 115, switching hardware 109, one or more controllers in the computing system 103, or some combination thereof.

[0078] At 503, the computing system 103 can monitor the total bandwidth of all ports 106 or a subset of ports of the computing system 103. Monitoring the total bandwidth can include, in some embodiments, monitoring the individual bandwidth of each of the multiple ports and aggregating the individual bandwidths to calculate the total bandwidth. Monitoring the bandwidth of one or more ports can be performed in one or more of a variety of ways, such as by counting the number of data packets and / or bytes sent from each port over a particular time range, measuring the rate of data packets sent over a particular time range (e.g., packets per second or bits per second), monitoring the length of one or more queues associated with one or more ports, using a flow analysis mechanism, a traffic sampling mechanism, monitoring the amount of data in one or more shared buffers serving multiple ports, or any other conceivable way of tracking port bandwidth. The monitored bandwidth can be ingress bandwidth or can be egress bandwidth.

[0079] At 506, the computing system 103 can determine whether the total bandwidth is less than a threshold. The threshold can be a user-configurable data point stored in the memory 118 of the computing system 103. The threshold can represent a number of bits, number of data packets, bit rate, packet rate, or other quantifiable metric. Determining whether the total bandwidth is less than the threshold can be performed periodically at a particular rate, such as, for example, once every 10 milliseconds, or continuously.

[0080] In this document, various comparisons are made between the bandwidth state and a predetermined threshold. It should be understood that any reference to determining whether a variable is "less than" a threshold may alternatively or additionally include "less than or equal to" the threshold. Similarly, for determining whether a variable is "less than or equal to" a threshold, it should be understood that such determination may also or alternatively be made based only on whether the variable is "less than" the threshold. This interpretation also applies to the determination of a variable being "greater than" or "greater than or equal to" a threshold.

[0081] In some embodiments, the total bandwidth threshold may be set by a user or system administrator, while in other embodiments, the total bandwidth threshold may be automatically set by the computing system 103, such as by calculating the total bandwidth threshold. In some embodiments, calculating the total bandwidth threshold may involve determining the network topology, determining the number of backbone switches in the network, or otherwise making determinations related to possible alternative routes of traffic.

[0082] In some embodiments, the total bandwidth threshold may be directly or indirectly related to the network topology, the number of backbone switches in the network, or other network design considerations. The thresholds described herein may be at least partially related to or dependent on the total bandwidth, the total network capacity, or the amount of traffic across the computing system network. For example, when the network of which the computing system 103 is a part occupies a relatively high total bandwidth, method 500 may require the total bandwidth to be lower than a relatively low threshold before shutting down the switching hardware of the computing system 103. Similarly, when the network of which the computing system 103 is a part occupies a relatively low total bandwidth, method 500 may only require the total bandwidth to be lower than a relatively high threshold before shutting down the switching hardware of the computing system 103.

[0083] Method 500 may be useful in scenarios where the total network bandwidth (including the bandwidth across multiple backbone switches) is low relative to the total network capacity of the network. When the total network bandwidth is low relative to the total network capacity of the network, consolidating the flow of traffic by reducing the number of active but underutilized ports can be used to reduce the total power consumption of the network without affecting performance.

[0084] At 509, if the total bandwidth is less than the threshold, the computing system 103 may generate an instruction packet, such as an adaptive routing notification data packet. Generating the instruction packet may be an optional step. For example, it may not be necessary to generate the instruction packet, and the computing system 103 may be enabled to transmit the instruction packet without generating it. The instruction packet may include information that the node 200 can use to consider rerouting traffic.

[0085] In some embodiments, the computing system 103 may be configured to send instruction packets at a specific maximum rate, such as a limit of one instruction packet per port per second, but it should be understood that this limit may not be necessary in some embodiments.

[0086] At 512, the computing system may transmit instruction packets via each port 106. Transmitting the instruction packets may include sending the instruction packets as part of an acknowledgement or reply to a data packet received from port 106. The transmission may also or alternatively be performed without responding to any particular data packet and may be executed without any prompting from the source node.

[0087] If it is determined at 506 that the total bandwidth is greater than or equal to the bandwidth threshold, method 500 may include determining whether method 500 should continue at 515 and end or return to the above-mentioned monitoring step 503 at 518. Method 500 may be continuously executed during the operation of the computing system or may be run based on a specific schedule and / or application. For example, certain applications may require maximum performance while sacrificing power efficiency, and when such applications are executed, the computing system may end power efficiency methods such as method 500.

[0088] At any point in time before, during, or after the above method 500, the processor may also execute method 700, as described in more detail below, wherein when no traffic is received through a particular port or no traffic is received by the entire computing system 103, the switching hardware and / or other power-consuming components of the computing system 103 may be turned off, enter a sleep mode, or otherwise operate in a low-power mode.

[0089] As described above, Figure 4 An example implementation is shown where the first computing system 103a determines that the bandwidth of the traffic received at port 106a is below the per-port bandwidth threshold, while the second computing system 103b determines that the bandwidth of the traffic received at port 106e is below the per-port bandwidth threshold. In response to such determination, the computing system 103a has sent an adaptive routing notification data packet to node 200a, while the computing system 103b has sent an adaptive routing notification data packet to node 200b. After sending the adaptive routing notification, the computing system 103a has stopped receiving traffic from node 200a, while the computing system 103b has stopped receiving traffic from node 200e. After not receiving traffic from node 200a for a period of time, the computing system 103a has disabled port 106a or otherwise reduced the amount of power consumption. After not receiving traffic from node 200b for a period of time, the computing system 103b has disabled port 106e or otherwise reduced a certain amount of power consumption. This implementation may be as Figure 6 method 600 of Figure 7 and method 700 of

[0090] As Figure 6 shown and in accordance with as Figure 1The computing system 103 shown and described herein can execute method 600 to cause the computing system 103 to indicate to a node sending data to the computing system 103 to reroute or be more likely to reroute traffic away from the computing system 103 in the case of relatively low traffic received through a particular port. Accordingly, the computing system 103 can reduce power consumption by turning off or reducing the power of switching hardware and / or other power-consuming components associated with the particular port. Although the description of method 600 provided herein describes the steps of method 600 as being performed by processor 115 of computing system 103, the steps of method 600 can be performed by one or more of processors 115, switching hardware 109, one or more controllers in computing system 103, or some combination thereof. As described below, Figure 6 method 600 can be performed separately for each of a plurality of ports. For example, method 600 can be performed in parallel for each port or can be performed serially for each port.

[0091] At 603, the computing system 103 can monitor the bandwidth of each of one or more ports 106 of the computing system 103. Monitoring the bandwidth of one or more ports can be performed in one or more of a variety of ways, such as by counting the number of data packets and / or bytes sent from each port over a particular time range, measuring the rate of data packets sent over a particular time range (e.g., packets per second or bits per second), monitoring the length of one or more queues associated with one or more ports, using a flow analysis mechanism, a traffic sampling mechanism, monitoring the amount of data in one or more shared buffers serving multiple ports, or any other conceivable way of tracking port bandwidth. The monitored bandwidth can be ingress bandwidth or can be egress bandwidth.

[0092] At 606, the computing system 103 can determine whether the bandwidth of any port is less than a threshold. The threshold can be a user-configurable data point stored in the memory 118 of the computing system 103. The threshold can represent a number of bits, a number of data packets, a bit rate, a packet rate, or other quantifiable metric. Determining whether the bandwidth of any particular port is less than the threshold can be performed periodically at a particular rate, such as once every 10 milliseconds, or continuously.

[0093] In some embodiments, the per-port bandwidth threshold may be set by a user or system administrator, while in other embodiments, the per-port bandwidth threshold may be automatically set by the computing system 103, such as by calculating the per-port bandwidth threshold. In some embodiments, calculating the per-port bandwidth threshold may involve determining the network topology, determining the number of backbone switches in the network, or otherwise making determinations related to possible alternative routes of traffic. For example, the per-port bandwidth threshold may be directly or indirectly related to the network topology, the number of backbone switches in the network, or other network design considerations. In some embodiments, the thresholds described herein may be at least partially related to or dependent on the total bandwidth or traffic volume of the computing system network or the total bandwidth of each port of the computing system 103. For example, when the network to which the computing system 103 belongs occupies a relatively high total bandwidth or when the ports of the computing system 103 generally occupy a relatively high bandwidth, method 600 may require that the per-port bandwidth of a port be below a relatively low threshold before shutting down the switching hardware associated with that port. Similarly, when the network to which the computing system 103 belongs occupies a relatively low total bandwidth or when the ports of the computing system 103 generally occupy a relatively low bandwidth, method 600 may require that the per-port bandwidth of a port be below a relatively high threshold before shutting down the switching hardware associated with that port.

[0094] At 609, if the bandwidth of the port is less than the threshold, the computing system 103 may generate an instruction packet, such as an adaptive routing notification data packet. Generating the instruction packet may be an optional step. For example, it may not be necessary to generate the instruction packet, and the computing system 103 may be able to transmit the instruction packet without generating it. The instruction packet may include information that can be used by the node 200 to consider rerouting traffic.

[0095] In some embodiments, the computing system 103 may be configured to send the instruction packet at a specific maximum rate, such as a limit of one instruction packet per port per second, but it should be understood that such a limit may not be necessary in some embodiments.

[0096] At 612, the computing system may transmit the instruction packet via the port 106 whose bandwidth is less than the threshold. Transmitting the instruction packet may include sending the instruction packet as part of an acknowledgment or reply to a packet received from the port 106. The transmission may also or alternatively be performed without responding to any specific data packet and may be carried out without any prompting from the source node.

[0097] If it is determined at 606 that the individual bandwidth of each port is greater than or equal to the bandwidth threshold, method 600 may include determining whether method 600 should continue at 615 and end or return to the above-mentioned monitoring step 603 at 618. Method 600 may be continuously executed during the operation of the computing system, or may be run based on a specific schedule and / or application. For example, certain applications may require maximum performance while sacrificing power efficiency, and when such applications are executed, the computing system may end power efficiency methods such as method 600.

[0098] In some embodiments, method 600 may be executed for one port at a time. At the start of method 600, the computing system 103 may select a first port to determine whether its bandwidth is less than the threshold. After the determination is made at 606 (and if the bandwidth of the first port is less than the threshold, an instruction packet is sent at 612), the computing system 103 may select a second port to determine whether its bandwidth is less than the threshold.

[0099] In some embodiments, selecting the second port may involve using a polling schedule or another selection protocol or algorithm. For example, in at least one embodiment, at the start of method 700, the computing system 103 may activate a polling schedule iterator. The polling schedule iterator may be a software or hardware algorithm designed to generate and maintain an ordered sequence of identifiers. A dynamic pointer may be used to select the first port for which to monitor the bandwidth. After determining whether the bandwidth of the first port is less than the threshold, the pointer may advance to the next port. With each successive iteration, the pointer may advance to the subsequent port. Once each port has been monitored, the pointer may return to the first port, and the loop may continue.

[0100] At any point in time before, during, or after the above-described method 600, the processor may also execute method 700, described in more detail below, where when no traffic is received through a particular port or no traffic is received by the overall computing system 103, the switching hardware and / or other power-consuming components of the computing system 103 may be turned off, enter a sleep mode, or otherwise operate in a low-power mode.

[0101] Figure 7 An example method 700 for deactivating switching hardware components due to the detection of no activity at one or more ports during a particular time period is shown. The method may start at 703, where the computing system 103 may monitor the usage of one or more ports of the computing system 103. Monitoring the ports may include, in some embodiments, tracking the amount of data sent and / or received from the ports, monitoring the data in one or more queues associated with the ports, or otherwise tracking port usage. The monitoring of the ports may be performed by a processor or dedicated or general-purpose circuitry of the computing system 103.

[0102] At 706, it can be determined whether any activity has occurred on a port within a certain time period. The computing system 103 can continuously and individually monitor each port in parallel, or can serially and continuously check the usage of each port. The computing system 103 can be enabled to determine whether any data has been sent and / or received on each port within a specific previous amount of time. The amount of time may vary depending on configuration settings and can be set by a user or an administrator or automatically by the computing system 103.

[0103] If for any particular port no data is detected to be sent and / or received within a specific previous amount of time, the computing system 103 can deactivate one or more switching hardware elements of that particular port at 709. Deactivating the switching hardware elements can include deactivating one or more SerDes circuits or other circuitry that consumes power when idle. For example, one or more circuitry systems or circuitry associated with a port or with packet forwarding may consume power even when not actively participating in packet forwarding. Although the power consumed by these hardware elements when idle may be less than or equal to the power consumed by these hardware elements when actively participating in packet forwarding, the power consumed may not be negligible. By deactivating these hardware elements, the computing system 103 can reduce power consumption when the port associated with these circuits or circuitry systems is not in use. By deactivating the hardware elements described herein, a reduction in power consumption can be achieved. In some embodiments, the computing system 103 can be enabled to put a port into a sleep mode or otherwise reduce the power consumption associated with a port where no data transmission and / or reception is detected within a specific time. If all ports of the computing system 103 are deactivated, the computing system 103 can also enter a sleep mode or shut down.

[0104] If at 706 it is detected that no particular port has not sent and / or received data within a specific previous amount of time, or after deactivating the switching hardware elements at 709, method 700 can include determining at 712 whether to continue. Method 700 can be continuously executed during the operation of the computing system, or can be run based on a specific schedule and / or application. For example, certain applications may require maximum performance while sacrificing power efficiency, and when such applications are being executed, the computing system can terminate power efficiency methods such as method 700.

[0105] If the computing system 103 determines to continue method 700, method 700 can include returning to monitoring one or more ports at 703. If the computing system 103 determines not to continue method 700, method 700 can end at 715.

[0106] This disclosure encompasses Figures 5 to 7 methods with fewer than all the steps identified in (and the corresponding method descriptions), and includingFigures 5 to 7 (and corresponding method descriptions) for methods with additional steps beyond those identified. The present disclosure also encompasses methods that include one or more steps of the methods described herein and one or more steps of any other method described herein.

[0107] Embodiments of the present disclosure include a system for providing adaptive routing, the system including one or more circuits for: receiving data from a network via a first port; determining a bandwidth; determining that the bandwidth is below a threshold; generating an instruction packet in response to determining that the bandwidth is below the threshold; and sending the instruction packet via the first port.

[0108] Aspects of the above system include where the bandwidth includes the bandwidth of the first port, and where the one or more circuits are further configured to determine a second bandwidth associated with a second port.

[0109] Aspects of the above system include where the bandwidth includes the total bandwidth of a plurality of ports including the first port, and where determining the bandwidth includes determining the total bandwidth of each port in the plurality of ports.

[0110] Aspects of the above system include where the bandwidth includes an egress bandwidth, and the first port includes an egress port.

[0111] Aspects of the above system include where the one or more circuits are further configured to enter a sleep mode after the link has been idle for a period of time.

[0112] Aspects of the above system include where entering the sleep mode includes disabling one or more serializer / deserializer (SerDes) circuits associated with the link.

[0113] Aspects of the above system include where, after sending the instruction packet, traffic at the first port is stopped.

[0114] Aspects of the above system include where the total network bandwidth is low relative to the total network capacity.

[0115] Aspects of the above system include where the first port includes a port of a backbone switch, and where the data is received from a top-of-rack (TOR) switch.

[0116] Aspects of the above system include where the first port includes a port of an L2 switch, and where the data is received from an L3 switch, and where the data is forwarded to an L1 switch.

[0117] Embodiments of the present disclosure also include a system for providing adaptive routing, the system including one or more circuits for: receiving data from a network; determining a total bandwidth associated with the received data; determining that the total bandwidth associated with the received data is below a threshold; generating an instruction packet in response to determining that the total bandwidth is below the threshold; and sending the instruction packet via a plurality of ports.

[0118] Aspects of the above system include that the total bandwidth includes a total egress bandwidth.

[0119] Aspects of the above system include that sending the instruction packet includes sending data packets to a plurality of destinations via a plurality of ports.

[0120] Aspects of the above system include that the one or more circuits are further for entering a sleep mode after a link associated with the plurality of ports has been idle for a period of time.

[0121] Aspects of the above system include that entering the sleep mode includes disabling one or more serializer / deserializer (SerDes) circuits associated with the link.

[0122] Aspects of the above system include that after sending the instruction packet, traffic at the plurality of ports is stopped. Embodiments of the present disclosure also include a system for providing adaptive routing, the system including one or more circuits for: receiving data via a first port among the plurality of ports; determining a bandwidth associated with the received data; determining that the bandwidth associated with the received data is below a threshold; generating an instruction packet in response to determining that the bandwidth is below the threshold; and sending the instruction packet through the first port.

[0123] Aspects of the above system include that before determining the bandwidth associated with the first port, the one or more circuits are for selecting the first port using a selection process.

[0124] Aspects of the above selection process include that the selection process includes round-robin scheduling.

[0125] Aspects of the above system include that the one or more circuits are further for selecting a second port using the selection process after sending the instruction packet.

[0126] It should be understood that any feature described herein can be claimed in combination with any other feature described herein, whether or not those features are from the same embodiment described.

[0127] Specific details are given in the specification to provide a thorough understanding of the embodiments. However, those of ordinary skill in the art will understand that the embodiments can be practiced without these specific details. In other instances, well-known circuits, processes, algorithms, structures, and techniques have not been shown in unnecessary detail to avoid obscuring the embodiments.

[0128] Although illustrative embodiments of the present disclosure are described in detail herein, it should be understood that the inventive concept may be embodied and used in other various ways, and the appended claims are intended to be construed to include such variations, unless limited by the prior art.

Claims

1. A system for providing adaptive routing, the system comprising one or more circuits, wherein the one or more circuits are configured to: receiving data from the network through the first port; Determine bandwidth; determining that the bandwidth is below a threshold; generating an instruction packet in response to determining that the bandwidth is below the threshold; and The instruction packet is sent through the first port.

2. The system of claim 1, wherein the bandwidth comprises a bandwidth of the first port, and wherein the one or more circuits are further configured to determine a second bandwidth associated with a second port. 3 . The system of claim 1 , wherein the bandwidth comprises a total bandwidth of a plurality of ports including the first port, and wherein determining the bandwidth comprises determining a total bandwidth of each of the plurality of ports.

4. The system of claim 1, wherein the bandwidth comprises an egress bandwidth and the first port comprises an egress port.

5. The system of claim 1, wherein the one or more circuits are further configured to enter a sleep mode after the link has been idle for a period of time.

6. The system of claim 5, wherein entering the sleep mode comprises: One or more serializer / deserializer (SerDes) circuits associated with the link are disabled. The system of claim 1 , wherein after sending the instruction packet, traffic at the first port is stopped.

8. The system of claim 7, wherein the total network bandwidth is low relative to the total network capacity.

9. The system of claim 1, wherein the first port comprises a port of a spine switch, and wherein the data is received from a top-of-rack (TOR) switch.

10. The system of claim 1, wherein the first port comprises a port of an L2 switch, and wherein the data is received from an L3 switch, and wherein the data is forwarded to an L1 switch.

11. A system for providing adaptive routing, the system comprising one or more circuits, the one or more circuits being configured to: Receive data from the network; determining a total bandwidth associated with the received data; determining that the total bandwidth associated with the received data is below a threshold; generating an instruction packet in response to determining that the total bandwidth is below the threshold; as well as The instruction packet is sent through multiple ports.

12. The system of claim 11, wherein the total bandwidth comprises a total egress bandwidth.

13. The system of claim 11, wherein sending the instruction packet comprises: The instruction packet is sent to multiple destinations through the multiple ports.

14. The system of claim 11, wherein the one or more circuits are further configured to enter a sleep mode after links associated with the plurality of ports have been idle for a period of time.

15. The system of claim 14, wherein entering the sleep mode comprises: One or more serializer / deserializer (SerDes) circuits associated with the link are disabled.

16. The system of claim 11, wherein after sending the instruction packet, traffic at the plurality of ports is stopped.

17. A system for providing adaptive routing, the system comprising one or more circuits for: receiving data through a first port of the plurality of ports; determining a bandwidth associated with the received data; determining that the bandwidth associated with the received data is below a threshold; generating an instruction packet in response to determining that the bandwidth is below the threshold; as well as The instruction packet is sent through the first port.

18. The system of claim 17, wherein prior to determining the bandwidth associated with the first port, the one or more circuits are configured to select the first port using a selection process.

19. The system of claim 18, wherein the selection process comprises a round-robin schedule.

20. The system of claim 18, wherein the one or more circuits are further configured to select a second port using the selection process after sending the instruction packet.