Load-Aware ECMP with Flow Tables
By implementing central units and state machines on semiconductor chips, dynamically evaluating path selection is solved, and the problem of failure to dynamically adjust paths in existing ECMP routing technology is achieved, and dynamic load balancing and efficient path selection of network performance is achieved.
Patent Information
- Application Number
- CN202210207207.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-11-30
- Filing Date
- 2022-03-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-03-04
AI Technical Summary
The existing ECMP routing technology fails to dynamically consider network load and congestion when selecting paths, resulting in network performance degradation. The existing software solutions have a long response time and are unable to effectively group and reorder.
The semiconductor chip is used to realize load-aware ECMP, dynamically evaluate path selection through central units and state machines, combined with upper and lower-layer network information, dynamically adjust path selection to reduce packet reordering.
Dynamic load balancing of network performance is realized, congestion and tail delay are reduced, path selection accuracy and response speed are improved, and chip area and power costs are reduced.
Smart Images

Figure CN115208816B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application is a partial continuation application of U.S. Patent Application No. 17 / 230,940, filed Apr. 14, 2021, which is incorporated herein by reference in its entirety. Technical Field
[0003] This description generally relates to Ethernet communications, and more particularly, to load-aware equal-cost multi-path (ECMP) routing implementations with flow tables. Background Art
[0004] Equal-cost multi-path (ECMP) routing is a routing policy in which packet forwarding to a single destination can occur over multiple best paths with equal routing priority. Multi-path routing can be used in conjunction with most routing protocols because it is a per-hop decision made independently at each router. It can significantly increase bandwidth by load-balancing traffic over multiple paths; however, there can be significant limitations when deploying multi-path routing in practice. For example, ECMP route selection is essentially fixed to the flow hash % ECMP group size, which returns the same result across different network nodes. Hash algorithms are usually not perfect; for example, there may be biases and the distribution may be uneven. A more significant problem with ECMP is the lack of a concept of checking for instantaneous load or congestion while selecting paths. Paths are fixed based on the flow hash and are pre-programmed statistically. Even when someone tries to program the desired path via software (S / W), there are severe response time limitations and no mechanism for reordering packets, making it practically useless. Summary of the Invention
[0005] In one aspect, the present application provides a semiconductor chip for implementing load-aware equal-cost multi-path (ECMP), the semiconductor chip comprising: a plurality of ports; a plurality of pipelines, each pipeline coupled to a portion of the plurality of ports; and a central unit including a state machine and a plurality of databases, wherein: the plurality of databases are configured to contain information about a communication network including an upper network and a lower network, and the state machine is implemented in hardware and configured to determine at least one characteristic of corresponding path groups within the upper network and the lower network to overcome dynamic network conditions.
[0006] In another aspect, the present application provides a method of implementing load-aware Equal-Cost Multi-Path (ECMP) routing, the method comprising: configuring a plurality of databases to store information about a communication network comprising an upper network and a lower network; and configuring a state machine implemented in hardware to determine at least one characteristic of corresponding path groups within the upper network and the lower network to overcome dynamic network situations; implementing the plurality of databases and the state machine on a semiconductor chip, the semiconductor chip comprising a plurality of pipelines and a plurality of ports; and coupling each of the plurality of pipelines to a portion of the plurality of ports on the semiconductor chip.
[0007] In another aspect, the present application provides a system comprising: a memory; one or more processors coupled to the memory and configured to execute instructions to perform the following actions: programming a group table associated with an upper network and a lower network; and programming the members and ports of each group to a Next-Hop (NH) mapping. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Certain features of the technology are set forth in the appended claims. However, for explanatory purposes, several embodiments of the technology are set forth in the drawings.
[0009] Figure 1 is a block diagram illustrating an example of an abstract diagram of a pipeline of a load-aware Equal-Cost Multi-Path (ECMP) routing system according to various aspects of the present technology.
[0010] Figure 2 is a block diagram illustrating an example of a mapping for dynamic evaluation of desired paths for ECMP according to various aspects of the present technology.
[0011] Figure 3A and 3B is a block diagram illustrating an example of a static ECMP mapping and a load-aware dynamic ECMP mapping to be read by a packet according to various aspects of the present technology.
[0012] Figure 4 is a schematic diagram illustrating an example of a per-chip diagram of load-aware ECMP according to some aspects of the present technology.
[0013] Figure 5 is a flowchart illustrating an example of a software (S / W) process according to some aspects of the present technology.
[0014] Figure 6 is a flowchart illustrating an example of a central module process according to some aspects of the present technology.
[0015] Figure 7 is a flowchart illustrating an example of a pipeline process according to some aspects of the present technology.
[0016] Figure 8An electronic system in which some aspects of the present technology are implemented. Detailed Description
[0017] The detailed description set forth below is intended as a description of various configurations of the present technology and is not intended to represent the only configurations in which the present technology may be practiced. The accompanying drawings are incorporated herein and constitute a part of the detailed description, which includes specific details for providing a thorough understanding of the present technology. However, the present technology is not limited to the specific details set forth herein and may be practiced without one or more of the specific details. In some instances, structures and components are shown in block diagram form in order to avoid obscuring the concepts of the present technology.
[0018] The present technology relates to methods and systems for load-aware equal-cost multi-path (ECMP) routing implementations with flow tables. The disclosed ECMP techniques help improve network performance factors such as reducing network congestion and latency that effectively minimize packet loss. In existing solutions, ECMP groups and corresponding members are programmed statistically and do not take into account dynamic network conditions such as instantaneous load / congestion on a port or corresponding queue and the available / unavailable state of the port and / or link indicated by the next-hop index. If an output port is found to be unavailable in a given router, the protection switching logic provides pre-programmed alternative second and (if the second is also unavailable, then) third choices without knowledge of the dynamic congestion / load on those alternative ports. The entire table structure may need to be reprogrammed to address the unavailable port, which may take a significant amount of time to detect and then correct via software (S / W).
[0019] The load-aware ECMP technology of the present technology enables dynamic load balancing of the entire ECMP infrastructure with relatively very small chip area and / or power cost. Load-aware ECMP also helps reduce congestion in the network and significantly reduce tail latency. Biases, weights, and / or vectors can help account for parameters affecting additional hops in the underlay and overlay networks for desired path selection, which can dynamically improve network performance. The overlay network (overlay) is a virtual network built on top of the underlying network infrastructure and / or network layer (underlay network or underlay). Existing solutions require updating the entire underlay and overlay programming in response to dynamic network conditions, which is very time-consuming and can disrupt the entire network. The same desired selection also applies to the second and third protection switching logic alternatives to avoid congestion on those links in the case of unavailable ports, thus avoiding network bottlenecks.
[0020] In the load-aware ECMP technology of the present technology, each pipeline (e.g., the data communication pipeline between two nodes) replicates the minimum amount of information because there is no need to have deviation / weight / vector programming inside the pipeline and no need for port / queue quality metrics in each pipeline. All programming and port / queue quality metrics are performed by the central module (only one upper copy for the whole chip), thus saving significant area, power, and cost for chips with higher bandwidth and thus more pipelines. The area, power, latency, and cost of the load-aware ECMP solution are a small fraction of any other per-pipeline dynamic load balancing solution. This is because path selection is not done locally for each packet in the pipeline, but is done centrally by a single state machine for all pipelines; there is no need for an expensive timestamp mechanism to avoid reordering; it also supports dynamic desired selection for the upper layer based on the corresponding deviation / weight / vector without any packet reordering; it supports not only per-port but also per-port per-queue and quality metrics, which improves the selection granularity to be as low as possible; and it dynamically improves the accuracy of the selection of the desired port / queue and significantly expands the application scope of this load balancing. The disclosed central state machine is implemented in hardware, which is significantly faster and more beneficial for any exception handling, update of unavailable ports / links, and protection switching compared to the traditional implementation in S / W.
[0021] Figure 1FIG. is a block diagram illustrating an example of an abstract diagram of a pipeline 100 for load-aware ECMP according to various aspects of the present technology. The abstract diagram shows a pipeline 100 for load-aware ECMP routing that includes: level 1, an upper network 110; level 2, a lower network 120; protection switching logic 130; and a central module (also referred to as a central unit) 140. The upper network 110 includes an ECMP group table 112, an ECMP member table 114, and a flow table 116. The lower network 120 includes an ECMP group table 122, an ECMP member table 124, and a flow table 126. The ECMP group table 112 receives a level 1 group index 102 from a forwarding lookup of a packet or from an access control list and generates a group property 103. A group in the context of the present disclosure is a set of paths with a given characteristic. For example, from a source location (e.g., San Jose) to a destination (e.g., New York), there may be several (e.g., 10) paths. The group property 103 is passed to the ECMP members of the ECMP member table 114. The index of an ECMP member is given by a group base provided by the ECMP group table 112 + a group size (hash % group_size) that is a hash value calculated based on packet characteristics. The ECMP group table 122 receives an upper layer next hop 104 from the ECMP member table 114 and generates a group property 105. The group property 105 is forwarded to the ECMP members of the ECMP member table 124. The index of an ECMP member is given by group_base + hash % group_size. The output of the ECMP member table 124 is passed through the protection switching logic 130 for a lower layer next hop (NH) to reach the NH table to derive the outgoing property of the packet. NH refers to the nearest router / switch through which a packet can pass in a network.
[0022] The central module 140 is coupled to all pipelines on a given chip and implements many important features of the present technology. The central module 140 remedies the drawbacks of the existing ECMP solutions discussed above. The central module 140 includes databases for level 1 (upper layer) and level 2 (lower layer) group tables (e.g., ECMP group tables 112 and 124) and corresponding members (e.g., ECMP member tables 114 and 124), NH to port mapping, upper layer NH to lower layer group mapping, and dynamic port states. In the central module 140, there is also a database for any deviation, weight, and / or vector and threshold for paths beyond the current hop, as well as instantaneous load and congestion information of ports and queues per port. The state machine in the central module 140 enters level 1 groups and / or members and level 2 groups entry by entry, and processes the load, deviation, and port state information to select the desired each destination NH and / or group to be filled in the level 1 and / or level 2 member tables. The desired NH and / or group at any level is the NH and / or group with less load compared to other NHs and / or groups. Since the state machine has complete information about the device ports for each NH, the load and / or congestion on each of those ports, and information about remote ports (on different devices) through the deviation, it can determine which are the NHs and / or groups with the least load. The central module 140 provides periodic updates 142 to the upper layer network 110 and the lower layer network 120. When the hit bit corresponding to the index is 1, it is set to 0 and the member table is not updated. When the hit bit is 0, the member table is updated with the desired NH and / or group. The programmed inactive period of the central module 140 determines the update frequency.
[0023] Flow tables 116 and 126 are provided by the present technology to perform more fine-grained control in conjunction with the controls described above. The level 1 and level 2 group tables (e.g., ECMP group tables 112 and 124) can be enabled respectively to use the flow tables 116 and 126. The flow tables 116 and 126 can record any number of flows per group and are only limited by the flow table size. The flow does not change the desired selection unless the minimum inactive period for the previous selection is reached using the hit bit mechanism described herein. This will avoid packet reordering and interruption to the network. A configurable override for this behavior is also provided. If the flow remains inactive within the programmed period (no packets of the flow are received within that period), then the flow will expire (be deleted from the flow table) to reduce the flow table size requirement.
[0024] Figure 2It is a block diagram showing an example of a mapping 200 for dynamic evaluation of desired paths for load-aware ECMP according to various aspects of the present technology. In the mapping 200 for dynamic evaluation, a first host A is connected to a first router R1, and the first router R1 is in turn connected to routers R2 and R3 of a lower-layer ECMP group 210 and routers R4 and R5 of a lower-layer ECMP group 212 through paths (links) P5, P7, P15, and P19. Routers R2 and R3 are connected to a router R6 of an upper-layer ECMP group 214 through paths P1 and P9, and routers R4 and R5 are connected to a router R7 of the upper-layer ECMP group 214 through paths P21 and P2. Routers R6 and R7 are connected to a second host B through paths P25 and P22.
[0025] The router only has information about the load on the port and / or the queue on its own port. The quality of the link beyond the current hop needs to be programmed as a deviation / weight / vector in the group table. For example, when links (paths) P9 and P22 are instantaneously congested due to bandwidth limitations, for example, router R1 may or may not see the resulting load and / or congestion on its port P7 and the corresponding queue. However, the deviation for the R3-to-R6 link (P9) needs to be programmed in the lower-layer group table (e.g., Figure 1 122), and moreover, the deviation for the R7-to-B (P22) link must be programmed in the upper-layer group table (e.g., Figure 1 112). To address this situation, Figure 1 the central state machine of the central module 140 can select R6 for the upper layer and R2 for the lower layer as needed at that time. When the R7-to-B link (P22) is finally cleared, the state machine selects R7 as the upper layer as the desired choice at that time. However, the hit-bit mechanism will ensure that the update to the pipeline table will wait for the programmed minimum inactive period with respect to the previous selection. This in turn ensures that the packets sent on that path in a given pipeline and a given flow are always earlier than the packets that will be sent on the newly updated path, thus avoiding unwanted reordering of the packets.
[0026] To ensure that there is no packet reordering when the desired port selection changes for a given packet flow, there is a programmed minimum inactivity period before the change is actually updated into the member tables in the pipeline. The central state machine sets the hit bit corresponding to entry 0, and if there is no packet hitting the entry, the hit bit is set to 1 when the state machine returns to read it again after a specific period. The minimum inactivity period will cause all previous packets from the same flow that always used the previous path to be earlier than the packets that will see the updated path. The minimum inactivity for a given member entry is implemented as follows. Each entry in the level 2 member table maintains a hit bit that is set to 1 when the entry is referenced by a packet. This bit is checked during the periodic state machine update of the member table. If the checked hit bit is 1, then it is set to 0 and the entry is not updated. If the hit bit is 0, then the entry is updated. If the update rate of the state machine for each entry corresponds to the programmed minimum activity period, then that period will be automatically implemented for the minimum inactivity check.
[0027] Figure 3A and 3B is a block diagram illustrating examples of static ECMP mapping in pipeline 300A and load-aware dynamic ECMP mapping in pipeline 300B according to various aspects of the present technology. The static ECMP mapping is performed in the pipeline data path of pipeline 300A to be read by packets. As Figure 3A shown, the ECMP group table 310 provides group base and size information 302 to the ECMP member table 320. The size information 302 identifies the group base 322 and the group size 4, which includes NH NH1, NH2, NH3, and NH4. In pipeline 300A, all applicable members are pre-programmed. In the protection switching logic 330, a second NH is randomly selected from the group using the flow hash, and a third NH is statically selected from the group.
[0028] The dynamic ECMP mapping of the present technology is performed in the pipeline data path of pipeline 300B to be read by packets. As Figure 3BAs shown, the ECMP group table 310 similarly provides the group base and size information 302 to the ECMP members and the hit bit table 340. The size information 302 identifies the group base 342 and the group size 4. The desired selection can be extended to per-queue per-port at the lowest granularity. In an instance of two queues per port, the group base 342 contains the first desired selection (NH2, NH1, NH2, and NH1) and the second desired selection (NH5, NH1, NH2, and NH8) for each destination. The central state machine takes into account the instantaneous load on the ports and / or queues corresponding to the NHs and is programmed to account for any deviation / weight / vector in the parameters beyond this hop before selecting the desired NH for each destination and the second and third NHs. The protection switching logic 350 can select the desired second and third NHs. The first desired selection is for incoming packets with service class 0, and the second desired selection is for incoming packets with service class 1.
[0029] Figure 4 FIG. 400 is a schematic diagram illustrating an example of a per-chip diagram of load-aware ECMP according to some aspects of the present technology. The per-chip diagram 400 shows four pipelines including pipeline 410 (pipeline 1), pipeline 420 (pipeline 2), pipeline 430 (pipeline 0), and pipeline 440 (pipeline 3). Each of the pipelines 410, 420, 430, and 440 is connected to one quarter of the ports on the chip. The central ECMP database and state machine 450 receives port status and load information 452 and communicates with the pipelines 410, 420, 430, and 440 to provide periodic updates to the ECMP members (e.g., Figure 1 114) of the pipelines 410, 420, 430, and 440, as discussed above with respect to Figure 1 The load-aware desired path of the present disclosure is set in advance by the state machine 450 rather than selecting a path packet by packet. In addition, the disclosed solution should account for a minimum programmed inactive period on the path before updating a given path to a new path to avoid unwanted reordering of packets. These features are among the differentiators for implementing a centralized and aggregated implementation of the entire chip of the present technology (without any additional cost). These differentiators also support other advantageous features described above, such as a small fraction of the area, power, latency, and cost of using any other per-pipeline dynamic load balancing solution, or being faster and more advantageous compared to traditional S / W implementations due to the use of the state machine 450.
[0030] Figure 5 FIG. 500 is a flowchart illustrating an example of an S / W process 500 according to some aspects of the present technology. The S / W process 500 begins at operation block 502, where the S / W programs the lower and / or upper group tables (e.g., Figure 1 122 and 112) in the pipeline database and the members of each group (e.g.,Figure 1 of 124 and 114). At operation box 504, the S / W programs the lower and / or upper group tables in the central module database and the members of each group, the port-to-NH mapping, deviation, weight, and / or vector, and threshold values of the group. At operation box 506, the S / W continues to update the central module (e.g., Figure 1 of 140) the deviation, weight, and / or vector of the path in. At this time, the control of the S / W process 500 is passed to operation box 502 to continue the process.
[0031] Figure 6 is a flowchart illustrating an example of a central module process 600 according to some aspects of the present technology. The central module process 600 starts at operation box 602, where the central module receives real-time updates and processes them to calculate a quality metric. At operation box 604, the central module goes through each entry into level 1 group 1, and at control operation box 606, the central module checks whether the upper NH and the corresponding lower group are still favorable. If the answer is yes, then at control operation box 608, the central module continues to check whether the lower NH and the second and third choices are still favorable. If the answer checked in control operation box 606 is no, then at operation box 612, the desired choice is changed, and the control is passed to operation box 610. If at control operation box 608, the lower NH and the second and third choices are favorable, then the control is passed to operation box 610, where an automatic update for all pipelines is initiated.
[0032] Figure 7 is a flowchart illustrating an example of a pipeline process 700 according to some aspects of the present technology. The pipeline process starts at operation box 702, where the pipeline processor performs group and member lookups and derives destinations, as typically done in ECMP. At operation box 704, the pipeline processor sets the hit bit corresponding to the reference entry, and at operation box 706, the second and third options are selected in case of a port failure. At control operation box 708, the pipeline processor checks the corresponding hit bit when updated from the central state machine. If the hit bit is equal to zero, then at operation box 710, the update from the central state machine is accepted and the control is passed to operation box 702. If the hit bit is equal to 1, then at operation box 712, the update from the central state machine is ignored and the control is passed to operation box 702.
[0033] Figure 8An electronic system 800 in which some aspects of the present technology are implemented. The electronic system 800 can be a network switch of a data center or an enterprise network and / or a part of the network switch. The electronic system 800 can include various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 800 includes a bus 808, one or more processing units 812, a system memory 804 (and / or buffer), a ROM 810, a permanent storage device 802, an input device interface 814, an output device interface 806, and one or more network interfaces 816 or subsets and variants thereof.
[0034] The bus 808 collectively represents all system, peripheral, and chipset buses that communicatively connect numerous internal devices of the electronic system 800. In one or more embodiments, the bus 808 communicatively connects one or more processing units 812 with the ROM 810, the system memory 804, and the permanent storage device 802. One or more processing units 812 retrieve instructions to be executed and data to be processed from these various memory units in order to execute the processes of the present disclosure. In different embodiments, one or more processing units 812 can be a single processor or a multi-core processor. In one or more aspects, one or more processing units 812 can be used to implement Figure 5 , 6 and / or the processes of 7.
[0035] The ROM 810 stores static data and instructions required by one or more processing units 812 of the electronic system 800 and other modules. On the other hand, the permanent storage device 802 can be a read-write memory device. The permanent storage device 802 can be a non-volatile memory unit that stores instructions and data even when the electronic system 800 is turned off. In one or more embodiments, a mass storage device (such as a magnetic disk or an optical disk and its corresponding disk drive) can be used as the permanent storage device 802.
[0036] In one or more embodiments, a removable storage device (such as a floppy disk, a flash drive, and its corresponding disk drive) can be used as the permanent storage device 802. Like the permanent storage device 802, the system memory 804 can be a read-write memory device. However, different from the permanent storage device 802, the system memory 804 can be a volatile read-write memory, such as random access memory (RAM). The system memory 804 can store any of the instructions and data that one or more processing units 812 may need during operation. In one or more embodiments, the processes of the present disclosure are stored in the system memory 804, the permanent storage device 802, and / or the ROM 810. One or more processing units 812 retrieve instructions to be executed and data to be processed from these various memory units in order to execute the processes of one or more embodiments.
[0037] The bus 808 is also connected to input and output device interfaces 814 and 806. The input device interface 814 enables a user to convey information and selection commands to the electronic system 800. Input devices that may be used with the input device interface 814 may include, for example, an alphanumeric keyboard and a pointing device (also referred to as a "cursor control device"). The output device interface 806 may implement, for example, the display of images generated by the electronic system 800. Output devices that may be used with the output device interface 806 may include, for example, a printer and a display device, such as a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a flexible display, a flat panel display, a solid state display, a projector, or any other device for outputting information. One or more embodiments may include a device that functions as both an input device and an output device, such as a touch screen. In these embodiments, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input may be received from the user in any form, including acoustic, voice, or tactile input.
[0038] Finally, as Figure 8 shown in, the bus 808 also couples the electronic system 800 to one or more networks and / or one or more network nodes via one or more network interfaces 816. In this way, the electronic system 800 can be part of a computer network (such as a local area network (LAN), a wide area network (WAN), an intranet, or a network of networks, such as the Internet). Any or all components of the electronic system 800 may be used in conjunction with the present disclosure.
[0039] Embodiments within the scope of the present disclosure may be implemented in part or in whole using a tangible computer-readable storage medium (or multiple tangible computer-readable storage media of one or more types) encoding one or more instructions. The nature of the tangible computer-readable storage medium may also be non-transitory.
[0040] A computer-readable storage medium may be any storage medium that can be read, written to, or otherwise accessed by a general or special purpose computing device, including any processing electronics and / or processing circuitry capable of executing instructions. By way of example, and not limitation, a computer-readable medium may include any volatile semiconductor memory, such as RAM, DRAM, SRAM, T-RAM, Z-RAM, and TTRAM. A computer-readable medium may also include any non-volatile semiconductor memory, such as ROM, PROM, EPROM, EEPROM, NVRAM, flash, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, racetrack memory, FJG, and Millipede memory.
[0041] In addition, a computer-readable storage medium may include any non-semiconductor memory, such as an optical disc storage device, a magnetic disk storage device, magnetic tape, other magnetic storage devices, or any other medium capable of storing one or more instructions. In one or more embodiments, a tangible computer-readable storage medium may be directly coupled to a computing device, while in other embodiments, a tangible computer-readable storage medium may be indirectly coupled to a computing device, e.g., via one or more wired connections, one or more wireless connections, or any combination thereof.
[0042] Instructions may be directly executable or may be used to develop executable instructions. For example, instructions may be implemented as executable or non-executable machine code or as instructions in a high-level language that can be compiled to generate executable or non-executable machine code. In addition, instructions may also be implemented as data or may include data. Computer-executable instructions may also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, etc. As those skilled in the art will recognize, details including (but not limited to) the number, structure, sequence, and organization of instructions may vary significantly without changing the underlying logic, functionality, processing, and output.
[0043] Although the foregoing discussion mainly relates to a microprocessor or multi-core processor that executes software, one or more embodiments are executed by one or more integrated circuits such as an ASIC or FPGA. In one or more embodiments, this integrated circuit executes instructions stored on the circuit itself.
[0044] Those skilled in the art will understand that the various illustrative blocks, modules, elements, components, methods, and algorithms described herein may be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, the various illustrative blocks, modules, elements, components, methods, and algorithms have been generally described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each particular application. The various components and blocks may be arranged differently (e.g., in a different order or divided in a different manner), all of which do not depart from the scope of the present technology.
[0045] It should be understood that any particular order or hierarchy of blocks in the disclosed processes is an illustration of example methods. Based on design preferences, it should be understood that the particular order or hierarchy of blocks in a process may be rearranged or that all illustrated blocks of the process may be executed. Any one of the blocks may be executed simultaneously. In one or more embodiments, multitasking and parallel processing may be advantageous. In addition, the separation of the various system components in the embodiments described above should not be understood as required in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0046] As used in the specification and any claims of this application, the terms "base station", "receiver", "computer", "server", "processor", and "memory" all refer to electronic or other technical devices. These terms exclude a person or group of people. For the purposes of the specification, the term "display / displaying" means displaying on an electronic device.
[0047] As used herein, the phrase "at least one" before a series of items (where the terms "and" or "or" separate any of the items in the series) modifies the entire list, rather than each member of the list (i.e., each item). The phrase "at least one of..." does not require selection of each of each item listed; rather, the phrase allows the meaning of including at least one of any of the items and / or at least one of any combination of the items and / or at least one of each of the items. By way of example, the phrase "at least one of A, B, and C" or "at least one of A, B, or C" each refers to only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.
[0048] The predicates "configured to", "operable to", and "programmed to" do not imply any particular tangible or intangible modification of the subject, but are intended to be used interchangeably. In one or more embodiments, a processor configured to monitor and control an operation or component also means that the processor is programmed to monitor and control the operation or that the processor is operable to monitor and control the operation. Similarly, a processor configured to execute code can be constructed as a processor programmed to execute code or operable to execute code.
[0049] Phrases such as "in one aspect", "the aspect", "in another aspect", "some aspects", "one or more aspects", "one embodiment", "the embodiment", "another embodiment", "some embodiments", "one or more embodiments", "one example", "the example", "another example", "some examples", "one or more examples", "one configuration", "the configuration", "another configuration", "some configurations", "one or more configurations", "the present technology", "the disclosure / the present disclosure", and other variations and the like are for convenience and do not imply that the disclosure associated with such phrases is necessary for the present technology or that such disclosure applies to all configurations of the present technology. The disclosure associated with such phrases may apply to all configurations or one or more configurations. The disclosure associated with such phrases may provide one or more examples. For example, the phrase "in one aspect" or "some aspects" may refer to one or more aspects, and vice versa, and this similarly applies to the other aforementioned phrases.
[0050] As used herein, the term "exemplary" is used to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or superior to other embodiments. Additionally, to the extent that the terms "comprising," "having," or the like are used in the description or claims, such terms are intended to be inclusive in a manner similar to the term "including" as that term is interpreted when used as a transitional word in a claim.
[0051] All structures and functions of elements throughout various aspects of the present disclosure that are equivalent to those of elements known or later to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Additionally, nothing disclosed herein is intended to be dedicated to the public, whether or not the disclosure is explicitly recited in the claims. No element is to be construed under 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase "means for" or, in the case of a method claim, the element is recited using the phrase "step for."
[0052] The foregoing description is provided to enable those of ordinary skill in the art to practice the various aspects described herein. Those of ordinary skill in the art will readily appreciate various modifications to these aspects, and the general principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein but are to be accorded the full scope consistent with the language of the claims, where reference to an element in the singular is not intended to mean "one and only one" unless explicitly so stated but rather means "one or more." Unless otherwise explicitly stated, "some" means one or more. Masculine pronouns (e.g., "his") include feminine and neuter (e.g., "her" and "its"), and vice versa. Headings and subheadings (if any) are used for convenience only and do not limit the disclosure.
Claims
1. A semiconductor chip for implementing load-aware Equal-Cost Multi-Path (ECMP), the semiconductor chip comprising: a plurality of ports; a plurality of pipelines, each pipeline coupled to a portion of the plurality of ports; and a central unit including a state machine and a plurality of databases, wherein: the plurality of databases are configured to contain information about a communication network including an upper network and a lower network, and the state machine is configured to determine at least one characteristic of corresponding path groups within the upper network and the lower network to overcome dynamic network conditions, wherein the dynamic network conditions include network congestion, and wherein the state machine is configured to check the upper network and lower network group tables and member tables entry by entry to determine a set of desired ones of the plurality of ports for the path groups to reduce the network congestion, and wherein the state machine is further configured to impose a minimum inactivity period on entries in the member table by: setting a hit bit maintained for each entry in the member table of the lower network to 1 when the entry is referenced by a packet, checking the hit bit during periodic state machine updates to the member table and updating the entry when the value of the hit bit is 0, and automatically imposing a programmed minimum activity period for minimum inactivity checking when the update rate of the state machine for each entry corresponds to the programmed minimum activity period.
2. The semiconductor chip according to claim 1, wherein the information includes the nature of a plurality of paths between a first host and a second host.
3. The semiconductor chip according to claim 1, wherein the information includes next-hop (NH) to port mapping information for the upper network and the lower network.
4. The semiconductor chip according to claim 3, wherein the information includes upper NH to lower group mapping information.
5. The semiconductor chip according to claim 1, wherein the information includes dynamic port status and deviation, weight, and / or vector for paths beyond the current hop.
6. The semiconductor chip according to claim 1, wherein the information includes thresholds, instantaneous load, and congestion information for the plurality of ports and queues per port.
7. The semiconductor chip according to claim 1, wherein the state machine is configured to process load, deviation, and port status information.
8. The semiconductor chip according to claim 1, wherein the state machine is configured to evaluate the set of desired ones of the plurality of ports for the path groups based on at least one of: local port available / unavailable status, local instantaneous load and congestion and corresponding queues of the plurality of ports, deviation, weight, and / or vector on the path beyond the current hop, and thresholds.
9. A method for implementing load-aware ECMP routing, the method comprising: configuring a plurality of databases to store information about a communication network including an upper network and a lower network; and configuring a hardware-implemented state machine to determine at least one characteristic of corresponding path groups within the upper network and the lower network to overcome dynamic network conditions; implementing the plurality of databases and the state machine on a semiconductor chip including a plurality of pipelines and a plurality of ports; and Couple each of the plurality of pipelines to portions of the plurality of ports on the semiconductor chip, wherein the dynamic network condition includes network congestion, and wherein the method further comprises configuring the state machine to check the upper network and the lower network group and member tables entry by entry and evaluate a set of desired ones of the plurality of ports of the path group to reduce the network congestion, and wherein the method further comprises configuring the state machine to impose a minimum inactivity period on a given entry of the member table to avoid unwanted reordering of packets by: setting the hit bit maintained for each entry of the member table of the lower network to 1 when the entry is referenced by a packet, checking the hit bit during periodic state machine updates to the member table and updating the entry when the value of the hit bit is 1, and automatically imposing a programmed minimum activity period for minimum inactivity checking when the update rate of the state machine for each entry corresponds to the programmed minimum activity period.
10. The method of claim 9, further comprising configuring the state machine to pre-set load-aware desired paths and handle load, deviation, and port status information.
11. The method of claim 9, further comprising configuring the state machine to evaluate the set of desired ones of the plurality of ports of the path group based on at least one of: local port available / unavailable status, local instantaneous load and congestion of the plurality of ports and corresponding queues, deviation, weight, and / or vector on the path beyond the current hop, and thresholds.
12. The method of claim 9, wherein the information includes: the nature of a plurality of paths between a first host and a second host of the network, the NH to port mapping of the upper network and the lower network, the upper NH to lower group mapping, the dynamic port status and deviation, weight, and / or vector of the path beyond the current hop, and the thresholds, instantaneous load, and congestion information of the plurality of ports and the queue per port.
Citation Information
Patent Citations
A method for dynamic load balancing of network flows on lag interfaces
CN104756451A
Traffic management for high-bandwidth switching
CN110417670A