Method and equipment for realizing equivalent multipath group load sharing
By calculating the bandwidth weights of ECMP group member paths and writing them sequentially into ECMP entries, the problem of traffic imbalance caused by the ECMP group hash algorithm is solved, achieving a more balanced traffic distribution and improved link utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-07
AI Technical Summary
In the Spine-Leaf network architecture, the existing ECMP group's hash algorithm leads to uneven traffic distribution. Especially in scenarios with complex topologies and large differences in link bandwidth, traffic is easily concentrated on certain links, causing congestion and underutilizing the bandwidth of other links.
By calculating the bandwidth weights of ECMP group member paths and writing them into ECMP entries in turn, traffic is dynamically allocated to match link quality, avoiding hash polarization, improving link utilization, and reducing hotspot congestion.
It enables a more balanced distribution and forwarding of traffic based on link quality, improves link utilization, reduces hotspot congestion, and has the advantages of being simple to implement and easy to upgrade in the existing network.
Smart Images

Figure CN121814673A_ABST
Abstract
Description
Technical Field
[0001] This application relates to communication technology, specifically a method and device for implementing load sharing of equal-cost multipath groups. Background Technology
[0002] ECMP (Equal-Cost Multi-Path) is a routing strategy that load-balances traffic forwarding across multiple paths with equal costs to a destination IP address, thereby improving link utilization, increasing bandwidth, and enhancing network redundancy and reliability.
[0003] When a communication device encounters a failure of an equivalent next hop in the ECMP group leading to a destination IP address, it modifies its hash algorithm based on the number of working equivalent next hops in the ECMP group. This reloads traffic destined for the destination IP address onto different ECMP members, resulting in different equivalent paths. However, this can cause traffic load-sharing on the equivalent path containing the unaffected equivalent next hop to be switched to other equivalent paths. To address this, the communication device determines the equivalent next hops for each equivalent multipath and sequentially fills them into the ECMP table of the ECMP group until it is full. For example, if the communication device's routing learning indicates that the ECMP group leading to the destination IP address contains 6 equivalent next hops, these 6 equivalent next hops are sequentially filled into the 32 entries of the ECMP table until it is full. When one of these equivalent next hops fails, the other normal equivalent next hops sequentially replace the failed next hop's ECMP entry, preventing traffic from normal ECMP members from being switched to other ECMP members.
[0004] In existing Spine-Leaf network architectures, the hash value calculated based on packet characteristic parameters of service flows for Leaf nodes is used to select the next hop in the ECMP group. However, there is a lack of a dynamic allocation mechanism that considers the actual available bandwidth of the links. This approach can easily lead to uneven traffic distribution in scenarios with complex topologies and significant differences in link bandwidth. For example, when multiple Leaf nodes communicate with the same destination Leaf node simultaneously, the hash algorithm may concentrate most of the traffic on the same Spine node and link, causing congestion on that link while other link bandwidth resources are not fully utilized. Summary of the Invention
[0005] The purpose of this application is to provide a method and device for implementing load balancing of equal-cost multipath groups, which load balances traffic forwarded through the equivalent next hop of the ECMP group based on bandwidth weight calculated by link quality.
[0006] To achieve the above objectives, this application provides a method for implementing load balancing in an equal-cost multipath group. The method includes: calculating the bandwidth weight of each member path in the ECMP group; calculating the number of ECMP entries allocated to each member path based on the bandwidth weights of all member paths; and writing the equivalent next hop of each member path into the ECMP table of the ECMP group in turn until the number of allocated ECMP entries for each member path is reached.
[0007] To achieve the above objectives, this application also provides an apparatus for implementing load balancing in an equal-cost multipath group. The apparatus includes a processor, a machine-readable storage medium, a switching chip, and a network interface. The processor executes machine-executable instructions recorded on the machine-readable storage medium to perform the following operations: calculate the bandwidth weight of each member path in the ECMP group; calculate the number of ECMP entries allocated to each member path based on the bandwidth weights of all member paths; and write the equivalent next hop of each member path into the ECMP table of the ECMP group in turn until the number of allocated ECMP entries for each member path is reached.
[0008] The beneficial effect of this application is that it dynamically allocates traffic based on the weight of the actual available bandwidth of the link, so that the traffic can be distributed and forwarded more evenly according to the link quality bandwidth, avoiding hash polarization. This not only improves the link utilization and reduces hot spot congestion, but also has the advantages of being simple to implement and easy to upgrade and deploy in the existing network. Attached Figure Description
[0009] Figure 1 This application provides an embodiment of a method for implementing load balancing of equivalent multipath groups; Figure 2 A schematic diagram of the network architecture for calculating the bandwidth weight of each member path of an ECMP group, provided for this application; Figure 3 A schematic diagram of an embodiment of the allocation of ECMP entries provided for this application; Figure 4 A schematic diagram of Embodiment 2 for allocating ECMP entries provided for this application; Figure 5 A schematic diagram of Embodiment 3 for allocating ECMP entries provided for this application; Figure 6 This is a schematic diagram of an embodiment of a device for implementing load sharing of equivalent multipath groups provided in this application. Detailed Implementation
[0010] The following detailed description will be provided with reference to several examples illustrated in the accompanying figures. In this detailed description, numerous specific details are used to provide a comprehensive understanding of the present application. Known methods, steps, components, and circuits are not described in detail in the examples to avoid obscuring their meaning.
[0011] In the terminology used, the term "including" means including but not limited to; the term "containing" means including but not limited to; the terms "above," "within," and "below" include the number itself; the terms "greater than" and "less than" mean not including the number itself. The term "based on" means based on at least a portion of them.
[0012] Figure 1 The illustration shows an embodiment of a method for implementing load balancing across equal-cost multipath groups. This embodiment includes: Step 101: Calculate the bandwidth weight of each member path in the ECMP group; Step 102: Based on the bandwidth weights of all member paths, calculate the number of ECMP entries allocated to each member path. Step 103: Write the equivalent next hop of each member path into the ECMP table of the ECMP group in turn until the number of assigned ECMP entries for each member path is reached.
[0013] Figure 1 The beneficial effect of the embodiment is that the dynamic allocation based on the weight of the actual available bandwidth of the link enables the traffic to be distributed and forwarded more evenly according to the link quality bandwidth, avoiding hash polarization. This not only improves the link utilization and reduces hotspot congestion, but also has the advantages of being simple to implement and easy to upgrade and deploy on the existing network.
[0014] Figure 2 A schematic diagram of the network architecture for calculating the bandwidth weight of each member path of an ECMP group, provided for this application; Figure 2 In the Spine-Leaf architecture of the data center, Spine nodes 21-24 and Leaf nodes 11-14 are both Layer 3 switching devices. Through routing protocols such as OSPF or EBGP, equal-cost multipath load balancing and link backup are achieved between Spine devices and Leaf devices.
[0015] Leaf nodes 11-14 and Spine nodes 21-24 can measure link quality within the Spine-Leaf architecture using in-band telemetry or the Two-Way Active Measurement Protocol (TWAMP). Taking the measurement of ECMP link quality from Leaf node 11 to Leaf node 14 as an example...
[0016] The next hops of the four ECMP member paths from Leaf node 11 to Leaf node 14 are Spine nodes 21-24. Leaf node 11 measures the link quality of the direct links L1.1, L1.2, L1.3, and L1.4 of the four ECMP member paths, and measures the link quality of the direct links LS1.4, LS2.4, LS3.4, and LS4.4 of the equivalent next hop Spine nodes 21, 22, 23, and 24 on the four ECMP member paths.
[0017] In this application, Leaf node 11 can choose the measured link quality of direct links L1.1, L1.2, L1.3, and L1.4 as the bandwidth weights QA, QB, QC, and QD of the four member paths, respectively; or, it can choose the link quality of the direct links LS1.4, LS2.4, LS3.4, and LS4.4 of the four equivalent next-hop Spine nodes 21, 22, 23, and 24 as the bandwidth weights QA, QB, QC, and QD of the four member paths, respectively; or Leaf node 11 compares the link quality of link L1.1 and link LS1.4, and selects the highest link quality among the two links as the bandwidth weight QA of the first member path based on the link quality priority principle; Leaf node 11 selects the bandwidth weights QB, QC, and QD of the other three member paths in the same way, based on the link quality priority principle.
[0018] Other Leaf nodes calculate the bandwidth weights of each member path in the ECMP group from this device to other Leaf nodes in the same way. In this embodiment, QA, QB, QC, and QD represent the calculated bandwidth weight values. The specific bandwidth weight values will be calculated based on the service network, and this application does not provide specific examples.
[0019] Figure 3 A schematic diagram of an embodiment of the allocation of ECMP entries provided in this application.
[0020] Leaf node 11 calculates the bandwidth weight ratio among all member paths of the ECMP group reaching Leaf node 14 as QA:QB:QC:QD, and assigns the same number of ECMP entries according to the ratio of the bandwidth weight of each member path in the proportional relationship.
[0021] In one example, Leaf node 11 calculates the ratio QA:QB:QC:QD=2:4:5:7, then the number of ECMP entries allocated to each member path is 2, 4, 5, and 7.
[0022] Leaf node 11 writes the equivalent next hops S1, S2, S3, and S4 of each member path in turn in sequence. Figure 3The ECMP table 30 for this ECMP group is shown.
[0023] When the number of ECMP entries for the first member path written to ECMP table 30 reaches 2, the equivalent next hop S1 of that member path will no longer be written. Instead, the equivalent next hops S2, S3, and S4 of the other member paths will be written to ECMP table 30 in turn.
[0024] If Leaf node 11 determines that the number of ECMP entries for the second member path written to ECMP table 30 has reached 4, then it will no longer write the equivalent next hop S2 for that member path, and will write the equivalent next hops S3 and S4 for other member paths to ECMP table 30 in turn.
[0025] Leaf node 11 determines that the number of ECMP entries for the third member path written to ECMP table 30 has reached 5. If so, it stops writing the equivalent next hop S3 for that member path and continues writing the equivalent next hop S4 for the fourth member path to ECMP table 30 until the number of equivalent next hop S4 entries written to ECMP table 30 reaches 7.
[0026] Figure 4 A schematic diagram of an embodiment two for assigning ECMP entries to this application.
[0027] Figure 4 In the ECMP Table 40, the total number of entries is 32.
[0028] In one example, Leaf node 11 calculates the bandwidth weights of each member path of the ECMP group leading to Leaf node 14 as 2, 4, 4, and 6, respectively.
[0029] Leaf node 11 calculates the total bandwidth weight of all member paths: QA + QB + QC + QD = 2 + 4 + 4 + 6 = 16.
[0030] Leaf node 11 distributes the total number of entries 32 in ECMP table 40 equally according to the total bandwidth weight of the member paths, and calculates the number of entries allocated to each weight unit: .
[0031] Leaf node 11 modulo the total number of entries 32 in ECMP table 40 according to the total bandwidth weight 32 of all member paths: .
[0032] Leaf node 11 calculates that the number of unsplitting entries in ECMP table 40 is zero. Based on the bandwidth weight of each member path and the number of entries allocated to each weight unit, the number of ECMP entries for each member path is calculated.
[0033] Leaf node 11 calculates the number of ECMP entries for the first member path to be 4 based on the bandwidth weight 2 of the first member path and the number of entries allocated to each weight unit 2.
[0034] Leaf node 11 calculates the number of ECMP entries for the second member path to be 8 based on the bandwidth weight 2 of the second member path and the number of entries assigned to each weight unit 4.
[0035] Leaf node 11 calculates the number of ECMP entries for the third member path to be 8 based on the bandwidth weight 2 of the third member path and the number of entries allocated to each weight unit 4.
[0036] Leaf node 11 calculates the number of ECMP entries for the fourth member path to be 12 based on the bandwidth weight 2 of the fourth member path and the number of entries allocated to each weight unit 6.
[0037] Leaf node 11 writes the equivalent next hops S1, S2, S3, and S4 of each member path in turn in sequence. Figure 4 The ECMP table 40 for this ECMP group is shown.
[0038] When Leaf node 11 determines that the number of ECMP entries for the first member path written to ECMP table 40 reaches 4, it will no longer write the equivalent next hop S1 of that member path, and will write the equivalent next hops S2, S3, and S4 of other member paths to ECMP table 40 in turn.
[0039] Leaf node 11 determines that the number of ECMP entries for the second and third member paths written to ECMP table 30 has reached 8, then it will no longer write the equivalent next hops S2 and S3 for these two member paths, and will write the equivalent next hop S4 for the remaining fourth member path to ECMP table 40.
[0040] Figure 5 A schematic diagram of Embodiment 3 for assigning ECMP entries to this application.
[0041] Figure 5 In this context, ECMP Table 50 contains a total of 32 entries.
[0042] In one example, Leaf node 11 calculates the bandwidth weights of each member path of the ECMP group leading to Leaf node 14 as 2, 2, 3, and 3, respectively.
[0043] Leaf node 11 calculates the total bandwidth weight of all member paths: QA + QB + QC + QD = 2 + 2 + 3 + 3 = 10.
[0044] Leaf node 11 distributes the total number of entries 32 in ECMP table 50 equally according to the total bandwidth weight of the member paths, and calculates the number of entries allocated to each weight unit: .
[0045] Leaf node 11 modulo the total number of entries 32 in ECMP table 50 according to the total bandwidth weight 32 of all member paths: .
[0046] Leaf node 11 calculates that the number of unsplitting entries in ECMP table 50 is 3. Based on the bandwidth weight of each member path and the number of entries allocated to each weight unit, the number of ECMP entries for each member path is calculated.
[0047] Leaf node 11 calculates the number of ECMP entries for the first member path to be 6 based on the bandwidth weight 2 of the first member path and the number of entries assigned to each weight unit 3.
[0048] Leaf node 11 calculates the number of ECMP entries for the second member path to be 6 based on the bandwidth weight 2 of the second member path and the number of entries allocated to each weight unit 3.
[0049] Leaf node 11 calculates the number of ECMP entries for the third member path to be 9 based on the bandwidth weight of 3 for the third member path and the number of entries allocated to each weight unit of 3.
[0050] Leaf node 11 calculates the number of ECMP entries for the fourth member path to be 9 based on the bandwidth weight of the fourth member path (9) and the number of entries allocated to each weight unit (3).
[0051] Leaf node 11, following the link quality priority principle, allocates the two ECMP entries that cannot be evenly distributed to the fourth member path, which has the highest link quality. That is, the fourth member path has 11 ECMP entries.
[0052] Leaf node 11 writes the equivalent next hops S1, S2, S3, and S4 of each member path in turn in sequence. Figure 4 The ECMP table 50 for this ECMP group is shown.
[0053] When Leaf node 11 determines that the number of ECMP entries for the first and second member paths written to ECMP table 50 reaches 6, it will stop writing the equivalent next hop S1 of that member path and write the equivalent next hops S3 and S4 of other member paths to ECMP table 50 in turn.
[0054] Leaf node 11 determines that if the number of ECMP entries written to the third member path in ECMP table 30 reaches 9, then it will no longer write the equivalent next hop S3 of the third member path; then Leaf node 11 determines to write the remaining 2 ECMP entries to the equivalent next hop S4 of the fourth member path.
[0055] Figure 6 This is a schematic diagram of an embodiment of a device for implementing equivalent multipath group load balancing provided in this application. Device 60 includes a processor 61, a machine-readable storage medium 62, a switching chip 63, and a network interface 64.
[0056] The processor 61 performs the following operations by executing machine-executable instructions recorded in the machine-readable storage medium 62: calculating the bandwidth weight of each member path in the ECMP group; calculating the number of ECMP entries allocated to each member path based on the bandwidth weights of all member paths; and writing the equivalent next hop of each member path into the ECMP table of the ECMP group in turn until the number of allocated ECMP entries for each member path is reached.
[0057] The processor 61 executes machine-executable instructions recorded in the machine-readable storage medium 62 to perform the operation of calculating the bandwidth weight of each member path in the ECMP group, including: calculating the link quality of the direct links of each member path as the bandwidth weight of each member path; or, calculating the link quality of the direct links of the equivalent next hop of each member path as the bandwidth weight of each member path; or, calculating the link quality of the direct links of each member path and the link quality of the direct links of the equivalent next hop of each member path, and selecting the highest link quality as the bandwidth weight of each member path based on the link quality priority principle.
[0058] The processor 61 executes machine-executable instructions recorded in the machine-readable storage medium 62 to perform the operation of calculating the number of ECMP entries allocated to each member path based on the bandwidth weights of all member paths, including: calculating the proportional relationship of bandwidth weights among all member paths; and allocating the same number of ECMP entries according to the ratio of the bandwidth weights of each member path in the proportional relationship.
[0059] The processor 61 executes machine-executable instructions recorded in the machine-readable storage medium 62 to perform the operation of calculating the number of ECMP entries allocated to each member path based on the bandwidth weights of all member paths, including: dividing the total number of ECMP entries into the total number of entries of the ECMP table according to the total bandwidth weights of all member paths, and calculating the number of entries allocated to each weight unit; taking the remainder of the total number of ECMP entries into the total number of entries of the ECMP table according to the total bandwidth weights of all member paths, and calculating the number of entries that cannot be evenly divided to be zero; and calculating the number of ECMP entries for each member path according to the bandwidth weight of each member path and the number of entries allocated to each weight unit.
[0060] The processor 61 executes machine-executable instructions recorded in the machine-readable storage medium 62 to perform the operation of calculating the number of ECMP entries allocated to each member path based on the bandwidth weights of all member paths. This includes: dividing the total number of ECMP entries into the total bandwidth weights of all member paths equally, and calculating the number of entries allocated to each weight unit; taking the remainder of the total number of ECMP entries into the total bandwidth weights of all member paths, and calculating the number of entries that cannot be evenly distributed; calculating the number of ECMP entries evenly distributed to each member path according to the bandwidth weight of each member path and the number of entries allocated to each weight unit; and allocating the number of entries that cannot be evenly distributed to the member path with the highest link quality according to the link quality priority principle.
[0061] In this disclosure, a machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device used to store or contain information (such as executable instructions, data, etc.). For example, any machine-readable storage medium herein can be any type of random access memory (RAM), volatile memory, non-volatile memory, flash memory, storage drive (such as a hard disk drive), solid-state drive, any type of optical disc (such as an optical disc, DVD, etc.), and similar devices, or combinations thereof. Furthermore, any machine-readable storage medium described herein can be a non-transitory machine-readable storage medium.
[0062] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for implementing load sharing in an equivalent multipath group, characterized in that, The method includes, Calculate the bandwidth weight for each member path in the ECMP group; Based on the bandwidth weights of all the member paths, calculate the number of ECMP entries allocated to each member path; The equivalent next hop for each member path is written sequentially and in turn into the ECMP table of the ECMP group until the number of assigned ECMP entries for each member path is reached.
2. The method according to claim 1, characterized in that, The bandwidth weight for each member path in the ECMP group is calculated as follows: Calculate the link quality of the directly connected links for each member path as the bandwidth weight for each member path; or, Calculate the link quality of the directly connected link with the equivalent next hop for each member path as the bandwidth weight for each member path; or, Calculate the link quality of the direct link of each member path and the link quality of the direct link of the equivalent next hop of each member path. Based on the link quality priority principle, select the highest link quality as the bandwidth weight of each member path.
3. The method according to claim 1, characterized in that, Based on the bandwidth weights of all the member paths, the number of ECMP entries allocated to each member path is calculated, including: Calculate the proportional relationship of bandwidth weights among all the aforementioned member paths; The same number of ECMP entries are assigned according to the ratio of the bandwidth weight of each member path in the proportional relationship.
4. The method according to claim 1, characterized in that, Based on the bandwidth weights of all the member paths, the number of ECMP entries allocated to each member path is calculated, including: The total number of entries in the ECMP table is divided equally according to the total bandwidth weight of all the member paths, and the number of entries allocated to each weight unit is calculated. The total number of entries in the ECMP table is moduloed by the total bandwidth weight of all member paths, and the number of entries that cannot be evenly distributed is zero. The number of ECMP entries for each member path is calculated based on the bandwidth weight of each member path and the number of entries allocated to each weight unit.
5. The method according to claim 1, characterized in that, Based on the bandwidth weights of all the member paths, the number of ECMP entries allocated to each member path is calculated, including: The total number of entries in the ECMP table is divided equally according to the total bandwidth weight of all the member paths, and the number of entries allocated to each weight unit is calculated. The number of entries that cannot be evenly distributed is calculated by taking the remainder of the total bandwidth weight of all the member paths of the ECMP table. Calculate the average number of ECMP entries allocated to each member path based on the bandwidth weight of each member path and the number of entries allocated to each weight unit. Based on the link quality priority principle, the number of entries that cannot be evenly distributed is allocated to the member path with the highest link quality.
6. A device for implementing load balancing across equal-cost multipath groups, the device comprising a processor, a machine-readable storage medium, a switching chip, and a network interface; characterized in that, The processor performs the following operations by executing machine-executable instructions recorded on the machine-readable storage medium. Calculate the bandwidth weight for each member path in the ECMP group; Based on the bandwidth weights of all the member paths, calculate the number of ECMP entries allocated to each member path; The equivalent next hop for each member path is written sequentially and in turn into the ECMP table of the ECMP group until the number of assigned ECMP entries for each member path is reached.
7. The device according to claim 6, characterized in that, The processor performs the operation of calculating the bandwidth weight of each member path of the ECMP group by executing machine-executable instructions recorded on the machine-readable storage medium, including: Calculate the link quality of the directly connected links for each member path as the bandwidth weight for each member path; or, Calculate the link quality of the directly connected link with the equivalent next hop for each member path as the bandwidth weight for each member path; or, Calculate the link quality of the direct link of each member path and the link quality of the direct link of the equivalent next hop of each member path. Based on the link quality priority principle, select the highest link quality as the bandwidth weight of each member path.
8. The device according to claim 6, characterized in that, The processor performs the operation of calculating the number of ECMP entries allocated to each of the member paths based on the bandwidth weights of all the member paths by executing machine-executable instructions recorded on the machine-readable storage medium. Calculate the proportional relationship of bandwidth weights among all the aforementioned member paths; The same number of ECMP entries are assigned according to the ratio of the bandwidth weight of each member path in the proportional relationship.
9. The device according to claim 6, characterized in that, The processor performs the operation of calculating the number of ECMP entries allocated to each of the member paths based on the bandwidth weights of all the member paths by executing machine-executable instructions recorded on the machine-readable storage medium. The total number of entries in the ECMP table is divided equally according to the total bandwidth weight of all the member paths, and the number of entries allocated to each weight unit is calculated. The total number of entries in the ECMP table is moduloed by the total bandwidth weight of all member paths, and the number of entries that cannot be evenly distributed is zero. The number of ECMP entries for each member path is calculated based on the bandwidth weight of each member path and the number of entries allocated to each weight unit.
10. The device according to claim 6, characterized in that, The processor performs the operation of calculating the number of ECMP entries allocated to each of the member paths based on the bandwidth weights of all the member paths by executing machine-executable instructions recorded on the machine-readable storage medium. The total number of entries in the ECMP table is divided equally according to the total bandwidth weight of all the member paths, and the number of entries allocated to each weight unit is calculated. The number of entries that cannot be evenly distributed is calculated by taking the remainder of the total bandwidth weight of all the member paths of the ECMP table. Calculate the average number of ECMP entries allocated to each member path based on the bandwidth weight of each member path and the number of entries allocated to each weight unit. Based on the link quality priority principle, the number of entries that cannot be evenly distributed is allocated to the member path with the highest link quality.