A path fault processing method and apparatus, and related device

CN122226588APending Publication Date: 2026-06-16NEW H3C TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610379765.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-26
Publication Date
2026-06-16

Smart Images

  • Figure CN122226588A_ABST
    Figure CN122226588A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of network communication, in particular to a path fault processing method and device and related equipment. The method is applied to a network device, and a priority path link list of a plurality of path sets ordered from high to low priority is maintained on the network device for each destination network prefix; the currently highest priority available path set in a priority path link list is a current primary path set corresponding to the destination network prefix; the method comprises the following steps: monitoring the connectivity state of each path included in the current primary path set corresponding to the target destination network prefix in real time; if it is determined based on the monitoring result that each path included in the current primary path set is invalid, marking the current path set as an unavailable path set; determining the currently highest priority available path set from the target priority path link list, and taking the newly determined currently highest priority available path set as the current primary path set corresponding to the target destination network prefix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network communication technology, and in particular to a path fault handling method, apparatus and related equipment. Background Technology

[0002] Modern data communication networks widely employ dynamic routing protocols (such as RIP, OSPF, IS-IS, BGP, etc.) to automatically discover and select paths. Under normal circumstances, these protocols select an optimal route for forwarding to each destination network, while keeping other candidate paths as inactive backups. When the primary path fails, the router needs to select an alternative path to maintain communication.

[0003] However, the traditional "one primary, one backup" failover mechanism has significant shortcomings: Slow convergence speed: When there is only a single backup path, if the primary path fails, the router often needs to wait for the routing protocol to recalculate or periodically update in order to obtain a backup path, resulting in long service interruption time.

[0004] Low resource utilization: There are often multiple suboptimal paths in the network, but traditional mechanisms only select one as a backup, leaving the rest idle, which fails to make full use of network resources.

[0005] Weak resistance to multiple failures: With only one layer of backup, when the primary path and the only backup path fail one after another, the entire route convergence process must be carried out again, making it difficult to guarantee business continuity.

[0006] Therefore, there is an urgent need for a general routing enhancement mechanism that can support multi-level backup, make full use of network redundant paths, and achieve fast and smooth convergence. Summary of the Invention

[0007] This application provides a path fault handling method, apparatus, and related equipment.

[0008] Firstly, this application provides a path failure handling method applied to a network device, wherein the network device maintains a corresponding priority path list for each destination network prefix, wherein the target priority path list corresponding to the target destination network prefix includes several path sets sorted from high to low priority, and the set of available paths with the highest current priority in the target priority path list is the set of currently used paths corresponding to the target destination network prefix; the method includes: Real-time monitoring of the connectivity status of each path in the current primary path set corresponding to the target network prefix; If, based on monitoring results, it is determined that all paths in the current primary path set are invalid, the current path set will be marked as an unavailable path set. The set of available paths with the highest current priority is determined from the target priority path list, and the newly determined set of available paths with the highest current priority is used as the set of current primary paths corresponding to the target destination network prefix.

[0009] Optionally, the method further includes: Monitor the connectivity status of each path in the target priority path list, which includes a set of paths sorted from high to low priority. If it is determined that all paths in any set of paths are invalid, then that set of paths is marked as an unavailable set of paths.

[0010] Optionally, the steps for monitoring the connectivity of each path include: Millisecond-level fault detection is achieved by establishing a bidirectional forwarding detection (BFD) session for the next hop of each path.

[0011] Optionally, the method further includes: For each destination network prefix, the metric or administrative distance of each candidate path corresponding to that destination network prefix is ​​calculated according to a preset routing protocol; Multiple paths with the same metric or administrative distance are grouped into a path set of the same priority level as equivalent paths. Paths with smaller metric or administrative distance belong to a path set with higher priority.

[0012] Optionally, if the current primary path set includes multiple equivalent paths, then the traffic of the target destination network prefix is ​​load-balanced by the multiple equivalent paths.

[0013] Optionally, for each path set, if a connectivity failure is detected in one of the paths included in the path set, the faulty path is removed from the path set. If the path set is the current primary path set, the traffic of the target destination network prefix is ​​load-balanced by the other equivalent paths included in the path set after removing the faulty path.

[0014] Optionally, the method further includes: Once a set of higher-priority paths (which have a higher priority than the current primary path set) is detected to be available again, a switchback delay timer is started. If the back-switch delay timer times out and the higher priority path set remains available, the higher priority path set is used as the current primary path set corresponding to the target destination network prefix, so as to switch the traffic of the target destination network prefix back to the higher priority path set.

[0015] Optionally, based on the connectivity status of each path in the target priority path linked list, which includes several path sets sorted from high to low priority, the available path set of the next priority level of the current primary path set is determined; and the available path set of the next priority level is set as the backup next hop of the current primary path set in the routing information base (RIB) or the forwarding information base (FIB).

[0016] Secondly, this application provides a path failure handling device applied to a network device, wherein the network device maintains a corresponding priority path list for each destination network prefix, wherein the target priority path list corresponding to the target destination network prefix includes several path sets sorted from high to low priority, and the set of available paths with the highest current priority in the target priority path list is the set of currently used paths corresponding to the target destination network prefix; the device includes: The monitoring unit is used to monitor in real time the connectivity status of each path in the current primary path set corresponding to the target network prefix; If the determining unit determines, based on the monitoring results, that all paths included in the current primary path set are invalid, it marks the current path set as an unavailable path set. The switching unit is used to determine the set of available paths with the highest current priority from the target priority path list, and to use the newly determined set of available paths with the highest current priority as the set of current primary paths corresponding to the target destination network prefix.

[0017] Optionally, the monitoring unit is used to monitor the connectivity status of each path in the set of paths sorted from high to low priority included in the target priority path list; If the determining unit determines that all paths in any path set are invalid, it marks the path set as an unavailable path set.

[0018] Optionally, when monitoring the connectivity of each path, the monitoring unit is specifically used for: Millisecond-level fault detection is achieved by establishing a bidirectional forwarding detection (BFD) session for the next hop of each path.

[0019] Optionally, the device further includes: The calculation unit is used to calculate the metric value or administrative distance of each candidate path corresponding to each destination network prefix according to a preset routing protocol. A partitioning unit is used to divide multiple paths with the same metric or administrative distance into a path set of the same priority level as equivalent paths. The path set to which the path with the smaller metric or administrative distance belongs has a higher priority.

[0020] Optionally, if the current primary path set includes multiple equivalent paths, then the traffic of the target destination network prefix is ​​load-balanced by the multiple equivalent paths.

[0021] Optionally, for each path set, if the monitoring unit detects a connectivity failure in one of the paths included in the path set, the faulty path is removed from the path set. If the path set is the current primary path set, the traffic of the target destination network prefix is ​​load-balanced by the other equivalent paths included in the path set after removing the faulty path.

[0022] Optionally, the device further includes: The back-switch unit is used to start a back-switch delay timer after detecting that a higher priority path set with a higher priority than the current primary path set has become available again; if the back-switch delay timer expires and the higher priority path set remains available, the higher priority path set is used as the current primary path set corresponding to the target destination network prefix, so as to switch the traffic of the target destination network prefix back to the higher priority path set.

[0023] Optionally, the device further includes: The setting unit is used to determine the available path set of the next priority level of the current primary path set based on the connectivity status of each path in the target priority path linked list, which includes several path sets sorted from high to low priority; and to set the available path set of the next priority level as the backup next hop of the current primary path set in the routing information base (RIB) or the forwarding information base (FIB).

[0024] Thirdly, embodiments of this application provide a path fault handling device, which includes: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the steps of the method as described in any one of the first aspects above, according to the obtained program instructions.

[0025] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the steps of the method as described in any of the first aspects above.

[0026] In summary, the path failure handling method provided in this application is applied to a network device. The network device maintains a corresponding priority path list for each destination network prefix. The target priority path list corresponding to the target destination network prefix includes several path sets sorted from highest to lowest priority. The set of available paths with the highest current priority in the target priority path list is the current primary path set corresponding to the target destination network prefix. The method includes: real-time monitoring of the connectivity status of each path included in the current primary path set corresponding to the target destination network prefix; if, based on the monitoring results, it is determined that all paths included in the current primary path set are invalid, marking the current path set as an unavailable path set; determining the set of available paths with the highest current priority from the target priority path list, and using the newly determined set of available paths with the highest current priority as the current primary path set corresponding to the target destination network prefix.

[0027] The path failure handling method provided in this application enables network devices to maintain a multi-level priority backup path list for each destination network prefix, and quickly switch between paths of different priorities when the primary path fails. It also supports equivalent multi-path forwarding at the same level to ensure the continuity of data forwarding. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings of the embodiments of this application.

[0029] Figure 1 A detailed flowchart of a path fault handling method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a path fault handling device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the hardware architecture of a path fault handling device provided in an embodiment of this application. Detailed Implementation

[0030] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “the,” and “the” as used in this application and claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations comprising one or more of the associated listed items.

[0031] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when," "when," or "in response to a determination."

[0032] Modern data communication networks widely employ dynamic routing protocols (such as RIP, OSPF, IS-IS, BGP, etc.) to automatically discover and select paths. Under normal circumstances, these protocols select an optimal route for each destination network for forwarding, while other candidate paths are used as inactive backups. When the primary path fails, the router needs to select an alternative path to maintain communication. However, in traditional implementations, only a single backup route is typically used for failover at any given time. For example, in a typical route backup configuration, an administrator can pre-configure a primary route and a secondary route as backups; when the primary link fails, the device disables the primary route and enables the backup route to forward data, switching back to the primary route upon recovery. While this "one primary, one backup" failover mechanism improves reliability, it suffers from insufficient flexibility and convergence delays.

[0033] Firstly, regarding convergence speed: When only a single backup path exists, if the primary path fails, the router often needs to recalculate the backup path using routing protocols or wait for route updates. For example, the RIP protocol has a default update cycle of 30 seconds; after the primary route fails, it needs to wait for periodic updates to learn the backup route, resulting in a very slow convergence time. Even in link-state protocols (OSPF / IS-IS), when the primary path failure triggers SPF recalculation, the network may take hundreds of milliseconds to several seconds to converge. During this period, services may be interrupted.

[0034] Secondly, regarding path utilization and flexibility: Large networks often have multiple suboptimal paths with different costs and topology independence. If only one backup is selected, the remaining available paths remain idle, failing to fully utilize network resources. For example, a site may have multiple alternative WAN links for access. If only one is configured as a backup, when the primary link and the backup link fail one after the other, the other links cannot immediately step in to handle the traffic and must wait for the routing protocol to recalculate, making it difficult to guarantee service continuity. Furthermore, when multiple suboptimal paths with the same metric exist (e.g., two paths in an OSPF topology with slightly higher costs than the primary path but equal costs to each other), traditional routers will select one for forwarding after the primary path fails, rather than simultaneously utilizing these equivalent backup lines for load balancing. Existing routing software such as FRR typically only supports one backup next hop, which limits bandwidth utilization and redundancy in backup mode.

[0035] Secondly, there's the convergence issue in multi-level failure scenarios: As network scales, simultaneous link or node failures are not uncommon. Under the current mechanism, routers only prepare one primary / backup layer for each prefix. If consecutive link interruptions or multiple failures occur, routing must go through multiple "failure-recalculation-switching" cycles. For example, the primary path fails, switching to the sole backup A, but then the link containing backup A breaks again. The router then has to recalculate the next path, causing further interruptions to service traffic. Such multiple convergence processes are extremely detrimental to real-time services. If multiple levels of backup paths could be pre-prepared, the next level could be activated directly when the first backup fails, eliminating the need for a complete protocol convergence process and significantly improving network stability.

[0036] Finally, existing solutions have limitations: some attempts have tried storing all available routes in the router to accelerate failover, but this significantly increases CPU and memory overhead due to the need to maintain multiple next hops for each route, making implementation complex. Early solutions were not widely adopted due to limited equipment resources. However, with the performance improvements of modern network equipment and the increasing demands for high availability, it is necessary to re-examine the feasibility of multi-path backup mechanisms. In summary, current solutions lack support for multi-level backup and equal-cost routing in dynamic route convergence scenarios, failing to fully utilize multiple candidate paths for fast and reliable failover and recovery.

[0037] This application provides a dynamic routing convergence mechanism that enables routers to maintain a multi-level priority backup path list for each destination network prefix and quickly switch between paths of different priorities when the primary path fails (becomes ineffective), while also supporting equal-cost multipath forwarding at the same level.

[0038] For example, see Figure 1The diagram shows a detailed flowchart of a path failure handling method provided in this application embodiment. This method is applied to a network device, which maintains a corresponding priority path list for each destination network prefix. The target priority path list corresponding to the target destination network prefix includes several path sets sorted from high to low priority. The set of available paths with the highest current priority in the target priority path list is the set of currently used paths corresponding to the target destination network prefix. The method includes the following steps: Step 100: Monitor in real time the connectivity status of each path in the current primary path set corresponding to the target network prefix.

[0039] In the implementation of this application, the priority path linked lists maintained by the network device for each destination network prefix can be obtained in the following ways: For each destination network prefix, the metric or administrative distance of each candidate path corresponding to that destination network prefix is ​​calculated according to a preset routing protocol; Multiple paths with the same metric or administrative distance are grouped into a path set of the same priority level as equivalent paths. Paths with smaller metric or administrative distance belong to a path set with higher priority.

[0040] Specifically, in this embodiment, the routing calculation process of the network device needs to be able to store multiple candidate paths in the routing table according to priority (rather than storing only one best path). In practical applications, this can be achieved by extending the routing table entry structure, adding a priority path linked list or array for storing multiple next hops and priority attributes. For example, for each destination network prefix (destination routing prefix), a next-hop list sorted by Cost (metric value, or administrative distance) is maintained, and the currently active item (current primary path set) and the backup item are marked. The forwarding plane (FIB) needs to support multiple next hops based on priority: usually the primary next hop is in an active forwarding state, and the backup next hop can also be pre-installed but marked as alternative. When the primary next hop fails, the control plane triggers a switch to make its state active.

[0041] It should be noted that, to avoid routing loops and black holes, all backup paths should be loop-free and legitimate routes. In this embodiment, in link-state protocols, since each network device (e.g., a router) possesses the entire network topology, loops can be verified during calculation; in environments such as BGP, attributes such as AS-PATH can be combined to ensure that backup paths are loop-free. Thus, during path switching, traffic will not create new routing loops due to the adoption of suboptimal paths.

[0042] For example, suppose the candidate paths corresponding to the target network prefix 1 are path 1, path 2, ..., path 7 and path 8, where the metric value of path 1 and path 3 is 3; the metric value of path 2, path 4 and path 5 is 5; the metric value of path 6 is 8; and the metric value of path 7 and path 8 is 10.

[0043] Therefore, based on the metrics of each path, the candidate paths can be divided into four path sets, where: Path set 1 is {path 1, path 3}; its priority is the highest, first priority. Path set 2 is {path 2, path 4 and path 5}; its priority is the second highest priority. Path set 3 is {path 6}; its priority is the third highest priority. Path set 4 is {path 7, path 8}; its priority is the lowest, the fourth priority.

[0044] In this embodiment of the application, a path set includes at least one path.

[0045] In this embodiment of the application, a preferred method for monitoring the connectivity of each path is as follows: Millisecond-level fault detection is achieved by establishing a bidirectional forwarding detection (BFD) session for the next hop of each path.

[0046] In other words, this embodiment relies on rapid state monitoring to detect the availability of paths at all levels. A typical implementation involves establishing BFD sessions with the next-hop neighbors of each path, or utilizing the protocol's built-in Keepalive / Hello mechanism. When the BFD session of the primary path fails within a short period (e.g., 10 milliseconds), the primary link is determined to be down, and a route switch is immediately triggered without waiting for OSPF LSA flooding or BGP Hold Timer expiration. This rapid detection, combined with prepared backup paths, reduces traffic interruption time after a link failure to the millisecond level.

[0047] In this embodiment of the application, the above-mentioned path fault handling method may further include the following steps: Monitor the connectivity status of each path in the target priority path linked list, which includes several path sets sorted from high to low priority; if it is determined that all paths in any path set are invalid, mark that path set as an unavailable path set.

[0048] In other words, BFD monitoring can also be performed on backup paths. That is, BFD monitoring is performed on each path included in all path sets. When it is determined that the connectivity of a path included in a path set is abnormal, the path is removed from the path set so that all paths currently included in each path set are available. This allows the system to detect the status of the backup before it is actually effective (as the current primary path) and avoids switching to a path set that includes paths that have already failed.

[0049] In this embodiment of the application, if the current primary path set includes multiple equivalent paths, then the traffic of the target destination network prefix is ​​load-balanced by the multiple equivalent paths.

[0050] For example, assuming the current primary path set is path set 1, which includes path 1 and path 3, then the traffic of destination network prefix 1 can be load-shared by path 1 and path 3.

[0051] Therefore, for each path set, if a connectivity failure is detected in one of the paths included in the path set, the faulty path is removed from the path set. If the path set is the current primary path set, the traffic of the target destination network prefix is ​​load-balanced by the other equivalent paths included in the path set after removing the faulty path.

[0052] For example, suppose the current primary path set is path set 1, which includes path 1 and path 3. If a connectivity failure is detected in path 1, then path 1 will be removed from path set 1. In this case, path set 1 will only include path 3, and path 3 will independently handle the traffic of destination network prefix 1.

[0053] Step 110: If, based on the monitoring results, it is determined that all paths included in the current primary path set are invalid, mark the current path set as an unavailable path set.

[0054] As shown above, the system monitors the connectivity status of each path in the current primary path set in real time. When a path is found to have abnormal connectivity, it is determined to be invalid and removed from the current primary path set. If all paths in the current primary path set are determined to be invalid, the current primary path set is determined to be unavailable. At this point, the current primary path set can be marked as an unavailable path set.

[0055] At this point, it is necessary to select a new set of paths from the backup path set that currently has the highest priority and is available, as the current primary path set for the target network prefix.

[0056] Step 120: Determine the set of available paths with the highest current priority from the target priority path list, and use the newly determined set of available paths with the highest current priority as the set of current primary paths corresponding to the target destination network prefix.

[0057] Specifically, a set of available paths with the highest current priority is determined from the target priority path list, and the determined set of available paths with the highest current priority is used as the set of current primary paths corresponding to the target destination network prefix. The traffic corresponding to the target destination network prefix is ​​switched to the paths included in the newly determined set of current primary paths for load balancing.

[0058] For example, suppose the current primary path set is path set 1, which is the highest priority path set in the target priority path list. The target priority path list also includes path set 2, which is the second highest priority path set. Then, after all paths in path set 1 become unavailable, the currently highest priority available path set in the target priority path list becomes path set 2, which is the second highest priority path set. Path set 2 is then designated as the primary path set corresponding to the target destination network prefix. Traffic corresponding to the target destination network prefix is ​​then switched to the paths included in path set 2 for load balancing.

[0059] In this embodiment of the application, the connectivity status of each path in a set of paths sorted from high to low priority included in the target priority path list is monitored, and based on the connectivity status of each path in the set of paths sorted from high to low priority included in the target priority path list, the available path set of the next priority level of the current primary path set is determined; and the available path set of the next priority level is set as the backup next hop of the current primary path set in the routing information base (RIB) or forwarding information base (FIB).

[0060] In practical applications, the current primary path set can be marked as active, while other backup path sets can be marked as inactive. Then, when all equivalent paths of the highest priority (current primary path) fail, the router immediately switches the active path (current primary path) in the routing table to the next priority level. This switchover can be achieved through pre-installed backup next hops in the FIB, enabling instantaneous failover and ensuring service continuity. If the second-level backup contains multiple equivalent paths (e.g., two suboptimal paths with the same cost value), these paths will collectively serve as the new active forwarding paths, allowing the router to load balance traffic. Thus, after the primary path fails, service traffic is not limited to a single path but can be transmitted in parallel on multiple suboptimal lines, improving bandwidth utilization during the failure period.

[0061] After detecting that the current path set (e.g., first priority) has 0 available paths and is unavailable, and the next priority (the path set with the highest priority in the current available path set, i.e., second priority) becomes the current active path set, the router continues to monitor the connectivity status of each actual forwarding path within that priority path set. If there are multiple equivalent paths within that priority path set, then if one of these paths is detected to be interrupted (failed), the router removes that path from the priority path set, but does not change the forwarding status of the priority path set. That is, as long as the priority path set still includes at least one available path, traffic continues to be forwarded on the remaining paths included in the priority path set (if there are multiple remaining paths, load balancing continues). Only when all paths within the priority path set fail will a further switchover operation of the current primary path set be triggered.

[0062] If all paths in the second priority path set (the current primary path set) also fail (e.g., multiple equal-cost suboptimal paths may fail simultaneously or sequentially due to shared link failures), and the first priority path set has not been restored to availability, the router will sequentially use the third priority path set as the new forwarding route. Similarly, the third priority path set can be a single path or a combination of multiple equal-cost paths. If all paths in the third priority path set also fail, the router continues searching down the next priority backup path set. This process is similar to an ordered "relay": at any given time, the highest priority available path set is always selected as the current primary path. Once the previous level recovers, the router switches back according to the failback policy. By maintaining multi-level linked lists, the router can automatically undergo a smooth transition through each level in the event of multiple failures at different times, without having to rely on the routing protocol to calculate new paths from scratch each time.

[0063] Furthermore, in this embodiment of the application, after detecting that a higher priority path set with a priority higher than the current primary path set has become available again, a switchback delay timer is started; if the switchback delay timer expires and the higher priority path set remains available, the higher priority path set is used as the current primary path set corresponding to the target destination network prefix, so as to switch the traffic of the target destination network prefix back to the higher priority path set.

[0064] In other words, when a higher-priority path set recovers from a failure and becomes available again, this embodiment supports a failback strategy. Specifically, the recovered high-priority path set will not immediately preempt traffic; instead, an observation / stabilization timer is typically set. After confirming that the path is stable, available, and performing well, the router then promotes it to an active path (the current primary path set) and migrates traffic back to the primary high-priority available path set. Preferably, traffic can be gradually redirected during failback to prevent oscillations. For example, after the primary path AB recovers, it can first be allowed to perform BFD probing for a period to confirm link quality, and then some traffic can begin to be migrated back to AB, eventually fully restoring the high-priority path set as the sole active path set. This ensures that the network can reuse the optimal path for transmission after a failure, while avoiding routing jitter caused by frequent switching.

[0065] The path failure handling process provided in this application embodiment will be described in detail below with reference to a specific application scenario. For example, assume there are four alternative paths from node A to destination node J: Primary path set (first priority): A–B–J path. This path is assumed to have the lowest cost and highest bandwidth, therefore it is selected as the primary route. Normally, traffic sent from A to J is forwarded via ABJ.

[0066] The second-level backup path set (second priority) contains two equivalent suboptimal paths, A–C–D–J and A–E–F–J. These two paths have slightly higher metrics than the primary path, but they are identical and do not share critical nodes, thus belonging to the second priority level. They serve as parallel backups and will share traffic (ECMP load balancing) when the primary path becomes unavailable.

[0067] The third-level backup path set (third priority): the A–G–H–I–J path, has the highest overhead (most hops or lowest bandwidth) and serves as the last backup path. This path is only activated when all higher-priority paths fail.

[0068] During normal operation, Node A forwards data using the primary path A–B–J, while links C / D, E / F, and G / H / I are all in standby monitoring mode. When Node B or link AB fails, causing the primary path to be interrupted, Node A immediately detects the link break via BFD detection and immediately stops the A–B–J route. At this time, Node A selects a second-level backup path set from its backup list, namely A–C–D–J and A–E–F–J, as the new active routes. Since these two paths have the same priority and metric, Node A can use the ECMP strategy to distribute traffic on the two paths, achieving seamless switching and load balancing. In this way, even if the primary path suddenly fails, communication between Node A and Node J can continue, and due to the rapid switching, there is virtually no packet loss.

[0069] If one of the paths in the second-level backup path set (the current primary path set) (e.g., A–C–D–J) subsequently fails, but another path (A–E–F–J) remains operational, then node A does not need further failover: the remaining A–E–F–J path will carry all traffic alone. Only when all paths in the second-level backup path set become unavailable (e.g., CDJ and EFJ are both down, possibly due to a failure at node J) will node A activate the third-level backup path set A–G–H–I–J as a last resort. Since the third-level backup path set consists of only this single path, there is no load balancing, but it at least ensures that there is still a working path from node A to node J, preventing a complete outage.

[0070] As the network failure gradually recovers, assuming higher-level paths become reachable again: Once one of the paths in the second-level backup path set (e.g., CDJ) recovers and becomes stable, node A can reintegrate it into the second-level backup path set to share traffic with the existing EFJ. If the A–B–J path in the primary path set also recovers connectivity and is of good quality, then after a period of observation, node A will prioritize switching back to the primary path, redirecting traffic back to the A–B–J line. The backup path set is then released into standby mode (or dial-up backups are disconnected to conserve resources). The entire process ensures that the best path is always prioritized, optimizing forwarding efficiency while ensuring rapid convergence.

[0071] Based on the same inventive concept as the above-described embodiments, see, for example, the following: Figure 2 The diagram shown is a schematic representation of a path fault handling device provided in an embodiment of this application. This device is applied to a network device, which maintains a corresponding priority path list for each destination network prefix. The target priority path list corresponding to the target destination network prefix includes several path sets sorted from high to low priority. The set of available paths with the highest current priority in the target priority path list is the set of currently used paths corresponding to the target destination network prefix. The device includes: Monitoring unit 20 is used to monitor in real time the connectivity status of each path in the current primary path set corresponding to the target destination network prefix; If the determining unit 21 determines, based on the monitoring results, that all paths included in the current primary path set are invalid, it marks the current path set as an unavailable path set. The switching unit 22 is used to determine the set of available paths with the highest current priority from the target priority path list, and to use the newly determined set of available paths with the highest current priority as the set of current primary paths corresponding to the target destination network prefix.

[0072] Optionally, the monitoring unit 20 is used to monitor the connectivity status of each path in the target priority path list, which includes a set of paths sorted from high to low priority. If the determining unit 21 determines that all paths included in any path set are invalid, it marks the path set as an unavailable path set.

[0073] Optionally, when monitoring the connectivity of each path, the monitoring unit 20 is specifically used for: Millisecond-level fault detection is achieved by establishing a bidirectional forwarding detection (BFD) session for the next hop of each path.

[0074] Optionally, the device further includes: The calculation unit is used to calculate the metric value or administrative distance of each candidate path corresponding to each destination network prefix according to a preset routing protocol. A partitioning unit is used to divide multiple paths with the same metric or administrative distance into a path set of the same priority level as equivalent paths. The path set to which the path with the smaller metric or administrative distance belongs has a higher priority.

[0075] Optionally, if the current primary path set includes multiple equivalent paths, then the traffic of the target destination network prefix is ​​load-balanced by the multiple equivalent paths.

[0076] Optionally, for each path set, if the monitoring unit 20 detects a connectivity failure in one of the paths included in the path set, the faulty path is removed from the path set. If the path set is the current primary path set, the traffic of the target destination network prefix is ​​load-sharing by the other equivalent paths included in the path set after removing the faulty path.

[0077] Optionally, the device further includes: The back-switch unit is used to start a back-switch delay timer after detecting that a higher priority path set with a higher priority than the current primary path set has become available again; if the back-switch delay timer expires and the higher priority path set remains available, the higher priority path set is used as the current primary path set corresponding to the target destination network prefix, so as to switch the traffic of the target destination network prefix back to the higher priority path set.

[0078] Optionally, the device further includes: The setting unit is used to determine the available path set of the next priority level of the current primary path set based on the connectivity status of each path in the target priority path linked list, which includes several path sets sorted from high to low priority; and to set the available path set of the next priority level as the backup next hop of the current primary path set in the routing information base (RIB) or the forwarding information base (FIB).

[0079] These units can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more digital signal processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when one of these units is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these units can be integrated together to form a system-on-a-chip (SOC).

[0080] Furthermore, regarding the path fault handling device provided in this application embodiment, from a hardware perspective, the hardware architecture schematic diagram of the path fault handling device can be found in [reference needed]. Figure 3 As shown, the path fault handling device may include: a memory 30 and a processor 31. The memory 30 is used to store program instructions; the processor 31 calls the program instructions stored in the memory 30 and executes the above method embodiment according to the obtained program instructions. The specific implementation method and technical effect are similar, and will not be described again here.

[0081] Optionally, this application also provides a network device, including at least one processing element (or chip) for performing the above method embodiments.

[0082] Optionally, this application also provides a program product, such as a computer-readable storage medium storing computer-executable instructions for causing the computer to perform the above-described method embodiments.

[0083] Here, a machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, a machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0084] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0085] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0086] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0087] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0088] Furthermore, these computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0089] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0090] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A path fault handling method, characterized in that, The method is applied to network devices, wherein the network devices maintain corresponding priority path lists for each destination network prefix, wherein the target priority path list corresponding to the target destination network prefix includes several path sets sorted from high to low priority, and the set of available paths with the highest current priority in the target priority path list is the set of currently used paths corresponding to the target destination network prefix; the method includes: Real-time monitoring of the connectivity status of each path in the current primary path set corresponding to the target network prefix; If, based on monitoring results, it is determined that all paths in the current primary path set are invalid, the current path set will be marked as an unavailable path set. The set of available paths with the highest current priority is determined from the target priority path list, and the newly determined set of available paths with the highest current priority is used as the set of current primary paths corresponding to the target destination network prefix.

2. The method as described in claim 1, characterized in that, The method further includes: Monitor the connectivity status of each path in the target priority path list, which includes a set of paths sorted from high to low priority. If it is determined that all paths in any set of paths are invalid, then that set of paths is marked as an unavailable set of paths.

3. The method as described in claim 1 or 2, characterized in that, The steps for monitoring the connectivity of each path include: Millisecond-level fault detection is achieved by establishing a bidirectional forwarding detection (BFD) session for the next hop of each path.

4. The method as described in claim 1 or 2, characterized in that, The method further includes: For each destination network prefix, the metric or administrative distance of each candidate path corresponding to that destination network prefix is ​​calculated according to a preset routing protocol; Multiple paths with the same metric or administrative distance are grouped into a path set of the same priority level as equivalent paths. Paths with smaller metric or administrative distance belong to a path set with higher priority.

5. The method as described in claim 1 or 2, characterized in that, If the current primary path set includes multiple equivalent paths, then the traffic of the target destination network prefix is ​​load-balanced by the multiple equivalent paths.

6. The method as described in claim 5, characterized in that, For each path set, if a connectivity failure is detected in one of the paths in the path set, the faulty path is removed from the path set. If the path set is the current primary path set, the traffic of the target destination network prefix is ​​load-balanced by the other equivalent paths in the path set after removing the faulty path.

7. The method as described in claim 1, characterized in that, The method further includes: Once a set of higher-priority paths (which have a higher priority than the current primary path set) is detected to be available again, a switchback delay timer is started. If the back-switch delay timer times out and the higher priority path set remains available, the higher priority path set is used as the current primary path set corresponding to the target destination network prefix, so as to switch the traffic of the target destination network prefix back to the higher priority path set.

8. The method as described in claim 2, characterized in that, Based on the connectivity status of each path in the target priority path linked list, which includes several path sets sorted from high to low priority, determine the available path set of the next priority level of the current primary path set; and set the available path set of the next priority level as the backup next hop of the current primary path set in the routing information base (RIB) or forwarding information base (FIB).

9. A path fault handling device, characterized in that, Applied to network devices, the network devices maintain corresponding priority path lists for each destination network prefix, wherein the target priority path list corresponding to the target destination network prefix includes several path sets sorted from high to low priority, and the set of available paths with the highest current priority in the target priority path list is the set of currently used paths corresponding to the target destination network prefix; the device includes: The monitoring unit is used to monitor in real time the connectivity status of each path in the current primary path set corresponding to the target network prefix; If the determining unit determines, based on the monitoring results, that all paths included in the current primary path set are invalid, it marks the current path set as an unavailable path set. The switching unit is used to determine the set of available paths with the highest current priority from the target priority path list, and to use the newly determined set of available paths with the highest current priority as the set of current primary paths corresponding to the target destination network prefix.

10. A path fault handling device, characterized in that, The path fault device includes: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the steps of the method as described in any one of claims 1-8 according to the obtained program instructions.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing the computer to perform the steps of the method as described in any one of claims 1-8.