A cascading failure vulnerability testing system, method, and related devices

By constructing a cascading failure vulnerability testing system in a virtualized environment, simulating routing link bandwidth scaling and node function reconfiguration, the problem of cascading failure vulnerability assessment in large-scale routing and switching systems is solved, achieving efficient simulation of cascading failures and vulnerability location.

CN116707865BActive Publication Date: 2026-05-19BEIJING DINGNIU TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING DINGNIU TECH CO LTD
Filing Date
2023-05-09
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies struggle to reproduce cascading failures in large-scale routing and switching systems, making it impossible to effectively assess the vulnerability of routing and switching systems to cascading failures, especially under the complex interactions between routing nodes at scales of thousands or tens of thousands.

Method used

A cascading failure vulnerability testing system is adopted, including a routing switching module, a fault/attack event simulation module, and a vulnerability testing module. It simulates routing link bandwidth scaling, routing node function reconstruction, and fault/attack events in a virtualized environment, and performs tests and evaluations, including simulations of node failure, link failure, routing update storm, and congestion attacks.

Benefits of technology

It enables cascading failure vulnerability testing of large-scale routing and switching systems in a virtualized environment. It can simulate destructive factors such as routing update storms and traffic redirection, locate vulnerabilities, evaluate the performance degradation of the system under failure/attack, and support high-fidelity and large-scale network testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116707865B_ABST
    Figure CN116707865B_ABST
Patent Text Reader

Abstract

The application discloses a kind of cascading failure vulnerability test system, method set related device, including: routing exchange module, fault / attack event simulation module, vulnerability test module, wherein the routing exchange module is used to set the bandwidth reduction of current routing exchange system's routing link and reconstruct to routing node function;The fault / attack event simulation module is used to simulate the current routing exchange system based on preset fault / attack event after reconstruction is completed;The vulnerability test module is used to test and evaluate the cascading failure vulnerability of the current routing exchange system after simulation is completed.The above-mentioned system, routing exchange, fault / attack simulation and vulnerability test are carried out based on virtualization environment, and large-scale network can be constructed based on virtual environment, and large-scale routing node can be reproduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of system vulnerability detection technology, and in particular to a cascading failure vulnerability testing system, method and related apparatus. Background Technology

[0002] Cascading failures in routing and switching systems refer to the phenomenon where, when a small number of routing nodes fail, the routing algorithm causes the load of the failed nodes to be distributed to neighboring nodes. This places additional load on neighboring nodes and other surrounding nodes. When the load exceeds the capacity of these nodes, a new "node failure-load transfer" process occurs, leading to the continuous spread of the failure and even widespread routing paralysis. Cascading failures can be caused by occasional node failures or configuration errors, or by deliberate attacks (such as the "digital cannon" or CXPST attack). Given the severe consequences of causing widespread routing paralysis, testing and evaluating the vulnerability of routing and switching systems to cascading failures is extremely necessary.

[0003] Traditional vulnerability discovery and vulnerability analysis methods primarily target real physical devices, employing forged routers to inject spoofed LSU messages into target routers, thereby monitoring the target router's status. While these testing methods provide reliable measurements of device capacity, they are not suitable for building large-scale tested networks. They cannot reproduce the vulnerability in real-world routing and switching systems with thousands or tens of thousands of routing nodes, where the complex interactions between numerous routing nodes are precisely the key to route cascading failures. Summary of the Invention

[0004] In view of this, the present invention provides a cascading failure vulnerability testing system, method, and related apparatus to address the problem that existing vulnerability discovery and vulnerability analysis methods primarily target real physical devices, using forged routers to inject forged LSU messages into target routers, thereby monitoring the target router's status. While such testing methods provide reliable measurements of device capacity, they are not suitable for building large-scale initial test networks and cannot reproduce the vulnerability on thousands or tens of thousands of routing nodes in actual routing and switching systems. The specific solution is as follows:

[0005] A cascading failure vulnerability testing system includes: a routing and switching module, a fault / attack event simulation module, and a vulnerability testing module, wherein...

[0006] The routing and switching module is used to set the bandwidth scaling ratio of the current routing and switching system and to reconstruct the functions of the routing nodes;

[0007] The fault / attack event simulation module is used to simulate the current routing and switching system based on preset fault / attack events after the reconstruction is completed;

[0008] The vulnerability testing module is used to test and evaluate the cascading failure vulnerability of the current routing and switching system after the simulation is completed.

[0009] Optionally, in the above system, the routing switching module includes: a routing link bandwidth scaling unit and a routing node function reconfiguration unit, wherein,

[0010] The routing link bandwidth reduction unit is used to reduce the routing bandwidth of the current routing switching system according to a preset reduction ratio;

[0011] The routing node function reconfiguration unit is used to simulate the crash-restart characteristics when receiving update messages and the packet loss characteristics when the router is congested.

[0012] Optionally, in the aforementioned system, the fault / attack event simulation module includes: a node failure unit, a link failure unit, a route update unit, and a congestion attack unit, wherein...

[0013] The node failure unit is used to control the Agent on the first node to disable the BGP and OSPF of the first node based on the no bgp and no ospf instructions;

[0014] The link failure unit is used to disconnect the currently pending link;

[0015] The routing update unit is used to control the Agent on the second node to send a first preset number of packets to the neighboring nodes, causing the neighboring nodes and the remote nodes to fall into a routing update storm, resulting in the paralysis of the neighboring nodes and the remote nodes.

[0016] The congestion attack unit is used to control the virtual node to inject a second preset number of packets into the target link, causing the target link to become congested, or to cause the routing session to be disconnected by causing the routing heartbeat packet to be lost.

[0017] Optionally, the vulnerability testing module in the aforementioned system includes: an attack target selection unit, a first evaluation unit, a second evaluation unit, and a vulnerability location unit, wherein...

[0018] The attack target selection unit is used to select a target node or target link as an attack target based on the type of fault / attack event.

[0019] The first evaluation unit is used to measure the overall network throughput, overall network path performance, and critical link performance, and to locate the paralyzed nodes.

[0020] The second evaluation unit is used to evaluate the traffic forwarding function and traffic forwarding performance, and to determine cascading failures.

[0021] The vulnerability location unit is used to locate vulnerability points in routing updates and vulnerability points in congestion.

[0022] Optionally, in the above system, the attack target selection unit includes: a first attack target selection subunit and a second attack target selection subunit, wherein,

[0023] The first attack target selection subunit is used to select target nodes whose betweenness reaches a third preset number threshold as attack targets when the fault / attack event is a node failure or a route update.

[0024] The second attack target selection subunit is used to select a target link whose betweenness reaches a fourth preset threshold as the attack target when the fault / attack event is a link failure or a congestion attack.

[0025] Optionally, in the aforementioned system, the first evaluation unit includes: a throughput measurement subunit, a network-wide path performance testing subunit, a critical link performance testing subunit, and a paralyzed node location subunit, wherein...

[0026] The throughput measurement subunit is used to traverse all nodes in the current routing and switching system, and to count the first and second traffic at each point before and during a fault / attack event within a preset time period. Based on the preset time period, the first and second traffic are used to determine the overall network throughput degradation coefficient.

[0027] The critical link performance testing subunit is used to select a third preset number of high betweenness links, perform ping probing on the third preset number of high betweenness links, and count the first average latency and the second average latency before and during the failure / attack event. Based on the third preset number, the first average latency and the second average latency, the critical link performance degradation coefficient is determined.

[0028] The paralyzed node location subunit is used to report process shutdown / restart events to the control terminal when the current node's BGP and OSPF processes are shut down based on the no BGP and no OSPF instructions, and when the routing process is restarted, so as to locate the paralyzed node.

[0029] Optionally, in the above system, the second evaluation unit includes: a traffic forwarding function evaluation subunit, a traffic forwarding performance evaluation subunit, and a cascading failure determination subunit, wherein,

[0030] The traffic forwarding function subunit is used to obtain the network throughput degradation coefficient. If the network throughput degradation coefficient is within the first preset range, it is determined that the current network traffic forwarding function is normal. If the network throughput coefficient is within the second preset range, the traffic forwarding performance evaluation subunit is executed.

[0031] The traffic forwarding performance evaluation subunit is used to obtain the network path performance degradation coefficient. If the network path performance degradation coefficient is within the second preset range, it is determined that the network fault is local and the cascading failure determination subunit is executed. Otherwise, it is determined that the network fault is global.

[0032] The cascading failure determination subunit is used for all routing node process shutdown / restart events. If the number of process shutdown / restart events exceeds a preset threshold, a cascading failure is determined to have occurred; otherwise, no cascading failure has occurred.

[0033] Optionally, in the aforementioned system, the vulnerability location unit includes: a route update vulnerability testing subunit and a congestion vulnerability location subunit, wherein,

[0034] The routing update vulnerability testing subunit is used to perform attack tests and determine cascading failures. If there are no cascading failures, a list of paralyzed nodes is obtained. If there are cascading failures, the routing update tolerance threshold of the paralyzed nodes is increased by a preset ratio and the attack test is repeated until there are no cascading failures and a list of paralyzed nodes is obtained. The nodes in the list of paralyzed nodes are vulnerable nodes to cascading failures.

[0035] The congestion vulnerability location subunit is used to locate congested links based on critical link performance testing, and to identify the congested links as congestion vulnerability points.

[0036] A method for testing cascading failure vulnerability includes:

[0037] Pre-set the bandwidth scaling ratio of the routing links in the current routing and switching system;

[0038] After the settings are completed, the routing node functions will be refactored for both route update storm and traffic redirection.

[0039] After reconstruction, a preset fault / attack event is selected, and the current routing and switching system is simulated based on the preset fault / attack event;

[0040] After the simulation is completed, the vulnerability of the current routing and switching system to cascading failures is tested and evaluated.

[0041] A testing device for cascading failure vulnerability includes:

[0042] The configuration module is used to pre-set the bandwidth scaling ratio of the routing links in the current routing and switching system;

[0043] The refactoring module is used to refactor the routing node functionality for both route update storms and traffic redirection after the settings are completed.

[0044] The simulation module is used to select preset fault / attack events after reconstruction and simulate the current routing and switching system based on the preset fault / attack events;

[0045] The test and evaluation module is used to test and evaluate the vulnerability of the current routing and switching system to cascading failures after the simulation is completed.

[0046] Compared with the prior art, the present invention has the following advantages:

[0047] This invention discloses a cascading failure vulnerability testing system and related apparatus, including: a routing switching module, a fault / attack event simulation module, and a vulnerability testing module. The routing switching module is used to set the routing link bandwidth scaling of the current routing switching system and reconstruct the functions of routing nodes. The fault / attack event simulation module is used to simulate the current routing switching system based on preset fault / attack events after reconstruction. The vulnerability testing module is used to test and evaluate the cascading failure vulnerability of the current routing switching system after simulation. In the above system, routing switching, fault / attack simulation, and vulnerability testing are all performed in a virtualized environment. The virtual environment facilitates the construction of large-scale networks and allows for the reproduction of large-scale routing nodes. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a block diagram of a real cascading failure vulnerability testing system disclosed in an embodiment of the present invention;

[0050] Figure 2 This is a flowchart illustrating the capacity assessment of a routing and switching system to withstand fault / attack events, as disclosed in an embodiment of the present invention.

[0051] Figure 3 This is a flowchart of a route update vulnerability location method disclosed in an embodiment of the present invention;

[0052] Figure 4 This is a flowchart of a method for testing cascading failure vulnerability disclosed in an embodiment of the present invention;

[0053] Figure 5 This is a structural block diagram of a cascading failure vulnerability testing device disclosed in an embodiment of the present invention. Detailed Implementation

[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] This invention discloses a testing system, method, and apparatus for cascading failure vulnerability, applied to the testing process of cascading failures in routing and switching systems. During cascading failures in routing and switching systems, two main destructive factors drive routing node failure: First, a routing update storm. When a neighboring node discovers a failed node, it sends update messages to all its friendly neighbors, informing them that routes related to the failed node are invalid and announcing new routes. This routing update storm continues to spread until all nodes in the network are aware of the route update. Many nodes will be busy processing update messages and unable to handle routing forwarding services, and may even crash due to resource overload. Second, traffic redirection. Traffic originally flowing through the failed node flows through neighboring nodes, causing congestion on their corresponding links, continuous loss of routing heartbeat packets, and misjudging normal nodes as failed, thus triggering new routing updates and traffic redirection. The key to testing the vulnerability of routing and switching systems to cascading failures is to reproduce these two destructive factors and their effects on the routing system, thereby observing the degree to which the routing and switching system is affected by these destructive factors.

[0056] Traditional vulnerability discovery and vulnerability analysis methods primarily target real physical devices, employing forged routers to inject spoofed LSU packets into target routers, thereby monitoring the target router's status. While these testing methods provide reliable measurements of device capacity, they are unsuitable for building large-scale initial test networks and cannot reproduce the vulnerability of thousands or tens of thousands of routing nodes in real-world routing and switching systems. However, the complex interactions between numerous routing nodes in reality are crucial to cascading failures. Cascading failure vulnerability testing in routing and switching systems demands high fidelity in simulating the interactions between routing nodes, while the fidelity of the simulation of the routing nodes themselves is relatively low. Therefore, achieving a suitable trade-off between high fidelity and large scale is essential. This invention provides a cascading failure vulnerability testing system, the structural block diagram of which is shown below. Figure 1 As shown, it includes: a routing and switching module 101, a fault / attack event simulation module 102, and a vulnerability testing module.

[0057] in,

[0058] The routing and switching module 101 is used to set the bandwidth scaling of the current routing and switching system and to reconstruct the functions of the routing nodes;

[0059] In this embodiment of the invention, the cascading failure vulnerability testing system is not located in a cloud environment, for example, it is not a Quagaa soft router on the vSphere platform. The routing switching module includes: a routing link bandwidth scaling unit and a routing node function reconfiguration unit, wherein...

[0060] The routing link bandwidth reduction unit is used to reduce the routing bandwidth in the current routing switching system according to a preset reduction ratio. The preset reduction ratio can be set based on experience or specific circumstances. In this embodiment of the invention, no specific limitation is made. After selecting the reduction ratio, the reduced uplink and downlink bandwidth can be set for the selected link. In the VMware vSphere platform, this can be achieved by setting the uplink and downlink bandwidth of the selected port of the distributed virtual switch. The purpose of performing routing bandwidth reduction simulation is to simulate high-performance links in reality using low-performance links in the cloud platform. This can be achieved based on the link performance configuration function of the cloud platform. Considering that the simulation objects are mainly 10GE fiber, 40GE fiber and 100GE fiber, a preset reduction ratio of 1 / 100000 or 1 / 10000 can be used, that is, using 100Kbps / 1Mbps, 400Kbps / 4Mbps and 1000Kbps / 10Mbps to simulate the above fiber links respectively.

[0061] The routing node function reconfiguration unit is used to simulate the crash-restart characteristics when receiving update messages and the packet loss characteristics when the router is congested. The specific processing procedure is as follows:

[0062] By making the following modifications and configurations to the soft router node, it can simulate the characteristics of a real router node facing two destructive factors: route update storms and traffic redirection (the specific configuration should be based on the test results of a single physical device):

[0063] To simulate the crash-restart characteristic when receiving update messages, an Agent is deployed on a soft router node, and a routing update tolerance threshold is set for that node. This threshold depends on the object being simulated by the soft router node. The threshold is obtained by performing routing update stress tests on the physical devices to be simulated. The Agent intercepts the UPDATE and LSU messages received by the node and counts the number of updates received per unit time. When the node receives a specified number of update entries within a unit time, the BGP and OSPF processes are shut down using the `no bgp` and `no ospf` commands, and then restarted later. This simulates the crash-restart characteristic of a router receiving excessive update messages. The BGP and OSPF processes are responsible for handling inter-domain and intra-domain routing information, respectively. Both BGP and OSPF routing processes run simultaneously on inter-domain routers to achieve inter-domain network communication.

[0064] To simulate packet loss characteristics when a router experiences congestion, tools such as ethtool are used to configure the network interface card (NIC) buffer of the soft router node. This allows packet loss to occur within a specified time after the corresponding link becomes congested, simulating the packet loss characteristics of a router during congestion. ethtool is a network interface card (NIC) management tool running on Linux that can configure NIC parameters. For example, the command ethtool -G eth0 rx 2048 sets the receive buffer of the eth0 NIC to 2048 bytes.

[0065] The fault / attack event simulation module 102 is used to simulate the current routing and switching system based on preset fault / attack events after the reconstruction is completed;

[0066] In this embodiment of the invention, the fault / attack event simulation module includes: a node failure unit, a link failure unit, a route update unit, and a congestion attack unit, wherein...

[0067] The node failure unit is used to control the Agent on the first node to shut down the BGP and OSPF processes of the first node using the `no bgp` and `no ospf` commands. The simulation method is as follows: control the Agent on the first node to use the `no bgp` and `no ospf` commands to shut down the BGP and OSPF processes of the first node. Applicable scenarios: routing and switching equipment failure, or routing and switching nodes being attacked and rendered inoperable.

[0068] The link failure unit is used to disconnect the currently unconnected link using the virtual link management function. Use case: Link failure. The virtual link management function is provided by a cloud platform; for example, the vSphere platform provides the function to close a specific port on a virtual switch, thus disconnecting the virtual link associated with that port. The route update unit is used to control the Agent on the second node to send a first preset number of packets to neighboring nodes, causing the neighboring nodes and remote nodes to fall into a route update storm, resulting in their paralysis. The simulation method is as follows: control the Agent on certain nodes (the second node) to send a first preset number of UPDATE / LSU packets to neighboring nodes, causing the neighboring nodes and remote nodes to fall into a route update storm and become paralyzed. In this embodiment, the first preset number of packets is not specifically limited. Applicable scenario: Routing switching nodes are controlled for malicious attacks.

[0069] The congestion attack unit is used to control virtual nodes to inject a second preset number of packets into the target link, causing congestion on the target link, or to cause the routing session to disconnect by causing the loss of routing heartbeat packets. The simulation method is as follows: control virtual nodes to inject the second preset number of packets into the target link, causing congestion on the corresponding link and resulting in a decrease in routing performance, or to cause the routing session to disconnect by causing the loss of routing heartbeat packets (such as HELLO packets of the OSPF protocol and KEEPALIVE packets of the BGP protocol). In this embodiment of the invention, the second preset number is not specifically limited; the second preset number is sufficient to cause congestion. Furthermore, for a single target link, the attack traffic should be increased as much as possible to achieve a higher packet loss rate. However, when facing multiple target links, the planning and allocation of attack traffic needs to be considered, but at least it should be ensured that the attack traffic exceeds the available bandwidth of the link, causing congestion and packet loss on the target link. Applicable scenarios: Attacks on the data plane or control plane of a routing switching network.

[0070] The vulnerability testing module 103 is used to test and evaluate the cascading failure vulnerability of the current routing and switching system after the simulation is completed.

[0071] In this embodiment of the invention, the vulnerability testing module 103 includes: an attack target selection unit, a first evaluation unit, a second evaluation unit, and a vulnerability location unit, wherein,

[0072] The attack target selection unit selects a target node or target link as the attack target based on the type of fault / attack event. Further, the attack target selection unit includes: a first attack target selection unit and a second attack target selection unit, wherein...

[0073] The first attack target selection unit is used to select target nodes with a betweenness index reaching a third preset threshold as attack targets when the fault / attack event is a node failure or route update. The third preset threshold can be set based on experience or specific circumstances. In this embodiment of the invention, no specific limitation is made. For example, the nodes with the top 5% betweenness index can be used as targets. Such targets are related to a large number of routes and have high attack value.

[0074] The specific selection process is as follows: Calculate the node betweenness number BN of all nodes in the network. i =∑ j,k∈N n j,k (i) / n j,k In this process, j and k traverse all nodes, and n... j,k n represents the number of routes between nodes j and k. j,k(i) represents the number of nodes i that pass through in the route between nodes j and k. All nodes in the network are sorted in descending order of their betweenness numbers. Target nodes whose betweenness numbers reach the third preset threshold are selected and defined as high betweenness number nodes. These high betweenness number nodes are used as attack targets. High betweenness number nodes are nodes associated with a large number of routes. The failure of such nodes will lead to a large number of route failures and traffic redirection.

[0075] Furthermore, the above selection process is applicable to targets in two types of failure / attack events: node failure and route update storm.

[0076] The second attack target selection unit is used to select target links whose betweenness reaches a fourth preset threshold as attack targets when the fault / attack event is a link failure or a congestion attack. The fourth preset threshold can be set based on experience or specific circumstances, and is not specifically limited in this embodiment of the invention.

[0077] The specific selection process includes: calculating the link betweenness coefficient BL of the entire network links. i =∑ j,k∈N l j,k (i) / l j,k In this process, j and k traverse all nodes, and l... j,k The number of routes between nodes j and k, l j,k (i) represents the number of links i traversed in the route between nodes j and k. All links in the network are sorted in descending order of their betweenness numbers. Target links whose betweenness numbers reach a fourth preset threshold are selected as high betweenness links, and these high betweenness links are used as attack targets. Similar to high betweenness nodes, attacking high betweenness links can affect as many routing entries and traffic paths as possible.

[0078] The first evaluation unit is used to measure the overall network throughput, overall network path performance, and critical link performance, and to locate the paralyzed nodes. The first evaluation unit includes: a throughput measurement subunit, an overall network path performance testing subunit, a critical link performance testing subunit, and a paralyzed node location subunit.

[0079] The throughput measurement subunit is used to traverse all nodes in the current routing and switching system, and to count the first and second traffic at each point before and during a fault / attack event within a preset time period. Based on the preset time period, the first and second traffic are used to determine the overall network throughput degradation coefficient.

[0080] In this embodiment of the invention, the overall network throughput coefficient measures the damage to the entire routing and switching network during a fault / attack event. A traffic collection program is deployed on each terminal node to collect the traffic received by the node per unit time. The overall network throughput is calculated before and during the fault / attack event, and the overall network throughput degradation coefficient is calculated using the following formula:

[0081]

[0082] Where i traverses all n terminal nodes in the entire network, F i and F′ i These are the first and second traffic flows (in Mb) acquired within time T before and during a fault / attack event, respectively. N If the value is close to 1, the failure / attack event has little impact on the routing and switching network; otherwise, it is much smaller. N The higher the value, the more severely the overall transmission capacity of the entire network is damaged during a failure / attack.

[0083] The network-wide performance testing subunit is used to count the number of paths in the current routing and switching system, and to perform ping probing on each path. Both Linux and Windows provide the ping command, which can perform connectivity probing and latency calculation on selected targets, and count the first latency and the second latency before and during the fault / attack event. Based on the number of paths, the first latency and the second latency determine the full path throughput degradation coefficient.

[0084] In this embodiment of the invention, the network-wide path performance degradation coefficient is used to roughly assess the latency performance of the entire network. All n terminal nodes in the network correspond to... Each path is pinged using nodes at both ends. Before and during a failure / attack event, latency is measured based on ping probing along each path. The overall network path performance degradation coefficient is calculated using the following formula:

[0085]

[0086] Where i iterates through the selected m paths, and Delay... i and Delay′ i These are the first and second delays (in milliseconds, ping timeout is uniformly recorded as 1ms) before and during the fault / attack event, respectively. A constant value is added to the first and second delays to avoid invalid data where the numerator or denominator is 0. p If the value is close to 1, the failure / attack event has little impact on the overall network latency performance; otherwise, it has little impact. l A larger value indicates that the latency across the entire network is significantly affected by faults / attacks.

[0087] The critical link performance testing subunit is used to select a third preset number of high betweenness links, perform ping probing on the third preset number of high betweenness links, and count the first average latency and the second average latency before and during the failure / attack event. Based on the third preset number, the first average latency and the second average latency, the critical link performance degradation coefficient is determined.

[0088] In this embodiment of the invention, the critical link performance degradation coefficient is used to further locate vulnerable links. n high betweenness numbers links are selected and their performance is measured. The connectivity and latency of the links are obtained by pinging each other between the nodes at both ends of the links. Measurements are performed before and during the failure / attack event, and the selected critical link performance degradation coefficient is calculated using the following formula:

[0089]

[0090] In this context, j iterates through the selected n key links, and Delay... j and Delay′ j These are the first and second average delays (in milliseconds, ping timeouts are uniformly recorded as 100ms) of the j-th link before and during the failure / attack event, respectively. A constant value is added to the forward and reverse average delays to avoid invalid data where the numerator or denominator is 0. l If the value is close to 1, the failure / attack event has little impact on the critical link; otherwise, it has little impact. i The higher the value, the more severely the selected set of critical links is damaged in a failure / attack event.

[0091] The paralyzed node location subunit is used to report process shutdown / restart events to the control terminal when the current node's BGP and OSPF processes are shut down based on the no BGP and no OSPF instructions, and when the routing process is restarted, so as to locate the paralyzed node.

[0092] The second evaluation unit is used to evaluate the traffic forwarding function and traffic forwarding performance, and to determine cascading failure; further, the second evaluation unit includes: a traffic forwarding function evaluation subunit, a traffic forwarding performance evaluation subunit, and a cascading failure determination subunit.

[0093] The traffic forwarding function evaluation subunit is used to obtain the network throughput degradation coefficient. If the network throughput degradation coefficient is within a first preset range, it is determined that the current network traffic forwarding function is normal. If the network throughput coefficient is within a second preset range, the traffic forwarding performance evaluation subunit is executed. The first preset range and the second preset range are set based on experience or specific circumstances, and are not specifically limited in this embodiment of the invention.

[0094] The traffic forwarding performance evaluation subunit is used to obtain the network path performance degradation coefficient. If the network path performance degradation coefficient is within the second preset range, the network fault is determined to be local and the cascading failure determination subunit is executed. Otherwise, the network fault is determined to be global.

[0095] The cascading failure determination subunit is used for all routing node process shutdown / restart events. If the number of process shutdown / restart events exceeds a preset threshold, a cascading failure is determined to have occurred; otherwise, no cascading failure has occurred.

[0096] In this embodiment of the invention, the specific processing procedures of the target selection unit, the first evaluation unit, and the second evaluation unit are illustrated with examples. Assuming the first preset interval is 1.0-1.2 and the second preset interval is 1.2-3.0, the specific processing procedures are as follows: Figure 2 As shown:

[0097] First, attack targets are selected, including high betweenness nodes and high betweenness links. Specifically, a node failure route update storm is performed on the high betweenness nodes, and a link failure congestion attack is performed on the high betweenness links. Traffic forwarding functionality is evaluated during the above process, and the overall network throughput degradation coefficient (Dec) is measured. N , such as Dec N If the value is <1.2 and very close to 1 (1.0-1.2), the overall network function is only slightly impaired, and it can be determined that the tested network can withstand the current fault / attack well. In this case, there is no need to perform traffic forwarding performance evaluation and cascading failure determination; if Dec N In the range of 1.3-3.0, the overall network function is moderately impaired. To determine this moderate impairment, a traffic forwarding performance assessment is needed to further determine the type of impairment; for example, Dec... N If the value is greater than 3, the overall network function is severely damaged. It can be determined that the tested network cannot withstand the current fault / attack, and there is a high probability that a large-scale cascading failure has occurred. It is necessary to perform a cascading failure judgment and conduct further verification.

[0098] Furthermore, regarding traffic forwarding performance evaluation, applicable to Dec NWithin the range of 1.3-3.0, this is used to further determine whether the overall network functionality impairment is due to a localized problem or a systemic issue. The entire network path performance degradation coefficient (Dec) is measured. N , such as Dec N If the value is very close to 1 (below 1.2), it can be determined that the overall network function impairment is due to the failure of a small number of backbone links / nodes or a small number of critical routing errors, and the network problem is localized. Conversely, if the value is much higher, it can be determined that the overall network function impairment is due to excessive network load, and the network problem is global. If the network problem is localized, cascading failure determination is required to further determine the nature of the local problem. If the number of times the routing process is shut down / restarted is less than a given first threshold, it means that the number of times the routing process is shut down / restarted is extremely low, and the damage is determined to be caused by the failure of a small number of links / nodes. If the number of times the routing process is shut down / restarted is greater than a given second threshold, it means that the number of times the routing process is shut down / restarted is frequent, and a layout cascading failure is determined. The first and second thresholds can be set based on experience or specific circumstances, and are not specifically limited in this embodiment of the invention.

[0099] Furthermore, regarding the determination of cascading failures, the number of shutdown / restart events of all routing nodes is counted. If the number of shutdown / restart events of the routing process is greater than a given second threshold, it indicates that the number of shutdown / restart events of the routing process is frequent, and it is determined that a cascading failure has occurred in the network, and the entire network or a part of it is in a state of oscillation. Conversely, if the number of shutdown / restart events is less than a given threshold, then there is no cascading failure phenomenon, and the network function impairment may be due to backbone link / node failure.

[0100] In practical applications, the frequency of specific events is generally judged in conjunction with network performance. For example, if shutdown / restart events are so frequent that they severely impact network throughput to an unacceptable level, the network can be judged as "failed." In this case, the shutdown / restart events are obviously "frequent." Conversely, if they are infrequent, they are acceptable. Furthermore, in practice, the judgment of cascading failures generally requires a joint analysis of routing process shutdown / restart events and throughput changes. Analyzing a single phenomenon may lead to incorrect conclusions. For example, a CXPST attack, even if it fails to trigger routing oscillations, may still cause a significant drop in throughput due to link congestion. Conversely, repeated restarts of routing processes on certain routing nodes may not necessarily cause global damage and cannot be considered a "failure" for the network as a whole.

[0101] The vulnerability location unit is used to locate vulnerability points in routing updates and congestion. Vulnerability location in the routing and switching system serves two purposes: first, when a cascading failure occurs in the routing and switching network, leading to overall paralysis / degradation, it locates the "first domino to fall," identifying the vulnerability points in the network that need reinforcement; second, when the routing and switching network has not experienced a cascading failure and can still perform traffic transmission functions, it analyzes which nodes / links are in a critical state, providing a basis for assessing the network's resilience to higher-intensity attacks. Therefore, the vulnerability location unit includes: a routing update vulnerability testing subunit and a congestion vulnerability testing subunit, wherein...

[0102] The routing update vulnerability testing subunit is used to perform attack testing and cascading failure judgment. If there is no cascading failure, a list of paralyzed nodes is obtained. If there is a cascading failure, the routing update tolerance threshold of the paralyzed node is increased by a preset ratio, and the attack test is repeated until there is no cascading failure, and a list of paralyzed nodes is obtained. The nodes in the list of paralyzed nodes are vulnerable nodes to cascading failure. The preset ratio is set based on experience or specific circumstances, and is not specifically limited in this embodiment of the invention.

[0103] Further testing procedures for the aforementioned route update vulnerability testing unit are as follows: Figure 3 As shown, preferably, the preset ratio is 10%. During the attack test, a cascading failure determination is performed. In the event of a cascading failure, the specific processing procedure is as follows:

[0104] ① Locate the paralyzed nodes and obtain a list of routing nodes that have experienced routing process shutdown / restart events.

[0105] ②Increase the routing update threshold of the routing nodes in step ① by 10%.

[0106] ③ Re-perform the attack test and obtain the list of routing nodes that experienced routing process shutdown / restart events once again.

[0107] ④ Return to step ② until cascading failures no longer occur.

[0108] In the absence of cascading failures, the last obtained list of routing nodes that experienced routing process shutdown / restart events represents vulnerable nodes that are sensitive to cascading failures.

[0109] The congestion vulnerability location subunit is used to locate congested links based on critical link performance testing, and to identify the congested links as congestion vulnerability points.

[0110] Furthermore, the cascading failure vulnerability testing system proposed in this invention, compared with traditional methods, has the following characteristics: ① It can be deployed in a virtualized environment, facilitating the simulation of large-scale routing and switching networks, and the performance of routing nodes and links is configurable; ② It can simulate two damaging factors, routing update storms and traffic redirection, and their effects on routing nodes; ③ It can be deployed based on mainstream network test ranges, and when necessary, it can be connected to physical devices or simulators like GNS3, using multiple methods to jointly carry out testing tasks.

[0111] This invention discloses a cascading failure vulnerability testing system, comprising: a routing switching module, a fault / attack event simulation module, and a vulnerability testing module. The routing switching module is used to set the routing link bandwidth scaling of the current routing switching system and reconstruct the functions of routing nodes. The fault / attack event simulation module is used to simulate the current routing switching system based on preset fault / attack events after reconstruction. The vulnerability testing module is used to test and evaluate the cascading failure vulnerability of the current routing switching system after simulation. In the above system, routing switching, fault / attack simulation, and vulnerability testing are all performed in a virtualized environment. The virtual environment facilitates the construction of large-scale networks and allows for the reproduction of large-scale routing nodes.

[0112] In this embodiment of the invention, based on the aforementioned testing system, a cascading failure vulnerability testing method is also provided. The execution flow of the testing method is as follows: Figure 4 As shown, the steps include:

[0113] S101. Pre-set the bandwidth reduction ratio of the routing links in the current routing and switching system;

[0114] In this embodiment of the invention, the process of determining the routing link bandwidth scaling ratio is the same as the processing process of the routing link bandwidth scaling ratio unit in the routing switching module of the cascade failure vulnerability test system, and will not be described again here.

[0115] S102. After the setup is complete, refactor the routing node functionality for both route update storm and traffic redirection.

[0116] In this embodiment of the invention, the routing node function reconstruction process is the same as the routing node function reconstruction unit of the routing switching module in the cascading failure vulnerability test system, and will not be described again here.

[0117] S103. After reconstruction, select a preset fault / attack event and simulate the current routing and switching system based on the preset fault / attack event;

[0118] In this embodiment of the invention, the process of simulating the current routing and switching system for preset fault / attack events is the same as the process of the fault / attack event simulation module in the cascaded failure vulnerability attack system, and will not be repeated here. The preset faults include: node failure, link failure, route update and congestion attack.

[0119] After the simulation in S104 is completed, the vulnerability of the current routing and switching system to cascading failures is tested and evaluated.

[0120] In this embodiment of the invention, the process for testing and evaluating the cascading failure vulnerability of the current routing and switching system is the same as the process for the vulnerability testing module in the cascading failure vulnerability attack system, and will not be described again here.

[0121] This invention discloses a method for testing cascading failure vulnerability, comprising: pre-setting the routing link bandwidth scaling ratio in the current routing and switching system; after setting, reconstructing the routing node functions for routing update storms and traffic redirection respectively; after reconstruction, selecting preset fault / attack events, and simulating the current routing and switching system based on the preset fault / attack events; after simulation, testing and evaluating the cascading failure vulnerability of the current routing and switching system. In the above method, setting the routing link bandwidth scaling ratio, reconstructing the routing node functions, simulating faults / attacks, and testing vulnerability are all based on a virtualized environment. The virtual environment facilitates the construction of large-scale networks and allows for the reproduction of large-scale routing nodes.

[0122] Based on the above-described cascading failure vulnerability testing method, this embodiment of the invention provides a cascading failure vulnerability testing device, the structural block diagram of which is shown below. Figure 5 As shown, it includes:

[0123] The module consists of module 201, module 202, module 203, and module 204 for setting up and evaluating the system.

[0124] in,

[0125] The setting module 201 is used to preset the bandwidth reduction ratio of the routing links in the current routing switching system;

[0126] The reconstruction module 202 is used to reconstruct the routing node function for routing update storm and traffic redirection respectively after the settings are completed;

[0127] The simulation module 203 is used to select a preset fault / attack event after reconstruction is completed, and simulate the current routing and switching system based on the preset fault / attack event.

[0128] The test and evaluation module 204 is used to test and evaluate the vulnerability of the current routing and switching system to cascading failures after the simulation is completed.

[0129] This invention discloses a cascading failure vulnerability testing device, comprising: pre-setting the routing link bandwidth scaling ratio in the current routing and switching system; after setting, reconstructing the routing node functions for routing update storms and traffic redirection respectively; after reconstruction, selecting preset fault / attack events, and simulating the current routing and switching system based on the preset fault / attack events; after simulation, testing and evaluating the cascading failure vulnerability of the current routing and switching system. In the above device, setting the routing link bandwidth scaling ratio, reconstructing the routing node functions, simulating faults / attacks, and testing vulnerability are all performed in a virtualized environment. The virtual environment facilitates the construction of large-scale networks and allows for the reproduction of large-scale routing nodes.

[0130] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are fundamentally similar to method embodiments, the descriptions are relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0131] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0132] The cascading failure vulnerability testing system, method, and related apparatus provided by this invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A cascading failure vulnerability testing system, characterized in that, include: The module includes a routing and switching module, a fault / attack event simulation module, and a vulnerability testing module. The routing and switching module is used to set the bandwidth scaling ratio of the current routing and switching system and to reconstruct the functions of the routing nodes; The fault / attack event simulation module is used to simulate the current routing and switching system based on preset fault / attack events after the reconstruction is completed; The vulnerability testing module is used to test and evaluate the cascading failure vulnerability of the current routing and switching system after the simulation is completed; The vulnerability testing module includes a vulnerability location unit, which is used to locate routing update vulnerabilities and congestion vulnerabilities. The vulnerability location unit includes: a route update vulnerability testing subunit and a congestion vulnerability location subunit, wherein... The routing update vulnerability testing subunit is used to perform attack tests and determine cascading failures. If there are no cascading failures, a list of paralyzed nodes is obtained. If there are cascading failures, the routing update tolerance threshold of the paralyzed nodes is increased by a preset ratio and the attack test is repeated until there are no cascading failures and a list of paralyzed nodes is obtained. The nodes in the list of paralyzed nodes are vulnerable nodes to cascading failures. The congestion vulnerability location subunit is used to locate congested links based on critical link performance testing, and to identify the congested links as congestion vulnerability points. In cases of cascading failures, the attack test is repeated after increasing the routing update threshold of the paralyzed node by a preset percentage until no cascading failures exist. The specific process is as follows: ① Locate the paralyzed nodes and obtain a list of routing nodes that have experienced routing process shutdown / restart events; ②Increase the routing update threshold of the routing nodes in step ① by 10%; ③ Re-perform the attack test and obtain the list of routing nodes that experienced routing process shutdown / restart events again; ④ Return to step ② until cascading failures no longer occur.

2. The system according to claim 1, characterized in that, The routing switching module includes: a routing link bandwidth scaling unit and a routing node function reconfiguration unit, wherein... The routing link bandwidth reduction unit is used to reduce the routing bandwidth of the current routing switching system according to a preset reduction ratio; The routing node function reconfiguration unit is used to simulate the crash-restart characteristics when receiving update messages and the packet loss characteristics when the router is congested.

3. The system according to claim 1, characterized in that, The fault / attack event simulation module includes: a node failure unit, a link failure unit, a route update unit, and a congestion attack unit, wherein... The node failure unit is used to shut down the bgp and ospf processes of the first node by controlling the Agent on the first node based on the no bgp and no ospf instructions; The link failure unit is used to disconnect the currently pending link; The routing update unit is used to control the Agent on the second node to send a first preset number of packets to the neighboring nodes, causing the neighboring nodes and the remote nodes to fall into a routing update storm, resulting in the paralysis of the neighboring nodes and the remote nodes. The congestion attack unit is used to control the virtual node to inject a second preset number of packets into the target link, causing the target link to become congested, or to cause the routing session to be disconnected by causing the routing heartbeat packet to be lost.

4. The system according to claim 1, characterized in that, The vulnerability testing module further includes: an attack target selection unit, a first evaluation unit, and a second evaluation unit, wherein... The attack target selection unit is used to select a target node or target link as an attack target based on the type of fault / attack event. The first evaluation unit is used to measure the overall network throughput, overall network path performance, and critical link performance, and to locate the paralyzed nodes. The second evaluation unit is used to evaluate the traffic forwarding function and traffic forwarding performance, and to determine cascading failure.

5. The system according to claim 4, characterized in that, The attack target selection unit includes: a first attack target selection subunit and a second attack target selection subunit, wherein... The first attack target selection subunit is used to select target nodes whose betweenness reaches a third preset number threshold as attack targets when the fault / attack event is a node failure or a route update. The second attack target selection subunit is used to select a target link whose betweenness reaches a fourth preset threshold as the attack target when the fault / attack event is a link failure or a congestion attack.

6. The system according to claim 4, characterized in that, The first evaluation unit includes: a throughput measurement subunit, a network-wide path performance testing subunit, a critical link performance testing subunit, and a paralyzed node location subunit, wherein, The throughput measurement subunit is used to traverse all nodes in the current routing and switching system, and to count the first and second traffic at each point before and during a fault / attack event within a preset time period. Based on the preset time period, the first and second traffic are used to determine the overall network throughput degradation coefficient. The critical link performance testing subunit is used to select a third preset number of high betweenness links, perform ping probing on the third preset number of high betweenness links, and count the first average latency and the second average latency before and during the failure / attack event. Based on the third preset number, the first average latency and the second average latency, the critical link performance degradation coefficient is determined. The paralyzed node location subunit is used to report process shutdown / restart events to the control terminal when the current node's BGP and OSPF processes are shut down based on the no BGP and no OSPF instructions, and when the routing process is restarted, so as to locate the paralyzed node.

7. The system according to claim 4, characterized in that, The second evaluation unit includes: a traffic forwarding function evaluation subunit, a traffic forwarding performance evaluation subunit, and a cascading failure determination subunit, wherein, The traffic forwarding function evaluation subunit is used to obtain the network throughput degradation coefficient. If the network throughput degradation coefficient is within the first preset range, it is determined that the current network traffic forwarding function is normal. If the network throughput degradation coefficient is within the second preset range, the traffic forwarding performance evaluation subunit is executed. The traffic forwarding performance evaluation subunit is used to obtain the network path performance degradation coefficient. If the network path performance degradation coefficient is within the second preset range, it is determined that the network fault is local and the cascading failure determination subunit is executed. Otherwise, it is determined that the network fault is global. The cascading failure determination subunit is used for all routing node process shutdown / restart events. If the number of process shutdown / restart events exceeds a preset threshold, a cascading failure is determined to have occurred; otherwise, no cascading failure has occurred.

8. A method for testing cascading failure vulnerability, characterized in that, include: Pre-set the bandwidth scaling ratio of the routing links in the current routing and switching system; After the settings are completed, the routing node functions will be refactored for both route update storm and traffic redirection. After reconstruction, a preset fault / attack event is selected, and the current routing and switching system is simulated based on the preset fault / attack event; After the simulation is completed, the vulnerability of the current routing and switching system to cascading failures is tested and evaluated. The testing and evaluation of the vulnerability of the current routing and switching system to cascading failures includes: Attack tests are conducted, and cascading failures are determined. If no cascading failures occur, a list of paralyzed nodes is obtained. If cascading failures occur, the paralyzed node's route update tolerance threshold is increased by a preset percentage, and the attack test is repeated until no cascading failures occur, and a list of paralyzed nodes is obtained. The nodes in the list of paralyzed nodes are vulnerable nodes to cascading failures. The congestion vulnerability location subunit is used to locate congested links based on critical link performance testing, and to identify the congested links as congestion vulnerability points. In cases of cascading failures, the attack test is repeated after increasing the routing update threshold of the paralyzed node by a preset percentage until no cascading failures exist. The specific process is as follows: ① Locate the paralyzed nodes and obtain a list of routing nodes that have experienced routing process shutdown / restart events; ②Increase the routing update threshold of the routing nodes in step ① by 10%; ③ Re-perform the attack test and obtain the list of routing nodes that experienced routing process shutdown / restart events again; ④ Return to step ② until cascading failures no longer occur.

9. A testing device for cascading failure vulnerability, characterized in that, include: The configuration module is used to pre-set the bandwidth scaling ratio of the routing links in the current routing and switching system; The refactoring module is used to refactor the routing node functionality for both route update storms and traffic redirection after the settings are completed. The simulation module is used to select preset fault / attack events after reconstruction and simulate the current routing and switching system based on the preset fault / attack events; The test and evaluation module is used to test and evaluate the vulnerability of the current routing and switching system to cascading failures after the simulation is completed. The routing update vulnerability testing subunit is specifically used to perform attack tests and determine cascading failures. If there are no cascading failures, a list of paralyzed nodes is obtained. If there are cascading failures, the routing update tolerance threshold of the paralyzed nodes is increased by a preset ratio and the attack test is repeated until there are no cascading failures and a list of paralyzed nodes is obtained. The nodes in the list of paralyzed nodes are vulnerable nodes to cascading failures. The congestion vulnerability location subunit is used to locate congested links based on critical link performance testing, and to identify the congested links as congestion vulnerability points. In cases of cascading failures, the attack test is repeated after increasing the routing update threshold of the paralyzed node by a preset percentage until no cascading failures exist. The specific process is as follows: ① Locate the paralyzed nodes and obtain a list of routing nodes that have experienced routing process shutdown / restart events; ②Increase the routing update threshold of the routing nodes in step ① by 10%; ③ Re-perform the attack test and obtain the list of routing nodes that experienced routing process shutdown / restart events again; ④ Return to step ② until cascading failures no longer occur.