A Network Resilience Analysis Method and Device Based on Dynamic Fault Trees
By building a dynamic fault tree and using historical operation and maintenance data to calculate the probability and key importance of basic events, targeted maintenance strategies are generated, and the problem of low network elasticity analysis efficiency in the existing technology is solved, and the rapid evaluation and effective maintenance of the network system are achieved.
Patent Information
- Application Number
- CN202510412267.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The network elastic analysis method in the prior art has low evaluation efficiency and cannot realize abnormal detection and maintenance during the operation of the actual network system.
A network elasticity analysis method based on a dynamic fault tree is adopted. By taking the elastic target of the network system as the top event, a dynamic fault tree is built, and the probability and key importance of basic events are calculated using historical operation and maintenance data, and targeted maintenance strategies are generated to ensure the effectiveness of network elasticity.
It realizes rapid evaluation and maintenance of network elasticity, and can conduct quantitative analysis of network systems with timing dependence and cold spare parts logic, improves evaluation efficiency and effectively maintains them in actual network systems.
Smart Images

Figure CN119966836B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly to a network resilience analysis method and device based on a dynamic fault tree. Background Art
[0002] Network resilience refers to the ability of a network to prevent, withstand, recover, and adapt when faced with adverse conditions, stress, attacks, or compromised components, so as to maintain the stability of the system function and structure, achieve an orderly and effective response to major network security incidents, and ensure the stable operation of critical services.
[0003] In related technologies, the network resilience analysis method for a network system is to construct a network system and attack behaviors through network simulation, so as to evaluate the network resilience of the network system. However, this method not only has low evaluation efficiency, but also cannot achieve the abnormal detection and maintenance of network resilience during the operation of an actual network system. Summary of the Invention
[0004] The present invention provides a network resilience analysis method and device based on a dynamic fault tree. The technical solutions are as follows:
[0005] On the one hand, an embodiment of the present invention provides a network resilience analysis method based on a dynamic fault tree, the method comprising:
[0006] Taking the resilience target of the network system as the top event, and constructing a network resilience dynamic fault tree for the top event; the dynamic fault tree includes a top event, intermediate events, and basic events, and the logical relationships between events at different levels include at least one or more of a timing dependence gate, a cold spare part gate, and a basic logic gate;
[0007] Obtaining historical operation and maintenance data of the network system, and calculating the occurrence probability of each basic event according to the historical operation and maintenance data;
[0008] Based on the occurrence probability of the basic events and the logical relationships in the dynamic fault tree, calculating the occurrence probability of the top event and the critical importance of each basic event;
[0009] When the occurrence probability of the top event exceeds a preset threshold, based on the critical importance of each basic event and the occurrence probability of each basic event, analyzing to obtain the weak points of the network system, and generating a targeted network resilience maintenance strategy to ensure the effectiveness of network resilience.
[0010] In a possible implementation manner, the historical operation and maintenance data includes at least one of: device parameters, historical fault records, attack logs, and current operation data of the network system;
[0011] Calculating the occurrence probability of each basic event according to the historical operation and maintenance data, including:
[0012] For each basic event, the following operations are performed: Based on the occurrence threshold set in advance for this basic event, the proportion of the historical operation and maintenance data that meets the occurrence threshold is counted, and the occurrence probability of this basic event is determined based on this proportion.
[0013] In a possible implementation manner, the determining the occurrence probability of this basic event based on this proportion includes:
[0014] According to the logical relationship in the dynamic fault tree, it is determined whether this basic event depends on the occurrence of other basic events; if so, based on the proportion of this basic event and the occurrence probabilities of the other basic events it depends on, the occurrence probability of this basic event is calculated using conditional probability.
[0015] In a possible implementation manner, when the network system detects that the system function corresponding to the elastic target fails, it is used to start the system recovery process, and the starting method of the system recovery process is: first start the main system recovery process, and if the main system recovery fails, then start the redundant system recovery process;
[0016] At least two intermediate events of main system recovery failure and redundant system recovery failure are included in the dynamic fault tree, and the logical relationship between these two intermediate events and the top event is expressed through a cold spare part gate.
[0017] In a possible implementation manner, if the basic events of main system recovery failure include a detection delay with a dependent time sequence and main system data recovery, and the basic events of redundant system recovery failure include a redundant system startup failure and a redundant system data synchronization failure with an OR relationship, then:
[0018] The occurrence probability of the redundant system recovery failure is calculated through the following conditional probability formula:
[0019] P(redundant system recovery failure | main system recovery failure) = P(BE3) + P(BE4) - P(BE3 ∩ BE4)
[0020] Among them, P(redundant system recovery failure | main system recovery failure) is the occurrence probability of the redundant system recovery failure, P(BE3) is the occurrence probability of the redundant system startup failure, and P(BE4) is the occurrence probability of the redundant system data synchronization failure;
[0021] The occurrence probability of the main system recovery failure is calculated through the following formula:
[0022] P(main system recovery failure) = P(BE1)P(BE2)
[0023] Wherein, P(Main system recovery fails) is the occurrence probability of the main system recovery failure, P(BE1) is the occurrence probability of the detection delay, and P(BE2) is the occurrence probability of the main system data recovery failure;
[0024] The occurrence probability of the top event is calculated by the following formula:
[0025] P(Top) = P(Main system recovery fails) × P(Redundant system recovery fails | Main system recovery fails)
[0026] Wherein, P(Top) is the occurrence probability of the top event, and P(Main system recovery fails) is the occurrence probability of the main system recovery failure.
[0027] In a possible implementation, the critical importance of each basic event is calculated in the following way:
[0028] For each basic event, the following is performed:
[0029] Calculate the probability importance of this basic event; the probability importance is used to characterize the influence degree of the change in the occurrence probability of this basic event on the occurrence probability of the top event;
[0030] Calculate the occurrence probability of this basic event divided by the occurrence probability of the top event, and multiply the calculated quotient value by the probability importance of this basic event, and use the product as the critical importance of this basic event.
[0031] On the other hand, an embodiment of the present invention provides a network resilience analysis device based on a dynamic fault tree. The device includes:
[0032] A construction unit, configured to use the resilience target of the network system as the top event, and construct a network resilience dynamic fault tree for the top event; the dynamic fault tree includes a top event, intermediate events, and basic events, and the logical relationships between events at different levels include at least one or more of a timing dependence gate, a cold spare part gate, and a basic logic gate;
[0033] An acquisition unit, configured to acquire the historical operation and maintenance data of the network system;
[0034] A calculation unit, configured to calculate the occurrence probability of each basic event according to the historical operation and maintenance data; and configured to calculate the occurrence probability of the top event and the critical importance of each basic event based on the occurrence probability of the basic event and the logical relationships in the dynamic fault tree;
[0035] An analysis unit, configured to when the occurrence probability of the top event exceeds a preset threshold, analyze the weak points of the network system based on the critical importance of each basic event and the occurrence probability of each basic event, and generate a targeted network resilience maintenance strategy to ensure the effectiveness of network resilience.
[0036] On the other hand, a computer device is provided, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the steps of the above-mentioned network resilience analysis method based on a dynamic fault tree.
[0037] On the other hand, a computer-readable storage medium is provided. A computer program is stored in the storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned network resilience analysis method based on a dynamic fault tree are implemented.
[0038] On the other hand, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the above-mentioned network resilience analysis method based on a dynamic fault tree are implemented.
[0039] The technical solution provided by the present invention can at least bring the following beneficial effects:
[0040] By taking the resilience goal of the network system as the top event, constructing a network resilience dynamic fault tree for the top event, using the historical operation and maintenance data obtained during the operation of the network system to calculate the occurrence probability of each basic event, and then using the occurrence probability of the basic event and the logical relationship of the dynamic fault tree to calculate the occurrence probability of the top event and the critical importance of the basic event, the weak points of network resilience are analyzed. Finally, targeted maintenance strategies are generated for the weak points to ensure the effectiveness of network resilience. It can be seen that this solution can achieve rapid evaluation of network resilience during the operation of the network system, can quantitatively analyze and calculate the resilience failure behavior of a network system with time-dependent and cold spare part logics, and does not require network simulation. It not only has high evaluation efficiency but also can realize the maintenance of network resilience during the operation of the actual network system. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 is a flowchart of a network resilience analysis method based on a dynamic fault tree provided by an embodiment of the present invention;
[0043] Figure 2 is a schematic diagram of a network resilience dynamic fault tree of a network system provided by an embodiment of the present invention;
[0044] Figure 3 It is a structural diagram of a network resilience analysis device based on a dynamic fault tree provided by an embodiment of the present invention;
[0045] Figure 4 It is a hardware architecture diagram of a computer device provided by an embodiment of the present invention. Specific embodiments
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] Please refer to Figure 1 , a network resilience analysis method based on a dynamic fault tree provided by an embodiment of the present invention, the method includes:
[0048] Step 100, taking the resilience objective of the network system as the top event, and constructing a network resilience dynamic fault tree for the top event; the dynamic fault tree includes a top event, intermediate events, and basic events, and the logical relationships between events at different levels include at least one or more of a timing dependence gate, a cold spare part gate, and a basic logic gate;
[0049] Step 102: Obtain the historical operation and maintenance data of the network system, and calculate the occurrence probability of each basic event according to the historical operation and maintenance data;
[0050] Step 104: Based on the occurrence probability of the basic events and the logical relationships in the dynamic fault tree, calculate the occurrence probability of the top event and the critical importance of each basic event;
[0051] Step 106, analyze the weak points of the network system according to the occurrence probability of the top event and the critical importance of each basic event, and generate a targeted network resilience maintenance strategy to ensure the effectiveness of network resilience.
[0052] In the embodiments of the present invention, by taking the elasticity target of the network system as the top event, a network elasticity dynamic fault tree is constructed for the top event. Using the historical operation and maintenance data obtained during the operation of the network system, the occurrence probability of each basic event is calculated. Then, the occurrence probability of the top event and the critical importance of the basic events are calculated using the occurrence probability of the basic events and the logical relationship of the dynamic fault tree, thereby analyzing the weak points of network elasticity. Finally, a targeted maintenance strategy is generated for the weak points to ensure the effectiveness of network elasticity. It can be seen that this solution can achieve rapid evaluation of network elasticity during the operation of the network system, can quantitatively analyze and calculate the elastic failure behavior of a network system with time-dependent and cold spare part logic, and does not require network simulation. It not only has a high evaluation efficiency but also can realize the maintenance of network elasticity during the operation of the actual network system.
[0053] The following describes Figure 1 the execution methods of the respective steps shown.
[0054] First, for step 100, take the elasticity target of the network system as the top event, and construct a network elasticity dynamic fault tree for the top event.
[0055] In the embodiments of the present invention, for a network system, the elasticity targets it cares about can include multiple. Since the events associated with different elasticity targets are different, the elasticity target can be used as the top event to construct a corresponding network elasticity dynamic fault tree for the top event.
[0056] A fault tree is an inverted tree-shaped logical causal relationship diagram that uses event symbols, logical gate symbols, and transfer symbols to describe the causal relationship between various events in the system; among them, the input events of the logical gate are the "causes" of the output event, and the output event of the logical gate is the "result" of the input event.
[0057] In the embodiments of the present invention, when constructing a network elasticity dynamic fault tree for the top event, it specifically may include: determining the intermediate events that affect the network elasticity of the top event in the network system, and determining the logical relationship between the intermediate events and the top event, and determining the basic events that cause the network elasticity failure of the top event based on the logical relationship; generating a network elasticity fault tree according to the top event, intermediate events, basic events, and logical relationship. Among them, the logical relationship between different levels of events includes at least one or more of a time-dependent gate (PAND), a cold spare part gate (CSP), and a basic logic gate (AND / OR). Among them, the intermediate events and basic events are determined using different functional structure levels associated with the top event.
[0058] The method provided by the embodiments of the present invention can be applied to at least the following scenarios: when the network system detects that the system function corresponding to the elastic target fails, it is used to start the system recovery process, and the starting method of the system recovery process is: first start the main system recovery process, and if the main system recovery fails, then start the redundant system recovery process;
[0059] Then, the constructed dynamic fault tree includes at least two intermediate events: main system recovery failure and redundant system recovery failure, and the logical relationship between these two intermediate events and the top event is expressed through a cold spare gate. Please refer to Figure 2 for the schematic diagram of the dynamic fault tree constructed for this scenario. Among them, network service a1 is located in the critical asset A of the network system. The interruption of network service a1 in critical asset A (i.e., network elasticity failure) is used as the top event, and a network elasticity dynamic fault tree can be generated for this top event. According to Figure 2 It can be known that the intermediate events affecting the network elasticity of the top event include main system recovery failure and redundant system recovery failure; the basic events of main system recovery failure include detection delay and main system data recovery failure; the basic events of redundant system recovery failure include redundant system startup failure and redundant system data recovery failure. The redundant system and the main system are in a cold spare relationship, and the redundant system is activated to take over the critical business when the main system recovery fails.
[0060] Figure 2 In
[0061] In one embodiment, the definitions and logical relationship configurations of the events in the dynamic fault tree are as follows:
[0062] Definition of the top event: Interruption of network service a1 (response duration > 10 minutes).
[0063] Among them, the occurrence of the top event can be caused by a network attack or a system failure.
[0064] Dynamic logic gate configuration:
[0065] PAND gate (sequential dependence gate)
[0066] Input event 1: Detection delay (this event is represented by BE1, detection time > 5 minutes);
[0067] Input event 2: Main system data recovery failure (this event is represented by BE2, recovery time > 5 minutes);
[0068] Logical relationship: BE1 must occur before BE2.
[0069] CSP gate (cold spare gate)
[0070] Main event: The main system recovery fails (BE2 = 1);
[0071] Standby event: The redundant system recovery fails, that is, the redundant system fails to start (this event is represented by BE3) or the redundant system data synchronization fails (this event is represented by BE4);
[0072] Logical relationship: After the main system recovery fails, the redundant system recovery process is activated.
[0073] Then, for step 102, obtain the historical operation and maintenance data of the network system, and calculate the occurrence probability of each basic event according to the historical operation and maintenance data.
[0074] In an embodiment of the present invention, the historical operation and maintenance data may include at least one of device parameters, historical fault records, attack logs, and current operation data of the network system. The historical operation and maintenance data is the data actually generated by the network system, which can truly reflect the network resilience of the network system.
[0075] It should be noted that the network resilience analysis method in the embodiment of the present invention can be applied to the running network system. During the analysis process, only the historical operation and maintenance data of the network system needs to be obtained, without pausing the operation of the network system, and it will not interfere with the normal operation of the network system.
[0076] In an embodiment of the present invention, the occurrence probability of each basic event is calculated as follows: For each basic event, perform: Based on the occurrence threshold preset for the basic event, count the proportion of the historical operation and maintenance data that meets the occurrence threshold, and determine the occurrence probability of the basic event based on the proportion.
[0077] Among them, the proportion can be obtained by statistical methods. Taking Figure 2 the basic event of detection delay as an example, according to the above definition, its occurrence threshold is 5 minutes. Count the total number of detection events in the historical operation and maintenance data and the total number of detection events with a detection time exceeding 5 minutes. Determine the proportion of the total number of detection events with a detection time exceeding 5 minutes to the total number of detection events as the proportion that meets the occurrence threshold.
[0078] If the basic event is an independent event, then after obtaining the proportion for the basic event through statistics, the proportion can be directly determined as the occurrence probability of the basic event;
[0079] If the basic event is not an independent event and its occurrence is related to other basic events, then the proportion obtained through statistics for the basic event cannot be determined as the occurrence probability of the basic event.
[0080] Based on this, in the embodiments of the present invention, determining the occurrence probability of the basic event based on the proportion may include:
[0081] According to the logical relationship in the dynamic fault tree, determine whether the basic event depends on the occurrence of other basic events; if so, based on the proportion of the basic event and the occurrence probability of the other basic events it depends on, use conditional probability to calculate the occurrence probability of the basic event.
[0082] Next, for Figure 2 each basic event in, calculate the occurrence probability.
[0083] Detection delay (BE1): Statistically count the proportion of attack events where the detection time exceeds the occurrence threshold (5 minutes). For example, based on historical logs, it is statistically obtained that among 100 attacks, 10 have detection delays exceeding the occurrence threshold. Then the occurrence probability of this basic event is: P(BE1)=0.10;
[0084] Main system data recovery failure (BE2): Since the occurrence of this basic event BE2 has a temporal dependence on the detection delay BE1, conditional probability needs to be used for calculation; for example, according to historical attack logs, in the scenario where the detection delay (BE1) occurs, the main system recovery failure rate is 25% (10 out of 40 delayed attacks result in main system recovery failure), that is, the occurrence probability of this basic event is: P(BE2∣BE1)=0.25;
[0085] Redundant system startup failure (BE3): Statistically count the number of times the redundant system fails to start after the main system recovery fails. For example, among 10 primary - backup switches, 2 fail. Then the occurrence probability of this basic event is: P(BE3)=0.20;
[0086] Redundant system data synchronization failure (BE4): Statistically count the number of times data synchronization fails between the main system and the redundant system. For example, among 20 data synchronizations, 5 fail. Then the occurrence probability of this basic event is: P(BE4)=0.25.
[0087] Step 104, based on the occurrence probability of the basic event and the logical relationship in the dynamic fault tree, calculate the occurrence probability of the top event and the critical importance of each basic event.
[0088] In the embodiments of the present invention, in order to analyze whether the network resilience of the elastic target is effective, the occurrence probability of the top event can be used for analysis. After calculating the occurrence probability of the basic event, in order to calculate the occurrence probability of the top event, it is necessary to first calculate the occurrence probability of each intermediate event;
[0089] Continue to take Figure 2 the two intermediate events in the shown dynamic fault tree as an example.
[0090] For the intermediate event of the failure of the main system recovery, its occurrence probability is the product of the occurrence probabilities of two basic events (BE1 and BE2), that is, P(Failure of main system recovery) = P(BE1)P(BE2) = 0.1×0.25 = 0.025.
[0091] Among them, P(Failure of main system recovery) is the occurrence probability of the failure of the main system recovery, P(BE1) is the occurrence probability of detection delay, and P(BE2) is the occurrence probability of the failure of the main system data recovery.
[0092] For the intermediate event of the failure of the redundant system recovery, after the failure of the main system recovery (BE2 = 1), the redundant system recovery process is activated, and the occurrence probability of the failure of the redundant system recovery is jointly determined by the following basic events: the failure of the redundant system startup (BE3) and the failure of the redundant system data synchronization (BE4).
[0093] Then, the occurrence probability of the failure of the redundant system recovery is calculated by the following conditional probability formula:
[0094] P(Failure of redundant system recovery | Failure of main system recovery) = P(BE3) + P(BE4) - P(BE3 ∩ BE4)
[0095] Among them, P(Failure of redundant system recovery | Failure of main system recovery) is the occurrence probability of the failure of the redundant system recovery, P(BE3) is the occurrence probability of the failure of the redundant system startup, and P(BE4) is the occurrence probability of the failure of the redundant system data synchronization.
[0096] Assume that BE3 (Failure of redundant system startup) and BE4 (Failure of redundant system data synchronization) are independent events, that is, there is no direct dependence between the two, then:
[0097] P(Failure of redundant system recovery | Failure of main system recovery) = P(BE3) + P(BE4) - P(BE3)P(BE4) = 0.20 + 0.25 - 0.20×0.25 = 0.40.
[0098] In the embodiment of the present invention, after calculating the occurrence probabilities of the two intermediate events, considering that among the two intermediate events, the failure of the main system recovery and the failure of the redundant system recovery are cold spare part logics, the core of the cold spare part logic is that the spare part will be activated only after the main part fails, but the success or failure of the standby event does not directly depend on the state of the main event, but depends on the running state of the standby event itself and the resource availability. That is:
[0099] Main event: Failure of main system recovery (BE2 = 1);
[0100] Standby event: The redundant system recovery fails, i.e., the redundant system startup fails (BE3) or the redundant system data synchronization fails (BE4);
[0101] Logical relationship: After the primary system recovery fails, the redundant system recovery process is activated. The probability of the redundant system recovery failure is determined by the joint probability P(redundant system recovery failure | primary system recovery failure) of BE3 and BE4.
[0102] The top event (service interruption) occurs when the primary system recovery fails and the redundant system recovery fails. Both need to occur simultaneously, and neither can be missing.
[0103] Then, the occurrence probability of the top event can be calculated through the following formula:
[0104] P(Top) = P(primary system recovery failure) × P(redundant system recovery failure | primary system recovery failure) = 0.025 × 0.40 = 0.01 (i.e., 1%)
[0105] Among them, P(Top) is the occurrence probability of the top event.
[0106] After calculating the occurrence probability of the top event, the critical importance of each basic event can be calculated using the occurrence probability of the top event.
[0107] In one implementation, the critical importance of each basic event is calculated as follows:
[0108] Calculate the probability importance of this basic event; the probability importance is used to characterize the influence degree of the change in the occurrence probability of this basic event on the occurrence probability of the top event;
[0109] Calculate the quotient of the occurrence probability of this basic event divided by the occurrence probability of the top event, and multiply the calculated quotient value by the probability importance of this basic event. Take the product as the critical importance of this basic event.
[0110] Among them, the probability importance of the i-th basic event is calculated through the following formula:
[0111] Calculate the probability importance I p (BE i ), that is, the influence of the probability change of the i-th basic event BE i on the top event probability. The calculation formula is:
[0112]
[0113] Among them, I p (BE i ) is the probability importance of the i-th basic event, and P(BE i ) is the occurrence probability of the i-th basic event.
[0114] Continuing with the dynamic fault tree shown in Figure 2 as an example, calculate the critical importance of each basic event. Among them, for the cold spare part design scenario, the calculation of probability importance needs to consider conditional probability. Then:
[0115] For the basic event BE1 of detection delay, its probability importance is:
[0116]
[0117] =P(BE2|BE1)×P(Redundant system recovery fails | Primary system recovery fails)
[0118] =0.25×0.40 = 0.10
[0119] For the basic event BE2 of primary system data recovery failure, its probability importance is:
[0120]
[0121] =P(BE1)×P(Redundant system recovery fails | Primary system recovery fails)
[0122] =0.10×0.40 = 0.04
[0123] For the basic event BE3 of redundant system startup failure, its probability importance is:
[0124]
[0125] =P(Primary system recovery fails)×(1 - P(BE4))
[0126] =0.025×0.75 = 0.01875
[0127] For the basic event BE4 of redundant system data synchronization failure, its probability importance is:
[0128]
[0129] =P(Primary system recovery fails)×(1 - P(BE3))
[0130] =0.025×0.80 = 0.02
[0131] Then, based on the probability importance, calculate the critical importance, that is, the degree of influence of the change in the occurrence probability of the basic event BE i on the occurrence probability of the top event, which represents the contribution of the basic event to the occurrence probability of the top event and is the basis for the priority ranking of the basic events.
[0132] In the embodiment of the present invention, the critical importance degree of the i-th basic event is calculated by using the following calculation formula:
[0133]
[0134] wherein, is the critical importance degree of the i-th basic event.
[0135] Then for Figure 2 each basic event shown, its critical importance degree is:
[0136] For the basic event BE1 of detection delay, its critical importance degree is:
[0137]
[0138] For the basic event BE2 of main system data recovery failure:
[0139] First, calculate the conditional probability: P(BE2∣BE1)=0.25P(BE2∣BE1)=0.25
[0140] Then the unconditional probability is: P(BE2)=P(BE1)×P(BE2∣BE1)=0.10×0.25=0.025
[0141] The critical importance degree is:
[0142]
[0143] For the basic event BE3 of redundant system startup failure, its critical importance degree is:
[0144]
[0145] For the basic event BE4 of redundant system data synchronization failure, its critical importance degree is:
[0146]
[0147] In this way, the critical importance degree of each basic event is obtained.
[0148] Step 106, when the occurrence probability of the top event exceeds the preset threshold, then based on the critical importance degree of each basic event and the occurrence probability of each basic event, analyze and obtain the weak points of the network system, and generate a targeted network resilience maintenance strategy to ensure the effectiveness of network resilience.
[0149] In the embodiments of the present invention, during the operation of the network system, the occurrence probability of the top event of the network system can be calculated by using the dynamic fault tree and historical operation and maintenance data. That is to say, if the occurrence probability of the top event determined based on the current historical operation and maintenance data is greater, it indicates that the network elasticity of the network system is worse and the network elasticity is more likely to fail. Therefore, a threshold can be preset in advance, and the occurrence probability of the top event is compared with the preset threshold to determine the vulnerability degree of the network elasticity. For example, assuming that the preset threshold is 0.008, since the calculated occurrence probability of the top event 0.01 > 0.008, therefore, the risk of network system elasticity failure should be concerned, and the weak points of network elasticity should be further analyzed.
[0150] Specifically, it is necessary to utilize the critical importance of each basic event and the occurrence probability of each basic event to analyze and obtain the weak points of the network system.
[0151] First, sort the critical importance of each basic event from large to small. Taking Figure 2 the critical importance of each basic event shown as an example, it can be known that I c (BE1) = 1.0, I c (BE2) = 0.10, I c (BE3) = 0.375, I c (BE4) = 0.50. Then the order of the critical importance of the basic events from large to small is:
[0152] I c (BE1) > I c (BE4) > I c (BE3) > I c (BE2)
[0153] Thus, it can be seen that in the embodiments of the present invention:
[0154] Detection delay (BE1) is the primary weak point and needs to be optimized first;
[0155] Secondly, it is the failure of redundant system data synchronization (BE4) and the failure of redundant system startup (BE3);
[0156] The critical importance of the failure of the main system data recovery (BE2) is the lowest, so the optimization priority is the lowest.
[0157] Then, after sorting according to the critical importance, whether the basic event with the maximum critical importance needs to be maintained can be evaluated in combination with its occurrence probability.
[0158] Specifically, the same maintenance threshold can be set for basic events, or corresponding maintenance thresholds can be set separately for different basic events. If the occurrence probability of a basic event exceeds its maintenance threshold, the basic event needs to be maintained.
[0159] Among them, the maintenance strategy is generated for reducing the occurrence probability of basic events, and different basic events adopt different maintenance strategies. The maintenance strategies include but are not limited to: preferentially allocating resources, increasing the maintenance frequency, and adding standby components.
[0160] It should be noted that if the occurrence probability of the top event can be reduced below the preset threshold after generating the maintenance strategy for some basic events, then the maintenance strategy does not need to be generated for all basic events.
[0161] Specifically, in S1, among several basic events whose occurrence probabilities exceed the maintenance threshold, determine the first basic event with the maximum critical importance, and generate a corresponding maintenance strategy for the first basic event;
[0162] In S2, after generating the current maintenance strategy, determine whether the current strategy can reduce the occurrence probability of the top event below the preset threshold. If so, end; otherwise, among the basic events whose occurrence probabilities exceed the maintenance threshold and for which no maintenance strategy has been generated, determine the second basic event with the maximum critical importance, generate a corresponding maintenance strategy for the second basic event, and continue to execute S2.
[0163] Through the analysis of the weak points of network resilience in the embodiments of the present invention, a maintenance strategy for reducing the probability of network resilience failure and improving the overall network resilience is generated, which can realize the maintenance of network resilience during the operation of the actual network system.
[0164] Please refer to Figure 3 , the embodiments of the present invention provide a network resilience analysis device based on a dynamic fault tree. The device includes:
[0165] A construction unit 300, configured to use the resilience target of the network system as the top event, and construct a network resilience dynamic fault tree for the top event; the dynamic fault tree includes a top event, intermediate events, and basic events, and the logical relationship between different levels of events includes at least one or more of a timing dependence gate, a cold spare part gate, and a basic logic gate;
[0166] An acquisition unit 302, configured to acquire the historical operation and maintenance data of the network system;
[0167] A calculation unit 304, configured to calculate the occurrence probability of each basic event according to the historical operation and maintenance data; and configured to calculate the occurrence probability of the top event and the critical importance of each basic event based on the occurrence probability of the basic event and the logical relationship in the dynamic fault tree;
[0168] An analysis unit 306, which is configured to, when the occurrence probability of the top event exceeds a preset threshold, analyze weak points of the network system based on the critical importance of each basic event and the occurrence probability of each basic event, and generate a targeted network resilience maintenance strategy to ensure the effectiveness of network resilience.
[0169] In an embodiment of the present invention, the historical operation and maintenance data includes at least one of device parameters, historical fault records, attack logs, and current operation data of the network system.
[0170] When the calculation unit calculates the occurrence probability of each basic event according to the historical operation and maintenance data, it specifically includes: for each basic event, execute: based on the occurrence threshold preset for the basic event, count the proportion of the historical operation and maintenance data that meets the occurrence threshold, and determine the occurrence probability of the basic event based on the proportion.
[0171] In an embodiment of the present invention, when the calculation unit determines the occurrence probability of the basic event based on the proportion, it specifically includes: according to the logical relationship in the dynamic fault tree, determine whether the basic event depends on the occurrence of other basic events; if so, calculate the occurrence probability of the basic event using conditional probability based on the proportion of the basic event and the occurrence probability of the other basic events it depends on.
[0172] In an embodiment of the present invention, when the network system detects that the system function corresponding to the resilience target fails, it is used to start the system recovery process, and the starting method of the system recovery process is: first start the main system recovery process, and if the main system recovery fails, then start the redundant system recovery process.
[0173] The dynamic fault tree includes at least two intermediate events of main system recovery failure and redundant system recovery failure, and the logical relationship between the two intermediate events and the top event is expressed through a cold spare part gate.
[0174] In an embodiment of the present invention, if the basic events of main system recovery failure include detection delay with dependent timing and main system data recovery, and the basic events of redundant system recovery failure include redundant system startup failure and redundant system data synchronization failure with an OR relationship, then:
[0175] The occurrence probability of the redundant system recovery failure is calculated through the following conditional probability formula:
[0176] P(redundant system recovery failure | main system recovery failure) = P(BE3) + P(BE4) - P(BE3 ∩ BE4)
[0177] Among them, P(redundant system recovery failure | primary system recovery failure) is the occurrence probability of redundant system recovery failure, P(BE3) is the occurrence probability of redundant system startup failure, and P(BE4) is the occurrence probability of redundant system data synchronization failure;
[0178] The occurrence probability of the primary system recovery failure is calculated by the following formula:
[0179] P(primary system recovery failure) = P(BE1)P(BE2)
[0180] Among them, P(primary system recovery failure) is the occurrence probability of the primary system recovery failure, P(BE1) is the occurrence probability of detection delay, and P(BE2) is the occurrence probability of primary system data recovery failure;
[0181] The occurrence probability of the top event is calculated by the following formula:
[0182] P(Top) = P(primary system recovery failure) × P(redundant system recovery failure | primary system recovery failure)
[0183] Among them, P(Top) is the occurrence probability of the top event, and P(primary system recovery failure) is the occurrence probability of the primary system recovery failure.
[0184] In an embodiment of the present invention, the critical importance of each basic event is calculated in the following manner:
[0185] For each basic event, the following operations are performed: calculate the probability importance of the basic event; the probability importance is used to characterize the influence degree of the change in the occurrence probability of the basic event on the occurrence probability of the top event; calculate the quotient of the occurrence probability of the basic event divided by the occurrence probability of the top event, and multiply the calculated quotient value by the probability importance of the basic event, and take the product as the critical importance of the basic event.
[0186] It should be noted that: the network elasticity analysis device based on the dynamic fault tree provided in the above embodiment is only illustrated by the division of the above functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the network elasticity analysis device of the network system provided in the above embodiment and the embodiment of the network elasticity analysis method of the network system belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0187] The embodiment of the present application also provides a computer device. Please refer to Figure 4, the computer device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the network resilience analysis method based on a dynamic fault tree provided by each of the above method embodiments.
[0188] An embodiment of the present application further provides a computer-readable storage medium, and at least one instruction, at least one program, a code set or an instruction set is stored on the computer-readable storage medium. The at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the network resilience analysis method based on a dynamic fault tree provided by each of the above method embodiments.
[0189] An embodiment of the present application further provides a computer program product, which includes a computer program. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the network resilience analysis method based on a dynamic fault tree described in any one of the above embodiments.
[0190] For the convenience of description, when describing the above system or device, various modules or units are described separately according to functions. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.
[0191] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.
[0192] Finally, it should also be noted that in this text, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the said element.
[0193] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A network resilience analysis method based on dynamic fault trees, characterized in that, For network resilience analysis of a running network system, the method includes: Taking the resilience goal of the network system as the top event, and constructing a network resilience dynamic fault tree for the top event; the dynamic fault tree includes a top event, intermediate events, and basic events, and the logical relationships between events at different levels include at least one or more of a timing dependence gate, a cold spare part gate, and a basic logic gate; Obtaining the historical operation and maintenance data of the network system, and calculating the occurrence probability of each basic event according to the historical operation and maintenance data; Based on the occurrence probability of the basic event and the logical relationships in the dynamic fault tree, calculating the occurrence probability of the top event and the critical importance of each basic event; When the occurrence probability of the top event exceeds a preset threshold, based on the critical importance of each basic event and the occurrence probability of each basic event, analyzing to obtain the weak points of the network system, and generating a targeted network resilience maintenance strategy to ensure the effectiveness of network resilience; The critical importance of each basic event is calculated in the following way: for each basic event, perform: calculating the probability importance of the basic event; the probability importance is used to characterize the influence degree of the change in the occurrence probability of the basic event on the occurrence probability of the top event; calculating the quotient of the occurrence probability of the basic event divided by the occurrence probability of the top event, and multiplying the calculated quotient value by the probability importance of the basic event, and taking the product as the critical importance of the basic event; The generating a targeted network resilience maintenance strategy to ensure the effectiveness of network resilience includes: S1. Determining the first basic event with the maximum critical importance among several basic events whose occurrence probabilities exceed the maintenance threshold, and generating a corresponding maintenance strategy for the first basic event; S2. After generating the current maintenance strategy, determining whether the occurrence probability of the top event can be reduced below the preset threshold at present. If so, end; otherwise, among several basic events whose occurrence probabilities exceed the maintenance threshold and among the basic events for which no maintenance strategy has been generated, determining the second basic event with the maximum critical importance, generating a corresponding maintenance strategy for the second basic event, and continuing to execute S2.
2. The method according to claim 1, wherein The historical operation and maintenance data includes at least one of equipment parameters, historical fault records, attack logs, and current operation data of the network system; Calculating the occurrence probability of each basic event according to the historical operation and maintenance data includes: For each basic event, perform: based on a preset occurrence threshold for the basic event, counting the proportion of the historical operation and maintenance data that meets the occurrence threshold, and determining the occurrence probability of the basic event based on the proportion.
3. The method according to claim 2, wherein The determining the occurrence probability of the basic event based on the proportion includes: According to the logical relationships in the dynamic fault tree, determining whether the basic event depends on the occurrence of other basic events; if so, calculating the occurrence probability of the basic event using conditional probability based on the proportion of the basic event and the occurrence probabilities of the other basic events it depends on.
4. The method according to claim 1, characterized in that When the network system detects the failure of the system function corresponding to the elastic target, it is used to start the system recovery process, and the starting method of the system recovery process is as follows: first start the main system recovery process, and if the main system recovery fails, then start the redundant system recovery process; The dynamic fault tree at least includes two intermediate events of the failure of the main system recovery and the failure of the redundant system recovery, and the logical relationship between these two intermediate events and the top event is expressed through a cold spare gate.
5. The method according to claim 4, wherein If the basic events of the failure of the main system recovery include the detection delay with dependent time sequence and the main system data recovery, and the basic events of the failure of the redundant system recovery include the failure of the redundant system startup with an OR relationship and the failure of the redundant system data synchronization, then: The occurrence probability of the failure of the redundant system recovery is calculated through the following conditional probability formula: P(Failure of redundant system recovery | Failure of main system recovery) = P(BE3) + P(BE4) - P(BE3 ∩ BE4) Among them, P(Failure of redundant system recovery | Failure of main system recovery) is the occurrence probability of the failure of the redundant system recovery, P(BE3) is the occurrence probability of the failure of the redundant system startup, and P(BE4) is the occurrence probability of the failure of the redundant system data synchronization; The occurrence probability of the failure of the main system recovery is calculated through the following formula: P(Failure of main system recovery) = P(BE1) × P(BE2) Among them, P(Failure of main system recovery) is the occurrence probability of the failure of the main system recovery, P(BE1) is the occurrence probability of the detection delay, and P(BE2) is the occurrence probability of the failure of the main system data recovery; The occurrence probability of the top event is calculated through the following formula: P(Top) = P(Failure of main system recovery) × P(Failure of redundant system recovery | Failure of main system recovery) Among them, P(Top) is the occurrence probability of the top event, and P(Failure of main system recovery) is the occurrence probability of the failure of the main system recovery.
6. A network resilience analysis device based on dynamic fault trees, characterized in that, The device includes: A construction unit, which is used to use the elastic target of the network system as the top event to construct a network elastic dynamic fault tree for the top event; the dynamic fault tree includes a top event, intermediate events and basic events, and the logical relationship between events at different levels includes at least one or more of a time sequence dependence gate, a cold spare gate and a basic logic gate; An acquisition unit, which is used to acquire the historical operation and maintenance data of the network system; A calculation unit, which is used to calculate the occurrence probability of each basic event according to the historical operation and maintenance data; and is used to calculate the occurrence probability of the top event and the critical importance of each basic event based on the occurrence probability of the basic event and the logical relationship in the dynamic fault tree; An analysis unit, which is used to when the occurrence probability of the top event exceeds the preset threshold, then analyze the weak points of the network system based on the critical importance of each basic event and the occurrence probability of each basic event, and generate a targeted network elastic maintenance strategy to ensure the effectiveness of network elasticity; The critical importance of each basic event is calculated as follows: For each basic event, the following operations are performed: Calculate the probability importance of the basic event; the probability importance is used to characterize the degree of influence of the change in the occurrence probability of the basic event on the occurrence probability of the top event; Calculate the quotient of the occurrence probability of the basic event divided by the occurrence probability of the top event, and multiply the calculated quotient by the probability importance of the basic event, and take the product as the critical importance of the basic event; When the analysis unit executes the generation of a targeted network resilience maintenance strategy to ensure the effectiveness of network resilience, it specifically includes: S1. Determine the first basic event with the maximum critical importance among several basic events whose occurrence probabilities exceed the maintenance threshold, and generate a corresponding maintenance strategy for the first basic event; S2. After generating the current maintenance strategy, determine whether the current strategy can reduce the occurrence probability of the top event to below the preset threshold. If so, end; otherwise, among the several basic events whose occurrence probabilities exceed the maintenance threshold and the basic events for which no maintenance strategy has been generated, determine the second basic event with the maximum critical importance, generate a corresponding maintenance strategy for the second basic event, and continue to execute S2.
7. A computer device, characterized in that, The computer device includes a memory and a processor. The memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the steps of the method according to any one of claims 1-5 above.
8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-5.
9. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Data center network equipment intelligent operation and maintenance system, method and terminal
CN112187514A
Quantum communication network availability analysis method based on Bayesian network model
CN113872801A