A deep learning-based network security situation evolution analysis method

By embedding a gradient correction mechanism in a deep reinforcement learning model, the potential game-theoretic motivations of network attacks are extracted, key nodes are identified, and defense strategies are adjusted. This solves the problem of insufficient adaptability to the dynamics and complexity of network attacks in existing technologies and achieves efficient network security defense.

CN121098617BActive Publication Date: 2026-04-14深圳宸元网信科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
深圳宸元网信科技有限公司
Filing Date
2025-10-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing network security protection technologies are insufficient to effectively cope with the dynamic, complex and strategic nature of network attacks, making it difficult to predict and prevent unknown or highly strategic attacks in a timely manner. Existing technologies also suffer from insufficient model prediction accuracy and strong data dependence, making it difficult to meet the security protection requirements of the increasingly complex network environment.

Method used

A deep reinforcement learning model with an embedded gradient correction mechanism is used to extract the potential game motivations of attack behavior, generate an initial attack-defense relationship graph, reverse-engineer the asymmetric distribution of node game gains and losses, identify key nodes and deduce threat propagation paths, and implement defense strategy adjustments based on the attack-defense game equilibrium conditions to dynamically optimize network defense measures.

Benefits of technology

It significantly improves the accuracy of identifying cybersecurity attacks and the targeting of defense strategies, reduces the waste of defense resources, enhances the synergy and response speed of the network defense system, and effectively suppresses the spread of attack paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121098617B_ABST
    Figure CN121098617B_ABST
Patent Text Reader

Abstract

The application discloses a network security situation evolution analysis method based on deep learning, and relates to the technical field of network security, which comprises the following steps: a deep reinforcement learning model with an embedded gradient correction mechanism is used to extract potential game causes of attack behaviors, determine potential threat nodes and generate an initial attack-defense correlation graph; based on the asymmetric distribution state of game losses and gains of the potential threat nodes in the attack-defense correlation graph, the disturbance effect of each node on the global game state is deduced reversely, and key nodes are identified; according to the game equilibrium disturbance effect caused by the key nodes, the game state transition conditions of the key nodes are limited, and a threat propagation path set is deduced; the defense strategy adjustment of the key nodes is implemented based on the attack-defense game equilibrium condition with the game state transition condition as a constraint; through the deep reinforcement learning model and game analysis, the network threat nodes can be accurately identified and the defense strategy can be dynamically optimized, so that the network security situation prediction and active defense capability can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and specifically to a method for analyzing the evolution of network security situation based on deep learning. Background Technology

[0002] With the rapid development of information technology and internet applications, the scale of various network information systems is expanding and their structures are becoming increasingly complex, leading to a significant increase in cybersecurity risks. Cyberattacks are gradually shifting from single, random incidents to multi-stage, coordinated processes with clear intent, exhibiting more complex and covert attack paths and methods, posing a significant challenge to existing cybersecurity protection technologies.

[0003] Traditional cybersecurity protection methods primarily rely on pre-defined rules and static detection mechanisms to identify and prevent cyberattacks. This approach typically only passively detects known attack characteristics, making it difficult to effectively predict and respond to unknown or highly strategic attacks. In recent years, in response to the dynamic, strategic, and complex nature of cyberattacks, a series of new methods have emerged aimed at enhancing cybersecurity situational awareness. These methods utilize game theory and machine learning techniques to delve deeper into the inherent patterns of attack behavior and identify potential threats in advance.

[0004] However, current game theory-based cybersecurity analysis techniques mostly employ static game models, which are ill-suited to address the dynamic changes in cyberattacks. Meanwhile, machine learning or deep learning-based cybersecurity techniques still suffer from limitations in practical applications, including insufficient model prediction accuracy, strong data dependence, and difficulty in accurately reflecting network structural characteristics. These existing technical issues restrict the practical effectiveness and application scope of cybersecurity situational analysis, making it difficult to meet the increasingly complex security protection requirements of the network environment.

[0005] Therefore, it is necessary to propose a more effective method for network security situation analysis to better adapt to the dynamic, complex, and strategic development trends of network attack behavior, improve the ability to proactively predict and prevent network security risks, and ensure the safe and stable operation of network information systems. Summary of the Invention

[0006] The purpose of this invention is to provide a network security situation evolution analysis system and method based on deep learning to solve the problems in the background art mentioned above.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] This invention provides a deep learning-based method for analyzing the evolution of network security situation, including:

[0009] Based on historical attack events and network topology, a deep reinforcement learning model with embedded gradient correction mechanism is used to extract the potential game-theoretic motivations of attack behavior, identify several potential threat nodes, and generate an initial attack-defense relationship graph.

[0010] Based on the asymmetric distribution of game gains and losses of potential threat nodes in the attack-defense relationship graph, the perturbation effect of each node on the global game state is deduced in reverse, and key nodes are identified.

[0011] Based on the game equilibrium disturbance effect caused by key nodes, the game state transition conditions of key nodes are constrained, and the set of threat propagation paths is derived.

[0012] By taking the game state transition conditions of key nodes in the threat propagation path set as constraints, and implementing defense strategy adjustments for the key nodes based on the attack-defense game equilibrium conditions, the attack-defense relationship graph is reconstructed.

[0013] Furthermore, the method for obtaining the deep reinforcement learning model with the embedded gradient correction mechanism includes:

[0014] Input the actual attack behavior sequence corresponding to the historical attack events, and the simulated attack behavior sequence generated by the network topology;

[0015] The real attack behavior sequence and the simulated attack behavior sequence are processed by a pre-trained deep reinforcement learning model, and the corresponding potential game motivations are output.

[0016] Based on the differences between the potential game motives corresponding to the actual attack behavior sequence and the simulated attack behavior sequence, the prediction bias in attack path selection is determined.

[0017] Based on the specific structure of the prediction bias, the gradient update direction of the deep reinforcement learning model is adjusted, and the deep reinforcement learning model with embedded gradient correction mechanism is output.

[0018] Furthermore, determining the prediction bias in attack path selection includes:

[0019] Establish separate attack path structures for real attack behavior sequences and simulated attack behavior sequences;

[0020] Compare the topological differences between the two attack path structures to determine the location distribution of the attack path structure differences;

[0021] Based on the determined location distribution, analyze the deviation characteristics of the attack path;

[0022] Based on the attack path deviation characteristics, the predicted deviation in attack path selection is output.

[0023] Furthermore, the reverse deduction of the perturbation effect of each node on the global game state includes:

[0024] Input the potential threat nodes in the initial attack-defense relationship graph, as well as the historical attack benefits and historical attack costs associated with each potential threat node;

[0025] Based on the spatial distribution structure of the historical attack gains and costs of each potential threat node in the attack-defense relationship graph, the asymmetric distribution state of the node game gains and losses is determined.

[0026] By utilizing the node propagation paths in the network topology, a reverse spatial diffusion analysis is performed on the asymmetric distribution of game gains and losses of each potential threat node to obtain the perturbation effect of each potential threat node on the game state of other nodes in the network.

[0027] Based on the spatial concentration of the influence of each node on the game state disturbance of other nodes in the network, the key nodes are output.

[0028] Furthermore, determining the asymmetric distribution state of the game's profit and loss at each node includes:

[0029] The attack benefits directly obtained by each potential threat node in historical attack events are calculated separately, as well as the attack costs paid to obtain those benefits.

[0030] Establish a spatial relationship structure between the attack benefits and attack costs on the corresponding nodes in the attack-defense relationship graph;

[0031] Analyze the proportional distribution characteristics between attack gains and attack costs in the aforementioned spatial correlation structure;

[0032] Based on the aforementioned proportional distribution characteristics, the asymmetric distribution state of the node game's profit and loss is determined.

[0033] Furthermore, the inverse spatial diffusion analysis of the asymmetric distribution of game gains and losses for each potential threat node includes:

[0034] By utilizing the physical connections between nodes in the network topology, a spatial propagation path for the asymmetric distribution of node game gains and losses is constructed.

[0035] The propagation trajectory of the asymmetric distribution of the game gains and losses of the nodes is analyzed by backtracking along the spatial propagation path.

[0036] Based on the described step-by-step propagation trajectory, the area of ​​disturbance to the game state of other nodes in the network by the potential threat node is determined;

[0037] Based on the perturbation region of the game state of each potential threat node, the perturbation effect of each potential threat node on the game state of other nodes in the network is output.

[0038] Furthermore, the derived set of threat propagation paths includes:

[0039] Input the identified key nodes and their location distribution in the attack-defense relationship graph;

[0040] Based on the location distribution of the key nodes, analyze the changes in the game payoffs and costs to neighboring nodes after the key nodes undergo a game state transition.

[0041] Based on the combination of changes in game payoffs and game costs, determine the triggering conditions for game state transitions at each key node;

[0042] By utilizing the game state transition triggering conditions of the key nodes, the possible attack behaviors between each key node and its neighboring nodes are deduced, and the set of threat propagation paths is output.

[0043] Furthermore, the adjustment of the defense strategy for the key nodes based on the equilibrium conditions of the attack-defense game includes:

[0044] Input the set of threat propagation paths and the triggering conditions for the game state transition of key nodes;

[0045] Using the aforementioned game state transition triggering conditions, the mutual constraint characteristics among the defense strategies of different key nodes are determined;

[0046] Based on the aforementioned mutual constraint characteristics, the combination of defense strategies is limited to ensure that the combination of defense strategies simultaneously satisfies the equilibrium conditions of the offensive and defensive game.

[0047] Based on the defensive strategy combination that satisfies the equilibrium condition of the attack and defense game, the node defense capabilities of the initial attack and defense association graph are reconfigured, and the reconstructed attack and defense association graph is output.

[0048] Furthermore, determining the mutual constraints among different critical node defense strategies includes:

[0049] Based on the network topology, we analyze the impact range of the implementation of defense strategies by different key nodes on the game state of adjacent nodes.

[0050] Based on the scope of influence of the game state, determine the spatial constraints between nodes after the implementation of the defense strategy;

[0051] Based on the spatially constrained positions, analyze the changes in node defense capability configuration caused by the combined implementation of different defense strategies;

[0052] Based on the changes in the node defense capability configuration, the mutual constraint characteristics between the different critical node defense strategies are output.

[0053] The technical effects and advantages provided by the present invention in the above technical solution are as follows:

[0054] This invention extracts the potential game-theoretic motivations of attack behavior by using a deep reinforcement learning model with an embedded gradient correction mechanism. This can effectively reduce the prediction bias between real attack behavior sequences and simulated attack behavior sequences, and avoid overfitting or underfitting problems in the model during training, thereby significantly improving the accuracy and stability of identifying the motivations of network security attack behavior.

[0055] This invention reverse-engineers the perturbation effect of each node on the global game state based on the asymmetric distribution of node game gains and losses. This enables efficient identification and accurate positioning of key nodes in the network, significantly improving the pertinence and effectiveness of network security defense strategy formulation, avoiding the blind allocation and waste of defense resources, and enhancing the synergy of the overall network defense system.

[0056] This invention defines the conditions for the transition of game states at key nodes and derives a set of threat propagation paths. By combining these with the equilibrium conditions of the attack and defense game, it dynamically adjusts defense strategies, enabling proactive optimization and real-time updates of network defense measures. This effectively suppresses the spread of attack paths, reduces the potential damage to network systems caused by attacks, and improves the overall efficiency and response speed of network security defense. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0058] Figure 1 This is a flowchart of a deep learning-based network security situation evolution analysis method according to the present invention. Detailed Implementation

[0059] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided to make the description of this application more complete and comprehensive, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The drawings are merely illustrative illustrations of this application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.

[0060] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more exemplary embodiments. Numerous specific details are provided in the following description to give a full understanding of the exemplary embodiments disclosed in this application. However, those skilled in the art will recognize that the technical solutions disclosed in this application can be practiced with one or more specific details omitted, or other methods, components, steps, etc., can be employed. In other instances, well-known structures, methods, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the disclosure of this application.

[0061] Example 1

[0062] like Figure 1 As shown in the figure, this embodiment discloses a method for analyzing the evolution of network security situation based on deep learning, including:

[0063] S101: Based on historical attack events and network topology, a deep reinforcement learning model with embedded gradient correction mechanism is used to extract the potential game-theoretic motivations of attack behavior, identify several potential threat nodes and generate an initial attack-defense relationship graph.

[0064] In a specific implementation, the method for obtaining the deep reinforcement learning model with embedded gradient correction mechanism includes:

[0065] Input the actual attack behavior sequence corresponding to the historical attack events, and the simulated attack behavior sequence generated by the network topology;

[0066] It is understandable that the actual attack behavior sequences corresponding to historical attack events are obtained through actual network security monitoring systems, and these sequences include, but are not limited to, the following information:

[0067] The starting node and target node of the attack;

[0068] The timing and duration of the attack;

[0069] Specific characteristic data such as the security vulnerabilities used during the attack, the type of attack tools, and the attack intensity level.

[0070] The simulated attack behavior sequence is generated by performing virtual attack simulation on a real network topology. The specific process is as follows:

[0071] Based on real network topology, a digital simulation model of network nodes and network connections is constructed.

[0072] In network simulation models, starting and target nodes are selected randomly or based on a certain probability, and possible attack paths are determined.

[0073] Each simulated attack is assigned characteristics similar to those of a real attack, including vulnerability type, attack intensity, and duration, to generate multiple sets of simulated attack sequences.

[0074] The real attack behavior sequence and the simulated attack behavior sequence are processed by a pre-trained deep reinforcement learning model, and the corresponding potential game motivations are output.

[0075] In a specific implementation, the deep reinforcement learning model is constructed based on a combination of deep neural networks (DNNs) and reinforcement learning (RL):

[0076] Using the attack behavior sequence as input, the temporal and spatial features in the attack behavior sequence are first extracted through the convolutional layer (CNN) and recurrent neural network (such as LSTM or GRU) layers of the neural network;

[0077] The extracted feature vectors are fed into the policy network of reinforcement learning, and the output is the potential game motivation vector corresponding to each attack sequence, which represents the attacker's policy preference, goal orientation and expected benefit-cost comprehensive evaluation when carrying out the attack.

[0078] For example, the potential game motivation vector M can be represented as: ;in, The value represents the intensity of the motivation in each dimension, and n is the number of motivation dimensions (such as expected attack reward, attack cost sensitivity, risk aversion, etc.). The motivation vector M is obtained by inputting the attack behavior sequence into the deep reinforcement learning model.

[0079] Based on the differences between the potential game motives corresponding to the actual attack behavior sequence and the simulated attack behavior sequence, the prediction bias in attack path selection is determined.

[0080] Specifically, determining the prediction bias in attack path selection includes:

[0081] Establish separate attack path structures for real attack behavior sequences and simulated attack behavior sequences;

[0082] It is understood that the attack path structure in this embodiment is represented by a graph model, specifically defined as follows:

[0083]

[0084] Among them, the node set Represents network nodes and edge sets. The weight set represents the transition relationships between nodes during an attack. This indicates the attack strength or transition probability on each side of the attack path.

[0085] Compare the topological differences between the two attack path structures to determine the location distribution of the attack path structure differences;

[0086] In a specific implementation, the quantitative analysis of topological differences is achieved by calculating the graph structure similarity between the actual attack path structure and the simulated attack path structure:

[0087] First, we define the topological difference metric for the attack path structure as follows: In the formula: Represents a node The degree (i.e., the number of edges directly connected to a node) in the actual attack path structure, and node Degree in the simulated attack path structure;

[0088] Furthermore, the topological differences are calculated for all nodes to determine the set of attack path structural difference location distributions: In the formula: The preset topological difference threshold is usually determined based on historical data and experience, for example, a value of 2 or greater.

[0089] Based on the determined location distribution, analyze the deviation characteristics of the attack path;

[0090] In practice, the deviation characteristic analysis of the attack path includes the following steps:

[0091] First, based on the set of differential location distributions Calculate the clustering coefficient and centrality indices (such as degree centrality and betweenness centrality) of the set of differing nodes to determine the network structure characteristics of the differing locations;

[0092] Specifically, the clustering coefficient of the set of differing nodes is defined as: In the formula: For nodes The local clustering coefficient, representing the degree of connectivity between a node's neighbors, is calculated using the following formula: In the formula: For nodes The number of edges between adjacent nodes For nodes The number of neighboring nodes;

[0093] Secondly, analyze the centrality index of the set of differing nodes to assess the importance of nodes in different locations to the attack path;

[0094] Finally, the results of the above clustering coefficient and centrality index are used as the bias characteristics of the attack path for subsequent bias correction.

[0095] Based on the attack path deviation characteristics, the predicted deviation in attack path selection is output;

[0096] In a specific implementation, a prediction bias feature vector is defined. for: In the formula, This refers to the aforementioned differential node clustering coefficient; The average degree centrality of the set of differing nodes; The average betweenness centrality of the set of differing nodes.

[0097] The output of the prediction bias is specifically the feature vector of that bias. This indicates the degree of difference between the path structure of a real attack and a simulated attack, serving as the basis for the next step of gradient correction.

[0098] Based on the specific structure of the prediction bias, the gradient update direction of the deep reinforcement learning model is adjusted, and the deep reinforcement learning model with embedded gradient correction mechanism is output.

[0099] In a specific implementation, the method for gradient correction in this step is as follows:

[0100] First, define the gradient correction objective function as follows: ;in, This is the objective function of the original deep reinforcement learning model; This is the gradient correction impact factor, which represents the degree of influence of prediction bias on gradient updates. It is determined based on experimental data, for example, with a value of 0.1. The penalty loss function introduced to account for prediction bias can, for example, be in the form of squared error: Then, the back propagation algorithm is used to update the gradient of the deep reinforcement learning model parameters φ. ;in, This is the learning rate parameter, determined based on experimental data, for example, a value of 0.001;

[0101] Through the above gradient update process, a deep reinforcement learning model with an embedded gradient correction mechanism is obtained, which can be used to more accurately identify the potential game motivations of attack behavior.

[0102] In a specific implementation, the method for identifying potential threat nodes and generating an initial attack-defense correlation graph specifically includes:

[0103] Based on the potential game motivation vector output by the model Calculate the comprehensive game motivation score for each node. The specific formula is as follows:

[0104]

[0105] In the formula: This represents the weighting factor of the k-th dimensional driver, which is determined based on experimental data.

[0106] Then, based on the comprehensive game motivation score of the node. Set threshold The conditions will be met. The nodes were identified as potential threat nodes;

[0107] Furthermore, for any two nodes among the potential threat nodes... and Calculate the cosine similarity between their potential game-driving vectors. If the cosine similarity exceeds a preset threshold Then at node and nodes Establish relationships between them;

[0108] Finally, based on the set of potential threat nodes and the relationships between nodes, the initial attack and defense relationship graph is generated. .

[0109] S102: Based on the asymmetric distribution of game gains and losses of potential threat nodes in the attack-defense relationship graph, reverse the deduction of the perturbation effect of each node on the global game state and identify key nodes.

[0110] In a specific implementation, the reverse deduction of the perturbation effect of each node on the global game state includes:

[0111] Input the potential threat nodes in the initial attack-defense relationship graph, as well as the historical attack benefits and historical attack costs associated with each potential threat node;

[0112] In this specific implementation, the potential threat nodes in this step are obtained through the initial attack-defense correlation graph generated in step S101. The historical attack gains and costs corresponding to each potential threat node include, but are not limited to:

[0113] Attack benefits: Based on the magnitude of the impact generated in actual attack events, a standardized scoring index system is established to determine the scoring indexes. The scoring rules can be determined by domain experts based on the attack event characteristic data (such as the degree of privilege escalation, the amount of information leakage, the duration of service interruption, or the scale of economic loss) recorded by the actual network security monitoring system.

[0114] Attack cost: Calculated based on the resources invested by the attacker to complete the attack. These resources include, but are not limited to, the time required for the attack, the complexity of the attack tools, and the level of risk detected during the attack. The specific scoring method is obtained based on the records and statistics of historical attack events.

[0115] The historical attack gains are denoted as set. The cost of historical attacks is denoted as set. ,in This represents the i-th potential threat node.

[0116] Based on the spatial distribution structure of the historical attack gains and costs of each potential threat node in the attack-defense relationship graph, the asymmetric distribution state of the node game gains and losses is determined.

[0117] Specifically, determining the asymmetric distribution state of the game's profit and loss at each node includes:

[0118] The attack benefits directly obtained by each potential threat node in historical attack events are calculated separately, as well as the attack costs paid to obtain those benefits.

[0119] In practice, this step calculates the total attack benefit for each potential threat node. and total attack value The formula is expressed as:

[0120]

[0121] Where n represents a node Number of historical attack incidents participated in.

[0122] Establish a spatial relationship structure between the attack benefits and attack costs on the corresponding nodes in the attack-defense relationship graph;

[0123] In a specific implementation, the attack gains and costs are embedded as node attributes into the initial attack-defense relationship graph, forming a relationship graph structure with node attributes: ;in, For a set of nodes, It is a set of connections between nodes. For a collection of node attributes, specifically defined as: ;

[0124] The above structure reflects the spatial distribution characteristics of attack gains and costs.

[0125] Analyze the proportional distribution characteristics between attack gains and attack costs in the aforementioned spatial correlation structure;

[0126] In practice, the ratio between attack gains and attack costs is defined as the profit-loss ratio. The calculation method is as follows: ;in, It is a very small positive number (such as 0.0001) used to avoid division by zero.

[0127] Based on the aforementioned proportional distribution characteristics, the asymmetric distribution state of the node game's profit and loss is determined;

[0128] In a specific implementation, a node with a higher payoff ratio indicates a higher relative cost to its attack, thus a stronger motivation to attack and a more aggressive node state; conversely, a node with a lower payoff ratio tends to be more conservative. Therefore, by performing threshold analysis on the payoff ratio, the asymmetric distribution state of node game payoffs is defined as follows:

[0129] like Then the node It is in a state of high-yield, low-cost attack;

[0130] like Then the node It is in a state of benefit-cost equilibrium;

[0131] like Then the node It is in a conservative state of low returns and high costs.

[0132] Among them, threshold and Based on historical attack data analysis and expert experience, values ​​can be set to 1.5 and 0.5 respectively.

[0133] By utilizing the node propagation paths in the network topology, a reverse spatial diffusion analysis is performed on the asymmetric distribution of game gains and losses of each potential threat node to obtain the perturbation effect of each potential threat node on the game state of other nodes in the network.

[0134] Specifically, the reverse spatial diffusion analysis of the asymmetric distribution of game gains and losses for each potential threat node includes:

[0135] By utilizing the physical connections between nodes in the network topology, a spatial propagation path for the asymmetric distribution of node game gains and losses is constructed.

[0136] In this specific implementation, the network topology of this step is represented in graphical form, namely: ;in, For a set of network nodes, This refers to the set of physical or logical connections between nodes in a network. Based on the network topology described above, nodes... The spatial propagation path is defined as starting from the node A set of connection paths to other nodes: .

[0137] The propagation trajectory of the asymmetric distribution of the game gains and losses of the nodes is analyzed by backtracking along the spatial propagation path.

[0138] In practice, this step employs a reverse analysis approach, starting from potential threat nodes and tracing back along the spatial propagation path to analyze the step-by-step propagation trajectory of the asymmetric distribution of game gains and losses among nodes. The specific method is as follows:

[0139] First, identify potential threat nodes. The initial disturbance impact value is the game payoff ratio of the node, that is: ;

[0140] Then, for each potential threat node As a source of propagation, it propagates upstream along the spatial propagation path, level by level. When a node... Upward neighbor node During backpropagation, nodes The disturbance impact value is updated by accumulating using the following formula:

[0141]

[0142] in, For nodes The cumulative disturbance impact value during the reverse backtracking analysis; For nodes The disturbance impact value; The attenuation coefficient, which ranges from 0 to 1, is determined by the network structure and propagation characteristics. For nodes To the node The difference in propagation levels in the backtracking path reflects the nodes The degree to which the impact of disturbances decays during upward backtracking. For example, if a node It is a node If the direct next-level neighbor node is the hierarchical difference... If they are separated by one layer, then And so on.

[0143] The above process traces back upwards along the spatial propagation path, recording and updating the perturbation impact value of the node each time it propagates to the next level, thereby realizing the step-by-step propagation trajectory analysis of the asymmetric distribution of the node's game gains and losses.

[0144] Based on the described step-by-step propagation trajectory, the area of ​​disturbance to the game state of other nodes in the network by the potential threat node is determined;

[0145] In specific implementation, the game state disturbance region is defined as the set of network nodes whose disturbance value caused by a specific potential threat node significantly exceeds a set threshold, i.e.: ;middle, The threshold for disturbance impact is obtained through statistical analysis of historical disturbance data, for example, it is set to 1.5 times the average disturbance impact value of the node.

[0146] Based on the perturbation region of the game state of each potential threat node, the perturbation effect of each potential threat node on the game state of other nodes in the network is output.

[0147] In practice, the output of this step is the set of disturbance effects on potential threat nodes, represented as: ;in, Represents the set of all potential threat nodes; Represents a node The specific area of ​​disturbance to other nodes in the network.

[0148] Based on the spatial concentration of the influence of each node on the game state disturbance of other nodes in the network, the key nodes are output.

[0149] In practice, the spatial concentration index of the impact of node disturbances is first defined as follows: ;in, Represents a node The number of nodes in the disturbed region; This represents the total number of nodes in the network.

[0150] Then, by ranking all potential threat nodes according to their spatial concentration index, a threshold is determined based on preset key nodes. (For example, ranking in the top 5% or top 10%), output the nodes that meet the conditions as the set of key nodes: .

[0151] S103: Based on the game equilibrium disturbance effect caused by key nodes, the game state transition conditions of key nodes are constrained, and the set of threat propagation paths is derived.

[0152] In a specific implementation, the derivation of the threat propagation path set includes:

[0153] Input the identified key nodes and their location distribution in the attack-defense relationship graph;

[0154] In a specific implementation, the key node set output in step S102 The node's location information is used as input for this step, where location information is defined as the node's position in the attack-defense relationship graph. Topological positional relationships, such as the set of adjacent nodes of a key node. Defined as: .

[0155] Based on the location distribution of the key nodes, analyze the changes in the game payoffs and costs to neighboring nodes after the key nodes undergo a game state transition.

[0156] In practice, the game state transitions at key nodes are divided into two types: attack state transitions and defense state transitions.

[0157] Attack state transition: Critical nodes transition from an inactive state to an active attack state;

[0158] Defensive state transition: Critical nodes switch from an active attack state to a controlled or defensive state.

[0159] First, define the changes in the game payoffs of neighboring nodes when a key node transitions to a new state. Changes in game costs for:

[0160]

[0161] In the formula: and These represent the nodes before the state transition of the critical nodes. The gains and losses in the game, and The game's gains and losses after the state transition;

[0162] By simulating and analyzing the state transition process of key nodes and recording changes in payoffs and costs, the game change values ​​of neighboring nodes are calculated.

[0163] Specifically, the simulation analysis method involves: using the deep reinforcement learning model obtained in step S101, and taking the potential game motivation vectors of key nodes as input, performing forward simulation; in the simulation, simulating multiple game processes after the key node transitions from its current state to an attack state or a defense state, statistically analyzing the changes in the game payoffs and costs of neighboring nodes during the simulation, and averaging the changes obtained from multiple simulations to obtain the changes in the game payoffs of neighboring nodes. Changes in game costs .

[0164] Based on the combination of changes in game payoffs and game costs, determine the triggering conditions for game state transitions at each key node;

[0165] In this specific implementation, the game state transition trigger condition is defined as a state transition that occurs when the change in the payoff-cost ratio exceeds a set threshold, as follows:

[0166] Define the change in the revenue-cost ratio of neighboring nodes as follows:

[0167]

[0168] In the formula: To prevent extremely small positive numbers with a denominator of zero.

[0169] Set a threshold based on the magnitude of change in the benefit-cost ratio. (e.g., 0.2), when When the game state transition is triggered, the conditions are met.

[0170] Therefore, a set of game state transition triggering conditions is defined for each key node: .

[0171] Using the game state transition triggering conditions of the key nodes, the possible attack behaviors between each key node and its neighboring nodes are deduced, and the set of threat propagation paths is output.

[0172] In practice, based on the set of game state transition triggering conditions Starting from key nodes, the possible propagation paths of attack behaviors are deduced step by step along the attack-defense relationship graph:

[0173] Specifically, with key nodes Starting from the node, for neighboring nodes that satisfy the state transition triggering condition. Conduct step-by-step propagation simulations to identify potential threat propagation paths. Specifically, it is expressed as:

[0174]

[0175] The output of the set of threat propagation paths for all critical nodes is as follows:

[0176]

[0177] Through the above steps, the conditions for the transition of game states at key nodes and the set of threat propagation paths are defined, providing an effective decision-making basis for subsequent defense strategies.

[0178] S104: Using the game state transition conditions of key nodes in the threat propagation path set as constraints, adjust the defense strategy of the key nodes based on the attack and defense game equilibrium conditions, and reconstruct the attack and defense association graph.

[0179] In a specific implementation, the step of adjusting the defense strategy for the key node based on the equilibrium conditions of the attack-defense game includes:

[0180] Input the set of threat propagation paths and the triggering conditions for the game state transition of key nodes;

[0181] In a specific implementation, the input for this step is the set of threat propagation paths output in step S103. The set of triggering conditions for state transitions in a game with key nodes This is used to determine the constraints for adjusting the defense strategy.

[0182] Using the aforementioned game state transition triggering conditions, the mutual constraint characteristics among the defense strategies of different key nodes are determined;

[0183] Specifically, determining the mutual constraints among different critical node defense strategies includes:

[0184] Based on the network topology, we analyze the impact range of the implementation of defense strategies by different key nodes on the game state of adjacent nodes.

[0185] In practice, key nodes are first set. After the defensive strategy is implemented, the range of game state changes of its neighboring nodes is defined as the state influence set:

[0186]

[0187] Among them, "significant state change" refers to a change in the payoff-cost ratio of node games exceeding a predetermined threshold. (e.g., 0.1).

[0188] Based on the scope of influence of the game state, determine the spatial constraints between nodes after the implementation of the defense strategy;

[0189] In practical implementation, spatially constrained locations are defined as nodes whose influence areas overlap with those of multiple key node defense strategies, i.e.: .

[0190] Based on the spatially constrained positions, analyze the changes in node defense capability configuration caused by the combined implementation of different defense strategies;

[0191] In practice, node defense capability is defined as a quantitative indicator of a node's comprehensive ability to resist attacks. The data for each indicator comes from historical security data recorded by the network security monitoring system and asset management system, specifically including: node protection coverage. Node security patch update status Monitoring intensity of nodes The formula for calculating the node's defense capability is as follows: (Three indicators)

[0192]

[0193] in, The weighting factor is determined by the actual network environment;

[0194] The initial defense capability of a node is calculated using the formula above. The defense capability of a node after the implementation of a combination of defense strategies is recalculated using the same formula, and the difference between the two is defined as the change in defense capability. .

[0195] Based on the changes in the node defense capability configuration, output the mutual constraint characteristics between the different critical node defense strategies;

[0196] In practice, if the combined implementation of defense strategies causes changes in the node's defense capabilities exceeding a preset threshold... (e.g., 20%) indicates the existence of significant mutual constraint characteristics. The output set of mutual constraint characteristics is:

[0197]

[0198] In the formula: As a key node, These are the affected nodes.

[0199] Based on the aforementioned mutual constraint characteristics, the combination of defense strategies is limited to ensure that the combination of defense strategies simultaneously satisfies the equilibrium conditions of the offensive and defensive game.

[0200] In practice, the equilibrium condition of the offensive-defense game is defined as follows: after the implementation of a defensive strategy, the changes in the defensive capabilities of each node should tend to stabilize, and the defensive capabilities of no node should decrease significantly. Combinations of defensive strategies with mutual constraints are screened to exclude those that would cause a decrease in the defensive capabilities of a node exceeding a threshold. The combination of defensive strategies. The limited combination of defensive strategies is as follows: .

[0201] Based on the defensive strategy combination that satisfies the equilibrium condition of the attack and defense game, the node defense capabilities of the initial attack and defense association graph are reconfigured, and the reconstructed attack and defense association graph is output.

[0202] In practice, the formula for reconfiguring node defense capabilities is as follows:

[0203]

[0204] in, This indicates the specific numerical improvement in a node's defensive capabilities resulting from the combination of defensive strategies.

[0205] After the defense capabilities are configured, the updated defense capability values ​​are embedded into the initial attack-defense relationship graph to form the final reconstructed attack-defense relationship graph: .

[0206] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired or wireless network. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0207] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A method for analyzing the evolution of network security situation based on deep learning, characterized in that, include: Based on historical attack events and network topology, a deep reinforcement learning model with embedded gradient correction mechanism is used to extract the potential game-theoretic motivations of attack behavior, identify several potential threat nodes, and generate an initial attack-defense relationship graph. Based on the asymmetric distribution of game gains and losses of potential threat nodes in the attack-defense relationship graph, the perturbation effect of each node on the global game state is deduced in reverse, and key nodes are identified; including: Input the potential threat nodes in the initial attack-defense relationship graph, as well as the historical attack benefits and historical attack costs associated with each potential threat node; Based on the spatial distribution structure of the historical attack gains and costs of each potential threat node in the attack-defense relationship graph, the asymmetric distribution state of the node game gains and losses is determined. By utilizing the node propagation paths in the network topology, a reverse spatial diffusion analysis is performed on the asymmetric distribution of game gains and losses of each potential threat node to obtain the perturbation impact of each potential threat node on the game states of other nodes in the network; including: By utilizing the physical connections between nodes in the network topology, a spatial propagation path for the asymmetric distribution of node game gains and losses is constructed. The propagation trajectory of the asymmetric distribution of the game gains and losses of the nodes is analyzed by backtracking along the spatial propagation path. Based on the described step-by-step propagation trajectory, the area of ​​disturbance to the game state of other nodes in the network by the potential threat node is determined; Based on the perturbation region of the game state of each potential threat node, output the perturbation effect of each potential threat node on the game state of other nodes in the network. Based on the spatial concentration of the influence of each node on the game state disturbance of other nodes in the network, the key nodes are output. Based on the game equilibrium disturbance effect caused by key nodes, the game state transition conditions of key nodes are constrained, and the set of threat propagation paths is derived; including: Input the identified key nodes and their location distribution in the attack-defense relationship graph; Based on the location distribution of the key nodes, analyze the changes in the game payoffs and costs to neighboring nodes after the key nodes undergo a game state transition. Based on the combination of changes in game payoffs and game costs, determine the triggering conditions for game state transitions at each key node; Using the game state transition triggering conditions of the key nodes, the possible attack behaviors between each key node and its neighboring nodes are deduced, and the set of threat propagation paths is output. By taking the game state transition conditions of key nodes in the threat propagation path set as constraints, and implementing defense strategy adjustments for the key nodes based on the attack-defense game equilibrium conditions, the attack-defense relationship graph is reconstructed.

2. The method according to claim 1, characterized in that, The method for obtaining the deep reinforcement learning model with embedded gradient correction mechanism includes: Input the actual attack behavior sequence corresponding to the historical attack events, and the simulated attack behavior sequence generated by the network topology; The real attack behavior sequence and the simulated attack behavior sequence are processed by a pre-trained deep reinforcement learning model, and the corresponding potential game motivations are output. Based on the differences between the potential game motives corresponding to the actual attack behavior sequence and the simulated attack behavior sequence, the prediction bias in attack path selection is determined. Based on the specific structure of the prediction bias, the gradient update direction of the deep reinforcement learning model is adjusted, and the deep reinforcement learning model with embedded gradient correction mechanism is output.

3. The method according to claim 2, characterized in that, The determination of prediction bias in attack path selection includes: Establish separate attack path structures for real attack behavior sequences and simulated attack behavior sequences; Compare the topological differences between the two attack path structures to determine the location distribution of the attack path structure differences; Based on the determined location distribution, analyze the deviation characteristics of the attack path; Based on the attack path deviation characteristics, the predicted deviation in attack path selection is output.

4. The method according to claim 3, characterized in that, The determination of the asymmetric distribution of the game's payoffs at each node includes: The attack benefits directly obtained by each potential threat node in historical attack events are calculated separately, as well as the attack costs paid to obtain those benefits. Establish a spatial relationship structure between the attack benefits and attack costs on the corresponding nodes in the attack-defense relationship graph; Analyze the proportional distribution characteristics between attack gains and attack costs in the aforementioned spatial correlation structure; Based on the aforementioned proportional distribution characteristics, the asymmetric distribution state of the node game's profit and loss is determined.

5. The method according to claim 4, characterized in that, Based on the equilibrium conditions of the offensive and defensive game, the defense strategy for the key nodes is adjusted, including: Input the set of threat propagation paths and the triggering conditions for the game state transition of key nodes; Using the aforementioned game state transition triggering conditions, the mutual constraint characteristics among the defense strategies of different key nodes are determined; Based on the aforementioned mutual constraint characteristics, the combination of defense strategies is limited to ensure that the combination of defense strategies simultaneously satisfies the equilibrium conditions of the offensive and defensive game. Based on the defensive strategy combination that satisfies the equilibrium condition of the attack and defense game, the node defense capabilities of the initial attack and defense association graph are reconfigured, and the reconstructed attack and defense association graph is output.

6. The method according to claim 5, characterized in that, The determination of the mutual constraints among different critical node defense strategies includes: Based on the network topology, we analyze the impact range of the implementation of defense strategies by different key nodes on the game state of adjacent nodes. Based on the scope of influence of the game state, determine the spatial constraints between nodes after the implementation of the defense strategy; Based on the spatially constrained positions, analyze the changes in node defense capability configuration caused by the combined implementation of different defense strategies; Based on the changes in the node defense capability configuration, the mutual constraint characteristics between the different critical node defense strategies are output.

Citation Information

Patent Citations

  • Network security evaluation system and method based on dynamic attack and defense game model

    CN119544307A