A multi-stage network defense method and system based on signal game

Through the multi-stage network defense method of signal game, defenders and attackers dynamically optimize strategies, which solves the problem of lack of initiative and flexibility in existing network defense methods and improves the robustness and effectiveness of network defense.

CN118827214BActive Publication Date: 2025-09-19STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411022268.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-09-19
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

Existing network defense methods lack initiative and flexibility, making it difficult to cope with complex and changeable attack behaviors. They also ignore the complexity of signal characteristics and the learning ability of attackers, resulting in difficulty in maintaining long-term defense effects.

Method used

A multi-stage network defense method based on signal game is adopted. The defender releases multiple signals to deceive the attacker. The attacker updates the posterior probability and combines the machine learning algorithm to determine the optimal attack strategy. The defender monitors the network status in real time and optimizes the signal release strategy. Both parties use dual optimization algorithms and dynamic game theory to optimize strategies, and combine complex network analysis methods to evaluate and adjust.

Benefits of technology

It achieves dynamic optimization of defense strategies and effective response to attack strategies, improves the initiative and flexibility of network defense, and enhances the robustness and effectiveness of the defense system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118827214B_ABST
    Figure CN118827214B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of network security defense technology, and specifically to a multi-stage network defense method and system based on signal game. The method includes: the defender releases multiple signals to deceive the attacker, and the attacker forms an initial prior judgment through detection; the attacker updates the posterior probability based on the defense signal and determines the optimal attack strategy; the defender monitors the network status in real time, optimizes the signal release strategy, and dynamically adjusts the signal frequency and type; the attacker and defender optimize their respective strategies using a dual optimization algorithm, and the defender adjusts the game strategy in real time using a reinforcement learning algorithm; the attacker updates the second posterior probability based on the profit and loss functions in the multi-stage game; the game is iterated and judged whether the termination condition is met, the results are analyzed and the strategy is adjusted, and the game ends or continues to the next round. Through the present invention, dynamic optimization of defense strategies and effective response to attack strategies are achieved, the initiative and flexibility of network defense are improved, and the overall robustness and effectiveness of the defense system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security defense technology, and in particular to a multi-stage network defense method and system based on signal game. Background Art

[0002] Currently, with the continuous escalation and diversification of cyberattacks, network security faces severe challenges. Traditional network defense methods, which are mostly passive, lack initiative and flexibility, making it difficult to effectively respond to complex and changing attack behaviors. As attackers use increasingly advanced techniques, the effectiveness of existing defense measures is gradually weakening. To improve the effectiveness of network defense, researchers have begun to focus on defense strategies based on game theory, among which signaling games are considered a promising approach.

[0003] Signaling games are a type of game model in which two parties make strategic choices and decisions by releasing and interpreting signals in an asymmetric information environment. Defenders protect network security by releasing various deceptive signals to induce attackers to misjudge their decisions. Simultaneously, attackers detect and analyze these signals to update their judgment of the defender's strategy and select corresponding attack strategies. This signaling game-based approach can improve the proactiveness and flexibility of defense, but it also faces a series of technical challenges. Existing research has explored the application of game theory in network defense, but most focus on single-stage games and lack in-depth research on multi-stage dynamic games. Furthermore, existing methods often overlook the complexity of signal characteristics and the attacker's learning ability in practical applications, making sustained defensive effectiveness difficult. Summary of the Invention

[0004] The present invention provides a multi-stage network defense method and system based on signal game, thereby effectively solving the problems pointed out in the background technology.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] A multi-stage network defense method based on signaling game, the method comprising:

[0007] The defender releases multiple signals to deceive the attacker, and the attacker forms an initial prior judgment through detection;

[0008] The attacker updates the posterior probability based on the defense signal and combines machine learning algorithms to analyze signal characteristics and historical attack data to determine the optimal attack strategy.

[0009] The defender monitors the network status in real time, optimizes the signal release strategy, and dynamically adjusts the signal frequency and type through topological structure analysis based on graph neural networks and multi-level game decision-making models;

[0010] The attacker and defender use dual optimization algorithms to optimize their respective strategies. The dual optimization algorithms used by the defender are simulated degradation algorithm and differential evolution algorithm, and the dual optimization algorithms used by the attacker are genetic algorithm and particle swarm optimization algorithm.

[0011] The attacker combines the profit and loss functions in the multi-stage game, uses dynamic game theory and fuzzy logic rules, and updates the second posterior probability;

[0012] The two parties in the game iterate and determine whether the termination conditions are met. They analyze the results and adjust the strategy to end or continue the next round of the game. After each round of the game, the two parties use complex network analysis methods and nonlinear dynamic system models to evaluate and adjust the game process and optimize the strategy for the next round of the game. The complex network analysis method includes:

[0013] Representing the elements in the game process as network nodes and the relationships between the network nodes as edges to construct a complex network model, wherein the elements include but are not limited to defense signals, attack strategies, and benefits;

[0014] Calculate basic network indicators, analyze the structural characteristics of the network, calculate network centrality indicators, and identify key nodes and important strategies in the game process;

[0015] Use community discovery algorithms to identify community structures in game networks and analyze the relationships between different strategies and signals.

[0016] Furthermore, the posterior probability calculation process adopts a solution method based on Monte Carlo simulation, including:

[0017] Set the initial signal strategy set and the attacker strategy set;

[0018] Generate a number of combinations of signals and attack strategies through random sampling, wherein the combinations simulate various possible situations in actual games;

[0019] For each sample, calculate the attacker and defender's profit functions according to their corresponding signal strategy and attack strategy;

[0020] Based on the generated samples and the calculated benefit function, the posterior probability distribution of the attacker and the defender is updated using the Bayesian formula.

[0021] Furthermore, the posterior probability is solved by a refined Bayesian Nash equilibrium, and the refined Bayesian Nash equilibrium satisfies the following conditions:

[0022]

[0023] Among them, a *(m) is the optimal attack strategy chosen by the attacker given the signal strategy m; For the attack strategy set S A Find the expression Maximizing attack strategy a h ; Is for all defenders of type t i Perform weighted summation; is the defender type t under a given signal strategy m i The posterior probability of U A (m,a h ,t i ) indicates that in the signal strategy m, attack strategy a h and defender type t i The attacker's profit under m * (t) is the optimal signaling strategy chosen by the defender given the defender type t; To find the expression in the signal strategy set M Maximize the signaling strategy m j ; is for all attack strategies a h Perform weighted summation; p D (a h ) Select attack strategy a for the attacker h The probability of U D (m j ,a h ,t) represents the signal strategy m j , attack strategy a h and the defender's payoff under defender type t.

[0024] Furthermore, at the beginning of each game, the attacker considers the defender's type misjudgment and calculates the probabilities of different types of misjudgments to obtain the misjudgment rate matrix, which is:

[0025]

[0026] Among them, ε is the misjudgment rate, k is the number of stages of the game, and ε ij is the probability that the attacker misjudges the defender's true type.

[0027] Furthermore, the losses caused by the attacker adopting different attack strategies when the defender type is misjudged are expressed by a loss matrix, which is:

[0028]

[0029] Among them, λ(t r ,a x) is the real defender type t r When the attacker chooses attack strategy a x The resulting losses, t1, t2…t n For all possible defender types, a1, a2…a n for all possible attack strategies.

[0030] Furthermore, the risk-return of the attacker is calculated based on the risks and benefits generated when the attacker misjudges the defender type as different and adopts different strategies. The formula is:

[0031]

[0032] Among them, Rr(t i ,a h ) means that when the attacker misjudges the defender as t i And take the attack strategy a h The risk return at time λ(t i ,a h ) means that when the defender type is t i And the attacker adopts attack strategy a h The loss at time ε(t i ,t g ) indicates that the attacker misjudges the defender type as t i The actual type is t g , where n is the number of defender types.

[0033] Furthermore, the optimization method is used to solve the game process between the two parties in order to maximize their own benefits. When the defender adopts a pure strategy, the attacker's benefit for each strategy is U A (m,a) is calculated as follows:

[0034]

[0035] Among them, U A (m,a) is the expected benefit of the attacker when he adopts strategy a for the defender’s signal selection strategy m, P(t i |m) is the attacker's response to defender type t when the defender chooses signal strategy m. i The posterior probability, U A (a,m,t i ) is when the attacker adopts strategy a, the defender chooses signal strategy m and the defender type is t i In the case of , the attacker's profit.

[0036] Furthermore, when the defender has a pure strategy, the expected payoff of each attacker's strategy forms the attacker's expected payoff matrix U A, the expression is:

[0037]

[0038] Among them, U A The matrix element of row i and column j is U A (m i ,a j ) is the defender's choice of signal strategy m i , the attacker chooses strategy a j The expected profit of the attacker in the case of .

[0039] Furthermore, the complex network analysis method includes:

[0040] Representing the elements in the game process as network nodes and the relationships between the network nodes as edges to construct a complex network model, wherein the elements include but are not limited to defense signals, attack strategies, and benefits;

[0041] Calculate basic network indicators, analyze the structural characteristics of the network, calculate network centrality indicators, and identify key nodes and important strategies in the game process;

[0042] Use community discovery algorithms to identify community structures in game networks and analyze the relationships between different strategies and signals.

[0043] A multi-stage network defense system based on signaling game, the system comprising:

[0044] Initial prior judgment formation module: The defender releases multiple signals to deceive the attacker, and the attacker forms an initial prior judgment through detection;

[0045] The optimal attack strategy determination module updates the posterior probability of the attacker based on the defense signal and combines the machine learning algorithm to analyze the signal characteristics and historical attack data to determine the optimal attack strategy;

[0046] Release strategy optimization and adjustment module: The defender monitors the network status in real time, optimizes the signal release strategy, and dynamically adjusts the signal frequency and type through topological structure analysis based on graph neural networks and multi-level game decision-making models;

[0047] In the two-party strategy optimization module, the attacker and defender use a dual optimization algorithm to optimize their respective strategies. The dual optimization algorithms used by the defender are simulated degradation algorithm and differential evolution algorithm, and the dual optimization algorithms used by the attacker are genetic algorithm and particle swarm optimization algorithm.

[0048] In the second posterior probability update module, the attacker combines the profit and loss functions in the multi-stage game, uses dynamic game theory and fuzzy logic rules to update the second posterior probability;

[0049] In the game termination judgment optimization module, both parties iterate and determine whether the termination conditions are met, analyze the results and adjust the strategy to end or continue the next round of the game. After each round of the game, both parties use complex network analysis methods and nonlinear dynamic system models to evaluate and adjust the game process and optimize the strategy for the next round of the game. The complex network analysis method includes:

[0050] Representing the elements in the game process as network nodes and the relationships between the network nodes as edges to construct a complex network model, wherein the elements include but are not limited to defense signals, attack strategies, and benefits;

[0051] Calculate basic network indicators, analyze the structural characteristics of the network, calculate network centrality indicators, and identify key nodes and important strategies in the game process;

[0052] Use community discovery algorithms to identify community structures in game networks and analyze the relationships between different strategies and signals.

[0053] The technical solution of the present invention can achieve the following technical effects:

[0054] Through the present invention, dynamic optimization of defense strategies and effective response to attack strategies are achieved, the initiative and flexibility of network defense are improved, and the overall robustness and effectiveness of the defense system are enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0056] Figure 1 The figure is a flowchart of a multi-stage network defense method based on signaling game;

[0057] Figure 2 A schematic diagram of the game tree model. DETAILED DESCRIPTION

[0058] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0060] Example 1

[0061] like Figure 1 As shown, the present invention provides a multi-stage network defense method based on signal game, the method comprising:

[0062] S1: The defender releases multiple signals to deceive the attacker, and the attacker forms an initial prior judgment through detection;

[0063] Specifically, defenders release a variety of signals designed to confuse and deceive attackers, making it difficult for them to accurately identify real network traffic and system behavior. These signals may be simulated network traffic, false service responses, or other false information. By releasing multiple signals, defenders attempt to increase attacker confusion and the likelihood of misjudgment, thereby enhancing network security. Attackers form an initial priori judgment of the target system by probing the network and analyzing system vulnerabilities. This judgment, based on their technical knowledge, experience, and detected information, helps them formulate subsequent attack strategies and action plans.

[0064] S2: The attacker updates the posterior probability based on the defense signal and combines machine learning algorithms to analyze signal characteristics and historical attack data to determine the optimal attack strategy.

[0065] First, the attacker uses network detection tools (such as traffic monitoring, intrusion detection systems, etc.) to detect the signals released by the defender. These signals may include forged traffic packets, deceptive network behaviors, and misleading security warnings. After detecting the signal, the attacker uses data preprocessing technology to perform preliminary processing on the signal to ensure the quality and availability of the data, including noise removal, standardization, and feature extraction. Next, the attacker combines the prior probability with the observed defense signal to update the posterior probability of the defender type. Through the Bayesian update method, the possibility of the signal under different defender types is calculated based on the defense signal, and then the posterior probability is updated. Then, the attacker extracts features from the defense signal and performs dimensionality reduction and optimization through feature engineering technology. The historical data is used to train machine learning models (such as decision trees, random forests, support vector machines, and neural networks, etc.), input signal features, and output the predicted probability of the defender type. The trained model is used to classify the currently detected signal and confirm The attacker determines the possible types and strategies of the defender. At the same time, the attacker collects and cleans historical attack data, including past attack types, targets, defense signals, and attack success rates. Through pattern recognition techniques (such as cluster analysis and association rule mining), common attack patterns and defense strategies are identified from historical data, the correlation between defense signals and attack success rates is analyzed, and which signals are likely to lead to attack success or failure are identified. On this basis, the attacker constructs a strategy optimization model based on the benefit function and the loss function, comprehensively considering the attack benefit, cost, and risk. The optimization goal is to maximize the attacker's expected benefit. The optimal attack strategy is searched through optimization algorithms (such as genetic algorithms, particle swarm optimization algorithms, or gradient descent, etc.), and the updated posterior probability and signal characteristics are input to output the optimal attack strategy. Finally, the attacker simulates and verifies the determined optimal attack strategy, evaluates its effectiveness and robustness in different scenarios, and ensures that the selected strategy has the expected effect in the actual network environment through multiple simulations.

[0066] S3: The defender monitors the network status in real time and optimizes the signal release strategy and dynamically adjusts the signal frequency and type through topological structure analysis based on graph neural networks and a multi-level game decision model.

[0067] First, the defender uses network monitoring tools (such as traffic analysis, intrusion detection systems, etc.) to collect network status data in real time. These data include network traffic, node connection status, attack events, and defense measures. Through continuous monitoring, the defender can obtain the network topology and dynamic changes in a timely manner; next, the defender uses graph neural network (GNN) to analyze the network topology. GNN is a deep learning model that can process graph structure data. Through the characteristics of nodes and the information of neighboring nodes, GNN can learn the high-dimensional representation of network topology. The defender inputs the network status data collected in real time into GNN, performs topological structure analysis, extracts the characteristics and connection relationships of network nodes, and forms a global representation of the network. Based on the topological structure analysis results of GNN, the defender constructs a multi-level game decision model. The model contains multiple levels, each level corresponds to a different game stage and strategy combination. At each level, the defender selects the optimal signal release strategy based on the current network topology and the attacker's strategy. With multiple levels of game decision-making, the defender can dynamically adjust the frequency and type of signals at different stages of the game to respond to the attacker's different strategies. To optimize the signal release strategy, the defender uses a reinforcement learning algorithm. Reinforcement learning is a method of learning the optimal strategy by interacting with the environment. In this process, the defender uses the signal release strategy as the reinforcement learning action, the network state as the state, and the attacker's response and the network's security status as the reward. By continuously interacting with the network environment, the defender can gradually optimize the signal release strategy and improve the defense effect. Finally, the defender dynamically adjusts the signal frequency and type based on the optimized strategy. The defender can adjust the signal sending interval, signal type, and signal strength according to the real-time status of the network and the attacker's behavior, thereby effectively deceiving the attacker and protecting the network security.

[0068] S4: The attacker and defender use dual optimization algorithms to optimize their respective strategies. The dual optimization algorithms used by the defender are simulated degradation algorithm and differential evolution algorithm, and the dual optimization algorithms used by the attacker are genetic algorithm and particle swarm optimization algorithm.

[0069] Specifically, the attacker and defender each construct a policy optimization model based on a reward and loss function, comprehensively considering benefits, costs, and risks. The attacker's goal is to maximize the attack success rate and revenue, while the defender's goal is to minimize the attack losses and protect network security. The policy optimization model takes the current game state and historical data as input and outputs the optimal policy combination. Next, the attacker and defender each optimize their policies using a dual optimization algorithm, which combines global and local optimization methods to improve the efficiency and accuracy of policy optimization.

[0070] For attackers, the dual optimization algorithm includes genetic algorithm and particle swarm optimization algorithm:

[0071] Genetic Algorithm: A genetic algorithm is a global optimization algorithm based on natural selection and heredity. The attacker first generates an initial population, where each individual represents an attack strategy. Through selection, crossover, and mutation operations, the genetic algorithm continuously generates new strategy combinations, gradually approaching the optimal solution.

[0072] Particle Swarm Optimization Algorithm: The particle swarm optimization algorithm simulates the foraging behavior of bird flocks. It quickly finds the optimal strategy through information sharing and collaborative search between individuals. The attacker uses the particle swarm optimization algorithm to search for the optimal attack strategy in the strategy space and gradually improves the attack effect through iterative updates.

[0073] For the defender, its dual optimization algorithm includes simulated annealing algorithm and differential evolution algorithm:

[0074] Simulated Annealing: This is a global optimization algorithm based on the thermodynamic annealing process. The defender first sets an initial temperature and randomly selects an initial policy from the policy space. Through the simulated annealing process, the defender explores the policy space as the temperature gradually decreases, gradually converging to the globally optimal policy.

[0075] Differential Evolution: This algorithm is an evolutionary optimization algorithm based on population differences. Defenders generate an initial population and, through differential mutation and crossover operations, generate new strategy combinations. The algorithm leverages individual differences within the population to continuously optimize defense strategies and improve effectiveness.

[0076] At the same time, the defender uses a reinforcement learning algorithm to adjust the game strategy in real time. Reinforcement learning is a method of learning optimal strategies through interaction with the environment. In this process, the defender uses the signal release strategy as the reinforcement learning action, the network state as the state, and the attacker's response and network security status as the reward. By continuously interacting with the network environment, the defender can gradually optimize the signal release strategy and improve the defense effect. The specific implementation steps are as follows:

[0077] Status monitoring and data collection: Defenders monitor network status in real time and collect attacker behavior data and network feedback.

[0078] Strategy Update: Based on the collected data, the defender uses reinforcement learning algorithms to update its strategy. Through algorithms such as Q-learning and Deep Q Network (DQN), the defender can gradually optimize its strategy through continuous learning.

[0079] Real-time adjustment: During the actual game, the defender adjusts signal release and defense measures based on the latest strategy to cope with the dynamic changes of the attacker.

[0080] By implementing the above detailed steps, attackers and defenders can use dual optimization algorithms and reinforcement learning algorithms to optimize and adjust their respective strategies, thereby achieving the best results in the game process.

[0081] S5: The attacker combines the profit and loss functions in the multi-stage game, uses dynamic game theory and fuzzy logic rules, and updates the second posterior probability;

[0082] Specifically, first, after each stage of the game, the attacker collects and records data such as the strategy choices, gains and losses of both parties. These data include the attacker's actual gains, the defender's reactions, and various events that occurred during the game. Through this data, the attacker can construct gain and loss functions, which are used to evaluate the effectiveness of different strategy combinations. Next, the attacker uses dynamic game theory to analyze multi-stage games. Dynamic game theory can handle multi-stage decision-making problems. By recursively analyzing the strategy choices and gain changes in each stage, the attacker predicts possible strategy combinations and corresponding gains in future stages based on the game results of the current stage, and then optimizes the strategy choices of the current stage. On this basis, the attacker introduces fuzzy logic rules to deal with the uncertainty and ambiguity in the game process. Fuzzy logic can handle uncertainty that is difficult to express in traditional logic. By defining fuzzy sets and fuzzy rules, the attacker can fuzzify the gains and losses of different strategy combinations, thereby better describing the complex game environment. In order to update the second posterior probability, the attacker combines dynamic game theory and fuzzy logic rules to calculate the posterior probability of each possible defender type in the current stage. This process includes the following steps:

[0083] Application of fuzzy rules: Based on the game data, define fuzzy rules, fuzzify gains and losses, and form fuzzy sets;

[0084] Fuzzy reasoning: Calculate the likelihood of defender types under different strategy combinations using fuzzy reasoning methods (such as fuzzy inference systems or fuzzy neural networks);

[0085] Posterior probability update: Based on the fuzzy inference results and combined with dynamic game analysis, the second posterior probability of each defender type is updated.

[0086] Through the above steps, after each stage of the game, the attacker can combine the profit and loss functions, use dynamic game theory and fuzzy logic rules to dynamically adjust and optimize its strategy selection. At the same time, the attacker can continuously update the posterior probability of the defender type, improve the accuracy of the prediction of the defender's behavior, and thus select the optimal strategy in subsequent games.

[0087] S6: The two parties in the game iterate and determine whether the termination conditions are met. They analyze the results and adjust the strategy to end or continue the next round of the game. After each round of the game, the two parties use complex network analysis methods and nonlinear dynamic system models to evaluate and adjust the game process and optimize the strategy for the next round of the game. The complex network analysis method includes:

[0088] Representing the elements in the game process as network nodes and the relationships between network nodes as edges to construct a complex network model, wherein the elements include but are not limited to defense signals, attack strategies, and benefits;

[0089] Calculate basic network indicators, analyze the structural characteristics of the network, calculate network centrality indicators, and identify key nodes and important strategies in the game process;

[0090] Use community discovery algorithms to identify community structures in game networks and analyze the relationships between different strategies and signals.

[0091] Specifically, in the multi-stage network defense method based on signal game, the two parties in the game decide whether to end or continue the next round of game through cyclic iteration, judging the termination conditions, analyzing the results and adjusting the strategies. After each round of game begins, the attacker and the defender respectively execute their respective strategies and monitor the network status and strategy effects during the game in real time. The two parties interact through a series of game steps (such as signal release, attack path selection, profit calculation, etc.). After each round of game, the two parties need to determine whether the preset termination conditions are met. The termination conditions may include reaching the preset number of game rounds, the attacker's or defender's profit reaching a certain threshold, or the strategies of both parties no longer change significantly, that is, the strategies converge.

[0092] After each round of the game, the attacker and defender analyze the results. This analysis includes the actual effectiveness of their strategies, changes in network state, and the opponent's behavior. This analysis helps them understand the performance of their current strategies in that round, including attack success rate, defense effectiveness, gains, and losses. They also identify new attack and defense patterns and understand the opponent's potential types and strategy adjustments. To better evaluate and adjust their strategies, both parties utilize complex network analysis methods to conduct an in-depth analysis of the game process. Based on this complex network analysis, both parties use nonlinear dynamic system models to model and evaluate the game process. Specifically, these steps include constructing a nonlinear dynamic system model based on game data to describe the dynamic relationship between the strategies and gains of both parties. Using methods such as least squares and maximum likelihood estimation, the model parameters are estimated to ensure that the model accurately describes the actual game process. Finally, linearization methods or the Lyapunov method are used to analyze the model's stability and determine equilibrium points and periodic solutions during the game. Based on the evaluation results of the complex network analysis and nonlinear dynamic system model, both parties optimize their strategies for the next round of the game.

[0093] By implementing the above detailed steps, the two parties in the game can evaluate and adjust the game process through complex network analysis methods and nonlinear dynamic system models after each round of the game, thereby optimizing the strategy for the next round of the game and ensuring the best results in a dynamically changing network environment.

[0094] Through the present invention, dynamic optimization of defense strategies and effective response to attack strategies are achieved, the initiative and flexibility of network defense are enhanced, and the overall robustness and effectiveness of the defense system are improved. As a preferred embodiment of the above embodiment, the verification probability calculation process adopts a solution method based on Monte Carlo simulation, including:

[0095] Set the initial signal strategy set and the attacker strategy set;

[0096] Generate a number of combinations of signals and attack strategies through random sampling, and simulate various possible situations in actual games;

[0097] For each sample, calculate the attacker and defender's profit functions according to their corresponding signal strategy and attack strategy;

[0098] Based on the generated samples and the calculated payoff function, the posterior probability distributions of the attacker and defender are updated using the Bayesian formula.

[0099] Specifically, first, set the initial signal strategy set and attacker strategy set. The signal strategy set includes various signal release strategies that the defender may adopt, while the attacker strategy set includes various attack strategies that the attacker may adopt. The setting of these strategy sets should be as comprehensive as possible, covering various possible situations to ensure the comprehensiveness and representativeness of the simulation; next, generate several combinations of signals and attack strategies through random sampling. Each combination simulates a possible situation in the actual game, reflecting the scenario in which the attacker chooses the corresponding attack strategy when the defender releases a specific signal. These random combinations should be sufficient to ensure the reliability and statistical significance of the simulation results; then, for each sample, according to its corresponding signal strategy and attack strategy, The payoff function is calculated for each strategy. The payoff function is an important tool for evaluating the effectiveness of each strategy combination. It is typically based on predefined payoff and loss parameters, reflecting factors such as the strategy's success rate, cost, and risk. By calculating the payoff function for each sample, the effectiveness of different strategy combinations can be quantified. After all samples are calculated, the posterior probability distributions of the attacker and defender are updated using the Bayesian formula based on the generated samples and the calculated payoff function. The Bayesian formula combines prior probabilities with observed data to calculate posterior probabilities, thereby updating the judgment of the defender's type and behavior. The core of this step is to use simulation data to gradually correct and update the basis for strategy selection, thereby making better decisions in future games. By implementing the detailed steps above, attackers can continuously adjust and optimize their attack strategies in a dynamically changing network environment, thereby improving the success rate and effectiveness of their attacks. The Monte Carlo simulation-based method provides a flexible and efficient means for attackers to optimize their strategies using large amounts of simulated data in uncertain and complex game environments, thereby gaining an advantageous position in actual network attacks and defenses.

[0100] As a preferred embodiment of the above, the posterior probability is solved by a refined Bayesian Nash equilibrium, and the refined Bayesian Nash equilibrium satisfies the following conditions:

[0101]

[0102] Among them, a * (m) is the optimal attack strategy chosen by the attacker given the signal strategy m; For the attack strategy set S A Find the expression Maximizing attack strategy a h ; Is for all defenders of type t i Perform weighted summation; is the defender type t under a given signal strategy m i The posterior probability of U A (m,a h,t i ) indicates that in the signal strategy m, attack strategy a h and defender type t i The attacker's profit under m * (t) is the optimal signaling strategy chosen by the defender given the defender type t; To find the expression in the signal strategy set M Maximize the signaling strategy m j ; is for all attack strategies a h Perform weighted summation; p D (a h ) Select attack strategy a for the attacker h The probability of U D (m j ,a h ,t) represents the signal strategy m j , attack strategy a h and the defender's payoff under defender type t.

[0103] Specifically, The attacker is based on the prior judgment P A , observed signal m and attack strategy a * (m) is calculated using the Bayesian rule. The first formula represents the attacker's optimal strategy towards the defender after obtaining the posterior probability. The second formula represents the defender's optimal signaling strategy towards the attacker based on its belief in the attacker's choice of attack type under rational conditions.

[0104] As a preferred embodiment of the above, at the beginning of each game, the attacker considers the defender's type misjudgment and calculates the probabilities of different types of misjudgment to obtain a misjudgment rate matrix, which is:

[0105]

[0106] Among them, ε is the misjudgment rate, k is the number of stages of the game, and ε ij is the probability that the attacker misjudges the defender's true type.

[0107] Specifically, at the beginning of each game, the attacker will consider the defender's type misjudgment and calculate the probabilities of different types of misjudgment to obtain the misjudgment rate matrix M k (ε), the elements of this matrix represent the probability that the attacker misjudges the defender’s true type as other types, M k (ε) is an n×n matrix, where n represents the number of defender types. Each element in the matrix ∈ ijIndicates that in the kth game stage, the attacker will change the defender's true type t i Mistakenly identified as type t j The probability of false positives is obtained from the attacker's detection and information collection of the defender. The attacker will use various means to obtain information about the defender, and then estimate the probability of false positives of different types based on this information.

[0108] As a preferred embodiment of the above embodiment, the losses caused by the attacker adopting different attack strategies when the defender type is misjudged are represented by a loss matrix. The loss matrix is:

[0109]

[0110] Among them, λ(t r ,a x ) is when the real defender type is t r When the attacker chooses attack strategy a x The resulting losses, t1, t2…t n For all possible defender types, a1, a2…a n for all possible attack strategies.

[0111] Specifically, L(a) is an n×m matrix, where n represents the number of all possible defender types and m represents the number of all possible attack strategies. Each element λ(tr,ax) in the matrix represents the number of possible defender types when the true defender type is t. r When the attacker chooses attack strategy a x These losses can be estimated or obtained through experiments by the attacker based on his own goals and strategy choices. Usually, the attacker will consider the possible consequences of different attack strategies, such as the probability of attack success, the degree of damage caused, the risk of exposure, and other factors. Based on these considerations, the attacker can evaluate the possible losses caused by different attack strategies for each defender type. The element λ(t r ,a x ) represents the actual defender type is t r When the attacker chooses attack strategy a x The loss caused by . Calculate each element λ(t r ,a x ) are as follows: Defender type t r It can be classified according to the configuration, strategy and behavior pattern of the defense system. For example, the defender type can be the type of firewall used, the setting of the intrusion detection system, the response strategy, etc.; the attack strategy a xIncluding different attack methods, such as DDoS attacks, SQL injections, social engineering attacks, etc. Each strategy will be selected and implemented according to different goals and purposes. The loss λ(t r ,a x ) is mainly determined by the following factors:

[0112] Success probability: attack strategy ax on defender type t r The probability of success Ps;

[0113] Extent of damage: If the attack is successful, the specific damages such as economic losses, degree of data leakage, service interruption time, etc.; Exposure risk: If the attack fails, the risk of the attacker's identity being revealed and the possible legal penalties.

[0114] The specific calculation formula can be expressed as: λ(t r ,a x )=Ps(t r ,a x )×D(t r ,a x )+(1-Ps(t r ,a x ))×R(t r ,a x ) Among them, Ps(t r ,a x ) is the attack strategy a under defender type tr x The probability of success, D(t r ,a x ) is the degree of damage caused by a successful attack, R(t r ,a x ) is the loss of exposure risk after a failed attack. This loss matrix can help attackers make decisions when formulating attack strategies, thereby maximizing their attack benefits or minimizing potential losses.

[0115] As a preferred embodiment of the above, the risk-benefit of the attacker is calculated based on the risks and benefits generated when the attacker misjudges the defender type as different and adopts different strategies. The formula is:

[0116]

[0117] Among them, Rr(t i ,a h ) means that when the attacker misjudges the defender type as t i And adopt attack strategy a h The risk return at time λ(t i ,a h ) means that when the defender type is t i And the attacker adopts attack strategy ah The loss at time ε(t i ,t g ) indicates that the attacker misjudges the defender type as t i The actual type is t g , where n is the number of defender types.

[0118] Specifically, considering the loss factor and the probability of misjudgment, this formula calculates the risk-benefit when the attacker misjudges the defender as a certain type and adopts the corresponding attack strategy. The attacker can use this formula to weigh different attack strategies and choose the best action plan to maximize their own benefits or minimize possible losses. The loss factor measures the degree of loss that the attacker may incur when choosing a certain attack strategy under different defender types. The cumulative part of the formula means adding the risks and benefits of each possible defense type, where i represents the index value of the defense type, each i corresponds to a defense type, and λ(t i ,a h ) and ε(t i ,t g ) is the specific calculated value for this defense type. The misjudgment probability reflects the misjudgment that the attacker may make when evaluating the target system, that is, the deviation between the attacker's perception of the target system and the actual situation. When making an attack decision, the attacker will choose the corresponding attack strategy based on their assessment of the defender's type. The attacker will first judge the defender's type (which may result in misjudgment), and then choose the optimal attack strategy based on this judgment. h , therefore, the attacker misjudges the defender type as t i When , the corresponding attack strategy a will be selected h to maximize their gains or minimize their losses.

[0119] As a preferred embodiment of the above, the two-party game process is solved by the optimization method to maximize the benefits. When the defender adopts a pure strategy, the attacker's benefit for each strategy is U A (m,a) is calculated as follows:

[0120]

[0121] Among them, U A (m,a) is the expected benefit of the attacker when he adopts strategy a for the defender’s signal selection strategy m, P(t i |m) is the attacker's response to defender type t when the defender chooses signal strategy m. i The posterior probability, U A (a,m,t i ) is when the attacker adopts strategy a, the defender chooses signal strategy m and the defender type is ti In the case of , the attacker's profit.

[0122] As a preferred embodiment of the above, when the defender has a pure strategy, the expected return of each strategy of the attacker forms the attacker's expected return matrix U A , the expression is:

[0123]

[0124] Among them, U A The matrix element of row i and column j is U A (m i ,a j ) is the defender's choice of signal strategy m i , the attacker chooses strategy a j The expected profit of the attacker in the case of .

[0125] Specifically, in addition to the above strategies, this embodiment also provides the following strategy combinations:

[0126] The attacker adopts mixed strategy y and the defender adopts pure strategy D, then the attacker's expected payoff is U A =D·U A ·y T The defender adopts mixed strategy x, and the attacker adopts pure strategy A. Since the defender knows his own defense type, the defender already has the payoff matrix

[0127] The defender adopts mixed strategy x, and the attacker adopts pure strategy A, then the attacker's payoff is U D =x·U D ·A T

[0128] When the defender adopts mixed strategy x and the attacker adopts mixed strategy y, the expected benefits of both the attacker and the defender are as follows:

[0129] U A =x·U A ·y T

[0130] U D =x·U D ·y T

[0131] As a preferred embodiment of the above, the complex network analysis method includes:

[0132] The elements in the game process are represented as network nodes, and the relationship between network nodes is represented as edges, to build a complex network model. The elements include but are not limited to defense signals, attack strategies, and benefits.

[0133] Calculate basic network indicators, analyze the structural characteristics of the network, calculate network centrality indicators, and identify key nodes and important strategies in the game process;

[0134] Use community discovery algorithms to identify community structures in game networks and analyze the relationships between different strategies and signals.

[0135] Specifically, first, the elements in the game process are represented as network nodes, and the relationships between network nodes are represented as edges, and a complex network model is constructed. The elements include but are not limited to defense signals, attack strategies, and benefits. By representing these elements as nodes and the interactions and relationships between nodes as edges, the various associations and interactions in the game process can be intuitively displayed; next, the basic indicators of the network are calculated and the structural characteristics of the network are analyzed. These basic indicators include node degree, clustering coefficient, average path length, etc. Through these indicators, the overall structure and connectivity of the network can be understood. In addition, it is also necessary to calculate the centrality indicators of the network, such as degree centrality, betweenness centrality, and eigenvector centrality, to identify the game process. The centrality index helps determine which nodes play a key role in the network and are the focus of strategy optimization. Then, the community discovery algorithm is used to identify the community structure in the game network. The community discovery algorithm can divide the network into several communities. The nodes in each community have strong internal connections, while the connections between communities are relatively sparse. Commonly used community discovery algorithms include the Louvain algorithm and the Girvan-Newman algorithm. Through these algorithms, the community structure in the game network can be identified and the relationship between different strategies and signals can be analyzed. For example, certain defense signals and attack strategies may appear frequently in the same community, indicating that there is a close relationship between them.

[0136] Example 2

[0137] Based on the same inventive concept as the multi-stage network defense method based on signal game in the aforementioned embodiment, the present invention further provides a multi-stage network defense system based on signal game, the system comprising:

[0138] Initial prior judgment formation module: The defender releases multiple signals to deceive the attacker, and the attacker forms an initial prior judgment through detection;

[0139] The optimal attack strategy determination module updates the posterior probability of the attacker based on the defense signal and combines the machine learning algorithm to analyze the signal characteristics and historical attack data to determine the optimal attack strategy;

[0140] Release strategy optimization and adjustment module: The defender monitors the network status in real time, optimizes the signal release strategy, and dynamically adjusts the signal frequency and type through topological structure analysis based on graph neural networks and multi-level game decision-making models;

[0141] In the two-party strategy optimization module, the attacker and defender use dual optimization algorithms to optimize their respective strategies. The dual optimization algorithms used by the defender are simulated degradation algorithm and differential evolution algorithm, while the dual optimization algorithms used by the attacker are genetic algorithm and particle swarm optimization algorithm.

[0142] In the second posterior probability update module, the attacker combines the profit and loss functions in the multi-stage game, uses dynamic game theory and fuzzy logic rules to update the second posterior probability;

[0143] The game termination judgment optimization module iterates the game cycle and determines whether the termination conditions are met, analyzes the results and adjusts the strategy to end or continue the next round of the game. After each round of the game, both parties use complex network analysis methods and nonlinear dynamic system models to evaluate and adjust the game process and optimize the strategy for the next round of the game. Complex network analysis methods include:

[0144] Representing the elements in the game process as network nodes and the relationships between network nodes as edges to construct a complex network model, wherein the elements include but are not limited to defense signals, attack strategies, and benefits;

[0145] Calculate basic network indicators, analyze the structural characteristics of the network, calculate network centrality indicators, and identify key nodes and important strategies in the game process;

[0146] Use community discovery algorithms to identify community structures in game networks and analyze the relationships between different strategies and signals.

[0147] The above-mentioned defense system in the present invention can effectively implement a multi-stage network defense method based on signal game, and the technical effects that can be achieved are as described in the above embodiments and will not be repeated here.

[0148] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and drawings are merely illustrative of the present application as defined herein and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, the present application is intended to include such modifications and variations as fall within the scope of the present application and its equivalents.

Claims

1. A multi-stage network defense method based on signaling game, characterized in that: include: The defender releases multiple signals to deceive the attacker, and the attacker forms an initial prior judgment through detection; The attacker updates the posterior probability based on the defense signal and combines machine learning algorithms to analyze signal characteristics and historical attack data to determine the optimal attack strategy. The defender monitors the network status in real time, optimizes the signal release strategy, and dynamically adjusts the signal frequency and type through topological structure analysis based on graph neural networks and multi-level game decision-making models; The attacker and defender use dual optimization algorithms to optimize their respective strategies. The dual optimization algorithms used by the defender are simulated degradation algorithm and differential evolution algorithm, and the dual optimization algorithms used by the attacker are genetic algorithm and particle swarm optimization algorithm. The attacker combines the profit and loss functions in the multi-stage game, uses dynamic game theory and fuzzy logic rules, and updates the second posterior probability; The two parties in the game iterate and determine whether the termination conditions are met. They analyze the results and adjust the strategy to end or continue the next round of the game. After each round of the game, the two parties use complex network analysis methods and nonlinear dynamic system models to evaluate and adjust the game process and optimize the strategy for the next round of the game. The complex network analysis method includes: Representing the elements in the game process as network nodes and the relationships between the network nodes as edges to construct a complex network model, wherein the elements include but are not limited to defense signals, attack strategies, and benefits; Calculate basic network indicators, analyze the structural characteristics of the network, calculate network centrality indicators, and identify key nodes and important strategies in the game process; Use community discovery algorithms to identify community structures in game networks and analyze the relationships between different strategies and signals.

2. The multi-stage network defense method based on signal game according to claim 1 is characterized in that: The posterior probability calculation process adopts a solution method based on Monte Carlo simulation, including: Set the initial signal strategy set and the attacker strategy set; Generate a number of combinations of signals and attack strategies through random sampling, wherein the combinations simulate various possible situations in actual games; For each sample, calculate the attacker and defender's profit functions according to their corresponding signal strategy and attack strategy; Based on the generated samples and the calculated benefit function, the posterior probability distribution of the attacker and the defender is updated using the Bayesian formula.

3. The multi-stage network defense method based on signal game according to claim 1 is characterized in that: The posterior probability is solved by a refined Bayesian Nash equilibrium, and the refined Bayesian Nash equilibrium satisfies the following conditions: Among them, a * (m) is the optimal attack strategy chosen by the attacker given the signal strategy m; For the attack strategy set S A Find the expression Maximizing attack strategies; Is for all defenders of type t i Perform weighted summation; is the defender type t under a given signal strategy m i The posterior probability of U A (m,a h ,t i ) indicates that in the signal strategy m, attack strategy a h and defender type t i The attacker's profit under m * (t) is the optimal signaling strategy chosen by the defender given the defender type t; To find the expression in the signal strategy set M Maximize the signaling strategy m j ; is for all attack strategies a h Perform weighted summation; p D (a h ) Select attack strategy a for the attacker h The probability of U D (m j ,a h ,t) represents the signal strategy m j , attack strategy a h and the defender's payoff under defender type t.

4. The multi-stage network defense method based on signal game according to claim 1 is characterized in that: At the beginning of each game, the attacker considers the defender's misjudgment of the type and calculates the probabilities of different types of misjudgments to obtain the misjudgment rate matrix, which is: Among them, ε is the misjudgment rate, k is the number of stages of the game, and ε ij is the probability that the attacker misjudges the defender's true type.

5. The multi-stage network defense method based on signal game according to claim 1 is characterized in that: The loss caused by the attacker adopting different attack strategies when the defender's type is misjudged is expressed by the loss matrix, which is: Among them, λ(t r ,a x ) is the real defender type t r When the attacker chooses attack strategy a x The resulting losses, t1, t2…t n For all possible defender types, a1, a2…a n for all possible attack strategies.

6. The multi-stage network defense method based on signal game according to claim 5 is characterized in that: The risk-reward of the attacker is calculated based on the risks and benefits of misjudging the defender's type and adopting different strategies. The formula is: Among them, Rr(t i ,a h ) means that when the attacker misjudges the defender type as t i And adopt attack strategy a h The risk return at time λ(t i ,a h ) means that when the defender type is t i And the attacker adopts attack strategy a h The loss at time ε(t i ,t g ) indicates that the attacker misjudges the defender type as t i The actual type is t g , where n is the number of defender types.

7. The multi-stage network defense method based on signal game according to claim 1 is characterized in that: The optimization method is used to solve the game process between the two parties to maximize their own benefits. When the defender adopts a pure strategy, the attacker's benefit for each strategy is U A (m,a) is calculated as follows: Among them, U A (m,a) is the expected benefit of the attacker when the defender chooses the signal strategy m, i |m) is the attacker's response to defender type t when the defender chooses signal strategy m. i The posterior probability, U A (a,m,t i ) is when the attacker adopts attack strategy a, the defender chooses signal strategy m and the defender type is t i In the case of , the attacker's profit.

8. The multi-stage network defense method based on signal game according to claim 1 is characterized in that , when the defender has a pure strategy, the expected return of each strategy of the attacker forms the attacker's expected return matrix U A , the expression is: Among them, U A The matrix element of row i and column j is U A (m i ,a j ) is the defender's choice of signal strategy m i , the attacker chooses attack strategy a j The expected profit of the attacker in the case of .

9. A multi-stage network defense system based on signaling game, characterized in that: The system comprises: Initial prior judgment formation module: The defender releases multiple signals to deceive the attacker, and the attacker forms an initial prior judgment through detection; The optimal attack strategy determination module updates the posterior probability of the attacker based on the defense signal and combines the machine learning algorithm to analyze the signal characteristics and historical attack data to determine the optimal attack strategy; Release strategy optimization and adjustment module: The defender monitors the network status in real time, optimizes the signal release strategy, and dynamically adjusts the signal frequency and type through topological structure analysis based on graph neural networks and multi-level game decision-making models; In the two-party strategy optimization module, the attacker and defender use dual optimization algorithms to optimize their respective strategies. The defender uses simulated degradation algorithm and differential evolution algorithm, while the attacker uses genetic algorithm and particle swarm optimization algorithm. In the second posterior probability update module, the attacker combines the profit and loss functions in the multi-stage game, uses dynamic game theory and fuzzy logic rules to update the second posterior probability; In the game termination judgment optimization module, both parties iterate and determine whether the termination conditions are met, analyze the results and adjust the strategy to end or continue the next round of the game. After each round of the game, both parties use complex network analysis methods and nonlinear dynamic system models to evaluate and adjust the game process and optimize the strategy for the next round of the game. The complex network analysis method includes: Representing the elements in the game process as network nodes and the relationships between the network nodes as edges to construct a complex network model, wherein the elements include but are not limited to defense signals, attack strategies, and benefits; Calculate basic network indicators, analyze the structural characteristics of the network, calculate network centrality indicators, and identify key nodes and important strategies in the game process; Use community discovery algorithms to identify community structures in game networks and analyze the relationships between different strategies and signals.

Citation Information

Patent Citations

  • Network defense method based on multi-stage attack and defense signals

    CN114024738A

  • Power system regulation method based on deep reinforcement learning

    WO2024092954A1