Defense strategy self-generation method and system for intelligent device cluster
By constructing a self-generation method for defense strategies based on partially observable Markov games and deep reinforcement learning, the problem of insufficient coordination of defense resources in smart device cluster networks is solved, dynamic security protection and resource optimization are achieved, and network security and stability are improved.
Patent Information
- Application Number
- CN202510963075.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-14
Smart Images

Figure CN120768612A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security for smart device clusters, and specifically refers to a method and system for self-generating defense strategies for smart device clusters based on partially observable Markov games (POMG) and deep reinforcement learning. Background Art
[0002] With the rapid development of intelligent technology, smart device clusters are widely used in key sectors such as industry, transportation, and energy. However, these clusters face increasingly severe network security challenges, facing complex and diverse threats such as Advanced Persistent Threats (APTs), DoS attacks, and virus intrusions. These attacks can lead to data leaks, system crashes, and significant losses. The complex structure and interconnectedness of smart device clusters require the coordinated efforts of multiple security defense resources to mitigate these threats. However, existing defense mechanisms have significant shortcomings: traditional fixed-strategy defense approaches are inadequate to address the variability of attacks; independent defense resources lack coordination; and existing defense resource allocation cannot flexibly adapt to the dynamic changes of smart device clusters. Therefore, it is imperative to build a system that can adapt to the dynamic changes of smart device clusters and enable autonomous optimization of defense strategies. Using innovative technologies to intelligently generate defense strategies and efficiently allocate resources is crucial for ensuring the security, safety, and stability of smart device cluster networks. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention conducts in-depth research on network security defense of smart device clusters, and proposes a self-generation method and system for defense strategies for smart device clusters. Based on partially observable Markov games and deep reinforcement learning technology, it solves the problems of poor adaptability, insufficient coordination, and unreasonable resource scheduling in traditional defense technologies for smart device cluster networks, thereby greatly improving the overall effectiveness of network security defense.
[0004] The present invention provides a defense strategy self-generation system for a smart device cluster, comprising a network modeling module, an attack strategy integration module and a defense strategy dynamic generation module.
[0005] The network modeling module constructs a network attack and defense game model based on a partially observable Markov game for the currently studied smart device cluster; the network attack and defense game model includes: a participant space N, an action space A, a probability space P, a strategy space S, a reward space R and a state space M; wherein the probability space P quantifies the probability distribution of action choices of the attacker and the defender under different strategies, the strategy space S includes the attacker's attack strategy and the defender's dynamic defense strategy, the reward space R includes the rewards obtained by the attacker and the defender after each round of confrontation, and the state space M includes a defense matrix and an attack matrix; each network node cluster in the smart device cluster is regarded as a row of the matrix, and the types of network node clusters include data centers, communication node clusters and peripherals. Each column of the matrix corresponds to a network node in the network node cluster, the element value in the defense matrix corresponds to the defense value of the node, and the element value in the attack matrix corresponds to the attack value of the node; the defense value of the node is determined by quantifying the defense resources of the node, characterizing the defense capability of the node; the attack value of the node represents the intensity of the attack on the node, which is determined by quantifying the attack resources.
[0006] The attack strategy integration module is used to set the attack strategy and provide it to the network attack and defense game model, and select the attack strategy and action path according to the network topology of the smart device cluster.
[0007] When a smart device cluster engages in an attack and defense game, corresponding attack behaviors are executed on the smart device cluster according to the attack strategy to obtain the current attack matrix of the smart device cluster; the defense strategy dynamic generation module obtains the current defense matrix of the smart device cluster and the quantified detection capabilities of each network node cluster, and uses a deep reinforcement learning model to generate a defense strategy.
[0008] The system also includes a strategy execution and feedback module that executes the defense strategy generated by the defense strategy dynamic generation module, updates the network state space and calculates the reward calibration matrix based on the attack and defense confrontation results, and feeds back to the defense strategy dynamic generation module to drive the defense strategy update; the reward calibration matrix is calculated based on the reward function of the action, and the reward function of the action is calculated based on the attack and defense results during the confrontation process and the resources required to execute the defense action.
[0009] The attack strategies set by the attack strategy integration module include: denial of service DoS attack, virus propagation, advanced persistent threat APT attack, vulnerability exploitation attack, detection elimination attack, and random control attack; the detection elimination attack refers to the execution of behaviors that increase attack or reduce target detection capabilities; the random control attack refers to the first execution of zombie host control behavior, and the execution of DoS attack when it is not the first time.
[0010] Accordingly, the present invention provides a method for self-generating a defense strategy for a smart device cluster, comprising the following steps:
[0011] Step 1: The network modeling module constructs a network attack and defense game model to characterize the network attack and defense scenario of the current smart device cluster; the network attack and defense game model is represented as a sextuple (N, A, P, S, R, M), where the participant space N includes attackers and defenders, the action space A includes the attacker's attack actions and the defender's defense actions, the probability space P quantifies the probability distribution of the attacker and defender's action choices under different strategies, the strategy space S includes the attacker's attack strategy and the defender's dynamic defense strategy, the reward space R includes the reward value obtained by the attacker and defender based on the attack and defense results after each round of confrontation, and the state space M includes the defense matrix and the attack matrix; each network node cluster in the smart device cluster is taken as a row of the matrix, and the network node cluster types include data centers, communication node clusters, and peripherals. Each column of the matrix corresponds to a network node in the network node cluster, the element value in the defense matrix corresponds to the node's defense value, and the element value in the attack matrix corresponds to the node's attack value; the node's defense value is determined by quantifying the node's defense resources, characterizing the node's defense capability; the node's attack value represents the intensity of the node's attack, which is determined by quantifying the attack resources;
[0012] Step 2: Obtain the network topology of the smart device cluster and select the attack strategy and action path based on the connection characteristics of the nodes in the network topology;
[0013] Step 3: The smart device cluster conducts an attack and defense game, including: executing corresponding attack behaviors on the smart device cluster according to the attack strategy, and outputting the current attack matrix of the smart device cluster; the defense strategy dynamic generation module obtains the current defense matrix of the smart device cluster and the quantified detection capabilities of each network node cluster, uses the deep reinforcement learning model to generate the defense strategy, uses the reward calibration matrix to drive the defense strategy update, and outputs the optimal defense action and defense matrix; the reward calibration matrix is calculated based on the reward function of the action, and the reward function of the action is calculated based on the attack and defense results during the confrontation process and the resources consumed to execute the defense action.
[0014] In step 3, the deep reinforcement learning models set in the defense strategy dynamic generation module include DDQN, REINFORCE, and Actor-Critic; if the attack strategy is a multi-action stage attack, the DDQN algorithm is called to generate the defense strategy; if the attack strategy is a single-action regular attack, the Actor-Critic or REINFORCE algorithm is called to generate the defense strategy.
[0015] The advantages and positive effects of the present invention are:
[0016] (1) The method and system of the present invention construct a multi-dimensional network attack and defense game model, integrating multi-dimensional information such as network architecture and attack and defense resource distribution with typical attack strategies. Using a deep reinforcement learning algorithm, dynamic environment modeling and multi-stage attack identification enable autonomous defense strategy decision-making and optimization of defense status. The method and system of the present invention can provide dynamic security protection for smart device cluster systems, resolving the issues of lagging defense strategies and excessive resource consumption.
[0017] (2) The method and system of the present invention innovatively integrate multiple typical attack strategies: by deeply analyzing the characteristics of virus attacks, DoS attacks, APT attacks, etc. in smart device cluster networks, and combining them with the actual network topology, the present invention constructs a comprehensive and realistic attack strategy set, which makes the training and optimization of defense strategies more targeted and can effectively deal with diverse attack scenarios.
[0018] (3) The method and system of the present invention utilize deep reinforcement learning to construct a dynamic generation mechanism for defense strategies: using POMG as a theoretical framework, the interaction between the strategies of the attacker and defender is depicted in a network attack and defense scenario with incomplete information; with the help of deep reinforcement learning algorithms such as DDQN, REINFORCE, and Actor-Critic, the defense strategy is continuously optimized based on real-time network state feedback and a carefully designed reward mechanism; the reward mechanism comprehensively considers factors such as the success or failure of defense and resource consumption, and guides the deep reinforcement learning algorithm to generate a strategy that is more adaptable to complex attack environments. The method and system of the present invention rely on deep reinforcement learning to construct a defense strategy generation mechanism with excellent dynamic adaptability. In network attack and defense scenarios, information is often incomplete, and traditional methods are difficult to effectively handle this uncertainty. The present invention uses partially observable Markov game (POMG) as a solid theoretical framework to accurately depict the strategic interaction process between the attacker and defender in a situation with incomplete information. The key elements such as participants, actions, probabilities, strategies, rewards, and states in network attack and defense are carefully modeled, clearly presenting the complex dynamic process of the attack and defense game. At the same time, the system incorporates advanced deep reinforcement learning algorithms such as DDQN (Double-DQN), REINFORCE, and Actor-Critic. These algorithms continuously learn and optimize based on real-time network status feedback and a carefully designed reward mechanism. This reward mechanism comprehensively considers multiple factors, including defense success or failure, resource consumption, and attack severity, providing the algorithm with a clear and rational optimization guide. Through continuous iterative training, deep reinforcement learning algorithms can generate defense strategies that are highly adaptable to complex and changing attack environments, achieving a transition from static defense to dynamic, adaptive defense, significantly enhancing the defense strategy's ability to respond to diverse attack scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is an application example diagram of the defense strategy self-generation system for the smart device cluster of the present invention;
[0020] Figure 2 This is an example diagram of constructing a state matrix for network topologies of different complexity levels according to an embodiment of the present invention;
[0021] Figure 3 The corresponding relationship between attack strategies and actions in the embodiment of the present invention;
[0022] Figure 4 This is a diagram of the interaction model between deep reinforcement learning and a network attack and defense game environment according to an embodiment of the present invention;
[0023] Figure 5 This is the input and output graph of deep reinforcement learning in an embodiment of the present invention. Figure 6 This is a flowchart of the attack and defense game of the smart device cluster participated in by the method of the present invention. DETAILED DESCRIPTION
[0024] Exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings so that those skilled in the art can better understand the technical features and effects of the present invention.
[0025] This paper presents an autonomous defense strategy generation system for smart device clusters, based on partially observable Markov games (POMG) and deep reinforcement learning. The system primarily comprises an attack strategy integration module, a defense strategy dynamic generation module, and a network modeling module. The network modeling module, based on POMG, constructs a network attack and defense scenario model for the smart device cluster under investigation. This model includes abstract modeling of network elements, matrix mapping, and network topology analysis. The attack strategy integration module analyzes typical attack characteristics, designs and stores multiple attack strategies, and provides strategic support for simulated attacks. The defense strategy dynamic generation module utilizes a deep reinforcement learning model to perform environmental modeling, identify attack patterns, and generate and optimize defense strategies.
[0026] like Figure 1As shown, the network modeling module uses the six-tuple model NADGM to abstract network attack and defense scenarios. It constructs the six-tuple elements of the network attack and defense game model (NADGM) using the participant space building unit, action space building unit, probability space building unit, policy space building unit, reward space building unit, and state space building unit. Specifically, the participant space N includes the attacker and defender; the action space A defines the attacker's attack actions and the defender's defense actions; the probability space P quantifies the probability distribution of the attacker and defender's action choices under different strategies; the policy space S includes the attacker's different attack strategies and the defender's dynamic defense strategy combinations; the reward space R sets the reward value based on the attack and defense results of the attacker and defender, including the rewards received by the attacker and defender after each round of confrontation; and the state space M reflects the distribution of the attacker's and defender's attack and defense network resources, such as server load, through matrix mapping. The state space M of the present invention includes a defense matrix and an attack matrix. Each network node cluster in the smart device cluster is represented as a row in the matrix. Network node cluster types include data centers, communication node clusters, and peripherals. Each column in the matrix corresponds to a network node in the cluster. The elements in the defense matrix correspond to the node's defense value, while the elements in the attack matrix correspond to the node's attack value. A node's defense value is determined by quantifying its defense resources, representing its defensive capabilities. A node's attack value represents the attack strength of the node and is determined by quantifying its attack resources.
[0027] The attack strategy integration module in this embodiment of the present invention is responsible for storing and invoking six attack strategies, including DoS attacks (which consume bandwidth through traffic flooding) and APT attacks (which attempt to infiltrate key nodes over a long period of time). It also designs attack node selection rules based on network topology. For example, attackers can only infect the system through connected paths and cannot attack disconnected cluster nodes. The attack strategy integration module selects attack strategies and attack action paths based on the connectivity characteristics of nodes in the network topology of the intelligent device cluster.
[0028] The defense strategy dynamic generation module of the embodiment of the application integrates deep reinforcement learning algorithms such as DDQN (Double Deep Q-Network), REINFORCE and Actor-Critic, generates a defense strategy through dynamic environment modeling and multi-stage attack identification. When the intelligent device cluster performs attack and defense game, the intelligent device cluster performs corresponding attack behaviors according to the attack strategy, obtains the current attack matrix of the intelligent device cluster, the defense strategy dynamic generation module obtains the defense matrix of the current intelligent device cluster and the quantified detection capability of each network node cluster, and generates a defense strategy. The system of the application also includes a strategy execution and feedback module, which executes the defense strategy generated by the defense strategy dynamic generation module, updates the network state space according to the attack and defense confrontation result, and delivers feedback information to the defense strategy dynamic generation module to realize continuous optimization of the defense strategy. The application drives the defense strategy update by using a reward calibration matrix, and outputs the optimal defense action and defense matrix. The reward calibration matrix is obtained according to the reward function of the action, and the reward function of the action is obtained by calculating the attack and defense results in the confrontation process and the resources consumed by executing the defense action.
[0029] The defense strategy self-generation method for the intelligent device cluster realized by the embodiment of the application based on POMG and deep reinforcement learning mainly includes four steps.
[0030] Step 1: Construct a network attack and defense scene model of the intelligent device cluster based on POMG.
[0031] The network modeling module of the application uses a network attack and defense game model, constructs network attackers, defenders, actions, probabilities, strategies, rewards and network states into six tuples (N, A, P, S, R, M), abstracts the intelligent device cluster network through a multi-dimensional matrix, establishes a coordinate mapping of the attack and defense matrix and the server, and designs different configured matrices according to the network topology structure to study the influence of network complexity on attack and defense.
[0032] As Figure 1As shown, the attack behaviors of the attacker set in the action space A of the embodiment of the application include four kinds: enhance attack, attack detection, reduce detection and control zombie hosts. The embodiment of the application constructs six attack strategies including denial of service (DoS) attack, virus propagation, advanced persistent threat (APT) attack, vulnerability exploitation attack, elimination detection attack and random control attack by analyzing typical attack characteristics in the intelligent device cluster, sets attack node selection rules in combination with actual network attack scenes and clearly defines specific actions of each attack strategy. A user can set the attack strategy according to actual conditions.
[0033] In the embodiment of the application, the DoS attack strategy is set as is expressed as
[0034] ;
[0035] wherein, num_attack_actions represents the optional range of attack strength, num_attack_positions represents the optional range of reachable attack nodes, and Random represents a random selection operation.
[0036] The elimination detection attack strategy is set as is expressed as
[0037] ;
[0038] wherein, select represents selection, represents increasing attack behavior, represents reducing target detection behavior.
[0039] The random control attack strategy is set as is expressed as
[0040] ;
[0041] wherein, represents controlling zombie host behavior, represents DoS attack behavior.
[0042] The APT attack strategy is set as is expressed as
[0043] ;
[0044] in, Reconnaissance and Detection Operations, Represents the minimum cost action in the attack phase.
[0045] Set Virus attack strategy Expressed as:
[0046] ;
[0047] in, Represents the currently achievable virus attack actions, Represents the last selected virus attack action.
[0048] Setting Vulnerability Exploitation Attack Strategy Expressed as:
[0049] ;
[0050] in, Stands for Detect Vulnerability Action, Represents a virus attack action, which performs a virus attack based on the detected vulnerability.
[0051] Traditional network defenses are often limited to simulating common and known attack strategies, making it difficult to cope with the ever-evolving complex attack methods in smart device cluster networks. This paper thoroughly and comprehensively analyzes the unique attack characteristics of various attacks, such as virus attacks, DoS attacks, and APT attacks, in the context of smart device cluster networks. It not only meticulously studies the propagation mechanism, infection path, and destruction methods of virus attacks, but also conducts in-depth analysis of the traffic flooding pattern, attack source distribution, and impact on network bandwidth and server performance in DoS attacks. For APT attacks, it also disassembles the entire process from long-term lurking, gradual penetration, to precise strikes.
[0052] This embodiment of the present invention configures a defender's defensive action space to include three types of actions: defend, repair, and enhance defense. Defensive actions are executed through isolation, blocking, and reallocation of defense resources. Repair actions increase the quantitative strength of defense resources. Enhance defense increases detection capabilities, improving the probability of detecting an attacker in each round of actions. A defender's defense strategy is a dynamically composed sequence of defense actions.
[0053] The state space M reflects the different state distributions of attack and defense resources in the intelligent device cluster system network, and the network constitutes a zero-sum game, satisfying , hour, ;in represents the defense matrix, represents the attack matrix, () indicates the reward obtained by the defender, () represents the reward obtained by the attacker, is the defender state space, is the attacker's state space.
[0054] The attack and defense of the smart device cluster system constitutes a zero-sum game. The present invention models the smart device cluster network through a multi-dimensional matrix, constructs an attack matrix and a defense matrix based on the network architecture, and establishes a matrix-server coordinate mapping. Figure 2 As shown, the embodiment of the present invention illustrates the states of a single-layer network topology architecture and a multi-layer network topology architecture. Figure 2 The single-layer network topology architecture of example (a) includes data center 1, node cluster 2, and peripheral device 3. The established state matrix includes an attack matrix and a defense matrix. Each row in each matrix corresponds to a network node cluster in the network topology architecture of the intelligent device cluster system. For example, the defense matrix contains 3 rows, which correspond to the network resources of data center 1, node cluster 2, and peripheral device 3 respectively. Figure 2 The multi-layer network topology architecture illustrated in (b) above comprises data center 1, node clusters 2 through 7, and peripheral device 8. The established state matrix includes an attack matrix and a defense matrix. Each row in each matrix corresponds to a network node cluster in the network topology. For example, the attack matrix contains eight rows, corresponding in sequence to the network resources of data center 1, node clusters 2 through 7, and peripheral device 8. Each column of the matrix corresponds to a different network node in the same network node cluster, and the element values correspond to the node's attack value or defense value. The rows and columns of the defense matrix and the attack matrix correspond to the same physical quantities: network node clusters and network nodes. The element values in the defense matrix and the attack matrix are derived based on physical quantities such as the defense resources and attack resources of real-world network nodes. In this embodiment of the present invention, a node's defense value represents the node's defense strength, while a node's attack value represents the attack strength against the node. The specific defense and attack network resources of interest can be set based on actual conditions, such as server load, computing resources, or network bandwidth. Attack strength is a quantitative representation of the attack resources applied and the attack effect, such as the degree to which the node's defenses are breached. Defense Strength is a quantitative representation of a node's defense resources. Higher values indicate stronger defense capabilities. Defense resources are quantified in the same way as attack resources. For defense and attack values of the same magnitude, the required resources are typically equal or nearly equal.
[0055] Step 2: Obtain the network topology of the smart device cluster. The attacker selects attack nodes based on the single-layer or multi-layer nature of the smart device cluster's network topology, prioritizing highly accessible or high-value nodes as targets. For example, in a single-layer topology, an attacker might infect all connected nodes through broadcasting, while in a multi-layer topology, the attacker would gradually infiltrate core nodes along a hierarchical path. Specifically, the attacker assesses the node's connectivity and defense resource distribution, such as targeting server nodes with high bandwidth or edge nodes with low protection levels. This target node is then determined based on pre-defined attack strategies, such as prioritizing bandwidth consumption in DoS attacks or long-term latency in APT attacks. This selection logic can be mapped using an attack matrix. The attacker's actions can also dynamically influence the node's defense resource status. For example, attacking a node can reduce its defense value from 2 to 1, driving the evolution of the attack-defense game. The defense strategy module must be aware of this node selection behavior in real time, for example, by isolating the attacked node or reallocating defense resources to high-risk areas to block the attack path and minimize losses.
[0056] Figure 3 The paper demonstrates the mapping relationship between attack strategies and specific actions, and how these actions dynamically influence the state evolution of the network attack-defense game matrix. After selecting a target strategy from six strategies, the attacker deploys a series of sequential actions to manipulate the game matrix. For example, the action of "reducing detection" adjusts the defender's detection capability matrix value from 2 to 1, thereby weakening its defense capabilities. Each attack strategy corresponds to a different combination of actions. For example, a DoS attack strategy consumes bandwidth resources through traffic flooding, while an APT attack strategy gradually infiltrates key nodes through long-term lurking. These actions reconfigure payoff parameters, such as the probability of attack success and the cost of defense, through strategic interaction, and establish an adaptive competitive environment. For example, a virus attack strategy may spread to multiple nodes through infection paths, while an Elimination Detection attack prioritizes high-value nodes to undermine the defense system. Figure 3 The correlation between attack strategies and network topology is also reflected in the figure. For example, attackers need to choose action paths based on the connection characteristics of single-layer or single-layer / multi-layer topologies. This strategy-action-matrix linkage mechanism supports multi-dimensional scenario analysis, quantifies the impact of different attack behaviors on network status, and provides a data basis for defense strategy optimization. Figure 3 As shown in Figure 2, there are four attack actions and six attack strategies of the attacker in this example. Figure 3 In this example, when attack strategy 1 (DoS) is executed and action 1 (enhanced attack) is used to attack a node in a cluster, the corresponding row in the attack matrix represents the situation in which the cluster node is attacked by this attack strategy-action combination, such as increasing the attack intensity of a node to 1.
[0057] Step 3: Execute the attack and defense game of the smart device cluster and generate a defense strategy through deep reinforcement learning.
[0058] Figure 4 This paper demonstrates a model for generating defense strategies using deep reinforcement learning (DRL). By modeling dynamic environments and identifying multi-stage attack patterns, it automatically generates adaptive defense strategies, enabling defenders to effectively respond to evolving attack methods and maintain real-time responsiveness. In the figure, a deep reinforcement learning model is used to dynamically generate strategies, providing defenders with adaptive strategies. A game theory-based network attack and defense model forms the foundation for deep reinforcement learning to participate in these attacks and defenses.
[0059] The dynamic defense strategy generation module is based on the nature of zero-sum games. Following the principle of maximizing the defender's minimum benefit and the attacker's maximum benefit, it constructs a multi-algorithm framework that includes DDQN, REINFORCE, and Actor-Critic algorithms, and sets the objective function for each algorithm. It designs a reward function based on the attack and defense resources in the network matrix and the resource consumption of defense actions, and optimizes the defense strategy through matrix transformation with reward calibration.
[0060] In the embodiment of the present invention, the objective function of the DDQN algorithm is set for:
[0061] ;
[0062] in, An immediate reward for the defender after performing an action; A discount factor used to balance the weights of current rewards and future rewards, usually ranging from 0 to 1; The status of the smart device cluster, including the defense matrix and attack matrix, represents the current network environment; For defensive behavior, i.e., the defensive actions performed by the defender, it is represented as a matrix with the same dimensions as the defense matrix; are the parameters of the current Q network, are the parameters of the target Q network; the Q function Q() describes the state-action value relationship.
[0063] When using the REINFORCE algorithm, the defense strategy is output through the policy network, and the policy network parameters are updated. Optimize the defense strategy, the objective function is as follows:
[0064] ;
[0065] in, are the updated policy network parameters, is the current strategy network parameter; The learning rate controls the step size of each parameter update and determines the magnitude of the strategy adjustment. is the policy network parameter The gradient operator is used to calculate the parameter update direction; is the policy network parameter at time t; is the action value function, indicating that at time t, the state Execute defensive behavior After that, the expected sum of future rewards is used to measure the benefit of the action-state combination; Is the policy function, which means that in a given state and parameters Take defensive action in the event of probability.
[0066] When using the Actor-Critic algorithm to optimize defense strategies, the value network evaluates the action value, and the strategy network outputs the probability of a defense action. Setting the objective function of the Actor-Critic algorithm includes the following:
[0067] ;
[0068] ;
[0069] in, are the parameters of the updated value network Critic; The parameters of the current Critic; is the learning rate of the Critic parameter; is the time series difference error at time t, reflecting the deviation between the value estimate and the actual return; is the gradient operator of the value network parameter w; is the action value function, evaluated in state Execute defensive behavior expected return; The updated policy network actor parameters; are the parameters of the current policy network; is the learning rate of the policy network parameters; is the estimated action value at time t; is the policy function, which means that in the state and parameters Take defensive action probability.
[0070] The three deep reinforcement learning algorithms used in this embodiment of the present invention are invoked as needed based on the characteristics of the attack strategy, whether it is a multi-action attack or a single-action pattern attack. For example, DDQN is used for complex multi-stage attacks, while Actor-Critic or REINFORCE are used for single-action pattern attacks. The corresponding algorithm is invoked as needed to generate a defense strategy.
[0071] The equilibrium solution in an incomplete information game is achieved through minimax optimization, where the two adversaries strategically maximize the attacker's minimum payoff and the defender's optimal strategy under the constraints of a zero-sum game. The defender's payoff is as follows:
[0072] ;
[0073] According to the nature of zero-sum game, the attacker's payoff is as follows:
[0074] ;
[0075] Then, we can derive the defender's strategy selection scheme , as shown below:
[0076] ;
[0077] Schemes for attackers to choose strategies It can be expressed by the following formula:
[0078] ;
[0079] By continuously optimizing the choices of the attacker and defender in the game process, the optimal strategy under the current state and design is achieved. When the following formula is satisfied, the strategy This constitutes a two-player zero-sum game Nash equilibrium.
[0080] .
[0081] The reward function is a key factor in evaluating the effectiveness of a strategy in a reinforcement learning algorithm. In the adversarial process, the reward function of an action is As shown in the following formula:
[0082] ;
[0083] in, Represents the defense value of the node in the nth row and mth column of the defense matrix, Represents the attack value of the node in the nth row and mth column in the attack matrix, Represents the resources consumed to perform a defense action. When the defense value is less than the attack value, the attack is successful; otherwise, the defense is successful. If a node successfully defends, a +1 reward is given for the defense action; if the attack is successful, a -1 penalty is imposed on the defender. The reward function is calculated for each node in the smart device cluster. The reward calibration matrix R of the smart device cluster under the current defense action is obtained.
[0084] The multi-algorithm deep reinforcement learning framework established in this paper generates autonomous defense strategies by detecting the perception matrix M and discrete actions A, and drives the defense strategy update with the reward calibration matrix R, outputting the optimal defense action and defense matrix. Figure 5 As shown, the input of the deep reinforcement learning model of the embodiment of the present invention is a matrix composed of the current defense matrix of the smart device cluster and the detection capability. The detection capability is the quantitative value of the attacker's behavior detected by the defender in each round. The stronger the detection capability, the greater the probability of discovering the attacker. In the embodiment of the present invention, a detection capability value is quantified for each network node cluster in the smart device cluster, and merged into the last column of the defense matrix to obtain the input matrix. The detection capability of the network node cluster is generally related to the attack detection means possessed by the cluster device. The detection capability value can be quantified according to the attack behavior that can be detected by the attack detection means. Then, in the attack and defense game of the smart device cluster, for the current attack matrix, it can be judged according to the detection capability value whether the network node cluster has detected the attacker.
[0085] The deep reinforcement learning algorithm decides the defensive action based on the input and defense action space. The game environment judges the effect of the action and outputs the adjusted defense matrix result as the defense state matrix for the next round of attack and defense game. Figure 6 As shown, the attack and defense game process of the smart device cluster using the method of the present invention is as follows: during the initial game, the attacker initiates an attack and the defender selects a defensive action, and both interact in an attack and defense environment. During the game process, it is first determined whether the attacker is detected by the network node cluster. If not detected, the attacker performs an attack action, and then determines whether the data center is breached. If breached, the attacker's gain -1 and the defender's gain +1 are fed back, and the action reward is calculated based on the current defense matrix and the attack matrix. After that, the defender selects a defensive action and updates the defense matrix. If not breached, the defender selects a defensive action to execute, updates the defense matrix, and then determines the attack and defense results based on the updated defense matrix and the attack matrix, and calculates the action reward. If an attacker is detected, the defender selects a defensive action to execute, updates the defense matrix, and determines the attack and defense results based on the updated defense matrix and the attack matrix, and calculates the action reward. The attack and defense actions act cyclically on the attack and defense environment, and the gain is determined based on the detection and breach results. This embodiment of the present invention determines the outcome of attack and defense based on defense and attack values. If the attack is successful, the attacker receives a +1 reward and the defender a -1 reward. If the defense is successful, the attacker receives a +1 reward and the defender a -1 reward. This defense strategy generation and optimization method incorporates defender action selection logic and leverages algorithms such as reinforcement learning to allow defenders to dynamically select optimal defensive actions within an attack-defense environment and participate in competitive games.
[0086] The foregoing description is merely one embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Except for the technical features described in the specification, all other technical features are known to those skilled in the art. Descriptions of well-known components and technologies are omitted in this disclosure to avoid redundancy and unnecessary limitation of the present invention.
Claims
1. A defense strategy self-generation system for smart device clusters, characterized by: It includes network modeling module, attack strategy integration module and defense strategy dynamic generation module; The network modeling module constructs a network attack and defense game model based on partially observable Markov game for the smart device cluster currently under study; the network attack and defense game model includes: participant space N, action space A, probability space P, strategy space S, reward space R and state space M; wherein, the probability space P quantifies the probability distribution of action choices of attackers and defenders under different strategies, the strategy space S includes the attacker's attack strategy and the defender's dynamic defense strategy, the reward space R includes the rewards obtained by the attacker and defender after each round of confrontation, and the state space M includes a defense matrix and an attack matrix; each network node cluster in the smart device cluster is taken as a row of the matrix, the network node cluster types include data centers, communication node clusters and peripherals, each column of the matrix corresponds to a network node in the network node cluster, the element value in the defense matrix corresponds to the node's defense value, and the element value in the attack matrix corresponds to the node's attack value; the node's defense value is determined by quantifying the node's defense resources, characterizing the node's defense capability; the node's attack value represents the intensity of the node's attack, which is determined by quantifying the attack resources; The attack strategy integration module is used to set the attack strategy and provide it to the network attack and defense game model, and select the attack strategy and action path according to the network topology of the smart device cluster; When a smart device cluster engages in an attack and defense game, corresponding attack behaviors are executed on the smart device cluster according to the attack strategy to obtain the current attack matrix of the smart device cluster; the defense strategy dynamic generation module obtains the current defense matrix of the smart device cluster and the quantified detection capabilities of each network node cluster, and uses a deep reinforcement learning model to generate a defense strategy.
2. The system according to claim 1, wherein: The attack strategies set by the attack strategy integration module include: denial of service DoS attack, virus propagation, advanced persistent threat APT attack, vulnerability exploitation attack, detection elimination attack, and random control attack; the detection elimination attack refers to the execution of behaviors that increase attack or reduce target detection capabilities; the random control attack refers to the first execution of zombie host control behavior, and the execution of DoS attack when it is not the first time.
3. The system according to claim 1, wherein: The system also includes a strategy execution and feedback module that executes the defense strategy generated by the defense strategy dynamic generation module, updates the network state space and calculates the reward calibration matrix based on the attack and defense confrontation results, and feeds back to the defense strategy dynamic generation module to drive the defense strategy update; The reward calibration matrix is obtained by calculating the reward function of the action, and the reward function of the action is obtained by calculating the attack and defense results in the confrontation process and the resources consumed by executing the defense action.
4. A method for self-generating defense strategies for smart device clusters, characterized in that: The steps include: Step 1: The network modeling module builds a network attack and defense game model to characterize the network attack and defense scenario of the current smart device cluster; The network attack and defense game model is represented by a sextuple (N, A, P, S, R, M), where the participant space N includes attackers and defenders, the action space A includes the attacker's attack actions and the defender's defense actions, the probability space P quantifies the probability distribution of the attacker's and defender's action choices under different strategies, the strategy space S includes the attacker's attack strategy and the defender's dynamic defense strategy, the reward space R includes the reward values obtained by the attacker and defender from the attack and defense results after each round of confrontation, and the state space M includes a defense matrix and an attack matrix; each network node cluster in the smart device cluster is taken as a row of the matrix, and the types of network node clusters include data centers, communication node clusters, and peripherals. Each column of the matrix corresponds to a network node in the network node cluster, the element value in the defense matrix corresponds to the node's defense value, and the element value in the attack matrix corresponds to the node's attack value; the node's defense value is determined by quantifying the node's defense resources, characterizing the node's defense capability; the node's attack value represents the intensity of the node's attack, which is determined by quantifying the attack resources; Step 2: Obtain the network topology of the smart device cluster and select the attack strategy and action path based on the connection characteristics of the nodes in the network topology; Step 3: The smart device cluster conducts an attack and defense game, including: executing corresponding attack behaviors on the smart device cluster according to the attack strategy, and outputting the current attack matrix of the smart device cluster; the defense strategy dynamic generation module obtains the current defense matrix of the smart device cluster and the quantified detection capabilities of each network node cluster, uses the deep reinforcement learning model to generate the defense strategy, uses the reward calibration matrix to drive the defense strategy update, and outputs the optimal defense action and defense matrix; the reward calibration matrix is calculated based on the reward function of the action, and the reward function of the action is calculated based on the attack and defense results during the confrontation process and the resources consumed to execute the defense action.
5. The method according to claim 4, characterized in that In the step 1, in the action space A constructed, the attacker's attack behaviors are set to include four types: enhanced attack, attack detection, defense reduction, and zombie host control; In the constructed strategy space S, six attack strategies are set: denial of service DoS attack, virus propagation, advanced persistent threat APT attack, vulnerability exploitation attack, detection elimination attack, and random control attack; the detection elimination attack refers to the execution of behaviors that increase attack behavior or reduce target detection capability; the random control attack refers to the execution of zombie host control behavior for the first time, and the execution of DoS attack is not the first time.
6. The method according to claim 4, characterized in that In the above step 1, in the action space A, the defender's defense behaviors are set to include three types: defense, repair, and enhanced defense. Among them, the repair behavior increases the quantitative strength of the defense resources, and the enhanced defense behavior increases the probability of detecting the attacker by increasing the node detection capability. The defense strategy is obtained by dynamically combining the defense behaviors.
7. The method according to claim 4, characterized in that In step 2, when selecting attack nodes in the action path, high-accessibility or high-value nodes are preferentially selected as attack targets, the connectivity and defense strength of the nodes are evaluated, and the nodes to be attacked are determined in combination with the attack strategy.
8. The method according to claim 4, characterized in that In step 3, the deep reinforcement learning models set in the defense strategy dynamic generation module include DDQN, REINFORCE, and Actor-Critic; if the attack strategy is a multi-action stage attack, the DDQN algorithm is called to generate the defense strategy; If the attack strategy is a single-action regular attack, call the Actor-Critic or REINFORCE algorithm to generate a defense strategy.
9. The method according to claim 4, characterized in that In step 3, the reward function for the defensive action is calculated as follows: ; in, Represents the defense value of the mth node in the nth network node cluster in the defense matrix, Represents the attack value of the mth node in the nth network node cluster in the attack matrix; Represents the resources consumed to perform a defensive action; when the defense value is less than the attack value, it indicates a successful attack; otherwise, it indicates a successful defense; Calculate the reward function for each node in the smart device cluster The reward calibration matrix R of the current defense action is obtained by taking the value of 10. The method according to claim 8, characterized in that In step 3, setting the objective function of each deep reinforcement learning model includes: (1) Objective function of the DDQN algorithm ; in, For the reward after performing a defensive action, is the discount factor, Defense matrix and attack matrix for smart device clusters, is the defensive behavior matrix, are the parameters of the current Q network, are the parameters of the target Q network; Q is the Q function; (2) The strategy network of the REINFORCE algorithm is used to output the defense strategy, and the defense strategy is optimized by updating the strategy network parameters. ; in, is the current strategy network parameter; is the learning rate; is the policy network parameter The gradient operator of is the policy network parameter at time t; is the action value function, indicating that at time t, the state Execute defensive behavior After that, the expected total future rewards; Is the policy function, which means that in a given state and parameters Take defensive action in the event of probability; (3) The Actor-Critic algorithm is used to optimize the defense strategy. The objective functions set include: ; ; in, is the updated value network parameter, is the parameter of the current value network, is the learning rate of the value network parameters, is the timing difference error at time t, is the gradient operator of the value network parameter w, is the action-value function; are the updated policy network parameters, are the parameters of the current policy network, is the learning rate of the policy network parameters, is the estimated action value at time t; is the policy function, which means that in the state and parameters Take defensive action probability.
Citation Information
Patent Citations
Opponent behavior strategy modeling method and system based on deep reinforcement learning
CN116205298A
Unmanned system cluster multi-target game confrontation method
CN118068703A
Attention network migration-based multi-agent reinforcement learning air combat decision-making method
CN118917171A
Network adversarial decision-making method based on reinforcement learning
CN119966697A
Honey array defense resource allocation optimization method based on self-game reinforcement learning
CN120090880A
Cited By
Multi-agent dynamic defense game method and system based on federal reinforcement learning
CN121864500A
Multi-agent dynamic defense game method and system based on federal reinforcement learning
CN121864500B