Network security dynamic defense system and method based on multi-agent joint game and mobile target defense

The dynamic network security defense system, which utilizes multi-agent joint game theory and mobile target defense, solves the problem of existing technologies being unable to cope with dynamic network attacks. It enables real-time adjustment of defense strategies and resource optimization, improves the efficiency of threat perception and response in network security, and is suitable for large-scale network environments.

CN121508950APending Publication Date: 2026-02-10BEIJING UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511645320.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing network security defense systems are ill-equipped to deal with dynamic and intelligent network attacks, such as APT attacks. Defenders often struggle to obtain complete information about attackers, leading to reactive responses. Traditional MTD technology is resource-intensive and lacks adaptive capabilities. Multi-agent defense systems also lack efficient information sharing and collaborative decision-making mechanisms among agents.

Method used

A dynamic network security defense system based on multi-agent joint game theory and mobile target defense is adopted, including an environment perception module, a multi-agent decision-making module, an MTD execution module, and a utility evaluation module. It utilizes Bayesian game models and deep reinforcement learning to achieve real-time defense strategy adjustment and information sharing, and combines collaborative defense at the software and network layers.

Benefits of technology

It enables the prediction of attackers' strategies in an environment with incomplete information, reduces response latency, lowers resource consumption, increases attack costs, ensures the continuity of critical business operations, and achieves a user experience latency of less than 5%, thus achieving an optimal balance between security and feasibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121508950A_ABST
    Figure CN121508950A_ABST
Patent Text Reader

Abstract

The invention discloses a network security dynamic defense system and method based on a multi-agent joint game and mobile target defense, and the system comprises an environment sensing module which is used for monitoring network state data in real time, and the network state data comprise traffic features, vulnerability information and attack behaviors; and the multi-agent decision module is composed of a plurality of distributed agents, and each agent performs game strategy decision based on local observation information and shares key data with other agents through a secure communication protocol. Through the combination of the multi-agent joint game and the MTD strategy, the system can adjust the defense strategy in real time and effectively cope with complex attacks such as advanced persistent threats, and based on the Bayesian game model, the defense system can predict possible strategies of attackers in an incomplete information environment, deploy defense measures in advance, reduce response delay and improve the security of the attackers. And meanwhile, the software layer MTD and the network layer MTD are adopted, so that an attacker is difficult to accurately detect system vulnerabilities, and the attack cost is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, specifically to a dynamic network security defense system and method based on multi-agent joint game theory and moving target defense. Background Technology

[0002] With the rapid development of artificial intelligence technology and its widespread application in various fields, many industries have ushered in unprecedented changes, and almost every field has found matching intelligent solutions. However, technological advancements have inevitably been used in some gray areas, with cybersecurity being one of them. Traditional defense methods are no longer effective, and it is necessary to rely on technologies with similar automation and intelligence characteristics to effectively deal with these highly autonomous attacks.

[0003] Current research on network attack and defense game theory focuses on optimizing adversarial strategies. The introduction of machine learning and big data enables defense strategies to adapt and adjust adaptively, demonstrating significant advantages in threat prediction and response efficiency. For example, Chinese Patent 202011610220.8 proposes a dynamic network security defense system and method based on big data. This system uses a big data module to locate and analyze abnormal behaviors in data packets. Combined with active and passive defense, it establishes a dynamic defense model from the perspectives of both attackers and defenders. By setting up honeypots, it actively deceives attackers, disrupts their vision, and lures them into launching attacks, thereby extending the attack time and providing opportunities for the defense model to implement its defense strategy. Ultimately, this achieves a dynamic, real-time, and proactive defense system, enhancing its defensive effectiveness.

[0004] However, current network security defenses mostly employ static strategies, such as firewalls and intrusion detection systems, which are ill-equipped to counter dynamic and intelligent network attacks, such as APT attacks. Existing technologies have the following shortcomings:

[0005] 1. Information asymmetry: The defender has difficulty obtaining complete information about the attacker, leading to a passive response.

[0006] 2. Limited strategy: Traditional MTD technology consumes a lot of resources and lacks adaptability.

[0007] 3. Insufficient collaboration: In multi-agent defense systems, there is a lack of efficient information sharing and collaborative decision-making mechanisms among agents.

[0008] Therefore, a dynamic network security defense system and method based on multi-agent joint game theory and mobile target defense is proposed to solve the above problems. Summary of the Invention

[0009] To address the shortcomings of existing technologies, this invention provides a dynamic network security defense system and method based on multi-agent joint game theory and moving target defense. It offers advantages such as improved threat perception and response efficiency, and enhanced robustness of the defense system against unknown threats. It solves the problems of current network security defenses, which often employ static strategies such as firewalls and intrusion detection systems, struggling to cope with dynamic and intelligent network attacks, such as APT attacks. Defenders often struggle to obtain complete information about attackers, leading to passive responses. Traditional MTD technology is resource-intensive and lacks adaptive capabilities. Furthermore, multi-agent defense systems often lack efficient information sharing and collaborative decision-making mechanisms among agents.

[0010] To achieve the above objectives, the present invention provides the following technical solution: a dynamic network security defense system based on multi-agent joint game theory and moving target defense, comprising:

[0011] The environment awareness module is used to monitor network status data in real time, including traffic characteristics, vulnerability information, and attack behavior;

[0012] The multi-agent decision-making module consists of multiple distributed agents. Each agent makes game strategy decisions based on local observation information and shares key data with other agents through a secure communication protocol.

[0013] The MTD execution module is used to dynamically adjust network configuration parameters, including service port switching, virtual topology reconstruction, and honeypot deployment, and supports collaborative defense between the software layer and the network layer.

[0014] The effectiveness evaluation module is used to quantify the defense effect, including attack interception rate, system resource consumption and user experience indicators, and to provide feedback on optimization decision-making strategies.

[0015] Furthermore, the multi-agent decision-making module adopts a Bayesian game model, specifically including:

[0016] Define the strategy space and utility function of both the attacker and defender, where the defender's utility function integrates network security status, resource overhead, and service continuity.

[0017] The agent dynamically modifies its beliefs about attacker behavior and network status through a Bayesian probability update mechanism to optimize its defense strategy.

[0018] Furthermore, the multi-agent decision-making module adopts a centralized training and distributed execution framework, specifically including:

[0019] During the training phase, the agent learns a cooperative strategy based on global information and introduces noise to simulate an incomplete information environment.

[0020] During the execution phase, the agent makes independent decisions based solely on local observation information and achieves limited information sharing through encrypted channels.

[0021] Furthermore, the multi-agent decision-making module is trained using a multi-agent deep deterministic policy gradient algorithm, where each agent's Actor network outputs action policies based on local observations; the Critic network uses global information to evaluate the value of actions during the training phase to improve collaboration efficiency.

[0022] Furthermore, the MTD execution module includes:

[0023] The software layer MTD unit is used to dynamically switch application service interfaces, inject fake processes, or adjust runtime configurations to increase the difficulty for attackers to detect.

[0024] The network layer MTD unit, combined with network function virtualization technology, enables lightweight IP address hopping, port randomization, and dynamic route adjustment;

[0025] The MTD policy optimizer, based on deep reinforcement learning, is used to adaptively select the optimal defensive action according to the attack pattern.

[0026] A dynamic network security defense method based on multi-agent joint game theory and moving target defense network security dynamic defense system includes the following steps:

[0027] 1) Collect network status data in real time through the environmental awareness module to identify potential attack behaviors;

[0028] 2) Multi-agent systems initiate a Bayesian game model based on local observation information to predict attacker strategies and generate defensive actions;

[0029] 3) The MTD execution module dynamically adjusts the network configuration based on the decision results and implements the software layer or network layer MTD strategy;

[0030] 4) The utility evaluation module calculates the defense effect and feeds it back to the multi-agent decision-making module to optimize subsequent strategies.

[0031] Furthermore, the Bayesian game model in step 2) includes:

[0032] Construct the strategy trees of attackers and defenders, and define the game equilibrium conditions under incomplete information;

[0033] The agent updates rules using Bayesian methods and dynamically adjusts the prior probability distribution of attacker types by combining historical interaction data.

[0034] Furthermore, the MTD strategy generation in step 3) includes:

[0035] For small-scale networks, software-layer MTD should be enabled first to reduce resource overhead through dynamic service masquerading.

[0036] For large-scale networks, the combination of network layer MTD and NFV technologies is used to achieve efficient IP hopping and topology reconstruction.

[0037] Furthermore, the utility evaluation indicators in step 4) include:

[0038] Security metrics: attack interception success rate, average attack detection time;

[0039] Performance metrics: System CPU / memory utilization, service request response latency;

[0040] Economic indicator: The ratio of the cost of implementing a defense strategy to the potential losses caused by an attack.

[0041] Compared with the prior art, the technical solution of this application has the following beneficial effects:

[0042] 1. This invention combines multi-agent joint game theory with MTD strategy, enabling the system to adjust its defense strategy in real time and effectively respond to complex attacks such as advanced persistent threats. Based on the Bayesian game model, the defense system can predict the attacker's possible strategies in an environment with incomplete information, deploy defense measures in advance, and reduce response delay. At the same time, the use of software-layer MTD and network-layer MTD makes it difficult for attackers to accurately detect system vulnerabilities and increases the cost of attack.

[0043] 2. The software layer dynamic adjustment proposed in this invention can significantly reduce resource consumption in small-scale or critical systems. The MTD policy optimizer based on deep reinforcement learning can automatically adjust the defense strength according to the attack mode, avoiding resource waste caused by over-defense. The multi-agent adopts a local decision-making and limited information sharing mode, reducing the computational and communication burden of centralized defense systems, and is suitable for large-scale network environments.

[0044] 3. This invention adopts a distributed intelligent agent architecture, which can execute defense decisions quickly locally, avoiding the single-point bottleneck problem of traditional centralized security systems. The MTD policy is transparent to normal users, ensuring the continuity of critical business. The user experience latency increases by only less than 5%. The utility evaluation module comprehensively considers security indicators, performance indicators and economic indicators, so that the system achieves the optimal balance between security and feasibility. Attached Figure Description

[0045] Figure 1 This is a framework diagram of the network security dynamic defense system based on multi-agent joint game and moving target defense of the present invention;

[0046] Figure 2 This is a flowchart of the network security dynamic defense method based on multi-agent joint game and moving target defense of the present invention;

[0047] Figure 3 This is a flowchart of the Bayesian game theory process of this invention. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] Please see Figure 1-3 The network security dynamic defense system based on multi-agent joint game and moving target defense in this embodiment includes:

[0050] The environment awareness module is used to monitor network status data in real time, including traffic characteristics, vulnerability information, and attack behavior;

[0051] The multi-agent decision-making module consists of multiple distributed agents. Each agent makes game strategy decisions based on local observation information and shares key data with other agents through a secure communication protocol.

[0052] The MTD execution module is used to dynamically adjust network configuration parameters, including service port switching, virtual topology reconstruction, and honeypot deployment, and supports collaborative defense between the software layer and the network layer.

[0053] The effectiveness evaluation module is used to quantify the defense effect, including attack interception rate, system resource consumption and user experience indicators, and to provide feedback on optimization decision-making strategies.

[0054] Furthermore, the multi-agent decision-making module adopts a Bayesian game model, specifically including:

[0055] Define the strategy space and utility function of both the attacker and defender, where the defender's utility function integrates network security status, resource overhead, and service continuity.

[0056] The agent dynamically modifies its beliefs about attacker behavior and network status through a Bayesian probability update mechanism to optimize its defense strategy.

[0057] Please refer to Figure 3 The agent continuously updates its policy weights through Bayesian game theory to respond to the opponent's actions. The process begins with an initial state; the agent chooses an action based on the current situation, then evaluates the effect of the action, determining whether to adjust the strategy or continue with the current strategy. Simultaneously, the agent is required to continuously update its beliefs about the intentions and strategies of other agents based on observed behavior during interactions. These belief updates rely on observational information and prior probabilities. Through Bayesian update rules, the agent gradually adjusts its expectations of others' strategies, thereby optimizing its own decisions under incomplete information.

[0058] The multi-agent decision-making module adopts a centralized training and distributed execution framework, specifically including:

[0059] During the training phase, the agent learns a cooperative strategy based on global information and introduces noise to simulate an incomplete information environment.

[0060] During the execution phase, the agent makes independent decisions based solely on local observation information and achieves limited information sharing through encrypted channels.

[0061] The multi-agent decision-making module is trained using a multi-agent deep deterministic policy gradient algorithm. In this algorithm, each agent's Actor network outputs action policies based on local observations. The Critic network evaluates the value of actions using global information during the training phase to improve collaboration efficiency.

[0062] When establishing a simulated multi-agent network attack and defense environment, the multi-agent decision-making module focuses on features such as incomplete information for the defender, noise in information transmission between agents, and the dynamics of the attacker and the setting of decoys. Randomization and noise interference are introduced into the environment modeling to realistically represent the information asymmetry of multi-agents in complex networks. When constructing the multi-agent game, the attacker and defender can be regarded as a turn-based game.

[0063] In this embodiment, the MTD execution module includes:

[0064] The software layer MTD unit is used to dynamically switch application service interfaces, inject fake processes, or adjust runtime configurations to increase the difficulty for attackers to detect.

[0065] The network layer MTD unit, combined with network function virtualization technology, enables lightweight IP address hopping, port randomization, and dynamic route adjustment;

[0066] The MTD policy optimizer, based on deep reinforcement learning, is used to adaptively select the optimal defensive action according to the attack pattern.

[0067] The MTD execution module dynamically adjusts the software service's configuration and interface by monitoring attacker behavior patterns. Utilizing deep reinforcement learning, it trains the defense system to automatically identify threats and respond in real-time based on changes in attack patterns.

[0068] A dynamic network security defense method based on multi-agent joint game theory and moving target defense network security dynamic defense system includes the following steps:

[0069] 1) Collect network status data in real time through the environmental awareness module to identify potential attack behaviors;

[0070] 2) A multi-agent system initiates a Bayesian game model based on local observation information to predict the attacker's strategy and generate defensive actions. The Bayesian game model includes:

[0071] Construct the strategy trees of attackers and defenders, and define the game equilibrium conditions under incomplete information;

[0072] The agent updates rules using Bayesian methods and dynamically adjusts the prior probability distribution of attacker types by combining historical interaction data.

[0073] 3) The MTD execution module dynamically adjusts the network configuration based on the decision results, implementing software-layer or network-layer MTD policies. MTD policy generation includes:

[0074] For small-scale networks, software-layer MTD should be enabled first to reduce resource overhead through dynamic service masquerading.

[0075] For large-scale networks, the combination of network layer MTD and NFV technologies is used to achieve efficient IP hopping and topology reconstruction.

[0076] 4) The utility evaluation module calculates the defense effectiveness and feeds it back to the multi-agent decision-making module to optimize subsequent strategies. Utility evaluation metrics include:

[0077] Security metrics: attack interception success rate, average attack detection time;

[0078] Performance metrics: System CPU / memory utilization, service request response latency;

[0079] Economic indicator: The ratio of the cost of implementing a defense strategy to the potential losses caused by an attack.

[0080] Example

[0081] Application scenarios

[0082] A fintech company's hybrid cloud platform suffered an advanced persistent threat (APT) attack. Attackers attempted to compromise the core database through vulnerability probing, lateral movement, and data theft. Traditional static defense solutions (such as WAF and IDS) are insufficient to cope with dynamic attacks, necessitating the deployment of the dynamic defense system of this invention.

[0083] System Deployment

[0084] Environment awareness module: Deploys traffic probes to monitor abnormal connections in real time (such as high-frequency port scanning and SQL injection attempts). Vulnerability scanner identifies unauthorized access vulnerabilities (CVE-2023-1234) in the database service (MySQL).

[0085] Multi-agent decision-making module: One defense agent is deployed on each cloud host, and local observations include: current service status, network traffic, and login logs. Agents share attack signatures via a TLS encrypted channel.

[0086] MTD Execution Module: Software Layer MTD: Dynamically injects fake query interfaces into the database service container. Network Layer MTD: Combines NFV technology to randomly switch the database virtual IP every 5 minutes (e.g., 10.0.1.100 → 10.0.1.201).

[0087] Utility evaluation module: Define the utility function: Defense utility = 0.6 × attack interception rate + 0.2 × (1 - CPU load) + 0.2 × (1 - user latency increase).

[0088] Defense process

[0089] Attack Detection Phase: Agent A detects an attacker's brute-force attempt on the MySQL service (e.g., 50 login requests per second). The attacker's next strategy is calculated using a Bayesian game model (probabilities: exploit 60%, lateral movement 30%, hibernation 10%).

[0090] Dynamic response phase

[0091] MTD policy trigger: The software layer MTD deploys a high-interaction honeypot next to the real MySQL service, returning a fake database structure. The network layer MTD migrates the database VIP to a standby node and closes the original port (3306 → random port).

[0092] Multi-agent collaboration: Agent B discovers the same attacking IP attempting SSH brute-force attacks and immediately notifies Agent A to update the game model (the attacker's strategy confidence is increased to 80%).

[0093] Effect verification phase

[0094] Attackers are lured to the honeypot, triggering an alert and logging the attack fingerprint (such as the exploit code used).

[0095] System metrics:

[0096] Attack interception rate: 95% (original static defense rate was 60%).

[0097] CPU load increase: 8% (25% for traditional IP switching schemes).

[0098] User query latency: Increased by 3ms (negligible).

[0099] In summary, this invention combines multi-agent joint game theory with MTD (Multi-Target Deployment) strategies, enabling the system to adjust its defense strategy in real time and effectively respond to complex attacks such as advanced persistent threats. Based on a Bayesian game model, the defense system can predict attackers' possible strategies in environments with incomplete information, deploying defensive measures in advance and reducing response latency. Simultaneously, the use of software-layer and network-layer MTD makes it difficult for attackers to accurately detect system vulnerabilities, increasing the cost of attacks. The proposed software-layer dynamic adjustment can significantly reduce resource consumption in small-scale or critical systems. The deep reinforcement learning-based MTD strategy optimizer can automatically adjust defense strength according to attack patterns, avoiding resource waste caused by over-defense. The multi-agent approach employs local decision-making and limited information sharing, reducing the computational and communication burden of centralized defense systems and making it suitable for large-scale network environments. The distributed agent architecture allows defense decisions to be executed quickly locally, avoiding the single-point bottleneck problem of traditional centralized security systems. The MTD strategy is transparent to normal users, ensuring the continuity of critical services, with user experience latency increasing by only less than 5%. The utility evaluation module comprehensively considers security, performance, and economic indicators, achieving an optimal balance between security and feasibility.

[0100] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0101] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A dynamic network security defense system based on multi-agent joint game theory and moving target defense, characterized in that, include: The environment awareness module is used to monitor network status data in real time, including traffic characteristics, vulnerability information, and attack behavior; The multi-agent decision-making module consists of multiple distributed agents. Each agent makes game strategy decisions based on local observation information and shares key data with other agents through a secure communication protocol. The MTD execution module is used to dynamically adjust network configuration parameters, including service port switching, virtual topology reconstruction, and honeypot deployment, and supports collaborative defense between the software layer and the network layer. The effectiveness evaluation module is used to quantify the defense effect, including attack interception rate, system resource consumption and user experience indicators, and to provide feedback on optimization decision-making strategies.

2. The network security dynamic defense system based on multi-agent joint game theory and moving target defense according to claim 1, characterized in that, The multi-agent decision-making module adopts a Bayesian game model, specifically including: Define the strategy space and utility function of both the attacker and defender, where the defender's utility function integrates network security status, resource overhead, and service continuity. The agent dynamically modifies its beliefs about attacker behavior and network status through a Bayesian probability update mechanism to optimize its defense strategy.

3. The network security dynamic defense system based on multi-agent joint game theory and moving target defense according to claim 1, characterized in that, The multi-agent decision-making module adopts a centralized training and distributed execution framework, specifically including: During the training phase, the agent learns a cooperative strategy based on global information and introduces noise to simulate an incomplete information environment. During the execution phase, the agent makes independent decisions based solely on local observation information and achieves limited information sharing through encrypted channels.

4. The network security dynamic defense system based on multi-agent joint game and moving target defense according to claim 3, characterized in that, The multi-agent decision-making module is trained using a multi-agent deep deterministic policy gradient algorithm. Each agent's Actor network outputs action policies based on local observations. The Critic network evaluates the value of actions using global information during the training phase to improve collaboration efficiency.

5. The network security dynamic defense system based on multi-agent joint game and moving target defense according to claim 1, characterized in that, The MTD execution module includes: The software layer MTD unit is used to dynamically switch application service interfaces, inject fake processes, or adjust runtime configurations to increase the difficulty for attackers to detect. The network layer MTD unit, combined with network function virtualization technology, enables lightweight IP address hopping, port randomization, and dynamic route adjustment; The MTD policy optimizer, based on deep reinforcement learning, is used to adaptively select the optimal defensive action according to the attack pattern.

6. A dynamic network security defense method based on the multi-agent joint game and moving target defense dynamic network security defense system according to any one of claims 1-5, characterized in that, Includes the following steps: 1) Collect network status data in real time through the environmental awareness module to identify potential attack behaviors; 2) Multi-agent systems initiate a Bayesian game model based on local observation information to predict attacker strategies and generate defensive actions; 3) The MTD execution module dynamically adjusts the network configuration based on the decision results and implements the software layer or network layer MTD strategy; 4) The utility evaluation module calculates the defense effect and feeds it back to the multi-agent decision-making module to optimize subsequent strategies.

7. The network security dynamic defense method based on multi-agent joint game and moving target defense according to claim 6, characterized in that, The Bayesian game model in step 2) includes: Construct the strategy trees of attackers and defenders, and define the game equilibrium conditions under incomplete information; The agent updates rules using Bayesian methods and dynamically adjusts the prior probability distribution of attacker types by combining historical interaction data.

8. The network security dynamic defense method based on multi-agent joint game and moving target defense according to claim 6, characterized in that, The MTD strategy generation in step 3) includes: For small-scale networks, software-layer MTD should be enabled first to reduce resource overhead through dynamic service masquerading. For large-scale networks, the combination of network layer MTD and NFV technologies is used to achieve efficient IP hopping and topology reconstruction.

9. The network security dynamic defense method based on multi-agent joint game and moving target defense according to claim 6, characterized in that, The utility evaluation indicators in step 4) include: Security metrics: attack interception success rate, average attack detection time; Performance metrics: System CPU / memory utilization, service request response latency; Economic indicator: The ratio of the cost of implementing a defense strategy to the potential losses caused by an attack.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the network security dynamic defense method as described in any one of claims 6-9.

Citation Information

Patent Citations

  • Network security dynamic defense system and method based on big data

    CN112788008A