APT attack defense method based on edge intelligence

By constructing an APT attack defense method based on optimal control and intelligent edge game theory, the shortcomings of existing technologies in identifying and responding to complex APT attacks are solved, efficient collaborative defense and resource optimization of edge nodes are achieved, and the security and stability of critical infrastructure are improved.

CN120639425APending Publication Date: 2025-09-12HANGZHOU NORMAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510931395.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing APT attack defense methods have limited identification capabilities when facing complex and covert attack variants, high management and maintenance costs, and insufficient applicability in resource-constrained environments. Traditional protection measures are difficult to effectively deal with advanced persistent threats.

Method used

An APT attack defense method based on optimal control and intelligent edge game theory is adopted. By constructing the edge node communication network topology adjacency matrix, the five states of the device life cycle are defined. Combined with optimal control theory and differential game model, the Nash strategy reinforcement learning mechanism of the multi-agent deep Q network is used to optimize the defense strategy, realize the coordinated defense of edge nodes and efficient resource utilization.

Benefits of technology

It improves the system's accuracy and timeliness in detecting new threats, reduces the risk of large-scale coordinated attacks, ensures the continuous and reliable operation of critical infrastructure, and achieves efficient security defense with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639425A_ABST
    Figure CN120639425A_ABST
Patent Text Reader

Abstract

The invention discloses an APT (Advanced Persistent Threat) attack defense method based on edge intelligence, which aims to improve the dynamic perception capability, strategy response capability and resource adaptability of an edge network to APT attacks, and comprises the following steps of: firstly, constructing a topological adjacency matrix of an edge node communication network and defining five states and state conversion processes of a life cycle of edge equipment; the method comprises the following steps: firstly, designing an attack defense model based on optimal control and differential game, then introducing a system evolution model based on hidden confrontation, and accurately simulating dynamic interaction of attack and defense in an actual edge network, secondly, designing an attack defense model based on optimal control and differential game, and finally, adopting a Nash strategy reinforcement learning mechanism based on a multi-agent deep Q network, optimizing an edge game strategy, and finally, obtaining an attack defense result. And the attack detection performance is improved. According to the method disclosed by the invention, the modeling precision and defense effectiveness of the system in a complex APT attack scene are remarkably improved, and the method is suitable for key infrastructure network environments with relatively high requirements on safety, timeliness and expandability, such as industrial Internet and Internet of Things.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of critical infrastructure network security and relates to an APT attack defense method based on edge intelligence, particularly to security defense technology for edge nodes, as well as a defense method and system against APT attacks. Background Art

[0002] Critical infrastructure networks (CINs), as a crucial component supporting the operation of modern society, are crucial for social stability and economic development. However, due to their complex structure and high degree of interdependence, CIN systems face increasingly complex and pervasive cybersecurity threats, particularly from Advanced Persistent Threat (APT) attacks. APT attacks, through gradual and covert means, can remain lurking within networks for extended periods, posing a serious threat to the data and operations of critical infrastructure. Edge nodes, as crucial components of CIN systems, not only provide critical communication and data processing functions but also meet the system's critical requirements for real-time, low-latency, and highly reliable communication. However, their accessibility, limited computing resources, and lack of real-time updates make them potential targets for APT attacks. Therefore, researching and implementing security defense technologies for edge nodes is urgent and necessary.

[0003] While traditional security measures such as firewalls and intrusion detection systems can detect and block known attack patterns to a certain extent, their effectiveness and adaptability are significantly insufficient in the face of increasingly complex and diverse APTs. Therefore, academia and industry urgently need innovative security defense technologies to address evolving threats. In recent years, researchers have proposed a variety of novel APT defense methods and technologies. For example, the HOLMES system analyzes real-world case studies to extract and identify the coordinated behavioral patterns that constitute APT activities, enabling high-precision attack detection with low false positive rates. The ATLAS framework utilizes natural language processing and sequential model learning techniques to restore attack steps from network audit logs, helping security analysts understand and respond to APT incidents. Furthermore, blockchain-based APT detection systems leverage the immutable nature of ledgers to ensure the protection and traceability of critical data, such as those in the Industrial Internet of Things.

[0004] However, these existing methods still face several challenges and limitations, such as limited ability to identify new and highly stealthy attack variants, high management and maintenance costs, and practical applicability issues in resource-constrained environments. Summary of the Invention

[0005] The purpose of this invention is to propose an APT attack defense method based on optimal control and intelligent edge game theory, aiming to improve the security of edge nodes in critical infrastructure networks and effectively defend against advanced persistent threats. In view of this, the present invention proposes a novel APT attack defense system and method based on edge intelligence. By simulating the strategic interaction between attackers and defenders, it studies how to optimize defense strategies and effectively respond to different types of APT attacks. In addition, combined with optimal control theory, while ensuring attack detection performance, it maximizes the utilization of the limited resources of edge nodes, thereby improving the security and reliability of the system. These innovative technologies are expected to not only improve the performance and effectiveness of existing defense systems, but also effectively respond to new threats and attack modes that may emerge in the future.

[0006] To achieve the above object, the technical solution of the present invention is as follows: An APT attack defense system and method based on edge intelligence includes the following steps: Construct a topological adjacency matrix of the edge node communication network to represent the set of edge devices in the system and their connection relationships; the set of edge devices is represented as , the connection relationship between edge devices is through the matrix To express it, where N is the total number of edge devices, Representation device and There is a communication channel between them. Indicates non-existence; Define the five states of the edge device life cycle, establish a differential equation to describe the state transition process, and simulate the state change process of the edge device under APT attack; wherein, the five states are: susceptible to infection state , exposure state , patch status , infection status and isolation , No. t time slots Represents the status of edge devices in the system, where Represents edge devices In a susceptible state , exposure state , infection status , Isolation status or patch status ; The state transition process is described as follows:

[0007] in is the initial system state; the five-state evolution can more finely reflect the behavioral evolution of edge nodes during an APT attack, and accurately depict the hidden evolution characteristics of the attack chain during an APT attack.

[0008] Based on the five states and their transition processes, an attack-defense model based on optimal control and differential game is constructed to monitor the behavior of attackers and defenders. Specifically, the following steps are included: (1) Constructing a covert countermeasure system evolution model , where L represents the set of participants in the APT attack and defense game, D represents the action space of the participants, u represents the control strategy space of the participants, X represents the state variables of the edge devices, F represents the system state evolution function, and J represents the set of participants' benefits; the participants include attackers and defenders; (2) Introducing attack and defense timing, attack and defense strategies, and cost functions into the evolutionary model of the covert adversarial system, a spatiotemporal strategy decision model based on differential games is constructed, in which the system state transition is controlled by the strategies of the participants, forming a dynamic attack and defense game process; (3) Apply optimal control theory to solve the spatiotemporal strategy decision model based on differential games and obtain the optimal attack and defense strategies under Nash equilibrium; The attack and defense model is optimized using the Nash strategy reinforcement learning mechanism based on the multi-agent deep Q network, and the attack and defense strategy is optimized by learning sequence features and environmental feedback.

[0009] Compared with the prior art, the present invention has the following advantages: (1) The present invention applies adversarial game theory to enable the system to evaluate and flexibly adjust defense strategies in real time, and make dynamic adjustments based on real-time monitored attack intelligence and changes in the network environment. This not only improves the accuracy and timeliness of the system's detection of new threats, but also ensures the long-term effective operation of the system in a complex network environment, and has high practicality and adaptability.

[0010] (2) This invention introduces intelligent edge game theory to enable the system to achieve collaborative defense and attack detection between edge nodes. The collaborative combat strategy not only improves the system's overall ability to resist attack threats, but also effectively reduces the risk of the system being subjected to large-scale, coordinated attacks, thereby ensuring the continuous and reliable operation of critical infrastructure.

[0011] (3) This invention not only introduces the traditional game state-strategy-reward triple modeling framework, but also further integrates attack and defense timing modeling with resource constraint optimization mechanisms, constructing the Hamilton-Jacobi-Bellman (HJB) equations based on optimal control theory. The application of optimal control theory helps the system effectively allocate and utilize the computing and storage capabilities of edge nodes under limited computing resources. Therefore, based on the comprehensive consideration of attack detection performance and resource utilization efficiency, this invention achieves a balance between system security and operability, providing reliable protection for the safe operation of critical infrastructure.

[0012] In summary, this invention possesses significant technical advantages and practical application value for defending against APTs in critical infrastructure. By incorporating APT attack modeling, edge collaboration mechanisms, optimal control optimization, and resource efficiency considerations into a game-theoretic model, this invention not only enhances the system's ability to respond to diverse and complex attack threats, but also achieves efficient security defense with limited resources, bringing new technological breakthroughs and development directions to the field of network security. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a node state transition model diagram.

[0014] Figure 2 Nash strategy reinforcement learning architecture diagram.

[0015] Figure 3 Network topology diagram of the simulation system.

[0016] Figure 4 Game state transition diagram.

[0017] Figure 5 Graph of the exact Nash equilibrium solution.

[0018] Figure 6 Graph of the approximate Nash equilibrium solution. DETAILED DESCRIPTION

[0019] The present invention will be further analyzed below in conjunction with the accompanying drawings.

[0020] The present invention proposes a method for defending against APT attacks based on optimal control and intelligent edge game theory. The technical solution adopted is: first, the optimal control theory is used to optimize the defense strategy of the system to ensure that under limited computing resources, security incidents of edge nodes can be effectively monitored and responded to. This method can not only reduce the operating cost of the system, but also maximize the security and detection performance of the system. Secondly, the intelligent edge game theory is introduced, and the Nash strategy reinforcement learning mechanism based on the multi-agent deep Q network is used to realize collaborative defense and attack detection between edge nodes. By combining optimal control theory and intelligent edge game theory, the present invention can effectively respond to dynamic and complex APT attacks and ensure the security and stable operation of critical infrastructure networks.

[0021] The implementation of the APT attack defense method based on optimal control and intelligent edge game theory includes the following steps and technical details: Step 1: Construct a topological adjacency matrix of the edge node communication network to represent the set of edge devices and their connection relationships; Step 2: Define the five states of the edge device lifecycle and establish differential equations to describe the state transition process. This is used to simulate the state changes of edge devices under APT attacks. In other words, a system evolution model based on covert adversarial attack is constructed. The five states of the device lifecycle and their state transition processes are defined, including susceptible state, exposed state, patched state, infected state, and isolated state. This generates a dynamic evolution model of attack and defense in the edge network, capturing system state changes and potential risks of APT attacks in real time. Step 3: Based on the five states and their transitions, an attack-defense model based on optimal control and differential game theory is constructed. By monitoring the behavior of attackers and defenders in real time, the defense strategy is dynamically adjusted to balance attack detection performance and resource utilization efficiency. Step 4: Use the Nash strategy reinforcement learning mechanism based on the multi-agent deep Q network to optimize the attack defense model. By learning sequence features and environmental feedback, the attack defense strategy is optimized to improve the system's adaptability and detection capabilities to APT attacks.

[0022] In step 1, the APT attacker uses a multi-stage, progressive attack method to gradually penetrate the target system, resulting in multiple changes in the state of edge devices. Therefore, the topological adjacency matrix is ​​used to represent the topology of the industrial equipment communication network. The set of edge devices is represented as , the connection relationship between edge devices is through the matrix To express, Representation device and There is a communication channel between them. Indicates that it does not exist.

[0023] In step 2, during the APT attack based on stealth, the device life cycle will go through five state transitions, including the vulnerable state , exposure state , patch status , infection status and isolation , No. t time slots Indicates the status of system devices, where Respectively represent devices In a susceptible state , exposure state , infection status , Isolation status and patch status The state transition process is as follows: Figure 1 shown.

[0024] APT attackers use deceptive information to induce device users to perform malicious operations, based on the probability The device falls into a vulnerable state. Nodes in the vulnerable state have potential threats but have not yet been actually infected. Once the device is infected according to the probability Entering the exposed state, attackers will exploit known vulnerabilities or zero-day vulnerabilities to infect the device, according to the probability Transform it into an infected state. The device in the infected state becomes the attacker's control node and actively moves laterally to infect other devices. When the detection system finds an infected node, it disconnects it from the network and isolates the device using a dedicated security gateway. The probability of an infected node being isolated is At this time, the exposed state node has a probability Fix known vulnerabilities through patching. After patching, nodes will Transformed into a vulnerable state. The differential equations of the lateral movement model of the state transition of the system node are expressed as shown in formula (1), where is the initial system state. The state function of the differential system is .

[0025] (1) In step three, the specific implementation process includes the following steps: (1) Evolutionary Model of Covert Countermeasure System The construction of ① represents the set of participants in the APT attack and defense game, where Indicates defender, Indicates attacker; ② represents the action space of the participants, where is the defender's behavior set, represents the set of attacker behaviors; ③ express t The control strategy space of participants at each moment, where yes t A hybrid control strategy for constant defenders, express t The defender chooses a defensive action at any time The probability of .at the same time, for t The attacker's hybrid control strategy at any moment, express t Attacker chooses attack action The probability of ; ④ It is the status variable of the device node; ⑤ represents the system state evolution function, where , , , , ; ⑥ is the payoff set of the participants, where represents the attacker’s profit, is the defender's payoff, and the payoff function and Determined jointly by the strategies of all participants; To obtain the optimal strategy, we need to solve the coupled Hammer-Jacobi-Bellman equations to find the Nash equilibrium. A Nash equilibrium is when each participant chooses the most advantageous strategy for themselves given the strategies of the other participants, known as a saddle point strategy.

[0026] The strategy pairs composed of the optimal strategies of the participants in the game equilibrium in the spatiotemporal strategy model based on differential games This is called the saddle point strategy of the model, and it satisfies the following conditions:

[0027] (2) The differential game strategy of the timing optimization mechanism, based on the evolutionary model of the covert adversarial system, introduces the triggering timing of the attacker and defender's behavior into the strategy space and defines a "timing-action" joint strategy structure to more realistically depict the dynamic interaction process of attack latency, sudden intrusion and defense response: ① Construction of attack and defense timing. In order to effectively deal with the resource consumption and the risk of being detected by attackers caused by frequent attacks, an exponential time interval strategy and a defensive behavior time strategy will be adopted. and aggressive behavior timing strategies The parameters are respectively and Exponential distribution and Given the memoryless nature of the exponential distribution, when an attacker receives defense feedback, they only learn the overall distribution of defense cycles, but not specific information about the next defense action. Furthermore, while the defender understands the distribution of attack intervals, they cannot precisely determine when the attacker executed their last attack action.

[0028] ② Construction of attack and defense strategies. In the evolutionary model of the covert adversarial system, the defender selects appropriate actions to prevent the attacker’s behavior, while the attacker attempts to select strategies that can bypass the defense, forming a dynamic attack and defense game. The attacker can control the parameters 、 and To manipulate the attack process in order to maximize the attack effect and increase the persistence of the attack. At the same time, the defender can control the parameters and To implement defense strategies, the system responds more quickly and accurately to exposure and infection states, thereby reducing the possibility of attackers gaining persistent control within the network.

[0029] Given a defender's behavior set: May include monitoring and analyzing network traffic , strengthen authentication , upgrade vulnerability patches and fixes and deploying industrial firewalls The attacker's strategy set is , which may include reconnaissance and tracking , inject malicious code , Patch Bypass , as well as deceptive information and social engineering etc. Indicates the time period All segmented continuous defined within N The set of dimensional functions, where t0 represents the start time of the time period and T represents the end time of the time period. Then the control strategies of the defender and the controller are and ,in . And the state transition probability is limited by resources, that is , , , , .

[0030] ③ Cost Function Definition. In the attack-defense process, the attacker's goal is to infiltrate the target software and its related systems at the lowest cost to gain access to the target organization. The defender's goal is to clean up the infected system at the lowest cost, repair the damage, and take measures to prevent the attacker's actions and restore the system to normal operation.

[0031] assumed Indicates that the attacker infected the node The reward obtained from the infected node may be sensitive information, access rights or control rights. The reward obtained by the attacker during the game is .

[0032] Successfully patching a node can effectively defend against malicious attacks. represents the reward for preventing potential attacks, and the reward for the defender to patch the exposed node is expressed as .

[0033] Quickly and effectively isolating infected nodes helps prevent the spread of threats. Indicates that the defender is isolated from the Therefore, the reward obtained by the defender isolating the node during the game is .

[0034] Considering the resource consumption caused by APT attack and defense, it includes four aspects: computing resource consumption, network bandwidth consumption, storage resource consumption, and service time. , , , They represent four different types of resource consumption of different nodes, among which resource consumption, network bandwidth consumption and storage resource consumption are calculated by CPU computing time, bandwidth usage time and memory usage within a period of time.

[0035] The cost of the defender is evaluated by the resource consumption of monitoring, patching and the service time lost when the node is in the infected state. The cost function of the defender is defined as:

[0036] Therefore, the defender's total payoff function is expressed as follows:

[0037] The attacker’s cost is evaluated by the average cost during the duration of the attack process. The cost of the attacker’s lateral movement is Re-infecting patched nodes requires more cost, but the rewards that come with it are also more generous. When the attacker's reward exceeds the cost, the attacker has the motivation to re-attack the system. In this case, the attacker is more inclined to perform reinfection operations to obtain greater potential benefits. The cost of the attacker performing reinfection operations is The attacker’s total profit function is expressed as follows:

[0038] (3) Optimization of Nash equilibrium strategy based on optimal control.

[0039] The defender may predict the attacker's type and possible strategies through the game history. At this time, the defender hopes to choose a strategy To maximize , and when the defender selects the optimal defense strategy , the attacker must choose an attack strategy To maximize the profit, the strategy is It is called a saddle point strategy and satisfies:

[0040] According to the optimal control theory, the saddle point strategy problem can be transformed into the Hamilton function as shown in the formula, where is the associated adjoint function.

[0041]

[0042]

[0043] Given is the optimal control strategy, at this time , ; Then when the accompanying function When, for ,have:

[0044] The Nash equilibrium strategy solution of the differential game model based on optimal control is transformed into the solution of a system of ordinary differential equations. The specific steps of the solution are as follows: (1) Initialize the differential game model , set the initial iteration counter ; (2) Constructing the system state evolution equations , forward calculation evolution state ; (3) Iterative optimization to obtain the optimal strategy: ① Forward calculation system status ; ② Construct Hamiltonian function and ; ③ Backward calculation of adjoint function ; ④ Calculate new strategies ; ⑤ Check the convergence conditions. If the convergence error meets the conditions, exit the iteration loop. (4) Return the optimal control strategy .

[0045] In step 4, in order to effectively solve the problem of strategic interaction between participants, the Nash strategy reinforcement learning mechanism based on the multi-agent deep Q network is adopted, such as Figure 2 As shown in Figure 2. The multi-agent deep Q network consists of two neural networks, the actor Q network (estimation network) and the target Q network (target network), where are the parameters of the estimated network, and is the parameter of the target network. In the APT game model, the environment is formed by the interaction between the participants and the outside world. Participants perceive the network environment by observing the proportion of nodes in a certain state, and form an observation state based on their own state. , Will be controlled by the policy parameters and Subsequently, Input into the LSTM neural network to learn sequence features. The output of LSTM will be used as the input of the FC layer and the fully connected layer. The fully connected layer combines the softmax operation to obtain the Q value corresponding to different actions of the participants. and , thereby guiding the participation in the combination -greed strategy selects the optimal game strategy The update formula of the participant's Q function is as follows, where It is a common reward pursued by participants, is the weight factor, It is learning efficiency, is the discount factor representing the decay of future rewards, is the control strategy for the next time step is the defender's payoff function, is the attacker’s payoff function.

[0046]

[0047] Participants execute the chosen strategy After that, the two parties interact to obtain the latest status and current rewards . Participants' training data It is stored in the experience pool, providing a data basis for the subsequent learning process. During the training process, historical data is randomly extracted from the experience pool. Form a minibatch as training data and gradually seek the optimal strategy through learning.

[0048] Multi-agent deep Q-network optimizes the loss function through the temporal difference error of environmental feedback During the online learning process, participants use gradient backpropagation to continuously update the parameters of the estimated network according to the loss function. The loss function compares the estimated value with the value predicted by the neural network to obtain a difference, which is used to update the network parameters. The specific loss function of the participant is shown in the following formula:

[0049] That and It is the estimated network output Q value, and is the output of the target network Q value.

[0050] The specific steps of the Nash strategy reinforcement learning algorithm based on the multi-agent deep Q network are as follows: (1) Parameter initialization, setting initial time k = 0 and parameters , LSTM parameters, target network parameters, initial state , initial strategy ; (2) Iterative storage of experience value, first update the time , then observe the current state and instant rewards , then save the current experience ; (3) Iterative update of network parameters, first randomly sampling from the experience pool BS samples, for each sample , calculate its importance sampling weight Time difference error and , , Then accumulate the changes in network parameters, including the defender's parameter changes Parameter changes with attackers , , The estimated network parameters are then updated and target network parameters ,

[0051] (4) Obtaining attack and defense strategies .

[0052] Example: The performance of the proposed differential game model was evaluated on a simulated, simplified ethanol distillation system testbed. The ethanol distillation system testbed is designed and built based on an actual production environment, recreating the ethanol distillation production process in a realistic scenario. Figure 3 The simulation system network topology is demonstrated. The system's production edge gateway uses a Siemens S7-300 PLC. Edge device sensors transmit analog values ​​such as temperature and liquid level to the PLC, which then issues commands such as opening and start / stop to actuators such as the solenoid valve (Actuator 1), condenser (Actuator 1), and reactor (Actuator 3). The PLC communicates with the host computer via the S7COMM protocol, uploading status information and feedback on parameter adjustment success or failure. It also receives commands from the host computer for start / stop, parameter adjustment, and status queries. The host computers, including Management Host 1 (MH1) and Management Host 2 (MH2), monitor the entire industrial control process through a human-machine interface (HMI).

[0053] Next, the specific implementation steps of the present invention will be described in detail with reference to the accompanying drawings: Step 1: Build a system evolution model based on covert adversarial technology, define the five states of the device lifecycle and their state transition process, including susceptible state, exposed state, patched state, infected state, and isolated state, generate a dynamic evolution model of attack and defense in the edge network, and capture the potential risks of system state changes and APT attacks in real time.

[0054] The set of attack strategies available in each game state is , , The corresponding defense policy set is: . This gives us the game state transition diagram, as Figure 4 As shown, They represent susceptible state, exposed state, infected state, isolated state and patched state respectively.

[0055] Step 2: Build an attack-defense model based on optimal control and differential game theory. By monitoring the behavior of attackers and defenders in real time, dynamically adjust the defense strategy to balance attack detection performance and resource utilization efficiency.

[0056] The precise Nash equilibrium solution of the APT attack defense model based on differential game is as follows: Figure 5 As shown, The approximate Nash equilibrium solution is shown in Figure 6. By comparing the Nash equilibrium solutions, it is found that the optimal attack strategy solution is less than , the best defense strategy solution is less than As a result, the policy reinforcement learning algorithm based on the multi-agent deep Q network can find the Nash equilibrium solution of the APT attack defense model with high accuracy.

[0057] Obviously, the above embodiments are provided as examples for the APT attack detection method and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications can be made based on the above description. Any modifications or variations made to the present invention remain within the scope of protection of the present invention.

Claims

1. A method for defending against APT attacks based on edge intelligence, characterized in that: The method comprises the following steps: Construct a topological adjacency matrix of the edge node communication network to represent the set of edge devices in the system and their connection relationships; Define five states in the edge device lifecycle and establish a differential equation to describe the state transition process, which is used to simulate the state change process of edge devices under APT attacks. The five states are: susceptible state, exposed state, patched state, infected state, and isolated state. Based on the five states and their transition processes, an attack-defense model based on optimal control and differential game theory is constructed to monitor the behavior of attackers and defenders. The attack and defense model is optimized using the Nash strategy reinforcement learning mechanism based on the multi-agent deep Q network, and the attack and defense strategy is optimized by learning sequence features and environmental feedback.

2. The method according to claim 1, characterized in that The set of edge devices is represented as , the connection relationship between edge devices is through the matrix To express it, where N is the total number of edge devices, Representation device and There is a communication channel between them. Indicates that it does not exist.

3. The method according to claim 1, characterized in that The five states of the device life cycle include the vulnerable state , exposure state , patch status , infection status and isolation , No. t time slots Represents the status of edge devices in the system, where Represent edge devices In a susceptible state , exposure state , infection status , Isolation status or patch status ; The state transition process is described as follows: in This is the initial system state.

4. The method according to claim 1, characterized in that The construction process of the attack defense model includes the following steps: Constructing a covert countermeasure system evolution model , where L represents the set of participants in the APT attack and defense game, D represents the action space of the participants, u represents the control strategy space of the participants, X represents the state variables of the edge devices, F represents the system state evolution function, and J represents the set of participants' benefits; the participants include attackers and defenders; By introducing attack and defense timing, attack and defense strategies, and cost functions into the evolutionary model of the covert adversarial system, a spatiotemporal strategy decision model based on differential games is constructed, in which the system state transition is controlled by the strategies of the participants, forming a dynamic attack and defense game process. Optimal control theory is applied to solve the spatiotemporal strategy decision model based on differential games, and the optimal attack and defense strategies under Nash equilibrium are obtained.

5. The method according to claim 4, characterized in that: The cost function is specifically: assumed Indicates that the attacker infected the node The reward obtained by the attacker during the game is , represents the reward for preventing potential attacks, and the reward for the defender to patch the exposed node is expressed as , Indicates that the defender is isolated from the The reward obtained by the defender isolating the node during the game is .

6. Assumptions , , , Represents four types of resource consumption of different nodes, among which resource consumption, network bandwidth consumption and storage resource consumption are calculated by CPU computing time, bandwidth usage time and memory usage within a period of time; The defender’s cost is evaluated by the resource consumption of monitoring, patching, and the service time lost when the node is in the infected state. The defender’s cost function is defined as: Therefore, the defender's total payoff function is expressed as follows: The attacker's cost is evaluated by the average cost during the duration of the attack process: the attacker's cost of lateral movement is , the cost for the attacker to perform reinfection is , then the attacker’s total profit function is expressed as follows: 。 7. The method according to claim 4, characterized in that: The optimal attack and defense strategies under Nash equilibrium are as follows: (1) Initialize the differential game model , set the initial iteration counter ; (2) Constructing the system state evolution equations , forward calculation evolution state ; (3) Iterative optimization to obtain the optimal strategy: ①Forward calculation system status ; ②Construct Hamiltonian function and ; ③ Backward calculation of adjoint function ; ④Calculate new strategies ; ⑤ Check the convergence conditions. If the convergence error meets the conditions, exit the iteration loop. (4) Return the optimal control strategy .

8. The method according to claim 1, characterized in that: The Nash strategy reinforcement learning mechanism based on the multi-agent deep Q network includes an estimation network and a target network. The environment is formed by the interaction between the participants and the outside world. The participants form the observation state through the observation state and their behavior. And input into the LSTM neural network to learn sequence features and environmental feedback.

9. The method according to claim 7, characterized in that: The output of the LSTM neural network is used as the input of the fully connected layer. The fully connected layer combines the softmax operation to obtain the Q value corresponding to different actions. and , to guide participants based on -greed strategy selects the optimal game strategy .

10. The method according to claim 7, characterized in that: The participant’s Q function is updated by the following formula, in It is a common reward pursued by participants, is the weight factor, It is learning efficiency, is the discount factor representing the decay of future rewards.

11. The method according to claim 7, characterized in that: The loss function of the participant is, in and is the Q value of the estimated network output, and is the Q value output by the target network, E is the expected value function.

Citation Information

Cited By

  • A social internet of things malware defense method and system based on coalition game

    CN122457307A