A Dynamic Defense Strategy Iterative Optimization Method

By constructing an offensive-defense game framework and optimizing defense strategies using reinforcement learning algorithms, the shortcomings of existing technologies in optimizing MTD deployment strategies are addressed. This achieves synergistic optimization of defense strategies in both time and space dimensions, thereby improving the long-term defense effectiveness of the defense system.

CN121841864BActive Publication Date: 2026-05-26GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU UNIVERSITY
Filing Date
2026-03-13
Publication Date
2026-05-26

Smart Images

  • Figure CN121841864B_ABST
    Figure CN121841864B_ABST
Patent Text Reader

Abstract

This invention provides a dynamic defense strategy iterative optimization method, comprising: defining a heterogeneous resource set of candidate configurations for the defense system and defining core physical parameters; defining an attack and defense strategy space based on the core physical parameters, reconstructing the mapping relationship between attack resources and time, and modeling the time accumulation effect of successful attacks to construct an attack and defense game framework; solving for the Nash equilibrium point of the switching path to obtain the optimal dwell period and the corresponding unit-time equilibrium utility potential energy; mapping the reward signal of the reinforcement learning environment and iteratively updating the state-action value table; calculating the action selection probability, and determining the iteratively optimized defense strategy tuple based on the action selection probability. The innovative dual-layer coupling architecture of micro-game evaluation and macro-intelligent planning constructed using this method directly maps the Nash equilibrium utility potential energy calculated at the bottom layer to the environment reward signal of the upper-layer reinforcement learning, realizing the full-element collaborative optimization of the defense strategy in both time and space dimensions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and in particular to a dynamic defense strategy iterative optimization method. Background Technology

[0002] With the increasing complexity and automation of cyberattacks, traditional static defense systems based on perimeter protection and signature matching (such as firewalls and IDS) are no longer sufficient to cope with advanced persistent threats (APTs) and attacks exploiting unknown vulnerabilities. Cyberspace presents an asymmetrical situation where attacks are easy to launch but difficult to defend, giving attackers ample time to discover vulnerabilities in static systems, while defenders are in a passive and defensive position. To reverse this situation, Moving Target Defense (MTD) technology has emerged. MTD increases the randomness and uncertainty of a system by continuously and dynamically changing its attack surface (such as IP address, port, operating system version, software configuration, etc.), disrupting the attacker's cyber kill chain and making it difficult for them to establish a stable attack foothold within a limited time.

[0003] Existing technologies for optimizing MTD deployment strategies have several drawbacks:

[0004] Existing technologies typically employ traditional game theory models for MTD strategy deployment optimization, often simplistically assuming that attack time is inversely proportional to the resources invested. This leads to the physical paradox that "when attack resources approach infinity, attack time approaches zero." It ignores the actual impact of factors such as network physical latency (RTT), protocol handshakes, and information gathering on attack time, which can easily result in overly aggressive or ineffective defense strategies.

[0005] Existing MTD deployment switching strategies (such as random jumps and fixed-sequence migrations) typically assume that all alternative configurations are homogeneous in terms of security and migration costs, or perform one-time optimizations based solely on static vulnerability scores (CVSS). However, in real-world single-node heterogeneous systems, different configurations (such as different OS kernels or web service versions) exhibit natural heterogeneity, with inherently different static non-heterogeneous attack surface standard intrusion times and migration deployment times. Traditional static game theory models struggle to perform long-term, adaptive path planning for such fine-grained heterogeneous differences in dynamically changing environments, easily getting trapped in local optima and failing to maximize long-term defense benefits.

[0006] Furthermore, existing technologies struggle to couple the optimization of time parameters and spatial path planning in MTD deployment strategies. They often have to solve the two issues of switching cycle selection and switching path planning as two separate problems, or use highly complex genetic algorithms for offline calculations. Lacking a unified metric, path planning cannot be aware of physical time constraints, and cycle settings ignore the differences in the quality of target configurations, making it difficult to achieve full-element collaborative optimization of MTD deployment strategies.

[0007] Therefore, it is necessary to provide a dynamic defense strategy iterative optimization method to achieve coordinated iterative optimization of the defense strategy in both time and space dimensions. Summary of the Invention

[0008] The purpose of this invention is to provide a dynamic defense strategy iterative optimization method to achieve a deep integration of refined evaluation of switching cycle selection and global optimization of switching path planning, so that the optimized defense strategy can meet both physical stability requirements and global defense effectiveness.

[0009] Firstly, the dynamic defense strategy iterative optimization method provided by this invention includes: defining a heterogeneous resource set for candidate configurations of the defense system; defining core physical parameters for mobile target defense scenarios based on historical attack and defense data, including defining the standard intrusion time of a non-heterogeneous attack surface, the attacker's resource input coefficient, the defender's physical migration time, and the rigid time limit of attack events; defining the defender's strategy space, the attacker's strategy space, reconstructing the mapping relationship between attack resources and time, and modeling the time accumulation effect of successful attacks based on the core physical parameters; defining the attack and defense game sequence; and constructing an attack and defense game framework based on the attack and defense game sequence and the time accumulation effect of successful attacks. The attack-defense game framework solves for the Nash equilibrium points of both the attacker and defender for each pair of switching paths in the heterogeneous resource set to obtain the optimal dwell time and the corresponding unit-time equilibrium utility potential under each switching path. The unit-time equilibrium utility potential is normalized to map as a reward signal of the reinforcement learning environment. Based on the reward signal, the state-action value table is updated iteratively by interacting with the virtual environment using a reinforcement learning algorithm. Based on the updated state-action value table and the Softmax probability mechanism, the action selection probability is calculated. The next target configuration for the defense system switching is determined according to the action selection probability. The corresponding optimal period is retrieved according to the next target configuration to form the iteratively optimized defense strategy tuple.

[0010] The beneficial effects of the dynamic defense strategy iterative optimization method provided by this invention are as follows: The innovative dual-layer coupling architecture of micro-game evaluation and macro-intelligent planning directly maps the Nash equilibrium utility potential energy calculated at the bottom layer to the environmental reward signal of the upper layer reinforcement learning, breaking the decision-making separation between "switching cycle selection" and "heterogeneous path planning" in traditional methods, and realizing the full-element collaborative optimization of defense strategy in time and space dimensions.

[0011] In one possible embodiment, core physical parameters for mobile target defense scenarios are defined based on historical attack and defense data, including: defining the standard intrusion time of a non-heterogeneous attack surface as the average expected time required for an attacker to break through the current static configuration under standard attack intensity; defining the attacker's resource investment coefficient as the multiple of the attacker's actual resource investment relative to the baseline resource investment, where the attacker's baseline resource investment is the resource investment under standard attack intensity; the defender's physical migration time as the physical time required for the defense system to complete a complete attack surface switching operation; and the rigid time limit of the attack event as the minimum theoretical time required for the attacker's attack process.

[0012] In another possible embodiment, the defender's policy space, attacker's policy space, the mapping relationship between attack resources and time, and the time accumulation effect of a successful attack are defined based on core physical parameters. This includes defining the defender's behavior as performing system migration with a switching cycle, requiring the switching cycle to be longer than the defender's physical migration time, and defining the defender's policy space as follows: Define a nonlinear stability penalty term. Used to constrain the choice of defenders, where... Represents the defender's strategy space. Indicates the switching cycle. Indicates the physical migration time of the defender. This represents the nonlinear stability penalty term. This represents the stability cost coefficient. The penalty index is represented; the attacker's action is defined as choosing a multiple of resource allocation to accelerate the attack process, and the attacker's policy space is... ,in, Represents the attacker's policy space. This represents the attacker's resource investment coefficient; the mapping relationship between attack resources and time satisfies the following formula: ,in, A scale parameter indicating the time it takes for an attack to succeed. This indicates the standard intrusion time for non-heterogeneous attack surfaces. This indicates a rigid time limit for an attack event. This represents the logarithmic decay adjustment coefficient; the cumulative distribution of successful attacks satisfies the following formula: ,in, Indicates the attack time is The attacker's resource investment coefficient is The probability of a successful attack at that time. Indicates the attack time. Indicates shape parameters.

[0013] In other possible embodiments, defining the attack-defense game timing includes: defining that within a single switching cycle, the attack first enters the defense migration window and then enters the effective defense window; and determining the time when both the attacker and defender gain benefits by comparing the attack time with the effective defense window within the corresponding switching cycle.

[0014] A game theory framework based on the temporal progression of attack and defense and the cumulative effect of successful attacks is constructed, including: constructing the defender's payoff function that satisfies the following formula: ,in, This indicates that the attacker's resource investment coefficient is... In the case of the defender switching cycles Internal benefits, This represents the value of defense operations per unit of time. Indicates the switching period The expected duration for which the internal system is in a secure service state. The nonlinear stability penalty term is represented; the attacker's payoff function is constructed to satisfy the following formula: ,in, This indicates that when the resource input coefficient is In the case of attackers switching cycles Internal benefits, This indicates that the attacker has complete control over the system's gains per unit of time. Indicates during the switching cycle The expected duration for which the internal system is under attacker control. Indicates the expected reward for the process. This indicates that the attacker's resource investment coefficient is [value missing]. The required linear resource cost rate.

[0015] Based on the core physical parameters and the attack-defense game framework, the Nash equilibrium points of the attacking and defending sides are solved for each pair of switching paths in the heterogeneous resource set to obtain the optimal period and corresponding unit-time equilibrium utility potential energy under each switching path. This includes: traversing all possible switching paths in the heterogeneous resource set, substituting the physical migration time of switching from the current configuration to the target configuration according to the switching path, and the standard intrusion time of the attacker breaking through the non-heterogeneous attack surface of the target configuration under the standard attack strength into the attack-defense game framework; assuming that both the attacking and defending sides are trying to maximize their own benefits, performing Nash equilibrium solutions to obtain the optimal period and unit-time equilibrium utility potential energy under each switching path; and generating the optimal period matrix and utility potential energy matrix based on the Nash equilibrium solution results of all switching paths in the heterogeneous resource set.

[0016] The optimal period and unit-time equilibrium utility potential energy under each switching path are obtained by solving the Nash equilibrium, including defining the attacker's optimal response function. satisfy Calculate the given switching period Attack strength that maximizes attack benefits ,in, This represents the attacker's payoff function. This represents the maximum resource investment coefficient by the attacker; it defines the optimal response function for the defender. satisfy Calculate the resource input coefficient for a given attacker. The switching cycle that maximizes the benefits of defense. ,in, The defender's payoff function, Indicates the physical migration time for switching paths. This indicates the maximum value of the switching cycle; initialize the attacker's resource input coefficient and the switching cycle, and perform alternating iterative optimization of the defense strategy and the attack strategy based on the attacker's best response function and the defender's best response function, and calculate the strategy offset in each iteration; when the strategy offset is less than the preset convergence threshold or the number of iterations reaches the set maximum value, output the Nash equilibrium point to obtain the optimal cycle and unit time equilibrium utility potential of the switching path.

[0017] The unit-time equilibrium utility potential energy is normalized and mapped to the reward signal of the reinforcement learning environment. Based on the reward signal, the state-action value table is iteratively updated using a reinforcement learning algorithm in interaction with the virtual environment. This includes: the unit-time equilibrium utility potential energy normalization calculation satisfies the following formula: ,in, This indicates a reward signal that reinforces the learning environment. This represents the instant reward function. Indicates time The system is running at the 1st Number configuration, Indicates at time The system selects to change the environment from configuration. Migrate to configuration , This represents the utility potential matrix, which represents the unit-time equilibrium utility potential energy of each switching path. Indicates switching paths The unit-time equilibrium utility potential energy, This represents the minimum value in the utility potential matrix. This represents the maximum value in the utility potential matrix. Represents a non-negative incentive bias; the state-action value table is iteratively updated based on the reward signal from the reinforcement learning environment. ,in, Indicates time State-action value Indicates the learning rate. Indicates the discount factor. Indicates time The maximum estimated value of all possible actions under the given state.

[0018] Based on the updated state-action value table and the Softmax probability mechanism, the action selection probability is calculated. The next target configuration for the defense system switching is determined according to the action selection probability, including: based on the state... The probability of selecting the target configuration action is calculated according to the following formula: ,in, Indicates based on state Select through action The probability of selecting the action to switch target configuration. Indicates the state Next action State-action value Indicates the temperature coefficient. This represents the action space for switching the target configuration in the system. Indicates the state Next action The state-action value is calculated; the calculated action selection probabilities are compared, and the next target configuration for the defense system to switch is determined based on the action corresponding to the maximum action selection probability. Attached Figure Description

[0019] Figure 1 A flowchart illustrating an iterative optimization method for dynamic defense strategies provided in an embodiment of the present invention;

[0020] Figure 2 This is a schematic diagram of an electronic device structure provided in an embodiment of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed following the word and its equivalents, but do not exclude other elements or objects.

[0022] This embodiment provides a method for iterative optimization of dynamic defense strategies.

[0023] See the instruction manual appendix Figure 1 The dynamic defense strategy iterative optimization method includes:

[0024] S101: Defines a heterogeneous resource set for candidate configurations of the defense system, and defines the core physical parameters for mobile target defense scenarios based on historical attack and defense data, including defining the standard intrusion time of non-heterogeneous attack surfaces, the attacker's resource investment coefficient, the defender's physical migration time, and the rigid time limit of attack events.

[0025] In one possible embodiment, core physical parameters for mobile target defense scenarios are defined based on historical attack and defense data, including: defining the standard intrusion time of a non-heterogeneous attack surface as the average expected time required for an attacker to break through the current static configuration under standard attack intensity; defining the attacker's resource investment coefficient as the multiple of the attacker's actual resource investment relative to the baseline resource investment, where the attacker's baseline resource investment is the resource investment under standard attack intensity; the defender's physical migration time as the physical time required for the defense system to complete a complete attack surface switching operation; and the rigid time limit of the attack event as the minimum theoretical time required for the attacker's attack process.

[0026] In one specific embodiment, a heterogeneous resource set for candidate configurations of the defense system is defined, and core physical parameters for the mobile target defense scenario are defined based on historical attack and defense data. This quantifies the adversarial process between the attacker and defender in the mobile target defense scenario, constructing the physical environment foundation for the attack and defense game. Specifically, in the mobile target defense scenario, the defender attempts to eliminate potential threats through periodic system migration (cleaning), while the attacker attempts to invest computational resources to breach the system before the defense cycle ends.

[0027] The initialization of the heterogeneous resource set is as follows: Define the defense system to have a set containing... A heterogeneous set of candidate configurations. When the defender performs a move operation, that is, when switching from the current configuration to a new configuration, the new configuration is selected from the heterogeneous set of configurations.

[0028] Standard intrusion time for non-heterogeneous attack surfaces For: A standardized benchmark value established to measure the inherent security of a system. (Static Non-Heterogeneous Attack Surface Standard IntrusionTime): This represents the average expected time required for an attacker to breach the current defense configuration under standard attack strength (i.e., the attacker's baseline resource investment). Game Theory Context: This parameter reflects the underlying code complexity and vulnerability density of the system, serving as a baseline for the game. An attacker must invest additional resources to breach the system faster than this time; a defender must switch over before this time expires to prevent an attack.

[0029] attacker resource input coefficient The attacker's core decision variable is not to directly choose the attack time, but to choose how much computing power or tool resources to invest to accelerate the attack process. (Attacker Invested Resources): Indicates that the attacker's actual resource investment is a multiple (dimensionless) relative to the baseline resource investment. When At that time, the corresponding standard intrusion time. This indicates that the attacker is accelerating the attack by increasing parallel computing power or using advanced exploit tools. This reflects the attacker's economic cost. The more intense the attack ( The larger the value, the faster the attack can theoretically be achieved, but the cost of the attack also increases linearly or non-linearly. Attackers need to find a balance between the benefits of a rapid breakthrough and the high cost of resources.

[0030] Defender Physical Migration Time This refers to the time window during which a defender switches from its current configuration to a new one, inevitably resulting in service unavailability or performance degradation due to physical processes such as image loading, service restart, and data synchronization. The physical meaning of the standard intrusion time for a non-heterogeneous attack surface represents the physical time required for the defense system to complete a full attack surface switchover operation. For the defender, This constitutes the time cost of defense. The more frequent the switching (the shorter the cycle), the greater the time cost. The higher the percentage of total time spent on switching, the more compromised the system's availability becomes, which prevents defenders from shortening the switching cycle indefinitely.

[0031] Rigid time limit for attack events This indicates that limitations include network physical latency (RTT), the number of TCP / IP handshake interactions, attack stealth requirements, and information gathering processes, regardless of the resources invested. The rigid time limit of an attack event introduces a physical boundary to the game, meaning that even with unlimited resources, an attacker cannot achieve an instantaneous breakthrough. This provides a favorable physical barrier for the defender. The introduction of this rigid time limit resolves the physical paradox in traditional models where "unlimited resources mean zero time."

[0032] S102: Define the defender's strategy space and the attacker's strategy space based on the core physical parameters, reconstruct the mapping relationship between attack resources and time, and model the time accumulation effect of successful attacks. Define the attack and defense game sequence, and construct the attack and defense game framework based on the attack and defense game sequence and the time accumulation effect of successful attacks.

[0033] In one possible embodiment, the defender's policy space, attacker's policy space, the mapping relationship between attack resources and time, and the time accumulation effect of a successful attack are defined based on core physical parameters. This includes defining the defender's behavior as performing system migration with a switching cycle, requiring the switching cycle to be longer than the defender's physical migration time, and defining the defender's policy space as follows: Define a nonlinear stability penalty term. Used to constrain the choice of defenders, where... Represents the defender's strategy space. Indicates the switching cycle. Indicates the physical migration time of the defender. This represents the nonlinear stability penalty term. This represents the stability cost coefficient. This represents the penalty index. The attacker's action is defined as choosing a multiple of resource allocation to accelerate the attack process; the attacker's policy space is... ,in, Represents the attacker's policy space. This represents the attacker's resource investment coefficient. The mapping relationship between attack resources and time satisfies the following formula: ,in, A scale parameter indicating the time it takes for an attack to succeed. This indicates the standard intrusion time for non-heterogeneous attack surfaces. This indicates a rigid time limit for an attack event. This represents the logarithmic decay adjustment coefficient. The cumulative distribution of successful attacks satisfies the following formula: ,in, Indicates the attack time is The attacker's resource investment coefficient is The probability of a successful attack at that time. Indicates the attack time. Indicates shape parameters.

[0034] The timing of the attack-defense game is defined as follows: within a single switching cycle, the player first enters the defense migration window and then the effective defense window; based on the attack time and the effective defense window, the attack-defense game results within the corresponding switching cycle are compared to determine the time when both the attacker and defender gain benefits.

[0035] A game theory framework based on the temporal progression of attack and defense and the cumulative effect of successful attacks is constructed, including: constructing the defender's payoff function that satisfies the following formula: ,in, This indicates that the attacker's resource investment coefficient is... In the case of the defender switching cycles Internal benefits, This represents the value of defense operations per unit of time. Indicates the switching period The expected duration for which the internal system is in a secure service state. This represents the nonlinear stability penalty term. The attacker's payoff function is constructed to satisfy the following formula: ,in, This indicates that when the resource input coefficient is In the case of attackers switching cycles Internal benefits, This indicates that the attacker has complete control over the system's gains per unit of time. Indicates during the switching cycle The expected duration for which the internal system is under attacker control. Indicates the expected reward for the process. This indicates that the attacker's resource investment coefficient is [value missing]. The required linear resource cost rate.

[0036] In one specific embodiment, the defender's actions are defined as occurring at fixed intervals. Perform system migration (cleanup). Defender strategy space ( ) represents the set of all switching cycles that meet service availability requirements, i.e. Define engineering constraints: Unlike traditional models that only require... Therefore, this embodiment takes into account the switching cycle. Approaching the migration time infinitely At this time, the system will be in a state of frequent restarts and oscillations, leading to serious engineering problems such as data loss and connection interruptions. Therefore, a nonlinear stability penalty term is defined as follows: ,in As a penalty index, This is the stability cost coefficient, used to quantify the extremely high losses caused by such impractical operations in subsequent revenue functions.

[0037] An attacker's action is defined as selecting the resource input coefficient. To accelerate the breach process. Attacker's strategy space ( Considering the economic cost and effectiveness of the attack, the attack intensity must be greater than or equal to the benchmark value, i.e. The reset perception mechanism assumes that the attacker has continuous detection capabilities, and that the attacker will reset their perception after the defender completes the migration (after...). (Duration) and after service recovery, attackers immediately detect and exploit the resources invested. Initiate a new attack process.

[0038] Amdahl's Law is introduced to physically model the attack process, establishing a mapping relationship between attack resources and time. The mapping logic posits that the attack process consists of "accelerable computational components" (such as brute-force password cracking and vulnerability scanning) and "non-accelerable rigid components" (such as network RTT, protocol handshakes, and stealthy waiting). The scale parameter is reconstructed as a characteristic value (scale parameter) representing the attack's success time. Defined as attack strength The function satisfies the following formula: , A scale parameter indicating the time it takes for an attack to succeed. This represents the logarithmic decay adjustment coefficient. The formula for the scaling parameter of attack success time indicates that as attack resources... With the increase of the attack time, the attack time can only approach the rigid time limit of the attack event infinitely, but cannot break through the limit, thus ensuring the physical reality of the model.

[0039] Modeling the APT attack process based on Weibull distribution to characterize the time accumulation effect of "reconnaissance-weaponization-exploitation" in APT attacks: Assuming the attack success time follows a Weibull distribution, the cumulative distribution of attack time is as follows: ,in, Indicates the attack time is The attacker's resource investment coefficient is The probability of a successful attack at that time. Indicates the attack time. Represents shape parameters, and This indicates that as the attack continues, the attacker's understanding of the target system gradually deepens, and the conditional probability density of a successful attack shows a monotonically increasing trend, which is consistent with the practical characteristics of APT attacks: slow incubation in the early stage and rapid breakthrough in the later stage.

[0040] In one possible implementation, a single switching cycle is used. The timeline is divided into two phases: the defense migration window and the effective defense window. Two possible game outcomes are defined based on the attack duration. The specific attack-defense game sequence logic is as follows:

[0041] The defense migration window is the beginning phase of each switching cycle ( The defense migration window is the time window during which the defender performs system resets or heterogeneous environment switching operations. Within this window, the defender is unavailable or in a high-overhead state, performing image loading, IP switching, or service restarts. While there is no risk of attack during the migration window, inherent migration costs are incurred. Meanwhile, due to drastic changes in target characteristics or temporary service interruptions, the attacker's original attack chain becomes invalid, the attack progress is forcibly reset to zero, and the attacker is in a reconnaissance and waiting state.

[0042] The effective service window is: at any time Once the system has completed the migration and is open to external services, it enters an effective defense window. The duration of this effective service window is: At that moment Once the attacker detects service recovery, they immediately allocate resources to launch a new attack. Whether the attacker can breach the system before the window switches ends depends on the attack duration. With effective service window duration The speed-based relationship is as follows: a successful attack must be completed within the entire switching cycle, while a failed attack indicates that the attacker failed to breach the defense within the entire switching cycle, and the defender's execution of the MTD (Mean Time To Divide) interrupts the attacker's attack process. The attack time follows a Weibull distribution.

[0043] Depending on the attack duration, the switching cycle will evolve into two mutually exclusive physical states, directly determining the gains for both the attacker and defender: When If the attack time exceeds the effective service window, the defense is successful. This means that the attacker continuously attempted the attack (accumulating progress) throughout the entire effective service window, but had not yet completed the final exploitation or privilege acquisition. The switching cycle ended, and the defender successfully defended the cycle, preventing the system from being compromised. In other words, an attack is considered successful if its execution time is less than the duration of the effective service window. This means that the attacker can achieve the desired attack time at any given moment. Successfully breached the system and gained control, from [time period]. From now until the end of the switching cycle ( The system is under control, the defenders lose control in the latter half of the cycle, and the attackers gain resource control benefits.

[0044] When the duration of a switching cycle reaches At this time, regardless of whether the current system is in a secure or controlled state, the defender will forcibly trigger the next round of migration operations. The system's memory, processes, and network status will be reset, and the controlled window established by the attacker will immediately end due to the change in environment. The game will enter the next cycle of defense migration window, realizing the forced switch of resource occupation status.

[0045] In a specific embodiment, the time when the attacker and defender gain benefits in each switching cycle can be determined based on the game sequence. Based on this, the average benefit per unit time for both the attacker and the defender can be defined.

[0046] The defender's goal is to maximize the effective security service duration per unit time while avoiding system instability caused by excessively frequent switching. The defender's payoff function is constructed as follows: . This indicates that the attacker's resource investment coefficient is... In the case of the defender switching cycles Internal benefits, This represents the value of defensive operations per unit of time (benchmark revenue). Indicates the switching period The expected duration for which the internal system remains in a secure service state; physically, this means the secure time would be [duration if an attack were successful]. If the attack fails, the safe period is the entire effective service window, based on game theory timing analysis. The calculation formula is: , express. This represents a nonlinear stability penalty term, used to mathematically constrain the choice of defenders. This unrealistic behavior ensures that when the effective window approaches zero, the defense cost approaches infinity, thus forcibly preserving the service buffer period.

[0047] The attacker's goal is to obtain the maximum control gain with the minimum resource cost. This embodiment introduces a process reward mechanism based on the kill chain to construct the attacker's reward function, as follows: . This indicates that when the resource input coefficient is In the case of attackers switching cycles Internal benefits, This indicates that the attacker has complete control over the system's gains per unit of time. Indicates during the switching cycle The expected duration for which the internal system is under attacker control. The calculation formula is: . This indicates the expected reward for the process / intelligence, even if the attack does not completely breach the system. The model posits that information accumulated by attackers during the early reconnaissance phase (such as port fingerprints and vulnerability intelligence) remains valuable. This breaks the zero-gain deadlock, forcing attackers to maintain [a certain level of awareness / capacity]. The active state can solve the problem of attackers giving up attacks when faced with a short window in existing technologies. This leads to the problem of trivial deadlock resolution. This indicates that the attacker's resource investment coefficient is [value missing]. The required linear resource cost rate.

[0048] S103: Based on the core physical parameters and the attack-defense game framework, solve the Nash equilibrium point of the attacking and defending sides for each pair of switching paths in the heterogeneous resource set to obtain the optimal dwell period and the corresponding unit time equilibrium utility potential energy under each switching path.

[0049] In one possible embodiment, all possible switching paths in the heterogeneous resource set are traversed, and the physical migration time of each switching path from the current configuration to the target configuration, and the standard intrusion time of the attacker breaking through the non-heterogeneous attack surface of the target configuration under the standard attack strength, are substituted into the attack and defense game framework; it is assumed that both the attacker and the defender are trying to maximize their own benefits, and the optimal period and unit time equilibrium utility potential energy under each switching path are obtained by performing Nash equilibrium solution; the optimal period matrix and utility potential energy matrix are generated based on the Nash equilibrium solution results of all switching paths in the heterogeneous resource set.

[0050] In one possible embodiment, the optimal period and unit-time equilibrium utility potential energy under each switching path are obtained by solving the Nash equilibrium, including: defining the attacker's optimal response function. satisfy Calculate the given switching period Attack strength that maximizes attack benefits ,in, This represents the attacker's payoff function. This represents the maximum resource investment coefficient by the attacker; it defines the optimal response function for the defender. satisfy Calculate the resource input coefficient for a given attacker. The switching cycle that maximizes the benefits of defense. ,in, The defender's payoff function, Indicates the physical migration time for switching paths. This indicates the maximum value of the switching cycle; initialize the attacker's resource input coefficient and the switching cycle, and perform alternating iterative optimization of the defense strategy and the attack strategy based on the attacker's best response function and the defender's best response function, and calculate the strategy offset in each iteration; when the strategy offset is less than the preset convergence threshold or the number of iterations reaches the set maximum value, output the Nash equilibrium point to obtain the optimal cycle and unit time equilibrium utility potential of the switching path.

[0051] In a specific embodiment, the core physical parameters of the defined moving target defense scenario are substituted into the constructed attack-defense game framework to solve for the optimal strategy for a specific switching path. The switching paths for which the optimal strategy is to be solved are all possible switching paths in the heterogeneous resource set, that is, all possible switching paths from the current configuration determined by the heterogeneous resource set. Switch to target configuration The switching path. The process of finding the optimal strategy for a switching path is as follows:

[0052] Incorporating path-specific physical parameters into the attack-defense game framework, specifically the physical migration time for switching the path from the current configuration to the target configuration, we can define this as the physical migration time. When replacing the definition, the one applied Substituting this into an attack-defense game framework, the standard intrusion time for an attacker to breach a target's non-heterogeneous attack surface configuration under standard attack strength will be... The baseline attack time applied when the definition is replaced Substitute it into the framework of attack and defense game theory.

[0053] Within the non-cooperative game framework obtained after substituting specific physical parameters, both rational attackers and defenders attempt to maximize their own gains. and Then, perform the Nash equilibrium solution. The obtained Nash equilibrium point The following partial derivative extremum conditions must be met: Due to the offensive and defensive payoff function and The inclusion of Weibull distribution integrals and Amdahl logarithmic terms leads to a highly nonlinear partial derivative system of equations, making analytical solutions difficult to obtain. Therefore, an iterative optimal response algorithm is employed for solving the problem. The solution process first decouples the two-layer coupled optimization problem into two univariate optimization subproblems: under a given defense period... Under the given conditions, find the attack strength that maximizes the attack benefit. , and , at a given attack strength Under these conditions, find the switching cycle that maximizes the benefits of defense. Specifically, define the attacker's optimal response function. satisfy Calculate the given switching period Attack strength that maximizes attack benefits ,in, This represents the attacker's payoff function. This represents the maximum resource investment coefficient by the attacker; it defines the optimal response function for the defender. satisfy Calculate the resource input coefficient for a given attacker. The switching cycle that maximizes the benefits of defense. ,in, The defender's payoff function, Indicates the physical migration time for switching paths. This indicates the maximum switching cycle value.

[0054] During the solution process, the attacker's resource input coefficient is first initialized. (For example, set as a baseline value of 1.0) and switch cycles before entering the iteration loop again (let the current iteration number be...). Attackers update their strategies based on observed defense tactics. Adjust your strategy to the optimal response point The defender is updated based on observed attack strategies. Adjust your strategy to the optimal response point .

[0055] After each iteration, calculate the policy offset for that iteration. and If the policy offset is less than the preset convergence threshold (i.e. and If the system reaches a steady state, then the strategy combination at this point is considered to be... This is the Nash equilibrium point (NE) under this path. If the policy offset still has not converged when the number of iterations reaches the set maximum value, then the current historical best solution is output or the heuristic default policy is triggered as the Nash equilibrium point under this path.

[0056] After solving for the Nash equilibrium point for all possible switching paths, the optimal periodic matrix and utility potential matrix are obtained, with dimensions of [missing information]. Among them, the optimal periodic matrix Record from configuration arrive The optimal dwell time for the defender should be set. Utility potential matrix. Record from configuration Switch to Under Nash equilibrium, the defender's net gain per unit time ( ), The higher the value, the higher the safety and cost-effectiveness of the corresponding switching path.

[0057] S104: Normalize the unit time equilibrium utility potential energy to map it as a reward signal of the reinforcement learning environment, and use the reinforcement learning algorithm to iteratively update the state-action value table with the virtual environment based on the reward signal.

[0058] In one possible embodiment, the normalized calculation of the equilibrium utility potential energy per unit time satisfies the following formula: ,in, This indicates a reward signal that reinforces the learning environment. This represents the instant reward function. Indicates time The system is running at the 1st Number configuration, Indicates at time The system selects to change the environment from configuration. Migrate to configuration , This represents the utility potential matrix, which represents the unit-time equilibrium utility potential energy of each switching path. Indicates switching paths The unit-time equilibrium utility potential energy, This represents the minimum value in the utility potential matrix. This represents the maximum value in the utility potential matrix. Represents a non-negative incentive bias; the state-action value table is iteratively updated based on the reward signal from the reinforcement learning environment. ,in, Indicates time State-action value Indicates the learning rate. Indicates the discount factor. Indicates time The maximum estimated value of all possible actions under the given state.

[0059] In one specific embodiment, a Markov decision process model is established at the macro-decision level. By defining a specific reward mapping function, the unit-time equilibrium utility potential energy obtained through the attack-defense game framework is transformed into a driving signal for reinforcement learning, thereby constructing a hierarchical coupling architecture that supports macro-planning through micro-evaluation.

[0060] For example, the dynamic defense process of a single-node heterogeneous system is modeled as a discrete-time step. Markov decision processes. The state space of a Markov decision process ( The state is defined as the configuration environment in which the defense system currently resides. The heterogeneous configuration set contains... Given several candidate configurations, the state space is represented as follows: ,state Indicates at time The system is running at the 1st Number configuration; Action space ( Define the target configuration that the defense system is preparing to switch to in the next moment as the action, and the action space. Isomorphic with state space, action This indicates that the decision-making system will choose to configure the environment. Migrate to configuration .

[0061] The process of converting the unit-time equilibrium utility potential energy obtained through the attack-defense game framework into a driving signal for reinforcement learning is specifically as follows: calling the utility potential matrix. Since the physical dimensions of game utility values ​​may vary considerably, to ensure the stability of reinforcement learning convergence, the Min-Max normalization method is used to map the utility potential values ​​to... Interval, construct an instant reward function : Among them, non-negative incentive bias terms This is used to ensure the positive nature of basic rewards and avoid exploration stagnation caused by negative returns.

[0062] The Q-Learning algorithm is used as the core of macro-level decision-making, and a state-action value table (Q-Table) is maintained, denoted as... Its size is Elements in the State-Action Value Table Characterization in the current configuration Select "Switch to Configuration" below. The long-term cumulative expected value of this action. At each decision step, the agent responds to the reward signals from the reinforcement learning environment. Iteratively update the Q value: Among them, the learning rate Control newly acquired experience (rewards) (And future projections) the speed at which old experience is covered. Discount factor. Weighing the immediate game gains against the long-term survival value. A higher [weight / weight] Value-driven defense systems focus on the connectivity and long-term security of configuration paths, rather than short-sightedly pursuing high returns from a single switch. This represents the maximum estimated value of all possible actions in the next state, reflecting the offline strategy characteristics of Q-Learning, which always estimates value in the direction of global optimum.

[0063] S105: Calculate the action selection probability based on the updated state-action value table and the Softmax probability mechanism. Determine the next target configuration for the defense system switching based on the action selection probability. Retrieve the corresponding optimal period based on the next target configuration to form the iteratively optimized defense strategy tuple.

[0064] In one possible embodiment, based on the state The probability of selecting the target configuration action is calculated according to the following formula: ,in, Indicates based on state Select through action The probability of selecting the action to switch target configuration. Indicates the state Next action State-action value Indicates the temperature coefficient. This represents the action space for switching the target configuration in the system. Indicates the state Next action The state-action value is calculated; the calculated action selection probabilities are compared, and the next target configuration for the defense system to switch is determined based on the action corresponding to the maximum action selection probability.

[0065] In one specific embodiment, based on the iteratively converged state-action value table, a probabilistic decision-making mechanism can be used to generate target configuration instructions that combine high returns with unpredictability, and combined with the optimal periodic matrix to generate an iteratively optimized spatiotemporal coordinated defense strategy tuple.

[0066] For example, the Q-values ​​in the iteratively converged state-action value table are transformed into action selection probabilities. Specifically, for the current state... Execute actions The probability of selecting the next target configuration is The calculation formula is as follows: ,in, Represents the temperature coefficient, when When the strategy tends to be completely random (uniformly distributed), it maximizes entropy; when At this point, the strategy tends to become completely greedy (selecting only the action with the highest Q value). In this embodiment, a moderate... The value of the state-action value table causes the defense system to favor high-value actions while retaining a non-zero probability of selecting suboptimal actions, thus maximizing long-term gains while deceiving attackers. After obtaining the action selection probability from the state-action value table, the next target configuration for the defense system's switching is determined based on the action corresponding to the maximum action selection probability, denoted as [the value of the state-action value table]. .

[0067] In one specific embodiment, the temperature coefficient ranges from [0.01, 1.0], and a suitable temperature coefficient range is set to [0.05, 0.1].

[0068] The optimal period for the corresponding switching path is retrieved from the optimal period matrix based on the next target configuration, denoted as . By merging spatial and temporal decisions, a complete, iteratively optimized spatiotemporal coordinated defense strategy tuple is output. Defense strategy tuple This has two physical implications. The first is Spatial Decision: the system will be reconfigured in the next migration. This is to gain a macro-level advantage in path planning. The second meaning is "temporal decision": how long the system will reside under the current configuration. This is to ensure that the cost-effectiveness ratio is optimal under the Nash equilibrium condition at the micro level.

[0069] In one possible embodiment, the iteratively optimized defense strategy principles can be parsed by the system's defense control center, which then issues instructions to the underlying execution units accordingly. The system maintains its current operating state until the timer expires. This triggers a cleaning and migration operation to switch to After the migration is complete, the system status will be updated to... Using this as a new starting point, the defense strategy generation process of repeatedly updating and iteratively optimizing the state-action value table is carried out to form a continuously evolving, dynamically closed-loop proactive defense cycle.

[0070] In a specific embodiment, to further explain the dynamic defense strategy iterative optimization method, a heterogeneous resource set is used. Heterogeneous configurations (denoted as) Taking a single-node defense system environment as an example, the iterative optimization process of the defense strategy is explained. Specifically:

[0071] During the system initialization phase, the physical attributes of the six heterogeneous configuration nodes are first globally calibrated to form boundary constraints for subsequent calculations. Using historical attack and defense exercise data, the standard intrusion time of the static non-heterogeneous attack surface for each configuration is determined. This parameter characterizes the baseline security of the configuration, and in this embodiment, it is calibrated as the following vector. (Unit: seconds): This data indicates the configuration. ( ) is a high-security node, and ( The node with low security exhibits significant security differences. The physical migration time between configurations was measured. ),form Agility Input Matrix The first in the matrix Line number Column elements represent configurations from the source. Switch to target configuration Time taken (in seconds): The matrix shows the distribution of migration costs across... Within the specified range. This embodiment defines a rigid lower limit for attack time based on network physical latency characteristics. Second.

[0072] For matrix For each switching path, an improved FlipIt game theory model is constructed. Key economic and engineering parameters and specific calculation logic defined in the model include: Amdahl resource mapping: setting acceleration factors. For any path (let its base time be...), Attackers invest resources Average attack time after The calculation is as follows: Weibull distribution characteristics: setting shape parameters To describe the cumulative effect of the attack success rate accelerating over time, the probability density function is: .

[0073] Based on the actual defense scenario, the following unit parameters are set: Unit defense value Unit attack benefit Attack cost rate Stability penalty coefficient The penalty index for nonlinear stability calculation Based on the above parameters, for any switching period and physical migration time (Taken from ), Defender's Profit Function Build as: ,in The integral represents the expected control duration after a successful attack.

[0074] Solve for the Nash equilibrium for all paths and output the micro-optimal defense cycle matrix. (Unit: seconds): Typical path example analysis: High-security path ( ):Target High safety (3523s), the model calculates the optimal period. s. High-frequency cleaning path ( ):Target It is easily compromised (1766s) and migrates quickly (20.8s), so the model compresses the cycle to... s. At this time Maintaining the leverage ratio within the range of 3.0 to 4.0 times avoids blind investment.

[0075] Extract the utility potential energy under the above equilibrium state, and perform Min-Max normalization mapping to... Range as environmental reward Set the learning rate. Discount factor The agent interacts with the environment for 2500 episodes, updating the Q-value table using the Bellman equation. After training convergence, the temperature coefficient is used... The Boltzmann (Softmax) policy generates a macroscopic policy probability matrix. : .

[0076] Assume the system is currently in configuration. Macro-level decision-making (where to go): View the matrix In line 2, the system selects to switch to the target configuration with a maximum probability of 0.78. This is because It has the highest level of security. and lower migration costs This is the globally optimal solution. Micro-decision (how long to wait): Consult the matrix. Row 2, column 4: The optimal dwell time is determined to be 244 seconds. Execution command: The defense system issues the command: "Maintain the current configuration for 244 seconds, then perform system cleaning and migration to the new configuration." ".

[0077] The above strategy satisfies Under the Nash equilibrium condition, the global maximization of the defense cost-effectiveness ratio is achieved through spatiotemporal collaborative optimization.

[0078] The dynamic defense strategy iterative optimization method provided by this invention designs a novel two-layer coupled architecture that integrates micro- and macro-scale iterative optimization of dynamic defense strategies. It introduces Amdahl's law and a rigid time limit to reconstruct the physical mapping of attack and defense, thus correcting the physical paradox of "infinite resources leading to instantaneous breach" in traditional models. Furthermore, it utilizes the Weibull distribution to accurately characterize the time accumulation effect of APT attacks during the kill chain construction phase, significantly improving the success rate and physical realism of defense strategies against higher-order threats. The innovatively constructed two-layer coupled architecture of micro-game evaluation and macro-intelligent planning directly maps the Nash equilibrium utility potential energy calculated at the bottom layer to the environmental reward signal of the upper-layer reinforcement learning. This invention breaks down the decision-making disconnect between "switching cycle selection" and "heterogeneous path planning" in traditional methods. By establishing a mapping mechanism based on Nash equilibrium utility potential energy, it cleverly avoids the challenges of high-dimensional computation. It solves the problem that existing technologies, due to the inherent heterogeneity between continuous parameter optimization in the time dimension and discrete path planning in the spatial dimension, often fall into the computational complexity curse of mixed-integer nonlinear programming (MINLP). Furthermore, it addresses the lack of a unified utility metric to bridge the gap between "instantaneous safety return" and "long-term path cumulative value," making it difficult to couple time parameter optimization with spatial path planning. This invention achieves comprehensive optimization of defense strategies across both time and space dimensions. In addition, by combining a nonlinear stability penalty term and a Boltzmann probabilistic decision-making mechanism, this invention, while ensuring system service continuity and availability, endows the defense system with the ability to continuously and adaptively evolve in heterogeneous environments and possesses high unpredictability. This effectively solves the core problem of balancing security benefits and engineering costs in the practical application of mobile target defense technology.

[0079] In other embodiments of this application, an electronic device is disclosed, such as... Figure 2As shown, the electronic device 300 may include: one or more processors 301; a memory 302; a display 303; one or more application programs (not shown); and one or more computer programs 304. These devices can be connected via one or more communication buses 305. The one or more computer programs 304 are stored in the memory and configured to be executed by the one or more processors 301. The one or more computer programs 304 include instructions that can be used to perform actions such as... Figure 1 And the steps in the corresponding embodiments.

[0080] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0081] In the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0082] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.

[0083] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.

Claims

1. A dynamic defense policy iterative optimization method, characterized in that, include: Define a heterogeneous resource set for candidate configurations of the defense system, and define core physical parameters for mobile target defense scenarios based on historical attack and defense data. These parameters include defining the standard intrusion time of a non-heterogeneous attack surface, the attacker's resource investment coefficient, the defender's physical migration time, and the rigid time limit for attack events. The standard intrusion time of a non-heterogeneous attack surface represents the average expected time required for an attacker to break through the current static configuration under standard attack intensity. The attacker's resource investment coefficient represents the multiple of the attacker's actual resource investment relative to the baseline resource investment, where the attacker's baseline resource investment is the resource investment under standard attack intensity. The defender's physical migration time represents the physical time required for the defense system to complete a full attack surface switching operation. The rigid time limit for attack events represents the minimum theoretical time required for the attacker's attack process. Based on the core physical parameters, the defender's strategy space and the attacker's strategy space are defined, the mapping relationship between attack resources and time is reconstructed, and the time accumulation effect of successful attacks is modeled. The attack and defense game sequence is defined, and the attack and defense game framework is constructed based on the attack and defense game sequence and the time accumulation effect of successful attacks. Based on the core physical parameters and the attack-defense game framework, the Nash equilibrium points of the attacking and defending sides are solved for each pair of switching paths in the heterogeneous resource set to obtain the optimal dwell period and the corresponding unit time equilibrium utility potential under each switching path. The unit-time equilibrium utility potential energy is normalized to map it into a reward signal of the reinforcement learning environment. Based on the reward signal, the state-action value table is updated iteratively by the reinforcement learning algorithm interacting with the virtual environment. The action selection probability is calculated based on the updated state-action value table and the Softmax probability mechanism. The next target configuration for the defense system switching is determined based on the action selection probability. The corresponding optimal period is retrieved based on the next target configuration to form the iteratively optimized defense strategy tuple.

2. The method of claim 1, wherein, Based on the core physical parameters, the defender's strategy space, the attacker's strategy space, the mapping relationship between attack resources and time are reconstructed, and the time accumulation effect of a successful attack is modeled, including: Defining the defender's behavior as performing system migration with a switching period, requiring the switching period to be greater than the defender's physical migration time, the defender's strategy space is , defining a nonlinear stability penalty term for constraining the defender's choice, where, represents the defender's strategy space, represents the switching period, represents the defender's physical migration time, represents the nonlinear stability penalty term, represents the stability cost coefficient, represents the penalty exponent; The action of the attacker is defined as the multiple of the resources chosen to accelerate the attack process, the attacker strategy space is wherein, represents the attacker strategy space, represents the attacker resource input coefficient; The mapping relationship between attack resources and time satisfies the following formula: ,in, A scale parameter indicating the time it takes for an attack to succeed. This indicates the standard intrusion time for non-heterogeneous attack surfaces. This indicates a rigid time limit for an attack event. This represents the logarithmic decay adjustment coefficient; The cumulative distribution of successful attacks satisfies the following formula: ,in, Indicates the attack time is The attacker's resource investment coefficient is The probability of a successful attack at that time. Indicates the attack time. Indicates shape parameters.

3. The method according to claim 1, characterized in that, The timing of attack and defense games includes: Define that within a single switching cycle, the defense migration window is entered first, followed by the effective defense window; By comparing the attack time with the effective defense window and the attack and defense game results within the corresponding switching cycle, the time when both sides gain benefits can be determined.

4. The method according to claim 1, characterized in that, An attack-defense game framework is constructed based on the timing of attack-defense games and the cumulative effect of successful attacks over time, including: The payoff function for constructing a defender satisfies the following formula: ,in, This indicates that the attacker's resource investment coefficient is... In the case of the defender switching cycles Internal benefits, This represents the value of defense operations per unit of time. Indicates the switching period The expected duration for which the internal system is in a secure service state. This represents the nonlinear stability penalty term; The attacker's payoff function is constructed to satisfy the following formula: ,in, This indicates that when the resource input coefficient is In the case of attackers switching cycles Internal benefits, This indicates that the attacker has complete control over the system's gains per unit of time. Indicates during the switching cycle The expected duration for which the internal system is under attacker control. Indicates the expected reward for the process. This indicates that the attacker's resource investment coefficient is [value missing]. The required linear resource cost rate.

5. The method according to claim 1, characterized in that, Based on the core physical parameters and the attack-defense game framework, the Nash equilibrium points of both the attacker and defender are solved for each pair of switching paths in the heterogeneous resource set to obtain the optimal cycle and the corresponding unit-time equilibrium utility potential energy under each switching path, including: Traverse all possible switching paths in the heterogeneous resource set, and incorporate the physical migration time of each switching path from the current configuration to the target configuration, and the standard intrusion time of the attacker breaking through the non-heterogeneous attack surface of the target configuration under the standard attack strength into the attack and defense game framework. Assuming that both the attacker and defender are trying to maximize their own gains, the optimal period and unit time equilibrium utility potential energy under each switching path are obtained by solving the Nash equilibrium. The optimal periodic matrix and utility potential matrix are generated based on the Nash equilibrium solution of all switching paths in the heterogeneous resource set.

6. The method according to claim 5, characterized in that, The optimal period and unit-time equilibrium utility potential energy under each switching path are obtained by solving the Nash equilibrium, including: Define the attacker's optimal response function satisfy Calculate the given switching period Attack strength that maximizes attack benefits ,in, This represents the attacker's payoff function. This represents the maximum resource investment coefficient by the attacker. Define the defender's optimal response function satisfy Calculate the resource input coefficient for a given attacker. The switching cycle that maximizes the benefits of defense. ,in, The defender's payoff function, Indicates the physical migration time for switching paths. This indicates the maximum value of the switching cycle; Initialize the attacker's resource input coefficient and switching cycle, and perform alternating iterative optimization of defense and attack strategies based on the attacker's best response function and the defender's best response function, and calculate the strategy offset in each iteration. When the policy offset is less than the preset convergence threshold or the number of iterations reaches the set maximum value, the Nash equilibrium point is output to obtain the optimal period and unit time equilibrium utility potential of the switching path.

7. The method according to claim 1, characterized in that, The unit-time equilibrium utility potential energy is normalized to map to a reward signal of the reinforcement learning environment. Based on the reward signal, the state-action value table is iteratively updated using a reinforcement learning algorithm in interaction with the virtual environment, including: The normalized calculation of the equilibrium utility potential energy per unit time satisfies the following formula: ,in, This indicates a reward signal that reinforces the learning environment. This represents the instant reward function. Indicates time The system is running at the 1st Number configuration, Indicates at time The system selects to change the environment from configuration. Migrate to configuration , This represents the utility potential matrix, which represents the unit-time equilibrium utility potential energy of each switching path. Indicates switching paths The unit-time equilibrium utility potential energy, This represents the minimum value in the utility potential matrix. This represents the maximum value in the utility potential matrix. This represents a non-negative incentive bias term; Iteratively update the state-action value table based on the reward signals from the reinforcement learning environment: ,in, Indicates time State-action value Indicates the learning rate. Indicates the discount factor. Indicates time All possible actions in the state The maximum estimated value, This represents the action space for switching the target configuration in the system.

8. The method according to claim 1, characterized in that, Based on the updated state-action value table and the Softmax probability mechanism, the action selection probability is calculated, and the next target configuration for the defense system switching is determined according to the action selection probability, including: According to the status The probability of selecting the target configuration action is calculated according to the following formula: ,in, Indicates based on state Select through action The probability of selecting the action to switch target configuration. Indicates the state Next action State-action value Indicates the temperature coefficient. This represents the action space for switching the target configuration in the system. Indicates the state Next action State-action value; By comparing the calculated action selection probabilities, the next target configuration for the defense system to switch is determined based on the action corresponding to the maximum action selection probability.