Dynamic defense switching period optimization method and device, medium and electronic equipment

By constructing a game theory framework and using particle swarm optimization to optimize the defense switching cycle, the problem of imbalance between defense effectiveness and system stability in dynamic attack and defense confrontation is solved, the optimal defense switching cycle is scientifically set, and the effectiveness of network defense and system availability are improved.

CN120979835AActive Publication Date: 2025-11-18GUANGZHOU UNIVERSITY

Patent Information

Application Number
CN202511486188.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-11-18
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing technologies make it difficult to scientifically and quantitatively set the optimal defense switching cycle in dynamic attack and defense scenarios, resulting in an imbalance between defense effectiveness and system stability, and making it impossible to effectively deal with advanced persistent threat (APT) attacks.

Method used

A game theory framework is constructed, defining the policy spaces of defenders and attackers. The defense switching cycle is optimized using the particle swarm optimization algorithm, taking into account the randomness of the attack success time and the availability of system resources. The payoff functions of both attackers and defenders are quantified, and the Nash equilibrium strategy is iteratively solved to determine the optimal switching cycle.

Benefits of technology

It achieves the optimal defense switching cycle selection that balances defense effectiveness and system stability in dynamic offensive and defensive confrontation, improves the scientific nature and effectiveness of defense strategies, and reduces system resource waste and service interruption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979835A_ABST
    Figure CN120979835A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic defense switching period optimization method and device, a medium and electronic equipment, and the method comprises the steps: initializing defense resources and attack surface reset time, and defining a strategy space constraint of a defense switching period; defining a game framework; constructing an attack and defense revenue function; initializing a particle swarm, calculating the fitness value of each particle, iteratively updating the speed and position of the particle, calculating the fitness value of the updated particle, and judging whether to update the individual optimal solution and the global optimal solution according to the updated particle fitness value, the original individual optimal solution and the global optimal solution; when a preset iteration termination condition is met, iteration is terminated, and a Nash equilibrium strategy of the attacker and the defender is obtained; and extracting a switching period from the Nash equilibrium strategy as an optimal dynamic defense switching period. By applying the method, the dynamic counterbalance relationship between attacker and defender strategies can be comprehensively analyzed, and the optimal dynamic defense switching period capable of giving consideration to the defense effect and the system stability can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network security, and in particular to a dynamic defense switching period optimization method. BACKGROUND

[0002] In the field of network security, advanced persistent threat (APT) is a challenging type of attack, which is highly covert, long in attack period, and the attacker usually has sufficient resources and clear goals, and specifically targets key targets for precise penetration, making it difficult for traditional static defense methods to effectively respond. To respond to APT threats, network defense shifts to the idea of active defense, and mobile target defense (MTD) emerges as a core technology. MTD breaks the stability of the attacker's reconnaissance and attack chain by dynamically adjusting the attack surface (such as IP address, port, etc.) of the system, increasing the difficulty of penetration, and becomes an important means to respond to APT. When using mobile target defense (MTD) technology to respond to advanced persistent threat (APT), the scientific setting of the defense switching period is the core problem to ensure the effectiveness of the defense.

[0003] As a key parameter of MTD strategy, the defense switching period directly determines the frequency of attack surface dynamic adjustment: the length of the period will directly affect the balance of two core dimensions - on the one hand, the period needs to be short enough to compress the attacker's effective attack window, avoiding the establishment of a stable attack link through continuous reconnaissance and penetration; on the other hand, the period cannot be too short, otherwise it will cause system resource overload, service interruption or stability decline due to frequent attack surface migration, increasing the system's own running cost and usability loss.

[0004] The influence of the defense switching period on the core dimensions of MTD technology directly presents the core contradiction of the setting of the defense switching period. At the same time, the characteristics of APT attacks further exacerbate the complexity of period setting: APT attacks have high concealment and persistence, and the time of attack success has significant randomness (i.e. the time from the initiation of the attack to the successful control of the target by the attacker is not a fixed value, but a random variable affected by attack intensity, target system defense state and other factors). This means that the setting of the defense period must fully consider the uncertainty of the attack process in order to accurately match the attack rhythm and effectively reset the attack process. Therefore, how to comprehensively balance the randomness of attack success time and the system usability loss of defense switching in a dynamic attack-defense confrontation scene, and establish a scientific and quantitative theoretical model to determine the optimal defense switching period, has become a key problem to be solved in the practical application of mobile target defense technology.

[0005] Therefore, it is necessary to provide a dynamic defense switching period optimization method that can balance defense effectiveness and system stability. SUMMARY

[0006] The application aims to provide a dynamic defense switching period optimization method for selecting an optimal dynamic defense switching period considering defense effect and system stability.

[0007] In a first aspect, the application provides a dynamic defense switching period optimization method, which comprises: initializing defense resources and attack surface reset time to define strategy space constraints of defense switching period; defining a game framework, including defining strategy space of defenders and attackers, attack effective time delay, attack start mechanism and attacker control resource time length; constructing an attack-defense benefit function according to effective time proportion of actual defense resource control in a unit period and net benefit of successful attack resource control of the attacker; initializing a particle swarm, calculating fitness value of each particle, setting each particle as an individual optimal solution, comparing and determining a global optimal solution, iteratively updating speed and position of the particle and calculating fitness value of the updated particle, judging whether to update the individual optimal solution and the global optimal solution according to the updated particle fitness value, the original individual optimal solution and the global optimal solution, terminating iteration when a preset iteration termination condition is met, and obtaining a Nash equilibrium strategy of the attack and defense sides; and extracting the switching period from the Nash equilibrium strategy as an optimal dynamic defense switching period.

[0008] The dynamic defense switching period optimization method provided by the application has the beneficial effects that: when performing dynamic defense switching period optimization design, the migration time (i.e. attack surface reset time) of defense switching is considered as a defense cost, the assumption of "switching completed instantaneously" in the traditional method is broken, and the cost calculation of the defense strategy is more in line with the actual system service interruption scenario. Moreover, the randomness of the time spent by the attacker in successful attack is considered , which is more in line with the multi-stage randomness characteristics of attack behavior in the network kill chain. The constructed game framework can comprehensively analyze the dynamic balance relationship of the strategies of the attack and defense sides when applied, and provide more accurate theoretical guidance for selection of the optimal defense period, and obtain an optimal dynamic defense switching period that can consider defense effect and system stability.

[0009] In a possible embodiment, defining the strategy space of the defenders and the attackers comprises: defining the action of the defenders as performing attack surface switching at a set switching period, requiring the switching period to be greater than the attack surface reset time, and the strategy space of the defenders being a set of legal switching periods; and defining the action of the attackers as selecting an attack rate to launch an attack, and the strategy space of the attackers being a set of effective attack rates.

[0010] In another possible embodiment, the calculation of the attacker control resource time length satisfies the following formula: , wherein, represents the switching period, represents the migration time, represents the time used by the attacker to successfully attack the system, A probability density function representing a time of success of the attack.

[0011] In other possible embodiments, constructing the attack-defense benefit function according to the effective time proportion of the defender actually controlling the resource in a unit period and the net benefit of the attacker successfully controlling the resource includes: defining a defender benefit function as a proportion of the effective time of the defender controlling the resource in a period to the switching period, and the calculation of the defender benefit function satisfies the following formula: wherein, represents the switching period, represents the duration of the attacker controlling the resource, represents the attack surface reset time; defining an attacker benefit function as a net benefit of the attacker successfully controlling the resource in the effective defense period, and the calculation of the attacker benefit function satisfies the following formula: = - wherein, represents a unit benefit of the controlled resource, represents a unit rate cost of the attack, represents an attack rate.

[0012] The calculation of the fitness value of each particle includes: calculating the fitness value of the particle according to the attack-defense benefit function and the attack-defense strategy solution of the particle; and the fitness value calculation satisfies the following formula: F( )= wherein, represents the i th particle, represents the defender benefit function, represents the attack-defense strategy solution of the i th particle, represents the attacker benefit function. According to the updated particle fitness value, the original individual optimal solution and the global optimal solution, it is judged whether to update the individual optimal solution and the global optimal solution, including: comparing the updated particle fitness value with the fitness value of the original individual optimal solution, when the updated particle fitness value is greater, updating the individual optimal solution; and comparing the updated particle fitness value with the fitness value of the original global optimal solution, when the updated particle fitness value is greater, updating the global optimal solution.

[0013] When the preset iteration termination condition is met, the iteration is terminated, including: when the number of iterations reaches a preset maximum number of iterations or the fitness value of the global optimal solution changes less than a preset threshold, the iteration is terminated.

[0014] When the preset iteration termination condition is met, the iteration is terminated, including: when the number of iterations reaches a preset maximum number of iterations or the fitness value of the global optimal solution changes less than a preset threshold, the iteration is terminated.

[0015] ​In a second aspect, the present application also provides a dynamic defense switching period optimization device, comprising: an initialization unit configured to initialize defense resources and attack surface reset time to define policy space constraints of the defense switching period; a game framework construction unit configured to define a game framework, including defining policy spaces of the defender and the attacker, attack effective time delay, attack start mechanism and attacker control resource duration; a payoff function construction unit configured to construct an attack-defense payoff function according to the effective time proportion of the actual control resources of the defender in a unit period and the net income of the successful control resources of the attacker; a solution unit configured to initialize a particle swarm, calculate the fitness value of each particle, set each particle as an individual optimal solution, compare and determine a global optimal solution, iteratively update the speed and position of the particle and calculate the fitness value of the updated particle, judge whether to update the individual optimal solution and the global optimal solution according to the updated particle fitness value, the original individual optimal solution and the global optimal solution, terminate the iteration when the preset iteration termination condition is met, and obtain the Nash equilibrium strategy of the attack-defense sides; and a switching period extraction unit configured to extract the switching period from the Nash equilibrium strategy as an optimal dynamic defense switching period.

[0016] In a third aspect, the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the dynamic defense switching period optimization method.

[0017] In a fourth aspect, the present application also provides an electronic device, comprising: a processor and a memory; the memory is configured to store a computer program; and the processor is configured to execute the computer program stored in the memory, so that the electronic device executes the dynamic defense switching period optimization method.

[0018] The beneficial effects of the second aspect to the fourth aspect can be referred to the description of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 A flowchart of a dynamic defense switching period optimization method provided by an embodiment of the present application; Figure 2 A dynamic defense model provided by an embodiment of the present application; Figure 3 A schematic diagram of a dynamic defense switching period optimization device provided by an embodiment of the present application; Figure 4 A schematic diagram of an electronic device structure provided by an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the objects, technical solutions and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application. Unless otherwise defined, the technical terms or scientific terms used herein should be understood as the common meanings of the technical terms or scientific terms for those having ordinary skills in the art to which the present application belongs. The similar words such as “comprise” used herein mean that the elements or objects before the words cover the elements or objects listed after the words and their equivalents, and do not exclude other elements or objects.

[0021] The embodiment provides a dynamic defense switching period optimization method. Referring to the drawings accompanying the description Figure 1 The method comprises the following steps. S101: initializing defense resources and attack surface reset time, and defining policy space constraints of the defense switching period.

[0022] In a possible embodiment, system resources are reset through a mobile target defense (MTD) technology, the attack surface conversion time (i.e. service interruption time during defense switching) is determined, and the policy space constraints of the defense switching period are defined, that is, the period needs to be a positive real number greater than the conversion time (only the feasible range is set, and no specific period value is preset).

[0023] Exemplarily, for the host resources that can be occupied by both the attacker and the defender (i.e. the host resources that can be controlled by both the attacker and the defender through attack or defense behaviors), the defender performs an initial resource reset operation before the start of the game process, that is, selects the IP hopping, system version switching and other mobile target defense technologies to comprehensively reset the host resources, clears the old system information and updates to a new state, and ensures that the core features (such as network identification, system configuration, etc.) of the host after the reset are refreshed, and finally the game process starts with the defender stably occupying the control right of the resource.

[0024] Since the process of switching and resetting the attack surface of the host resources by the defender through the mobile target defense technology is not instantaneous, a certain time is needed, and during this period, the defender cannot obtain any benefit because the host resources cannot provide services to the outside, and the attacker cannot attack and probe the host resources. The reset time is determined by the specific mobile target defense technology adopted, and is usually set to a fixed value according to actual experience Taking "dynamic port switching" as an example: To prevent attacks, the server periodically switches its external service ports from old ports to new ports (e.g., from port 8080 to port 9090). During the switch, the old port listening must first be closed, and the new port service must be configured and started. During the period from when the old port is closed to when the new port provides a stable response, the server cannot receive client requests, and attackers cannot perform scans or attacks through the old or new ports. If the process of switching and resetting the host resource attack surface takes 15 seconds, then the set attack surface reset time... .

[0025] After completing the initial resource reset, the policy space constraints for the defense switching cycle need to be defined. The defense switching cycle refers to the time interval between two adjacent host resource attack surface reset operations. Its policy space must satisfy two constraints: the cycle must be a positive real number greater than 0, and it must be strictly greater than the switching time of a single reset. This is to avoid the period being less than or equal to If the previous reset is not completed before the next operation is triggered, the host resources will remain unavailable. At the same time, during the reset, the host is in a state where it cannot provide services or has a temporarily interrupted network connection because it needs to shut down old services and configure new features (such as changing IP, port, etc.). During this time, the defender cannot obtain service benefits through the host.

[0026] S102: Define the game theory framework, including defining the strategy space of the defender and the attacker, the attack activation delay, the attack initiation mechanism, and the duration of the attacker's control over resources.

[0027] In one possible embodiment, defining the policy spaces of the defender and the attacker includes: defining the defender's action as performing attack surface switching at a set switching period, requiring the switching period to be greater than the attack surface reset time, and the defender's policy space is a set of legal switching periods; defining the attacker's action as selecting an attack rate to launch an attack, and the attacker's policy space is a set of valid attack rates.

[0028] In one possible embodiment, the calculation of the duration of the attacker's control over the resources satisfies the following formula: ,in, Indicates the switching cycle. Indicates the migration time. This indicates the time it took for an attacker to successfully compromise the system. The probability density function represents the time it takes for an attack to succeed.

[0029] For example, see the appendix to the specification. Figure 2 The defenders use a fixed period Take action to switch the system's attack surface; after each switch, resource ownership reverts to the defender. The defender's actions are defined as periodic... The attack surface should be switched periodically, requiring... The defender strategy space is the set of all legitimate switching periods, mathematically denoted as = {t | t > 0} ∈ , }, where is a positive real number, and the constraint ensures that the switching period is greater than the transition time to guarantee the feasibility of the strategy.

[0030] The attack effective delay is defined as: when the attack surface transition is completed, the system service is restored, at which time the attacker can launch an attack on the host resources. After the attacker launches an attack, the control of the resources is not instantaneous, but there is a delay. Specifically, if the attack is launched at time t, the time at which the attacker successfully establishes control and obtains ownership of the resources is t , where is a random time variable extracted from the probability distribution .

[0031] In one possible embodiment, the probability distribution is set to comply with the Weibull probability distribution. The Weibull probability distribution conforms to the non-negativity of the attack time, and when the shape parameter of the probability distribution is greater than 1, it can depict the increasing success rate of the attack over time (“the later the easier it is to succeed”), which conforms to the characteristics of the attack.

[0032] The attack initiation mechanism is defined as: since the resource reset process of the defender will cause the host to be unable to provide services externally for time (manifesting as network connection interruption, service non-response, or invalidation of the old access path), and the attacker can perceive the resource state change through continuous network scanning, service detection, etc.: when the target resource is detected to switch from a non-responsive state to an accessible state (such as a new IP / port responding to a request, service characteristics reappearing), the attacker judges that the resource reset operation of the defender has been completed, and the attacker initiates the attack action after the resource switching is completed.

[0033] The action of the attacker is defined as selecting to launch an attack at an attack rate r, where the attack rate r is a core parameter of the probability distribution form of the attack effective delay , and its value directly determines ​​​​The specific characteristics of the probability distribution. When the attack rate *r* is low, it indicates a weaker attack intensity. At this time, the probability density of the probability distribution within the small value range (i.e., effective in the short term) is significantly reduced, meaning the attacker is less likely to successfully establish resource control in a short period. A lower attack rate results in the following shape of the probability distribution: the peak of the probability density curve shifts towards a greater delay and the peak becomes lower; the probability of short-term effectiveness decreases significantly, while the probability of long-term effectiveness increases significantly, resulting in a longer distribution "tail." Conversely, an increased attack rate results in the following shape: the peak of the probability density curve shifts towards a smaller delay and the peak becomes higher; the probability of short-term effectiveness increases significantly, while the probability of long-term effectiveness decreases significantly, resulting in a shorter distribution "tail." Adjusting *r* directly controls the probability density distribution characteristics for different delay values. The attacker's policy space is defined as the set of all effective attack rates, i.e. ={r|r∈ This set encompasses all strategy choices that can influence the probability distribution of attack effectiveness delay by adjusting the attack rate.

[0034] In one possible implementation, the attack rate r is defined to satisfy r a / rc, where r a `r` represents the number of requests sent per second, while `rc` represents the target system's defense criticality rate, i.e., the attack rate threshold that triggers the alarm system (assuming this parameter is known to both parties). The value of `r` ranges from [0, 1]. When `r ∈ [0, 0.3]`, it is defined as a low attack rate, where the attack intensity is far below the target system's defense criticality, and the attacker is less likely to successfully establish resource control in a short period of time. When `r ∈ [0.8, 1]`, it is defined as a high attack rate, where the attack intensity is close to or reaches the target system's defense criticality, and the attacker is more likely to successfully establish resource control in a short period of time.

[0035] Attacker controls resource duration This refers to the expected duration for which an attacker successfully controls resources within each game (i.e., one defense switching cycle). When the defender uses a cycle... ( > When switching defenses, or when an attacker launches an attack at an attack rate r, Specifically, it refers to the effective defense time during this switching cycle ( - Within a given period, this represents the average duration for which an attacker successfully controls a resource. The duration of resource control is calculated using the following formula: ,in, Indicates the switching cycle. Indicates the migration time. This indicates the time it took for the attacker to successfully compromise the system. a probability honeycomb function representing the attack success time, for the corresponding cumulative distribution function, r represents the attack rate, represents the attack surface security metric. The calculation of the attacker control resource duration quantifies the average control duration of the attacker within the defense period, taking into account the attack rate and the influence of the attack surface security metric .

[0036] S103: Constructing an attack-defense benefit function according to the effective time proportion of the defender actually controlling the resource within a unit period and the net benefit of the attacker successfully controlling the resource.

[0037] In one possible embodiment, the defender benefit function is defined as the proportion of the effective time of the defender controlling the resource within a period to the switching period, and the calculation of the defender benefit function satisfies the following formula: , wherein, represents the switching period, represents the attacker control resource duration, represents the attack surface reset time. The attacker benefit function is defined as the net benefit of the attacker successfully controlling the resource within the effective defense period, and the calculation of the attacker benefit function satisfies the following formula: = , wherein, represents the unit benefit of controlling the resource, represents the unit rate cost of attack, represents the attack rate.

[0038] The defender benefit function characterizes the net benefit of the defender controlling the resource in the dynamic defense process, and is defined as the proportion of the effective time of the defender actually controlling the resource to the switching period. The game calculates the average benefit within a switching period, assuming that the defender executes the system attack surface switching (migration time is included in the cost of the defender), and the duration of the attacker successfully controlling the resource within a period is , then the time of the defender controlling the resource is , and the expression of the defender benefit function is: .

[0039] The attacker benefit function characterizes the net benefit of the attacker successfully controlling the resource in the attack-defense game, and is defined as the benefit of the attack successfully controlling the resource minus the attack cost. If the attacker chooses to launch an attack at an attack rate r, the unit rate cost is , the unit benefit of controlling the resource is , and the attack success duration is , then in the effective defense period The expression of the attacker's benefit function is: .

[0040] S104: initialize the particle swarm, calculate the fitness value of each particle, set each particle as the individual optimal solution, compare and determine the global optimal solution, iteratively update the speed and position of the particle and calculate the fitness value of the updated particle, judge whether to update the individual optimal solution and the global optimal solution according to the updated particle fitness value, the original individual optimal solution and the global optimal solution, terminate the iteration when the preset iteration termination condition is met, and obtain the Nash equilibrium strategy of the attack and defense sides.

[0041] In a possible embodiment, the fitness function of the particle is defined as the sum of the benefits of the attack and defense sides, and specifically satisfies the following formula: F( )= , wherein, represents the i-th particle, represents the defense benefit function, represents the attack and defense strategy solution of the i-th particle, represents the attack benefit function.

[0042] According to the updated particle fitness value, the original individual optimal solution and the global optimal solution, whether to update the individual optimal solution and the global optimal solution, including: comparing the updated particle fitness value with the fitness value of the original individual optimal solution, updating the individual optimal solution when the updated particle fitness value is greater, comparing the updated particle fitness value with the fitness value of the original global optimal solution, and updating the global optimal solution when the updated particle fitness value is greater.

[0043] The preset iteration termination condition includes the maximum iteration number and the fitness value change threshold, that is, when the iteration number reaches the maximum iteration number or the fitness value change of the global optimal solution is less than the set threshold, the iteration is terminated.

[0044] Exemplarily, the application particle swarm algorithm iteratively optimizes the strategy solution, including: setting the maximum iteration number and the fitness value change threshold.

[0045] Initialize the particle swarm: set the particle swarm size to N, each particle represents a set of attack and defense strategy solutions, and the attack and defense strategy solution includes the switching period of the defense and the attack rate r of the attacker. The i-th particle is denoted as ( , ), ​​​​​​=1,2,⋯,N, and ∈ , ∈ Randomly initialize the particle's position in the solution space, and simultaneously initialize the particle's velocity. =( , ),in, Let represent the initial velocity of the i-th particle in the strategy dimension of "defender switching cycle". Let represent the initial velocity of the i-th particle in the policy dimension of "attacker's attack rate". Define the global optimal solution. =( , and the individual optimal solution for each particle. =( , In this process, the initial position of each particle is set as its individual optimal solution. The fitness values ​​of all particles are compared, and the position of the particle with the highest fitness value is set as the global optimal solution. The fitness value of a particle is calculated using the following formula: F( )= ,in, Indicates the first One particle, Represents the defender's profit function. Indicates the first Solution to the attack and defense strategies of individual particles. This represents the attacker's profit function.

[0046] Particle updates and optimal solution iterations include updating the particle velocity and position in each iteration according to the following formula: ,in, This represents the updated velocity of the i-th particle in the (k+1)-th iteration, corresponding to the "defense switching cycle" strategy dimension. Indicates inertia weight, Indicates the number of iterations. and Represents the learning factor. and express Random numbers within a range Indicates the first In each iteration, the switching cycle of the individual optimal solution corresponding to the "individual optimal solution" of the i-th particle. Indicates the first In the next iteration, the switching cycle corresponding to the "global optimal solution" of the i-th particle. Indicates the first In each iteration, the switching period corresponding to the "global optimal solution" for all particles. denotes the updated speed of the i-th particle corresponding to the strategy dimension of "attacker attack rate" in the k+1-th iteration, denotes the attacker attack rate corresponding to the "individual optimal solution" of the i-th particle in the k-th iteration, denotes the current attacker attack rate (the strategy value before updating) of the i-th particle in the k-th iteration, denotes the attacker attack rate corresponding to the "global optimal solution" of all particles in the k-th iteration. After updating, if the particle position exceeds the range of the strategy space (i.e. or ), boundary correction is performed. The fitness value of the updated particle is calculated , if , the individual optimal solution is updated . At the same time, if there is a particle whose fitness value is better than the current global optimal solution, the global optimal solution is updated .

[0047] When the number of iterations reaches the maximum number of iterations, or the fitness value of the global optimal solution changes by less than a set threshold in continuous iterations, the iteration is terminated.

[0048] When the iteration is terminated, the global optimal solution is obtained as the Nash equilibrium strategy of the attacker and defender, i.e., the optimal strategy combination is obtained. The optimal strategy combination obtained by the method satisfies the following conditions: ( , )≥ ( , ), , ( , )≥ ( , ), r .

[0049] S105: Extract the switching period from the Nash equilibrium strategy as the optimal dynamic defense switching period.

[0050] In a possible embodiment, the switching period extracted from the Nash equilibrium strategy is applied as the obtained optimized dynamic defense switching period, which can balance the defense effect of the mobile target defense technology and the system stability.

[0051] In a specific embodiment, the specific process of optimal dynamic defense switching period selection by applying the dynamic defense switching period optimization method is as follows: in the initialization stage, the defender selects IP hopping technology as the initial strategy, and the security metric of this technology is 98, for example, by switching the server IP address from 192.168.1.100 to 192.168.1.150, the resource state initialization is completed. The migration time t m is set to 3 seconds, that is, the attack surface reset time. The defender's strategy space is defined as all switching period sets that satisfy > 15 seconds.

[0052] After the attack surface conversion is completed (that is, after t m = 15 seconds), the attack can be launched, and it is assumed that the attack success time t a obeys the Weibull distribution, then its probability density function and the cumulative distribution function P are: = , P = , where represents the scale parameter.

[0053] According to the actual situation, the attacker's unit rate cost , the unit benefit of the controlled resource 1.5, the shape parameter k=2 in the probability distribution (this parameter describes the trend of the attack success probability changing with time, and when k=2, the attack success rate increases with time, which conforms to the actual law of "accumulating attack resources and gradually improving the success probability" in most network attacks), and the scale parameter = .

[0054] The average time of the attacker controlling the resource in one game period is calculated as: ). .

[0055] The payoff functions of the defender and the attacker are defined as and : = t c - t m -[ ( t c - t m ) - ∫ 0 t c - t m e -( r · t a T AS ¯ ) k d t a ] t c , .

[0056] In the particle swarm optimization algorithm iteration stage of the attack and defense game, the system sets the maximum iteration number Max_iter=1000, the fitness change threshold ε=0.01, the particle swarm size N=30, and each particle represents a set of attack and defense strategy solutions ( r). Initially, particle positions are randomly generated, where defender switching period ∈(15,100) seconds (ensure t m =15 seconds), attacker attack rate r∈[0,1], particle velocity is initialized as a random value in the range of [0,1]. The algorithm uses the standard particle swarm algorithm update formula, inertia weight w=0.7, learning factor c1=c2=1.5, in each iteration, the particle adjusts the speed and position according to the current global optimal solution and individual historical optimal solution . The fitness function is defined as the sum of the benefits of both attack and defense, that is, F( )= .

[0057] The optimal strategy obtained by joint solution is , and the Nash equilibrium strategy combination of attack and defense satisfies the following conditions: ( , )≥ ( , ), , ( , )≥ ( , ), r .

[0058] The dynamic defense switching period optimization method provided by the application abstracts the attack and defense dynamic confrontation into a game process, and clearly defines the strategy space and interaction rules of both parties. The defender selects the defense switching period (dynamically adjusts the frequency of the attack surface) as the core strategy, and the period needs to be greater than the attack surface conversion time to ensure feasibility; the attacker selects the attack rate as the core strategy, and the attack rate directly affects the time characteristics of its successful control of resources. Through this framework, the attack and defense behaviors are converted into a quantifiable strategy interaction model. The definition of the game framework fully considers the uncertainty of the attack process, defines the time from the initiation of the attack to the successful control of the resources by the attacker as a random variable, describes the randomness with a probability distribution, and reflects the successful time law under different attack intensities. At the same time, the service interruption time (attack surface conversion time) when the defense switches is taken into account in the defense cost, ensuring that the model is consistent with the actual scenario of the impact on system availability. The problems of experience-dependent strategy rigidity, key assumptions of theoretical models deviating from reality, and imbalance between system availability and defense effect in the design scheme of the dynamic defense switching period in the prior art are solved, which leads to the problems of lack of scientificity and effectiveness in the selection of the defense period.

[0059] The method is used for dynamic defense switching period optimization design, the migration time (i.e., attack surface reset time) of the defense switching is considered as the defense cost, the assumption of "switching instantaneous completion" in the traditional method is broken, the cost calculation of the defense strategy is more in line with the actual system service interruption scene. Moreover, considering the randomness of the time consumed by the attacker to attack successfully , the uncertainty of the attack process is described through the probability distribution , compared with the traditional fixed attack time setting, the multi-stage randomness characteristics of the attack behavior in the network kill chain are more in line with the network kill chain. In addition, the attack rate r of the attacker is taken as a strategy variable, the double influence of the attack strength on the success probability and the attack cost is quantified, so that the constructed game framework can comprehensively analyze the dynamic balance relationship between the strategies of the attack and defense sides when applied, and provide more accurate theoretical guidance for the selection of the optimal defense period, and the optimal dynamic defense switching period which can balance the defense effect and system stability is obtained.

[0060] Referring to the accompanying drawings Figure 3 in the description, the embodiment also provides a dynamic defense switching period optimization device, which is used for implementing the method embodiment. The device comprises: An initialization unit 201 is configured to initialize the defense resources and the attack surface reset time, and define the strategy space constraints of the defense switching period.

[0061] A game framework construction unit 202 is configured to define a game framework, including defining the strategy space of the defender and the attacker, the attack effective delay, the attack start mechanism and the attacker control resource duration.

[0062] A revenue function construction unit 203 is configured to construct an attack and defense revenue function according to the effective time proportion of the actual control resources of the defender in a unit period and the net revenue of the successful control resources of the attacker.

[0063] A solving unit 204 is configured to initialize a particle swarm, calculate the fitness value of each particle, set each particle as an individual optimal solution, compare and determine a global optimal solution, iteratively update the speed and position of the particle and calculate the fitness value of the updated particle, judge whether to update the individual optimal solution and the global optimal solution according to the updated particle fitness value, the original individual optimal solution and the global optimal solution, terminate the iteration when the preset iteration termination condition is met, and obtain the Nash equilibrium strategy of the attack and defense sides.

[0064] A switching period extraction unit 205 is configured to extract the switching period from the Nash equilibrium strategy as the optimal dynamic defense switching period.

[0065] All related contents of each step involved in the above method embodiment can be referred to the function description of the corresponding function module, and will not be repeated here.

[0066] In some embodiments of the present application, the embodiments of the present application disclose an electronic device, such as Figure 4 As shown in the figure, the electronic device 300 can include one or more processors 301, a memory 302, a display 303, one or more application programs (not shown), and one or more computer programs 304, which can be connected through one or more communication buses 305. The one or more computer programs 304 are stored in the memory and configured to be executed by the one or more processors 301, and the one or more computer programs 304 include instructions that can be used to perform various steps in the embodiments of the present application and corresponding embodiments. Figure 1

[0067] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0068] The functional units in each of the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0069] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application essentially or the parts that make contributions to the prior art or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a flash memory, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, and various media that can store program codes.

[0070] The above is only a specific implementation of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the embodiments of the present application should be covered in the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.​

Claims

1. A method for optimizing dynamic defense switching cycles, characterized in that, The method comprises the following steps: initializing defense resource and attack surface reset time, defining strategy space constraint of defense switching period; defining game framework, including defining strategy space of defender and attacker, attack effective time delay, attack start mechanism and attacker control resource time length; constructing attack-defense benefit function according to effective time proportion of actual control resource of defender in unit period and net benefit of successful control resource of attacker; initializing particle swarm, calculating fitness value of each particle, setting each particle as individual optimal solution, comparing to determine global optimal solution, iteratively updating speed and position of particle and calculating fitness value of updated particle, judging whether to update individual optimal solution and global optimal solution according to fitness value of updated particle, original individual optimal solution and global optimal solution, terminating iteration when preset iteration termination condition is met, and obtaining Nash equilibrium strategy of attack and defense; extracting switching period from the Nash equilibrium strategy as optimal dynamic defense switching period.

2. The method of claim 1, wherein, Defining strategy space of defender and attacker comprises: defining action of defender as performing attack surface switching with set switching period, requiring switching period to be greater than attack surface reset time, and strategy space of defender being set as collection of legal switching periods; defining action of attacker as selecting attack rate to launch attack, and strategy space of attacker being set as collection of effective attack rates.

3. The method of claim 1, wherein, The calculation of the attacker control resource duration satisfies the following formula: wherein denotes the switching period, denotes the migration time, denotes the time taken by the attacker to successfully attack the system, denotes the probability density function of the attack success time.

4. The method of claim 1, wherein, Constructing attack-defense benefit function according to effective time proportion of actual control resource of defender in unit period and net benefit of successful control resource of attacker comprises: The defender benefit function is defined as the ratio of the effective time of the defender controlling the resource in a period to the switching period, and the calculation of the defender benefit function satisfies the following formula: wherein, represents the switching period, represents the duration of the attacker controlling the resource, represents the attack surface reset time; The attacker benefit function is defined as the net benefit of the attacker successfully controlling the resource in the effective defense period, and the calculation of the attacker benefit function satisfies the following formula: = - , wherein, represents the unit benefit of the controlled resource, represents the unit rate cost of the attack, represents the attack rate.

5. The method of claim 1, wherein, calculating fitness value of each particle comprises: calculating fitness value of particle according to attack-defense benefit function and attack-defense strategy solution of particle; Fitness values ​​are calculated according to the following formula: F( )= ,in, Indicates the first One particle, Represents the defender's profit function. Indicates the first Solution to the attack and defense strategies of individual particles. This represents the attacker's profit function.

6. The method of claim 1, wherein, judging whether to update individual optimal solution and global optimal solution according to fitness value of updated particle, original individual optimal solution and global optimal solution, comprising: comparing fitness value of updated particle with fitness value of original individual optimal solution, updating individual optimal solution when fitness value of updated particle is greater, and comparing fitness value of updated particle with fitness value of original global optimal solution, updating global optimal solution when fitness value of updated particle is greater.

7. The method of claim 1, wherein, terminating iteration when preset iteration termination condition is met comprises: terminating iteration when iteration number reaches preset maximum iteration number or fitness value of global optimal solution changes by less than preset threshold.

8. A dynamic defense handoff period optimization apparatus, comprising: The device comprises: an initialization unit configured to initialize defense resource and attack surface reset time, and define strategy space constraint of defense switching period; a game framework construction unit configured to define game framework, including defining strategy space of defender and attacker, attack effective time delay, attack start mechanism and attacker control resource time length; a benefit function construction unit configured to construct attack-defense benefit function according to effective time proportion of actual control resource of defender in unit period and net benefit of successful control resource of attacker; A solving unit is configured to initialize a particle swarm, calculate a fitness value of each particle, set each particle as an individual optimal solution, compare and determine a global optimal solution, iteratively update a velocity and a position of each particle and calculate a fitness value of the updated particle, judge whether to update the individual optimal solution and the global optimal solution according to the fitness value of the updated particle, the original individual optimal solution and the global optimal solution, terminate the iteration when a preset iteration termination condition is met, and obtain a Nash equilibrium strategy of the attack and defense sides. A switching period extraction unit is configured to extract a switching period from the Nash equilibrium strategy as an optimal dynamic defense switching period.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the dynamic defense switching period optimization method in any one of claims 1 to 7.

10. An electronic device, comprising: The computer program is executed by the processor to implement the dynamic defense switching period optimization method in any one of claims 1 to 7. The computer program is executed by the processor to implement the dynamic defense switching period optimization method in any one of claims 1 to 7. The computer program is executed by the processor to implement the dynamic defense switching period optimization method in any one of claims 1 to 7. ​

Citation Information

Patent Citations

  • Mobile target defense decision selection method, device and system based on Markov time game

    CN110300106A

  • Network spoofing defense decision-making method and system based on Flipit intelligent game

    CN116962050A

  • Host scanning and defense strategy collaborative optimization method based on random game

    CN120389879A

  • Game-based optimal resource allocation method for cyber-physical power system (cpps) to defend against false data injection (fdi) attack

    GB2621417A

Cited By

  • Dynamic defense strategy iterative optimization method

    CN121841864A