A cross-domain attack defense method based on game strategy
By constructing a joint state space and game strategy and designing a cross-layer coupled state transfer model, the problem of cross-domain attack defense in the cyber-physical fusion environment is solved, the optimal defense under incomplete information conditions is achieved, and the defense effectiveness and robustness of the system are improved.
Patent Information
- Application Number
- CN202510896599.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-07-01
AI Technical Summary
Existing technologies are difficult to effectively defend against cross-domain attacks, especially in the cyber-physical fusion environment. The detection and response mechanisms of cross-domain attacks are difficult to capture attack residuals, and there is a lack of quantifiable joint state assessment standards and response strategies, which makes active defense difficult.
Construct a joint state space, design game strategies, discretize the residuals of the information domain and the physical domain, design the action spaces of attackers and defenders, establish a cross-layer coupled state transfer model, solve the Nash equilibrium to output the defense strategy, and realize the dynamic coupling interaction between the information domain and the physical domain.
Under conditions of incomplete information, through continuous interactive updates, we gradually approach the optimal defense strategy, thereby improving the cross-layer defense effectiveness and robustness of complex physical information systems, and being able to clearly quantify the security status in cascade reactions and guide optimal defense decisions.
Smart Images

Figure CN120455164B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a cross-domain attack defense method based on game strategy. Background Art
[0002] Industrial control systems (ICS) face increasingly complex security threats in a cyber-physical convergence environment. Attackers can covertly intrude by manipulating the command flow in the cyber domain or influencing the state in the physical domain. Cross-domain attacks, in particular, manifest as cross-domain interference. This occurs when attackers precisely tamper with commands or forge physical feedback, causing system states to err. Executing these erroneous commands can further exacerbate system state deviations.
[0003] To counter these attacks, researchers have attempted to establish detection mechanisms in both the cyber and physical domains. However, due to significant cross-layer coupling, response mechanisms struggle to capture the full picture of attack residuals. Furthermore, the lack of quantifiable joint state assessment criteria and response strategies has limited the development of intelligent active defense mechanisms.
[0004] Therefore, it is necessary to provide an active defense method for cross-domain attacks to solve the problem of difficulty in active defense for cross-domain attacks in the prior art. Summary of the Invention
[0005] The purpose of the present invention is to provide a cross-domain attack defense method based on game strategy to achieve active defense against cross-domain attacks.
[0006] In the first aspect, the cross-domain attack defense method based on game strategy provided by the present invention includes: constructing a joint state space, the joint state space includes information domain residuals and physical domain residuals, and discretizing the joint norm of the joint state space; designing the action space of the attacker and the defender, and based on the joint state space design, factorizing the action space of the attacker and the defender to design a strategy; designing a benefit function for calculating the benefit obtained by the defender for defense; establishing a cross-layer coupled state transition model based on the change trend of the joint residual of the joint state space to simulate the dynamic coupling interaction between the information domain and the physical domain; formalizing the state transition model into a zero-sum game, solving the Nash equilibrium of the attacker's benefit and the defender's benefit, and outputting the defender's strategy at the Nash equilibrium as the defense strategy.
[0007] The beneficial effects of the cross-domain attack defense method based on game strategy provided by the present invention are: while considering cascade reactions, unified state modeling clearly quantifies the security status of the information-physical system, the design of defense strategy and action space, and reasonably designs the benefit function quantification to guide the optimal defense decision, so that the defense strategy can gradually approach the optimal confrontation through continuous interactive updates under conditions of incomplete information, and ultimately improve the effectiveness and robustness of cross-layer defense of complex physical information systems (CPS).
[0008] In a possible embodiment, a joint state space is constructed, where the joint state space includes information domain residuals and physical domain residuals, including: ;in, The calculation satisfies the following formula: , The calculation satisfies the following formula: , represents the information domain residual, represents the physical domain residual, , represents the residual value of the information layer, Represents the residual value of the physical layer, represents the sensor observation vector at time 𝑘, Represents the observation matrix, which is used to predict the state vector Mapped to the observation space, Indicates based on Always The estimated value of the state at time t, represents the actual physical state vector at time 𝑘, Represents the expected state value after executing the instruction at time 𝑘.
[0009] In another possible embodiment, the joint norm of the joint state space is discretized, including: using statistical confidence intervals to classify the residual abnormality of the information domain residual and the physical domain residual, and the classification design satisfies the following formula: ,in, represents the standard deviation of the residuals, express, 、 、 They correspond to the low-risk, medium-risk and high-risk states of the residual classification respectively.
[0010] In another possible embodiment, the action spaces of the attacker and defender are designed, and based on the joint state space design, a factorization strategy design is performed on the action spaces of the attacker and defender, including: designing the action space of the defender, based on the joint state space design, and designing a factorized hybrid defense strategy that satisfies the following formula: π x (s, A X )= ∏ i=1 4 [px i (s )] a x i · [1- px i (s) ] 1- a x i ,in, Pointing to the state Lower Defender Action subset The mixed strategy representation is: represents the probability that the defender chooses the defense method, The calculation satisfies the following formula: , represents the state feature vector, () indicates that the probability calculation falls within The sigmoid function within the range; design the attacker's action space, based on the joint state space design, and design a factorized hybrid attack strategy that satisfies the following formula: π y (s, A Y h )= ∏ j=1 2 py i (s) a y i · [1-( py i (s)] 1- a y i ,in, Pointing to the state Next attacker Action subset The mixed strategy representation is: Indicates the random probability of the attacker choosing an attack method.
[0011] Designing a profit function for calculating the defender's profit from defense, including: the calculation of the defender's profit function satisfies the following formula: ,in, represents the immediate defensive benefit, represents the discount factor, Indicates long-term benefits, Represents the defense cost.
[0012] Based on the joint residual change trend of the joint state space, a cross-layer coupled state transition model is established to simulate the dynamic coupling interaction between the information domain and the physical domain, including: calculating the attack and defense actions and Under the joint action of Transition to the successor state The conditional transition probability is calculated as follows: ,in, Represents the quantized residual from the state To status The inherent tendency of the state coupling potential, Indicates the action chosen by the defender The influence coefficient on state transition, Indicates the attacker's actions The influence coefficient on state transition, Indicates a situation in the subsequent state of the state transition; the computing system The overall transition probability satisfies the following formula: ,in, Indicates status Lower Defender Action The mixed strategy representation is: Indicates that the status Next attacker Action The mixed strategy representation of .
[0013] Solving the Nash equilibrium of the attacker's and defender's benefits includes: establishing the Q-value iterative update formula based on the Minimax-Q learning method: ,in, represents the Q value of the next step, represents the learning rate, represents the immediate reward at the current time step, represents the discount factor, Indicates at time , the defender chooses an action , the attacker chooses an action When the Q value is When the solution converges, the mixed strategy obtained at the time of convergence is extracted as the solution result of Nash equilibrium.
[0014] In the second aspect, the present invention also provides a cross-domain attack defense device based on game strategy, including: a state space construction unit, used to construct a joint state space, the joint state space includes information domain residuals and physical domain residuals, and the joint norm of the joint state space is discretized; a strategy design unit, used to design the action space of the attacker and the defender, and based on the joint state space design, the action space of the attacker and the defender is factorized to design a strategy; a benefit function design unit, used to design a benefit function for calculating the benefit obtained by the defender for defense; a model construction unit, used to establish a cross-layer coupled state transition model based on the change trend of the joint residual of the joint state space to simulate the dynamic coupling interaction between the information domain and the physical domain; a solving unit, used to formalize the state transition model into a zero-sum game, solve the Nash equilibrium of the attacker's benefit and the defender's benefit, and output the defender's strategy at the Nash equilibrium as the defense strategy.
[0015] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned cross-domain attack defense method based on game strategy is implemented.
[0016] In a fourth aspect, the present invention also provides an electronic device comprising: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned cross-domain attack defense method based on game strategy.
[0017] For the beneficial effects of the second to fourth aspects, please refer to the description of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A flowchart of a cross-domain attack defense method based on a game strategy provided by an embodiment of the present invention;
[0019] Figure 2 A schematic diagram of a cross-domain attack defense device based on a game strategy provided by an embodiment of the present invention;
[0020] Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be the common meanings understood by people with ordinary skills in the field to which the present invention belongs. The words "including" and similar words used in this article mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0022] This embodiment provides a cross-domain attack defense method based on game strategy. Figure 1 , the method comprising:
[0023] S101: Construct a joint state space, where the joint state space includes an information domain residual and a physical domain residual, and discretize the joint norm of the joint state space.
[0024] In one possible embodiment, the design of the joint state space includes designing an information domain state estimation residual that reflects the deviation between the current observation and the prediction. and the physical domain state residual, which measures the difference between the state result of the executed instruction and the expected result. . Information domain state estimation residual The meaning of is to quantify the difference between "observation-prediction" as residual, which is used to evaluate the accuracy of the information layer's understanding of the state of the physical process. Physical domain state residual The meaning of is to quantify the deviation of the execution result of the control instruction by the physical layer by comparing the state after the actual execution of the instruction with the expected state, and to capture the deviation caused by actuator failure or physical attack.
[0025] In a specific embodiment, the joint state space includes: ;in, The calculation satisfies the following formula: , The calculation satisfies the following formula: , represents the information domain residual, represents the physical domain residual, , represents the residual value of the information layer, Represents the residual value of the physical layer, represents the sensor observation vector at time 𝑘, Represents the observation matrix, which is used to predict the state vector Mapped to the observation space. The observation matrix reflects the state variables and their mapping relationships that can be perceived by the information layer in the system, and is a key component in calculating the state estimation error. Indicates based on Always The estimated value of the state at time t, represents the actual physical state vector at time 𝑘, represents the expected state value after executing the instruction at time 𝑘. The joint residual state drives the game strategy selection, and strategy execution generates new residual feedback, forming a closed loop of "state estimation-strategy decision-physical execution-residual update", thereby enabling cross-layer active defense.
[0026] In one possible embodiment, the joint norm is discretized into three risk levels, using the 3σ principle. Based on the statistical properties of Kalman filter residuals, when the residuals satisfy the Gaussian distribution assumption, approximately 99.7% of the samples fall within the ±3σ interval. Therefore, the "3σ principle" can be used to classify the degree of residual abnormality using statistical confidence intervals, with the boundaries set at 2σ and 4σ to strike a balance between false positive rate and false negative rate.
[0027] In a specific embodiment, the joint norm of the joint state space is discretized, including: using statistical confidence intervals to classify the residual abnormality of the information domain residual and the physical domain residual, and the classification design satisfies the following formula: ,in, represents the standard deviation of the residuals, express, 、 、 They correspond to the low-risk, medium-risk and high-risk states of the residual classification respectively.
[0028] S102: Design the action spaces of attackers and defenders. Based on the joint state space design, factorize the action spaces of attackers and defenders to design strategies.
[0029] In one possible embodiment, based on the above-mentioned joint state space design, a factorization strategy design is adopted for the action spaces of the defender and the attacker, decomposing the joint action space into a combination distribution of several actions, assuming that the actions are approximately conditionally independent at the strategy level. Designing the action spaces of the attacker and defender, and factorizing the action spaces of the attacker and defender based on the joint state space design, includes: designing the defender's action space, and designing a factorized hybrid defense strategy based on the joint state space design to satisfy the following formula: π x (s, A X )= ∏ i=1 4 [px i (s )] a x i · [1- px i (s) ] 1- a x i ,in, Pointing to the state Lower Defender Action subset The mixed strategy representation is: represents the probability that the defender chooses the defense method, The calculation satisfies the following formula: , represents the state feature vector, () indicates that the probability calculation falls within The sigmoid function within the range; design the attacker's action space, based on the joint state space design, and design a factorized hybrid attack strategy that satisfies the following formula: π y (s, A Y h )= ∏ j=1 2 py i (s) a y i · [1-( py i (s)] 1- a y i ,in, Pointing to the state Next attacker Action subset The mixed strategy representation is: Indicates the random probability of the attacker choosing an attack method.
[0030] In a specific embodiment, the action space of the defender is designed to have a total of 4 actions {Traffic rerouting, injecting fake vulnerabilities, device isolation, deploying decoy nodes}. That is, the action space can achieve 24 combinations: The factorized hybrid defense strategy is designed to satisfy the following formula: π x ( s , A X )= ∏ i =1 4 [ px i ( s )] a x i · [1- px i ( s ) ] 1- a x i , where the probability The specific design is: , the strategy chooses to accept the budget constraint: ,in, Indicates execution The cost, Represents the total budget.
[0031] Designing the attacker's action space for: , respectively, a DoS attack on the information layer and interference with the equipment at the physical layer. Design a hybrid strategy for the attacker for: π y (s, A Y h )= ∏ j=1 2 py i (s) a y i · [1-( py i (s)] 1- a y i .
[0032] S103: Design a benefit function for calculating the benefits gained by the defender from defending.
[0033] In one possible embodiment, the profit function refers to the profit that can be obtained by selecting a strategy for defense in the current state and quantifying the profit that can be obtained through defensive actions. According to the designed action space, the overall defense profit is designed to include the immediate defense profit and long-term benefits The design of the profit function is to guide the iterative selection of defense strategies through the profit function. Among them, the immediate defense profit Profit from blocking attacks and misleading earnings The calculation of the defender's payoff function satisfies the following formula: ,in, Represents the discount factor, which adjusts the ratio of current efficiency to long-term expected benefits.
[0034] Specifically, the cost of defense The calculation satisfies the following formula: , Indicates execution of an action Unit cost; direct benefits Calculation r dir (s, a x , a y )= r imm (s, a x , a y )+ r mis (s, a x , a y )= ∑ j∈ T v j (s) · [ p j no (s, a y )- p j def (s, a x , a y )]+η · ∑ h∈ T 1 { a x h =1} · 1 { a y h =1} · v h (s) , Represents the set of nodes in the system that may be attacked. Indicates the node under attack In state The value of Indicates the current status The probability of attack success without defense, Indicates the previous state The probability of attack success under defensive action, represents a misleading return adjustment factor, Indicates that a decoy node is deployed at node h. Indicates that the attacker hits node h, Indicates that the status The value of the next node h.
[0035] Only when "deployment" and "attack" occur simultaneously can the defender truly obtain misleading benefits, thereby encouraging the reasonable deployment of bait and improving the robustness of the strategy without excessive deployment caused by the deployment reward.
[0036] ,in, It represents the total benefit that the defender can obtain in state s.
[0037] The above formula is used to calculate a zero-sum game strategy update mechanism based on value iteration, which is a form of Minimax dynamic programming. Max refers to the number of possible actions. In the , the optimal strategy is selected to maximize the expected benefit; min means that under the premise that the defender chooses a strategy, the attacker tries to choose a strategy that minimizes the defender's benefit; Refers to the weighted expected value of considering all possible successor states s' from the current state s.
[0038] How the defender adjusts his actions when considering the worst-case attacker strategy; the future value is transmitted back to the current state to complete the value update of the strategy; the entire game model approaches the Nash equilibrium strategy during the dynamic evolution process.
[0039] S104: Based on the joint residual change trend of the joint state space, a cross-layer coupled state transition model is established to simulate the dynamic coupling interaction between the information domain and the physical domain.
[0040] In one possible embodiment, a cross-layer coupled state transition model is established based on the joint residual variation trend of the joint state space to simulate the dynamic coupling interaction between the information domain and the physical domain, including: calculating the attack and defense actions and Under the joint action of Transition to the successor state The conditional transition probability is calculated by the following formula: ,in, Represents the quantized residual from the state To status The inherent tendency of the state coupling potential, Indicates the action chosen by the defender The influence coefficient on state transition, Indicates the attacker's actions The influence coefficient on state transition, Indicates a situation in the subsequent state of the state transition; the computing system The overall transition probability satisfies the following formula: ,in, Indicates status Lower Defender Action The mixed strategy representation is: Indicates that the status Next attacker Action The mixed strategy representation of .
[0041] In a specific embodiment, the state transition probability is designed using a cross-layer coupled Softmax model, which can be used to calculate the attack and defense actions. and Under the joint action of Transition to the successor state The conditional transition probability is calculated as follows: ,in, The state coupling potential quantized residual from the state To status The inherent tendency of the system means that the system coupling preference is considered in the state transfer, that is, the impact of cross-layer interaction is taken into account to reflect the real evolution of the system; The action selected by the defender The influence coefficient on state transition reflects the defense effectiveness. Responding to the attacker's actions The impact coefficient on state transition reflects the degree of attack damage; refers to the set of successor states "Upgrade one level", "Downgrade one level" and "Keep" are the three possible situations of the subsequent state. In summary, the system has a The overall transition probability of is the global expectation of the action space: State transition probability Strictly follow the total probability formula to ensure the accuracy and consistency of the subsequent state distribution under any strategy mixture.
[0042] S105: Formalize the state transition model as a zero-sum game, solve the Nash equilibrium of the attacker's payoff and the defender's payoff, and output the defender's strategy at the Nash equilibrium as the defense strategy.
[0043] In one possible embodiment, solving the Nash equilibrium of the attacker's payoff and the defender's payoff includes: establishing a Q-value iterative update formula based on the Minimax-Q learning method: ,in, represents the Q value of the next step, that is, the estimated profit function after updating through game learning, represents the learning rate, represents the immediate reward at the current time step, calculated by the compound reward function, represents the discount factor, Indicates at time , the defender chooses an action , the attacker chooses an action When the Q value is Converges when , extract the mixed strategy obtained at convergence as the solution of Nash equilibrium, where, Indicates the preset Q value change threshold.
[0044] The defender's strategy achieves the goal of maximizing expected benefits, while the attacker's strategy achieves the goal of minimizing the defender's benefits. That is, neither party can obtain better results by unilaterally changing the strategy. At this time, the system reaches an "optimally stable" game equilibrium state, namely the Nash equilibrium state.
[0045] Specifically, when Time convergence means that the change in Q value meets the threshold change (that is, the change in Q value is relatively small and tends to be flat), and it is considered that the game has reached an approximate Nash equilibrium. At this time, the Nash equilibrium solution of the game can be obtained.
[0046] After obtaining the solution of Nash equilibrium, the defender strategy at the Nash equilibrium is output as the defense strategy.
[0047] The cross-domain attack defense method proposed in this paper, based on a game strategy, takes a stochastic game perspective and integrates state space and cross-layer reward functions to achieve integrated quantification of residuals at the information and physical layers. This method accurately assesses defense benefits at different risk levels while balancing immediate blocking with long-term returns. During the game process, the cascading reactions between the information and physical domains are integrated into a coupled state transition model, realistically simulating cross-layer evolution and enabling dynamic strategy adaptation in an incomplete information environment. A mechanism for deploying fake high-value nodes is also introduced to incorporate misleading benefits into cross-layer rewards. Overall, the cross-domain attack defense method based on a game strategy, while taking into account cascading reactions, unifies state modeling to clearly quantify the security state of cyber-physical systems. Active induction capabilities are considered in the design of defense strategies and action spaces. A rational reward function is designed to guide optimal defense decisions. This allows defense strategies to gradually approach the optimal adversarial strategy through continuous interactive updates under conditions of incomplete information, ultimately improving the effectiveness and robustness of cross-layer defense for complex physical cyber systems (CPSs).
[0048] The cross-domain attack defense method based on game strategy of the present invention is applicable to cross-domain attack and defense scenarios in information-physical systems, and can defend against the characteristics of cross-domain attacks in information-physical systems, such as dynamics and concealment. In terms of state space design, the residual norm of the Kalman filter is used as the core indicator, and the state estimation residuals of the information domain and the physical domain are uniformly discretized into three risk levels, and a two-dimensional combined state space is constructed, which can simultaneously characterize the estimation deviation and security situation, and design a cross-layer reward function that takes into account both immediate benefits (including immediate blocking defense and misleading defense) and long-term benefits. In terms of strategy space design, a factorized hybrid strategy modeling method is proposed. This method assumes that each atomic action is approximately conditionally independent at the strategy level. By modeling the execution probability of each atomic action as a state-dependent Sigmoid function, the joint strategy distribution is represented as the product of each sub-strategy, which can significantly reduce the strategy dimension and improve the scalability and computational efficiency of strategy learning. This approach considers the ability of defenders to execute immediate blocking actions, effectively interrupting attacks and curbing their spread. Furthermore, by deploying fake nodes with high decoy value, attackers are misled about key system targets, causing their attacks to deviate from their true, high-value targets. This wastes attack resources while further improving the protection of core resources, thereby enhancing the effectiveness of defense strategies. Regarding state transition probability design, a logarithmic linear softmax model, incorporating the residual coupling coefficient and the influence of attack and defense actions, is constructed to accurately simulate the dynamic coupling interaction between the information layer and the physical layer. The subsequent state is designed to be either a first-level rise or fall or a hold, effectively reflecting the cross-layer response process. This model is formalized as a zero-sum game and the Nash equilibrium is solved, enabling optimal adversarial defense.
[0049] See the instructions attached Figure 2 This embodiment also provides a cross-domain attack defense device based on a game strategy, which is used to implement the above method embodiment. The device includes:
[0050] The state space construction unit 201 is used to construct a joint state space, where the joint state space includes information domain residual and physical domain residual, and discretize the joint norm of the joint state space.
[0051] The strategy design unit 202 is used to design the action space of the attacker and the defender, and to perform factorization strategy design on the action space of the attacker and the defender based on the joint state space design.
[0052] The profit function design unit 203 is used to design a profit function for calculating the profit obtained by the defender from performing defense.
[0053] The model building unit 204 is configured to establish a cross-layer coupled state transition model based on the joint residual variation trend of the joint state space to simulate the dynamic coupling interaction between the information domain and the physical domain.
[0054] The solving unit 205 is configured to formalize the state transition model into a zero-sum game, solve the Nash equilibrium of the attacker's payoff and the defender's payoff, and output the defender's strategy at the Nash equilibrium as the defense strategy.
[0055] All relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0056] In other embodiments of the present application, the present application discloses an electronic device, such as Figure 3 As shown, the electronic device 300 may include: one or more processors 301; a memory 302; a display 303; one or more applications (not shown); and one or more computer programs 304. The above components may be connected via one or more communication buses 305. The one or more computer programs 304 are stored in the above memory and configured to be executed by the one or more processors 301. The one or more computer programs 304 include instructions, which may be used to execute the following: Figure 1 、 Figure 2 and each step in the corresponding embodiment.
[0057] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0058] The functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0059] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: flash memory, mobile hard disk, read-only memory, random access memory, disk or optical disk, and other media that can store program code.
[0060] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A cross-domain attack defense method based on game strategy, characterized in that: include: Constructing a joint state space, the joint state space including an information domain residual and a physical domain residual, and discretizing a joint norm of the joint state space; Designing the action spaces of the attacker and defender, and designing factorization strategies for the action spaces of the attacker and defender based on the joint state space design; Design a payoff function to calculate the payoff to the defender from defending; A cross-layer coupled state transition model is established based on the joint residual variation trend of the joint state space to simulate the dynamic coupling interaction between the information domain and the physical domain; The state transition model is formalized as a zero-sum game, the Nash equilibrium of the attacker's payoff and the defender's payoff is solved, and the defender's strategy at the Nash equilibrium is output as the defense strategy.
2. The method according to claim 1, characterized in that Constructing a joint state space, wherein the joint state space includes information domain residuals and physical domain residuals, including: The joint state space includes: ; in, The calculation satisfies the following formula: , The calculation satisfies the following formula: , represents the information domain residual, represents the physical domain residual, , represents the residual value of the information layer, Represents the residual value of the physical layer, represents the sensor observation vector at time k, Indicates based on Always The estimated value of the state at time t, Represents the observation matrix, which is used to predict the state vector Mapped to the observation space, represents the actual physical state vector at time k, Indicates the expected state value after executing the instruction at time k.
3. The method according to claim 1, characterized in that Discretizing the joint norm of the joint state space includes: The statistical confidence interval classifies the residual abnormality of the information domain residual and the physical domain residual, and the classification design satisfies the following formula: ,in, represents the standard deviation of the residuals, 、 、 They correspond to the low-risk, medium-risk and high-risk states of the residual classification respectively.
4. The method according to claim 1, wherein Design the action spaces of the attacker and defender. Based on the joint state space design, factorize the action spaces of the attacker and defender to design strategies, including: Design the defender's action space. Based on the joint state space design, design a factorized hybrid defense strategy that satisfies the following formula: ,in, Pointing to the state Lower Defender Action subset The mixed strategy representation is: represents the probability that the defender chooses the defense method, The calculation satisfies the following formula: , represents the state feature vector, () indicates that the probability calculation falls within sigmoid function within the range; Design the attacker's action space. Based on the joint state space design, design a factorized hybrid attack strategy that satisfies the following formula: ,in, Pointing to the state Next attacker Action subset The mixed strategy representation is: Indicates the random probability of the attacker choosing an attack method.
5. The method according to claim 4, characterized in that Design a reward function to calculate the defender's benefits from defending, including: The calculation of the defender's payoff function satisfies the following formula: ,in, represents the immediate defensive benefit, represents the discount factor, Indicates long-term benefits, Represents the defense cost.
6. The method according to claim 4, characterized in that A cross-layer coupled state transition model is established based on the joint residual variation trend of the joint state space to simulate the dynamic coupling interaction between the information domain and the physical domain, including: Calculate the attack and defense actions and Under the joint action of Transition to the successor state The conditional transition probability is calculated by the following formula: ,in, Represents the quantized residual from the state To status The inherent tendency of the state coupling potential, Indicates the action chosen by the defender The influence coefficient on state transition, Indicates the attacker's actions The influence coefficient on state transition, Indicates a situation in the subsequent state of the state transition; Computing system state The overall transition probability satisfies the following formula: ,in, Indicates status Lower Defender Action The mixed strategy representation is: Indicates that the status Next attacker Action The mixed strategy representation of .
7. The method according to claim 4, characterized in that Solve the Nash equilibrium of the attacker's payoff and the defender's payoff, including: The Q value iterative update formula based on the Minimax-Q learning method is: ,in, represents the Q value of the next step, represents the learning rate, represents the immediate reward at the current time step, represents the discount factor, Indicates at time , the defender chooses an action , the attacker chooses an action Q value when when When the solution converges, the mixed strategy obtained at the time of convergence is extracted as the solution result of Nash equilibrium.
8. A cross-domain attack defense device based on game strategy, characterized in that: The device comprises: A state space construction unit is used to construct a joint state space, wherein the joint state space includes an information domain residual and a physical domain residual, and discretize a joint norm of the joint state space; The strategy design unit is used to design the action space of the attacker and defender. Based on the joint state space design, the action space of the attacker and defender is factorized to design the strategy. A profit function design unit, used to design a profit function for calculating the profit obtained by the defender from performing defense; A model building unit, configured to establish a cross-layer coupled state transition model based on a joint residual variation trend of the joint state space to simulate the dynamic coupling interaction between the information domain and the physical domain; The solving unit is used to formalize the state transition model into a zero-sum game, solve the Nash equilibrium of the attacker's payoff and the defender's payoff, and output the defender's strategy at the Nash equilibrium as the defense strategy.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cross-domain attack defense method based on game strategy according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory, so that the electronic device executes the cross-domain attack defense method based on game strategy according to any one of claims 1 to 7.
Citation Information
Patent Citations
Power distribution network defense strategy design method based on dynamic game under network attack
CN115348064A
Optimization control method for human-in-loop multi-agent system under FDI attack
CN119758816A