A game-based CPPS optimal resource allocation method against FDI attacks

By building a game model to simulate the state offset after the FDI attack, calculating the Nash equilibrium for resource allocation, the problem of inefficient defense in the information physical power system is solved, the optimal defense strategy under limited resources is realized, and the system's security and economics are improved.

CN115580423BActive Publication Date: 2025-08-12ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210962376.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2025-08-12
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

The prior art defense methods for counterfeit data injection attacks in information physics power systems fail to fully consider the balance between information physics layer fusion and reliability and economy, resulting in inefficient defense.

Method used

By simulating the state estimation offset after the FDI attack as profit, a game model is built, Nash equilibrium is calculated and resource allocation is carried out to achieve the optimal defense strategy, and multi-node resource allocation is considered in the case of limited resources.

Benefits of technology

On the basis of ensuring system security, defense efficiency is improved, flexible and efficient defense measures are provided, and various state estimation methods and defense means are adapted to.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115580423B_ABST
    Figure CN115580423B_ABST
Patent Text Reader

Abstract

This invention provides a game-based CPPS optimal resource allocation method for FDI attacks. The method first obtains the system state estimation results under normal conditions, then simulates the state estimation results after an FDI attack. The offset between the two is used as the participant's payoff. The corresponding payoff amount, the resources required by the attacker to launch the attack, and the resources required by the defender to provide defense are input into the game model. The Nash equilibrium is calculated to obtain the corresponding strategy and expected payoff. Finally, resources are allocated to multiple nodes under limited resources. This allocation method improves defense efficiency while ensuring system security from the defender's perspective, providing feasible suggestions and references for defenders to configure defensive measures. Furthermore, the method can adapt to multiple state estimation methods and different defense methods, showing high flexibility and scalability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of smart grid security and relates to a game-based cyber-physical power system (CPPS) optimal resource allocation method for false data injection (FDI) attacks. Background Art

[0002] With the development of China's power industry and the intelligentization of power systems, basic power systems are gradually transforming into complex cyber-physical power systems. This improves system efficiency and connection availability while also bringing more security risks.

[0003] FDI attacks, the most widespread type of attack with cyber-physical layer characteristics, pose a significant threat to cyber-physical power systems. FDI attacks involve 1. developing attack targets and strategies; 2. tampering with measurement results from control instruments, communication networks, and master stations; 3. transmitting forged measurements to the energy management system (EMS) and interfering with state estimation results; and 4. inducing the control center to take emergency measures, tripping critical lines and causing power loss and overvoltage or undervoltage. Attackers primarily target two types of attacks: 1. system state variables (voltage, phase angle, etc.); and 2. device measurements (voltage, active power, etc.). State estimation methods are primarily categorized into dynamic state estimation (such as the Kalman filter algorithm) and static state estimation (such as the least squares method).

[0004] Game theory, a proven and effective formal tool, quantifies the interaction between attackers and defenders and provides a sound theoretical framework to guide defenders in optimizing resource allocation strategies under limited resources. Game theory can be primarily categorized into cooperative and non-cooperative games based on the relationships between the players. Furthermore, it can be further categorized into dynamic and static games, complete and incomplete information games, and zero-sum and non-zero-sum games based on the number of moves, level of understanding, and payoffs between the attacker and defender, as shown in Table 1. The payoff of a strategy is a crucial basis for rational decision-making by each player in game theory. Game theory, through theoretical analysis and research, can select the most profitable decision-making strategy for each player. The validity of this strategy is primarily determined by the fact that all rational players will consciously adhere to the equilibrium strategy derived from game theory, with no player independently deviating from it. Under equilibrium strategies, each player's strategy is the optimal response to the strategies of the others. Currently, the general approach to applying game theory to cyber attack and defense strategies within the context of CPPS is to model attack and defense behaviors using game theory, quantitatively assessing attack and defense resources, consequences, and action strategies; then, to find equilibrium points and determine the optimal attack and defense strategy. Modeling is done from the defender's perspective, with the goal of minimizing the damage caused by the attack; or from the attacker's perspective, with the goal of maximizing the damage caused by the attacker. Ultimately, the optimal game strategy is developed.

[0005] Existing research, such as on firewalls, interface control, and permission checking, mostly focuses solely on the cyber or physical layers. At the cyber layer, some researchers have proposed using active defenses to protect critical nodes by inducing attackers to launch attacks; utilizing redundant detection systems to ensure data accuracy; and employing multiple keys to protect transmission security. At the physical layer, some research proposes adding detection equipment to provide more comprehensive protection. However, these approaches fail to fully consider the convergence of cyber and physical layers and the resulting trade-off between reliability and cost-effectiveness. Summary of the Invention

[0006] This invention provides a CPPS optimal resource allocation game method for FDI attacks. First, the system state estimation results under normal conditions are obtained. Then, the state estimation results after an FDI attack are simulated. The offset between the two is used as the participant's payoff. The corresponding payoff amount, the resources required by the attacker to launch the attack, and the resources required by the defender to provide defense are input into the game model. The Nash equilibrium is calculated, and the corresponding strategy and expected payoff are obtained. Finally, resources are allocated to multiple nodes under limited resources. This achieves the optimal allocation strategy under limited resources, taking into account economic efficiency while ensuring security.

[0007] The present invention provides a game-based CPPS optimal resource allocation method for FDI attacks, which includes the following steps:

[0008] S1: Obtain the power system configuration, which includes a power system topology diagram and parameters of each branch of the power system; the power system topology diagram is a network structure diagram composed of network node devices and communication media, and a network node in the power system topology diagram represents an active electronic device connected to the network in the power system, which can send, receive or forward information through a communication channel.

[0009] S2: Calculate the power state quantity of the power system under normal conditions, wherein the power state quantity includes the phase angle θ and the voltage V;

[0010] S3: Traverse all attacker actions and defender actions for network nodes 1,…,n and calculate the corresponding power state quantities;

[0011] S4: Calculate the power state quantity deviation to obtain the corresponding participant benefits;

[0012] S5: Input the participant benefits and the resource consumption required to perform their actions into the game model;

[0013] S6: Calculate the optimal defense resource allocation plan for a single network node;

[0014] S7: Under limited resources, traverse all network nodes to calculate the multi-node defense resource configuration plan that maximizes the total expected benefit.

[0015] As a preferred embodiment of the present invention, the S2 is specifically:

[0016] Based on the relationship between the power system instrument measurement value and the state variable z=h(x)+e, the power system estimated state variable x, i.e., the power state quantity, is obtained according to the instrument measurement value. The power state quantity includes the phase angle θ and the voltage V;

[0017] Where h(x)=[h1(x),...,h i (x),...,h m (x)] T is the measurement function, z is the value based on the instrument measurement, that is, the instrument measurement value, e is the independent random measurement noise; m is the number of measuring devices.

[0018] As a preferred solution of the present invention, the method for calculating the corresponding power state quantity in step S3 is the same as the method for calculating the power state quantity in step S2.

[0019] As a preferred solution of the present invention, the participant benefit in step S4 is a power state quantity offset, which is the difference between the power state quantity calculated in step S3 and the power state quantity under normal circumstances in step S2.

[0020] As a preferred embodiment of the present invention, in the game model described in step S5,

[0021] The attacker can choose a certain attack action A or not to attack, and the attacker needs to pay the resource consumption It is related to the attack action selected by the defender; the defender can choose not to defend or take the corresponding defensive action D, and the defender needs to pay the corresponding resources. To complete these actions;

[0022] The success rate of the attacker when launching an attack Pr(A i ,D j ) is related to the actions of the attacker and the defender; at the same time, the success rate Pr(A i ,D j ) and the vulnerability of devices on network nodes V Pr Also related to; Vulnerability V Pr Defined by Common Vulnerability Scoring System metrics: V Pr =2×S AV ×S AC ×S AU , where S AV is the attack path, S AC is the attack complexity, SAU It is the degree of certification;

[0023] For node n, define and are the collection of the probabilities of actions that the attacker and defender can take at the node,

[0024]

[0025]

[0026] when The attacker's strategy is considered a pure strategy; otherwise, the attacker's strategy is a mixed strategy.

[0027] As a preferred embodiment of the present invention, step S6 is specifically as follows:

[0028] For node n, given the attacker-defender strategy pair The expected return is expressed as:

[0029]

[0030]

[0031] in and The attacker and defender are in action The benefit is obtained by injecting false data into the physical layer equipment of the target power system to simulate the offset of the state quantity compared with the normal operation;

[0032] In the game model, both the attacker and the defender want to maximize their benefits. When they choose a strategy that neither party will change, it is called a Nash equilibrium. Assume that for any defender strategy All exist The attacker's expected profit Maximum, while for any attacker strategy All exist The defender's expected return Maximum, then the Nash equilibrium is reached;

[0033] For node n, the defender strategy that reaches Nash equilibrium is solved as the optimal defense resource allocation solution for a single network node n.

[0034] Compared to existing technologies, this invention utilizes state estimation methods to simulate the post-attack system state offset as the baseline payoff for participants. Taking into account their action costs, the proposed method calculates Nash equilibrium and expected payoff under equilibrium within a game theory framework. This allocation method, approached from the defender's perspective, improves defense efficiency while ensuring system security, providing feasible advice and reference for defenders in deploying defensive strategies. Furthermore, this method is adaptable to a variety of state estimation methods and defense strategies, offering high flexibility and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is an overview diagram of the system design targeted by the method of the present invention;

[0036] Figure 2 It is a schematic flow diagram of the method of the present invention;

[0037] Figure 3 This is a comparison chart of the expected benefits of the strategy proposed in this invention and the resource allocation based only on the vulnerability index. DETAILED DESCRIPTION

[0038] like Figure 1 and 2 As shown, this paper constructs a static game model based on incomplete information and develops a game-based CPPS optimal resource allocation method for FDI attacks. This method uses state estimation to simulate and obtain system state offsets to quantify the attack results as a profit indicator. This profit and the corresponding action costs are added to the game matrix and a Nash equilibrium is calculated. Ultimately, a hybrid optimal strategy is obtained to achieve reasonable allocation under limited resources.

[0039] The system design for which the method of the present invention is directed is summarized as follows: Figure 1 As shown, Figure 1 The physical layer in the power system is the physical layer of the power system. The physical layer equipment of the power system includes control equipment, execution equipment and measurement equipment. Under normal circumstances (i.e., without attack), there are only Figure 1 In steps ③, ④, and ⑤, the physical layer monitors system components based on instruments and reports their readings to the control center. The control center estimates the status of the power system based on these instrument measurements and sends corresponding commands to the physical layer devices.

[0040] The relationship between the power system instrument measurement value and the state variable is z = h(x) + e, where x is the estimated state variable, z is the value based on the instrument measurement, and e is the independent random measurement noise. The measurement function h(x) = [h1(x),...,h i (x),...,h m (x)] TThe subscript m is the number of measuring devices, which depends on the specific measurement type and the network topology and parameters of the power system.

[0041] The method commonly used for state estimation is weighted least squares, whose goal is to find a set of system states that minimizes the sum of squared errors F(x):

[0042] min F(x)=[zh(x)] T W[zh(x)]

[0043] Where W is the weighting matrix, which can be designed as the variance The reciprocal diagonal matrix of The higher σ is, the lower the confidence level of the equipment is. It is usually provided by the manufacturer. The measurement value with higher confidence level has higher weight.

[0044] Derivative this equation to make it zero:

[0045]

[0046] in is the least squares estimate of x, is the Jacobian matrix. Combining the Newton-Raphson iterative relationship, we get the formula:

[0047] [G(x k )]Δx k =H T (x k )W[zh(x k )]

[0048] G(x k )=H T (x k )WH(x k )

[0049] x k+1 =x k +Δx k

[0050] Where G(x k ) is the gain matrix, x k+1 =x k +Δx k Used to update the Jacobian matrix, repeat the iteration until Δx is less than the pre-set convergence condition, x k It is the estimated result of the system state.

[0051] When attacked by FDI, Figure 1As shown in steps ① and ②, the attacker tampers with the measurement results by controlling the instrument, communication network, and master station to interfere with the estimation results, inducing the control center to take emergency measures, as shown in step ⑥, causing the critical line to trip, resulting in power loss and overvoltage or undervoltage.

[0052] refer to Figure 1 and Figure 2 The steps of the game-based CPPS optimal resource allocation method for FDI attacks proposed in the present invention are as follows:

[0053] S1: Obtain the power system configuration from matpower, where the power system configuration includes a power system topology diagram and parameters of each branch of the power system; the power system topology diagram is a network structure diagram consisting of network node devices and communication media, and a network node in the power system topology diagram represents a device in the power system.

[0054] S2: Calculate the power state quantity of the power system under normal conditions, wherein the power state quantity includes the phase angle θ and the voltage V;

[0055] Based on the relationship between the power system instrument measurement value and the state variable z=h(x)+e, the power system estimated state variable x, i.e., the power state quantity, is obtained according to the instrument measurement value. The power state quantity includes the phase angle θ and the voltage V;

[0056] Where h(x)=[h1(x),...,h i (x),...,h m (x)] T is the measurement function, z is the value based on the instrument measurement, that is, the instrument measurement value, e is the independent random measurement noise; m is the number of measuring devices.

[0057] S3: Traverse all attacker actions and defender actions for network nodes 1, ..., n and calculate the corresponding power state quantities, which include phase angle θ and voltage V; the method for calculating the corresponding power state quantities in step S3 is the same as the method for calculating the power state quantities in step S2.

[0058] S4: Calculate the power state quantity offset to obtain the corresponding participant benefit; the participant benefit is the power state quantity offset, which is the difference between the power state quantity calculated in step S3 and the power state quantity under normal circumstances in step S2.

[0059] S5: Input the participant benefits and the resource consumption required to perform their actions into the game model;

[0060] The game model mainly includes concepts such as participants, actions, strategy pairs, benefits, and rewards.

[0061] Participants are attackers and defenders. The attacker's action A includes not attacking and attacking. In the case of attacking, he can choose the intensity of the attack (i.e. the range of injected data), the target of the attack (measurement value or state value), and the location of the attack (single point attack or multi-point coordinated attack). Obviously, the attacker needs to pay the resource consumption. It is related to the attack action it chooses. On the other hand, the defender can also choose not to defend or take corresponding defensive action D, including increasing the defense intensity (such as increasing the monitor scanning frequency, etc.), changing the defense plan (such as modifying the detection algorithm accuracy), expanding the defense nodes (such as setting up monitoring equipment on more nodes), etc. Similarly, the defender needs to spend corresponding resources. to complete these actions.

[0062] The success rate of the attacker when launching an attack Pr(A i ,D j ) is related to the actions of the attacker and the defender. At the same time, the success rate Pr(A i ,D j ) and the vulnerability of the device on the node V Pr Also related. Vulnerability V Pr Defined by Common Vulnerability Scoring System (CVSS) metrics: V Pr =2×S AV ×S AC ×S AU , where S AV is the attack path, S AC is the attack complexity, S AU It's the degree of certification.

[0063] For node n, the attacker's strategy and the defender's strategy can be expressed as follows:

[0064]

[0065]

[0066] in and are the set of probabilities of actions that the attacker and defender can take at the node. Obviously, the probability interval should be in [0,1], and the sum of the probabilities is 1. The attacker's strategy can be considered a pure strategy, otherwise it is a mixed strategy. The same is true for the defender.

[0067] S6: Calculate the optimal defense resource configuration plan for a single network node;

[0068] For node n, given the attacker-defender strategy pair is the attacker’s expected profit, is the defender's expected payoff; and It can be expressed as:

[0069]

[0070]

[0071] in and The attacker and defender are in action The benefit is obtained by comparing the offset of the state quantity during the false data injection simulation attack (such as step 126) and the normal operation (such as step 345) of the physical layer device of the target system.

[0072] In a game problem, both the attacker and the defender want to maximize their benefits. When they choose a strategy that neither party will change, it is called a Nash equilibrium. Assume that for any defender strategy All exist The attacker's expected profit Maximum, while for any attacker strategy All exist The defender's expected return The maximum value is the Nash equilibrium. For node n, the defender strategy is solved to achieve the Nash equilibrium, which is the optimal defense resource allocation solution for a single network node n.

[0073] S7: Due to attacker resources and defender resources It's not infinite. Under the constraints of limited resources, we traverse all network nodes and calculate a multi-node defense resource allocation plan that maximizes the expected benefit for the overall defender. That is, given limited resources, we allocate resources across n nodes. The resource allocation plan is an optimization problem that seeks to maximize the expected benefit for the defender. Using an iterative approach, we can find a multi-node defense resource allocation plan that maximizes the expected benefit for the overall defender.

[0074] The effectiveness and feasibility of the present invention are illustrated by simulation experiments. The IEEE 33 bus system is used to simulate the attack and defense game. The configuration information of the IEEE 33 bus system is obtained from the MATPOWER package, and the power flow is calculated using MATLAB software. The attack actions are classified from three aspects: 1. Attack intensity: 5%, 10%, 20%, 2. Attack location, selection of attack nodes and whether to conduct multi-point attacks, 3. Attack target: measurement type (voltage, current, strategy, etc.). The corresponding defense actions are divided here into: 1. Defense intensity: reflected in the accuracy requirements. The higher the intensity, the stricter the accuracy requirements. 2. Defense nodes: selection of defense nodes and whether to conduct collaborative defense.

[0075] Static defense is chosen because the defender cannot dynamically change its resource allocation and can only operate based on prior assumptions. This means that once the defender has configured its equipment, changes take considerable time. When resource investment exceeds a certain level, even a successful attack will not generate sufficient returns to offset the resources consumed. Therefore, considering the defender's perspective in relation to the attacker's resources, once the resources required for a successful attack exceed a certain threshold, the attacker (rationally) will not launch an attack, allowing the defender to choose not to defend.

[0076] like Figure 3 As shown, for the defender, the expected benefit result of resource allocation based on the strategy proposed by the present invention (i.e. ) is better than the expected benefit of allocating resources based only on vulnerability indicators (i.e. ) has increased significantly.

[0077] The above-described embodiments merely illustrate several implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that variations and modifications are possible within the scope of the present invention, as would be apparent to one skilled in the art. These variations and modifications fall within the scope of the present invention. Therefore, the scope of the present invention shall be determined by the appended claims.

Claims

1. A game-based CPPS optimal resource allocation method for FDI attacks, characterized by: The steps include: S1: Obtaining a power system configuration, wherein the power system configuration includes a power system topology diagram and parameters of each branch of the power system; the power system topology diagram is a network structure diagram consisting of network node devices and communication media; S2: Calculate the power state quantity of the power system under normal conditions, wherein the power state quantity includes the phase angle θ and the voltage V; S3: Traverse all attacker actions and defender actions for network nodes 1,…,n and calculate the corresponding power state quantities; S4: Calculate the power state quantity deviation to obtain the corresponding participant benefits; S5: Input the participant benefits and the resource consumption required to perform their actions into the game model; S6: Calculate the optimal defense resource configuration plan for a single network node; S7: Under limited resources, traverse all network nodes to calculate the multi-node defense resource configuration plan that maximizes the total expected benefit.

2. The game-based CPPS optimal resource allocation method for FDI attacks according to claim 1 is characterized in that: The S2 is specifically: Based on the relationship between the power system instrument measurement value and the state variable z=h(x)+e, the power system estimated state variable x, i.e., the power state quantity, is obtained according to the instrument measurement value. The power state quantity includes the phase angle θ and the voltage V; Where h(x)=[h1(x),...,h i (x),...,h m (x)] T is the measurement function, z is the value based on the instrument measurement, that is, the instrument measurement value, e is the independent random measurement noise; m is the number of measuring devices.

3. The game-based CPPS optimal resource allocation method for FDI attacks according to claim 1 is characterized in that: The method of calculating the corresponding power state quantity in step S3 is the same as the method of calculating the power state quantity in step S2.

4. The game-based CPPS optimal resource allocation method for FDI attacks according to claim 1 is characterized in that: The participant benefit in step S4 is the power state quantity offset, which is the difference between the power state quantity calculated in step S3 and the power state quantity under normal circumstances in step S2.

5. The game-based CPPS optimal resource allocation method for FDI attacks according to claim 1, characterized in that: In the game model described in step S5, The attacker can choose a certain attack action A or not to attack, and the attacker needs to pay the resource consumption It is related to the attack action chosen by the defender; the defender can choose not to defend or take the corresponding defensive action D, and the defender needs to pay the corresponding resources. To complete these actions; The success rate of the attacker when launching an attack Pr(A i ,D j ) is related to the actions of the attacker and the defender; at the same time, the success rate Pr(A i ,D j ) and the vulnerability of devices on network nodes V Pr Also related to; Vulnerability V Pr Defined by Common Vulnerability Scoring System metrics: V Pr =2×S AV ×S AC ×S AU , where S AV is the attack path, S AC is the attack complexity, S AU It is the degree of certification; For node n, define and are the collection of the probabilities of actions that the attacker and defender can take at the node, when The attacker's strategy is considered a pure strategy; otherwise, the attacker's strategy is a mixed strategy.

6. The game-based CPPS optimal resource allocation method against FDI attacks according to claim 5, characterized in that: The step S6 is specifically as follows: For node n, given the attacker-defender strategy pair The expected return is expressed as: in and The attacker and defender are in action The benefit is obtained by injecting false data into the physical layer equipment of the target power system to simulate the offset of the state quantity compared with the normal operation; In the game model, both the attacker and the defender want to maximize their benefits. When they choose a strategy that neither party will change, it is called a Nash equilibrium. Assume that for any defender strategy All exist The attacker's expected profit Maximum, while for any attacker strategy All exist The defender's expected return Maximum, then the Nash equilibrium is reached; For node n, the defender strategy that reaches Nash equilibrium is solved as the optimal defense resource allocation solution for a single network node n.

Citation Information

Patent Citations

  • Network safety optimum attacking and defending decision method for attacking and defending game

    CN103152345A

  • Information security analysis using game theory and simulation

    US20140157415A1