FDIA construction method based on data model fusion driving under incomplete information
By using measurement data equivalent and neural network fitting the current model under incomplete information conditions, a mixed integer programming problem is constructed, and an FDIA attack that triggers the limit of power grid lines under local information is realized, which solves the problem of unclear attack effects caused by insufficient information in the existing technology, ensuring the effectiveness and concealment of the attack.
Patent Information
- Application Number
- CN202510618213.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
AI Technical Summary
The existing FDIA model assumes that the attacker has mastered complete grid topology and line parameters, making it difficult to construct effective false data injection attacks in practical applications and cannot trigger clear physical consequences, such as line limits.
By collecting measurement data from the attack area, using Kirchoff's law to perform measurement value equivalents, training a single-layer neural network to fit the nonlinear mapping relationship of the current equation, and constructing a single-layer mixed integer linear programming problem. You only need to master the power injection of nodes and line flow information to construct an attack vector, trigger residual pollution and cause the line flow estimation result to exceed the limit.
Under incomplete information conditions, the FDIA attack was successfully constructed, affecting the status estimation results and triggering line limits, threatening the safe and stable economic operation of the power system.
Smart Images

Figure CN120498040A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electric power monitoring, and in particular relates to a FDIA construction method based on data model fusion drive under incomplete information. Background Art
[0002] With electricity as the core pillar of modern urban operations, the cybersecurity of its infrastructure has become a strategic priority for global attention in the digital age. Driven by the Internet of Everything (IoE), power monitoring systems are facing unprecedented security challenges. As the nerve center of power grid operations, the power supervisory control and data acquisition (SCADA) system collects network-wide measurement data in real time, performing key functions such as state estimation and the detection, identification, and elimination of bad data, providing high-precision data support for dispatching decisions. If this core system were to suffer a cyberattack, it would trigger a chain reaction, leading to distorted dispatching instructions and a reduction in grid stability margins. This could not only cause significant economic losses but also jeopardize the operation of critical infrastructure, ultimately impacting the normal operation of society and the safety of people.
[0003] False data injection attack (FDIA) is a cyberattack technique targeting power SCADA systems, proposed by Liu et al. It injects false data into uploaded measurement information to circumvent the bad data detection phase of state estimation, causing the dispatch center to misjudge the power grid state. FDIA is characterized by high concealment, high flexibility, wide attack range, and strong destructive power. Existing FDIA research mainly focuses on attack modeling and detection. Attack detection, from the defender's perspective, studies how to quickly and accurately detect FDIA. Attack modeling, on the other hand, takes the attacker's perspective, exploring system vulnerabilities at the operational level and studying attack mechanisms, applicable conditions, and optimal strategies.
[0004] In the early days of FDIA attack modeling, scholars generally assumed that attackers possessed sufficient attack resources. Liu et al. first derived the basic form of FDIA for a DC power flow model. Subsequently, Kosut et al. used graph theory to analyze the minimum set of FDIA measurement attacks that could render the network unobservable. Yuan et al. constructed a two-level optimization model for FDIA, maximizing power system operating costs as the objective function. Furthermore, Kim et al. took a unique approach, studying an FDIA model that accounts for residual contamination (where bad data points cause the residuals of other normal measurements to exceed a threshold and are incorrectly eliminated). This relaxed the FDIA's requirement for avoiding bad data detection and formulated the optimal attack strategy model as a convex quadratic constrained quadratic programming problem. All of these models assume that the attacker possesses complete knowledge of the grid topology and line parameters. However, this information is typically stored in highly secure dispatching centers, making it difficult for attackers to obtain, limiting the models' practical application value.
[0005] As academic research deepens, scholars have begun studying FDIA under incomplete information. Liu et al. first revealed the attack mechanism of FDIA with incomplete information, proposed a heuristic algorithm for finding the optimal attack area, and expanded the scenario to AC power flow models. Deng et al. further relaxed the requirement for the reactance of all lines in the attack area. Subsequently, Yu et al. proposed a blind FDIA model based on principal component analysis, which only requires measurement data from multiple sections to launch an attack. Building on this, several researchers have subsequently proposed blind FDIA methods based on robust principal component analysis, kernel principal component analysis, matrix reconstruction, and subspace estimation, enhancing robustness in measurement noise scenarios.
[0006] Most existing FIDA models assume that attackers have complete knowledge of the grid topology and line parameters. However, this information is usually stored in dispatching centers with a high level of security protection, making it difficult for attackers to obtain it, resulting in limited practical application value of the model. However, models that take into account real-world scenarios with insufficient attack resources intelligently implement arbitrary tampering of measurements without leading to clear physical consequences, such as line over-limit. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide an FDIA construction method based on data model fusion drive under incomplete information.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A FDIA construction method based on data model fusion drive under incomplete information, including:
[0010] Step S1: Collect measurement data of the attack area and form a measurement history database;
[0011] Step S2: Based on the measurement history database, the measurement values of the busbars located at the boundary of the local area are merged and equivalent, so that the measurement in the local area satisfies Kirchhoff's law;
[0012] Step S3: Extract the active power measurement data equivalent to the attack area measurement, use the node injected active power as input, and the line active power flow as output, train a single-layer neural network with a through layer, fit the nonlinear mapping relationship of the power flow equation in the area, and linearize the mapping relationship using the big M method;
[0013] Step S4: construct a single-level mixed integer linear programming problem with line crossing as the attack target, and use the CPLEX solver to solve it.
[0014] Preferably, the attacker can only obtain the measurement information of the node injection power and line flow in the local area of the target power grid, but does not know the topology structure and line parameter information of the area.
[0015] Preferably, in step S4, if the problem has a solution, the constructed attack vector is injected into the corresponding measuring instrument, which will trigger the residual contamination phenomenon and eventually lead to the deviation of the state estimation result; if the problem has no solution, the attacker chooses to lie low and jump to the first step to continue collecting the measurement data of the attack area at the next moment to find a suitable attack opportunity.
[0016] As a preference, in step S2, taking the boundary bus i as an example, according to the KCL and KVL laws, the node injects active power p i and reactive power q i The equivalent boundary bus i injection power measurement can be calculated by formulas (3)-(4) by satisfying formulas (1)-(2) respectively.
[0017]
[0018] Where S i,A and S i,N They represent the node sets of the attack area and non-attack area connected to the boundary bus i respectively; P ij and Q ij They represent the active power flow and reactive power flow of the line starting from busbar i and ending at busbar j respectively; i,B and q i,B They represent the active power and reactive power injected into the node of the equivalent boundary bus i respectively;
[0019] For the equivalent attack area, KCL can be expressed as
[0020]
[0021] That is, the node injection power is equal to the algebraic sum of the line power flows in all areas connected to the bus;
[0022] Similarly, KVL can be expressed as:
[0023]
[0024] Where, δ l Indicates the phase angle difference between the beginning and end of line l, S k represents a set of lines on path k; formula (7) indicates that the phase angle difference between any two buses in the area is equal to the sum of the phase angle differences at the beginning and end of all lines along path k, and there are always at least two paths between any two buses.
[0025] Preferably, in step S4, the active power flow constraint learned by the NN model is used to replace the following power flow equation, and a single-layer mixed integer linear programming model is constructed;
[0026] ΔP=-S@V@ΔD
[0027] P=S·U·GS·V·(D+ΔD)
[0028] Where D and ΔD represent the bus load vector and the FDIA vector in load measurement, respectively; S represents the transfer factor matrix of the power grid, U and V represent the bus-unit correlation matrix and the bus-load correlation matrix, respectively.
[0029] The present invention first makes the measurement data in the attack area conform to Kirchhoff's Current Law (KCL) and Kirchhoff's Voltage Law (KVL) by equivalent boundary bus measurement; then uses a shallow neural network (NN) with a through layer to learn the nonlinear power flow constraints of the attack area and linearizes them using the large M method; finally, solves a single-layer mixed integer linear programming problem. When the problem has no solution, the attacker lurks and maintains the channel concealment; when the problem has a solution, the attacker injects the attack vector into the corresponding measuring instrument, triggering residual pollution and causing the line power flow estimation result to exceed the limit. Using the technical solution of the present invention, it is only necessary to master the active power measurement information of all nodes and lines in the attack area to construct the attack vector, triggering residual pollution and causing the estimated result of the line active power flow in the state estimation program to exceed the limit, thereby affecting subsequent scheduling decisions and threatening the safe, stable and economical operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0031] Figure 1 This is a flow chart of the FDIA construction method based on data model fusion drive under incomplete information in an embodiment of the present invention.
[0032] Figure 2 Schematic diagram of active power flow constraint learning model;
[0033] Figure 3 It is a single-layer NN model with a pass-through layer. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0035] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] Example 1:
[0037] like Figure 1 As shown, an embodiment of the present invention provides an FDIA construction method based on data model fusion drive under incomplete information, including:
[0038] Step S1: Assume that the attacker only has access to the measurement information of node injection power and line flow in a local area of the target power grid, but does not know the topology and line parameter information of the area. The attacker continuously collects measurement data in the attack area and forms a measurement history database;
[0039] Step S2: The attacker performs a merged equivalent measurement on the busbars located at the boundary of the local area, so that the measurement in the local area satisfies Kirchhoff's laws (i.e., KCL and KVL);
[0040] Step S3: Extract the active power measurement data equivalent to the attack area measurement, use the node injected active power as input, and the line active power flow as output, train a single-layer neural network with a through layer, fit the nonlinear mapping relationship of the power flow equation in the area, and linearize the mapping relationship using the big M method;
[0041] Step S4: Create a single-level mixed-integer linear programming problem with line crossing as the attack target and solve it using the CPLEX solver. If the problem has a solution, the constructed attack vector is injected into the corresponding measurement instrument, triggering residual contamination and ultimately causing a shift in the state estimation results. If the problem has no solution, the attacker chooses to lie low and jump to the first step to continue collecting measurement data in the attack area at the next moment, looking for a suitable attack opportunity.
[0042] As an implementation method of the embodiment of the present invention, in step S2, the target power grid is divided into two parts: the attack area and the non-attack area. Figure 1 The solid line border is shown in the figure, and the thinner dotted line represents the transmission line connected to the boundary bus in the non-attack area. Taking the boundary bus i as an example, according to the KCL and KVL laws, the node injected active power p i and reactive power q i The equivalent boundary bus i injection power measurement can be calculated by formulas (3)-(4) by satisfying formulas (1)-(2) respectively.
[0043]
[0044] Where S i,A and S i,N They represent the node sets of the attack area and non-attack area connected to the boundary bus i respectively; P ij and Q ij They represent the active power flow and reactive power flow of the line starting from busbar i and ending at busbar j respectively; i,B and q i,B They represent the active power and reactive power injected into the node of the equivalent boundary bus i respectively.
[0045] For the equivalent attack area, KCL can be expressed as
[0046]
[0047] That is, the node injection power is equal to the algebraic sum of the line currents in all areas connected to the bus.
[0048] Similarly, KVL can be expressed as
[19]
[0049]
[0050] Where, δ l Indicates the phase angle difference between the beginning and end of line l, S k represents a set of lines on path k. Formula (7) indicates that the phase angle difference between any two busbars in the region is equal to the sum of the phase angle differences at the beginning and end of all lines along path k. Since the transmission network is usually a ring network structure, there are always at least two paths between any two busbars.
[0051] As one implementation of an embodiment of the present invention, in step S3, in the AC power flow model, due to line power loss, a nonlinear response relationship exists between the grid node injection power vector p and the line active power flow vector P. In operational optimization, this response relationship is essentially a constraint, which can be learned using a NN model with strong fitting performance in the field of machine learning.
[0052] Considering that the output P is a continuous variable, learning this constraint is essentially a regression problem. Therefore, a neural network based on the rectified linear unit (ReLU) activation function is constructed as follows: Figure 2 The model includes an input layer (active power injected by the node), a hidden layer (neurons based on ReLU) and an output layer (active power flow of the line). Its mathematical model is shown in formulas (8)-(11):
[0053]
[0054] Where t∈Γ represents the t-th time section (Γ={1,2,…,T}), k∈K represents the k-th hidden layer (K={1,2,…,K}), and x t Represents the input of the NN model, which is the active power p injected by the node at time t t , represents the output of the kth layer of the NN model, W k and b k Represent the weight matrix and bias of the kth layer of the NN model, represents the line active power flow at time t output by the NN model, formulas (9) and (10) represent the linear mapping and ReLU activation function of the kth hidden layer, respectively, and formulas (8) and (11) represent the input layer and output layer of the NN model, respectively.
[0055] Due to the strong nonlinear characteristics of neural networks, the data-driven power flow constraint model is difficult to be directly used to solve the convex optimization problem of attack vectors. Therefore, the constraint model needs to be linearized. Considering that the nonlinear part of the NN model is the ReLU activation function of the hidden layer neural unit, and the core of the ReLU function is to set the negative value of the input to 0 and output it, while the positive value remains unchanged, the big M method can be used to convert the ReLU into a mixed integer linear constraint as shown in formula (12):
[0056]
[0057] Where n k ∈Ν k represents the number of neurons in the kth hidden layer (N k ={1,2,…,N k}), and represents the auxiliary variable introduced in the kth layer, and M is a sufficiently large constant.
[0058] As can be seen from formula (12), the more hidden layers there are, the more auxiliary variables are introduced, and the lower the efficiency of solving the corresponding optimization problem; on the other hand, too few hidden layers (for example, only one) will lead to insufficient fitting accuracy. Therefore, how to balance computational efficiency and fitting accuracy and determine the number of hidden layers is a key challenge.
[0059] Improve the above model into a single-layer NN model with a pass-through layer, such as Figure 3 As shown, the improved mathematical model is shown in formulas (13)-(16)
[0060]
[0061]
[0062] Where x t 、z t 、h t and Same meaning as above, W (z) and b (z) Represents the weight matrix and bias of the first layer of the NN model, W (x) and b (x) Represents the weight matrix and bias of the NN model's pass-through layer.
[0063] Based on the above model, the loss function is used to quantify the distance between the neural network output value and the true value, and the neural network parameters W are optimized. (z) 、b (z) 、W (x) and b (x) , so that the predicted value is close to the true value. According to the nature of the regression problem, the mean square error (MSE) is used to construct the loss function
[0064]
[0065] Where NL is the number of lines in the attack area, and P t,i They represent the predicted value and true value of the active power flow of the i-th line at time t respectively.
[0066] Based on the neural network structure and loss function, to minimize the loss function L mse As the goal, the adaptive moment estimation (Adam) algorithm is used to optimize the neural network parameters and improve the prediction accuracy.
[0067] Since the nonlinear part of the above improved model is still the ReLU activation function, the linearization method is the same as formula (12).
[0068] It's worth noting that while the aforementioned model accurately fits power flow constraints between historical measurement data under normal operating conditions in the attack area and exhibits a certain degree of generalization, given the attacker's lack of measurement samples for abnormal operating conditions, the line power flow predictions for individual busbar load abrupt changes (equivalent to the measurement distribution after injecting the FDIA vector) may still be inaccurate, resulting in the measurement data for that section not fully conforming to the physical laws of KCL and KVL. If the FDIA is constructed based on this, it will be unable to avoid bad data detection, triggering the residual error search phase.
[0069] As an implementation method of the embodiment of the present invention, in step S4, the FDIA attack model with line crossing as the target can be represented as a two-layer mixed integer linear programming problem as follows:
[0070] Upper objective function:
[0071]
[0072] Upper level constraints:
[0073]
[0074] -τD≤ΔD≤τD (21)
[0075] 1 T ΔD=0 (22)
[0076] ΔP=-S·V·ΔD (23)
[0077]
[0078] Lower layer objective function:
[0079]
[0080] Lower level constraints:
[0081] 1 T G=1 T (D+ΔD) (26)
[0082] P=S·U·GS·V·(D+ΔD) (27)
[0083] G min ≤G≤G max (28)
[0084] -P max ≤P≤P max (29)
[0085] Where ΔD d and δ D,d represents the false data in the dth bus measurement and the corresponding 0-1 variable; ΔP l and δ P,l represents the false data in the measurement of the first line and the corresponding 0-1 variable; D and ΔD represent the bus load vector and the FDIA vector in the load measurement, respectively; G, G max and G min Represent the power plant output vector, the maximum output vector and the minimum output vector of the power plant respectively; P max and They represent the active power flow limit vector of the line and the active power flow limit of the kth line respectively; S represents the transfer factor matrix of the power grid, U and V represent the bus-unit correlation matrix and the bus-load correlation matrix respectively; τ represents the maximum percentage of load measurement change, usually 50%; f k represents the absolute value of the k-th line power flow; λ represents the line crossing threshold, which is usually 1.2; c g represents the power generation cost vector.
[0086] The above attack model requires a complete grid topology and line parameters (i.e., S, U, and V) as well as an optimal dispatching strategy (i.e., c g ), however, this information is usually stored in a dispatch center with a high level of security protection and is difficult to obtain, so its application in practical scenarios is limited.
[0087] This embodiment of the present invention replaces the power flow equations in the aforementioned model formulas (23) and (27) with active power flow constraints learned by the NN model and constructs a single-layer mixed integer linear programming model. This allows an attacker to initiate a FDIA that causes a line to exceed the limit even if they only have access to measurement information in the attack area. The line-exceeding FDIA optimization model driven by data-model fusion is as follows.
[0088] Objective function:
[0089]
[0090] Constraints:
[0091] Formula (24)-(27), Formula (29)
[0092]
[0093] Where Δa represents the complete FDIA vector. The other variables have been explained above and will not be repeated here.
[0094] In the above model, formulas (19) and (20) are logical constraints. By introducing additional 0-1 variables, they can be transformed into mixed integer linear forms as shown in formulas (32) and (33).
[0095]
[0096]
[0097] Where ε is a sufficiently small positive number, δ D+,d ,δ D-,d and δ D,d All are 0-1 variables. When ΔD d ≠0, δ D,d =1; when ΔD d >0, δ D+,d =1; when ΔD d <0, δ D-,d =1; when ΔD d =0, δ D,d =0. δ P+,l ,δ P-,l and δ P,l The meaning is similar and will not be repeated here. At this point, the problem can be easily solved using commercial solvers such as CPLEX.
[0098] It should be noted that since the majority of the historical measurement data collected by the attacker is normal operating condition data, the above optimization model may have no solution for certain load sections. However, this does not mean that the attack has failed, but rather reminds the attacker to lie low and wait for the opportunity. When the solution is successful, the attacker injects the constructed attack vector into the corresponding measurement instrument, which will trigger residual contamination and ultimately lead to a deviation in the state estimation results, causing the line to exceed the limit.
[0099] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A FDIA construction method based on data model fusion drive under incomplete information, characterized by: include: Step S1: Collect measurement data of the attack area and form a measurement history database; Step S2: Based on the measurement history database, the measurement values of the busbars located at the boundary of the local area are merged and equivalent, so that the measurement in the local area satisfies Kirchhoff's law; Step S3: Extract the active power measurement data equivalent to the attack area measurement, use the node injected active power as input, and the line active power flow as output, train a single-layer neural network with a through layer, fit the nonlinear mapping relationship of the power flow equation in the area, and linearize the mapping relationship using the big M method; Step S4: construct a single-level mixed integer linear programming problem with line crossing as the attack target, and use the CPLEX solver to solve it.
2. The FDIA construction method based on data model fusion drive under incomplete information according to claim 1 is characterized in that: The attacker can only obtain the measurement information of node injection power and line flow in the local area of the target power grid, but does not know the topology and line parameter information of the area.
3. The FDIA construction method based on data model fusion drive under incomplete information according to claim 1 is characterized in that: In step S4, if the problem has a solution, the constructed attack vector is injected into the corresponding measuring instrument, which will trigger the residual contamination phenomenon and eventually lead to the deviation of the state estimation result; If there is no solution to this problem, the attacker will choose to lie low and jump to the first step to continue collecting measurement data of the attack area at the next moment, looking for a suitable time to attack.
4. The FDIA construction method based on data model fusion drive under incomplete information according to claim 3 is characterized in that: In step S2, taking the boundary bus i as an example, according to the KCL and KVL laws, the node injects active power p i and reactive power q i The equivalent boundary bus i injection power measurement can be calculated by formulas (3)-(4) by satisfying formulas (1)-(2) respectively. Where S i,A and S i,N They represent the node sets of the attack area and non-attack area connected to the boundary bus i respectively; P ij and Q ij They represent the active power flow and reactive power flow of the line starting from busbar i and ending at busbar j respectively; i,B and q i,B They represent the active power and reactive power injected into the node of the equivalent boundary bus i respectively; For the equivalent attack area, KCL can be expressed as That is, the node injection power is equal to the algebraic sum of the line power flows in all areas connected to the bus; Similarly, KVL can be expressed as: Where, δ l It represents the phase angle difference between the two ends of line l. S k represents a set of lines on path k; formula (7) indicates that the phase angle difference between any two buses in the area is equal to the sum of the phase angle differences at the beginning and end of all lines along path k, and there are always at least two paths between any two buses.
5. The FDIA construction method based on data model fusion drive under incomplete information according to claim 4 is characterized in that: In step S4, the active power flow constraint learned by the NN model is used to replace the following power flow equation, and a single-layer mixed integer linear programming model is constructed; ΔP=-S·V@ΔD P=S@U@GS@V@(D+ΔD) Where D and ΔD represent the bus load vector and the FDIA vector in load measurement, respectively; S represents the transfer factor matrix of the power grid; U and V represent the bus-unit correlation matrix and the bus-load correlation matrix, respectively.