Power system data integrity attack detection method and system, computer equipment and storage medium
The GCN-DQN model addresses the challenge of detecting complex network relationships in electric power systems by using reinforcement learning to identify and counteract false data injection attacks, enhancing system security and stability.
Patent Information
- Application Number
- CN202510261146.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art cannot effectively capture the complex relationship between nodes and nodes and nodes in the power network, making it difficult to detect false data injection attacks, and the traditional methods have poor detection effects in smart grids.
Combining graph convolution network (GCN) and deep Q network (DQN), a GCN-DQN model is constructed to extract feature and output decision-making actions on the power system graph structure to realize detection and defense against false data injection attacks.
It realizes efficient detection and response to various types of attacks, dynamically identify data integrity attacks in smart grids, improves the security and stability of the power system, and reduces the false alarm rate.
Smart Images

Figure CN120320971A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power system network security, and specifically relates to a method, a system, a computer device, and a storage medium for detecting data integrity attacks in a power system. Background Art
[0002] With the wide application of smart grids, the power system is gradually developing towards digitalization and automation. By deeply integrating information technology with the power network, smart grids have improved the operation efficiency, reliability, and flexibility of the power system. However, with the increasing degree of informatization of the power system, the risk of cyber attacks has also increased. In particular, false data injection attacks (FDIA) are a typical form of cyber attack. Attackers tamper with the data of sensors or measurement devices, causing the control center to receive incorrect system status information, thereby affecting power dispatching and control decisions. Such attacks are not only difficult to detect by traditional intrusion detection systems but may also lead to serious economic losses and power system instability.
[0003] Traditional network security protection means mainly focus on rule-based detection methods. However, in the face of the complexity of the power system and large-scale data interaction, the detection accuracy of such methods is limited, and it is difficult to cope with different types of attacks. Existing machine learning and deep learning methods have also been applied to power system security, but they usually cannot fully capture the complex relationships between nodes and nodes, and between nodes and edges in the power network. Especially when facing high-dimensional and heterogeneous data in smart grids, the detection effect is poor. In addition, how to perform real-time and efficient attack detection under the complex topology of the power system remains an urgent problem to be solved. Summary of the Invention
[0004] Object of the Invention: To solve the problem that existing network security protection means cannot fully capture the complex relationships between nodes and nodes, and between nodes and edges in the power network, and to detect and respond to various types of attacks such as false data injection attacks, the present invention combines a graph convolutional network (GCN) and a deep Q network (DQN) to propose a method, a system, a computer device, and a storage medium for detecting data integrity attacks in a power system, solves the limitations of existing power system attack detection methods, and constructs an efficient method for detecting and defending data integrity attacks.
[0005] Technical Solution: A method for detecting data integrity attacks in a power system includes the following steps:
[0006] Collect grid information of the power system;
[0007] Construct a power system graph structure according to the collected grid information of the power system;
[0008] Input the constructed power system graph structure into the trained GCN-DQN model to output decision actions, where the decision actions include: keeping the current state of the power system unchanged when no attack is detected and taking preset defense measures when an attack is detected;
[0009] The GCN-DQN model includes a graph convolutional network model and a deep Q-learning network model. The graph convolutional network model is used to extract features from the power system graph structure to generate feature vectors. The deep Q-learning network model is used to output decision actions based on the feature vectors generated by the graph convolutional network model.
[0010] Further, the graph convolutional network model is used to extract features from the power system graph structure to generate feature vectors, expressed as:
[0011]
[0012] In the formula, ρ represents the generated feature vector, W i represents the weight parameter in the convolution operation, b i represents the bias term, ReLU(·) represents the activation function, represents the node features in the power system graph structure, represents the re-normalized adjacency matrix, and the adjacency matrix represents the connection relationship between node i and node j in the power system graph structure;
[0013] Among them, In the formula, Vm represents the voltage amplitude of the node, Va represents the voltage phase angle of the node, P i and Q i respectively represent the active power and reactive power outputs of node i;
[0014] Among them, the adjacency matrix is expressed as:
[0015]
[0016] In the formula, A i,j represents the connection relationship between node i and node j in the power system graph structure, L is the set of power connection lines, and f[l] represents the impedance value of the power connection line.
[0017] Further, the trained GCN-DQN model is obtained according to the following steps:
[0018] At time step t, based on the parameters of the current graph convolutional network model, the graph convolutional network model extracts features from the power system graph structure to generate a feature vector ρ(t);
[0019] Based on the parameters of the current deep Q-learning network model, the generated feature vector ρ(t) is used as the input state s of the deep Q-learning network model t ; Based on the input state s t , an action a is randomly selected t ;
[0020] By executing the action a t , the cumulative reward R is obtained t ;
[0021] According to the obtained cumulative reward R t , update the parameters of the graph convolutional network model and the parameters of the deep Q-learning network model; and generate a new feature vector ρ(t+1) at the next time step;
[0022] Only when the time step t = T, where T is the total number of time steps, the trained GCN-DQN model is obtained
[0023] Furthermore, the obtaining of the cumulative reward R by executing the action a t is expressed as: t
[0024]
[0025] where R t is the cumulative reward obtained by taking the action a from the input state s t , γ t is the discount factor, T is the total number of time steps, and r(s t ,a t ) is the immediate reward obtained by taking the action a in the input state s t , expressed as: t t
[0026]
[0027] Furthermore, the updating of the parameters of the graph convolutional network model and the parameters of the deep Q-learning network model according to the obtained cumulative reward R t includes:
[0028] Update the parameters of the graph convolutional network model every target number of steps;
[0029] Update the parameters of the deep Q-learning network model by minimizing the following loss function:
[0030] L(θ Q ) = E[(y t - Q(s t ,a t |θ Q )) 2
[0031] In the formula, L(θ Q ) represents the loss function, E(·) represents the expected value operator, θ Q represents the parameters of the deep Q-learning network model, y t = r(s t , a t ) + γQ(s t+1 , π(s t+1 )|θ Q ) represents the expected return, γ represents the discount factor, π(s t+1 ) represents the detection strategy at time step t + 1 for the input state s t+1 ; Q(s t , a t |θ Q ) represents the Q value at time step t, and the expression is:
[0032] Q(s t , a t ) = E[R t |s t , a t .
[0033] The present invention discloses a power system data integrity attack detection system, including:
[0034] An information collection module for collecting power grid information of the power system;
[0035] A power system graph structure construction module for constructing a power system graph structure according to the collected power grid information of the power system;
[0036] A GCN-DQN model module for inputting the constructed power system graph structure into a trained GCN-DQN model and outputting a decision action, where the decision action includes: keeping the current state of the power system unchanged when no attack is detected and taking a preset defense measure when an attack is detected;
[0037] The GCN-DQN model includes a graph convolutional network model and a deep Q-learning network model. The graph convolutional network model is used for feature extraction of the power system graph structure to generate a feature vector; the deep Q-learning network model is used for outputting a decision action according to the feature vector generated by the graph convolutional network model.
[0038] Further, the graph convolutional network model performs feature extraction on the power system graph structure to generate a feature vector, which is expressed as:
[0039]
[0040] In the formula, ρ represents the generated feature vector, Wi represents the weight parameter in the convolution operation, b i represents the bias term, and ReLU(·) represents the activation function, represents the node features in the power system graph structure, represents the re-normalized adjacency matrix, and the adjacency matrix represents the connection relationship between node i and node j in the power system graph structure;
[0041] Among them, In the formula, Vm represents the voltage amplitude of the node, Va represents the voltage phase angle of the node, P i and Q i respectively represent the active power and reactive power outputs of node i;
[0042] Among them, the adjacency matrix is expressed as:
[0043]
[0044] In the formula, A i,j represents the connection relationship between node i and node j in the power system graph structure, L is the set of power connection lines, and f[l] represents the impedance value of the power connection line.
[0045] Furthermore, the trained GCN-DQN model is obtained according to the following steps:
[0046] At time step t, based on the parameters of the current graph convolutional network model, the graph convolutional network model extracts features from the power system graph structure to generate a feature vector ρ(t);
[0047] Based on the parameters of the current deep Q-learning network model, the generated feature vector ρ(t) is used as the input state s of the deep Q-learning network model t ; Based on the input state s t , randomly select an action a t ;
[0048] By executing the action a t , obtain the cumulative reward R t ;
[0049] According to the obtained cumulative reward R t , update the parameters of the graph convolutional network model and the parameters of the deep Q-learning network model; and generate a new feature vector ρ(t + 1) at the next time step;
[0050] Only when the time step t = T, where T is the total number of time steps, the trained GCN-DQN model is obtained.
[0051] Furthermore, by executing the action a t , obtain the cumulative reward R t, expressed as:
[0052]
[0053] Wherein, R t is the cumulative reward obtained after taking action a t from the input state s t , γ t is the discount factor, T is the total number of time steps, and r(s t , a t ) is the immediate reward obtained by taking action a t in the input state s t , expressed as:
[0054]
[0055] Furthermore, according to the obtained cumulative reward R t , update the parameters of the graph convolutional network model and the parameters of the deep Q - learning network model, including:
[0056] Update the parameters of the graph convolutional network model every target number of steps;
[0057] Update the parameters of the deep Q - learning network model by minimizing the following loss function:
[0058] L(θ Q ) = E[(y t - Q(s t , a t |θ Q )) 2
[0059] Wherein, L(θ Q ) represents the loss function, E(·) represents the expected value operator, θ Q represents the parameters of the deep Q - learning network model, y t = r(s t , a t ) + γQ(s t+1 , π(s t+1 )|θ Q ) represents the expected return, γ represents the discount factor, π(s t+1 ) represents the detection policy at the input state s t+1 at time step t + 1; Q(s t , a t |θ Q ) represents the Q - value at time step t, and the expression is:
[0060] Q(s t , a t ) = E[R t |st , a t .
[0061] The present invention discloses a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of a method for detecting power system data integrity attacks disclosed above are implemented.
[0062] The present invention discloses a storage medium storing a power system data integrity attack detection program. When the power system data integrity attack detection program is executed by at least one processor, the steps of a method for detecting power system data integrity attacks disclosed above are implemented.
[0063] Advantageous effects: Compared with the prior art, the present invention has the following advantages:
[0064] (1) In the method of the present invention, the graph convolutional network model can extract complex feature information of power system nodes and edges from the graph structure of the power system and generate a high-dimensional system state representation. The deep Q-learning network model, through reinforcement learning, learns and selects the optimal detection and defense strategies according to the features extracted by the graph convolutional network model; when detecting false data injection attacks or other anomalies, the deep Q-learning network model selects corresponding defense actions (such as isolating the attacked node or adjusting grid operations), and through state-action interaction learning, maximizes the reward for detecting attacks, thereby realizing the dynamic adjustment of defense strategies and realizing the detection and response to various attack types such as false data injection attacks.
[0065] (2) Aiming at the characteristics of complex load changes and large data volume in the smart grid, the method of the present invention can maintain high detection ability under different load conditions and attack sparsities through training the model.
[0066] (3) The method of the present invention can dynamically identify various data integrity attack types in the smart grid, automatically adjust power dispatching and defense strategies, and ensure system stability and security; experimental results show that compared with other general detection methods, this method can achieve high-precision and low false-alarm-rate attack detection; therefore, the present invention not only improves the security of the smart grid, but also provides a new technical idea for the network protection of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a flowchart of a method for detecting power system data integrity attacks proposed by the present invention;
[0068] Figure 2 is a schematic diagram of the training process of the GCN-DQN model proposed by the present invention;
[0069] Figure 3 This is the topology diagram of the IEEE 39-node system of the present invention;
[0070] Figure 4 This is the comparison chart of the false alarm rates of each method of the present invention. Detailed implementation manners
[0071] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the power system data integrity attack detection method, system, computer device, and storage medium proposed by the present invention in combination with the accompanying drawings and embodiments in the present invention.
[0072] Embodiment 1:
[0073] This embodiment proposes a power system data integrity attack detection method based on GCN-DQN, which detects power system data integrity attacks by combining a graph convolutional network (hereinafter referred to as GCN) and a deep Q-learning network (hereinafter referred to as DQN). As Figure 1 shown, it mainly includes the following steps:
[0074] Step 1: Power grid information collection; in practical applications, there are often attack vectors in the collected power grid information;
[0075] Step 2: Create a power system graph structure according to the collected power grid information;
[0076] Step 3: Input the created power system graph structure into the trained GCN-DQN model to output a decision action, that is, when detecting a false data injection attack or other anomalies, select corresponding defense actions (such as isolating the attacked node or adjusting the power grid operation).
[0077] In the GCN-DQN model, GCN extracts complex features of nodes and edges from the power system graph structure; the specific feature extraction process includes:
[0078] Construct the adjacency matrix A of the power system:
[0079]
[0080] In the formula, A i,j represents the element in the i-th row and j-th column of the adjacency matrix, indicating the connection relationship between node i and node j in the power system. L is the set of power connection lines. f[l] represents the impedance value of the power connection line.
[0081] Define the node feature matrix as:
[0082]
[0083] Where, Vm represents the voltage amplitude (per unit value) of each node. Va represents the voltage phase angle of each node (in degrees). P i and Q i respectively represent the active power and reactive power outputs of the i-th node (in MW and MVAr). The GCN processes the input features through multiple convolutional operations to generate a feature vector, and the formula is as follows:
[0084]
[0085] Where, ρ represents the feature vector generated after being processed by the GCN, which contains voltage amplitude, phase angle, and power information, representing the current state of the power system. is the re-normalized adjacency matrix. W i is the weight matrix, representing the weight parameters in the convolutional operation. b i is the bias term, used to adjust the output features. ReLU(·) is the activation function, responsible for introducing non-linearity.
[0086] In the GCN-DQN model, after receiving the feature vector ρ generated by the GCN, the DQN makes decisions in the following way:
[0087] Define the input state s t of the DQN, which is the feature vector ρ(t) output by the GCN;
[0088] Define the decision-making action a t of the DQN as:
[0089] a t =[a c ,a s
[0090] Where, a t represents the action selected by the DQN at time step t, a c represents keeping the state unchanged when no attack is detected, and a s represents the defense measures taken when an attack is detected, such as isolating the attacked node, reallocating the load, etc.
[0091] The goal of the DQN is to optimize the policy through reinforcement learning to maximize the expected cumulative reward. The cumulative reward is defined as:
[0092]
[0093] Where, R t is the cumulative reward obtained after taking the action a t from the state s t , and r(s t ,a t ) is the reward obtained in the state s t Take action a below t The immediate reward obtained, γ t is the discount factor at time step t, with a value range between [0, 1], used to reduce the impact of future rewards, and T is the total number of time steps.
[0094] Among them, the immediate reward of DQN is designed as:
[0095]
[0096] DQN makes decisions by learning the value of each state-action pair, and its return value function is expressed as:
[0097] Q(s t ,a t )=E[R t |s t ,a t
[0098] In the formula, Q(s t ,a t ) represents the expected return of taking action a in state s t , and E represents the expected value operator. t
[0099] In this embodiment, the feature vector extracted by GCN provides real-time input for DQN, enabling it to identify potential data integrity attack patterns in the power system. For example, an attacker tampers with sensor data, affecting the decision-making of the power system. GCN extracts features such as voltage and power of nodes and their neighbors, and DQN selects the optimal detection action based on these features, such as deploying redundant measurement units. For example, by overloading communication links or the power network, data transmission or power dispatching is hindered. GCN monitors the traffic information of communication links by extracting edge features, and DQN takes response measures according to the abnormal features of communication links, such as isolating the link or reallocating communication resources.
[0100] In this embodiment, through the reinforcement training of the GCN-DQN model, the Q value is continuously updated, and DQN can learn the optimal policy π(s). π(s) is to select action a in state s, and the optimal action is selected at each time step to deal with attacks. As Figure 2 shown, the specific training steps are:
[0101] (1) Input the power system graph structure G=(N, E), which consists of the node set N and the line set E in the power system. N includes power generation nodes, load nodes, and substation nodes, and E includes AC lines and DC lines. The state s t =x i ,…,x N represents the voltage amplitude and phase angle of the bus. Initialize the GCN and deep Q learning network model;
[0102] (2) Train in the order of increasing time steps, including: The steps include:
[0103] 1) Initialize the time step t to 0;
[0104] 2) Obtain the adjacency matrix A and the node feature matrix
[0105] 3) Output the feature vector ρ(t) through the GCN as the input state s of the DQN t ;
[0106] 4) Generate a random policy π t , select an action a t ;
[0107] 5) Obtain the cumulative reward R(t) by executing the action;
[0108] 6) Use the action a t to train the DQN, minimize the following loss function, learn the optimal policy and update the DQN network parameters θ Q ;
[0109] L(θ Q ) = E[(y t - Q(s t , a t | θ Q )) 2
[0110] In the formula, L(θ Q ) represents the loss function, which is used to measure the gap between the predicted value and the true value. y t = r(s t , a t ) + γQ(s t+1 , a(s t+1 ) | θ Q ) is the expected return. Q(s t , a t | θ Q ) is the return value function Q(s Q under the condition of the DQN network parameters θ t , a t ).
[0111] 7) Update the GCN parameter W - ← W;
[0112] 8) The time step t = t + 1, repeat steps 2) - 7) until t = T, and output the GCN parameter W G and the DQN parameter θ Q .
[0113] Now, an experimental design is carried out for the power system data integrity attack detection method based on GCN-DQN proposed in this embodiment, and its effectiveness is verified. The data information of the standard IEEE 39-bus system is used, and the system topology is as Figure 3 shown. Power flow measurement data is generated through an AC power flow model. The voltage phase angle of each bus is used as the state variable x, while the active power flow of each branch and the active power injection of each bus are used as the measurement values z.
[0114] Due to the large amount of load data, it is difficult to directly calculate the load change range. Therefore, AC state estimation is used to generate normal load data. The collected data is first standardized through the min-max normalization technique. After the normalized active load data is calculated by the ratio of the original load data of the bus and then multiplied by the generator power.
[0115] To construct a false data injection attack, the ratio changes of active and reactive power data in the standard system are simulated. The voltages of different buses are obtained by running the AC power flow program, and Gaussian noise with a standard deviation of 1% is added. The modified measurement values include active and reactive power injection values, and the affected active power measurement values affect the target state variable. It is assumed that a -10% error is injected into the state variable θ2. The injection error c in the state vector is expressed as:
[0116] c = [0, -0.1θ2, 0, …, 0]
[0117] The state vector after the injection attack is:
[0118]
[0119] In the formula, is the vector of the original state estimation, x bad is the state estimation vector after the attack.
[0120] And the modified measurement values for the training process are generated:
[0121] z bad = h(x bad ) + e
[0122] The experiment implements the GCN-DQN framework using the OpenAI Gym library in the Python 3.6 environment. The DQN algorithm is trained in the way of a batch size of 32 and 2000 steps. Table 1 summarizes the simulation parameters.
[0123] Table 1 Simulation Parameters
[0124] Parameter Value Node embedding dimension 64 Discount factor 0.05 Learning rate 0.001 Batch size 512 Activation function ReLU Number of training epochs 100k Number of evaluation epochs 500 Number of GNN layers 4 Dataset IEEE 39-node data
[0125] The performance of the GCN-DQN framework was verified through the following parameters, and was compared and verified with methods such as Deep Q-Network Decomposition (DQND), State-Action-Reward-State-Action (SARSA), Local Outlier Factor (LOF), Isolation Forest (i-Forest), One-Class Support Vector Machine (One-Class), etc. Among them Figure 4 represents the false positive rate of each method, Table 2 represents the real-time detection efficiency of each method, and Table 3 represents the loss value, accuracy, and recall rate of each method.
[0126] Table 2 Detection time of each method
[0127]
[0128]
[0129] Table 3 Performance indicators of each method
[0130] Performance metric DQND SARSA LOF i-Forest One-Class GCN-DQN Loss value 0.4522 0.4407 0.337 0.41 0.386 0.2115 Accuracy 0.8017 0.8211 0.781 0.825 0.803 0.8571 Recall 0.7622 0.7751 0.7655 0.78 0.7827 0.8011
[0131] From the above indicators, it can be found that the GCN-DQN method shows obvious advantages in various performance indicators. It is superior to other traditional methods in terms of loss value, accuracy, recall rate, and false positive rate. Especially in terms of detection accuracy (85.71%) and recall rate (80.11%), the GCN-DQN method is significantly ahead. At the same time, the false positive rate of GCN-DQN is 0.012, which is much lower than other methods, meaning that it can not only accurately detect attacks, but also effectively reduce false positives and improve system stability.
[0132] In addition, the loss value of GCN-DQN is 0.2115, indicating that the model has a good convergence effect and shows a low error during the training process. The detection time is the shortest. In the previous detection time analysis, the detection time of GCN-DQN is also significantly shorter than other methods, only 0.00634 seconds, demonstrating the high efficiency of this method.
[0133] In summary, the GCN-DQN method based on the combination of GCN and DQN shows extremely high accuracy, efficiency, and stability in data integrity attack detection. It provides reliable and fast security guarantee in large-scale power grid systems and has important application prospects for the safe operation of smart grids.
[0134] This embodiment demonstrates the effectiveness of power system data integrity attack detection based on GCN-DQN. By constructing false data injection attacks and using GCN to extract features, DQN performs reinforcement learning to select the optimal actions. This framework can handle complex smart grid attacks and shows excellent detection performance in experiments.
[0135] Example 2:
[0136] Based on Example 1, this example proposes a power system data integrity attack detection system, including:
[0137] An information collection module for collecting power grid information of the power system;
[0138] A power system graph structure construction module for constructing a power system graph structure according to the collected power grid information of the power system;
[0139] A GCN-DQN model module for inputting the constructed power system graph structure into a trained GCN-DQN model and outputting a decision action, where the decision action includes: keeping the state unchanged when no attack is detected and defense measures taken when an attack is detected;
[0140] In the trained GCN-DQN model, a graph convolutional network model extracts features from the power system graph structure to generate feature vectors; a deep Q-learning network model outputs a decision action based on the generated feature vectors.
[0141] Example 3:
[0142] This example discloses a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps disclosed in any of the above examples are implemented.
[0143] Example 4:
[0144] This example discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps disclosed in any of the above examples are implemented.
[0145] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0146] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0147] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for detecting data integrity attacks in a power system, characterized in that: It includes the following steps: Collect power grid information of the power system; Construct a power system graph structure based on the collected power grid information of the power system; Input the constructed power system graph structure into the trained GCN-DQN model to output decision-making actions, where the decision-making actions include: keeping the current state of the power system unchanged when no attack is detected and taking preset defense measures when an attack is detected; The GCN-DQN model includes a graph convolutional network model and a deep Q-learning network model. The graph convolutional network model is used to extract features from the power system graph structure to generate feature vectors. The deep Q-learning network model is used to output decision-making actions based on the feature vectors generated by the graph convolutional network model.
2. The method for detecting data integrity attacks in a power system according to claim 1, wherein: The graph convolutional network model is used to extract features from the power system graph structure to generate feature vectors, expressed as: where ρ represents the generated feature vector, W i represents the weight parameter in the convolution operation, b i represents the bias term, ReLU(·) represents the activation function, represents the node features in the power system graph structure, represents the re-normalized adjacency matrix, and the adjacency matrix represents the connection relationship between node i and node j in the power system graph structure; Among them, where Vm represents the voltage magnitude of the node, Va represents the voltage phase angle of the node, P i and Q i respectively represent the active power and reactive power outputs of node i; Among them, the adjacency matrix is expressed as: Where, A i,j represents the connection relationship between node i and node j in the power system diagram structure, L is the set of power connection lines, and f[l] represents the impedance value of the power connection line.
3. The method for detecting data integrity attacks in a power system according to claim 2, characterized in that: The trained GCN-DQN model is obtained according to the following steps: At time step t, based on the parameters of the current graph convolutional network model, the graph convolutional network model extracts features from the power system graph structure to generate a feature vector ρ(t); Based on the parameters of the current deep Q - learning network model, the generated feature vector ρ(t) is used as the input state s of the deep Q - learning network model t ; Based on the input state s t , an action a is randomly selected t ; By performing action a t , the cumulative reward R is obtained t ; According to the obtained cumulative reward R t , update the parameters of the graph convolutional network model and the parameters of the deep Q-learning network model; and generate a new feature vector ρ(t + 1) at the next time step; Only when the time step t = T, where T is the total number of time steps, the trained GCN-DQN model is obtained.
4. A method for detecting data integrity attacks in a power system according to claim 3, characterized in that: By performing action a t , the cumulative reward R is obtained t , which is expressed as: where R t is the cumulative reward obtained after taking action a t from the input state s t , θ t is the discount factor, T is the total number of time steps, r(s t , a t ) is the immediate reward obtained by taking action a t in the input state s t , expressed as:
5. The method for detecting data integrity attacks in a power system according to claim 4, characterized in that: According to the obtained cumulative reward R t , update the parameters of the graph convolutional network model and the parameters of the deep Q-learning network model, including: Update the parameters of the graph convolutional network model every target number of steps; Update the parameters of the deep Q-learning network model by minimizing the following loss function: L(θ Q ) = E[(y t - Q(s t , a t |θ Q )) 2 ) where L(θ Q ) represents the loss function, E(·) represents the expected value operator, θ Q represents the parameters of the deep Q - learning network model, y t = r(s t , a t ) + γQ(s t+1 , π(s t+1 )|θ Q ) represents the expected return, γ represents the discount factor, π(s t+1 ) represents the action selected at time step t + 1 for the input state s t+1 ; Q(s t , a t |θ Q ) represents the Q - value at time step t, and the expression is: Q(s t ,a t ) = E[R t |s t ,a t .
6. A power system data integrity attack detection system, characterized in that: It includes: An information collection module for collecting power grid information of the power system; A power system graph structure construction module for constructing a power system graph structure based on the collected power grid information of the power system; A GCN-DQN model module for inputting the constructed power system graph structure into the trained GCN-DQN model to output decision-making actions, where the decision-making actions include: keeping the current state of the power system unchanged when no attack is detected and taking preset defense measures when an attack is detected; The GCN-DQN model includes a graph convolutional network model and a deep Q-learning network model. The graph convolutional network model is used to extract features from the power system graph structure to generate feature vectors. The deep Q-learning network model is used to output decision-making actions based on the feature vectors generated by the graph convolutional network model.
7. An integrity attack detection system for power system data according to claim 6, characterized in that: The graph convolutional network model is used to extract features from the power system graph structure to generate feature vectors, expressed as: where ρ represents the generated feature vector, W i represents the weight parameter in the convolution operation, b i represents the bias term, ReLU(·) represents the activation function, represents the node features in the power system graph structure, represents the re-normalized adjacency matrix, and the adjacency matrix represents the connection relationship between node i and node j in the power system graph structure; Among them, where Vm represents the voltage amplitude of the node, Va represents the voltage phase angle of the node, and P i and Q i respectively represent the active power and reactive power outputs of node i; Among them, the adjacency matrix is expressed as: Where, A i,j represents the connection relationship between node i and node j in the power system diagram structure, L is the set of power connection lines, and f[l] represents the impedance value of the power connection line.
8. The power system data integrity attack detection system according to claim 7, characterized in that: The trained GCN-DQN model is obtained according to the following steps: At time step t, based on the parameters of the current graph convolutional network model, the graph convolutional network model extracts features from the power system graph structure to generate a feature vector ρ(t); Based on the parameters of the current deep Q-learning network model, the generated feature vector ρ(t) is used as the input state s of the deep Q-learning network model t ; Based on the input state s t , randomly select an action a t ; By performing action a t , the cumulative reward R is obtained t ; According to the obtained cumulative reward R t , update the parameters of the graph convolutional network model and the parameters of the deep Q-learning network model; and generate a new feature vector ρ(t + 1) at the next time step; Only when the time step t = T, where T is the total number of time steps, the trained GCN-DQN model is obtained.
9. The power system data integrity attack detection system according to claim 8, wherein: By performing action a t , the cumulative reward R is obtained t , which is expressed as: where R t is the cumulative reward obtained after taking action a t from the input state s t , γ t is the discount factor, T is the total number of time steps, r(s t , a t ) is the immediate reward obtained by taking action a t in the input state s t , expressed as:
10. A power system data integrity attack detection system according to claim 9, characterized in that: According to the obtained cumulative reward R t , update the parameters of the graph convolutional network model and the parameters of the deep Q-learning network model, including: Update the parameters of the graph convolutional network model every target number of steps; Update the parameters of the deep Q-learning network model by minimizing the following loss function: L(θ Q ) = E[(y t - Q(s t , a t | θ Q )) 2 ) where \(L(\theta\ Q )\) represents the loss function, \(E(\cdot)\) represents the expected value operator, \(\theta\ Q \) represents the parameters of the deep Q - learning network model, \(y\ t = r(s\ t ,a\ t )+\gamma Q(s\ t+1 ,\pi(s\ t+1 )|\theta\ Q )\) represents the expected return, \(\gamma\) represents the discount factor, \(\pi(s\ t+1 )\) represents the action selected at time step \(t + 1\) for the input state \(s\ t+1 \); \(Q(s\ t ,a\ t )|\theta\ Q )\) represents the Q - value at time step \(t\), and the expression is: Q(s t ,a t ) = E[R t |s t ,a t .
11. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a method for detecting power system data integrity attacks according to any one of claims 1 to 5.
12. A storage medium, characterized in that, The storage medium stores a power system data integrity attack detection program, and when the power system data integrity attack detection program is executed by at least one processor, the steps of a power system data integrity attack detection method according to any one of claims 1 to 5 are implemented.