Power distribution network false data injection attack detection method based on vertical federal learning
By applying a vertical federal learning framework in the distribution network, we collaboratively build a false data injection attack detection model, which solves the problem of difficult detection of false data injection attacks in the distribution network, achieves efficient and accurate detection effects, and solves the problem of data sharing.
Patent Information
- Application Number
- CN202510208351.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
False data injection attacks in the distribution network are difficult to detect, and due to data privacy and security considerations, it is difficult for entities to share data, resulting in poor detection results.
Using a detection method based on vertical federated learning, a vertical federated learning framework for split learning methods is built to collaborate on the detection model of false data injection attack in the power distribution network. The framework includes a grid edge computing unit and a server, which deploys grid edge side models and server-side models respectively, and extracts spatial and temporal features using convolutional layers and Bi-LSTM models.
It improves the accuracy and efficiency of detection of false data injection attacks, solves data privacy and security issues, realizes data sharing and collaborative learning among different entities, and is suitable for complex distribution network environments.
Smart Images

Figure CN120074906A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distribution network security, and particularly to a method for detecting false data injection attacks in a distribution network based on vertical federated learning. Background Art
[0002] The integration of information and communication technologies has transformed the distribution network into a cyber-physical system. The increasing popularity of such technologies in the distribution network has brought a new problem, namely the emergence of cyberattacks. These cyberattacks pose a threat to the availability, integrity, or security of the distribution network. In recent years, a large number of cyberattack incidents have occurred in the energy industry. These incidents have highlighted the necessity of comprehensively understanding and correctly addressing the risks of grid cyberattacks. Although the forms and impacts of cyberattacks show considerable diversity, false data injection attacks (FDIA) are a formidable challenge because of their wide-ranging consequences and the difficulty of detection. False data injection attacks (FDIA), as a covert and highly destructive cyberattack means, interfere with the normal operation of the system by tampering with the measurement data in the power system, and may even lead to large-scale power failures and blackout incidents, bringing serious impacts to the social economy and people's lives.
[0003] The detection techniques of FDIA are mainly divided into two categories: model-based and data-driven methods. The former requires accurate system models and parameter information, and small changes or uncertainties in the parameters may lead to inaccurate detection performance. In contrast to model-based algorithms, data-driven algorithms are model-free and do not use system models or system parameters.
[0004] Traditional FDIA detection methods mainly focus on the transmission network level and use technical means such as state estimation, bad data detection, and identification for defense. However, with the continuous expansion of the scale of the distribution network and the continuous improvement of the intelligent level, the problem of FDIA in the distribution network has become increasingly prominent, while the research on FDIA for the distribution network is relatively less. The data in the distribution network is usually owned and managed by multiple different entities, such as power companies, distributed energy providers, users, etc. There is a certain correlation and complementarity between these data. However, due to data privacy and security considerations, these entities are often reluctant to directly share data, resulting in a serious data island phenomenon, which restricts the improvement of FDIA detection effect.
[0005] To address the data sharing issue in the distribution network, some researchers have attempted to adopt federated learning technology. Federated learning is a decentralized machine learning method that allows multiple participants to collaboratively learn by exchanging model parameters or intermediate representations while keeping the data local, thus achieving data sharing and knowledge integration. However, traditional federated learning mainly focuses on the horizontal federated learning scenario, where the data of the participants has the same feature space and different sample spaces. In the distribution network, the data of different entities often has different feature spaces and the same sample space, which belongs to the category of vertical federated learning.
[0006] Vertical federated learning has significant advantages in processing data with different feature spaces. It can split the model and learning tasks, enabling different entities to collaboratively train a global model while protecting their respective data privacy. However, applying vertical federated learning to FDIA detection in the distribution network still faces many challenges. First, the data in the distribution network has spatio-temporal correlations, and how to effectively extract and utilize these features for FDIA detection is a difficult problem. Second, the FDIA in the distribution network may have multiple attack modes and intensities, and how to design a general detection method to deal with different attack scenarios is a challenge. Finally, the data quality and quantity in the distribution network may be limited, and how to improve the generalization ability and robustness of the detection model is an important issue. Summary of the Invention
[0007] The objective of the present invention is to provide a method for detecting false data injection attacks in a distribution network based on vertical federated learning, aiming to improve the security of the distribution network, provide strong guarantees for the reliable operation of the smart grid, and at the same time solve the problems of insufficient data privacy and security, poor accuracy and robustness in different scenario detections, and difficulty in dynamic parameter adjustment.
[0008] To achieve the above objective, the present invention provides a method for detecting false data injection attacks in a distribution network based on vertical federated learning, including the following steps:
[0009] S1. Build a vertical federated learning framework based on the split learning method and collaboratively construct an FDIA detection model in the distribution network;
[0010] The vertical federated learning framework includes a grid edge computing unit and a server;
[0011] The FDIA detection model is split into a grid edge side model and a server side model based on the split learning method, and the grid edge side model and the server side model are respectively assigned to the grid edge computing unit and the server;
[0012] S2. Train the vertical federated learning framework;
[0013] S3. Implement the construction and application of the vertical federated learning framework through training and inference operations;
[0014] S4. Extract spatial features through the grid edge-side model. The grid edge-side model includes two convolutional layers, followed by a pooling layer and a convolutional block attention module after each convolutional layer. The convolutional block attention module includes a channel attention module and a spatial attention module;
[0015] S5. Use the input of the Bi-LSTM model of the server model to extract temporal features. The input includes the concatenated intermediate representation H of the extracted spatial features all of the concatenated intermediate representation H all which is processed by the server model.
[0016] Preferably, each grid edge computing unit holds its own grid edge-side model G p . The grid edge computing unit collects a set of measurement data from local sensors in its respective sub-network and processes it through forward propagation using its own grid edge-side model.
[0017] Preferably, the server has a server-side model G 0 and performs the following operations during training and real-time FDIA detection: The server aggregates the intermediate representations collected from the grid edge computing units; uses the aggregated intermediate representations to continue forward propagation; the server updates its own model and sends the calculated gradients to the split layer to each edge-side model to update the client models.
[0018] Preferably, S2 includes the following steps:
[0019] S21. Initialize the parameters of the server-side model G 0 and the client model G k in each grid edge computing unit; The parameters include the number of epochs E, learning rates η 0 and η p of the grid edge-side model and the server-side model, and the loss function L. At the beginning of each iteration, the participants randomly select a subset B t to reach a consensus, where t represents the iteration number;
[0020] S22. Each grid edge computing unit uses its own grid edge-side model G p and its own local dataset for forward propagation, and then obtains the intermediate representation from each grid edge-side model in the t-th round of communication. The calculation formula is as follows:
[0021]
[0022] where represents the grid edge-side model G of entity p in the t-th round of communicationp Parameter represents the mini-batch data selected by entity p during the t-th round of communication;
[0023] S23. After performing forward propagation in each grid edge computing unit, collect the intermediate representations on the server and make the following connections:
[0024]
[0025] where represents the connected intermediate representation;
[0026] S24. The server continues to perform forward propagation using the server-side model and the connected intermediate representations to obtain the server-side model output, then calculates the loss based on the server-side model output, and selects the loss function as the true value label y of the cross-entropy loss function t , and the formula for the obtained server-side model output is as follows:
[0027]
[0028] where represents the server-side model output and is also the predicted output;
[0029] S25. Update the server-side model parameters, and the calculation process is as follows:
[0030]
[0031] where is the gradient of the loss function with respect to the server-side model during the t-th round of communication, and η 0 is the learning rate of the server-side model;
[0032] S26. After updating the server model, the server continues to update each grid edge-side model in parallel. For each grid edge computing unit, calculate the gradient of the loss function with respect to each intermediate representation, and then send it to its respective grid edge computing unit to complete the backpropagation. The formula for the gradient of the loss function with respect to each intermediate representation is as follows:
[0033]
[0034] where represents the gradient of the loss function with respect to each intermediate representation, and y represents the true target value;
[0035] S27. Each grid edge computing unit calculates the gradient of its local model parameters using its own data and the gradients derived from the intermediate representations provided by the server, and applies the gradients to update the client model as follows:
[0036]
[0037] Wherein, is the gradient of the loss function with respect to the grid edge model G p ; After this step, one round of communication ends; this process iterates the second step until satisfactory performance is obtained; subsequently, the training of the FDIA collaborative detection model is completed.
[0038] Preferably, S3 includes the following steps:
[0039] S31. Based on the generated data, perform collaborative training between the grid edge model and the server-side model, and obtain the integrated grid edge model and the integrated server-side model after training, and utilize them in real time;
[0040] S32. Deploy the integrated grid edge model and the integrated server-side model on the grid edge computing unit and the server respectively, and perform FDIA detection collaboratively according to the measurement data collected from the subnet.
[0041] Preferably, S4 includes the following steps:
[0042] S41. Apply the input data of the grid edge model to the convolutional layer to obtain the feature map F as follows:
[0043] F ∈ R H×W×C ;
[0044] wherein, C represents the number of filters, also known as channels, which is specified by the convolutional layer, and H and W respectively represent the height and width of the input data;
[0045] S42. In the channel attention module, calculate the attention of the channel by compressing the spatial dimension;
[0046] S43. In the spatial attention module, create a spatial attention map by utilizing the spatial relationship between features.
[0047] Preferably, the input data in step S41 includes power injection, power flow, and voltage measurements at T data points collected from the subnetwork.
[0048] Preferably, step S42 includes:
[0049] Perform average pooling and max pooling operations simultaneously to generate average pooling and max pooling features respectively;
[0050] Use these pooled features by a multi-layer perceptron shared network to generate the channel attention map M c ∈ R C×1×1 , as shown in the following formula:
[0051] M c(F) = σ(MLP(AvgP.(F)) + MLP(MaxP.(F)));
[0052] Where, σ represents the sigmoid function;
[0053] Subsequently, multiply the obtained channel attention map with the input feature map F, as shown below:
[0054]
[0055] Where, represents element-wise multiplication, and F c represents the feature map processed by the channel attention module and serves as the input to the spatial attention module.
[0056] Preferably, step S43 includes:
[0057] Calculate the spatial attention map M s (F c ), perform average pooling and max pooling operations along the channel axis, and concatenate the results to create a valid feature descriptor;
[0058] Apply a convolutional layer to the concatenated feature descriptor, as shown in the following equation:
[0059] M s (F c ) = σ(f([AvgPool(F c ); maxPool(F c ))));
[0060] Where, f represents the convolutional operation;
[0061] Multiply the obtained spatial attention map M s by F c , to obtain the final refined output of the convolutional block attention module as shown in the following equation:
[0062]
[0063] Where, F cs represents the final refined output processed by CBAM.
[0064] Preferably, S5 includes the following steps:
[0065] S51. In the Bi-LSTM model structure, each LSTM block includes a storage unit, a forget gate, an input gate, and an output gate. Denote the activation vectors of the forget gate, input gate, output gate, and storage unit as f, i, o, and c respectively. Then the node connections and tensor operations of the forward LSTM hidden layer are defined as:
[0066] f t= Sigmoid(W f [h t-1 , X t + b f );
[0067] i t = Sigmoid(W i [h t-1 , X t + b i );
[0068]
[0069] o t = Sigmoid(W o [h t-1 , X t + b o );
[0070] h t = o t * tanh(c t );
[0071] Wherein, W f , W i , W c , W o are the input weight matrices of f, i, o, and c respectively, b f , b i , b c , b c are the bias vectors of f, i, o, and c respectively, Sigmoid and tanh represent the logistic Sigmoid and hyperbolic tangent activation functions respectively, represents the new state candidate vector, h t-1 represents the hidden state at the previous time step t - 1, X t represents the input vector at the current time step, h represents the hidden state, as the output of the LSTM hidden layer, and the concatenation operation is denoted by parentheses;
[0072] S52. When analyzing time series data, both forward and backward dependencies are considered. To comprehensively capture the forward and backward dependencies in the data, a bidirectional LSTM layer is incorporated into the LSTM configuration. Through the bidirectional operation in Bi - LSTM, the following two computational directions are involved:
[0073]
[0074] Wherein, the symbols → and ← represent the forward and backward operations respectively;
[0075] S53. Two different hidden state vectors and They are calculated independently respectively, and then form the final hidden state vector in the Bi-LSTM through concatenation, as follows:
[0076]
[0077] where h t represents the bidirectional hidden state vector.
[0078] Therefore, the present invention adopts the above-mentioned method for detecting false data injection attacks in a distribution network based on vertical federated learning, and has the following beneficial effects:
[0079] (1) The framework realizes real-time collaboration between the server and the grid edge side by allocating two models created by the split learning method applied to the proposed attention-based hybrid deep learning model;
[0080] (2) The vertical federated learning technology is applied to the FDIA detection in the distribution network, and a new collaborative learning framework is constructed. This framework allows the grid edge computing unit to perform efficient data feature exchange and model collaboration with the central server while protecting data privacy, thereby improving the accuracy and efficiency of FDIA detection;
[0081] (3) In view of the distributed characteristics of the distribution network, the present invention splits the FDIA detection model into two parts: a grid edge model and a server-side model, and deploys them on the grid edge computing unit and the central server respectively. This split deployment method not only reduces the resource requirements of the model for a single computing node, but also improves the scalability and flexibility of the system;
[0082] (4) The present invention also considers the influence of the unbalanced polyphase characteristics and limited measurement data of the distribution network on the detection performance. By designing a special feature extraction and fusion strategy, and dynamically adjusting the model parameters to adapt to network changes, the FDIA detection ability in a complex distribution network environment is improved;
[0083] (5) The method proposed by the present invention can realize data sharing and collaborative learning between different entities, and effectively solve the problem of data sharing in the distribution network; by considering the unbalanced polyphase characteristics and limited measurement data of the distribution network, the present invention has wide applicability and can work effectively in a complex distribution network environment.
[0084] Next, through the accompanying drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0085] Figure 1 is the collaborative FDIA detection framework based on vertical federated learning of the present invention;
[0086] Figure 2 Structural diagram of the FDIA detection model of the present invention;
[0087] Figure 3 Structural diagram of the convolutional block attention module of the present invention;
[0088] Figure 4 Test system diagrams of the IEEE37 and IEEE123 node distribution networks. Detailed implementation manners
[0089] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0090] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0091] Embodiment
[0092] The present invention provides a method for detecting false data injection attacks in a distribution network based on vertical federated learning, including the following steps:
[0093] S1. Propose a vertical federated learning framework based on the split learning method for the FDIA detection model in the distribution network.
[0094] The vertical federated learning framework consists of two main components: the grid edge computing unit and the server, as Figure 1 shown. In the vertical federated learning framework, the local dataset D of the entity p (p = {1,..., P}) in the grid edge computing unit is used for p co-training with the server for the FDIA detection model G f .
[0095] The FDIA detection model G fThe split learning method is split into two models: the grid edge-side model and the server-side model, which are assigned to the grid edge computing unit and the server respectively. The task of the grid edge-side model is to extract spatial features, while the server-side model is responsible for extracting temporal features from the data processed on the grid edge side. The grid edge-side model is designed by adopting an attention module integrated into the deep learning model, while the server-side model is designed based on the Bi-LSTM model.
[0096] Grid edge computing unit: Each grid edge computing unit holds its own grid edge-side model G p , and is responsible for collaborative FDIA detection of a certain part of the distribution network. The grid edge computing unit collects a set of measurement data from local sensors in its respective sub-network and processes it through forward propagation using its own grid edge-side model. The grid edge-side model actively participates in the model update process by exchanging the model output and gradients from the server-side model. Therefore, the grid edge-side model is an active participant in the learning process. To achieve this, it is assumed that each grid edge-side model has sufficient storage, capabilities, and computing power.
[0097] Server: The server has the server-side model G 0 , and undertakes multiple responsibilities during the training and real-time FDIA detection processes. First, the server aggregates the intermediate representations (outputs of the grid edge-side models) collected from the grid edge computing units. Second, it continues to propagate forward using the aggregated intermediate representations. Then, the server updates its own model (the server-side model) and sends the calculated gradients to the split layer to each edge-side model to update the client models. These operations are sequentially executed by the server during the training phase. After the training phase, the server also actively participates in the real-time generation of the final output.
[0098] S2. Train the vertical federated learning framework. The specific training process is as follows:
[0099] S21. Initialization: First, initialize the server-side model G of the server 0 and the client models G in each grid edge computing unit k parameters. Initialize the number of epochs E, learning rates η 0 and η p , and loss function L of the grid edge-side model and the server-side model. At the beginning of each iteration, the participants randomly select a subset B t to reach a consensus, where t represents the iteration number. Each entity p responsible for the sub-network has its own grid edge-side model G p and its respective model parameters θ p , while the server has the server-side model G 0 and its respective model parameters θ 0ownership
[0100] S22. Local data processing at the edge: After initialization, each grid edge computing unit uses its own grid edge-side model G p and its own local dataset for forward propagation. Then, as a result of the forward propagation, obtain the intermediate representation from each grid edge-side model in the t-th round of communication Calculate as follows:
[0101]
[0102] where represents the parameters of the grid edge-side model G of entity p in the t-th round of communication p parameters represents the mini-batch data selected by entity p during the t-th round of communication
[0103] S23. Collection of intermediate representations: After forward propagation in each grid edge computing unit, collect the intermediate representations on the server and perform the following connection:
[0104]
[0105] where represents the connected intermediate representations
[0106] S24. Server-side model execution: The server continues to use the server-side model G 0 and the connected intermediate representations to perform forward propagation to obtain the server-side model output, which is the predicted output. Then, calculate the loss based on the predicted output and select L as the true value label y of the cross-entropy loss function. The formula for the server-side model output is as follows: t The formula for the server-side model output is as follows:
[0107]
[0108] S25. Server-side model update: The server-side model parameters are updated as follows:
[0109]
[0110] where is the gradient of the loss function with respect to the server-side model during the t-th round of communication, and η 0 is the learning rate of the server-side model
[0111] S26. Broadcast Parameters: After updating the server model, the server continues to update each edge-side model of the grid in parallel. For each edge computing unit of the grid, calculate the gradient of the loss function with respect to each intermediate representation, and then send it to the respective edge computing unit of the grid to complete backpropagation. The formula for the gradient of the loss function with respect to each intermediate representation is as follows:
[0112]
[0113] where represents the gradient of the loss function with respect to each intermediate representation, and y represents the true target value.
[0114] S27. Edge-side Model Update of the Grid: Each edge computing unit of the grid calculates the gradient of its local model parameters using its own data and the gradient derived from the intermediate representation provided by the server. Subsequently, apply the gradient to update the client model as follows:
[0115]
[0116] In the formula, is the gradient of the loss function with respect to the edge-side model G p . After this step, one round of communication ends. This process iterates the second step until satisfactory performance is obtained. Subsequently, the training of the FDIA collaborative detection model is completed.
[0117] S3. Through training and inference, the construction and application of the vertical federated learning framework are realized. The specific processes of the two stages of training and inference are as follows:
[0118] S31. Training Stage: Based on the generated data, collaborative training is carried out between the edge-side model of the grid and the server-side model. After training, an integrated edge-side model of the grid and an integrated server-side model are obtained and utilized in real time.
[0119] S32. Inference Stage: Deploy the integrated edge-side model of the grid and the integrated server-side model on the edge computing unit of the grid and the server respectively, and perform FDIA detection collaboratively according to the measurement data collected from the subnet.
[0120] S4. Extract spatial features through the edge-side model of the grid, such as Figure 2As shown, the proposed grid edge-side model includes two convolutional layers, each followed by a pooling layer and a Convolutional Block Attention Module (CBAM) for spatial feature extraction. In a convolutional neural network, all channels are equally important, which may lead to the loss of key information. Therefore, a Convolutional Block Attention Module is added after the convolutional layer to enhance the model's ability to prioritize important features in the data obtained by the grid edge-side model. CBAM combines channel and spatial attention mechanisms, enabling the model to highlight important channel and spatial features, resulting in an enhanced feature representation and more efficient learning for the grid edge-side model.
[0121] S41. Apply the input data of the grid edge-side model (power injection, power flow, and voltage measurements at T data points collected from the sub-network) to the convolutional layer to obtain a feature map F ∈ R H×W×C , where C represents the number of filters, also known as channels, specified by the convolutional layer, and H and W represent the height and width of the input data respectively. CBAM consists of two modules, namely the Channel Attention Module (CAM) and the Spatial Attention Module (SAM).
[0122] Subsequently, the attention weighting operations in the channel and spatial dimensions are sequentially performed by CAM and SAM, and the specific process is as follows:
[0123] S42. Refer to Figure 3 , in CAM, calculate the channel attention by compressing the spatial dimension, which specifically includes the following steps:
[0124] First, perform average pooling and max pooling operations simultaneously to generate average pooling and max pooling features respectively.
[0125] Then, a shared network of a Multi-Layer Perceptron (MLP) uses these pooled features to generate a channel attention map M c ∈ R C×1×1 , as shown in the following formula:
[0126] M c (F) = σ(MLP(AvgP.(F)) + MLP(MaxP.(F)));
[0127] where σ represents the sigmoid function.
[0128] Subsequently, multiply the obtained channel attention map by the input feature map F, as follows:
[0129]
[0130] where, represents element-wise multiplication, and F c represents the feature map processed by CAM and serves as the input to SAM.
[0131] S43. In the SAM, a spatial attention map is created using the spatial relationships between features. Contrary to the CAM, the SAM emphasizes the importance of each part in the feature map and serves as a complement to the CAM. The specific steps are as follows:
[0132] To calculate the spatial attention map M s (F c ), average pooling and max pooling operations are performed along the channel axis, and their results are concatenated to create an effective feature descriptor.
[0133] Then, a convolutional layer is applied to the concatenated feature descriptor as shown in the following equation:
[0134] M s (F s ) = σ(f([AvgPool(F c ) ; maxPool(F c )]));
[0135] where f represents the convolutional operation.
[0136] Finally, the obtained spatial attention map M s is multiplied by F c to obtain the final refined output of CBAM as shown in the following equation:
[0137]
[0138] where F cs represents the final refined output processed by CBAM. Additionally, a second CBAM module is given as the last layer of the grid edge side model. F cs represents the intermediate representation transmitted from the grid edge side model to the server and represents the spatial features extracted from each grid edge side model.
[0139] S5. Temporal features are extracted from the input of the server model. In the vertical federated learning framework, the server model is split after the second convolutional block attention module, so that the grid edge side models are specified to extract spatial features from a set of measurement data collected from their respective subnets, while the server side model extracts temporal features from the aggregated intermediate representation. It should be noted that each intermediate representation H p represents the spatial features extracted by the grid edge side model.
[0140] As Figure 2 shown, by splitting the proposed model (what model) after the second CBAM module, the task of the server side model is to extract temporal features from its input. This input includes the concatenated intermediate representation H all of the extracted spatial features, the concatenated intermediate representation H allProcessed by the server model. The server-side model uses a Bi-LSTM model to extract temporal relationships because the Bi-LSTM model proficiently captures forward and backward dependencies, thereby improving the performance of the server-side model. In addition, it also addresses the common challenges of vanishing gradients and insufficient modeling of backward dependencies during backpropagation, which are prevalent in most recurrent models. Moreover, it also overcomes the limitations of unidirectional LSTMs, which can only handle dependencies in one direction. The concept of Bi-LSTM is derived from bidirectional RNNs, which use two different hidden layers to process sequence data in the forward and backward directions.
[0141] S51. In the Bi-LSTM model structure, each LSTM block includes a memory cell, a forget gate, an input gate, and an output gate. Represent the activation vectors of the forget gate, input gate, output gate, and memory cell as f, i, o, and c respectively. Then the node connections and tensor operations of the forward LSTM hidden layer are defined as:
[0142] f t = Sigmoid(W f [h t-1 , X t + b f );
[0143] i t = Sigmoid(W i [h t-1 , X t + b i );
[0144]
[0145] o t = Sigmoid(W o [h t-1 , X t + b o );
[0146] h t = o t * tanh(c t );
[0147] Among them, W f , W i , W c , W o are the input weight matrices of f, i, o, and c respectively, and b f , b i , b c , b cBias vectors f, i, o, and c respectively, where Sigmoid and tanh denote the logistic Sigmoid and hyperbolic tangent activation functions, represents the new state candidate vector, h t-1 represents the hidden state at the previous time step t - 1, X t represents the input vector at the current time step, h represents the hidden state, which is the output of the LSTM hidden layer, and the concatenation operator is denoted by parentheses.
[0148] S52. When analyzing time - series data, it is crucial to consider both forward and backward dependencies simultaneously. The backward dependencies obtained from data sorted in reverse chronological order provide unique and valuable insights that cannot be obtained solely from forward dependencies. Therefore, to comprehensively capture both types of dependencies present in the data, a bidirectional LSTM layer is incorporated into the LSTM configuration. The node connections and tensor computations in Bi - LSTM are very similar to those in unidirectional LSTM, with the main difference being the processing direction. Specifically, the operations in Bi - LSTM are bidirectional and involve two computational directions:
[0149]
[0150] where the symbols → and ← denote forward and backward operations respectively.
[0151] S53. Two different hidden state vectors and are calculated independently and then combined through concatenation to form the final hidden state vector within Bi - LSTM as follows:
[0152]
[0153] where h t represents the bidirectional hidden state vector.
[0154] The server - side model consists of dense layers following the Bi - LSTM model, as Figure 3 shown. Therefore, in the case of collaborating with the edge - side model of the grid, FDIA detection decisions can be obtained in the server - side model.
[0155] Refer to Figure 4, various robustness tests and comparative studies are first conducted to evaluate the effectiveness of the proposed model on the IEEE 123-bus and IEEE 37-bus test systems. All experiments are carried out using a Google Colab Pro Notebook with an NVIDIA A100-SXM4-40GB GPU. In the IEEE 123-bus test system, some switches along the branches are usually in the closed state, such as branches 13-152, 18-135, 60-160, and 97-197. The IEEE 123-bus test system is divided into 6 sub-networks, and the IEEE 37-bus test system is divided into 2 sub-networks.
[0156] Comparison with the state-of-the-art methods: The proposed model is compared with the current state-of-the-art methods (CNN, LSTM, and Bi-LSTM models). The method proposed in the present invention achieves an accuracy of 99.88%, an f1-score of 99.82%, a precision of 100%, and a recall of 99.75% in the IEEE 123-bus test system, and an accuracy of 99.06%, an f1-score of 98.85%, a precision of 99.45%, and a recall of 98.65% in the IEEE 37-bus test system, surpassing the state-of-the-art methods in various evaluation metrics. Among the two test systems, LSTM performs the worst, with an accuracy of 98.12%, an f1-score of 97.51%, a precision of 99.25%, and a recall of 96.91% in the IEEE 123-bus test system, and an accuracy of 98.19%, an f1-score of 97.88%, a precision of 98.77%, and a recall of 97.57% in the IEEE 37-bus test system. Similarly, the performance of Bi-LSTM and CNN is also relatively inferior to the method proposed in the present invention. These methods only focus on extracting temporal or spatial features from the measurements. In contrast, the proposed attention-based model effectively extracts both spatial and temporal features, thus providing the best detection performance for the two test systems.
[0157] Therefore, the present invention adopts the above-mentioned method for detecting false data injection attacks in a distribution network based on vertical federated learning. By constructing a collaborative learning framework, it uses vertical federated learning technology to achieve data sharing and collaborative learning among different entities, and at the same time combines an advanced deep learning model for efficient FDIA detection.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements do not make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for detecting false data injection attacks in distribution networks based on vertical federated learning, characterized in that: The following steps are involved: S1. Build a vertical federated learning framework based on the split learning method to collaboratively build a FDIA detection model in the distribution network; The vertical federated learning framework includes grid edge computing units and servers; The FDIA detection model is split into a grid edge side model and a server side model based on a split learning method. The grid edge side model and the server side model are assigned to the grid edge computing unit and the server respectively. S2, training the vertical federated learning framework; S3. Implement the application of vertical federated learning framework through training and reasoning operations; S4. Extracting spatial features through a mesh edge side model. The mesh edge side model includes two convolutional layers. Each convolutional layer is followed by a pooling layer and a convolutional block attention module. The convolutional block attention module includes a channel attention module and a spatial attention module. S5. The input of the Bi-LSTM model of the server model extracts the temporal features, and the input includes the concatenated intermediate representation H of the extracted spatial features. all , connect the middle to represent H all Handled by the server model.
2. According to claim 1, a method for detecting false data injection attacks in a power distribution network based on vertical federated learning is characterized in that: Each grid edge computing unit holds its own grid edge side model G p ,The Grid Edge Computing Unit collects a set of measurement data from the local sensors in their respective sub-networks and processes them through forward propagation using its own Grid Edge Side Model.
3. According to claim 2, a method for detecting false data injection attacks in a distribution network based on vertical federated learning is characterized in that: The server has the server-side model G0, which performs the following operations during training and real-time FDIA detection: the server aggregates the intermediate representations collected from the grid edge computing units; The aggregated intermediate representation is used to continue forward propagation; the server updates its own model and sends the gradient calculated to the split layer to each edge model to update the client model.
4. According to claim 3, a method for detecting false data injection attacks in a distribution network based on vertical federated learning is characterized in that: S2 includes the following steps: S21, initialize the server-side model G0 and the client model G in each grid edge computing unit k The parameters include the epoch number E, learning rate η0 and η of the grid edge model and server model. p , loss function L, at the beginning of each iteration, the participant randomly selects a subset B t A consensus is reached, where t represents the number of iterations; S22, each grid edge computing unit uses its own grid edge side model G p Perform forward propagation with its own local dataset, and then obtain the intermediate representation from each grid edge side model in the tth round of communication The calculation formula is as follows: in, The grid edge model G represents the entity p in the tth round of communication p parameter, represents the mini-batch data selected by entity p in the tth round of communication; S23. After forward propagation in each grid edge computing unit, collect the intermediate representation on the server And make the following connections: in, An intermediate representation representing a connection; S24. The server continues to perform forward propagation using the server-side model and the connected intermediate representation to obtain the server-side model output, and then calculates the loss based on the server-side model output and selects the loss function as the true value label y of the cross entropy loss function t , the formula for the server-side model output is as follows: in, Represents the server-side model output, which is also the predicted output; S25, update the server-side model parameters, the calculation process is as follows: in, is the gradient of the loss function relative to the server-side model at the tth round of communication, and η0 is the learning rate of the server-side model; S26. After updating the server model, the server continues to update each grid edge side model in parallel. For each grid edge computing unit, the gradient of the loss function relative to each intermediate representation is calculated, and then sent to the respective grid edge computing unit to complete back propagation. The gradient formula of the loss function relative to each intermediate representation is as follows: in, represents the gradient of the loss function with respect to each intermediate representation, and y represents the true target value; S27. Each grid edge computing unit calculates the gradients of its local model parameters using its own data and the gradients derived from the intermediate representation provided by the server, and applies the gradients to update the client model as follows: In the formula, is the loss function relative to the grid edge model G p The gradient of ; after this step, a round of communication ends; the process iterates the second step until satisfactory performance is obtained; then, the training of the FDIA collaborative detection model is completed.
5. According to claim 4, a method for detecting false data injection attacks in a distribution network based on vertical federated learning is characterized in that: S3 includes the following steps: S31. Based on the generated data, collaborative training is performed between the grid edge side model and the server side model, and after the training, a comprehensive grid edge side model and a comprehensive server side model are obtained, and they are used in real time; S32, deploying the integrated grid edge-side model and the integrated server-side model on the grid edge computing unit and the server respectively, and collaboratively performing FDIA detection based on the measurement data collected from the subnet.
6. According to claim 5, a method for detecting false data injection attacks in a distribution network based on vertical federated learning is characterized in that: S4 includes the following steps: S41, applying the input data of the mesh edge side model to the convolution layer, and obtaining the feature map F as follows: F∈R H×W×C ; Where C represents the number of filters, also called channels, specified by the convolutional layer, and H and W represent the height and width of the input data, respectively; S42, in the channel attention module, calculating the channel attention by compressing the spatial dimension; S43. In the spatial attention module, the spatial relationship between features is used to create a spatial attention map.
7. The method for detecting false data injection attacks in a power distribution network based on vertical federated learning according to claim 6 is characterized in that: The input data in step S41 includes power injection, power flow and voltage measurements at T data points collected from the sub-network.
8. A method for detecting false data injection attacks in a power distribution network based on vertical federated learning according to claim 7, characterized in that: Step S42 includes: Perform average pooling and maximum pooling operations simultaneously to generate average pooling and maximum pooling features respectively; A multi-layer perceptron shared network uses these pooled features to generate a channel attention map M c ∈R C×1×1 , as shown below: Among them, σ represents the sigmoid function; Subsequently, the resulting channel attention map is multiplied with the input feature map F as follows: in, represents element-wise multiplication, F c Represents the feature map processed by the channel attention module, which serves as the input of the spatial attention module.
9. A method for detecting false data injection attacks in a power distribution network based on vertical federated learning according to claim 8, characterized in that: Step S43 includes: Compute the spatial attention map M s (F c ), average pooling and maximum pooling operations are applied along the channel axis and their results are concatenated to create an effective feature descriptor; A convolutional layer is applied to the concatenated feature descriptors as shown below: M s (F c )=σ(f([AvgPool(F c );maxPool(F c )])); Among them, f represents the convolution operation; The obtained spatial attention map M s Multiply by F c , the final refined output of the convolutional block attention module is shown as follows: Among them, F cs Represents the final refined output of the CBAM process.
10. A method for detecting false data injection attacks in a power distribution network based on vertical federated learning according to claim 9, characterized in that: S5 includes the following steps: S51. In the Bi-LSTM model structure, each LSTM block includes a storage unit, a forget gate, an input gate, and an output gate. The activation vectors of the forget gate, input gate, output gate, and storage unit are represented as f, i, o, and c, respectively. The node connection and tensor operation of the forward LSTM hidden layer are defined as: f t =Sigmoid(W f [h t-1 ,X t ]+b f ); i t =Sigmoid(W i [h t-1 ,X t ]+b i ); o t =Sigmoid(W o [h t-1 ,X t ]+b o ); h t =o t *tanh(c t ); Among them, W f , W i , W c , W o are the input weight matrices of f, i, o and c respectively, and b f 、b i 、b c 、b c are the bias vectors of f, i, o and c respectively. Sigmoid and tanh represent the logistic sigmoid and hyperbolic tangent activation functions respectively. represents the new state candidate vector, h t-1 represents the hidden state of the previous time step t-1, X t represents the input vector of the current time step, h represents the hidden state, which is the output of the LSTM hidden layer, and the connection operator is represented by brackets; S52. When analyzing time series data, both forward and backward dependencies are considered. In order to fully capture the forward and backward dependencies in the data, a bidirectional LSTM layer is merged into the LSTM configuration. Through the bidirectional operation in Bi-LSTM, the following two calculation directions are involved: Among them, the symbols → and ← represent forward and reverse operations respectively; S53, two different hidden state vectors and They are calculated independently and then combined in series to form the final hidden state vector within the Bi-LSTM, as shown below: Among them, h t represents the bidirectional hidden state vector.