A power grid internal network traffic prediction and anomaly detection method and system based on federated learning and generative AI

By combining federated learning with generative AI, synthetic samples are generated and model parameters are asynchronously aggregated, which solves the data privacy and latency issues in power grid network monitoring, achieves efficient and secure flow anomaly detection, and improves detection accuracy and response speed.

CN120434061BActive Publication Date: 2025-10-10ZHANGZHOU POWER SUPPLY COMPANY STATE GRID FUJIANELECTRIC POWER +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510940369.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-10
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing power grid network monitoring solutions rely on centralized data collection, which has high data transmission delays and privacy leakage risks. In addition, traditional methods find it difficult to accurately distinguish between normal load fluctuations and malicious attack traffic.

Method used

By combining federated learning with generative AI, the generative adversarial network (GAN) is trained locally on edge nodes to generate synthetic samples, enhance the training set, and asynchronously aggregate model parameters on the central server. The deep autoencoder (AE) and LSTM model are combined for traffic prediction and anomaly detection, achieving lightweight and efficient anomaly detection.

Benefits of technology

It significantly improves the accuracy and response speed of anomaly detection, reduces data transmission latency, protects power grid data privacy, and reduces computation and memory usage while maintaining high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434061B_ABST
    Figure CN120434061B_ABST
Patent Text Reader

Abstract

The application discloses a power grid internal network flow prediction and anomaly detection method and system based on federated learning and generative AI. The system deploys local network flow collection and prediction models on each terminal node in the power grid. Through a federated learning mechanism, the system aggregates each node model on a central server, realizing collaborative learning of the whole network data without directly sharing the original data. The system integrates a generative AI module to generate synthetic network flow samples to enhance the model training data. The federated learning adopts an asynchronous aggregation strategy, dynamically adjusts the weights of each node, and performs lightweight processing on the model to adapt to the resource-limited power terminal environment. Through the global prediction model, real-time flow is predicted and detected. When the detected network flow deviates from the predicted value by more than a preset threshold, the system automatically starts a security response mechanism. The technical scheme can improve the anomaly detection accuracy and response speed, and protect the power grid data privacy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial control system security, and in particular to a power grid internal network traffic prediction and anomaly detection method and system based on federated learning and generative AI. BACKGROUND

[0002] The power grid internal communication network provides a basic support for power system monitoring and control, and its network traffic features contain power grid operation state information. With the wide deployment of smart meters, substation automation and Internet of Things devices, power grid communication traffic presents periodic changes and diversified scenarios. However, existing network monitoring solutions usually rely on centralized data collection and analysis, which has problems such as high data transmission delay and privacy leakage risk. At the same time, real abnormal traffic samples are often scarce, making it difficult for traditional statistical or machine learning methods to accurately distinguish between normal load fluctuations and malicious attack traffic.

[0003] Federated learning, as a distributed machine learning technology, can realize multi-source data collaborative modeling without exposing local data, which can alleviate the problems of power grid data privacy and communication overhead. However, the traditional federated learning commonly uses a synchronous aggregation mode, which needs to wait for all terminals to complete iteration, resulting in high delay and low reliability. On the other hand, generative AI technology (such as generative adversarial networks) has made breakthroughs in synthetic data enhancement and can be used to generate network traffic samples to expand the training set and improve the model's ability to identify abnormal situations. How to combine federated learning with generative AI to build an efficient and secure power grid network traffic prediction and anomaly detection system has become a problem to be solved. SUMMARY

[0004] Therefore, the purpose of the present application is to provide a power grid internal network traffic prediction and anomaly detection method and system based on federated learning and generative AI, which can improve the accuracy and response speed of anomaly detection while protecting the privacy of power grid data.

[0005] To achieve the above purpose, the present application adopts the following technical solution: a power grid internal network traffic prediction and anomaly detection method based on federated learning and generative AI, comprising the following steps:

[0006] S1: In the data collection and preprocessing stage, the edge node captures local area network packets in real time through a mirror port or a network TAP;

[0007] S2: Protocol analysis is performed on the original packets to extract basic five-tuple, packet length, timestamp statistical features, and event type, state change flag protocol features for IEC61850 MMS packets and GOOSE event packets;

[0008] S3: Normalize the time series features in the extracted quintuples and generate a traffic feature vector sequence x according to a fixed time window to prepare for subsequent model input:

[0009] S4: Locally at each edge node, a generative adversarial network (GAN) is trained using the previous 24 hours of normal traffic feature data to model the normal traffic feature distribution. A generator and a discriminator are set up. The generator takes random noise as input and learns to generate synthetic features consistent with real normal traffic. The discriminator distinguishes between real and synthetic samples, and the two are alternately optimized until convergence.

[0010] S5: After training convergence, the generator is used to produce synthetic samples. To balance the low-frequency patterns, the focus is on oversampling the low-frequency normal patterns in the training set to balance the data distribution. For each normal pattern cluster Calculate its frequency : Define sampling weights and randomly sample noise according to the weights Generate corresponding pattern samples;

[0011] S6: Mix the synthetic samples generated in S5 with the original normal samples to form an enhanced local training set;

[0012] S7: A deep autoencoder (AE) and an LSTM prediction model are used locally to form a combined model. The deep autoencoder (AE) has 1 to 3 encoder and decoder layers and 64 to 256 hidden units. The LSTM prediction model has a time step of 5 to 20 steps and 32 to 128 hidden units. If the edge node has more than 256MB of available memory and floating-point computing capabilities, a three-layer autoencoder and a 128-unit LSTM model are used. If the memory is less than 128MB, a single-layer structure and a 32-unit network are used.

[0013] S8: Locally train the combined model of S7 on the enhanced training set, optimizing the loss function of reconstruction error or prediction error until the validation set loss converges;

[0014] S9: Given the limited computing resources and memory of edge nodes, the combined model is pruned and quantized after local training. Model pruning evaluates the importance of each neuron or channel and removes those with low contributions.

[0015] S10: After local training is completed, only the combined model parameters and gradients are uploaded to the central server, without uploading any original or synthetic traffic data;

[0016] S11: Combined model parameter buffer ; When receiving any edge node Upload parameters When the buffer is first stored; when the buffer The cumulative amount reaches the preset threshold or timeout Upon arrival, an aggregation is performed immediately, and a new global model is calculated by weighted average of all combined model parameters in the buffer. , Represents the model parameter vector in the local model parameter set uploaded by the edge node:

[0017] ;

[0018] S12: Clear the buffer after aggregation , and Send it to each edge node and start a new round of local training;

[0019] S13: The edge node uses the latest global model to perform online inference on the real-time traffic feature sequence and calculate the reconstruction error and prediction error , Represents the vector obtained by normalizing the original input feature vector in the time window t, represents the hidden representation vector obtained after processing by the encoder module in the autoencoder, It is the output of the decoder module in the autoencoder after restoring the hidden representation vector. It represents the two-norm calculation, Represents the normalized feature vector predicted by the LSTM prediction model based on historical input. If and A large difference indicates that current behavior deviates from historical patterns:

[0020]

[0021] S14: Define the composite anomaly score and with dynamic threshold Compare, when When it is judged as abnormal,

[0022] Threshold Based on the historical normal score distribution

[0023]

[0024] in and are the mean and standard deviation of the scores during the normal period, is a constant to control the false alarm rate;

[0025] S15: The node automatically limits or isolates the traffic of suspected abnormal source IP according to the policy; at the same time, it reports the alarm and related feature summary to the operation and maintenance platform for manual review.

[0026] In a preferred embodiment, in S3, the original message is normalized to generate a flow feature vector sequence :

[0027] Each feature vector contains: the number of messages in the window , total bytes , the mean of packet length and standard deviation ;

[0028] Perform min–max normalization on each dimension feature to eliminate dimensional differences.

[0029] In a preferred embodiment, the sampling weights are defined in S5 Specifically:

[0030] ,

[0031] in is a small positive constant used to prevent division by zero.

[0032] In a preferred embodiment, in S7, the autoencoder reconstruction method: Encoder Mapping the input to the latent space , decoder Refactored to ; training to minimize the reconstruction error, It represents the input feature vector after min-max normalization, which is the original input of the model and represents the network traffic characteristics obtained by statistics within a time window. represents the encoder function in the autoencoder, which compresses the input into a lower-dimensional hidden vector. is the decoder function, which attempts to reconstruct the original input from the encoded hidden representation, and finally performs the square of the L2 norm to measure the size of the difference between the two vectors:

[0033] .

[0034] In a preferred embodiment, in S9, the combined model pruning adopts a sparse pruning method based on weight amplitude. After the training is completed, the absolute mean of the weight corresponding to each neuron is calculated, and the part below the set threshold will be set to zero or removed; the pruning ratio is set to 10% to 50%.

[0035] In a preferred embodiment, in S11, the asynchronous aggregation strategy includes receiving terminal node prediction model updates in batches within a preset time interval from the central server, without waiting for all nodes to complete training, and performing model aggregation immediately when the update amount or timeout condition is reached.

[0036] In a preferred embodiment, in S14, a comprehensive abnormality score is defined

[0037]

[0038] and with dynamic threshold Compare, when When it is judged as abnormal, is a weighting coefficient used to adjust the relative contribution of the autoencoder reconstruction error and the LSTM prediction error in the comprehensive anomaly score.

[0039] The present invention also provides a system for predicting and detecting network traffic within a power grid based on federated learning and generative AI, and runs the method for predicting and detecting network traffic within a power grid based on federated learning and generative AI.

[0040] Compared with existing technologies, this invention offers the following advantages: By introducing the FedBuff asynchronous aggregation strategy, it overcomes the bottleneck of slow nodes hindering overall training progress. Once some edge nodes complete local training and report model parameters, the central server can immediately aggregate and distribute the new global model, eliminating the need to wait for all nodes to synchronize data. This reduces the average latency of a single model iteration from several minutes to just over ten seconds. This significantly improves the response speed to sudden network attacks. Furthermore, this invention deeply integrates a generative enhancement module with federated learning: After the GAN locally augments normal samples and balances the pattern distribution at each node, the enhanced data participates in the global model update during federated training. This ensures that the final model not only recognizes the personalized traffic characteristics of each node but also integrates traffic patterns from a network-wide perspective. Finally, through configurable model pruning and quantization techniques, this invention reduces memory usage and computational complexity for autoencoder and LSTM models by nearly 60% and 50%, respectively, while maintaining over 90% detection accuracy. Combining the aforementioned asynchronous aggregation and generative enhancement, the present invention realizes "lightweight, high-precision, low-latency, and full-network collaborative" power grid flow anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Flowchart of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0042] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0043] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.

[0044] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form, and it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.

[0045] The basic concept of this invention is to deploy distributed, lightweight traffic prediction and detection agents within the power grid. Using a federated learning framework, local model updates from each node are aggregated into a global model. A generative AI module is then introduced at the central level to supplement and enhance normal and abnormal traffic samples. Specifically, edge nodes collect and preprocess local network traffic, train lightweight prediction models such as autoencoders and LSTMs locally, and asynchronously upload model parameters. The central server then integrates the updates from each node using the FedBuff asynchronous aggregation algorithm to form a global traffic prediction model, which is then quantized, pruned, and transmitted back to each node. Synthetic traffic samples are then generated using a generative adversarial network (GAN) or diffusion model to enhance the ability to identify rare and abnormal patterns. Finally, each node compares the prediction-observation error of real-time traffic and triggers automatic isolation or alarms if a threshold is exceeded. This invention, through the synergistic combination of generative AI, federated asynchronous aggregation, and lightweight models, protects the data privacy of each site while significantly improving the accuracy, real-time nature, and robustness of power grid traffic anomaly detection.

[0046] like Figure 1 As shown, the present invention provides a power grid internal network traffic prediction and anomaly detection system based on federated learning and generative AI, comprising the following steps:

[0047] S1: During the data collection and preprocessing phase, edge nodes (substation gateways or smart sensors) capture LAN messages in real time through mirror ports or network taps.

[0048] S2: Perform protocol analysis on the original message to extract statistical features such as the basic five-tuple (source / destination IP, port, protocol), packet length, and timestamp. It also extracts protocol features such as event type and state change flag for IEC61850 MMS messages and GOOSE event messages.

[0049] S3: Normalize the extracted time series features and generate a traffic feature vector sequence x according to a fixed time window (set to 5 seconds) to prepare for subsequent model input:

[0050]

[0051] in Indicates the The original statistical features in the time window The value inside, and Respectively represent the minimum and maximum values ​​of the feature within the set normalized time range; Represents the normalized eigenvalue. This normalization process is used to eliminate the dimensional differences between different features and improve the stability and convergence speed of model training.

[0052] S4: At each edge node, use the normal traffic feature data of the previous 24 hours to train a generative adversarial network (GAN) to model the normal traffic feature distribution. and the discriminator The generator takes random noise as input and learns to generate synthetic features consistent with real normal traffic; the discriminator distinguishes between real and synthetic samples, and the two are optimized alternately until convergence. are the parameters of the generator model, are the parameters of the discriminator model, It is the overall adversarial loss function, which consists of two parts: the true sample recognition score of the discriminator and the penalty term for the incorrect recognition of the generated sample. Represents the real sample , is the normal traffic sample collected, Represents a random vector From the prior noise distribution Medium sampling, indicates that the generator maps noise into fake samples, Represents the probability that the discriminator identifies the input sample as real data;

[0053]

[0054] in, Represents the generator parameters The minimization operation, Represents the discriminator parameters The maximization operation, is the adversarial loss function of GAN, Represents a sample From the real data distribution Sampling and obtaining the mathematical expectation of the output value, represents the noise vector From the prior noise distribution Take samples from the dataset and obtain the calculated mathematical expectation.

[0055] S5: After training converges, the generator is used to produce synthetic samples In order to balance the low-frequency patterns, we focus on oversampling the low-frequency normal patterns in the training set (such as specific control instruction message sequences) to balance the data distribution. Calculate its frequency :

[0056]

[0057] Define sampling weights and randomly sample noise according to the weights Generate corresponding mode samples so that the enhanced data set is approximately evenly distributed across each mode;

[0058] S6: Mix the synthetic samples with the original normal samples to form an enhanced local training set;

[0059] S7: Locally, a combination of a deep autoencoder (AE) and an LSTM prediction model is used. The autoencoder identifies anomalies based on reconstruction errors, while the LSTM identifies anomalies based on time series prediction errors. The model structure selects an appropriate number of layers and hidden units based on the computing power of edge nodes:

[0060] encoder Map the input to the latent space h, the decoder Refactored to . Train to minimize the reconstruction error, is the loss function of AE:

[0061]

[0062] LSTM prediction: based on the previous Window features predict the next window , minimize the prediction error, is the loss function of the LSTM model:

[0063]

[0064] The total loss is

[0065]

[0066] in is the weight hyperparameter.

[0067] S8: Perform local training on the enhanced training set to optimize the loss function of reconstruction error or prediction error until the validation set loss converges;

[0068] S9: Due to the limited computing resources and memory of edge nodes, the detection model is pruned and quantized after local training. Model pruning evaluates the importance of each neuron or channel and removes the ones with low contribution to reduce the number of parameters. represents the original model, Represents the pruned model:

[0069] Where r is the pruning rate (set to 0.5); quantization maps model weights from 32-bit floating-point numbers to 8-bit integers, reducing computational complexity by nearly 4x. This lightweight model reduces memory usage and inference latency on edge devices by approximately 60% and 50%, respectively, while maintaining detection accuracy at least 90% of the original model.

[0070] S10: After local training is completed, only the model parameters and gradients are uploaded to the central server, without uploading any original or synthetic traffic data;

[0071] S11: Central server maintains model parameter buffer When receiving any edge node Upload parameters When the buffer is The cumulative amount reaches the preset threshold (set to 50% of the total number of nodes) or timeout (set to 30 minutes) arrives, an aggregation is performed immediately, and a new global model is calculated by weighted average of all model parameters in the buffer. :

[0072] ;

[0073] S12: Clear the buffer after aggregation , and Send it to each edge node and start a new round of local training.

[0074] S13: Edge nodes use the latest global model to analyze real-time traffic feature sequences Perform online inference and calculate reconstruction error and prediction error :

[0075]

[0076] S14: Define a comprehensive anomaly score and combine it with a dynamic threshold Compare, when When it is judged as abnormal,

[0077] Threshold Based on the historical normal score distribution, it can be set as

[0078]

[0079] in and respectively the mean and standard deviation of scores during normal period, a constant to control false positive rate, here set to 0.5.

[0080] S15: The node automatically limits the flow rate or isolates the suspected abnormal source IP according to the strategy; meanwhile, the alarm and the related feature summary are reported to the operation and maintenance platform for manual review.

[0081] The above only describes the embodiments of the present application and is not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for predicting network traffic and detecting anomalies within a power grid based on federated learning and generative AI, characterized in that: The following steps are involved: S1: In the data collection and preprocessing stage, the edge node captures LAN messages in real time through the mirror port or network TAP; S2: Perform protocol parsing on the original message to extract basic five-tuple, packet length, and timestamp statistical features. It also extracts event type and state change flag protocol features for IEC61850 MMS messages and GOOSE event messages. S3: Normalize the time series features in the extracted quintuples and generate a traffic feature vector sequence x according to a fixed time window to prepare for subsequent model input: S4: Locally at each edge node, use the normal traffic feature data of the previous 24 hours to train a generative adversarial network (GAN) to model the normal traffic feature distribution; Suppose a generator and a discriminator. The generator takes random noise as input and learns to generate synthetic features consistent with real normal traffic; The discriminator distinguishes between real and synthetic samples, and the two are optimized alternately until convergence; S5: After training convergence, the generator is used to produce synthetic samples. To balance the low-frequency patterns, the focus is on oversampling the low-frequency normal patterns in the training set to balance the data distribution. For each normal pattern cluster Calculate its frequency : Define sampling weights and randomly sample noise according to the weights Generate corresponding pattern samples; S6: Mix the synthetic samples generated in S5 with the original normal samples to form an enhanced local training set; S7: A deep autoencoder (AE) and an LSTM prediction model are used locally to form a combined model. The deep autoencoder (AE) has 1 to 3 encoder and decoder layers and 64 to 256 hidden units. The LSTM prediction model has a time step of 5 to 20 steps and 32 to 128 hidden units. If the edge node has more than 256MB of available memory and floating-point computing capabilities, a three-layer autoencoder and a 128-unit LSTM model are used. If the memory is less than 128MB, a single-layer structure and a 32-unit network are used. S8: Locally train the combined model of S7 on the enhanced training set, optimizing the loss function of reconstruction error or prediction error until the validation set loss converges; S9: Due to the limited computing resources and memory of edge nodes, the combined model is pruned and quantized after local training. Model pruning evaluates the importance of each neuron or channel and removes those with low contribution. S10: After local training is completed, only the combined model parameters and gradients are uploaded to the central server, without uploading any original or synthetic traffic data; S11: Combined model parameter buffer ; When receiving any edge node Upload parameters When the buffer is The cumulative amount reaches the preset threshold or timeout Upon arrival, an aggregation is performed immediately, and a new global model is calculated by weighted average of all combined model parameters in the buffer. , Represents the model parameter vector in the local model parameter set uploaded by the edge node: ; S12: Clear the buffer after aggregation , and Send it to each edge node and start a new round of local training; S13: The edge node uses the latest global model to perform online inference on the real-time traffic feature sequence and calculate the reconstruction error and prediction error , Represents the vector obtained by normalizing the original input feature vector in the time window t, represents the hidden representation vector obtained after processing by the encoder module in the autoencoder, It is the output of the decoder module in the autoencoder after restoring the hidden representation vector. It represents the two-norm calculation, Represents the normalized feature vector predicted by the LSTM prediction model based on historical input. If and A large difference indicates that current behavior deviates from historical patterns: S14: Define the composite anomaly score and with dynamic threshold Compare, when When it is judged as abnormal, Threshold Based on the historical normal score distribution in and are the mean and standard deviation of the scores during the normal period, is a constant to control the false alarm rate; S15: The node automatically limits or isolates the traffic of suspected abnormal source IP according to the policy; at the same time, it reports the alarm and related feature summary to the operation and maintenance platform for manual review.

2. A method for predicting and detecting anomaly in a power grid based on federated learning and generative AI according to claim 1, characterized in that: In S3, the original message is normalized to generate a traffic feature vector sequence : Each feature vector contains: the number of messages in the window , total bytes , the mean of packet length and standard deviation ; Perform min–max normalization on each dimension feature to eliminate dimensional differences.

3. The method for predicting and detecting anomaly in a power grid based on federated learning and generative AI according to claim 1, characterized in that: Define sampling weights in S5 Specifically: , in is a small positive constant used to prevent division by zero.

4. The method for predicting and detecting anomaly in a power grid based on federated learning and generative AI according to claim 1, characterized in that: In S7, the autoencoder reconstruction method: encoder Mapping the input to the latent space , decoder Refactored to ; training to minimize the reconstruction error, It represents the input feature vector after min-max normalization, which is the original input of the model and represents the network traffic characteristics obtained by statistics within a time window. represents the encoder function in the autoencoder, which compresses the input into a lower-dimensional hidden vector. is the decoder function, which attempts to reconstruct the original input from the encoded hidden representation, and finally performs the square of the L2 norm to measure the size of the difference between the two vectors: 。 5. The method for predicting network traffic and detecting anomalies in a power grid based on federated learning and generative AI according to claim 1, characterized in that: In S9, the combined model pruning adopts a sparse pruning method based on weight amplitude. After training is completed, the absolute mean of the weight corresponding to each neuron is calculated, and the part below the set threshold will be set to zero or removed; the pruning ratio is set to 10% to 50%.

6. The method for predicting and detecting anomaly in a power grid based on federated learning and generative AI according to claim 1, characterized in that: In S11, the asynchronous aggregation strategy includes receiving terminal node prediction model updates in batches within a preset time interval from the central server, without waiting for all nodes to complete training, and performing model aggregation immediately when the update amount or timeout condition is reached.

7. The method for predicting network traffic and detecting anomalies in a power grid based on federated learning and generative AI according to claim 1, characterized in that: In S14, the comprehensive anomaly score is defined and with dynamic threshold Compare, when When it is judged as abnormal, is a weighting coefficient.

8. A system for predicting network traffic and detecting anomalies within a power grid based on federated learning and generative AI, characterized by A method for predicting network traffic and detecting anomalies within a power grid based on federated learning and generative AI as described in any one of claims 1-7 is run.

Citation Information

Patent Citations

  • Regional photovoltaic power probability prediction method based on federated learning and cooperative regulation and control system thereof

    CN111626506A

  • Power network flow anomaly detection method based on hierarchical federated learning

    CN116016110A