Dynamic optimization method for intelligent security convergence protection strategy of power system

By building a four-layer dynamic optimization framework of ‘perception-evaluation-decision-execution’, the problems of insufficient static defense, data fragmentation and real-time performance of complex multi-source attacks in the intelligent transformation of power systems are solved, efficient security protection strategy optimization is achieved, and detection capabilities and energy utilization efficiency are improved.

CN120498858APending Publication Date: 2025-08-15GUANGXI POWER GRID CORP
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510835937.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Power systems face complex multi-source attack threats in the process of intelligent transformation, and existing security protection technologies have problems such as static defense, data fragmentation, insufficient real-time and lack of adaptive capabilities.

Method used

Build a four-layer dynamic optimization framework of ‘perception-evaluation-decision-execution’, use lightweight edge agents to collect multi-source data at high frequency, perform data fusion through adversarial generation networks, improve threat assessment of spatio-temporal graph convolution networks, combine with deep deterministic policy gradient algorithms to generate dynamic protection strategies, and deploy policy through software-defined networks.

Benefits of technology

The attack detection rate was improved to 98.7%, the false alarm rate was reduced to 2.3%, the strategy optimization time was shortened to 457ms, and the energy loss was reduced by 19.8%, meeting the millisecond-level dynamic response needs of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498858A_ABST
    Figure CN120498858A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic optimization method for an intelligent security convergence protection strategy of a power system, and belongs to the technical field of power system security protection. According to the method, a'perception-evaluation-decision-execution 'four-layer dynamic optimization framework is constructed; a perception layer collects multi-source data at a sampling rate greater than or equal to 1kHz through a lightweight edge agent, and the multi-source data is fused through an adversarial generative network; the evaluation layer evaluates threats by using an improved space-time diagram convolutional network in combination with a dynamic threat index; the decision-making layer introduces a depth deterministic strategy gradient algorithm to generate a protection strategy; and the execution layer deploys and feeds back a result through a software defined network. The problems of static defense, data splitting, insufficient real-time performance and the like in the traditional technology are solved. Simulation verification shows that the attack detection rate is increased to 98.7%, the strategy optimization time consumption is shortened to 457ms, the energy loss is reduced by 19.8%, and the safety protection level of the power system is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power system security protection technology, and specifically to a dynamic optimization method for power system intelligent security convergence protection strategy, which is used to solve the complex multi-source attack threats faced in the intelligent transformation of power systems and realize dynamic optimization of security protection strategies. Background Art

[0002] As the intelligent transformation of power systems accelerates, their network architecture evolves towards cyber-physical systems (CPS), facing threats from complex multi-source attacks such as APT attacks and false data injection attacks. Existing security protection technologies have obvious limitations:

[0003] Static defense mechanism: Traditional firewalls and intrusion detection systems (IDS) rely on predefined rule bases and are difficult to deal with new attack patterns.

[0004] Multi-source data fragmentation: Data from power monitoring and control systems (SCADA), smart meters, and distributed energy resources (DER) have not yet been integrated and analyzed across domains.

[0005] Insufficient real-time performance: The existing security policy update cycle is long (usually >10 minutes), which cannot match the millisecond-level dynamic response requirements of the power grid.

[0006] Lack of adaptive capabilities: There is a lack of a closed-loop optimization mechanism for protection strategies based on the dynamic evolution of attack behaviors.

[0007] Full terms and abbreviations

[0008] CPS: Cyber-Physical System

[0009] IDS: Intrusion Detection System

[0010] SCADA: Supervisory Control And Data Acquisition (SCADA)

[0011] DER: Distributed Energy Resource

[0012] ST-GCN: Spatio-Temporal Graph Convolutional Network

[0013] DTI: Dynamic Threat Index

[0014] DDPG: Deep Deterministic Policy Gradient

[0015] SDN: Software Defined Network

[0016] GAN: Generative Adversarial Network Summary of the Invention

[0017] A dynamic optimization method for intelligent security convergence protection strategy of power system is provided, a dynamic optimization framework is constructed, and efficient dynamic optimization of power system security protection strategy is realized to solve the problems of static defense, data fragmentation, insufficient real-time performance and lack of adaptive capability in traditional technologies.

[0018] Technical Solution

[0019] To achieve the above object, the present invention adopts the following technical solutions:

[0020] In order to effectively deal with the complex security threats faced during the intelligent transformation of power systems, this paper constructs a four-layer dynamic optimization framework of "perception-assessment-decision-execution". Each layer works closely together to achieve efficient dynamic optimization of power system security protection strategies. The architecture diagram is shown below. Figure 1 shown.

[0021] 1. Perception layer

[0022] At this layer, a lightweight edge agent is deployed. This agent has powerful data collection capabilities, with a sampling rate of ≥1kHz, which can achieve millisecond-level data collection. The reason for choosing a lightweight design is to quickly collect various key data in the power system without adding too much system burden. These data come from a wide range of sources, covering multiple fields such as power monitoring systems (SCADA), smart meters, and distributed energy resources (DER). By collecting this data in real time, the perception layer provides a rich and timely information basis for subsequent analysis and decision-making, ensuring that the system can quickly capture potential security threat signals. The perception layer plays a key role in data collection and initial threat perception in the dynamic optimization method of intelligent security convergence protection strategy for power systems. Its implementation process is as follows:

[0023] (1) Edge Agent Deployment: Lightweight edge agents are deployed at key nodes in the power system, such as substations, distributed energy access points, and smart meter concentration areas. These locations are widely distributed and can cover different areas of the power system, ensuring comprehensive collection of various data. Lightweight designs are selected, and low-power, high-performance chips and devices are used in hardware selection. In software, the code structure is optimized and unnecessary functional modules are streamlined to reduce the use of computing and storage resources of the node devices and avoid increasing the system burden.

[0024] (2) Data collection startup: After deployment, the edge agent starts the data collection program. Its built-in collection module collects data at high frequency from the connected data source at a set sampling rate (≥1kHz), that is, at least 1000 times per second.

[0025] (3) Multi-source data collection: Real-time power system operating status data, such as voltage, current, and power, is obtained from the power monitoring and control system (SCADA); user electricity consumption data, including power consumption and time of use, is collected from smart meters; and power generation data, such as solar panel power generation and wind turbine power generation, is collected from distributed energy resources (DER). The edge agent establishes a stable data transmission link with these data sources, which can be connected through a wired network (such as Ethernet) or a wireless network (such as 4G / 5G, Wi-Fi) to ensure the reliability and real-time performance of data transmission.

[0026] (4) Data preprocessing and caching: Not all collected data can be used directly. There may be missing data, errors, or inconsistent formats. The data processing module of the edge agent preprocesses the collected data, using data filling and error correction algorithms to fill missing and erroneous data and unify the data format. The processed data is first temporarily stored in the local cache, and high-speed flash memory can be used as a cache medium to facilitate subsequent rapid reading and transmission.

[0027] Multi-source data fusion: The pre-processed data is fused using a multi-source data fusion model. Define a unified representation function for heterogeneous data:

[0028]

[0029] Among them, w i is the adaptive weight (dynamically adjusted by the Generative Adversarial Network (GAN)), λ is the topological correlation factor, E topo is the grid topology adjacency matrix. This function transforms multi-source data of different types and structures into a unified representation, providing more valuable data support for subsequent threat analysis and protection strategy generation.

[0030] The Generative Adversarial Network (GAN) structure is as follows:

[0031] GAN is mainly composed of a generator (G) and a discriminator (D). The generator G receives a random noise vector z as input, which is usually sampled from a standard normal distribution N(0,1). Its output is a set of weight vectors This weight vector will be used to weight different data sources in the multi-source data fusion formula.

[0032] The discriminator D has two input channels. One channel receives the weight vector output by the generator G The other channel receives the actual weight vector w real The true weight vector can be determined by evaluating and analyzing the fusion effect of historical multi-source data, using expert experience or traditional machine learning methods. The output of the discriminator D is a scalar value representing the probability that the input weight vector comes from the true data, with a value range between [0, 1].

[0033] The following is a schematic diagram of the corresponding GAN structure, such as Figure 2 shown.

[0034] Adversarial Generative Network Training Process:

[0035] The training process of GAN is a process of alternating training of the generator and the discriminator. The specific steps are as follows:

[0036] 1) Initialization parameters: Randomly initialize the network parameters θ of the generator G and the discriminator D G and θ D .

[0037] 2) Training the discriminator D:

[0038] Sample a batch of real weight vectors from the real dataset

[0039] · Sample a batch of noise vectors {z 1 ,z 2 ,…,z m}, get the generated weight vector through generator G

[0040] The discriminator D calculates the discrimination results between the real weight vector and the generated weight vector, and uses the binary cross entropy loss function to measure the performance of the discriminator:

[0041]

[0042] Update the parameters θ of the discriminator D through the back-propagation algorithm D , to minimize the loss function LD .

[0043] 3) Training the generator G:

[0044] · Sample a new batch of noise vectors {z 1 ,z 2 ,…,z m}, get the generated weight vector through generator G

[0045] The goal of the generator G is to make the discriminator D mistakenly classify the generated weight vector as real, using the following loss function:

[0046] Update the parameters θ of the generator G through the backpropagation algorithm G , to minimize the loss function L G .

[0047] 4) Iterative training: Repeat steps 2 and 3 multiple times until the weight vector generated by the generator G can well simulate the distribution of the real weight vector, making it difficult for the discriminator D to distinguish between the generated weight vector and the real weight vector. At this point, the weight vector output by the generator G is It can be used in multi-source data fusion formulas.

[0048] Data Transmission and Sharing: Following the established data transmission strategy, the edge agent sends pre-processed data to the assessment layer. A scheduled transmission mechanism can be set to package and transmit cached data at regular intervals (e.g., millisecond intervals). Event-triggered transmission can also be implemented, with data immediately sent upon detecting an unusual data change. Once the data is transmitted to the assessment layer via the network, it provides data support for subsequent threat propagation model analysis and protection strategy generation, enabling the system to quickly identify potential security threats and make effective decisions based on this data.

[0049] 2. Evaluation layer

[0050] In the dynamic optimization system of intelligent security convergence protection strategy for power systems, data interaction between layers follows specific protocol rules. These rules ensure efficient and accurate data transmission between the perception layer and the evaluation layer, providing a solid foundation for the evaluation layer to build a threat propagation model based on the improved spatiotemporal graph convolutional network (ST-GCN).

[0051] ·Construction of data interaction assurance model

[0052] The perception layer is responsible for acquiring data from multiple sources, including the power monitoring and control system (SCADA), smart meters, and distributed energy resources (DER). After data preprocessing and multi-source data fusion, the data is encapsulated in the Protobuf format. This format offers fast serialization and deserialization speeds and a small data size, making it suitable for millisecond-level data transmission requirements in power systems.

[0053] The perception layer transmits data to the evaluation layer at a fixed frequency of 10 milliseconds. This frequency matches the perception layer's sampling rate of ≥1kHz, ensuring that the evaluation layer receives the latest data in a timely manner. To ensure data transmission reliability, a data loss retransmission mechanism and data verification mechanism based on the TCP / IP protocol are implemented. If the perception layer does not receive an acknowledgment (ACK) from the evaluation layer within 50 milliseconds after sending data, the data is deemed lost and resent. Furthermore, a CRC code is calculated before data transmission. Upon receiving the data, the evaluation layer recalculates and compares the data. If any inconsistency occurs, a retransmission is requested to ensure data accuracy.

[0054] Improved ST-GCN model plays a role

[0055] The evaluation layer receives and parses Protobuf-formatted data from the perception layer. Traditional graph convolutional networks (GCNs) have limitations when processing graph-structured data, making it difficult to fully analyze the complex characteristics of power system data. The improved ST-GCN fully considers the spatiotemporal characteristics of power system data. Data in power systems not only exhibits spatial correlations, such as the electrical connections between different nodes, but also temporal dynamics, such as fluctuations in power load over time.

[0056] ST-GCN is capable of deeply mining this accurately transmitted and processed spatiotemporal data. Based on the unified representation of grid topology and timestamps received by the assessment layer, it accurately analyzes potential threats in the current system and predicts their propagation paths and impact range within the power system. This model provides a scientific basis for subsequent decision-makers to formulate effective and rational defense strategies.

[0057] To more intuitively demonstrate the performance advantages of the improved ST-GCN, a comparative experiment was conducted with the traditional GCN. Under the same simulation environment, various complex multi-source attack scenarios, including APT attacks and false data injection attacks, were set up to test the attack detection rate and false alarm rate of the traditional GCN and the improved ST-GCN. The experimental results showed that the attack detection rate of the traditional GCN was only 56.3%, with a false alarm rate as high as 15.7%. However, under the same attack scenario, the improved ST-GCN increased its attack detection rate to 91.2%, a 34.9% improvement compared to the traditional GCN; and reduced its false alarm rate to 5.8%, a decrease of nearly 10 percentage points. It can be clearly seen from these experimental data that the improved ST-GCN has significant advantages in attack detection accuracy and reducing false alarms by adopting K-order spatial convolution to consider multi-hop neighborhood information and using causal convolution to capture time series characteristics. It can more effectively analyze the complex characteristics of power system data, accurately analyze the potential threats in the current system, and predict the propagation path and impact range of threats in the power system, providing a more reliable scientific basis for subsequent decision-makers to formulate reasonable and effective protection strategies.

[0058] The specific implementation process of the evaluation layer is as follows:

[0059] 2.1 Data Preparation and Processing

[0060] Receive data from the perception layer after multi-source data fusion processing. The perception layer deploys a lightweight edge agent (Edge Agent) to achieve millisecond-level data collection with a sampling rate of ≥1kHz. Data sources include power monitoring and control systems (SCADA), smart meters, and distributed energy resources (DER). After preprocessing and multi-source data fusion, the collected data is formed into a data set that can be used for evaluation. Multi-source data fusion uses the following formula:

[0061]

[0062] Among them, w i is the adaptive weight (dynamically adjusted by the adversarial generation network), λ is the topological correlation factor, E topo is the grid topology adjacency matrix.

[0063] 2.2 Building the Graph Structure

[0064] Based on the physical topology of the power system, a graph G = (V, E) is constructed. V is a set of nodes, including power equipment such as generators, transformers, busbars, and lines; E is a set of edges, representing the electrical connection relationship between nodes. Define the adjacency matrix (N is the number of nodes), if nodes i and j are connected, A ij =1, otherwise A ij= 0. To reflect the connection density between nodes, A is normalized to obtain D is the degree matrix,

[0065] 2.3 Improved Spatiotemporal Graph Convolutional Network (ST-GCN) to Construct Threat Propagation Model

[0066] (1) Spatial convolution module

[0067] When processing graph-structured data, traditional graph convolution (GCN) only considers direct neighborhood information, which is difficult to meet the needs of complex spatial relationship analysis in power systems. The improved ST-GCN uses K-order spatial convolution to consider multi-hop neighborhood information. The formula is:

[0068]

[0069] in, is the feature matrix after spatial convolution of the l+1th layer, H (l) is the feature matrix of the l-th layer node, is the weight matrix associated with the k-order neighborhood, and σ is a nonlinear activation function (such as ReLU).

[0070] (2) Temporal Convolution Module

[0071] Power system data has temporal dynamics, and causal convolution is used to capture time series characteristics. Using causal convolution, let the input time series x t , convolution kernel w, causal convolution output n is the convolution kernel size. In ST-GCN, the temporal convolution layer convolves each node feature in the time dimension, and the formula is:

[0072]

[0073] in, is the node feature of the lth layer at time t, It is the time convolution kernel weight, which can capture the changing trend of power load, equipment status, etc. over time.

[0074] (3) Space-time fusion module

[0075] The results of spatial convolution and temporal convolution are fused. Here, a weighted summation method is used. The formula is:

[0076]

[0077] Among them, α∈[0,1] is a weight parameter that can be learned through training to balance the importance of spatiotemporal features.

[0078] 2.4 Real-time Threat Assessment Index Calculation

[0079] The Dynamic Threat Index (DTI) is constructed by comprehensively considering the time change rate of the attack and the matching degree of the attack characteristics. The formula is:

[0080]

[0081] In the formula, α, β are environmental sensitivity coefficients, Represents the degree of attack signature matching for k types. The environmental sensitivity coefficients α and β are used to balance the impact of the time change rate of the attack and the degree of attack signature matching on DTI. The following details the logic for determining their values:

[0082] Determine the initial value based on historical data regression

[0083] First, collect a large amount of operation data of the power system in the past period of time, including the attack probability P at different times. attack , attack characteristics And the corresponding actual threat level assessment value DTI real The actual threat level assessment value can be obtained through comprehensive evaluation by experts based on factors such as the actual losses after the system is attacked and business interruption.

[0084] The time rate of attack Matching degree with attack signature As independent variables x1 and x2, the actual threat level assessment value DTI real As the dependent variable y, a linear regression model is established:

[0085] y=αx1+βx2+∈

[0086] Where ∈ is the error term.

[0087] Use the least squares method to estimate the regression coefficients α and β so that the predicted value The mean square error between the actual value y Minimum. By solving the following normal equation:

[0088]

[0089] Get initial estimates of α and β.

[0090] Dynamic adjustment of online learning

[0091] As the power system operates in real time, new data is constantly generated. In order to make α and β adapt to the dynamic changes of the system environment, online learning methods are used to adjust them in real time. At each new time step t, the attack probability at the current moment is obtained. Attack characteristics and the latest actual threat level assessment Calculate the predicted value at the current moment And according to the prediction error To update α and β.

[0092] Update using the stochastic gradient descent (SGD) algorithm:

[0093]

[0094] Where η is the learning rate, which controls the step size of each update. Through continuous online learning and adjustment, α and β can reflect changes in the power system environment in real time, allowing DTI to more accurately assess the system's real-time threat status.

[0095] 2.5 Model Training and Optimization

[0096] (1) Loss function definition

[0097] If the model task is to predict the probability of threat occurrence, the cross entropy loss function is used:

[0098]

[0099] If the threat impact range or propagation path is predicted, the mean square error loss function is used:

[0100]

[0101] Where N is the number of samples, y i is the true value, is the predicted value.

[0102] (2) Optimization algorithm

[0103] Use the Adam optimization algorithm to update the model parameters θ, the formula is:

[0104]

[0105] Where η is the learning rate, m t is the first-order moment estimate, v t is the second-order moment estimate, and ∈ is a small constant to prevent the denominator from being zero.

[0106] 2.6 Threat Analysis and Prediction

[0107] Real-time power system data is fed into a trained model and combined with the Dynamic Threat Index (DTI) to analyze whether potential threats exist in the current system. If a potential threat is detected, the threat propagation path and impact range are predicted based on the power system diagram structure and learned associations. These predictions provide a scientific basis for decision-makers to formulate protection strategies.

[0108] 3. Decision-making level

[0109] The decision-making layer is primarily responsible for generating dynamic protection strategies based on the threat analysis and prediction results provided by the assessment layer. This process is closely integrated with the dynamic optimization objective function in the core algorithm model to achieve a comprehensive consideration of response time and operating costs while meeting relevant constraints to ensure the safe and stable operation of the power system. The following is a detailed implementation of the decision-making layer:

[0110] 3.1 Receive evaluation layer results

[0111] The decision-making layer first receives information from the assessment layer, including potential threat analysis, threat propagation path and impact range predictions derived from a threat propagation model built using an improved spatiotemporal graph convolutional network (ST-GCN), and the Dynamic Threat Index (DTI) calculated from real-time threat assessment metrics. This information provides a key basis for the subsequent formulation of protection strategies.

[0112] 3.2 Introducing the Deep Deterministic Policy Gradient (DDPG) algorithm

[0113] To generate dynamic protection strategies tailored to the real-time security status of power systems, the Deep Deterministic Policy Gradient (DDPG) algorithm is introduced at the decision-making level. As a policy gradient-based reinforcement learning algorithm, DDPG demonstrates significant advantages in decision-making problems in continuous action spaces. In power system security protection scenarios, the formulation of protection strategies requires dynamic selection from a multitude of feasible strategies, a requirement that falls squarely within the continuous action space decision-making problem, and DDPG is well-suited to this requirement.

[0114] In this power system security protection system, the DDPG algorithm uses the output of the assessment layer as environmental state input. Using an improved spatiotemporal graph convolutional network (ST-GCN), the assessment layer constructs a threat propagation model and calculates real-time threat assessment metrics, such as the dynamic threat index (DTI). This provides the DDPG algorithm with multi-dimensional information, including potential system threats, threat propagation paths, and impact ranges. The DDPG algorithm converts this information into a description of the environmental state, which it then learns and explores through continuous interaction with the environment.

[0115] Specifically, the DDPG algorithm consists of a policy network and a value network. The policy network generates specific defense strategy actions based on the current environment, while the value network evaluates the value of these actions in the current environment. During operation, the algorithm continuously interacts with the environment and collects reward feedback to adjust the parameters of the policy and value networks.

[0116] The reward function design here is closely integrated with the actual needs and constraints of the power system, and the formula is:

[0117]

[0118] Among them, γ1, γ2 and γ3 are weight coefficients used to balance the importance of different indicators in the reward function. response Indicates the time it takes for the system to respond to threats, T max is the maximum acceptable response time threshold, This item is used to encourage the system to respond to threats quickly. The shorter the response time, the higher the reward. defense Indicates the comprehensive defense effect of all defense means, I attack Indicates the impact of the attack. It reflects the ratio of defense effectiveness to attack impact. The larger the ratio, the better the defense effect and the higher the reward. It also indicates the probability of violating the service level agreement. Pr (SLAvionlation) is used to punish SLA violations. The lower the violation probability, the higher the reward. This ensures that decision-makers fully consider SLA constraints when formulating protection strategies.

[0119] During each interaction, the policy network generates a protective action and applies it to the power system. Based on the effectiveness of the action, the system generates a reward signal, reflecting the overall performance of the action in ensuring system security, meeting service-level agreements (SLAs), and controlling operating costs. The DDPG algorithm uses policy gradients and value assessment methods based on the reward signal to update network parameters.

[0120] To verify the convergence of the DDPG algorithm during training, multiple experiments were conducted. During the experiments, the changes in the parameters of the policy and value networks, as well as the changes in the reward values, were recorded. The training results show that with increasing training cycles, the parameters of the policy and value networks gradually stabilized, and the reward values gradually converged. For example, in one set of experiments, after 500 training cycles, the reward values remained stable at around 0.85 (the specific values varied across different experimental settings, but the overall trend remained consistent), with fluctuations less than 0.05. This demonstrates that the DDPG algorithm can effectively learn and gradually find a near-optimal protection strategy in this power system security protection scenario. Through continuous optimization, the policy network can gradually learn the optimal protection strategy under different environmental conditions, thereby achieving efficient and intelligent decision-making in power system security protection scenarios.

[0121] 3.3 Definition of dynamic optimization objective function

[0122] The decision-making process is guided by minimizing the following dynamic optimization objective function:

[0123] minγ1T response +γ2C operation

[0124] Among them, γ1 and γ2 are weight coefficients used to balance the response time Tresponse and operating costs C operation Two goals. Response time T response Measures the time it takes for the system to respond to threats, the operating cost C operation Indicates the resource costs required to implement protection strategies, such as computing resources, network bandwidth, and equipment maintenance.

[0125] 3.4 Considering Constraints

[0126] While optimizing the objective function, the decision-making process needs to meet the following constraints:

[0127] Service Level Agreement (SLA) constraints:

[0128] Pr(SLA violation)≤∈

[0129] This constraint ensures that the probability of violating the service-level agreement (SLA) is less than or equal to a given threshold ∈. SLAs typically specify the service levels that power systems should achieve in terms of availability and performance, such as the system's uptime percentage and data transmission latency. When formulating protection strategies, decision-makers must ensure that their implementation does not excessively impact the system's service quality, thereby meeting SLA requirements.

[0130] Defense effect constraints:

[0131]

[0132] in, It represents the defensive effect of the j-th defensive measure, such as the degree of preventing the attack, the degree of reducing the impact of the attack, etc.; m is the total number of defensive measures; I attack represents the impact of the attack, such as power loss or equipment damage caused by the attack; η is the defense effectiveness threshold. This constraint requires that the combined effectiveness of all defense measures must be at least η times the impact of the attack to ensure that the system can effectively respond to the attack.

[0133] 3.5 Strategy Generation Process

[0134] The DDPG algorithm continuously updates the parameters of the policy network and value network by interacting with the environment (i.e., the threat situation of the power system). The specific steps are as follows:

[0135] (1) Initialization: Initialize the policy network μ(s|θ μ ) and the value network Q(s,a|θ Q ) parameter θ μ and θ Q , and the target policy network μ'(s|θ μ′ ) and the target value network Q'(s,a|θ Q) parameter θ μ′ and θ Q′ , and initialize the parameters of the target network to be the same as the main network. At the same time, initialize the experience replay buffer R.

[0136] (2) Sampling: At each time step t, according to the current environment state s t , through the policy network μ(s|θ μ ) Generate an action a t , and add some noise for exploration. Execute action a t After that, the environment returns to the next state s t+1 and reward r t . The quadruple (s t ,a t ,r t ,s t+1 ) is stored in the experience replay buffer R.

[0137] (3) Learning: Randomly sample a small batch of samples from the experience replay buffer R For each sample, calculate the target value y i :

[0138] y i =r i +γQ'(s i+1 ,μ'(s i+1 |θ μ' )|θ Q' )

[0139] Where γ is the discount factor. Then, the mean squared error loss function is used to update the parameters θ of the value network Q :

[0140]

[0141] At the same time, the policy network parameters θ are updated using policy gradients μ :

[0142]

[0143] (4) Soft update target network: Regularly use the soft update method to update the parameters of the target policy network and the target value network:

[0144] θ μ' ←τθ μ +(1-τ)θ μ'

[0145] θ Q' ←τθ Q +(1-τ)θ Q'

[0146] Here, τ is a small hyperparameter that controls the speed at which the target network is updated.

[0147] 3.6 Strategy Selection and Output

[0148] At each decision moment, the trained policy network generates an optimal protection strategy based on the current environmental state. This strategy takes into account the dynamic optimization objective function and constraints, aiming to minimize response time and operational costs while ensuring system security and service quality. The generated protection strategy is then sent to the execution layer for implementation.

[0149] Through the above steps, the decision-making layer uses the formulas in the core algorithm model and the DDPG algorithm to achieve dynamic optimization of the power system security protection strategy, providing effective decision-making support to ensure the stable operation of the power system.

[0150] 4. Execution layer

[0151] The execution layer's primary task is to rapidly deploy the dynamic protection strategies generated by the decision layer across the entire network using software-defined networking (SDN), ensuring that the strategies take effect promptly to address power system security threats. This process is related to the dynamic optimization objective function and constraints in the core algorithm model. The following details its implementation:

[0152] 4.1 Receiving Decision-Making Strategy

[0153] The execution layer receives the dynamic protection strategy generated by the decision layer based on the deep deterministic policy gradient (DDPG) algorithm. The strategy is that the decision layer minimizes the objective function minγ1T response +γ2C operation Guided by the constraints Pr(SLA violation)≤∈ and Here, γ1 and γ2 are weight coefficients that balance the response time T response and operating costs C operation ; ∈ is the threshold of the service level agreement (SLA) violation probability; η is the defense effectiveness threshold.

[0154] 4.2 Strategy Analysis and Conversion

[0155] After receiving the policy, the execution layer first parses it. Because the policy generated by the decision layer may be an abstract representation based on an algorithm, the execution layer needs to convert it into instructions that the SDN controller can understand and execute. For example, a policy may require adjusting access rights to certain nodes to prevent a certain type of attack. The execution layer translates this requirement into specific flow table rules within the SDN, specifying matching conditions such as source address, destination address, and port number, as well as corresponding actions (e.g., allow, deny, forward, etc.).

[0156] 4.3SDN Controller Configuration

[0157] The execution layer communicates with the SDN controller and sends the converted policy instructions to the controller. The SDN controller is the core of the entire software-defined network, responsible for centralized management and control of network devices. Based on policy requirements, the execution layer instructs the controller to configure network devices such as switches and routers. This may involve updating flow tables, adjusting bandwidth allocation, and setting security access control lists.

[0158] 4.4 Network-wide deployment and real-time monitoring

[0159] After receiving the command, the SDN controller distributes the configuration information to each network device via a southbound interface (such as the OpenFlow protocol), implementing network-wide policy deployment. The execution layer monitors the deployment process in real time to ensure that the policy is fully configured across the entire network within minutes. During the deployment process, the execution layer collects feedback from network devices to verify the success of the configuration. If a device configuration fails, the execution layer promptly troubleshoots and retries to ensure policy consistency and effectiveness.

[0160] 4.5 Association with the core algorithm model and feedback

[0161] During the policy execution process, the execution layer continuously monitors the impact of the policy on the system, collects relevant data and feeds it back to the decision layer. This data includes the actual response time T response , operating costs C operation , SLA violations and defense effectiveness Etc. Based on this feedback information and in combination with the formulas in the core algorithm model, the decision layer can evaluate and adjust the objective function and constraints in real time. For example, if the actual operation cost is found to be too high, the decision layer can appropriately adjust the value of γ2 and regenerate a more optimized strategy. At the same time, if the SLA violation probability is close to or exceeds the threshold, or the defense effect is not satisfactory, According to the requirements, the decision-making layer will adjust the strategy in time to ensure the security and reliability of the system.

[0162] 4.6 Dynamic Strategy Adjustment and Optimization

[0163] The operating status and threats facing power systems are constantly changing. The execution layer needs to collaborate with the decision-making layer to dynamically adjust and optimize protection strategies based on real-time monitoring and feedback. When new threats emerge or the system's operating status changes significantly, the execution layer can quickly respond to the decision-making layer's updated strategies, re-analyzing, converting, and deploying them to ensure the power system remains secure and stable.

[0164] Through the above steps, the execution layer leverages the advantages of software-defined networks to efficiently deploy the strategies generated by the decision-making layer based on the core algorithm model to the entire network. Through real-time monitoring and feedback mechanisms, it dynamically adjusts and optimizes the strategies, providing strong guarantees for the safety protection of the power system.

[0165] Beneficial effects

[0166] The four-layer dynamic optimization framework of "perception-assessment-decision-execution" constructed by this invention effectively solves the security issues faced in the intelligent transformation of power systems and has multiple beneficial effects:

[0167] 1. Improved attack detection capabilities: The attack detection rate has increased to 98.7%, a 41% improvement compared to traditional methods. The lightweight edge agent in the perception layer collects multi-source data at high frequency, and the improved spatiotemporal graph convolutional network in the evaluation layer deeply analyzes this data, accurately identifying potential attacks.

[0168] 2. Reduced false alarm rate: The false alarm rate has been reduced to 2.3%. Multi-source data fusion and precise threat assessment indicator calculation reduce the possibility of normal data being misidentified as attacks, improving the reliability of detection results and avoiding resource waste.

[0169] 3. Reduced strategy optimization time: Strategy optimization time has been reduced from minutes to sub-seconds, averaging 457ms. The decision-making layer introduces the Deep Deterministic Policy Gradient (DDPG) algorithm, combining dynamic optimization objective functions and constraints to rapidly generate protection strategies that meet the millisecond-level dynamic response requirements of the power system.

[0170] 4. Reduced energy loss: Energy loss was reduced by 19.8%. By optimizing the start-up and shutdown strategies for protective equipment, the execution and decision-making layers collaborate to dynamically adjust strategies based on the real-time status of the system, preventing irrational operation of protective equipment, improving energy efficiency, and reducing operating costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0171] Figure 1 Schematic diagram of the dynamic optimization framework of the intelligent safety convergence protection strategy for the power system of the present invention.

[0172] Figure 2 This is the structure diagram of the Generative Adversarial Network (GAN). DETAILED DESCRIPTION

[0173] The present invention will be further described below in conjunction with specific embodiments:

[0174] This paper comprehensively verifies the proposed dynamic optimization method for intelligent power system security convergence protection strategy using the OPNET / Matlab co-simulation environment. The goal is to accurately evaluate the performance of this method in practical application scenarios. The following is a detailed implementation process and analysis of the results:

[0175] (1) Simulation environment construction

[0176] In the OPNET and Matlab co-simulation platform, a highly realistic power system model is constructed based on the actual power system topology, equipment parameters, and operating characteristics. This model includes key components such as the power monitoring and control system (SCADA), smart meters, and distributed energy resources (DER), and simulates their data interaction and operating logic.

[0177] When setting up various complex, multi-source attack scenarios, we simulated APT attacks targeting network vulnerabilities in areas where smart meters are concentrated within the power system. First, hackers exploited known vulnerabilities in the smart meter communication protocol to infect some smart meters with malware. These infected smart meters became the initial penetration nodes, where the malware lurked inside and collected meter data. The malware then exploited the communication links between the smart meters and the substation, attempting to breach the substation's network defenses and gradually infiltrate the power monitoring and control (SCADA) system. During this penetration process, the malware continuously scanned the power system's internal network, searching for additional exploitable vulnerabilities and attackable nodes to expand its attack range.

[0178] For false data injection attacks, data injection is targeted at distributed energy resource (DER) access points. Specifically, attackers compromise the DER access point's communication module and tamper with the power generation data it sends to the grid. For example, during the transmission of solar panel power data, the attacker can modify the actual power generation data, significantly increasing or decreasing the reported power generation value. This interferes with the grid's assessment of the energy supply and demand balance, impacting the normal scheduling and operation of the power system.

[0179] (2) Performance indicator setting and evaluation method

[0180] Attack Detection Rate: Calculate the attack detection rate by comparing the number of attacks detected by the present invention and traditional security protection technologies under the same attack scenario. Attack Detection Rate = (Number of Detected Attacks / Number of Actual Attacks) × 100%.

[0181] False alarm rate: Counts the number of normal data events that the system mistakenly identifies as attacks during the simulation process. False alarm rate = (number of false alarms / (number of detected attacks + number of false alarms)) × 100%.

[0182] Policy Optimization Time: Records the time from threat detection to the generation and implementation of optimized protection policies, accurately measuring the real-time nature of policy optimization.

[0183] Energy loss: Compare the energy consumption data before and after the adoption of the invention's optimized protective equipment start-stop strategy, and calculate the energy loss reduction ratio. Energy loss reduction ratio = ((traditional strategy energy loss - invention strategy energy loss) / traditional strategy energy loss) × 100%.

[0184] (3) Simulation hardware configuration and mapping relationship description

[0185] This simulation experiment was conducted on a high-performance computer with the following hardware configuration: The CPU was an Intel Xeon Platinum 8380 processor with 64 physical cores, a base frequency of 2.3 GHz, and a turbo frequency of up to 3.8 GHz, capable of meeting the computational demands of complex simulation models. The GPU was an NVIDIA A100 80GB, whose powerful graphics processing capabilities helped accelerate data processing and model calculations during the simulation. The high-speed, large-capacity 512GB DDR4 memory ensured rapid reading, writing, and storage of large amounts of data during the simulation.

[0186] In actual power systems, the computing power of the CPU can be compared to the processing power of data processing units (such as intelligent monitoring terminals and protection devices) in substations. These devices need to process large amounts of power data and control instructions in real time. The parallel computing capability of the GPU is similar to the collaborative processing capabilities of distributed computing nodes in power systems, playing an important role in large-scale data processing and complex algorithm calculations. Memory capacity is similar to the local cache capacity of each node device in the power system, used for temporary storage and rapid access to critical data to ensure efficient operation of the system. Through this mapping relationship, the results obtained in the simulation environment can, to a certain extent, reflect the performance of the actual power system under similar circumstances, providing a reference basis for practical applications.

[0187] (4) Simulation results and analysis

[0188] Attack Detection Rate: Simulations have shown that the attack detection rate of the proposed method has increased to 98.7%, a 41% improvement compared to traditional methods. This demonstrates that the perception and assessment layers constructed by the proposed method, through efficient data collection by lightweight edge agents and in-depth analysis of the improved spatiotemporal graph convolutional network (ST-GCN) threat propagation model, can more accurately and comprehensively identify potential attacks, significantly enhancing the power system's ability to detect complex attacks.

[0189] False alarm rate: The false alarm rate has been reduced to 2.3%. Thanks to multi-source data fusion processing and precise threat assessment indicator calculation, the misidentification of normal data as attacks has been effectively reduced, improving the reliability of detection results and avoiding unnecessary protection operations and resource waste caused by false alarms.

[0190] Policy Optimization Time: Policy optimization time has been significantly reduced from minutes to sub-seconds, averaging only 457 milliseconds. The Deep Deterministic Policy Gradient (DDPG) algorithm, introduced at the decision-making layer, combines dynamic optimization objectives and constraints to rapidly generate protection strategies, meeting the power system's millisecond-level dynamic response requirements and ensuring timely and effective protection measures when threats occur.

[0191] Energy loss: By optimizing the start-up and shutdown strategies for protective equipment, energy loss was reduced by 19.8%. Close collaboration between the executive and decision-making layers dynamically adjusted the protection strategy based on the system's real-time status, avoiding excessive operation or unreasonable start-up and shutdown of protective equipment. This effectively improved energy efficiency and reduced system operating costs.

[0192] In summary, through OPNET / Matlab joint simulation verification, the dynamic optimization method of the power system intelligent security convergence protection strategy of the present invention has shown significant advantages in attack detection, false alarm control, real-time strategy optimization, and energy loss reduction. It has good practical application value and can provide strong security protection for the intelligent transformation of the power system.

[0193] The embodiments of the present invention are not limited to the above description. The edge agent deployment location, ST-GCN model parameters and DDPG algorithm strategy can be adjusted according to the actual power system operation scenario. Such improvements fall within the scope of protection of the present invention.

Claims

1. A method for dynamic optimization of intelligent security convergence protection strategy for power system, characterized in that: The following steps are involved: The perception layer collects multi-source data of the power system at a sampling rate of ≥1kHz through a lightweight edge agent, and generates unified representation data after preprocessing and multi-source data fusion. The multi-source data fusion formula is: w i is the adaptive weight dynamically adjusted by the adversarial generation network, λ is the topological correlation factor, and E topo is the grid topology adjacency matrix; The evaluation layer uses the improved spatiotemporal graph convolutional network ST-GCN to build a threat propagation model and combines it with the dynamic threat index DTI to perform real-time threat assessment. The ST-GCN includes a spatial convolution module, a temporal convolution module, and a spatiotemporal fusion module. The calculation formula of the dynamic threat index DTI is: α and β are environmental sensitivity coefficients; The decision layer introduces the deep deterministic policy gradient DDPG algorithm, based on the dynamic optimization objective function minγ1T response +γ2C operation and constraints to generate a protection strategy, wherein the constraints include a service level agreement constraint Pr(SLAviolation)≤∈ and a defense effect constraint The execution layer converts protection strategies into executable instructions for the SDN controller through software-defined networking (SDN) and deploys them across the entire network, monitoring the execution of strategies in real time and providing feedback to the decision-making layer.

2. The method for dynamic optimization of intelligent security convergence protection strategy of power system according to claim 1 is characterized in that: The lightweight edge agent in the perception layer is deployed at key nodes of the power system such as substations, distributed energy access points, and smart meter concentration areas, and adopts a lightweight design with low power consumption, high-performance chips and optimized code structure.

3. The method for dynamic optimization of intelligent security convergence protection strategy of power system according to claim 1 is characterized in that: The improved ST-GCN in the evaluation layer adopts K-order spatial convolution to consider multi-hop neighborhood information, and the formula is Use causal convolution to capture time series features, the formula is Spatiotemporal fusion is achieved through weighted summation, and the formula is:

4. The method for dynamic optimization of intelligent security convergence protection strategy of power system according to claim 1 is characterized in that: The reward function of the DDPG algorithm in the decision layer is γ1, γ2, and γ3 are weight coefficients.

5. The method for dynamic optimization of intelligent security convergence protection strategy of power system according to claim 1 is characterized in that: The policy parsing and conversion in the execution layer converts the protection policy into flow table rules that can be executed by the SDN controller, including matching conditions such as source address, destination address, port number, and corresponding actions.

6. The method for dynamic optimization of intelligent security convergence protection strategy of power system according to claim 1 is characterized in that: It also includes key technological innovations: Joint spatiotemporal modeling: embedding grid physical equations into neural networks to predict cross-domain attack impacts. Dynamic strategy generation and an online optimizer based on deep reinforcement learning shorten the strategy update cycle to 30 seconds. Adaptive weight mechanism, using dual attention mechanism to dynamically adjust protection focus; Lightweight deployment and layered policy compression algorithm reduce the size of policy instructions by 78%.

Citation Information

Cited By

  • Network security threat detection method and device, equipment and storage medium

    CN121125348A

  • Network security threat detection method, device, equipment and storage medium

    CN121125348B