Simulation attack-defense method for power grid load frequency control system with DRL controller
Patent Information
- Application Number
- CN202610296664.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-12
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-03-12
AI Technical Summary
这种方法需要生成海量的攻击样本来“打补丁”,训练成本极高,并且可能以牺牲控制器在正常工况下的精度和响应速度为代价(即鲁棒性与性能的权衡)
[0024] In this application, the proposed simulated attack method only perturbs independent physical quantities and derives the perturbations of dependent quantities. The generated attack vectors are consistent with the perturbations in the real physical world, making the evaluation results more valuable. Furthermore, because it targets only key features, the attack efficiency is higher. The proposed simulated defense method monitors the direct result of the attack behavior (Q-value decrease) rather than indirect physical manifestations (frequency fluctuations), thus achieving more accurate detection and faster response. The provided controller switching strategy is a plug-and-play active defense that eliminates the need for expensive model retraining. The switching mechanism ensures a safety lower bound while maximizing the preservation of the DRL controller's high performance under normal operating conditions, achieving a balance between safety and performance.
Smart Images

Figure CN122225434B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the fields of power system security and artificial intelligence security, and specifically relates to a simulation attack and defense method for a power grid load frequency control system using a DRL controller. Background Technology
[0002] In modern power systems, load frequency control (LFC) is a crucial element in ensuring power quality and the safe and stable operation of the system. With the expansion of the power grid and the increase in the proportion of renewable energy integration, the system exhibits strong nonlinearity and uncertainty, making it difficult for traditional proportional-integral-derivative (PID) controllers to meet high-standard control requirements. Deep reinforcement learning (DRL), as an advanced intelligent control method, shows great potential in the LFC field due to its ability to operate without precise models, driven by data, and with adaptive learning capabilities. In practical deployments, DRL controllers often do not completely replace traditional controllers but rather work in conjunction with existing PID controllers, forming a "DRL-PID hybrid control architecture," which will be a long-term feature in the technological evolution process.
[0003] However, this hybrid architecture introduces new security challenges. At the heart of the DRL controller is a deep neural network, which is highly sensitive to minute perturbations in its input data. Attackers can exploit this vulnerability to design sophisticated adversarial attacks: injecting minute perturbations, imperceptible to the human eye, into the DRL controller's observed data (such as system frequency and tie-line power), thereby tricking the DRL controller into outputting erroneous control commands that could even lead to system instability. Even more dangerously, in a hybrid architecture, frequency fluctuations caused by an attack on the DRL controller in one area can propagate through the tie-lines to adjacent areas controlled by PID controllers, potentially triggering a chain reaction and causing widespread system instability.
[0004] Most existing research on adversarial attacks has not fully considered the strict physical coupling relationships between various state variables of the power system (such as the area control error ACE and frequency deviation). Power deviation of connecting lines The resulting attack (which may not conform to the laws of physics) is difficult to reproduce in reality or easily detected by traditional data verification mechanisms.
[0005] In terms of defense, the mainstream technique is "passive defense," such as enhancing the robustness of the model through adversarial training. This method requires generating massive amounts of attack samples to "patch" the system, resulting in extremely high training costs, and may come at the expense of the controller's accuracy and response speed under normal operating conditions (i.e., a trade-off between robustness and performance). Furthermore, for the specific risks of the DRL-PID hybrid architecture, there is currently a lack of an efficient, low-cost "active defense" strategy that does not affect normal performance.
[0006] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0007] To address or at least mitigate one or more of the above problems, a simulated attack and defense method for a power grid load frequency control system employing a DRL controller is provided.
[0008] To achieve the above objectives, according to the first aspect of this application, a simulation attack and defense method for a power grid load frequency control system employing a DRL controller is provided, comprising:
[0009] Conduct a simulated attack:
[0010] S1. Identify the system state variables that are most critical to the DRL controller's decisions:
[0011] Among them, the value network or policy network of the DRL controller is used to filter out key attack features;
[0012] S2. Set a simulated attack target and, based on the target's actions, calculate the initial perturbation amount targeting the key attack characteristics:
[0013] The simulated attack target is the attack target action that causes the DRL controller to output the worst decision;
[0014] S3. The initial perturbation amount is subjected to amplitude limiting processing to generate the final perturbation amount finally applied to the key attack feature;
[0015] S4. Based on the state-space equation and physical definition of the LFC system, the perturbation amount of the other dependent state features in the LFC system is obtained through the perturbation amount of the key attack feature.
[0016] Wherein, the dependent state feature is a state feature that changes with the key attack feature;
[0017] S5. Combine the final perturbation amount of the key attack feature with the perturbation amount of the dependent state feature to form a complete adversarial perturbation vector that conforms to physical constraints, and superimpose the adversarial perturbation vector onto the real-time state observation value of the DRL controller at the current moment to implement a simulated attack.
[0018] Based on the simulated attack, a simulated defense is performed:
[0019] Applied to a hybrid load frequency control system, the system comprising at least one deep reinforcement learning DRL controller and one backup proportional-integral-derivative PID controller, the defense methods include:
[0020] A1. In the operating mode where the system is primarily controlled by the DRL controller, the action value output by the evaluation network of the DRL controller is monitored and recorded in real time. Value, Formation Value time series;
[0021] A2. Using the Cumulative Sum (CUSUM) detection algorithm, the above... Perform statistical analysis on the time series values to determine the... Does the time series value exhibit an anomalous drift that consistently falls below the normal baseline?
[0022] A3. When the statistics of the CUSUM detection algorithm exceed the preset alarm threshold, it is determined that the hybrid system is under adversarial attack, and the control of the system is immediately switched from the DRL controller to the backup PID controller to interrupt the attack chain and maintain system stability.
[0023] By adopting the above technical solution, this application has the following beneficial effects compared with the prior art:
[0024] In this application, the proposed simulated attack method only perturbs independent physical quantities and derives the perturbations of dependent quantities. The generated attack vectors are consistent with the perturbations in the real physical world, making the evaluation results more valuable. Furthermore, because it targets only key features, the attack efficiency is higher. The proposed simulated defense method monitors the direct result of the attack behavior (Q-value decrease) rather than indirect physical manifestations (frequency fluctuations), thus achieving more accurate detection and faster response. The provided controller switching strategy is a plug-and-play active defense that eliminates the need for expensive model retraining. The switching mechanism ensures a safety lower bound while maximizing the preservation of the DRL controller's high performance under normal operating conditions, achieving a balance between safety and performance.
[0025] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. Attached Figure Description
[0026] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application. The illustrative embodiments and descriptions of the application are used to explain the application, but do not constitute an undue limitation of the application. Obviously, the drawings described below are merely some embodiments, and those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0027] In the attached diagram:
[0028] Figure 1 This is a flowchart illustrating the simulated attack and defense method of the power grid load frequency control system using a DRL controller in this specific embodiment.
[0029] Figure 2 This is a schematic diagram of three interconnected LFC system regions in this specific embodiment;
[0030] Figure 3 This is a comparison chart showing the system frequency deviation caused by attack methods targeting only key features and attack methods targeting all features in this specific embodiment.
[0031] Figure 4 This is a comparison chart of the detection accuracy of the Q-value CUSUM detector under different attack methods and different control architectures in this specific embodiment;
[0032] Figure 5 This is a comparison chart of the frequency deviation curves of region 1 before and after enabling the defense method of the present invention under the single-agent architecture in this specific embodiment. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0034] Please see Figure 1 This application provides a simulation attack and defense method for a power grid load frequency control system employing a DRL controller, comprising:
[0035] Conduct a simulated attack:
[0036] S1. Identify the system state variables that are most critical to the DRL controller's decisions:
[0037] Among them, the value network or policy network of the DRL controller is used to filter out key attack features;
[0038] S2. Set a simulated attack target and, based on the target's actions, calculate the initial perturbation amount targeting the key attack characteristics:
[0039] The simulated attack target is the attack target action that causes the DRL controller to output the worst decision;
[0040] S3. The initial perturbation amount is subjected to amplitude limiting processing to generate the final perturbation amount finally applied to the key attack feature;
[0041] S4. Based on the state-space equation and physical definition of the LFC system, the perturbation amount of the other dependent state features in the LFC system is obtained through the perturbation amount of the key attack feature.
[0042] Wherein, the dependent state feature is a state feature that changes with the key attack feature;
[0043] S5. Combine the final perturbation amount of the key attack feature with the perturbation amount of the dependent state feature to form a complete adversarial perturbation vector that conforms to physical constraints, and superimpose the adversarial perturbation vector onto the real-time state observation value of the DRL controller at the current moment to implement a simulated attack.
[0044] Based on the simulated attack, a simulated defense is performed:
[0045] Applied to a hybrid load frequency control system, the system comprising at least one deep reinforcement learning DRL controller and one backup proportional-integral-derivative PID controller, the defense methods include:
[0046] A1. In the operating mode where the system is primarily controlled by the DRL controller, the action value output by the evaluation network of the DRL controller is monitored and recorded in real time. Value, Formation Value time series;
[0047] A2. Using the Cumulative Sum (CUSUM) detection algorithm, the above... Perform statistical analysis on the time series values to determine the... A3. When the statistics of the CUSUM detection algorithm exceed the preset alarm threshold, it is determined that the hybrid system is under adversarial attack, and the control of the system is immediately switched from the DRL controller to the backup PID controller to interrupt the attack chain and maintain system stability.
[0048] The background of this invention is a multi-area interconnected LFC power system model. This model is well-known to those skilled in the art and is used to describe and analyze the dynamic response of a power system under load disturbances. The model will be described in detail below. A multi-area interconnected LFC system consists of multiple control areas interconnected via tie lines, such as... Figure 2 As shown.
[0049] Each region contains core components such as generators, speed governors, and steam turbines, forming a secondary control closed loop. For the i-th region in the system (i takes values of 1, 2, 3, corresponding to...), ... Figure 2 The dynamic behavior of regions 1, 2, and 3 in the equation can be described by a set of linearized differential equations:
[0050] ;
[0051] ;
[0052] ;
[0053] ;
[0054] Where, Δf i The frequency deviation of region i; This refers to the deviation in the mechanical power output of the generator. This refers to the load disturbance deviation within the region; This represents the net power deviation exchanged between region i and other regions via tie lines; the left side of the equals sign represents the derivative of the corresponding parameter with respect to time, i.e., the rate of change. D is the equivalent rotational inertia constant of region i; i Let be the load damping coefficient for region i; For governor valve position deviation; The time constant of the steam turbine; These are control commands issued by the LFC controller (i.e., the DRL or PID controller in this invention); This is the droop coefficient (or speed variation rate) of the speed governor. The time constant of the speed controller; Let be the tie-line synchronization coefficient between region i and region j. In LFC, the direct input signal to the controller is the region control error (ACE), which combines tie-line power deviation and frequency deviation.
[0055] No. The regional control error for each control area is:
[0056] ;
[0057] in, Let be the frequency offset coefficient for region i, and its calculation method is as follows:
[0058] ;
[0059] For ease of analysis and controller design, the above system of differential equations can be integrated into a standard linear continuous-time state-space form:
[0060] ;
[0061] ;
[0062] The vectors and matrices are defined as follows: State vector : =[ , , , ]ᵀ;Control Input : = Disturbance input : = Measurement output : =[ , ACE i System matrix Input matrix Perturbation matrix and output matrix The specific form is as follows:
[0063] ;
[0064] ;
[0065] ;
[0066] ;
[0067] Since the DRL controller in this invention is a digital controller that makes decisions in discrete time steps, the continuous-time model described above needs to be discretized using the zero-order hold (ZOH) method, with a sampling period of T. The discretized state-space model is shown below:
[0068] ;
[0069] Where t represents the discrete time step, and A, B, C are derived from the continuous system matrix. , , The discrete time constant matrix derived from the sampling period T. The DRL controller of this invention is based on this discretized model, using state vectors. Or, using the observations it contains as input, output control commands. .
[0070] Implementation and verification of physical constraint adversarial attack methods:
[0071] This embodiment elaborates on the specific technical details of the attack method of the present invention, and verifies its effectiveness through simulation experiments.
[0072] Step S1: Screening of key attack features
[0073] The purpose of this step is to identify the state features that have the greatest impact on the DRL controller's decisions.
[0074] Defining the loss function: First, we need a metric to measure how bad the decision is. We utilize the Critic Network in DRL. It can assess the state Take action below Good or bad. We define the "worst-case target action". In the current state The action with the lowest value is shown in formula (1). Then, the loss function is defined. For the current strategy and The cross-entropy is given by formula (2):
[0075] The larger the loss L is, the closer the current policy output is to the worst-case scenario.
[0076] Specifically, the simulated attack targets the evaluation network of the DRL controller. What is obtained, in the current state The lowest Value action The method of determination is as follows:
[0077] (1);
[0078] Where s is the current state vector, and A is the set of all possible actions. To evaluate the network's performance on state-action pairs under the current policy π The evaluation value;
[0079] The initial disturbance is calculated based on a current policy. and stated The cross-entropy is the loss function. Generated using gradient descent, it is represented as:
[0080] (2);
[0081] (3);
[0082] in, For the current policy network based on state The output action; For loss function State The gradient; Represents the cross-entropy loss function; It is a symbolic function; The perturbation step size; The initial disturbance vector is calculated.
[0083] Calculate gradient significance: Calculate loss For each input state feature gradient The larger the absolute value of the gradient, the better the understanding of the gradient. Applying a small perturbation effectively increases the loss L, meaning it leads the controller to make worse decisions. By accumulating and normalizing the absolute values of the gradients over a period of time, as shown in Equation (4), we obtain the final importance score for each feature. :
[0084] (4);
[0085] Experimental results show that the regional frequency deviation and its derivative, area control error ACE and its derivative, tie line power deviation Its derivative is the highest-scoring feature. Among these features, and It is an independent physical quantity, and therefore was selected as a key attack feature.
[0086] Step S2: Calculation of initial disturbance:
[0087] We use the Fast Signed Gradient Method (FGSM) to generate the initial perturbation vector. As shown in formula (4). The negative sign indicates that we want to reduce the loss, that is, to reduce the controller output. Maximize the probability. It's the perturbation step size, controlling the attack intensity. From... Extracting key attack features and The perturbation component is used, and the clip function is used to limit it to the range of [-ε, ε] to ensure the stealth of the attack.
[0088] Step S3: Disturbance Constraints and Concealment Processing
[0089] This step is crucial for achieving stealth in the attack. From Extracting key attack features and perturbation components and Then, the clip function is used to constrain it to a preset concealment budget. Within the range, the final attack disturbance amount is obtained. and As shown in formula (5). and It is a preset disturbance budget threshold; the smaller the value, the more covert the attack.
[0090] Specifically, the limiting process is performed through a... The function implementation is as follows:
[0091] (5);
[0092] in, and From the initial perturbation vector Extracted corresponding key attack features and The amount; and It is a preset perturbation budget threshold to ensure the stealth of the attack; and This is the final disturbance amount obtained after amplitude limiting.
[0093] Step S4: Derivation of feature-dependent perturbations
[0094] To ensure physical consistency, the perturbations of other dependent features must be determined by... and This is derived. For example, the perturbation of ACE is calculated according to formula (6). The perturbation quantity of the isodermied derivative characteristic is obtained by first-order backward difference approximation.
[0095] Specifically, the perturbation amount of the dependent state feature The perturbation amount of the key attack feature is based on the physical definition formula of the area control error ACE. and The linear combination derivation yields:
[0096] (6);
[0097] in, This is the regional bias coefficient. The governor droop coefficient is used; other derivative-dependent state-feature disturbances are obtained by first-order backward difference approximation of the basic feature disturbances.
[0098] Step S5: Simulate Attack Implementation
[0099] The final perturbations of all independent and dependent features are combined into an adversarial perturbation vector. This is then superimposed onto the actual observed value s to form adversarial examples. The input is given to the DRL controller.
[0100] Simulated attack effect verification:
[0101] To verify the effectiveness of the attack method of this invention, a comparative experiment was conducted under the full agent DRL architecture. Figure 3This is a comparison chart showing the system frequency deviation caused by the attack method targeting only key features and the attack method targeting all features, as described in this invention. From... Figure 3 As can be seen, the two curves represent the frequency deviation of region 2 under the two attack methods, while the other two curves represent the frequency deviation of regions 1 and 3. The key finding is that the frequency response curve caused by attacking only key features is highly consistent with the response curve caused by attacking all features in terms of dynamic pattern and deviation magnitude. This proves that step S1 (key feature selection) described in this invention is efficient and reasonable. Attackers do not need to perturb all states; by precisely attacking only a few key features, they can achieve almost the same destabilizing attack effect on the system.
[0102] To further verify the effectiveness of the attack method of this invention in different deployment scenarios, it was applied to three hybrid control architectures: single agent, partial agent, and full agent. Table 1 compares the key performance indicators of different architectures before and after being attacked by this invention, including maximum absolute frequency deviation (MAFD), integral absolute error (IAE), and root mean square error (RMSE).
[0103] Table 1: Performance Comparison of Different Control Architectures Before and After Being Attacked by This Invention
[0104]
[0105] As can be seen from the data in Table 1, the attack method of this invention poses a serious threat to all architectures. Particularly in the single-agent architecture, an attack on the DRL controller in region 1 causes a 315% deterioration in its MAFD metric, and this disturbance is rapidly propagated through the flyway to regions 2 and 3, which are controlled by PID controllers, leading to significant and sustained oscillations throughout the system. In the partial-agent architecture, the system vulnerability is most pronounced, with the IAE metric in region 1 deteriorating by more than 10 times (1094%), indicating that multiple independent DRL controllers may be more vulnerable to attack due to a lack of cooperation.
[0106] Figure 3Comparative curves of frequency deviation over time in three regions under a fully intelligent agent DRL control architecture, employing both critical feature attacks and full feature attacks, are presented. In the simulation, the attack was initiated at 240 s and continued for a period. It can be observed that once the attack began, the frequency deviations in the three regions changed almost instantaneously, with virtually no significant lag in the system response, indicating that the DRL controller is extremely sensitive to the observed input. Under both attack methods, the frequency deviation in region 2 increased and peaked near approximately 0.0025 pu, while the frequency deviations in regions 1 and 3 decreased to approximately -0.003 pu. Throughout the simulation period of 0–1400 s, the frequency response curves corresponding to the critical feature attack and the full feature attack showed a high degree of consistency in dynamic shape and amplitude, almost completely overlapping. This indicates that the deterioration in system stability almost entirely stems from the disturbance of a few critical vulnerable features, making the critical feature-based attack strategy both reasonable and efficient.
[0107] Active defense methods:
[0108] This embodiment details the specific technical aspects of the active defense method of the present invention and verifies its effectiveness through simulation experiments.
[0109] Steps A1 & A2: Q-value-based CUSUM anomaly detection
[0110] The core of this step is to perform continuous statistical tests on the Q-value sequence of the DRL controller within a sliding, fixed-size time window.
[0111] Q-value centralization: First, the mean of the Q-values is calculated on a large amount of attack-free data, serving as the normal baseline μ_Q. In real-time detection, we focus on a time window of size N preceding the current time t, i.e. For each Q value within this window Centralization is performed, as in formula (7).
[0112] Bidirectional CUSUM detection: To comprehensively monitor Q-value anomalies, this method simultaneously calculates both upward and downward cumulative sums to detect whether the Q-value is consistently too high or too low. The calculation formulas are shown in (8) and (9). It is an upward cumulative sum statistic used to detect a sustained positive drift in the Q value; These are downward cumulative sum statistics used to detect persistent negative drift in Q-values; their initial values... and All are 0; It is a preset positive number, called the drift or relaxation parameter, used to ignore random noise fluctuations when there is no attack to prevent false alarms.
[0113] Specifically, the specific calculation formula of the CUSUM detection algorithm is as follows:
[0114] First, within the sliding time window of value Centralized processing:
[0115] (7);
[0116] Then, the upward cumulative sum statistic is calculated simultaneously. With downward cumulative sum statistic :
[0117] (8);
[0118] (9);
[0119] in, The current moment; The size of the sliding time window; The discrete time step within the time window; for real time value; for The value is the baseline mean for normal operation; For centralized value; These are drift parameters; and initial value and All are 0;
[0120] when or Any one of them exceeds the alarm threshold At that time, it was determined that there was an abnormal drift.
[0121] Specifically, the attack method described in this invention aims to induce the controller to select actions with low Q values, thus primarily leading to The rapid accumulation of [something]. However, a robust detection system should simultaneously monitor [something]. This is to prevent other attack types that may cause abnormally high Q values.
[0122] Step A3: Controller Switching
[0123] We pre-set an alarm threshold h. At each time k, if... >h or >h, the system immediately determines that it has been attacked and executes the switch of the controller from DRL to PID.
[0124] Verification of defensive effectiveness:
[0125] Detection performance verification:
[0126] To verify the effectiveness of step A2 (CUSUM detection) in the defense method of the present invention, various gradient-based attacks were detected and tested in three different architectures. Figure 4 This is a comparison chart showing the detection accuracy of the Q-value CUSUM detector described in this invention under different attack methods and control architectures. From... Figure 4 As can be seen, the accuracy of this detector generally reaches over 90% across all architectures, regardless of whether it targets the attack proposed in this invention, PGD attacks, or Critic attacks. This demonstrates that this invention, by monitoring the Q-value, a core internal indicator, can reliably and accurately identify adversarial attacks.
[0127] Performance verification of relief and recovery:
[0128] To verify the effectiveness of step A3 (controller switching) in the defense method of this invention, the active defense method described in this invention was activated in a simulated attack scenario. Table 2 shows the various performance indicators of the system when attacked after activating the defense method of this invention.
[0129] Table 2: System performance after enabling the active defense method of the present invention
[0130]
[0131] Comparing the data in Tables 1 and 2, the superiority of the defense method of this invention can be clearly seen. Taking a single-agent architecture as an example, after enabling the defense, the IAE index decreased from 0.352606 to 0.188019, representing a performance improvement of 46.7%. Figure 5 This is a comparison of the frequency deviation curves in region 1 before and after enabling the defense method of this invention under a single-agent architecture. Figure 5 As can be clearly seen, after enabling the defense (blue curve), the fluctuations in system frequency during the attack (240s-600s) were significantly suppressed. More importantly, without the defense, it takes approximately 300 seconds for the system to gradually recover and stabilize after the attack stops. However, after enabling the defense method of this invention, because the attack is promptly cut off, the system can almost instantly return to normal after the attack stops, and the recovery performance is fundamentally improved.
[0132] In summary, through detailed embodiments and quantified experimental data, this invention demonstrates the effectiveness of its proposed physical constraint attack method and the superiority of its proposed active defense method.
[0133] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although this application has disclosed preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-mentioned technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. The implementation schemes in the above embodiments can also be further combined or replaced. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the content of the technical solution of this application shall still fall within the scope of this application.
Claims
1. A simulation attack and defense method for a power grid load frequency control system using a DRL controller, characterized in that, include: Conduct a simulated attack: S1. Identify the most critical system state variables for the decisions of the deep reinforcement learning DRL controller: Among them, the value network or policy network of the DRL controller is used to filter out key attack features; S2. Set a simulated attack target and calculate the initial perturbation amount for the key attack characteristics based on the target's actions: The simulated attack target is the attack target action that causes the DRL controller to output the worst decision; S3. The initial perturbation amount is subjected to amplitude limiting processing to generate the final perturbation amount finally applied to the key attack feature; S4. Based on the state-space equation and physical definition of the LFC system for grid load frequency control, the disturbance quantities of other dependent state features in the LFC system are obtained through the disturbance quantities of the key attack features. Wherein, the dependent state feature is a state feature that changes with the key attack feature; S5. Combine the final perturbation amount of the key attack feature with the perturbation amount of the dependent state feature to form a complete adversarial perturbation vector that conforms to physical constraints, and superimpose the adversarial perturbation vector onto the real-time state observation value of the DRL controller at the current moment to implement a simulated attack. Based on the simulated attack, a simulated defense is performed: Applied to a hybrid load frequency control system, the system comprising at least one deep reinforcement learning DRL controller and one backup proportional-integral-derivative PID controller, the defense methods include: A1. In the operating mode where the system is primarily controlled by the DRL controller, the action value output by the evaluation network of the DRL controller is monitored and recorded in real time. Value, Formation Value time series; A2. Using the Cumulative Sum (CUSUM) detection algorithm, the above... Perform statistical analysis on the time series values to determine the... Does the time series value exhibit an anomalous drift that consistently falls below the normal baseline? A3. When the statistics of the CUSUM detection algorithm exceed the preset alarm threshold, it is determined that the hybrid system is under adversarial attack, and the control of the system is immediately switched from the DRL controller to the backup PID controller to interrupt the attack chain and maintain system stability.
2. The method according to claim 1, characterized in that, The most critical system state variables for the decision-making of the deep reinforcement learning (DRL) controller include: Establish the state space of the load frequency control system, which contains N state characteristics; Based on the value network or policy network of the deep reinforcement learning DRL controller, the importance score of each of the N state features to the decision of the DRL controller is calculated by gradient saliency analysis. Select a predetermined number of independent physical state features with the highest scores as key attack features.
3. The method according to claim 1, characterized in that, The key attack characteristic is regional frequency deviation. Inter-regional tie line power deviation .
4. The method according to claim 2, characterized in that, The step of setting a simulated attack target and calculating the initial perturbation amount for the key attack characteristics based on the attack target's actions includes: The simulated attack targets the evaluation network of the DRL controller. What is obtained, in the current state The lowest Value action The method of determination is as follows: ; in, Let A be the current state vector, and A be the set of all possible actions. To evaluate the network's performance on state-action pairs under the current policy π The evaluation value; The initial disturbance is calculated based on a current policy. and stated The cross-entropy is the loss function. Generated using gradient descent, it is represented as: ; ; in, For the current policy network based on state The output action; For loss function State The gradient; Represents the cross-entropy loss function; It is a symbolic function; The perturbation step size; The initial disturbance vector is calculated.
5. The method according to claim 3, characterized in that, The limiting processing is achieved through a The function implementation is as follows: ; ; in, and From the initial perturbation vector Extracted corresponding key attack features and The amount; and It is a preset perturbation budget threshold to ensure the stealth of the attack; and This is the final disturbance amount obtained after amplitude limiting.
6. The method according to claim 5, characterized in that, The perturbation amount of the dependent state feature The perturbation amount of the key attack feature is based on the physical definition formula of the area control error ACE. and Linear combination calculations yielded: ; in, This is the regional bias coefficient. The governor droop coefficient is used. Other derivatives are derived from the disturbances of state characteristics by performing a first-order backward difference approximation on the disturbances of the basic characteristics.
7. The method according to claim 1, characterized in that, The specific calculation formula for the CUSUM detection algorithm is as follows: First, within the sliding time window of value Centralized processing: ; Then, the upward cumulative sum statistic is calculated simultaneously. With downward cumulative sum statistic : ; ; in, The current moment; The size of the sliding time window; The discrete time step within the time window; for real time value; for The value is the baseline mean for normal operation; For centralized value; These are drift parameters; and initial value and All are 0. , They are respectively The cumulative sum of time points upwards and the cumulative sum of time points downwards; when or Any one of them exceeds the alarm threshold At that time, it was determined that there was an abnormal drift.
8. The method according to claim 7, characterized in that, The method further includes step A4: after the LFC system switches to the PID controller, the system status is continuously monitored. When the attack threat is detected to have been eliminated and the system frequency has returned to stability, the control is returned from the PID controller to the DRL controller to restore the optimal control performance of the LFC system.
Citation Information
Patent Citations
Optimal load frequency control method for suppressing resonance attack of power system
CN118199104A
Load frequency control method and device based on deep reinforcement learning, and electronic equipment
CN119109079A