A method and device for training a security and economy collaborative optimization confrontation defense model of a power distribution network

By generating a set of fake attack scenarios and processing the topology using the GCNSAGE feature extractor, and training meta-defense and attack models, the problem of balancing security and economy in distribution network control methods is solved. This achieves proactive defense and rapid adaptation against fake data injection attacks, thereby improving the defense capabilities of the distribution network.

CN121809587BActive Publication Date: 2026-08-04TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-01-05
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing technologies, the control methods for power distribution networks have defects during the training process, resulting in the inability to fully control the power distribution network, making it difficult to simultaneously ensure both security and economy. Furthermore, the defense modes against spoofed data injection attacks are mostly passive responses, which are insufficient to cope with the challenges of uncertainty in renewable energy and changes in topology.

Method used

By generating a set of pseudo-attack scenarios, we train meta-defense and meta-attack models, optimize the overall performance of the defense and attack models using an objective function, introduce attack scenarios of different types and intensities to enhance the model's adaptability, and process topological information through the GCNSAGE feature extractor to construct an SA-MOMG model for active defense.

Benefits of technology

It improves the security and economy of power distribution network defense, can promptly identify and defend against fake data injection attacks, enhances the training efficiency and adversarial capabilities of models, and ensures rapid adaptation in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809587B_ABST
    Figure CN121809587B_ABST
Patent Text Reader

Abstract

This disclosure provides a training method and apparatus for an adversarial defense model for coordinated optimization of distribution network security and economy, which can be applied to the field of distribution system optimization and scheduling technology. The method includes: obtaining a trained meta-defense model and a trained meta-attack model; performing the following operations based on an objective function until the combined performance fluctuation of the obtained defense model and attack model meets preset conditions. The objective function is constructed based on the security operation indicators and economic operation indicators of the sample distribution network: training the meta-defense model in the first state of the sample distribution network to obtain an intermediate defense model; training the meta-attack model based on attack samples in the attack sample pool to obtain an attack model; updating the attack sample pool to obtain an updated attack sample pool; obtaining the defense model; and when the combined performance fluctuation of the defense model and attack model does not meet the preset conditions, using the defense model as the meta-defense model and the attack model as the meta-attack model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of power distribution system optimization and scheduling technology, and more specifically, to a method and apparatus for training an adversarial defense model for collaborative optimization of power distribution network security and economy. Background Technology

[0002] Renewable energy sources such as solar and wind power can be converted into electricity and connected to the distribution network through distributed photovoltaic (PV) and wind power equipment. However, renewable energy sources are inherently uncertain, causing random fluctuations on both the source and load sides of the distribution network, which affects its operational stability. Therefore, optimization of distribution network control is necessary. However, existing control methods for distribution networks still have shortcomings during the training process, resulting in control methods that cannot comprehensively control the distribution network. Summary of the Invention

[0003] In view of this, the present disclosure provides a method and apparatus for training an adversarial defense model for the coordinated optimization of power distribution network security and economy.

[0004] One aspect of this disclosure provides a method for training an adversarial defense model for coordinated optimization of distribution network security and economy, comprising: obtaining a trained meta-defense model and a trained meta-attack model using a set of pseudo-attack scenarios generated based on a sample distribution network, wherein the set of pseudo-attack scenarios includes attack scenarios of different types and intensities targeting the sample distribution network; performing the following operations based on an objective function until the combined performance fluctuation of the obtained defense model and attack model meets a preset condition, wherein the objective function is constructed based on the security operation indicators and economic operation indicators of the sample distribution network; training the meta-defense model in a first state of the sample distribution network to obtain an intermediate defense model, wherein the first state represents the sample distribution network as having no... Attack execution status; with the intermediate defense parameters of the intermediate defense model fixed, a meta-attack model is trained based on the attack samples in the attack sample pool to obtain the attack model. The attack samples in the attack sample pool are generated based on the pseudo-attack scenarios in the pseudo-attack scenario set. The attack sample pool is updated based on the confidence of the action state pairs generated by the intermediate defense model in the case of defending against the attack model. The intermediate defense model is trained with the attack parameters of the attack model fixed to obtain the defense model. If the combined performance fluctuation of the defense model and the attack model does not meet the preset conditions, the defense model is used as the meta-defense model and the attack model is used as the meta-attack model.

[0005] A second aspect of this disclosure provides a training device for a collaborative optimization of the security and economy of a distribution network, comprising: an acquisition module, configured to acquire a trained meta-defense model and a trained meta-attack model using a set of pseudo-attack scenarios generated based on a sample distribution network, wherein the set of pseudo-attack scenarios includes attack scenarios of different types and intensities targeting the sample distribution network; and an execution module, configured to perform the following operations based on an objective function until the combined performance fluctuation of the acquired defense model and attack model meets a preset condition, wherein the objective function is constructed based on the security operation indicators and economic operation indicators of the sample distribution network: training the meta-defense model in a first state of the sample distribution network to obtain an intermediate defense model, wherein the first state table The sample distribution network is in an attack-free operating state. With the intermediate defense parameters of the intermediate defense model fixed, a meta-attack model is trained based on the attack samples in the attack sample pool to obtain the attack model. The attack samples in the attack sample pool are generated based on the pseudo-attack scenarios in the pseudo-attack scenario set. The attack sample pool is updated based on the confidence of the action state pairs generated by the intermediate defense model in the case of defending against the attack model. The intermediate defense model is trained with the attack parameters of the attack model fixed to obtain the defense model. If the combined performance fluctuation of the defense model and the attack model does not meet the preset conditions, the defense model is used as the meta-defense model and the attack model is used as the meta-attack model.

[0006] According to embodiments of this disclosure, different types and intensities of attack scenarios are introduced through a set of pseudo-attack scenarios, making the training data more diverse and enhancing the rapid adaptability of the meta-defense model and meta-attack model to different types and intensities of attack scenarios, thereby improving training efficiency. Furthermore, an intermediate defense model is obtained in the first state; a meta-attack model is trained based on the intermediate defense model to obtain an attack model; and the intermediate defense model is then trained based on the attack model to obtain a defense model. The attack model and defense model are trained separately, allowing them to focus more on updating their own parameters. Moreover, during the intervals between the training of the attack model and defense model, the attack samples in the attack sample pool are updated, increasing the attack capability of the attack samples in the attack sample pool against the sample distribution network, enabling the model to converge more quickly and further improving training efficiency. Furthermore, the objective function includes safety operation indicators and economic operation indicators, enabling the defense model to ensure both security and economic efficiency. The defense model is a model for defending the distribution network under attack conditions, rather than the passive defense mode of "detect first, then correct" in related technologies. Therefore, when the distribution network is controlled through the defense model, security problems of the distribution network can be identified and defended more promptly, which can further improve the security of the distribution network. Attached Figure Description

[0007] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0008] Figure 1 An exemplary system architecture is shown in the embodiments of the present disclosure for training a counter-defense model that can be applied to the collaborative optimization of distribution network security and economy.

[0009] Figure 2 A flowchart is shown below illustrating a method for training an adversarial defense model for coordinated optimization of distribution network security and economy according to an embodiment of this disclosure;

[0010] Figure 3 The process of feature processing of the topology by the policy network and evaluation network according to embodiments of the present disclosure is illustrated.

[0011] Figure 4 A flowchart is shown for a method of training an adversarial defense model for coordinated optimization of distribution network security and economy according to yet another embodiment of the present disclosure;

[0012] Figure 5 A schematic diagram of a sample distribution network topology according to an embodiment of the present disclosure is shown in one example;

[0013] Figure 6 The diagram illustrates the performance of the training methods of this disclosure, methods 1, 2, and 3, in an attack-free environment.

[0014] Figure 7 The diagram illustrates the performance of the training methods of this disclosure, including methods 1, 2, and 3, under an attack environment.

[0015] Figure 8 A schematic diagram of the node voltage distribution before a defensive model obtained using the training method of this disclosure controls a distribution network is shown in one example.

[0016] Figure 9 A schematic diagram of the node voltage distribution after a defense model obtained using the training method of this disclosure is controlled in a distribution network is shown in one example.

[0017] Figure 10 A schematic diagram of 24-hour voltage distribution under different cases according to embodiments of the present disclosure is shown; and

[0018] Figure 11 A block diagram of a training apparatus for a counter-defense model of distribution network security and economic collaborative optimization according to an embodiment of the present disclosure is shown. Detailed Implementation

[0019] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0022] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0023] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security and network security.

[0024] Power dispatching in distribution networks faces multiple challenges. On the one hand, uncertainties on both the power source and load sides can easily lead to the accumulation of real-time dispatching deviations, affecting the overall operational efficiency of the distribution network. On the other hand, in the context of the integration of distribution networks and cyber-physical systems, threats such as False Data Injection Attacks (FDIA) disrupt the state estimation of the distribution network by tampering with its load data and topology, leading to erroneous operations and causing problems such as dispatching errors, line overloads, and voltage exceeding limits.

[0025] In related technologies, research on FDIA mainly focuses on detection methods and prevention and location, while research on distribution network optimization and scheduling is relatively insufficient.

[0026] First, distribution network optimization and dispatching in FDIA scenarios often employs single-objective optimization, making it difficult to simultaneously consider both the security and economy of the distribution network. Regarding defense mechanisms, related technologies are mostly focused on post-event detection, lacking proactive preventative defense strategies. Some FDIA detection methods can only respond after an attack on the distribution network has occurred. FDIA and the uncertainty of renewable energy in the distribution network can also have synergistic negative effects. For example, FDIA falsifying load data or photovoltaic forecast curves can further amplify dispatch deviations caused by renewable energy fluctuations, triggering cascading safety problems such as line overload and voltage exceeding limits. The passive defense mode of "detect first, then correct" in related technologies is insufficient to address such issues. Furthermore, the impact of FDIA on the distribution network is mainly reflected in two aspects: economic dispatch and security control. By tampering with the power generation equipment output plans or load forecast data of the distribution network, FDIA causes the distribution network to misjudge constraints such as load balance, power generation equipment operation, and grid security, thereby increasing the energy consumption of the distribution network and the losses of related power generation equipment. Therefore, the optimized dispatching of the distribution network requires not only economic efficiency but also enhanced security. Thus, rapid identification and defense against FDIA are necessary to achieve safe and economical coordinated dispatching of the distribution network.

[0027] Secondly, when processing information about the topology of distribution networks, methods based on Deep Reinforcement Learning (DRL) typically treat measurement data as Euclidean space vectors, leading to the loss of topological information. Furthermore, DRL methods exhibit policy instability and insufficient robustness when faced with adversarial attacks.

[0028] Furthermore, FDIA detection methods typically assume that the attacker possesses complete network information, which is inconsistent with reality. Methods based on Graph Convolutional Networks (GCNs) require information on the entire distribution network topology for training, making them ill-suited for real-time topology changes and limiting their ability to handle discrete-continuous mixed actions.

[0029] In view of this, embodiments of this disclosure provide a method for training an adversarial defense model for coordinated optimization of distribution network security and economy, comprising: obtaining a trained meta-defense model and a trained meta-attack model using a set of pseudo-attack scenarios generated based on a sample distribution network, wherein the set of pseudo-attack scenarios includes attack scenarios of different types and intensities targeting the sample distribution network; performing the following operations based on an objective function until the combined performance fluctuation of the obtained defense model and attack model meets a preset condition, wherein the objective function is constructed based on the security operation indicators and economic operation indicators of the sample distribution network; training the meta-defense model in a first state of the sample distribution network to obtain an intermediate defense model, wherein the first state represents the sample distribution network as follows: In the no-attack running state, with the intermediate defense parameters of the intermediate defense model fixed, a meta-attack model is trained based on the attack samples in the attack sample pool to obtain the attack model. The attack samples in the attack sample pool are generated based on the pseudo-attack scenarios in the pseudo-attack scenario set. The attack sample pool is updated based on the confidence of the action-state pairs generated by the intermediate defense model in the case of defending against the attack model. The intermediate defense model is trained with the attack parameters of the attack model fixed to obtain the defense model. If the combined performance fluctuation of the defense model and the attack model does not meet the preset conditions, the defense model is used as the meta-defense model and the attack model is used as the meta-attack model.

[0030] Figure 1 An exemplary system architecture for training a defense model that can be applied to the collaborative optimization of distribution network security and economy, according to embodiments of this disclosure, is shown. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0031] like Figure 1As shown, the system architecture 100 according to this embodiment may include a sample distribution network state 101, an attack node 102, a defense node 103, and an attack action 104. The sample distribution network state 101 can be given a first state to the attack node 102 through state estimation. A set of pseudo-attack scenarios is generated based on the sample distribution network for meta-training to obtain a trained meta-defense model and a trained meta-attack model. Under the constraints of the objective function, the following operations are performed until the combined performance fluctuation of the defense model and the attack model meets the preset conditions: Defense node 103 trains the meta-defense model in the first state of the sample distribution network to obtain the intermediate defense model, where the first state represents the sample distribution network in an attack-free operating state; Attack node 102 trains the meta-attack model based on the attack samples in the attack sample pool while fixing the intermediate defense parameters of the intermediate defense model to obtain the attack model, where the attack samples in the attack sample pool are generated based on the pseudo-attack scenarios in the pseudo-attack scenario set; The attack sample pool is updated based on the confidence of the action state pairs generated by the intermediate defense model in the case of defending against the attack model, to obtain the updated attack sample pool; Defense node 103 trains the intermediate defense model while fixing the attack parameters of the attack model to obtain the defense model; If the combined performance fluctuation of the defense model and the attack model does not meet the preset conditions, the defense model is used as the meta-defense model and the attack model is used as the meta-attack model.

[0032] Figure 2 A flowchart is shown for a method for training an adversarial defense model for coordinated optimization of distribution network security and economy according to an embodiment of the present disclosure.

[0033] like Figure 2 As shown, the method includes operations S210~S220.

[0034] In operation S210, a trained meta-defense model and a trained meta-attack model are obtained by using a set of pseudo-attack scenarios generated based on the sample distribution network.

[0035] In operation S220, based on the objective function, the following operations are performed until the combined performance fluctuation of the obtained defense model and attack model meets the preset conditions. The objective function is constructed based on the safety and economic operation indicators of the sample distribution network: A meta-defense model is trained in the first state of the sample distribution network to obtain an intermediate defense model, where the first state represents a no-attack operation state of the sample distribution network; With the intermediate defense parameters of the intermediate defense model fixed, a meta-attack model is trained based on attack samples in the attack sample pool to obtain an attack model, where the attack samples in the attack sample pool are generated according to pseudo-attack scenarios in the pseudo-attack scenario set; The attack sample pool is updated based on the confidence of the action-state pairs generated by the intermediate defense model in defending against attacks from the attack model, resulting in an updated attack sample pool; With the attack parameters of the attack model fixed, an intermediate defense model is trained to obtain a defense model; If the combined performance fluctuation of the defense model and attack model does not meet the preset conditions, the defense model is used as the meta-defense model, and the attack model is used as the meta-attack model.

[0036] A set of pseudo-attack scenarios can be generated based on the sample distribution network. This set includes attack scenarios of different types and intensities targeting the sample distribution network, covering various types of disturbances. Based on these pseudo-attack scenarios, pre-training can be performed to obtain trained meta-defense and meta-attack models. Both models converge after minor updates within the pseudo-attack scenarios. The meta-defense model is exposed to attack actions during the pre-training phase.

[0037] The first state represents the sample distribution network as operating safely. A safe operating state can be the state of the sample distribution network assuming it has not been attacked. The first state of the sample distribution network can be obtained through state estimation.

[0038] The objective function is constructed based on the safety and economic operation indicators of the sample distribution network. This allows the trained defense model to consider multiple objectives, balancing safety and economy.

[0039] A meta-defense model can be trained in the first state to obtain an intermediate defense model. With the parameters of the intermediate defense model fixed, the parameters of the meta-attack model are updated and optimized based on attack samples from the attack sample pool to obtain the attack model. Specifically, the intermediate defense model can be used for defense, while the meta-attack model is used to attack the sample distribution network. The meta-attack model is then updated based on the reward received by the meta-attack model to obtain the attack model.

[0040] After obtaining the attack model, the attack samples in the attack sample pool can be updated. The intermediate defense model generates corresponding action-state pairs when defending against attacks on the sample distribution network by the meta-attack model. The confidence level of these action-state pairs can be calculated to determine their robustness, i.e., the defense capability of the intermediate defense model against the sample distribution network. Attack samples corresponding to action-state pairs with strong defense capabilities can be processed to update the attack sample pool, resulting in an updated attack sample pool.

[0041] After obtaining the attack model, its parameters can be fixed, and an intermediate defense model can be trained to obtain the defense model. Specifically, the attack model can be used to attack a sample distribution network, the intermediate defense model can be used for defense, and the intermediate defense model can be updated based on the rewards obtained from the defense to obtain the defense model.

[0042] The overall performance fluctuation of the defense model and the attack model can be calculated. If the overall performance fluctuation does not meet the preset conditions, the defense model can be used as the meta-defense model and the attack model can be used as the meta-attack model for continued training. If the preset conditions are met, training can be stopped.

[0043] According to embodiments of this disclosure, different types and intensities of attack scenarios are introduced through a set of pseudo-attack scenarios, making the training data more diverse and enhancing the rapid adaptability of the meta-defense model and meta-attack model to different types and intensities of attack scenarios, thereby improving training efficiency. Furthermore, an intermediate defense model is obtained in the first state; a meta-attack model is trained based on the intermediate defense model to obtain an attack model; and the intermediate defense model is then trained based on the attack model to obtain a defense model. The attack model and defense model are trained separately, allowing them to focus more on updating their own parameters. Moreover, during the intervals between the training of the attack model and defense model, the attack samples in the attack sample pool are updated, increasing the attack capability of the attack samples in the attack sample pool against the sample distribution network, enabling the model to converge more quickly and further improving training efficiency. Furthermore, the objective function includes safety operation indicators and economic operation indicators, enabling the defense model to ensure both security and economic efficiency. The defense model is a model for defending the distribution network under attack conditions, rather than the passive defense mode of "detect first, then correct" in related technologies. Therefore, when the distribution network is controlled through the defense model, security problems of the distribution network can be identified and defended more promptly, which can further improve the security of the distribution network.

[0044] According to embodiments of this disclosure, the objective function includes a safety operation sub-function constructed based on safety operation indicators and an economic operation sub-function constructed based on economic operation indicators. The safety operation sub-function includes at least one of a voltage deviation rate term and a branch load rate term, wherein the voltage deviation rate term is determined based on the voltage amplitude and voltage over-limit boundary value of the nodes in the sample distribution network, and the branch load rate term is determined based on the actual current carrying capacity and upper limit of the current carrying capacity of the branches in the sample distribution network. The economic operation sub-function includes at least one of a grid loss cost term, a curtailment cost term, an energy storage charging and discharging cost term, a static var compensator (SVC) operating cost term, and a switching action cost term, wherein the grid loss cost term is determined based on the grid loss power in the sample distribution network, the curtailment cost term is determined based on the photovoltaic curtailment of the photovoltaic equipment in the sample distribution network, the energy storage charging and discharging cost term is determined based on the energy storage charging and discharging power of the energy storage equipment in the sample distribution network, the SVC operating cost term is determined based on the reactive power of the SVC in the sample distribution network, and the switching action cost term is determined based on the number of switching actions in the sample distribution network.

[0045] To quantify the ability of a sample distribution network to maintain normal operation in the face of attacks or faults, a safe operation sub-function can be constructed based on safe operation indicators. The safe operation sub-function includes a voltage deviation rate term and a branch current carrying rate term, as shown in the following formula.

[0046] (1);

[0047] (2);

[0048] in, This represents the voltage deviation rate corresponding to the voltage deviation rate term. This represents the voltage magnitude at node i at time t. Indicates the voltage exceeding the limit boundary value. This represents the set of nodes in the sample distribution network. This represents the branch current carrying rate corresponding to the branch current carrying rate term. Represents the branch at time t Actual carrying capacity Indicates a branch Maximum capacity, This represents the set of branches.

[0049] Based on the definition of the safe operation sub-function, the safe operation objective of the sample distribution network is to minimize the safe operation sub-function, as shown in the following formula.

[0050] (3);

[0051] in, This indicates a sub-function that runs safely. This indicates the safe operation objective of the safe operation sub-function. This indicates the scheduling cycle of the sample distribution network. This represents the voltage deviation rate at time t. This represents the branch current carrying rate at time t.

[0052] An economic operation sub-function is constructed based on economic operation indicators. The economic operation sub-function includes at least one of the following: grid loss cost, curtailment cost, energy storage charging and discharging cost, static var compensator operating cost, and switching operation cost, as shown in the following formula.

[0053] (4);

[0054] in, Represents the total cost of economic operation. These represent the grid loss cost corresponding to the grid loss cost item, the curtailment cost corresponding to the curtailment cost item, the energy storage charging and discharging cost corresponding to the energy storage charging and discharging cost item, the static var compensator operating cost corresponding to the static var compensator operating cost item, and the switching operation cost corresponding to the switching operation cost item, respectively. These represent the sets of nodes in the sample distribution network that are equipped with photovoltaic devices, energy storage devices, and static var compensators, respectively. , , , They represent in The amount of solar power curtailment from photovoltaic equipment at any given time, the energy storage charging and discharging power of energy storage equipment, the reactive power of static var compensators, and the number of switching operations.

[0055] Based on the definition of the economic operation sub-function, the economic operation objective of the sample distribution network is to minimize the economic operation sub-function, as shown in the following formula.

[0056] (5);

[0057] in, Represents the sub-function of economic operation. This represents the economic operation objective of the economic operation sub-function. This represents the economic operating cost at time t.

[0058] The objective function is constructed based on the safe operation sub-function and the economic operation sub-function, as shown in the following formula.

[0059] (6);

[0060] in, Describe the objective function. This represents the optimization objective of the objective function. This represents the control action vector of controllable devices such as switches, static var compensators, energy storage devices, and photovoltaic devices in the sample distribution network. This represents the state vectors of load demand, photovoltaic power, and energy storage state of charge in the sample distribution network. , This indicates a constraint condition.

[0061] The constraints in the training method of this disclosure may include power flow calculation equality and inequality constraints, as well as capacity and operation constraints of photovoltaic equipment, energy storage equipment, etc. in the sample distribution network, as shown in the following formula.

[0062] (7);

[0063] (8);

[0064] (9);

[0065] (10);

[0066] (11);

[0067] in, , Let represent the active power and reactive power of the photovoltaic device connected to node i at time t, respectively. This represents the charging and discharging power of the energy storage device connected to node i at time t. This represents the reactive power of the static var compensator connected to node i at time t. , Let represent the active power and reactive power of the load at node i at time t, respectively. , Representing branches The conductivity and susceptance, This represents the voltage phase angle between node i and node j. This represents the voltage at node i at time t. Let n represent the voltage at node j at time t, and n represent the number of nodes in the sample distribution network. , Let represent the maximum and minimum voltage values ​​at node i, respectively. Represents the branch at time t The current carrying capacity, Indicates a branch The maximum current carrying capacity, Let represent the topology of the sample distribution network after the switch is operated at time t, and satisfy the radial topology set of the sample distribution network. , Let represent the active power of the photovoltaic equipment connected to node i at time t. This represents the upper limit of active power of the photovoltaic equipment connected to node i. This represents the power capacity of the photovoltaic equipment connected to node i. This represents the reactive power of the static var compensator connected to node i at time t. This represents the maximum reactive power of the static var compensator connected to node i. This represents the charging and discharging power of the energy storage device connected to node i at time t. This indicates the upper limit of the charging and discharging power of the energy storage device connected to node i. , These represent the state of charge and change of the energy storage device connected to node i at time t, respectively. This represents the state of charge of the energy storage device connected to node i at time t+1. , These represent the upper and lower limits of the state of charge (SOC) of the energy storage device connected to node i, respectively. This indicates the capacity of the energy storage device connected to node i. This indicates the time interval for dispatching the sample distribution network.

[0068] The first attack action is subject to the attack constraints of the sample distribution network, which represent the conditions that prevent the first attack action from being detected by the default defense of the sample distribution network.

[0069] The default defense detection of the sample distribution network can be within the safe operating range of the sample distribution network, such as the topology error handling mechanism and bad data detection (BDD) mechanism. If the first attack action is detected by the default defense of the sample distribution network, the first attack action may fail. Attack constraints can be set for the first attack action to ensure its success. To successfully launch a topology attack, the measured values ​​of the modified nodes in the sample distribution network by the first attack action must pass the detection under the topology error handling mechanism and the BDD mechanism. A load redistribution attack (LRA) can be used to inject false power data into the sample distribution network to change the load measurement values, and the change ratio of the false power data must be within a safe range. Based on this, the LRA at time t can be expressed as follows.

[0070] (12);

[0071] (13);

[0072] (14);

[0073] in, This represents the spurious power injected into the sample distribution network at time t. Represents the transfer factor matrix. Represents the load correlation matrix. , These represent the spurious load and the power injection value of the photovoltaic equipment at time t, respectively. This is the transpose of the identity matrix whose diagonal elements are all 1s. , and Let represent the load power, photovoltaic power, and photovoltaic capacity at time t, respectively. This represents the change in load power at time t. This represents the change in power output of the photovoltaic device at time t. The formula (12) indicates the percentage change in the false power data. Formula (13) can enable the first attack action to bypass the state estimation of the sample distribution network and carry out the attack. Formula (14) can make the sum of the false load measurements injected into all nodes equal to zero. Formula (15) can make the power of the false load not exceed the safe range of the normal load to hide the attack.

[0074] Based on the application of robust adversarial reinforcement learning in the field of game theory, the attack and defense process of the sample distribution network in this embodiment can be regarded as two parties in a game, where the attacker is the attacker and the defender is the defender. The attacker's goal is to carry out a state adversarial attack to reduce the defender's reward, thereby reducing the safety and economic operation capability of the sample distribution network. The defense model and attack model in the training method can be described as a two-player discounted zero-sum Markov game (TDZMG), where the attacker and defender compete for limited resources and maximize their own rewards. The extension of TDZMG in a multi-objective game scenario is defined as a State Adversarial Multiobjective Markov Game (SA-MOMG) model to further refine the gains and losses of the attacker and defender.

[0075] The SA-MOMG model can be used with quintuples. express, Representing the state game space, , These represent the action spaces of the defender and the attacker, respectively. Represents the state transition function vector, Represents the reward function vector, The discount rate is indicated. Unlike the Markov Decision Process (MDP) in related technologies, in the SA-MOMG model, the state at time t... Below, attackers utilize attack strategies on the network. Execute attack actions Defenders utilize defense strategy networks Perform defensive actions After that, the state transitions to the next state at time t. The process can be described as The process of obtaining rewards at this time It can be described as , , These are the parameters for the attack policy network and the defense policy network, respectively. The goal of the SA-MOMG model is to learn a policy network that maximizes the total reward.

[0076] (15);

[0077] (16);

[0078] in, This represents the strategy network with the highest expected return. Indicates expected return, Represents the expectation function, Represents any policy, This represents the reward value at time t. This indicates the discount rate.

[0079] Therefore, in an adversarial environment, attackers and defenders learn the strategies shown in the following equation to maximize the total reward of the attack strategy network and the defense strategy network, respectively. .

[0080] (17);

[0081] (18);

[0082] in, This represents the defensive action with the highest expected reward performed by the defender at time t. This represents the expected reward of choosing the weakest defensive action given the attacker's strongest attack action at time t. This represents the attack action with the highest expected reward that the attacker can perform at time t. The expected return of the attacker's strongest attack action chosen under the weakest defensive action at time t.

[0083] Furthermore, before training, the topology of the sample distribution network can be processed. The topology of the sample distribution network is typically represented by non-Euclidean data, which is more difficult to process. Therefore, the topology of the sample distribution network can be represented by a set of nodes and branches and a feature matrix. , where the characteristic matrix Including node feature matrix ( For the number of nodes, (Number of input features for each node) and adjacency matrix The node feature matrix represents the operating state of the sample distribution network, and the adjacency matrix allows each node to aggregate the features of its neighboring nodes. When there is a branch between two nodes, then the elements... It is 1 if it is true, and 0 otherwise. Therefore, The state space of the sample distribution network at time step can be represented as: Node feature matrix It is expressed as the following formula.

[0084] (19);

[0085] in, These represent the active and reactive power of the load, the power of the photovoltaic equipment, and the state of charge of the energy storage equipment at time t, respectively.

[0086] GCN can directly extract features from graph-structured data using its unique convolutional aggregation mechanism, effectively integrating information from neighboring nodes to more accurately represent the operating status of the sample distribution network. However, this is a direct push learning method, requiring all nodes to participate in the training process to obtain an accurate node embedding model. To address the potential changes in the sample distribution network topology caused by FDIA, an inductive learning method, GraphSAGE (Graph Sample and Aggregate), is introduced based on the GCN framework to solve the problem of high computational cost. This model is called the GCNSAGE feature extractor. Its advantage lies in its ability to reduce complexity by sampling and aggregating information from neighboring nodes using GraphSAGE, while also capturing global information through the GCN aggregator.

[0087] Specifically, in this embodiment of the disclosure, the sample distribution network includes multiple nodes, and the node features of the sample distribution network include the node features of each node. For each node: the neighboring nodes of the node are sampled to obtain the local structure of the node; the features of the neighboring nodes in the local structure are aggregated to obtain aggregated features; the initial features of the node are updated based on the aggregated features to obtain the intermediate updated features of the node; the intermediate updated features of the node are updated according to the intermediate updated features of multiple nodes to obtain the node features of the node.

[0088] For each node, a fixed number of neighboring nodes can be randomly sampled from its set of neighboring nodes. This ensures that the sampled neighboring nodes represent the local structure of the sample distribution network at that node, reducing the number of nodes involved in the calculation and lowering the computational cost. Feature aggregation is performed on the neighboring nodes in the local structure to obtain the aggregated features of the node, which can be expressed as follows.

[0089] (20);

[0090] in, This represents an aggregation function, such as a GCN aggregator, a mean aggregator, etc. Let i represent the set of neighboring nodes of node i. Indicates that neighbor node v is at the th Aggregated features obtained by GraphSAGE aggregation of layers Represents the set of neighboring nodes of node i. In the Aggregated features obtained by GraphSAGE aggregation of layers .

[0091] The aggregated features can be concatenated with the initial features of the nodes to obtain the intermediate updated features of the nodes.

[0092] (twenty one);

[0093] in, Indicates the first The intermediate update features of node i after GraphSAGE layer processing. Represents the activation function, such as , Indicates the first Layer weight matrix, and , For cascading operations, Indicates the first The intermediate update features of node i after GraphSAGE layer processing. After layer GraphSAGE aggregation, the intermediate update features of node i are obtained. .

[0094] Based on the intermediate update features of multiple nodes, the GCN global aggregation mechanism can be used to enhance the expressive power of node embedding, thereby updating the intermediate update features of the nodes and obtaining the node features.

[0095] (twenty two);

[0096] in, Indicates that node i has passed through the first... Node features after layer GCN aggregation , Represents the adjacency matrix. Represents the identity matrix. express The weight matrix. After processing by the GCNSAGE feature extractor of the layer, the node features of node i are finally obtained. .

[0097] By combining feature processing and training through the GCNSAGE feature extractor, a joint representation of topology and electrical quantities is achieved, solving the problem of loss of topology information in related technologies.

[0098] Based on the FDIA characteristics of active power distribution networks, a set of pseudo-attack scenarios covering multiple disturbance types is constructed. ,in Let represent the parameterized representation of the k-th type of attack. Specifically, the state attack scenario targeting photovoltaic power output is represented as follows: , Let be the mean of the attack perturbation, and satisfy . , Let Variance be the perturbation variance. For load power constraint attack scenarios, it is represented as... , Let be the mean of the attack perturbation, and satisfy . , This represents the variance of the disturbance. All All of them meet the requirements of formula (12).

[0099] Model-independent meta-learning methods can be used to pre-train both the defender and the attacker to obtain the defender's meta-parameters. and attacker meta-parameters This ensures that the parameters converge with only a few updates under new attack actions, thus obtaining trained meta-attack and meta-defense models. The attacker's meta-parameters are the parameters of the meta-attack model, and the defender's meta-parameters are the parameters of the meta-defense model.

[0100] The attacker's meta-parameter optimization objective is as follows:

[0101] (twenty three);

[0102] in, This represents the loss function of the meta-attack model. This represents the expected outcome of a possible attack instance. These are attack instances from a set of pseudo-attack scenarios, i.e., attack actions in pre-training. These are the parameters of the meta-attack model. For the learning step size, This indicates that the gradient of the parameters of the meta-attack model is calculated.

[0103] The optimization objectives for the defender's meta-parameters are as follows.

[0104] (twenty four);

[0105] In the formula, This represents the loss function of the meta-defense model. These are the parameters of the meta-defense model. This represents the gradient of the parameters of the meta-defense model.

[0106] By incorporating meta-generalization constraints, the equilibrium strategy remains effective in unknown attack scenarios. The modified Nash equilibrium objective is as follows.

[0107] (25);

[0108] (26);

[0109] (27);

[0110] (28);

[0111] in, The generalization coefficient can be 0.1 to 0.2. For the meta-generalization constraint, for both the defender and the attacker... Satisfying equations (27) and (28) respectively, represents the loss function in the new attack scenario. This is a set of untrained attack action distributions. and These are the meta-policies for defenders and attackers, respectively. and The rewards are for the defender and the attacker, respectively.

[0112] According to embodiments of this disclosure, under the condition of fixing the intermediate defense parameters of the intermediate defense model, training a meta-attack model based on attack samples in the attack sample pool to obtain an attack model includes: under the condition of fixing the intermediate defense parameters of the intermediate defense model, using the meta-attack model to determine a first attack action based on attack samples in the attack sample pool; under the defense of the intermediate defense model, using the first attack action to attack a sample distribution network in a first state to obtain a second state of the sample distribution network and an attack reward in the second state; updating the parameters of the meta-attack model according to the attack reward to obtain an updated attack model; and obtaining an attack model if the increase in the attack success rate of the updated attack model is less than a preset attack increase threshold.

[0113] Training the meta-attack model can involve multiple rounds, with each round yielding an updated attack model. Training can be stopped based on the improvement in attack success rate of the updated model compared to the previous round. An attack improvement threshold can be set; if the improvement in attack success rate is less than this threshold (i.e., the improvement in the updated attack model is not significant), training can be stopped, and the updated attack model from the final round can be used as the attack model.

[0114] During the training process of the meta-attack model, attacks on the sample distribution network can be achieved by modifying the data within the sample distribution network. The first state of the sample distribution network can be obtained, and based on this first state, a first attack action can be derived. This first attack action could be, for example, interrupting transmission lines in the sample distribution network and then changing the power of photovoltaic devices or loads on those transmission lines. The meta-attack model can derive the first attack action based on the following assumptions: 1. Obtaining the topology of the sample distribution network, allowing for state estimation to acquire its state; 2. The ability to obtain status data on circuit breakers from the dispatch center by interrupting transmission lines in the sample distribution network topology and altering non-critical transmission lines; 3. Understanding historical data on load measurement and photovoltaic power generation dispatch, with load forecasting error being a normally distributed random variable.

[0115] The intermediate defense model can defend against the first attack action. For example, if the first attack action is to increase the power of photovoltaic equipment, the intermediate defense model can defend against it by reducing the power of photovoltaic equipment. Based on the intermediate defense model's defense, the second state of the sample distribution network and the attack reward in the second state are obtained. By determining the first attack action, the attack of the meta-attack model on the sample power supply network can be further refined.

[0116] An attacker's actions can include power fluctuations targeting photovoltaic (PV) devices and power fluctuations targeting the load. Since both power fluctuations targeting PV devices and power fluctuations targeting the load involve continuous power changes, the attacker's action at time t is... ,in, This represents the power change action performed on the photovoltaic equipment at time t. This represents the power variability action performed on the load at time t. The attacker's view of the state at time t. The attack is a direct adversarial attack, rather than directly affecting the sample power distribution network environment. The attacker's state after the attack can be represented as... Therefore, at time t, the attacker can modify the apparent power injected by the node by changing the output power of the photovoltaic equipment and the load, i.e. This disclosure also incorporates actual physical constraints on the attack actions to ensure that the FDIA mechanism is satisfied, as shown in the following formula.

[0117] (29);

[0118] in, This represents the average value calculation, where the attacker's actions satisfy formula (12). ,otherwise, .

[0119] Figure 3 The process of feature processing of the topology by the policy network and evaluation network according to embodiments of the present disclosure is illustrated.

[0120] like Figure 3 As shown, the GCNSAGE feature extractor can be combined with the SAC (Soft Actor and Critic) method, and this method is called GS-SAC. The policy network and the evaluation network can each process the topology features using the GCNSAGE feature extractor, allowing them to better process node features. The policy network predicts the next action at the current time step, while the evaluation network predicts the action after that at the next time step, avoiding biases in the policy network. The meta-attack model can include a meta-attack policy network and a meta-attack evaluation network. The intermediate defense model can include an intermediate defense policy network and an intermediate defense evaluation network.

[0121] The output of the policy network can be represented by a Gaussian distribution.

[0122] (30);

[0123] in, This indicates the action to be executed based on the output of the policy network in the state at time t. and Let represent the mean and variance of the output actions of the policy network, respectively. This indicates a Gaussian distribution.

[0124] According to embodiments of this disclosure, updating the parameters of the meta-attack model based on the attack reward to obtain the updated attack model may include: using a meta-attack evaluation network to determine a second attack action of the sample distribution network in a third state, wherein the third state represents the state obtained by defending the sample distribution network through an intermediate defense model when the sample distribution network is in the second state; obtaining the attack evaluation loss value of the meta-attack evaluation network based on the first state, the third state, the second attack action, and the attack reward; obtaining the attack strategy loss value of the meta-attack strategy network based on the first attack action, the first state, and the attack reward; and updating the parameters of the meta-attack strategy network and the meta-attack evaluation network based on the attack evaluation loss value and the attack strategy loss value to obtain the updated attack model.

[0125] During the update process of the meta-attack model, the first state can be represented as: The second state can be represented as The third state can be represented as The first attack action can be represented as The second attack action can be represented as The attack reward can be represented as .

[0126] The meta-attack strategy network can be a network used to determine the first attack action in the first state, and can be represented as follows: , as shown in the following formula.

[0127] (31);

[0128] in, This represents the maximum expected reward for performing the first attack action at time t, which is also the loss function of the meta-attack strategy network. This represents the reward obtained after taking the first attack action in response to a defensive action at time t. This represents the expectation of executing the first attack action in the state at time t. Policy entropy is a measure of the randomness in the choice of attack actions. It encourages exploration by penalizing the determinism in action selection. This represents a temperature parameter used to balance the relationship between strategy entropy and expected return. Let represent the logarithmic probability of the defensive action at time t. This represents the expectation of all possible defensive actions at time t.

[0129] The parameters of the meta-attack policy network and the meta-attack evaluation network can be updated using the gradient descent method based on the data processed by the GCNSAGE feature extractor. The loss function is as follows.

[0130] (32);

[0131] (33);

[0132] (34);

[0133] (35);

[0134] in, The loss function represents the meta-attack strategy network. This represents the expected values ​​of all possible states and attack actions. This indicates an attack sample pool that can store parameters from the training process. Let be the attack policy loss value of the meta-attack policy network at time t. This indicates calculating the gradient. This represents the parameters of the network based on the meta-attack strategy. For learning rate, This represents the attack evaluation loss value of the meta-attack evaluation network. Let represent the expected first state and first attack action in the experience pool at time t. This represents the expected value of all possible attack actions by the attacker in the next time step. This represents the attack evaluation loss value at time t. The objective Q-function value of the meta-attack evaluation network at time t is obtained based on the state prediction of the next time step. This represents the reward for executing the first attack action in the first state at time t. This represents the function value of the meta-attack evaluation network predicted at time t+1. and represent the parameters of the meta-attack evaluation network and the target attack evaluation network, respectively. This is the soft update coefficient.

[0135] Next, dynamic optimization is performed based on the meta-attack parameters, and the strategy parameters are initialized. At this point, the objective function is revised to the following formula.

[0136] (36);

[0137] in, The KL divergence weights constrain the deviation between the attack model and the meta-attack model, thus avoiding overfitting.

[0138] According to embodiments of this disclosure, when the improvement in the attack success rate of the updated attack model is less than a preset attack improvement threshold, obtaining the attack model may include: acquiring the attack success improvement rate of the updated attack model in each of two consecutive attack training windows, wherein the updated attack model is trained for a preset number of attack training times in each attack training window; and obtaining the attack model when the attack success improvement rate in each of the two consecutive attack training windows is less than the preset attack improvement threshold.

[0139] To avoid excessive resource consumption or insufficient training due to excessive iteration of attack strategies, dynamic termination conditions are set based on sliding window statistics. The attacker's attack success rate is expressed by the following formula.

[0140] (37);

[0141] in, This represents the attack success rate of the w-th attack training window, where w is the attack training window number. For example, each attack training window can be set to 30 training rounds. Let w be the total number of attack actions that satisfy the physical constraints within the w-th attack training window. The number of valid attack actions within the w-th attack training window is defined as a voltage exceeding the limit by more than ±5% or a network loss increase of more than 10% after the attack.

[0142] The attack success rate can be expressed by the following formula.

[0143] (38);

[0144] in, This represents the attack success rate improvement for the w-th attack training window. For a positive number close to 0, we can take 10. -5 This is set to avoid a denominator of 0, when two consecutive attack training windows meet the condition. In this case, attack strategy training can be stopped to obtain the attack model.

[0145] According to embodiments of this disclosure, the attack sample pool is updated based on the confidence level of action-state pairs generated by the intermediate defense model in the event of an attack by the attack model, resulting in an updated attack sample pool. This includes: determining that an action-state pair is a defense blind zone state when the confidence level of the action-state pair is less than a preset confidence threshold; performing layered perturbation on the action-state pairs in the defense blind zone state according to the physical constraints of the power distribution network to obtain adversarial perturbation samples; and mixing the adversarial perturbation samples with the attack samples in the attack sample pool at a preset ratio to obtain an updated attack sample pool.

[0146] Defense strategies based on the intermediate defense model The calculated actions are then used to calculate node features using the GCNSAGE feature extractor, thereby obtaining the confidence level of the state-action pair.

[0147] (39);

[0148] (40);

[0149] in, This represents the confidence level of the state-action pair. That is, the action state is correct. , where is the policy entropy of the intermediate defense model. The higher the confidence level, the stronger the robustness of the defense decision.

[0150] Screening confidence The state-action pair is used as the core feature set. To pre-set a confidence threshold, actions with confidence levels lower than the pre-set confidence threshold are identified as defense blind zone states, as expressed in the following formula.

[0151] (41);

[0152] (42);

[0153] in, This represents the set of states corresponding to state-action pairs with a confidence level greater than a preset confidence threshold. This represents the set of states corresponding to state-action pairs with a confidence level less than or equal to a preset confidence threshold, also known as defense blind zone states.

[0154] According to the physical constraints of the distribution network, the state of the defense blind zone can be perturbed in layers to obtain the perturbation state, as shown in the following formula.

[0155] (43);

[0156] in, This indicates a state of resistance to disturbance. Indicates a first-order perturbation. Indicates a second-order perturbation. The coefficients of the first-order perturbation are represented. This represents the coefficient of the second-order perturbation.

[0157] Based on the perturbation state, obtain adversarial samples. An updated attack sample pool can be obtained by mixing adversarial perturbation samples and attack samples from the attack sample pool in a certain proportion. For example, the proportion of perturbation samples can be 0.4 and the proportion of attack samples can be 0.6, as shown in the following formula.

[0158] (44).

[0159] By introducing a meta-learning framework and injecting "robust meta-parameters" during the pre-training phase, the meta-parameters are updated through the meta-learning objective after each round of attack and defense iterations. This enables the strategy to "quickly adapt to new scenarios," such as convergence with only a small number of gradient updates when facing unknown attacks. This can improve the defense success rate in new attack scenarios and overcome the bottleneck of "insufficient generalization" in traditional training.

[0160] Two key operations are inserted during the training interval between defender and attacker: policy distillation and sample reconstruction. Policy distillation extracts the defender's "core defense logic" (high-confidence decisions) and "knowledge blind spots" (low-confidence blind spot states), guiding the attacker to focus on essential weaknesses. Sample reconstruction generates targeted perturbation samples based on blind spot states, forcing the attacker to confront the defender's "core weaknesses" rather than instantaneous parameters from the early stages of training. Updating the attack sample pool makes the attacker's training more targeted, solving the problem of "asynchronous attack and defense strategies" in traditional alternating training, increasing the game intensity by more than 40%. Through core feature extraction and adversarial risk constraints, bidirectional flow of attack and defense knowledge is achieved. The defender's core strategy guides the attacker to avoid ineffective perturbations; the attacker's high-risk attack distribution forces the defender to prioritize optimizing key scenarios. This shifts the attack and defense strategy from "blind game" to "precise adversarial," accelerating convergence to a better Nash equilibrium point.

[0161] According to embodiments of this disclosure, training an intermediate defense model to obtain a defense model under fixed attack parameters of an attack model includes: obtaining a first defense action for a second state using the intermediate defense model under fixed attack parameters of the attack model, wherein the second state is the state after the attack model attacks the sample distribution network in the first state; defending the sample distribution network using the first defense action under the attack model's attack on the sample distribution network to obtain a fourth state of the sample distribution network and a defense reward in the fourth state; updating the parameters of the intermediate defense model according to the defense reward to obtain an updated defense model; and obtaining a defense model when the defense success rate of the updated defense model is greater than a preset defense threshold.

[0162] The training of the intermediate defense model can include multiple rounds. Each round can produce an updated defense model. The training can be stopped based on whether the defense success rate of the updated defense model in each round is greater than a preset defense threshold. If the defense success rate is greater than the preset defense threshold, that is, if the defense capability of the updated defense model meets the requirements, the training can be stopped and the updated defense model in the last round can be used as the defense model.

[0163] The intermediate defense model can defend the sample distribution network based on the second state. For example, it can change the power of the photovoltaic equipment in the sample distribution network or change the node switches in the topology. Based on the defense of the intermediate defense model, a fourth state of the sample distribution network can be obtained, along with the defense reward in the fourth state. When attacking the sample distribution network using the attack model, since the attack model is already trained, it can more comprehensively represent the attack scenarios of the sample distribution network, thus better revealing the shortcomings of the intermediate defense model.

[0164] According to embodiments of this disclosure, the intermediate defense model includes an intermediate attack strategy network and an intermediate attack evaluation network. Updating the parameters of the intermediate defense model based on defense rewards to obtain an updated defense model includes: using the intermediate defense evaluation network to determine a second defense action of the sample distribution network in a fifth state, wherein the fifth state represents the state obtained by attacking the sample distribution network through the attack model when the sample distribution network is in a fourth state; obtaining a defense evaluation loss value of the intermediate defense evaluation network based on the second state, the fifth state, the second defense action, and the defense reward; obtaining a defense strategy loss value of the intermediate defense strategy network based on the first defense action, the second state, and the defense reward; and updating the parameters of the intermediate defense strategy network and the intermediate defense evaluation network based on the defense evaluation loss value and the defense strategy loss value to obtain the updated defense model.

[0165] In the sample distribution network, photovoltaic (PV) devices can be considered continuously operating devices, while node switches in the topology can be considered discretely operating devices. Power fluctuations in PV devices will cause changes in energy storage devices and static var compensators (SVCs). Therefore, the defensive action space of the first defensive action can be defined. Represented as a continuous sub-action space and discrete action space The Cartesian product of the intermediate defense model at time t is given by the following formula.

[0166] (45);

[0167] in, These represent the actions of the photovoltaic device, energy storage device, and static var compensator at time t, respectively, with values ​​ranging from [value range missing]. , The action of the node switching in the topology is represented by the action of the intermediate defense model at time t, which satisfies the constraint of formula (6).

[0168] During the update process of the intermediate defense model, the second state can be represented as: The fifth state can be represented as Defense rewards can be represented as .

[0169] This disclosure relates to the problem of learning robust defense strategies for discrete and continuous mixed actions. By rationally scheduling each device, the scheduling cost is reduced while enhancing the robustness of the strategy under state adversarial attacks, thus addressing potential FDIA in the sample distribution network. The input to the intermediate defense model is the perturbation state obtained after the attack model, and the reward value obtained in the non-attack state is obtained. The Branching Dueling Q network (BDQ) is extended by introducing two parallel GCNSAGE feature extractors, referred to as GS-BDQ, which works in conjunction with GS-SAC to solve the robust defense strategy learning problem. GS-BDQ can be used to learn the defensive actions of discrete-action devices, while the GS-SAC algorithm can be used to learn the defensive actions of continuously-action devices. The following equation applies.

[0170] (46);

[0171] (47);

[0172] in, This represents the loss function of the intermediate defense strategy network corresponding to discrete action devices. Let represent the expected value of the defense strategy loss of the intermediate defense strategy network obtained by performing the discrete action of dimension d in the second state at time t, where d represents the dimension of the discrete action. This represents the discrete action of the d-th dimension executed at time t. This represents the d-th dimension discrete action executed at time t+1. This represents the intermediate defense evaluation network corresponding to discrete action devices. Let represent the defense evaluation loss value of the intermediate defense evaluation network corresponding to the discrete action device at time t. This represents the defense evaluation loss value of the intermediate defense evaluation network corresponding to the discrete action device, predicted based on time t+1. This indicates the number of samples during the training process. and These represent the parameters of the intermediate defense strategy network corresponding to discrete-action devices and the parameters of the intermediate defense strategy network corresponding to continuously-action devices, respectively. and Let represent the gradients of the intermediate defense strategy network corresponding to discrete-action devices and the intermediate defense strategy network corresponding to continuously-action devices, respectively. This represents the target defense evaluation loss value of the intermediate defense evaluation network corresponding to the discrete action device at time t, as predicted. This represents the reward for performing the first defensive action corresponding to the discrete action device in the second state at time t. This represents the loss function of the intermediate defense strategy network corresponding to continuously operating devices. This represents the continuous actions performed at time t. Let represent the expected first defense action of the intermediate defense strategy network corresponding to the continuously acting device at time t, representing all possible second states of the defense sample pool. This refers to samples in the defense sample pool. This represents the intermediate defense evaluation network corresponding to continuously operating equipment. This represents the defense evaluation loss value of the intermediate defense evaluation network corresponding to the continuously operating device at time t.

[0173] Optimize using GS-BDQ and GS-SAC training respectively and It is possible to obtain the intermediate defense policy network and intermediate defense evaluation network under the attack of the attack model. The parameters of the intermediate defense policy network and intermediate defense evaluation network can be updated by minimizing the value of the loss function.

[0174] (48);

[0175] (49);

[0176] in, This represents the loss function of the intermediate defense evaluation network. This represents the expectation of all possible second states and first defensive actions at time t. This represents the defense evaluation loss value of the intermediate defense evaluation network at time t. This represents the defense evaluation loss value of the intermediate defense evaluation network based on the prediction at time t+1. This represents the target defense evaluation loss value of the intermediate defense evaluation network at time t, as predicted. This represents the reward for performing the first defensive action in the second state at time t. This represents the expected value of all possible second defensive actions at time t+1. and These represent the parameters of the intermediate defense strategy network and the intermediate defense evaluation network, respectively. This represents the gradient of the parameters of the intermediate defense strategy network.

[0177] Next, a risk mitigation term is added to the loss function of the intermediate defense model, that is, an attacker distribution constraint is introduced, and the defense strategy parameters are initialized. The loss function is then modified as follows.

[0178] (50);

[0179] (51);

[0180] in, For risk-resistance items, To counteract the risk, i.e. the maximum impact of the attacker on the defender's defense evaluation value, the defense evaluation network is updated according to formulas (50) and (51), and the input features are replaced with the core feature set. To improve the robustness of the strategy.

[0181] According to an embodiment of this disclosure, obtaining a defense model when the defense success rate of the updated defense model is greater than a preset defense threshold includes: obtaining the defense success rate of the updated defense model in two consecutive defense training windows, wherein the updated defense model performs training for a preset number of defense training times in each defense training window; and obtaining a defense model when the defense success rate in two consecutive defense training windows is greater than the preset defense threshold.

[0182] To balance defense robustness and training efficiency, the defense success rate is used as the stopping condition for training intermediate defense models. The defense success rate within the defense training window is expressed by the following formula.

[0183] (52);

[0184] in, This represents the defense success rate of the w'-th defense training window. , These represent the number of effective defensive actions of the intermediate defense model and the number of effective attack actions of the attack model, respectively, for the w'-th defense training window.

[0185] Training can be stopped if the defense success rate is greater than the preset defense threshold for two consecutive defense training windows. The preset defense threshold can be 90%.

[0186] After completing one round of learning the defense robustness strategy, the meta-parameters are updated, that is, the defense model is used as the meta-defense model, and the attack model is used as the meta-attack model, as shown below:

[0187] (53);

[0188] (54);

[0189] in, For the pseudo-attack actions already covered in the current training, and These are the parameters for the current defense model and attack model, respectively.

[0190] The overall performance index for each round of attack and defense can be defined as follows:

[0191] (55);

[0192] in, Indicates the first The overall performance fluctuation of the wheel, , These are defense weights and attack weights, respectively. , The first The average defense success rate and average attack success rate of each round of global iteration.

[0193] The overall performance fluctuation for each round is expressed by the following formula.

[0194] (56);

[0195] in, Indicates the overall performance fluctuation. This represents the overall performance after three consecutive rounds of iterative training. It can be... In the case of [condition], it serves as the termination condition for global training.

[0196] According to embodiments of this disclosure, by using dynamic termination conditions, such as setting the attack success rate to be <1% and the defense success rate to >90%, the training of attack and defense strategies can be "stopped on demand." The attack strategy terminates only when it can no longer improve the effect, and the defense strategy stops only when robustness is restored or the parameters converge. This can reduce redundant training steps by 30%-50% while ensuring that each iteration achieves substantial improvement in the strategy, avoiding overfitting or underfitting.

[0197] Furthermore, for the objective function in this embodiment, which includes multiple objectives such as safety operation indicators and economic operation indicators, a penalty-based boundary intersection (PBI) method can be used to decompose the objective function. This mainly consists of two parts: objective function decomposition and Pareto optimal policy learning. A set of uniformly distributed weights can be generated to ensure that the optimal solutions of all sub-functions of the objective function uniformly cover the entire Pareto front. Using these preference vectors, the multi-objective Markov game model is transformed into m sub-models with scalar rewards using the Chebyshev method. Each sub-model has different weights, reflecting different degrees of trade-offs between voltage deviation and voltage regulation costs. Based on this, a meta-Pareto optimization term is introduced to ensure that the Pareto policy has cross-scenario adaptability. Therefore, the scalar reward of objective j at time t is expressed as...

[0198] (57);

[0199] In the formula, This represents the reward after performing an action in the state at time t. This represents the reference reward for the j-th subfunction at time t. This represents the reward of the j-th subfunction at time t. , , Indicates the number of sub-functions. This represents the weight of the j-th sub-function in the m-th sub-model. The Pareto generalization coefficient; Let be the average reward of objective function j in the pseudo-attack scenario.

[0200] Normalization can be used to process the reward function to solve the problem of different dimensions of reward values ​​for different objectives. The ARG-MODRL algorithm can be used to learn the Pareto policy for the sub-model to obtain the scheduling policy of each device.

[0201] According to embodiments of this disclosure, the reward obtained by the intermediate defense model in defending the sample distribution network with the weakest defense action under the strongest attack action is the same as the reward obtained by the meta-attack model in attacking the sample distribution network with the strongest attack action under the weakest defense action.

[0202] The reward can be the total reward obtained from defensive and offensive actions, therefore The reward for each moment is expressed as follows.

[0203] (58);

[0204] in, Let be a definite positive constant. If it satisfies the constraints of equations (7)-(11), then ,otherwise .

[0205] In equilibrium, both sides in a game will sacrifice their own interests to some extent to change their strategies. For TDZMG, Nash equilibrium represents the highest reward a defender can obtain when facing the strongest attacker, i.e., a minimax solution. Therefore, it is necessary to find the Nash equilibrium strategies for both the defender and the attacker, as shown in the following equation.

[0206] (59);

[0207] To directly solve the Nash equilibrium problem, the training method of this disclosure can consist of two phases, focusing on training the attacker and defender respectively to obtain attack and defense actions. During attacker training, the defense action is fixed, with the objective of minimizing the defender's cumulative reward, represented as a constrained minimization problem.

[0208] (60);

[0209] The meta-defense model can be pre-trained in an attack-free environment, establishing an initial exploration strategy for the attacker. The purpose of training the defender is to obtain a robust defense strategy that enhances adaptability and resilience against state attacks. At this point, the attacker's strategy remains fixed, with the goal of maximizing the defender's cumulative reward, which is represented as a constrained maximization problem.

[0210] (61);

[0211] The defense model of this disclosure integrates perception, decision-making, and defense to achieve proactive defense of the distribution network. It utilizes multi-objective deep reinforcement learning based on the GCNSAGE feature extractor to perform safe and economical scheduling of the distribution network. Furthermore, it uses flexible resources in the distribution network, such as switches, photovoltaic devices, energy storage devices, and static var compensators, as scheduling objects to achieve safe and economical coordinated scheduling of the distribution network's flexible resources.

[0212] Figure 4 A flowchart is shown for a method for training an adversarial defense model for coordinated optimization of distribution network security and economy according to yet another embodiment of the present disclosure.

[0213] like Figure 4 As shown, the method includes operations S401 to S410.

[0214] In operation S401, a trained meta-defense model and a trained meta-attack model are obtained by using a set of pseudo-attack scenarios generated based on the sample distribution network.

[0215] Based on the objective function, perform the following operations.

[0216] In operation S402, a meta-defense model is trained in the first state of the sample distribution network to obtain an intermediate defense model, wherein the first state represents the sample distribution network in an attack-free operating state.

[0217] Meta-defense parameters based on the meta-defense model are meta-adapted and trained in an attack-free environment to obtain an intermediate defense model.

[0218] In operation S403, with the intermediate defense parameters of the intermediate defense model fixed, the meta-attack model is used to determine the first attack action based on the attack samples in the attack sample pool.

[0219] In operation S404, under the defense of the intermediate defense model, the sample distribution network is attacked using the first attack action to obtain the second state of the sample distribution network and the attack reward in the second state. The parameters of the meta-attack model are updated according to the attack reward to obtain the updated attack model.

[0220] In operation S405, determine whether the attack success rate of the updated attack model is less than the preset attack enhancement threshold. If the increase in the attack success rate of the updated attack model is less than the preset attack enhancement threshold, obtain the attack model and execute operation S406. If the increase in the attack success rate of the updated attack model is greater than or equal to the preset attack enhancement threshold, execute operation S403.

[0221] In operation S406, the attack sample pool is updated based on the confidence of the action-state pairs generated by the intermediate defense model under the attack of the defense attack model, resulting in the updated attack sample pool.

[0222] In operation S407, with fixed attack parameters of the attack model, the intermediate defense model is used to obtain the first defense action against the second state. Under the attack model's attack on the sample distribution network, the first defense action is used to defend the sample distribution network, resulting in the fourth state of the sample distribution network and the defense reward in the fourth state. The parameters of the intermediate defense model are updated based on the defense reward to obtain the updated defense model. The second state is the state of the sample distribution network after the attack model attacks it in the first state.

[0223] In operation S408, determine whether the defense success rate of the updated defense model is greater than the preset defense threshold. If the defense success rate of the updated defense model is greater than the preset defense threshold, obtain the defense model and execute operation S409. If the defense success rate of the updated defense model is less than or equal to the preset defense threshold, execute operation S407.

[0224] In operation S409, determine whether the overall performance fluctuation of the defense model and the attack model meets the preset conditions. If the overall performance fluctuation of the defense model and the attack model meets the preset conditions, execute operation S410. If the overall performance fluctuation of the defense model and the attack model does not meet the preset conditions, use the defense model as the meta-defense model and the attack model as the meta-attack model, and execute operation S402.

[0225] When operating the S410, output defense and attack models.

[0226] Operations S401 to S410 can be described in other embodiments of this disclosure, and will not be repeated here.

[0227] By introducing a meta-learning framework and injecting "robust meta-parameters" during the pre-training phase, the meta-parameters are updated through the meta-learning objective after each round of attack and defense iterations. This enables the strategy to "quickly adapt to new scenarios," such as convergence with only a small number of gradient updates when facing unknown attacks. This can improve the defense success rate in new attack scenarios and overcome the bottleneck of "insufficient generalization" in traditional training.

[0228] Figure 5 A schematic diagram of a distribution network topology in one example according to an embodiment of the present disclosure is shown.

[0229] like Figure 5 As shown, the distribution network comprises 69 nodes, with high-load nodes and end nodes selected as the attacked nodes. Figure 5 In the sample distribution network shown, a total of 6 distributed photovoltaic (PV) devices are connected, each with an installed capacity of 0.9MW, located at nodes 19, 27, 35, 46, 49, and 61. Three distributed energy storage devices are connected at nodes 20, 45, and 48, each with an installed capacity of 2MWh and a rated power of 0.4MW. Two 0.4Mvar static var compensators are connected at nodes 24 and 64, respectively. The safe operating voltage range of the distribution network is [0.93, 0.97], and the maximum current carrying capacity of branches is 400A, set at 50%.

[0230] Three methods were selected for comparison and analysis with the training method of the present disclosure embodiment. Method 1 is a training method based on GS-SAC and GS-BDQ; Method 2 is a training method based on graph convolutional depth deterministic policy gradient and graph convolutional depth Q network; and Method 3 is a training method based on graph convolutional soft actor commentator.

[0231] Figure 6 The diagram illustrates the performance of the training methods of this disclosure, methods 1, 2, and 3, in an attack-free environment. Figure 7The diagram illustrates the performance of the training methods of this disclosure, including methods 1, 2, and 3, under an attack environment.

[0232] like Figure 6 As shown, the training method of this embodiment exhibits superior convergence stability. In the early training phase (0-1000 rounds), the reward curves of the three comparative methods fluctuate significantly, indicating insufficient adaptability to common disturbances such as random load changes and intermittent photovoltaic output. Although the training method of this embodiment experiences slight fluctuations in the very early stages (0-600 rounds) due to policy exploration, it converges quickly thanks to the robust policy pre-trained in the attack environment, demonstrating strong environmental adaptability. The final reward of the training method of this embodiment is slightly lower than that of Method 1. This is because introducing the attack environment pre-training mechanism slightly sacrifices training performance in exchange for robustness against state attacks. Method 1 maintains good exploration performance through the entropy regularized SAC algorithm, exhibits smaller reward fluctuations under state attacks, and uses SAC and BDQ to handle continuous and discrete actions respectively, achieving more refined device scheduling. Furthermore, the GCNSAGE feature extractor of the training method of this embodiment can quickly adapt to potential topological changes in the system, thus resulting in a higher cumulative reward compared to Method 2. Method 3 relies solely on GCN to extract spatial features, resulting in weaker recognition capabilities and inferior performance compared to the GCNSAGE feature extractor that embeds DRL, leading to lower rewards.

[0233] Depend on Figure 7 As can be seen, the reward curve of the training method in this embodiment exhibits a "fluctuation-stability" process. The decline originates from the attacker learning adversarial strategies to disrupt defense decisions, while the rise indicates that the defender has learned effective strategies, enhancing strategy quality and robustness. Finally, the small-range stable fluctuation demonstrates that the ability to resist attacks has been gradually mastered through adversarial training. In contrast, the comparative method, lacking robust training, suffers from the "pure optimization strategy" being impacted by false data, and the "state-action" mapping is disrupted, resulting in a continuously oscillating reward curve that fails to converge stably. In summary, the training method in this embodiment demonstrates significant advantages through adversarial robust alternating training.

[0234] Table 1 compares the scheduling security, economy, and solution speed of different methods. As can be seen from the table, under optimal voltage deviation, the training method of this embodiment has the second highest maximum branch current carrying capacity and operating cost after method 1, decreasing by 6.87% and 12.64 yuan compared to method 2, and by 7.15% and 14.44 yuan compared to method 3, respectively. This indicates that it is more effective in balancing voltage deviation, branch current carrying capacity, and operating cost, while simultaneously ensuring security and economy. This is because the training method of this embodiment uses a "state disturbance perception-robust decision output" mechanism based on adversarial training to capture voltage fluctuations caused by false disturbances in real time through the GCNSAGE feature extractor, enabling the distribution network to maintain voltage stability even under attacks. Secondly, the training method of this embodiment embeds branch current carrying capacity constraints into the reward function, strengthening the power balance of "load-power supply-branch" through adversarial training, avoiding branch overload due to excessive pursuit of economy. Its economic advantages also stem from the adoption of a hybrid action space optimization strategy. GS-SAC handles the continuous actions of PV, ESS, and SVC devices, while GS-BDQ handles the discrete actions of tie switches, achieving coordinated optimization of adjustable resources. Simultaneously, the GCNSAGE feature extractor adapts to topology changes in real time through neighbor sampling and mean aggregator, enabling more accurate calculation of network losses and operating costs compared to the static feature extraction methods 2 and 3. Although the training method in this embodiment requires simultaneous optimization of attack and defense dual-agent strategies, resulting in a longer training time, the testing time is only 0.365 seconds, still meeting the real-time requirements of distribution network dispatching.

[0235] Table 1

[0236]

[0237] Figure 8 A schematic diagram of the node voltage distribution before a defense model obtained using the training method of embodiments of this disclosure controls a distribution network is shown in one example. Figure 9 A schematic diagram of node voltage distribution after a defense model obtained using the training method of this disclosure is used to control a distribution network is shown in one example.

[0238] Figure 8 and Figure 9As shown, the intermittent nature and anti-peak-shaving characteristics of photovoltaic (PV) power generation significantly alter the power flow distribution of the distribution network. During the daytime surge in PV power generation before dispatch, local voltage rises and large voltage fluctuations occur. Conversely, after PV power is deactivated at night, insufficient reactive power support in the distribution network causes voltage drops exceeding permissible limits. Using the defense model obtained through the training method of this embodiment (all defense models described below are obtained through the training method of this embodiment), the voltage differences between nodes are reduced by optimizing reactive power and the charging and discharging of energy storage devices, alleviating the voltage exceedance problem. The node voltages at all times are within the specified range, indicating that the defense model effectively improves the power flow distribution of the distribution network under anti-attack conditions and reduces the local voltage distortion rate.

[0239] To verify the robustness of the defense model under state adversarial attack scenarios, five cases were set up for comparative analysis: 1) Case I: No scheduling measures in the adversarial attack environment; 2) Case To employ a defensive model to control the distribution network in an adversarial attack environment; 3) Case Study Method 1; 4) Case Study: Method 1 is used in an attack-free environment. 5) Case Study: Using a defensive model to control the power distribution network in an attack-free environment; Method 1 is used in a no-attack environment. The attack strength is set to medium interference, at 20%.

[0240] Figure 10 A schematic diagram of 24-hour voltage distribution under different cases according to embodiments of the present disclosure is shown.

[0241] like Figure 10 As shown, combined with Figure 9 Node voltage distribution, for Figure 5 The analysis of node 65 in the violin plots of 24-hour voltage distribution under different cases is shown in the figure. Figure 10 As shown in the figures, Case I exhibits a relatively concentrated voltage distribution, all below 0.91 pu, indicating that under adversarial attack conditions, the system, lacking a scheduling strategy, cannot correct measurement distortions caused by the attack, resulting in a long-term voltage deviation from the safe range and poor stability. Case II shows a voltage distribution within the allowable range (0.96 pu-0.98 pu), demonstrating that the defense model, through anomaly measurement data feature identification and pre-trained attack compensation actions, can effectively address voltage exceeding limits caused by attacks. Case III is similar to Case I, indicating that under adversarial attack conditions, Method 1, lacking robust training, cannot adapt to spurious states, executes incorrect scheduling actions, and lacks robustness, posing a safety hazard against FDIA. In Case IV, the voltage remains within the specified range at all times, indicating that under attack-free conditions, the defense model can effectively manage voltage issues.

[0242] Further experiments with varying degrees of attack were conducted to verify the robustness of the defense model. Table 2 shows that as the attack intensity increased from no attack (0%) to severe tampering (50%), the degradation of core indicators remained controllable and consistently met safety requirements: the minimum voltage decreased by only 2.60%, still above the safety threshold; the maximum branch current carrying capacity increased from 0.609 pu to 0.671 pu, below the thermal limit; and the operating cost increased by 28.9%, but still significantly outperformed Case III under the same attack intensity. This demonstrates that the defense model, through attack training and robust action compensation, can adaptively adjust its scheduling strategy according to the attack intensity, exhibiting a stable adaptability to attack uncertainties. This further proves the effectiveness of the defense model for multi-objective robust scheduling in distribution networks with FDIA.

[0243] Table 2

[0244]

[0245] This disclosure proposes a three-pronged training method of "perception-decision-defense," resulting in a defense model for the safe and economical scheduling of flexible resources based on GCNSAGE graph multi-objective deep reinforcement learning. It uses flexible resources such as network topology switches, photovoltaic equipment, energy storage devices, and static var compensators as scheduling objects to achieve safe and economical coordinated scheduling of distribution network flexible resources. The training method in this disclosure adds a robust distillation interface between the alternating training processes of attackers and defenders, accurately separating core defense features from robustness blind spots, providing targeted knowledge support for the attack-defense game. Through a collaborative mechanism of "dynamic iterative training + policy distillation + adversarial sample reconstruction," it replaces traditional static alternating training, improving attack targeting and correct defense efficiency. Combined with a single-index global training termination criterion, it simplifies the quantitative logic for engineering implementation, ensuring efficient training convergence. The training method of this disclosure has a fast solution speed, can be applied online in real time, and has good adaptability to complex scenarios such as topology changes, distributed photovoltaic power output fluctuations and adversarial attacks. It can effectively support the safe and economical collaborative scheduling of new power distribution systems with large-scale distributed photovoltaic penetration, significantly reduce system voltage deviation, reduce branch load rate, and improve operational economy and reliability.

[0246] Figure 11 A block diagram of a multi-target cooperative defense model training apparatus according to an embodiment of the present disclosure is shown.

[0247] like Figure 11 As shown, the device 1100 includes an acquisition module 1110 and an execution module 1120.

[0248] Module 1110 is used to obtain a trained meta-defense model and a trained meta-attack model by utilizing a set of pseudo-attack scenarios generated based on the sample distribution network, wherein the set of pseudo-attack scenarios includes different types of attack scenarios against the sample distribution network.

[0249] The execution module 1120 performs the following operations based on the objective function until the combined performance fluctuation of the obtained defense model and attack model meets preset conditions. The objective function is constructed based on the safety and economic operation indicators of the sample distribution network: A meta-defense model is trained in the first state of the sample distribution network to obtain an intermediate defense model, where the first state represents a no-attack operation state of the sample distribution network; With the intermediate defense parameters of the intermediate defense model fixed, a meta-attack model is trained based on attack samples in the attack sample pool to obtain an attack model, where the attack samples in the attack sample pool are generated from pseudo-attack scenarios in a pseudo-attack scenario set; The attack sample pool is updated based on the confidence of the action-state pairs generated by the intermediate defense model in defending against attacks from the attack model, resulting in an updated attack sample pool; With the attack parameters of the attack model fixed, an intermediate defense model is trained to obtain a defense model; If the combined performance fluctuation of the defense model and attack model does not meet preset conditions, the defense model is used as the meta-defense model, and the attack model is used as the meta-attack model.

[0250] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.

[0251] For example, any plurality of the obtaining module 1110 and the execution module 1120 can be combined into one module / unit / subunit, or any one of the modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits can be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of this disclosure, at least one of the obtaining module 1110 and the execution module 1120 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the obtaining module 1110 and the execution module 1120 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.

[0252] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0253] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. A training method for an adversarial defense model that promotes coordinated optimization of power distribution network security and economy, characterized in that, include: By using a set of pseudo-attack scenarios generated based on a sample distribution network, a trained meta-defense model and a trained meta-attack model are obtained, wherein the set of pseudo-attack scenarios includes attack scenarios of different types and intensities targeting the sample distribution network. Based on the objective function, perform the following operations until the combined performance fluctuation of the defense model and the attack model meets the preset conditions, wherein the objective function is constructed based on the safety operation indicators and economic operation indicators of the sample distribution network: The meta-defense model is trained in the first state of the sample distribution network to obtain the intermediate defense model, wherein the first state represents the sample distribution network as an attack-free operating state; With the intermediate defense parameters of the intermediate defense model fixed, the meta-attack model is trained based on the attack samples in the attack sample pool to obtain the attack model, wherein the attack samples in the attack sample pool are generated based on the pseudo-attack scenarios under the pseudo-attack scenario set. Based on the confidence level of the action-state pairs generated by the intermediate defense model in defending against the attack model, the attack sample pool is updated to obtain the updated attack sample pool. The intermediate defense model is trained with the attack parameters of the attack model fixed to obtain the defense model; If the combined performance fluctuation of the defense model and the attack model does not meet the preset conditions, the defense model is used as the meta-defense model, and the attack model is used as the meta-attack model.

2. The method according to claim 1, characterized in that, The attack sample pool is updated based on the confidence level of the action-state pairs generated by the intermediate defense model in defending against attacks from the attack model, resulting in an updated attack sample pool, including: If the confidence level of the action state pair is less than a preset confidence threshold, the action state pair is determined to be a defense blind zone state. For the action state pairs in the aforementioned defense blind zone state, hierarchical perturbation is performed according to the physical constraints of the distribution network to obtain counter-perturbation samples; The adversarial perturbation samples are mixed with the attack samples in the attack sample pool at a preset ratio to obtain an updated attack sample pool.

3. The method according to claim 1, characterized in that, The step of training the meta-attack model based on attack samples in the attack sample pool, while fixing the intermediate defense parameters of the intermediate defense model, to obtain the attack model, includes: With the intermediate defense parameters of the intermediate defense model fixed, the meta-attack model is used to determine the first attack action based on the attack samples in the attack sample pool; Under the defense of the intermediate defense model, the sample distribution network in the first state is attacked using the first attack action to obtain the second state of the sample distribution network and the attack reward in the second state. The parameters of the meta-attack model are updated based on the attack reward to obtain the updated attack model; The attack model is obtained when the increase in the attack success rate of the updated attack model is less than a preset attack improvement threshold.

4. The method according to claim 3, characterized in that, The attack model is obtained when the increase in the attack success rate of the updated attack model is less than a preset attack improvement threshold, including: The attack success rate of the updated attack model is obtained in each of two consecutive attack training windows, wherein the updated attack model is trained for a preset number of attack training times in each attack training window; The attack model is obtained when the attack success rate in each of two consecutive attack training windows is less than the preset attack improvement threshold.

5. The method according to claim 3 or 4, characterized in that, The meta-attack model includes a meta-attack policy network and a meta-attack evaluation network; The step of updating the parameters of the meta-attack model based on the attack reward to obtain the updated attack model includes: The second attack action of the sample distribution network in the third state is determined by the meta-attack evaluation network, wherein the third state represents the state obtained by defending the sample distribution network by the intermediate defense model in the second state; Based on the first state, the third state, the second attack action, and the attack reward, the attack evaluation loss value of the meta-attack evaluation network is obtained; Based on the first attack action, the first state, and the attack reward, the attack strategy loss value of the meta-attack strategy network is obtained; Based on the attack evaluation loss value and the attack strategy loss value, the parameters of the meta-attack strategy network and the meta-attack evaluation network are updated to obtain the updated attack model.

6. The method according to any one of claims 1 to 4, characterized in that, The step of training the intermediate defense model while fixing the attack parameters of the attack model to obtain the defense model includes: With the attack parameters of the attack model fixed, the intermediate defense model is used to obtain the first defense action against the second state, wherein the second state is the state after the attack model attacks the sample distribution network in the first state; Under the attack model on the sample distribution network, the first defense action is used to defend the sample distribution network, resulting in the fourth state of the sample distribution network and the defense reward in the fourth state. The parameters of the intermediate defense model are updated based on the defense reward to obtain the updated defense model; The defense model is obtained when the success rate of the updated defense model is greater than the preset defense threshold.

7. The method according to claim 6, characterized in that, The step of obtaining the defense model when the defense success rate of the updated defense model is greater than a preset defense threshold includes: The defense success rate of the updated defense model is obtained in two consecutive defense training windows, wherein the updated defense model is trained for a preset number of defense training times in each defense training window; The defense model is obtained when the defense success rate of each of the two consecutive defense training windows is greater than the preset defense threshold.

8. The method according to claim 7, characterized in that, The intermediate defense model includes an intermediate defense strategy network and an intermediate defense evaluation network; The step of updating the parameters of the intermediate defense model based on the defense reward to obtain the updated defense model includes: The intermediate defense evaluation network is used to determine the second defense action of the sample distribution network in the fifth state, wherein the fifth state represents the state obtained by attacking the sample distribution network through the attack model when the sample distribution network is in the fourth state; Based on the second state, the fifth state, the second defensive action, and the defensive reward, the defense evaluation loss value of the intermediate defense evaluation network is obtained; Based on the first defense action, the second state, and the defense reward, the defense strategy loss value of the intermediate defense strategy network is obtained; Based on the defense evaluation loss value and the defense strategy loss value, the parameters of the intermediate defense strategy network and the intermediate defense evaluation network are updated to obtain the updated defense model.

9. The method according to any one of claims 1 to 4, characterized in that, The objective function includes a safety operation sub-function constructed based on the safety operation indicators and an economic operation sub-function constructed based on the economic operation indicators; The safe operation sub-function includes at least one of a voltage deviation rate term and a branch load rate term, wherein the voltage deviation rate term is determined based on the voltage amplitude and voltage over-limit boundary value of the nodes of the sample distribution network, and the branch load rate term is determined based on the actual current carrying capacity and the upper limit of current carrying capacity of the branches of the sample distribution network. The economic operation sub-function includes at least one of the following: grid loss cost, curtailment cost, energy storage charging and discharging cost, static var compensator (SVC) operating cost, and switching operation cost. The grid loss cost is determined based on the grid loss power in the sample distribution network; the curtailment cost is determined based on the curtailment of photovoltaic (PV) equipment in the sample distribution network; the energy storage charging and discharging cost is determined based on the energy storage charging and discharging power of energy storage equipment in the sample distribution network; the SVC operating cost is determined based on the reactive power of the SVC in the sample distribution network; and the switching operation cost is determined based on the number of switching operations in the sample distribution network.

10. A training device for a collaborative optimization model of security and economic aspects in power distribution networks, characterized in that, include: The module is used to obtain a trained meta-defense model and a trained meta-attack model by using a set of pseudo-attack scenarios generated based on the sample distribution network, wherein the set of pseudo-attack scenarios includes attack scenarios of different types and intensities against the sample distribution network. The execution module performs the following operations based on the objective function until the combined performance fluctuation of the defense model and the attack model meets a preset condition. The objective function is constructed based on the safety and economic operation indicators of the sample distribution network. The meta-defense model is trained in the first state of the sample distribution network to obtain the intermediate defense model, wherein the first state represents the sample distribution network as an attack-free operating state; With the intermediate defense parameters of the intermediate defense model fixed, the meta-attack model is trained based on the attack samples in the attack sample pool to obtain the attack model, wherein the attack samples in the attack sample pool are generated based on the pseudo-attack scenarios under the pseudo-attack scenario set. Based on the confidence level of the action-state pairs generated by the intermediate defense model in defending against the attack model, the attack sample pool is updated to obtain the updated attack sample pool. The intermediate defense model is trained with the attack parameters of the attack model fixed to obtain the defense model; If the combined performance fluctuation of the defense model and the attack model does not meet the preset conditions, the defense model is used as the meta-defense model, and the attack model is used as the meta-attack model.