A robust adversarial training framework for multi-agent reinforcement learning energy system

By constructing adversary agents and a robust adversarial training framework, the anti-attack capability of the multi-agent reinforcement learning energy system is enhanced, solving the security problem of traditional systems under complex energy flow and network attacks, and achieving more stable operation.

CN116306903BActive Publication Date: 2025-11-28ZHEJIANG ZHENENG YUEQING POWER GENERATION CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211516697.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-11-28
Estimated Expiration
2042-11-30

AI Technical Summary

Technical Problem

Traditional integrated energy management systems struggle to achieve online assessment and real-time control when faced with complex energy flows and cyberattacks, and their communication networks are vulnerable to malicious attacks, highlighting significant security and vulnerability issues.

Method used

Construct an adversary agent and model it as an adversarial partially observable stochastic game system. Train the optimal deterministic adversarial strategy and enhance the robustness of the multi-agent reinforcement learning system model through robust adversarial training to resist adversarial attacks.

Benefits of technology

This improves the robustness and stability of the multi-agent reinforcement learning energy system in the face of potential attacks, reduces the oscillation of the load demand curve, and enhances the system's resistance to attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116306903B_ABST
    Figure CN116306903B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of robust confrontation training framework for multi-agent reinforcement learning energy system, comprising: constructing an adversarial agent to generate adversarial attack, and modeling as adversarial partially observable stochastic game system;Fixed pre-trained victim multi-agent strategy, train an optimal deterministic confrontation strategy to produce bounded disturbance;Fixed optimal confrontation attack strategy, improve the robustness of victim strategy under optimal attacker through adversarial training.The beneficial effects of the present application are: the present application models adversarial attack as an attack adversary based on single-agent reinforcement learning, and learns the strongest attack strategy considering attack constraints.Mathematically, the problem is constructed as a confrontation Markov game, and the performance of the integrated energy management system based on multi-agent reinforcement learning is improved through robust confrontation training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of power system security defense, more specifically, it relates to a robust adversarial training framework for multi-agent reinforcement learning energy system. BACKGROUND

[0002] With the development of social economy and the growth of energy demand, the power system is undergoing a fundamental revolution in planning and operation from fossil fuels to clean energy. Under the background of the rapid development of energy internet, the integrated energy system of electricity, gas, heat, cold and other multiple energy coupling and coordination can realize multi-energy complementation, promote renewable energy consumption, improve energy utilization efficiency and alleviate supply and demand imbalance. Compared with the traditional power system, the energy flow of the integrated energy system is more complex, and its operation and regulation involve more complex load demand, supply devices and operation modes. The new characteristics of high coupling between energy demand, supply and storage will cause the complexity of system operation mode and dynamic characteristics to increase, the source and load uncertainty to intensify, the simulation system mathematical model variables and dimensions to increase, and the safety and stability margin to decrease, thereby making it difficult for the traditional integrated energy management method based on mathematical model mechanism to meet the demand of online evaluation and real-time control. Therefore, the data-driven integrated energy management method based on multi-agent reinforcement learning emerges as the times require. With the integration of information and communication technology, the safety and vulnerability of the integrated energy management system based on multi-agent reinforcement learning cannot be underestimated. The communication network of the integrated energy management system, including the monitoring and data acquisition network and smart meters and other devices, is very vulnerable to attacks by malicious network actors. SUMMARY

[0003] The purpose of the present application is to overcome the deficiencies in the prior art, and provide a robust adversarial training framework for multi-agent reinforcement learning energy system. The present application enhances the resistance of the integrated energy management system based on multi-agent reinforcement learning to adversarial attacks through robust adversarial training. First, an adversary agent is constructed, whose goal is to cause the worst performance of the control system by formulating an adversarial attack, and the system is modeled as an adversarial partially observable stochastic game system; then the adversary agent is trained to learn an optimal deterministic adversarial attack strategy to generate a bounded disturbance; finally, the victim multi-agent reinforcement learning integrated energy management system is subjected to robust adversarial training to enhance model robustness.

[0004] In a first aspect, a robust adversarial training framework for multi-agent reinforcement learning energy system is provided, comprising:

[0005] Step 1, an adversarial agent is constructed to generate an adversarial attack, and is modeled as an adversarial partially observable stochastic game system;

[0006] Step 2: Fix the pre-trained victim multi-agent policy and train an optimal deterministic adversarial policy to generate bounded perturbations.

[0007] Step 3: Fix the optimal adversarial attack strategy and improve the robustness of the victim strategy under the optimal attacker through adversarial training.

[0008] Preferably, step 1 includes:

[0009] Step 1.1: The integrated energy management system based on multi-agent reinforcement learning is formulated as a partially observable stochastic game problem, where each agent controls one building. The goal is to optimize the policies of all agents to maximize the cumulative reward of the entire team.

[0010] <N,S,{A i} i∈N ,P,{R i} i∈N ,γ,{O i} i∈N ,Z>

[0011] Where N is the number of agents, S is the environmental state, and A is the number of agents. i It is the action space of the i-th agent, {A i} i∈N It is the joint action space, defined as A = A 1 ×…×A N P:S×A×S→Δ(S) is a given action at any time t. From state s t State s at the next time step t+1 t+1 The state transition probability; Is the i-th agent from (s) t ,a t ) to the next state s t+1 Timely feedback and rewards; γ is the discount factor; O i Let {O} be the observation space of the i-th agent, and let {O} be the joint observation space. i} i∈N Defined as O = O 1 ×…×O N Z:S×A→Δ(O) represents the joint observation at any time t. t ∈O in any action a t Under state s t The probability of observation;

[0012] At time t, each agent i, based on its observations... Through strategy Select Action Then, the environment moves to the next state according to the state transition probability P, st+1 ~P(·|s t ,a t Each agent i receives a reward. and new local observations

[0013] Step 1.2: Introduce an adversarial agent into the integrated energy management system. By generating the strongest adversarial attack to induce the worst performance of the model, model this system as an adversarial part of an observable stochastic game problem.

[0014] <N,S,A adv ,{A i} i∈N ,P,{R i} i∈N ,R adv ,γ,{O i} i∈N ,Z>

[0015] Where N is the number of victim agents, S is the environmental state, and A adv and R adv These are the attacker's action space and reward function, respectively; A i This is the action space of the Ath victim agent, {A i} i∈N It is the joint action space, defined as A = A 1 ×…×A N ;P:S×A adv ×A×S→Δ(S) is a given action at any time t. and A adv From state s t State s at the next time step t+1 The state transition probability; Is the i-th agent from (s) t ,a t ) to the next time step state s t+1 Timely feedback and rewards; γ is the discount factor; O i Let {O} be the observation space of the i-th agent, and let {O} be the joint observation space. i} i∈N Defined as O = O 1 ×…×O N Z:S×A→Δ(O) represents the joint observation at any time t. t ∈O in any action a t Under state s t The probability of observation.

[0016] Preferably, step 2 includes:

[0017] Step 2.1, fixing the pre-trained normal victim multi-agent system strategy parameters θ i represent the model parameters of each agent strategy, train an adversarial agent strategy u φ , φ is the strategy parameter of the attack agent to simulate the adversarial attack and threaten one of the agents, and the generated attack is:

[0018]

[0019] where δ t is the generated attack vector for the observation of a specific agent, is the observation of the agent to be attacked, B(o j ) is the boundary constraint of the disturbance; the input of the victim agent j is represented as:

[0020]

[0021] The victim strategy makes decisions based on the disturbed observation:

[0022]

[0023] where is the action made by the multi-agent integrated energy management system after being attacked;

[0024] Step 2.2, fixing the victim multi-agent system strategy π θ , the reward function of the attacker is defined as R adv = -∑R i , and its objective function is:

[0025]

[0026] where J(θ,φ) = ∑R i , the attack agent interacts with the multi-agent integrated energy management system for training to generate the optimal attack strategy

[0027] As an option, in step 3, the optimal attacker strategy trained in step 2.2 is fixed where φ * is the parameter of the optimal attack strategy, which is used to interact with the environment to generate attack vectors, and through adversarial training, the robustness of the victim strategy under the optimal attacker is improved, and its objective function is:

[0028]

[0029] where J(θ,φ) = ∑R i .

[0030] In a second aspect, a robust adversarial training device for a multi-agent reinforcement learning energy system is provided for executing the robust adversarial training framework for a multi-agent reinforcement learning energy system of the first aspect, comprising:

[0031] A construction module is configured to construct an adversarial agent to generate an adversarial attack and model an adversarial partially observable stochastic game system.

[0032] A first fixing module is configured to fix a pre-trained victim multi-agent strategy, and train an optimal deterministic adversarial strategy to generate a bounded disturbance.

[0033] A second fixing module is configured to fix the optimal adversarial attack strategy, and improve the robustness of the victim strategy under the optimal attacker through adversarial training.

[0034] The present application has the beneficial effects that: the present application designs a robust adversarial training framework for a multi-agent reinforcement learning energy system to cope with potential adversarial attacks. The adversarial attack is modeled as an attack adversary based on single-agent reinforcement learning, and the strongest attack strategy considering attack constraints is learned. Mathematically, the problem is constructed as an adversarial Markov game, and the performance of the multi-agent reinforcement learning-based integrated energy management system is improved through robust adversarial training. BRIEF DESCRIPTION OF DRAWINGS

[0035] Figure 1 A flowchart of a robust adversarial training framework for a multi-agent reinforcement learning energy system;

[0036] Figure 2 A structural schematic diagram of a robust adversarial training framework for a multi-agent reinforcement learning energy system. DETAILED DESCRIPTION

[0037] The present application will be further described below in conjunction with the embodiments. The following description of the embodiments is only to help understand the present application. It should be noted that for ordinary people in the technical field, some modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

[0038] In order to ensure the stability, reliability and efficient operation of the multi-agent reinforcement learning-based integrated energy management system as a whole, and to improve its robustness in response to malicious network attacks, the present application proposes a robust adversarial training framework for a multi-agent reinforcement learning energy system, which enhances its flexibility through adversarial training, and has important significance for realizing stable and safe operation of community integrated energy management system.

[0039] The following illustrates how to achieve robust reinforcement of a multi-agent reinforcement learning integrated energy management system with an experiment based on a community-level integrated energy management system containing nine buildings.

[0040] As shown in Figure 1 , the present application is a robust adversarial training framework for a multi-agent reinforcement learning energy system, which comprises the following steps:

[0041] (1) Express the multi-agent reinforcement learning-based integrated energy management system as a partially observable stochastic game problem, with each agent controlling a building, and optimize the strategies of all agents to maximize the cumulative reward of the entire team:

[0042] <N,S,{A i} i∈N ,P,{R i} i∈N ,γ,{O i} i∈N ,Z>

[0043] where N is the number of agents, S is the environment state, A i is the action space of the i-th agent, {A i} i∈N is the joint action space, defined as A=A 1 ×…×A N ; P:S×A×S→Δ(S) is the state transition probability from a given action s t at time t to the next state s t+1 at time t+1; is the i-th agent's immediate feedback reward from (s t ,α t ) to the next state s t+1 ; γ is the discount factor; O i is the observation space of the i-th agent, and the joint observation space is {O i} i∈N , defined as O=O 1 ×…×O N ; Z:S×A→Δ(O) is the observation probability of the joint observation o t ∈O under any action a t and state s t at time t. At time t, each agent i selects an action according to the observation and the strategy Then, the environment moves to the next state according to the state transition probability P. Each agent i obtains a reward and a new local observation The entire process is iterative, and for each agent i, a trajectory of observations, actions, and rewards can be obtained: The goal of agent i is to obtain a policy π i To maximize its cumulative discounted return, as shown in the following formula:

[0044] J=∑R i

[0045] Where -i represents all agents in set N except agent i. In this cooperative environment, the goal of the integrated energy management system based on multi-agent reinforcement learning is to optimize the agent policy parameters θ = {θ...} 1 ,θ 2 ,…,θ N Maximize the cumulative total team reward J:

[0046]

[0047] (2) Figure 2 As shown, an adversarial agent is introduced into this integrated energy management system based on multi-agent reinforcement learning. The aim is to induce the worst performance of the model by generating the strongest adversarial attack, thus modeling the system as a stochastic game problem with observable adversarial components.

[0048] <N,S,A adv ,{A i} i∈N ,P,{R i} i∈N ,R adv ,γ,{O i} i∈N ,Z>

[0049] Where N is the number of victim agents, S is the environmental state, and A adv and R adv These are the attacker's action space and reward function, respectively; A i It is the action space of the i-th victim agent, {A i} i∈N It is the joint action space, defined as A = A 1 ×…×A N ;P:S×A adv ×A×S→Δ(S) is a given action at any time t. and A adv From state s t State s at the next time step t+1 t+1 The state transition probability; Is the i-th agent from (s) t ,a t ) to the next time step state s t+1the timely feedback reward; γ is the discount factor; O i is the observation space of the ith agent, and the joint observation space is {O i} i∈N , defined as O = O 1 ×…×O N ; Z: S × A → Δ(O) is the observation probability of the state s t ∈ O under any action a t , state s t . Note that N, S, {A i} i∈N ,, γ, {O i} i∈N , Z are consistent with the definition of the partially observable stochastic game described above, but P, {R i} i∈N is affected by A adv .

[0050] (3) Fix the pre-trained normal victim multi-agent system policy parameters (θ i represents the model parameters of each individual agent policy), train an adversarial agent policy u φ (φ is the policy parameter of the attack agent) to simulate an adversarial attack and threaten one of the agents, and the generated attack is:

[0051]

[0052] where δ t is the generated attack vector on the observation of a specific agent, is the observation of the agent to be attacked, and B(o j ) is the boundary constraint of the disturbance. Then the input of the victim agent j is:

[0053]

[0054] The victim policy makes a decision based on the disturbed observation:

[0055]

[0056] where a is the action made by the multi-agent reinforcement learning integrated energy management system after being attacked. If the adversarial disturbance is within the physical constraints of the physical characteristics and amplitude range, such as stable increasing inflexible energy demand and energy storage within capacity, the defense mechanism cannot detect it. Therefore, the adversarial disturbance can be limited to B(o j ), thereby discovering the vulnerabilities of the multi-agent reinforcement learning integrated energy management system.

[0057] (4) Fixed victim multi-agent system strategy π θ , the reward function of the attacker is defined as R adv i , then the objective function is:

[0058]

[0059] where J(θ,φ) = ∑R i The attack agent interacts with the multi-agent integrated energy management system for training to generate the optimal attack strategy where φ * is the parameter of the optimal attack strategy. Compared with the random attack strategy generating random noise, the attack effect on the multi-agent integrated energy management system is as follows:

[0060] Table 1

[0061]

[0062] Cumulative ramp rate, average daily peak and maximum peak are the measurement indexes of the demand load curve of the multi-agent integrated energy management system. It can be found from the table that the optimal attack makes the cumulative ramp rate, average daily peak and maximum peak of the model increase by 38.61%, 8.77%, 16.42%, making the load demand curve of the multi-agent integrated energy management system more volatile, and the attack effect is better than the random attack, fully exploring the vulnerability of the integrated energy management system.

[0063] (5) Fixed training of the optimal attacker strategy Use it to generate an attack vector for a certain victim agent, and improve the robustness of the multi-agent integrated energy management system under the optimal attacker through adversarial training, and the objective function is:

[0064]

[0065] where J(θ,φ * ) = ∑R i

[0066] At the same time, the random attack strategy with random noise is also used for adversarial training as a comparison. The performance of the multi-agent integrated energy management system under different training modes under adversarial attack is as follows:

[0067] Table 2

[0068]

[0069] ​It can be found from the table that compared with no adversarial training, the optimal attack adversarial training makes the model cumulative climbing rate, average daily peak and maximum peak reduce by 13.24%, 4.78%, 6.96%, makes the load demand curve of the multi-agent integrated energy management system flat, and even under attack, it can maintain good performance, and improve the robustness of the integrated energy management system to attack. On the contrary, the model trained by random attack adversarial training does not maintain good performance.

[0070] In summary, the present application introduces an adversarial attacker based on single-agent reinforcement learning, which interacts with the multi-agent reinforcement learning integrated energy management system to generate the strongest attack, which is achieved by interfering with the observation of a certain victim agent; and the optimal adversarial attack strategy is fixed, and the victim multi-agent reinforcement learning integrated energy management system is trained against it, and through learning this adversarial experience, the resilience to adversarial attack is enhanced, and a robust control strategy is generated.

Claims

1. A robust adversarial training method for multi-agent reinforcement learning energy system, characterized in that, The application is applied to a comprehensive energy management system, comprising: Step 1, constructing an adversarial agent to generate an adversarial attack, and modeling as an adversarial partially observable stochastic game system; Step 1 includes: Step 1.1, expressing the comprehensive energy management system based on multi-agent reinforcement learning as a partially observable stochastic game problem, each agent controls a building, and the cumulative reward of the whole team is maximized by optimizing the strategy of all agents: < N, S, {A i} i∈N , P, {R i} i∈N , γ, {O i} i∈N , Z > where N is the number of agents, S is the environment state, A i is the action space of the i-th agent, {A i} i∈N is the joint action space, defined as A = A 1 ×…×A N ; P : S × A × S → Δ(S) is the state transition probability from a given action at time t, s t , to the next time t + 1, s t+1 ; is the immediate feedback reward of the i-th agent from (s t , a t ) to the next time state s t+1 ; γ is the discount factor; O i is the observation space of the i-th agent, and the joint observation space is {O i} i∈N , defined as O = O 1 ×…×O N ; Z : S × A → Δ(O) is the observation probability of the joint observation o t ∈ O at any action a t , state s t , at any time t. At time t, each agent i updates its belief By policy Selects action Then, the environment moves to the next state, s t+1 ~ P(·|s t ,a t ); each agent i gets reward and new local observation Step 1.2, introducing an opponent agent in the comprehensive energy management system, causing the worst performance of the model by generating the strongest adversarial attack, and modeling this system as an adversarial partially observable stochastic game problem: < N, S, A adv , {A i} i∈N , P, {R i} i∈N , R adv , γ, {O i} i∈N , Z > where N is the number of victim agents, S is the environment state, A adv and R adv are the action space and reward function of the attacker, respectively; A i is the action space of the ith victim agent, {A i} i∈N is the joint action space, defined as A = A 1 x... x A N ; P : S x A adv x A x S → Δ(S) is the state transition probability from state s and A adv to the next time state s t under action a t+1 ; is the immediate feedback reward of the ith agent from (s t , a t ) to the next time state s t+1 ; γ is the discount factor; O i is the observation space of the ith agent, and the joint observation space is {O i} i∈N , defined as O = O 1 x... x O N ; Z : S x A → Δ(O) is the observation probability of the joint observation o t ∈ O under any action a t , state s t , at any time t. Step 2, fixing the pre-trained victim multi-agent strategy, training an optimal deterministic adversarial strategy to generate bounded disturbance; Step 2 includes: Step 2.1, fix the pre-trained normal victim multi-agent system policy parameters θ i represent the model parameters of each agent policy, train an adversarial agent policy u φ φ is the policy parameter of the attack agent to simulate an adversarial attack and threaten one of the agents, and the attack generated by it is: where δ t is the generated attack vector for a particular agent observation, is the observation of the agent to attack, B(o j ) is the boundary constraint of the perturbation; the input of the victim agent J is represented as: The victim strategy makes decisions based on disturbance observation: wherein is the action made by the multi-agent integrated energy management system after being attacked; Step 2.2, fix the victim multi-agent system strategy π θ , define the reward function of the attacker as R adv = -∑R i , then its objective function is: where J(θ, φ) = ∑R i The intelligent agent interacts with the multi-agent integrated energy management system for training to generate an optimal attack strategy Step 3, fixing the optimal adversarial attack strategy, improving the robustness of the victim strategy under the optimal attacker through adversarial training; In step 3, the optimal attacker strategy obtained in step 2.2 is fixed where φ * is the parameter of the optimal attack strategy, which is used to generate attack vectors by interacting with the environment, and the robustness of the victim strategy under the optimal attacker is improved through adversarial training, and the objective function is: where J(θ, φ) = ∑R i .

2. A robust adversarial training device for multi-agent reinforcement learning energy system, characterized in that, A method for performing robust adversarial training of a multi-agent reinforcement learning energy system according to claim 1, comprising: A construction module for constructing an adversarial agent to generate an adversarial attack, and modeling as an adversarial partially observable stochastic game system; A first fixing module for fixing the pre-trained victim multi-agent strategy, training an optimal deterministic adversarial strategy to generate bounded disturbance; A second fixing module for fixing the optimal adversarial attack strategy, improving the robustness of the victim strategy under the optimal attacker through adversarial training.

Citation Information

Patent Citations

  • Disturbance award-oriented deep reinforcement learning confrontation defense method

    CN114925850A