Intelligent power generation control method for autonomous learning of regional reward function of interconnected power grid
By using trajectory sorting reward extrapolation and multi-agent deep stochastic policy gradient algorithm to autonomously learn the reward function, the problems of low accuracy of reward function and low efficiency of policy optimization in intelligent power generation control are solved, and efficient intelligent power generation control effect is achieved.
Patent Information
- Application Number
- CN202510988174.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-07
AI Technical Summary
Existing multi-agent deep reinforcement learning algorithms rely on reward functions set by human experience in intelligent power generation control, which lacks accuracy and has low policy optimization efficiency, especially in the early stages of training where there is exploration redundancy.
The system employs a trajectory sorting reward extrapolation algorithm and a multi-agent deep stochastic policy gradient algorithm. By autonomously learning the reward function, the system utilizes a suboptimal expert agent to generate the trajectory sorting reward extrapolation algorithm to autonomously learn the reward function of the current power grid area. Finally, it combines the multi-agent deep stochastic policy gradient algorithm to optimize the policy network, thereby achieving intelligent power generation control.
It achieves accuracy and efficiency of the reward function, reduces the complexity of manual configuration, improves the policy convergence speed and control effect, and can quickly adapt to and optimize control strategies in complex power systems.
Smart Images

Figure CN120914905A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automatic generation control, in particular to an intelligent generation control method based on autonomous learning of regional reward functions of interconnected power grids. BACKGROUND
[0002] Automatic Generation Control (AGC) of interconnected power grids is one of the most basic functions of power grid energy management systems and is a basic means to ensure active power balance and frequency stability of power systems. In recent years, deep reinforcement learning algorithms have been widely used in the field of AGC due to their strong decision-making ability and high environmental adaptability. For example, CN109217306A discloses an intelligent generation control method based on deep reinforcement learning with action self-optimization capability, which includes the following steps: 1. determining a state set; 2. determining an action set; 3. collecting real-time operation data of each regional power grid, such as frequency deviation and power deviation, and calculating the instantaneous value of the control error of each region and the instantaneous value of the control performance target; 4. determining the current state and the current internal state, and then obtaining a short-term reward function signal of a certain regional power grid according to the current state, the internal state and the reward function; 5. calculating the target Q value function and the loss function; 6. updating the weight value by calculation; 7. searching and evaluating new actions and updating the action set; 8. performing corresponding operations on all regional power grids; and 9. returning to step 3. This method can effectively obtain the optimal coordinated control of the power grid and solve the problem of strong random disturbance caused by the interconnection of large-scale new energy and distributed energy from the perspective of automatic generation control.
[0003] However, the reward function structure and its weight parameters relied on by the existing multi-agent deep reinforcement learning algorithm in intelligent generation control are usually set by artificial experience, which lacks sufficient accuracy and easily ignores important variables affecting policy optimization. This artificial design-dependent approach requires a large number of experimental iterations to determine reasonable parameter configurations, which is complex and inefficient. In addition, during the policy training process, traditional methods usually use random exploration of the generation control action space, especially in the early training stage, which has a lot of exploration redundancy, resulting in low efficiency of policy optimization. Therefore, how to realize automatic modeling of the reward function and efficient policy search in intelligent generation control has become a key problem to be solved. SUMMARY
[0004] The present application relates to the technical field of automatic generation control, in particular to an intelligent generation control method based on autonomous learning of regional reward functions of interconnected power grids.
[0005] The object of the present application can be achieved by the following technical solutions:
[0006] The method comprises the following steps:
[0007] S1, collecting power grid operation state quantities in real time;
[0008] S2, inputting the power grid operation state quantities into the agent corresponding to the region, the agent evaluating the probability distribution of each discrete action based on the trained policy network, and selecting the control action with the maximum probability from the discrete action space output by the policy network as the power generation adjustment instruction at the current time, wherein the agent comprises a control agent and a group of suboptimal expert agents, the control agent being trained by using a multi-agent deep stochastic policy gradient algorithm and being used for outputting a control strategy according to the input power grid operation state quantities, and the suboptimal expert agents autonomously learning the reward function of the current power grid region by using a trajectory ranking reward extrapolation algorithm during the training process and feeding back to the control agent;
[0009] S3, mapping the control action output by the agent into a corresponding power generation adjustment amount, and applying the power generation adjustment amount to the generator set in the control region to realize real-time adjustment of the frequency and update the power grid operation state quantities, thereby providing a basis for the control decision at the next time;
[0010] S4, cyclically executing the above steps S1-S3 at a fixed period to realize continuous, real-time and distributed intelligent collaborative power generation control.
[0011] The power grid operation state quantities comprise the frequency deviation, the regional control error, the derivative of the regional control error and the control instruction difference value of each region.
[0012] The suboptimal expert agent performs the following steps to self-learn the reward function:
[0013] Generating a non-optimal demonstration expert policy trajectory, training the non-optimal demonstration expert policy of the region, and generating a corresponding trajectory after a plurality of rounds;
[0014] Inputting the power grid operation state quantities contained in the trajectory into an artificially set reward function with prior knowledge, ranking the trajectory according to the output reward value, randomly extracting trajectories according to the ranking result, and generating a preferred trajectory pair by two-by-two combination;
[0015] Based on the preferred trajectory pair, the trajectory ranking reward extrapolation algorithm constructs a parameterized reward function through a multi-layer deep neural network Calculating the approximate reward value under the power system state, and performing reward inference:
[0016]
[0017] Wherein, s represents the agent observation state. represents a parameterized reward function; τ i represents the i-th trajectory, τ j represents the j-th trajectory.
[0018] The parameterized reward function is trained to optimize the reward inference process by a loss function as follows:
[0019]
[0020] where L(θ) represents the loss function, Π represents the trajectory distribution, τ i represents the i-th trajectory, τ j represents the j-th trajectory; ξ represents a binary classification loss function; P represents a probability function defined by a normalized distribution; represents the discounted reward calculated according to the parameterized reward function ; < represents an indication of the preference between the demonstration expert policy trajectories.
[0021] The meta-classification loss function ξ is instantiated using cross-entropy to obtain an instantiated loss function:
[0022]
[0023] During the training of the control agent, the policy network is u i (s|θ), u i (s|θ) represents the n-th discrete output value output by the control agent observing the grid operation state quantity of the i-th region, and the corresponding discrete action probability value is obtained through normalization processing and where k is the number of discrete actions.
[0024] When the agent policy network interacts with the environment, the discrete action index of the control agent of region i is :
[0025]
[0026] where g n is a random variable that obeys the Gumbel(0, 1) distribution, and obeys the following random variable distribution:
[0027] g n =-log(-log(ε)) ε∈U(0,1)
[0028] where ε is a uniform distribution random variable obeying the interval (0, 1), and 0 in U(0, 1) is the lower limit of the uniform distribution, and 1 is the upper limit of the uniform distribution.
[0029] Map the discrete action index to the discrete action space of the agent, and output the final control instruction value.
[0030] The purpose of the multi-agent deep stochastic policy gradient algorithm is to obtain the optimal parameters of the agent policy network, so that the total expected reward value J is maximum:
[0031]
[0032] Where T represents the total simulation time; γ represents the reward discount factor; r i t represents the self-learning reward function reward value at each time; s represents the state data sampled from the experience pool; D represents the experience pool; a is the real-time sampled action, and u is the policy network.
[0033] The gradient of the policy network parameters of the control agent of region i is calculated according to the chain rule:
[0034]
[0035] Where S is the observation value of the control agent of all regions, J i is the expected reward value of region i, and ▽ θ is the gradient operation on the parameters θ in the expected reward value, is the control action with the maximum probability selected from the discrete action space output by the policy according to the Gumbel-Max mechanism, Q i is the Q value network, a i is the output action of the i-th control agent, N is the total number of control agents, and ψ is the parameter of the critic network.
[0036] The calculation method of the control action with the maximum probability selected from the discrete action space output by the policy according to the Gumbel-Max mechanism is:
[0037]
[0038] Where η represents the temperature parameter, which controls the discreteness of softmax, and k is the number of discrete actions.
[0039] The update of the parameter ψ of the critic network of the control agent is realized by minimizing the loss function L(ψ), and the specific form is as follows:
[0040] L(ψ)=E S~D,a~u [(Q i (S,a1,···,a2,a N |ψ)-y i ) 2 ]
[0041]
[0042] wherein S' = {s' 1, s' 2, ···, s' i, ···, s' N} represents the set of observation states of all regional control agents at the next time point, s' i is the observation state of the i th regional control agent at the next time point; Q N} represents the set of observation states of all regional control agents at the next time point, s i ' is the observation state of the i th regional control agent at the next time point; Q i (·|ψ) represents a critic network for receiving observation states and action information of all control agents in the power system environment and evaluating the behavior of the current control agent; S represents observation values of all regional control agents; and represent a target critic network and a target policy network, respectively; y i represents a target value.
[0043] Compared with the prior art, the present application has the following beneficial effects:
[0044] (1) The reward function of the present application is autonomously learned by using a trajectory ranking reward extrapolation algorithm, without the need for artificial experience setting, and has high accuracy. In the iterative adaptive process, important variables affecting policy optimization can be automatically considered, and a large number of experimental iterations are not needed to determine reasonable parameter configurations, thereby reducing the complexity involved in the artificial reward function and improving efficiency.
[0045] (2) In the policy training process of the present application, the agent policy is not only updated according to the optimization ability of the algorithm itself, but also guided by the feedback signal provided by the self-learning reward function to optimize the policy gradient, thereby significantly improving the policy convergence speed. BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 is a flowchart of the method of the present application;
[0047] Figure 2 is a schematic diagram of the algorithm framework of the present application;
[0048] Figure 3 is a comparison diagram of the artificial design reward function curves of each algorithm in the pre-learning stage in an embodiment;
[0049] Figure 4 is a comparison diagram of the step response ACE of each algorithm in region 1 in an embodiment;
[0050] Figure 5 is a comparison diagram of the step response ACE of each algorithm in region 1 under random disturbance in an embodiment. DETAILED DESCRIPTION
[0051] The application will be described in detail below with reference to the drawings and specific embodiments. The embodiments are implemented on the premise of the technical solutions of the application, and detailed implementation modes and specific operation processes are given, but the protection scope of the application is not limited to the following embodiments.
[0052] The embodiment provides an intelligent power generation control method for autonomous learning of an interconnection power grid regional reward function, as shown in the formula (1), comprising the following steps: Figure 1
[0053] S1, collecting power grid operation state quantities in real time.
[0054] In the embodiment, the power grid operation state quantities include frequency deviation Δf i , area control error ACE i , derivative of the area control error dACE i / dt and control instruction difference ΔP ci .
[0055] S2, inputting the power grid operation state quantities to an agent corresponding to the region, and the agent evaluating the probability distribution of each discrete action based on a trained policy network, and selecting a control action with the maximum probability from a discrete action space output by the policy network as a power generation regulation instruction at the current time, wherein, as shown in the formula (2), the agent includes a control agent and a group of suboptimal expert agents, the control agent is trained by using a multi-agent deep stochastic policy gradient algorithm, and is used for outputting a control strategy according to the input power grid operation state quantities, and the suboptimal expert agents autonomously learn a reward function of the current power grid region by using a trajectory-ranked reward extrapolation (T-REX) algorithm during a training process, and feed back to the control agent. Figure 2
[0056] In the embodiment, the T-REX algorithm autonomously learns the reward function of the current power grid region according to trajectory data. Then, the control agent is trained by using a multi-agent deep stochastic policy gradient (MADSPG) algorithm. An interconnected power system model receives a total control power generation instruction generated by the agent, and calculates a unit power generation power distribution instruction by using a power distribution algorithm and inputs the unit power generation power distribution instruction to a generator unit in each region. Data generated in the interaction process are stored in an experience pool. A critic network of the agent samples power system state data of all regions from the experience pool to guide the agent to output corresponding actions at the next state, and synchronously updates a policy network. The process includes the following two parts:
[0057] 1, autonomously learning a reward function of an intelligent power generation control task
[0058] The sub-optimal expert agent performs the following steps to self-learn the reward function:
[0059] A1, generate a non-optimal demonstration expert policy trajectory, train the non-optimal demonstration expert policy for the region, and generate the corresponding trajectories (τ1, τ2, τ3, ···, τ T ) after T rounds.
[0060] A2, input the grid operating state quantity contained in the trajectory into the artificially set reward function with prior knowledge, sort the trajectories according to the output reward value, for example, arrange m trajectories (τ1, τ2, τ3, ···, τ T ) from worst to best, and randomly extract trajectories for pairwise combination according to the sorting result to generate a pair of preferred trajectories.
[0061] A3, based on the pair of preferred trajectories, the Trajectory-ranked Reward EXtrapolation (T-REX) algorithm constructs a parameterized reward function to calculate the approximate reward value under the power system state, and performs reward extrapolation:
[0062]
[0063] where s represents the observation state of the agent; represents the parameterized reward function; τ i represents the i-th trajectory, and τ j represents the j-th trajectory.
[0064] The parameterized reward function is trained by the following loss function to optimize the reward extrapolation process:
[0065]
[0066] where L(θ) represents the loss function, Π represents the trajectory distribution, τ i represents the i-th trajectory, and τ j represents the j-th trajectory; ξ represents the binary classification loss function; P represents the probability function defined by the normalized distribution; represents the discounted reward calculated according to the parameterized reward function ; < represents an indication of the preference between the demonstration expert policy trajectories.
[0067] Equation (2) expresses that the probability function P is defined by the normalized distribution. In an embodiment, the cross-entropy can also be used to instantiate the binary classification loss function ξ, obtaining the instantiated loss function:
[0068]
[0069] 2. Multi-agent Deep Stochastic Policy Gradient (MADSPG) algorithm
[0070] During the training of the control agent, the policy network outputs the control instruction value u i (s|θ) of the control agent in the i-th region. i (s|θ) represents the control agent in the i-th region observing the grid operating state quantity [ACE i ,Δf i ,dACE i / dt,ΔP ci ] in the region and outputting the n-th discrete output value, and the corresponding discrete action probability value is obtained through normalization processing and where k is the number of discrete actions.
[0071] When the agent policy network interacts with the environment, the discrete action index of the control agent in the i-th region is as shown in equation (4):
[0072]
[0073] where g n is a random variable that obeys the Gumbel (0,1) distribution, so that the probability distribution of the discrete action becomes differentiable, and is subject to the following random variable distribution:
[0074] g n = -log(-log(ε)) ε∈U(0,1) (5)
[0075] where ε is a uniform distribution random variable subject to the interval (0,1), and 0 in U(0,1) is the lower bound of the uniform distribution, and 1 is the upper bound of the uniform distribution.
[0076] Then, the discrete action index is mapped to the discrete action space of the agent, and the final control instruction value is output.
[0077] The purpose of the multi-agent deep stochastic policy gradient algorithm is to obtain the optimal parameters of the agent policy network, so that the total expected reward value J is maximum:
[0078]
[0079] where T represents the total simulation time; γ represents the reward discount factor; r i trepresents the self-learning reward function reward value at each time; s represents the state data sampled from the experience pool; D represents the experience pool; a is the real-time sampled action, and u is the policy network.
[0080] As shown in formula (7), the gradient of the policy network parameter of the control agent of region i is calculated according to the chain rule:
[0081]
[0082] Wherein, S is the observation value of the control agent of all regions, J i is the expected reward value of region i, and ▽ θ is the gradient operation on the parameter θ in the expected reward value, is the control action with the maximum probability selected from the discrete action space output by the policy according to the Gumbel-Max mechanism, Q i is the Q value network, a i is the output action of the i th control agent, N is the total number of control agents, and ψ is the parameter of the critic network.
[0083] In the policy network parameter updating process, since the discrete action space is not differentiable, the MADSPG algorithm needs an approximate gradient calculation method, in order to ensure the existence of the gradient of the action to the policy network parameter, the embodiment selects the control action with the maximum probability from the discrete action space output by the policy according to the Gumbel-Max mechanism The calculation method is as follows:
[0084]
[0085] Wherein, η represents the temperature parameter, controls the discrete degree of softmax, and k is the number of discrete actions.
[0086] The update of the parameter ψ of the critic network of the control agent is realized by minimizing the loss function L(ψ), and the specific form is as follows:
[0087] L(ψ)=E S~D,a~u [(Q i (S,a1,···,a2,a N |ψ)-y i ) 2 ] (9)
[0088]
[0089] Wherein, S'={s1',s'2,···,s' N}, represents the set of next time observation states of the control agent of all regions, s i ' is the next time observation state of the control agent of the i th region; Qi (·|ψ) represents a critic network to receive observation states and action information of all control agents in the power system environment and evaluate the behavior of the current control agent; S represents the observation values of all regional control agents; and represent a target critic network and a target policy network, respectively; y i represents a target value.
[0090] In the training process of the MADSPG algorithm, the agent policy is not only updated based on the optimization capability of the algorithm itself, but also guided by the feedback signal provided by the self-learning reward function. This method effectively reduces the complexity of the design of the artificial reward function and significantly improves the policy convergence speed and learning effect.
[0091] This embodiment uses the original training data set containing 1000 sets of random step disturbance data and 1000 sets of wind power random disturbance data for training. Through these data, the intelligent power generation controller can learn how to deal with different types of disturbances, thereby improving its adaptability and control effect in the actual power system.
[0092] Figure 3 The pre-learning phase of each algorithm is shown in the artificial design reward function curve comparison chart, wherein SRF-SGC is the algorithm adopted by the present application, according to Figure 3 It can be found that the SRF-SGC framework can infer a high-quality reward function from the control intention of the learned expert policy through the T-REX algorithm, and apply it to the MADPSG algorithm with better convergence performance. Therefore, the reward function value of SRF-SGC is significantly improved during 0 to 500 rounds, and the final convergence value is better than that of other algorithms, further enhancing the collaborative control effect of the multi-region agent controller. The CSAC algorithm uses a reward function that decouples the target reward and the frequency constraint, dynamically adjusting the balance between the target and the penalty during the training process. Therefore, the reward curve of the CSAC algorithm converges faster and has no obvious shock. MATD3 uses a double-target critic network to reduce the overestimation of Q values and introduces Gaussian noise to enhance the action exploration ability of the agent, so that its reward convergence value is higher than that of the MADDPG algorithm. However, the agents in the above-mentioned comparative algorithms are affected by the subjective bias in the artificially designed reward function and the low exploration efficiency in the high-dimensional continuous action space during the training process. Therefore, the learning speed of CSAC, MATD3 and MADDPG algorithms is slower than that of SRF-SGC, and the convergence value is also lower than that of SRF-SGC.
[0093] S3, mapping the control action output by the intelligent agent into a corresponding power generation adjustment amount, acting on the generator set in the control area, realizing real-time adjustment of the frequency, and updating the power grid operation state quantity to provide a basis for the control decision of the next moment.
[0094] S4, the above steps S1-S3 are executed in a fixed cycle to realize continuous, real-time and distributed intelligent collaborative power generation control.
[0095] Figure 4 For the step response ACE comparison chart of each algorithm in region 1, Figure 5 For the step response ACE comparison chart of each algorithm in region 1 under random disturbance. Figure 4 It can be seen that in the offline training process, SRF-SGC makes full use of the performance ranking relationship of the suboptimal expert strategy trajectory data in multiple regional scheduling centers, thereby deducing a more effective reward function. This reward function can better optimize the policy network of the intelligent agent under the SRF-SGC framework, so that the online deployment of the SRF-SGC controller can quickly restore the ACE frequency to 0 at a lower ACE overshoot, thereby obtaining a better control effect. In contrast, although the CSAC, MATD3 and MADDPG algorithms can also guide the intelligent agent to complete the automatic generation control task through the preset reward function, such a reward function usually fails to fully consider the complex changes of the environment and the task requirements, and is prone to cause the intelligent agent to fall into a suboptimal control strategy, thereby failing to obtain the best control action. From Figure 5 It can be seen that in the more complex strong random interconnected power system environment, SRF-SGC can learn the mapping relationship between the power system operation state and the reward function constructed by the neural network according to the operation characteristics of different power grid regions. Therefore, SRF-SGC is significantly better than other algorithms in the |ACE| index, and can effectively maintain ACE within a reasonable range. In contrast, the intelligent agent trained by the CSAC algorithm focuses more on optimizing the performance target, and the control effect is better than that of the MATD3 algorithm and the MADDPG algorithm, but since the reward function related to the control target still relies on artificial design, it lacks sufficient accuracy, resulting in slightly lower control performance than SRF-SGC.
[0096] The above detailed the preferred embodiments of the present application. It should be understood that those skilled in the art can make many modifications and changes without creative labor according to the concept of the present application. Therefore, any technical solution obtained by logical analysis, reasoning or limited experiment by those skilled in the art on the basis of the prior art according to the concept of the present application shall be within the protection scope determined by the claims.
Claims
1. A method for smart generation control with autonomous learning of interconnection grid regional reward functions, characterized in that, The method comprises the following steps: S1, collecting power grid operation state quantities in real time; S2, inputting the power grid operation state quantities into the agent of the corresponding region, the agent evaluating the probability distribution of each discrete action based on the trained policy network, and selecting the control action with the maximum probability from the discrete action space output by the policy network as the generation adjustment instruction at the current time, wherein the agent comprises a control agent and a group of suboptimal expert agents, the control agent is trained by using a multi-agent deep stochastic policy gradient algorithm, and is used for outputting a control strategy according to the input power grid operation state quantities, and the suboptimal expert agents autonomously learn the reward function of the current power grid region by using a trajectory ranking reward extrapolation algorithm during the training process, and feed back to the control agent; S3, mapping the control action output by the agent into a corresponding generation adjustment amount, and applying the generation adjustment amount to the generator set in the control region to realize real-time adjustment of the frequency, and updating the power grid operation state quantities to provide a basis for the control decision at the next time; S4, cyclically executing the above steps S1-S3 at a fixed period to realize continuous, real-time and distributed intelligent collaborative generation control.
2. The method of claim 1, wherein the method further comprises: The power grid operation state quantities comprise the frequency deviation, the regional control error, the derivative of the regional control error and the control instruction difference value of each region. 3.The smart generation control method of autonomous learning of interconnection grid regional reward function according to claim 1, characterized in that, The suboptimal expert agent performs the following steps to learn the reward function autonomously: generating a non-optimal demonstration expert strategy trajectory, training the non-optimal demonstration expert strategy of the region, and generating a corresponding trajectory after a plurality of rounds; inputting the power grid operation state quantities contained in the trajectory into an artificially set reward function with prior knowledge, ranking the trajectory according to the output reward value, randomly extracting the trajectory according to the ranking result to generate a preferred trajectory pair; A trajectory ranking reward extrapolation algorithm constructs a parameterized reward function through a multi-layer deep neural network based on pairs of preference trajectories An approximate reward value is calculated for a state of the power system, and reward extrapolation is performed: where s represents the agent observation state; represents the parameterized reward function; τ i represents the i-th trajectory, τ j represents the j-th trajectory.
4. The method of claim 3, wherein the method further comprises: The parameterized reward function is trained by using the following loss function to optimize the reward inference process: where L(θ) denotes the loss function, Π denotes the trajectory distribution, τ i denotes the i-th trajectory, τ j denotes the j-th trajectory; ξ denotes the binary classification loss function; P denotes the probability function defined with the normalized distribution; denotes the discounted reward computed from the parameterized reward function ; < denotes an indicator of the preference between the demonstrated expert policy trajectories.
5. The method of claim 4, wherein the method further comprises: The meta-classification loss function ξ is instantiated by using cross-entropy to obtain an instantiated loss function:
6. The method of claim 1, wherein, During the control agent training process, the policy network outputs the nth discrete output value u i (s|θ), u i (s|θ) represents the control agent observing the power grid operating state quantity of the region corresponding to the ith region and outputting the nth discrete output value, and the corresponding discrete action probability value is obtained through normalization processing And where k is the number of discrete actions. When the agent policy network interacts with the environment, the discrete action index of the control agent of region i is: where g n is a random variable following a Gumbel (0,1) distribution, subject to the following random variable distribution: g n = -log(-log(ε)) ε ∈ U(0, 1) wherein ε is a uniform distribution random variable in the interval (0, 1), 0 in the uniform distribution is the lower limit of the uniform distribution, and 1 is the upper limit of the uniform distribution; mapping the discrete action index to the discrete action space of the agent to output the final control instruction value.
7. The method of claim 6, wherein the method further comprises: The purpose of the multi-agent deep stochastic policy gradient algorithm is to obtain the optimal parameters of the agent policy network, so that the total expected reward value J is maximum: wherein T represents the total simulation time; and γ represents the reward discount factor; represents the self-learning reward function reward value at each time; s represents the state data sampled from the experience pool; D represents the experience pool; a is the real-time sampled action, and u is the policy network.
8. The method of claim 7, wherein the method further comprises: The gradient of the policy network parameters of the control agent of the region i is calculated according to the chain rule: where S is the observation value of the control agent of all regions, J i is the expected reward value of region i, is the gradient operation on the parameter θ in the expected reward value, is the control action with the maximum probability selected from the discrete action space output by the policy according to the Gumbel-Max mechanism, Q i is the Q value network, a i is the output action of the i th control agent, N is the total number of control agents, and ψ is the parameter of the critic network.
9. The method of claim 8, wherein the method further comprises: The calculation method for selecting the control action with the maximum probability from the discrete action space output by the policy according to the Gumbel-Max mechanism is as follows: wherein η represents a temperature parameter, controls the discrete degree of softmax, and k is the number of discrete actions.
10. The method of claim 8, wherein the method further comprises: The update of the parameter ψ of the critic network of the control agent is realized by minimizing the loss function L(ψ), and the specific form is as follows: L(ψ) = E S~D,a~u [(Q i (S,a1,···,a2,a N |ψ)-y i ) 2 ] where S' = {s'1, s'2, ···, s'N} represents the set of next time step observation states of all regional control agents, s'i is the next time step observation state of the i-th regional control agent; Q N i i (·|ψ) represents the critic network to receive the observation states and action information of all control agents in the power system environment and evaluate the behavior of the current control agent; S represents the observation values of all regional control agents; and represent the target critic network and the target policy network, respectively; y i represents the target value.
Citation Information
Patent Citations
An intelligent generation control method based on deep reinforcement learning with self-optimizing ability
CN109217306A