Micro-grid group energy management method and system
By building a cost-minimizing energy management model in microgrid group energy management and improving the SAC algorithm, the problem of low value estimation deviation and convergence accuracy in traditional methods is solved, and the optimal energy scheduling and cost optimization of microgrid group energy storage system is achieved.
Patent Information
- Application Number
- CN202510093296.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-23
AI Technical Summary
The traditional microgrid energy management method has a value estimation bias due to the reinforcement learning algorithm originating from Q learning and fails to effectively optimize sample selection, resulting in low algorithm convergence accuracy.
A microgrid group energy management method is proposed. By constructing an energy management model of energy storage system with the goal of cost minimization, and constructing an improved SAC algorithm, reducing the estimation deviation through priority sorting and triple critical mechanisms, optimizing actions, states and reward values to obtain optimal energy scheduling decisions.
Through the improved SAC algorithm, the estimation deviation is reduced, the algorithm convergence accuracy is improved, the optimal energy scheduling decision of the microgrid group energy storage system is realized, and the operating cost is reduced.
Smart Images

Figure CN120033674A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microgrid energy management, and more specifically to a microgrid group energy management method and system. Background Art
[0002] A microgrid cluster is a collection of multiple microgrids that can be interconnected and coordinated to form a larger and more complex energy system. By combining multiple microgrids to form a microgrid cluster, resource sharing and complementary utilization can be achieved, the efficiency and reliability of the overall system can be improved, and the problems faced by microgrids, such as energy waste and system instability, can be solved.
[0003] A microgrid group consisting of multiple microgrids can effectively absorb distributed power sources, improve power supply flexibility and reliability, reduce the rate of abandoned solar and wind power, and reduce economic costs. The energy management model of the microgrid energy storage system has become one of the key issues in the comprehensive control of microgrids.
[0004] At present, traditional microgrid energy management methods have value estimation biases due to the fact that reinforcement learning algorithms derived from Q-learning have no consideration of optimizing reinforcement learning samples and uniformly sampling with the same probability, which leads to large value estimation biases in reinforcement learning algorithms and affects the convergence accuracy of the algorithm. Summary of the invention
[0005] In response to the problems existing in the above-mentioned fields, the present invention proposes a microgrid group energy management method and system, which constructs an energy management model of an energy storage system with the goal of minimizing cost, and constructs an improved SAC algorithm to reduce the estimation deviation. The improved SAC algorithm is used to optimize the cost by determining the optimal action, state and reward value, and obtains the optimal energy scheduling decision of the microgrid group energy storage system corresponding to the minimum cost.
[0006] In order to solve the above technical problems, the present invention discloses a microgrid group energy management method, comprising the following steps:
[0007] Collect sample data of interconnected microgrid groups;
[0008] Construct an energy management model for microgrid energy storage systems with the goal of minimizing costs;
[0009] Prioritize the collected sample data of the interconnected microgrid group by priority scoring to obtain an experience review set with priority sorting, set small batch samples for the experience review set in batches, and prioritize the small batch samples to obtain a small batch sample set with priority sorting;
[0010] An actor-critic double Q learning model for underestimation bias is constructed, and the model is trained on a small batch of sample sets. The underestimation bias of double Q learning in the actor-critic double Q learning model is corrected by introducing a triple criticism mechanism; the triple criticism mechanism determines the target value of triple criticism through the single critic network of DDPG and the double critic network of TD3; the target value of triple criticism is updated by constructing a triple criticism network to obtain the target value of the updated policy network;
[0011] According to the target value of the updated strategy network, the optimal action, state and reward value are determined; according to the optimal action, state and reward value, the cost is optimized by determining the action function, state function and reward function of the microgrid group, and the optimal energy scheduling decision of the microgrid group energy storage system corresponding to the minimum cost is obtained.
[0012] Preferably, obtaining a small batch sample set with priority order comprises the following steps:
[0013] The collected sample data of the interconnected microgrid group are prioritized and the ρ e represents the recent score of an episode in which all individual transitions have the same most recent score,
[0014] A collection of individual transitions at this stage In the mini-batch update, the most recent score of set E is:
[0015]
[0016] Where N represents the capacity of the experience review buffer, η e ∈(0,1) represents the hyperparameter that determines the most recent score of the data, ρ min represents the minimum allowed value of ρ; the closer the sample, the higher its recent score;
[0017] Use the reward value ρ of the final state of the episode f As the value scores of all transitions in this episode, the total priority score of this episode is ρ, ρ = ρ e +ρ f ;
[0018] When collecting experiences reviewing mini-batches in the buffer to update the agent, the transformations in the mini-batches need to be prioritized based on the priority scores;
[0019] Add the total priority score ρ to the tuple In the conversion tuple Two batches, namely and Extract from the experience review buffer, where m is the batch size;
[0020] The two batches H 1 With H 2 Merge and set a batch element set G of size m;
[0021] According to the priority score ρ of each transition tuple, H 1 With H 2 The 2m transition tuples in are sorted, and the first m transition tuples are selected to form a new batch element set G for network training.
[0022] Preferably, the actor-critic double Q learning model constructed for underestimation bias is:
[0023]
[0024] Among them, r is the reward value, which represents the immediate reward of the environment feedback after the current state s performs action a; γ represents the discount factor, which is used to weigh the importance of current rewards and future rewards, and its value range is 0<γ<1; represents the soft Q-value function estimated by the critic network; π φ (s′) represents the policy function that maps the distribution of the state space to the action space, and s′ represents the current state.
[0025] Preferably, the method of correcting the underestimation bias of the double Q learning in the actor-critic double Q learning model by introducing a triple criticism mechanism comprises the following steps:
[0026] The triple critic mechanism includes the single critic network of DDPG and the double critic network of TD3. By combining the overestimation bias of the single critic network of DDPG and the underestimation bias of the double critic network of TD3, the estimation bias is between the two.
[0027] By parameterizing the function Q θ (s,a) and π φ (a|s) estimate the soft Q value and strategy respectively. The Q value of the single critic network of DDPG is:
[0028]
[0029] Among them, π φ (·|s′) represents the policy function used to estimate the current state s′, Indicates that through π φ (·|s′) is the new action taken after estimation;
[0030] The Q value of the dual-criticism network of TD3 is:
[0031]
[0032] The target value of triple criticism is updated to:
[0033]
[0034] Among them, λ∈(0,1) is the weight of a single critic, α is the learning rate, which indicates the step size of each update; Take action for the current state s' strategy;
[0035] The first update to the policy network is expressed as:
[0036]
[0037] in, represents the state s under the state distribution D t Expectation, π φ (a t |s t ) in state s t Take action a t The policy function, Q θ (s t ,a t ) is represented as in state s t and take action a t The Q value after , represents the expected total return;
[0038] The essence of SAC strategy network optimization is to re-parameterize the strategy through neural network conversion. The action function of the re-parameterized strategy is:
[0039]
[0040] Among them, ε t is a noise vector that follows a normal distribution, and is the re-parameter sampling in state s t The mean and variance of the output, is the Hadamard product;
[0041] Substitute the action function of the re-parameterized strategy into the first updated strategy network to obtain the second updated strategy network J π (φ) is:
[0042]
[0043] Where N is the number of parameters in the neural network, f φ (ε t ;s t ) is the action function of the reparameterized strategy;
[0044] The policy gradient of the second updated policy network is expressed as:
[0045]
[0046] in, is the policy network gradient of the second update; π φ (a t |s t ) is state s t Next action space a t strategy; represents the gradient of the policy network parameter φ; Represents the action space a t gradient.
[0047] Preferably, obtaining the target value of the updated policy network specifically includes:
[0048] The constructed three-criticism network includes the target criticism network and the current criticism network, where:
[0049] Update the Q network according to the minimized Bellman residual:
[0050]
[0051] in, Indicates that under the state distribution D, t Take action a t expectations, Q θ (s t ,a t ) corresponds to the minimized Bellman residual, The optimal Q value corresponding to the next action and state;
[0052] The Q network gradient is expressed as:
[0053]
[0054] The target critic network is soft updated, expressed as:
[0055] θ′←(1-τ)θ′+τθ
[0056] Among them, τ is the soft update parameter, θ is the original Q value, and θ′ is the current estimated Q value.
[0057] Preferably, the synchronous update of the policy network and the criticism network is also included, specifically including:
[0058] The state at time t is s t, the Actor network is connected through a t Get action; the system executes a t Then get feedback from the environment t And move to the next state s t+1 ;
[0059] The experience replay method stores the experience gained during the system movement into the experience pool (s t ,a t ,r t ,s t+1 ), after each action is executed, the Critic network evaluates the state s at the next moment t+1 , to ensure whether the strategy achieves the expected effect, the calculation formula of DT-error is:
[0060] δ t =r t+1 +γQ(s t+1 ,a t+1 )-Q(s t ,a t )
[0061] Among them, δ t Indicates TD-error;
[0062] When TD-error is positive, it means that the system chooses a t The trend needs to be gradually strengthened; when TD-error is negative, it means that the system chooses a t The trend needs to be gradually weakened;
[0063] Use TD-error as the deviation evaluation indicator of the target value, measure the deviation of the current strategy by calculating TD-error, and dynamically adjust the strategy network;
[0064] Use DT-error as deviation feedback to optimize the strategy network output and ensure that the energy management model of the microgrid energy storage system converges to the optimal solution;
[0065] The agent extracts sample actions from the policy network and applies them to the environment. The environment updates its state based on the actions and provides feedback on the reward value.
[0066] Collect state, action, reward, and next state transition data, and store them in the experience replay pool;
[0067] The parameters of the current critic network in the constructed three-critic network are synchronized to the target critic network regularly; by updating the loss functions of the policy network and the critic network, the agent's action selection tends to the global optimum.
[0068] Preferably, the construction of the energy management model of the microgrid group energy storage system specifically includes:
[0069] Taking the minimization of the operating cost and environmental cost of the microgrid group as the goal, the objective function and constraints are determined, and the objective function of the energy management model of the energy storage system of the interconnected microgrid group is established as follows:
[0070] C=C 1 +C 2
[0071]
[0072] in, represents the diesel generator cost, represents the cost of the microturbine, represents the maintenance cost of power generation equipment, represents the operating cost of the energy storage system, represents the electricity transaction cost, C k is the cost factor for treating type k pollutants, including CO 2 、NO 2 and NO x ;
[0073] i is a different microgrid, t is time, is the weight of diesel generator k-type pollutants, is the weight of k-type pollutants from microturbines, is the weight of type k pollutants in the electricity purchasing process, P i DG (t) is the power of diesel generator in microgrid i, P i MT (t) is the power of micro-turbine in microgrid i, P i buy (t) is the power purchased by microgrid i;
[0074] The set constraints include power balance constraints, power generation equipment limitations, and energy storage system constraints, among which:
[0075] The power balance constraint is expressed as:
[0076] P i WT (t)+P i PV (t)+P i DG (t)+P i MT (t)+P i ESS (t)+P i mg (t)+Pi grid (t) = P i load (t)
[0077] Among them, P i WT (t) is the wind turbine power in microgrid i, P i PV (t) is the power of the photovoltaic motor in the i microgrid, P i grid (t) is the power required by the i microgrid to purchase electricity, P i ESS (t) is the power of the energy storage system in the i microgrid, P i mg (t) is the power transaction power of microgrid i, P i load (t) is the load power of microgrid i;
[0078] The power generation equipment limitation is expressed as:
[0079]
[0080] in, is the minimum power generated by the diesel generator in the i microgrid per unit time, is the maximum power generated by the diesel generator in the i microgrid per unit time, is the minimum power generated by the micro-turbine in the i microgrid per unit time, is the maximum power generated by the microturbine in the i microgrid per unit time;
[0081] Energy storage system constraints include state of charge constraints and power limits for charging and discharging;
[0082] The charge state constraint is expressed as:
[0083]
[0084] in, is the minimum charge value of the i microgrid, SOC(t) is the charge state at time t, is the maximum value of the charge of the i microgrid;
[0085] The power limit for charging and discharging is expressed as:
[0086]
[0087] Among them, P i ch (t) is the charging power of microgrid i, P i max is the maximum power of microgrid i, P i dis(t) is the discharge power of microgrid i.
[0088] Preferably, the optimal energy dispatching decision of the microgrid group energy storage system corresponding to the minimum acquisition cost specifically includes:
[0089] According to the target value of the updated policy network, determine the optimal action a at the current moment t , status t and the reward value r t ;
[0090] The state function is used to describe the operating state of the microgrid group, reflecting the output power of each distributed energy source, the state parameters of the battery and the load demand. The state function is expressed as:
[0091] s t = {P WT (t),P PV (t),P MT (t),P DG (t),SOC(t),P load (t),Ctou(t)}
[0092] Where Ctou(t) is the electricity price in period t;
[0093] The action function represents the control strategy, that is, the energy dispatch decision of the system at the current moment, including the output adjustment of power generation equipment, energy storage charging and discharging strategy and power purchase decision. The action function is expressed as:
[0094] a t = {P i DG (t), P i MT (t),P i ch (t),P i dis (t),P i buy (t)}
[0095] The reward function is used to reflect the degree of optimization of the target, focusing on cost minimization and system performance improvement. At the same time, it is necessary to consider pollution emission constraints and system safety constraints. The reward value is set to a negative objective function value. The reward function is expressed as:
[0096] r=-C=-(C 1 +C 2 )
[0097] According to the state, action and reward function in reinforcement learning, the objective function of the energy management model of the energy storage system of the constructed interconnected microgrid group is solved.
[0098] Preferably, a microgrid group energy management system is also included, including:
[0099] A data collection module, used to collect sample data of the interconnected microgrid group;
[0100] Energy management model building module, used to build a microgrid group energy storage system energy management model with the goal of minimizing cost;
[0101] The SAC algorithm improvement module is used to prioritize the sample data of the interconnected microgrid group collected in the original SAC algorithm through priority scoring, obtain an experience review set with priority sorting, set small batch samples for the experience review set in batches, and prioritize the small batch samples to obtain a small batch sample set with priority sorting; construct an actor-critic double Q learning model for underestimation bias, train the model on the small batch sample set, and correct the underestimation bias of the double Q learning in the actor-critic double Q learning model by introducing a triple criticism mechanism; wherein the triple criticism mechanism is to determine the target value of the triple criticism through the single critic network of DDPG and the double criticism network of TD3; update the target value of the triple criticism by constructing a triple criticism network to obtain the target value of the updated strategy network;
[0102] The energy management model solving module is used to determine the optimal action, state and reward value according to the target value of the updated strategy network; according to the optimal action, state and reward value, the cost is optimized by determining the action function, state function and reward function of the microgrid group, and the optimal energy scheduling decision of the microgrid group energy storage system corresponding to the minimum cost is obtained.
[0103] Compared with the prior art, the present invention has the following beneficial effects:
[0104] The present invention proposes a microgrid group energy management method, which constructs an energy storage system energy management model with the goal of minimizing cost; through the constructed improved SAC algorithm, the action function, state function and reward function of the microgrid group are determined to optimize the cost, and the objective function of the microgrid group is solved to solve the energy strategy optimization scheduling problem of the microgrid group. Among them, the improved SAC algorithm is to prioritize the sample data of the interconnected microgrid group collected by the original SAC algorithm through priority scoring, obtain the experience review set with priority sorting and set small batch samples in batches, prioritize the small batch samples, and select better samples from the experience replay buffer through the priority sorting scheme. The actor-critic double Q learning model constructed for underestimation deviation can combine the overestimation deviation of DDPG and the underestimation deviation of TD3 by introducing a triple criticism mechanism, so that the estimation deviation can be between the two, thereby reducing the estimation error and improving the convergence accuracy of the algorithm. Through the constructed three-critic network, the triple criticism target value determined by the triple criticism mechanism is updated, and then the model is solved, and the optimal microgrid group energy scheduling strategy is output by updating the state, state and reward value. BRIEF DESCRIPTION OF THE DRAWINGS
[0105] Figure 1 This is a framework diagram of the microgrid group optimization scheduling method proposed by the present invention;
[0106] Figure 2 A framework diagram of a microgrid group optimization scheduling method based on deep reinforcement learning provided in an embodiment of the present invention;
[0107] Figure 3 A diagram showing actual energy scheduling results of a microgrid group using a microgrid group optimization scheduling method based on deep reinforcement learning provided in an embodiment of the present invention;
[0108] Figure 4 Training curves of four DRL algorithms (TCSAC, SAC, TD3, DDPG) corresponding to Table 1 provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0109] The following will be combined with the attached embodiment of the present invention Figure 1-Figure 4 , the technical solutions in the embodiments of the present invention are clearly and completely described. It should be understood that the terms described in the present invention are only used to describe specific implementation methods and are not used to limit the present invention.
[0110] like Figure 1 As shown in FIG. 1 , a flow chart of a microgrid group energy management method based on deep reinforcement learning proposed by the present invention is shown in FIG. 1 , and the method specifically includes the following steps:
[0111] S1: Collect sample data of interconnected microgrid groups;
[0112] S2: Construct an energy management model for microgrid energy storage systems with the goal of minimizing costs;
[0113] S3: In the original SAC algorithm, the collected sample data of the interconnected microgrid group are prioritized by priority scoring to obtain an experience review set with priority sorting, and small batch samples are set in batches for the experience review set, and the small batch samples are prioritized to obtain a small batch sample set with priority sorting;
[0114] S4: Construct an actor-critic double Q learning model for underestimation bias, train the model on a small batch of samples, and introduce a triple criticism mechanism to correct the underestimation bias of double Q learning in the actor-critic double Q learning model; the triple criticism mechanism determines the target value of triple criticism through the single critic network of DDPG and the double critic network of TD3; update the target value of triple criticism by constructing a triple critic network, and obtain the target value of the updated policy network;
[0115] S5: Determine the optimal action, state and reward value based on the target value of the updated strategy network; based on the optimal action, state and reward value, optimize the cost by determining the action function, state function and reward function of the microgrid group, and obtain the optimal energy scheduling decision of the microgrid group energy storage system corresponding to the minimum cost.
[0116] Specifically, in step S2, the objective function and constraints are determined with the goal of minimizing the operating cost and environmental cost of the microgrid group. The objective function of the energy management model of the energy storage system of the interconnected microgrid group is established as follows:
[0117] C=C 1 +C 2
[0118]
[0119] in, represents the diesel generator cost, represents the cost of the microturbine, represents the maintenance cost of power generation equipment, represents the operating cost of the energy storage system, represents the electricity transaction cost, C k is the cost factor for treating type k pollutants, including CO 2 、NO 2 and NO x .
[0120] i is a different microgrid, t is time, is the weight of diesel generator k-type pollutants, is the weight of k-type pollutants from microturbines, is the weight of type k pollutants in the electricity purchasing process, P i DG (t) is the power of diesel generator in microgrid i, P i MT (t) is the power of micro-turbine in microgrid i, P i buy (t) is the power purchased by microgrid i.
[0121] The set constraints include power balance constraints, power generation equipment limitations, and energy storage system constraints, among which:
[0122] The power balance constraint is expressed as:
[0123] P i WT (t)+P i PV (t)+P i DG (t)+P i MT (t)+P i ESS (t)+P i mg (t)+P i grid (t) = P i load (t)
[0124] Among them, P i WT (t) is the wind turbine power in microgrid i, P i PV (t) is the power of the photovoltaic motor in the i microgrid, P i grid (t) is the power required by the i microgrid to purchase electricity, P i ESS (t) is the power of the energy storage system in the i microgrid, P i mg (t) is the power transaction power of microgrid i, P i load (t) is the load power of microgrid i.
[0125] The power generation equipment limitation is expressed as:
[0126]
[0127] in, is the minimum power generated by the diesel generator in the i microgrid per unit time, is the maximum power generated by the diesel generator in the i microgrid per unit time, is the minimum power generated by the micro-turbine in the i microgrid per unit time, It is the maximum power generated by the micro-turbine in microgrid i per unit time.
[0128] Energy storage system constraints include state of charge constraints and charging and discharging power limits:
[0129] The charge state constraint is expressed as:
[0130]
[0131] in, is the minimum charge value of the i microgrid, SOC(t) is the charge state at time t, is the maximum charge of the i microgrid.
[0132] The power limit for charging and discharging is expressed as:
[0133]
[0134] Among them, P i ch (t) is the charging power of microgrid i, P i max is the maximum power of microgrid i, P i dis (t) is the discharge power of microgrid i.
[0135] It also includes obtaining the reward value reflecting the current action and establishing the state, action, and reward functions in reinforcement learning, where:
[0136] The state function is used to describe the operating state of the microgrid group, reflecting the output power of each distributed energy source, the state parameters of the battery and the load demand. The state function is expressed as:
[0137] s t = {P WT (t),P PV (t),P MT (t),P DG (t),SOC(t),P load (t),Ctou(t)}
[0138] Where Ctou(t) is the electricity price during period t.
[0139] The action function represents the control strategy, that is, the energy dispatch decision of the system at the current moment, including the output adjustment of power generation equipment, energy storage charging and discharging strategy and power purchase decision. The action function is expressed as:
[0140] a t = {Pi DG (t), P i MT (t),P i ch (t),P i dis (t),P i buy (t)}
[0141] The reward function is used to reflect the degree of optimization of the target, focusing on cost minimization and system performance improvement. At the same time, it is necessary to consider pollution emission constraints and system safety constraints. The reward value is set to a negative objective function value. The reward function is expressed as:
[0142] r=-C=-(C 1 +C 2 )
[0143] According to the state, action and reward function in reinforcement learning, the objective function of the energy management model of the energy storage system of the constructed interconnected microgrid group is solved.
[0144] Obtaining a prioritized mini-batch sample set includes the following steps:
[0145] The collected sample data of the interconnected microgrid group are prioritized and the ρ e represents the recent score of an episode in which all individual transitions have the same most recent score,
[0146] A collection of individual transitions at this stage In the mini-batch update, the most recent score of set E is:
[0147]
[0148] Where N represents the capacity of the experience review buffer, η e ∈(0,1) represents the hyperparameter that determines the most recent score of the data, ρ min represents the minimum allowed value of ρ; the more recent the sample, the higher its recent score.
[0149] Use the reward value ρ of the final state of the episode f As the value scores of all transitions in this episode, the total priority score of this episode is ρ, ρ = ρ e +ρ f .
[0150] When the collection experience looks back at the mini-batches within the buffer to update the agent, the transformations within the mini-batches need to be prioritized based on the priority scores.
[0151] Add the total priority score ρ to the tuple In the conversion tuple Two batches, namely and Extracted from the experience review buffer, where m is the batch size.
[0152] The two batches H 1 With H 2 Merge and set a batch element set G of size m.
[0153] According to the priority score ρ of each transition tuple, H 1 With H 2 The 2m transition tuples in are sorted, and the first m transition tuples are selected to form a new batch element set G for network training.
[0154] The actor-critic double Q-learning model for underestimation bias is constructed as:
[0155]
[0156] Among them, r is the reward value, which represents the immediate reward of the environment feedback after the current state s performs action a; γ represents the discount factor, which is used to weigh the importance of current rewards and future rewards, and its value range is 0<γ<1; represents the soft Q-value function estimated by the critic network; π φ (s′) represents the policy function that maps the distribution of the state space to the action space, and s′ represents the current state.
[0157] The underestimation bias of double Q learning in the actor-critic double Q learning model is corrected by introducing a triple criticism mechanism, which includes the following steps:
[0158] The triple-critic mechanism includes a single-critic network of DDPG and a double-critic network of TD3. By combining the overestimation bias of the single-critic network of DDPG and the underestimation bias of the double-critic network of TD3, the estimation bias is between the two.
[0159] By parameterizing the function Q θ (s,a) and π φ (a|s) estimate the soft Q value and strategy respectively. The Q value of the single critic network of DDPG is:
[0160]
[0161] Among them, π φ (·|s′) represents the policy function used to estimate the current state s′, Indicates that through π φ(·|s′) is the new action taken after estimation.
[0162] The Q value of the dual-criticism network of TD3 is:
[0163]
[0164] The target value of triple criticism is updated to:
[0165]
[0166] Among them, λ∈(0,1) is the weight of a single critic, α is the learning rate, which indicates the step size of each update; Take action for the current state s' strategy.
[0167] The first update to the policy network is expressed as:
[0168]
[0169] in, represents the state s under the state distribution D t Expectation, π φ (a t |s t ) in state s t Take action a t The policy function, Q θ (s t ,a t ) is represented as in state s t and take action a t The Q value after that represents the expected total return.
[0170] The essence of SAC strategy network optimization is to re-parameterize the strategy through neural network conversion. The action function of the re-parameterized strategy is:
[0171]
[0172] Among them, ε t is a noise vector that follows a normal distribution, and For reparameter sampling in state s t The mean and variance of the output, is the Hadamard product.
[0173] Substitute the action function of the re-parameterized strategy into the first updated strategy network to obtain the second updated strategy network J π (φ) is:
[0174]
[0175] Where N is the number of parameters in the neural network, f φ (ε t ;s t ) is the action function of the reparameterized strategy.
[0176] The policy gradient of the second updated policy network is expressed as:
[0177]
[0178] in, is the policy network gradient of the second update; π φ (a t |s t ) is state s t Next action space a t strategy; represents the gradient of the policy network parameter φ; Represents the action space a t gradient.
[0179] Get the target value of the updated policy network, including:
[0180] The constructed three-criticism network includes the target criticism network and the current criticism network, where:
[0181] Update the Q network according to the minimized Bellman residual:
[0182]
[0183] in, Indicates that under the state distribution D, t Take action a t expectations, Q θ (s t ,a t ) corresponds to the minimized Bellman residual, The optimal Q value corresponding to the next action and state.
[0184] The Q network gradient is expressed as:
[0185]
[0186] The target critic network is soft updated, expressed as:
[0187] θ′←(1-τ)θ′+τθ
[0188] Among them, τ is the soft update parameter, θ is the original Q value, and θ′ is the current estimated Q value.
[0189] The policy network and the critic network are updated synchronously, including:
[0190] The state at time t is s t , the Actor network is connected through a t Get action; the system executes a t Then get feedback from the environment t And move to the next state s t+1 ;
[0191] The experience replay method stores the experience gained during the system movement into the experience pool (s t ,a t ,r t ,s t+1 ), after each action is executed, the Critic network evaluates the state s at the next moment t+1 , to ensure whether the strategy achieves the expected effect, the calculation formula of DT-error is:
[0192] δ t =r t+1 +γQ(s t+1 ,a t+1 )-Q(s t ,a t )
[0193] Among them, δ t Indicates TD-error;
[0194] When TD-error is positive, it means that the system chooses a t The trend needs to be gradually strengthened; when TD-error is negative, it means that the system chooses a t The trend needs to be gradually weakened;
[0195] Use TD-error as the deviation evaluation indicator of the target value, measure the deviation of the current strategy by calculating TD-error, and dynamically adjust the strategy network;
[0196] Use DT-error as deviation feedback to optimize the strategy network output and ensure that the energy management model of the microgrid energy storage system converges to the optimal solution;
[0197] The agent extracts sample actions from the policy network and applies them to the environment. The environment updates its state based on the actions and provides feedback on the reward value.
[0198] Collect state, action, reward, and next state transition data, and store them in the experience replay pool;
[0199] Regularly synchronize the parameters of the current critic network in the three constructed critic networks to the target critic network; by updating the loss functions of the policy network and the critic network, make the action selection of the Agent tend to be globally optimal.
[0200] The present invention also proposes a microgrid group energy management system, including:
[0201] A data collection module for collecting sample data of the interconnected microgrid group;
[0202] An energy management model construction module for constructing an energy management model of the microgrid group energy storage system with the goal of minimizing cost;
[0203] An SAC algorithm improvement module, in the original SAC algorithm, through priority ranking of the collected sample data of the interconnected microgrid group by priority scoring, obtain an experience replay set with priority ranking, set small batch samples in batches for the experience replay set, and perform priority ranking on the small batch samples to obtain a small batch sample set with priority ranking; construct an actor-critic double Q learning model for underestimation bias, train the small batch sample set with this model, and correct the underestimation bias existing in the double Q learning in the actor-critic double Q learning model by introducing a triple critic mechanism; wherein, the triple critic mechanism is to determine the target value of the triple critic through the single critic network of DDPG and the double critic network of TD3; update the target value of the triple critic by constructing a three-critic network to obtain the target value of the updated policy network;
[0204] An energy management model solving module for determining the optimal action, state and reward value according to the target value of the updated policy network; optimize the cost by determining the action function, state function and reward function of the microgrid group according to the optimal action, state and reward value, and obtain the optimal energy scheduling decision of the microgrid group energy storage system corresponding to the minimum cost.
[0205] The improved SAC algorithm constructed by the microgrid group energy management method proposed by the present invention can reduce the estimation bias, optimize the cost by determining the action function, state function and reward function of the microgrid group, and obtain the optimal energy scheduling decision of the microgrid group energy storage system corresponding to the minimum cost.
[0206] Embodiment
[0207] To verify the proposed microgrid group energy management method, the present invention is verified with the following embodiments. This embodiment provides a microgrid group optimal scheduling method based on deep reinforcement learning, and the method framework of this method consists of an actor network, a three-critic network and a three-target critic network.
[0208] Such as Figure 2As shown in the figure, the actor network interacts with the environment through the agent, and the environment provides the agent with information about the current operating status of the system, including the voltage regulation power of the energy storage system, the voltage amplitude of the load bus, the status of power trading, etc. The actor network in the agent outputs the mean and variance of the action-related Gaussian distribution based on the data provided by the environment and the experience replay buffer through re-parameter sampling.
[0209] A sample is drawn from the policy distribution to generate an action value, which is applied to the environment, and the environment then gets a reward value reflecting the reward for the current action. After the environment gets the action command, it runs to the next state and gets a set of transitions, which are stored in the experience review buffer.
[0210] The actor network and the critic network are trained by selecting small batches from the experience replay buffer, and then the estimation bias is reduced by the three critic networks to obtain the TD-error, which is fed back to the actor network. This process is repeated until the agent can make the optimal decision, thus obtaining the final scheduling strategy.
[0211] like Figure 3 As shown, this is a diagram of the actual energy scheduling result of the microgrid group using the improved deep reinforcement learning provided by the present invention. Combined with the real-time status data of the microgrid group and the electricity price of the distribution network, it can be seen that when the power output is large, the microgrid group has a surplus of electricity. At this time, the energy storage system is in a charging state to balance the system power. At the same time, the demand-side response is also in a state of accepting load shifting, which can absorb renewable energy to the greatest extent, that is, the energy storage system achieves maximum economic benefits while meeting power requirements.
[0212] The energy storage system also responds significantly to changes in electricity prices. When the electricity price is in the lowest range, the energy storage continues to charge and remain fully charged. When the electricity price is high, it continues to discharge, alleviating the power shortage problem during peak hours and reducing the electricity purchase cost of the microgrid. When the electricity price is in the middle range, the power generation and load fluctuations of the microgrid are large. At this time, the energy management model optimizes the power configuration by finely controlling the energy storage system, thereby maximizing the economic benefits in real time.
[0213] As shown in Table 1, different indicators of four DRL algorithms (TSCAC, SAC, TD3, DDPG) are compared. The present invention uses four indicators to quantify the learning performance, namely, the final average reward, the final standard deviation, the maximum episode reward and the maximum cumulative reward.
[0214] Table 1 Comparison of different indicators of four DRL algorithms: TSCAC, SAC, TD3, and DDPG
[0215] index DDPG TD3 SAC TCSAC <![CDATA[I UAR ]]> 69.32 81.12 78.69 82.46 <![CDATA[I USD ]]> 2.94 3.64 3.61 3.53 <![CDATA[I MER ]]> 102.49 111.36 108.96 112.64 <![CDATA[I MCR ]]> 73.26 81.49 82.01 82.61 Training time 69874.2 76379.6 59888.6 80390.2
[0216] Compared with the SAC algorithm, the proposed microgrid group optimization scheduling method, denoted as (Based on ThreeCritics Soft Actor-Critic, TCSAC), referred to as TCSAC, has improved final cumulative rewards and standard deviations, indicating that the final performance and stability of the algorithm have been improved accordingly. In addition, the maximum episode reward and maximum cumulative reward of TCSAC are the highest, indicating that this method can explore scheduling solutions that are superior to other algorithms.
[0217] As for the training time, the TD3 algorithm takes the least time, which is 59,888.6 seconds, or 6 seconds per episode on average, for training one thousand episodes; the DDPG algorithm takes the most time, which is 80390.2 seconds. TCSAC takes the second least time, but because it has better training results, it has a stronger ability to make decisions later, and has better performance than TD3. It can be seen that TCSAC has stronger performance.
[0218] like Figure 4 As shown in Figure 1, the training curves of the four DRL algorithms (TCSAC, SAC, TD3, and DDPG) are corresponding to Table 1. It can be seen that SAC, TD3, and DDPG start to converge around 20 sets with a training set of 100, while TCSAC starts to converge around 40 sets. However, TCSAC has the highest cumulative reward value, and DDPG has the lowest average reward value. TCSAC has a better convergence effect, with a higher maximum episode reward and maximum cumulative reward. Although the training is slower, the convergence accuracy is higher.
[0219] It can be seen from this embodiment that the microgrid group optimization scheduling method proposed in the present invention can explore scheduling solutions that are superior to other algorithms, and the improved SAC algorithm constructed has higher convergence accuracy.
[0220] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
[0221] In addition, unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as commonly understood by those skilled in the art to which the present invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.
Claims
1. A microgrid group energy management method, characterized in that: The following steps are involved: Collect sample data of interconnected microgrid groups; Construct an energy management model for microgrid energy storage systems with the goal of minimizing costs; Prioritize the collected sample data of the interconnected microgrid group by priority scoring to obtain an experience review set with priority sorting, set small batch samples for the experience review set in batches, and prioritize the small batch samples to obtain a small batch sample set with priority sorting; An actor-critic double Q learning model for underestimation bias is constructed, and the model is trained on a small batch of sample sets. The underestimation bias of double Q learning in the actor-critic double Q learning model is corrected by introducing a triple criticism mechanism; the triple criticism mechanism determines the target value of triple criticism through the single critic network of DDPG and the double critic network of TD3; the target value of triple criticism is updated by constructing a triple criticism network to obtain the target value of the updated policy network; According to the target value of the updated strategy network, the optimal action, state and reward value are determined; according to the optimal action, state and reward value, the cost is optimized by determining the action function, state function and reward function of the microgrid group, and the optimal energy scheduling decision of the microgrid group energy storage system corresponding to the minimum cost is obtained.
2. The microgrid group energy management method according to claim 1, characterized in that: The step of obtaining a small batch sample set with priority sorting comprises the following steps: The collected sample data of the interconnected microgrid group are prioritized and the ρ e represents the recent score of an episode in which all individual transitions have the same most recent score, A collection of individual transitions at this stage In the mini-batch update, the most recent score of set E is: Where N represents the capacity of the experience review buffer, η e ∈(0,1) represents the hyperparameter that determines the most recent score of the data, ρ min represents the minimum allowed value of ρ; the closer the sample, the higher its recent score; Use the reward value ρ of the final state of the episode f As the value scores of all transitions in this episode, the total priority score of this episode is ρ, ρ = ρ e +ρ f ; When collecting experiences reviewing mini-batches in the buffer to update the agent, the transformations in the mini-batches need to be prioritized based on the priority scores; Add the total priority score ρ to the tuple In the conversion tuple Two batches, namely and Extract from the experience review buffer, where m is the batch size; Merge the two extracted batches H1 and H2, and set a batch element set G of size m; According to the priority score ρ of each transition tuple, the 2m transition tuples in H1 and H2 are sorted, and the first m transition tuples are selected to form a new batch element set G for network training.
3. The microgrid group energy management method according to claim 2, characterized in that: The actor-critic double Q-learning model constructed for underestimation bias is: Among them, r is the reward value, which represents the immediate reward of the environment feedback after the current state s performs action a; γ represents the discount factor, which is used to weigh the importance of current rewards and future rewards, and its value range is 0<γ<1; represents the soft Q-value function estimated by the critic network; π φ (s′) represents the policy function that maps the distribution of the state space to the action space, and s′ represents the current state.
4. The microgrid group energy management method according to claim 3, characterized in that: The method of correcting the underestimation bias of double Q learning in the actor-critic double Q learning model by introducing a triple criticism mechanism includes the following steps: The triple critic mechanism includes the single critic network of DDPG and the double critic network of TD3. By combining the overestimation bias of the single critic network of DDPG and the underestimation bias of the double critic network of TD3, the estimation bias is between the two. By parameterizing the function Q θ (s,a) and π φ (a|s) estimate the soft Q value and strategy respectively. The Q value of the single critic network of DDPG is: Among them, π φ (·|s′) represents the policy function used to estimate the current state s′, Indicates that through π φ (·|s′) is the new action taken after estimation; The Q value of the dual-criticism network of TD3 is: The target value of triple criticism is updated to: Among them, λ∈(0,1) is the weight of a single critic, α is the learning rate, which indicates the step size of each update; Take action for the current state s' strategy; The first update to the policy network is expressed as: in, represents the state s under the state distribution D t Expectation, π φ (a t |s t ) in state s t Take action a t The policy function, Q θ (s t ,a t ) is represented as in state s t and take action a t The Q value after , represents the expected total return; The essence of SAC strategy network optimization is to re-parameterize the strategy through neural network conversion. The action function of the re-parameterized strategy is: Among them, ε t is a noise vector that follows a normal distribution, and For reparameter sampling in state s t The mean and variance of the output, is the Hadamard product; Substitute the action function of the re-parameterized strategy into the first updated strategy network to obtain the second updated strategy network J π (φ) is: Where N is the number of parameters in the neural network, f φ (ε t ;s t ) is the action function of the reparameterized strategy; The policy gradient of the second updated policy network is expressed as: in, is the policy network gradient of the second update; π φ (a t |s t ) is state s t Next action space a t strategy; represents the gradient of the policy network parameter φ; Represents the action space a t gradient.
5. The microgrid group energy management method according to claim 4, characterized in that: The obtaining of the target value of the updated policy network specifically includes: The constructed three-criticism network includes the target criticism network and the current criticism network, where: Update the Q network according to the minimized Bellman residual: in, Indicates that under the state distribution D, t Take action a t expectations, Q θ (s t ,a t ) corresponds to the minimized Bellman residual, The optimal Q value corresponding to the next action and state; The Q network gradient is expressed as: The target critic network is soft updated, expressed as: θ′←(1-τ)θ′+τθ Among them, τ is the soft update parameter, θ is the original Q value, and θ′ is the current estimated Q value.
6. The microgrid group energy management method according to claim 5, characterized in that: It also includes the synchronous update of the strategy network and the criticism network, including: The state at time t is s t , the Actor network is connected through a t Get action; the system executes a t Then get feedback from the environment t And move to the next state s t+1 ; The experience replay method stores the experience gained during the system movement into the experience pool (s t ,a t ,r t ,s t+1 ), after each action is executed, the Critic network evaluates the state s at the next moment t+1 , to ensure whether the strategy achieves the expected effect, the calculation formula of DT-error is: δ t =r t+1 +γQ(s t+1 ,a t+1 )-Q(s t ,a t ) Among them, δ t Indicates TD-error; When TD-error is positive, it means that the system chooses a t The trend needs to be gradually strengthened; when TD-error is negative, it means that the system chooses a t The trend needs to be gradually weakened; Use TD-error as the deviation evaluation indicator of the target value, measure the deviation of the current strategy by calculating TD-error, and dynamically adjust the strategy network; Use DT-error as deviation feedback to optimize the strategy network output and ensure that the energy management model of the microgrid energy storage system converges to the optimal solution; The agent extracts sample actions from the policy network and applies them to the environment. The environment updates its state based on the actions and provides feedback on the reward value. Collect state, action, reward, and next state transition data, and store them in the experience replay pool; The parameters of the current critic network in the constructed three-critic network are synchronized to the target critic network regularly; by updating the loss functions of the policy network and the critic network, the agent's action selection tends to the global optimum.
7. The microgrid group energy management method according to claim 6, characterized in that: The construction of the energy management model of the microgrid energy storage system specifically includes: Taking the minimization of the operating cost and environmental cost of the microgrid group as the goal, the objective function and constraints are determined, and the objective function of the energy management model of the energy storage system of the interconnected microgrid group is established as follows: C=C1+C2 in, represents the diesel generator cost, represents the cost of the microturbine, represents the maintenance cost of power generation equipment, represents the operating cost of the energy storage system, represents the electricity transaction cost, C k is the cost factor for treating type k pollutants, including CO2, NO2 and NO x ; i is a different microgrid, t is time, is the weight of diesel generator k-type pollutants, is the weight of k-type pollutants from microturbines, is the weight of type k pollutants in the electricity purchasing process, P i DG (t) is the power of diesel generator in microgrid i, P i MT (t) is the power of micro-turbine in microgrid i, P i buy (t) is the power purchased by microgrid i; The set constraints include power balance constraints, power generation equipment limitations, and energy storage system constraints, among which: The power balance constraint is expressed as: P i WT (t)+P i PV (t)+P i DG (t)+P i MT (t)+P i ESS (t)+P i mg (t)+P i grid (t)=P i load (t) Among them, P i WT (t) is the wind turbine power in microgrid i, P i PV (t) is the power of the photovoltaic motor in the i microgrid, P i grid (t) is the power required by the i microgrid to purchase electricity, P i ESS (t) is the power of the energy storage system in the i microgrid, P i mg (t) is the power transaction power of microgrid i, P i load (t) is the load power of microgrid i; The power generation equipment limitation is expressed as: in, is the minimum power generated by the diesel generator in the i microgrid per unit time, is the maximum power generated by the diesel generator in the i microgrid per unit time, is the minimum power generated by the micro-turbine in the i microgrid per unit time, is the maximum power generated by the microturbine in the i microgrid per unit time; Energy storage system constraints include state of charge constraints and power limits for charging and discharging; The charge state constraint is expressed as: in, is the minimum charge value of the i microgrid, SOC(t) is the charge state at time t, is the maximum value of the charge of the i microgrid; The power limit for charging and discharging is expressed as: Among them, P i ch (t) is the charging power of microgrid i, P i max is the maximum power of microgrid i, P i dis (t) is the discharge power of microgrid i.
8. The microgrid group energy management method according to claim 7, characterized in that: The optimal energy dispatching decision of the microgrid group energy storage system corresponding to the minimum acquisition cost specifically includes: According to the target value of the updated policy network, determine the optimal action a at the current moment t , status t and the reward value r t ; The state function is used to describe the operating state of the microgrid group, reflecting the output power of each distributed energy source, the state parameters of the battery and the load demand. The state function is expressed as: s t ={P WT (t),P PV (t),P MT (t),P DG (t),SOC(t),P load (t),Ctou(t)} Where Ctou(t) is the electricity price in period t; The action function represents the control strategy, that is, the energy dispatch decision of the system at the current moment, including the output adjustment of power generation equipment, energy storage charging and discharging strategy and power purchase decision. The action function is expressed as: a t ={P i DG (t),P i MT (t),P i ch (t),P i dis (t),P i buy (t)} The reward function is used to reflect the degree of optimization of the target, focusing on cost minimization and system performance improvement. At the same time, it is necessary to consider pollution emission constraints and system safety constraints. The reward value is set to a negative objective function value. The reward function is expressed as: r=-C=-(C1+C2) According to the state, action and reward function in reinforcement learning, the objective function of the energy management model of the energy storage system of the constructed interconnected microgrid group is solved.
9. A microgrid group energy management system, characterized in that: include: A data collection module, used to collect sample data of the interconnected microgrid group; Energy management model building module, used to build a microgrid group energy storage system energy management model with the goal of minimizing cost; The SAC algorithm improvement module is used to prioritize the sample data of the interconnected microgrid group collected in the original SAC algorithm through priority scoring, obtain an experience review set with priority sorting, set small batch samples for the experience review set in batches, and prioritize the small batch samples to obtain a small batch sample set with priority sorting; construct an actor-critic double Q learning model for underestimation bias, train the model on the small batch sample set, and correct the underestimation bias of the double Q learning in the actor-critic double Q learning model by introducing a triple criticism mechanism; wherein the triple criticism mechanism is to determine the target value of the triple criticism through the single critic network of DDPG and the double criticism network of TD3; update the target value of the triple criticism by constructing a triple criticism network to obtain the target value of the updated strategy network; The energy management model solving module is used to determine the optimal action, state and reward value according to the target value of the updated strategy network; according to the optimal action, state and reward value, the cost is optimized by determining the action function, state function and reward function of the microgrid group, and the optimal energy scheduling decision of the microgrid group energy storage system corresponding to the minimum cost is obtained.
Citation Information
Cited By
Oxidation-reduction free radical synergistic electrochemical wastewater treatment method and system
CN121158908A
Micro-grid energy management method and system based on reinforcement learning
CN121689158A