Method and device for determining micro-grid energy storage operation scheduling strategy
Through the improved GRU neural network and Markov decision-making Q learning algorithm, the microgrid energy storage scheduling is optimized, and the traditional algorithms are low efficiency and insufficient prediction accuracy are solved, and efficient and intelligent scheduling is achieved, reducing operational costs and improving energy utilization efficiency.
Patent Information
- Application Number
- CN202510787322.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional algorithms have low efficiency in microgrid energy storage scheduling, insufficient deep learning prediction accuracy, and imperfect optimization of reinforcement learning strategies, making it difficult to cope with uncertainty and dynamic changes in photovoltaic power generation and power load.
The improved GRU neural network is adopted to combine Markov decision-making and Q-learning algorithms to improve prediction accuracy through deep learning, model energy storage scheduling into Markov decision-making process, and optimize the Q-learning algorithm with specific action exploration strategies to achieve intelligent scheduling.
It improves the prediction accuracy of photovoltaic power and power load demand, optimizes the scheduling strategy of the energy storage system, reduces operating costs, and improves the operating stability and energy utilization efficiency of the microgrid.
Smart Images

Figure CN120297773A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microgrid operation strategy scheduling, and more specifically, to a method and device for determining a microgrid energy storage operation scheduling strategy. Background Art
[0002] In the field of integrated energy microgrid (IEM), the intelligent scheduling optimization of energy storage systems is of great significance for achieving efficient energy utilization and reducing operating costs. In a microgrid, the photovoltaic power generation and power load demand are affected by various external factors such as temperature, humidity, light intensity, and user electricity consumption behavior, showing strong volatility and uncertainty, which greatly increases the difficulty of maintaining the energy supply-demand balance in the microgrid. For example, photovoltaic power generation is easily restricted by weather conditions and has periodic and linear correlation characteristics; the electricity load demand is closely related to user behavior and shows strong transient and nonlinear relationships.
[0003] Facing these problems, traditional algorithms expose many problems when dealing with microgrid energy storage scheduling. Rule algorithms such as genetic algorithms and stochastic programming are complex in modeling, low in computational efficiency, and prone to falling into local optimal solutions, making it difficult to adapt to the dynamic change scenarios of the uncertainty of the microgrid power source and load. Although exact algorithms such as mixed integer linear programming (MILP) can obtain theoretical optimal solutions, the calculation time is too long to meet the urgent needs of real-time energy storage scheduling in practical engineering.
[0004] In the aspect of prediction technology, deep learning models such as long short-term memory network (LSTM) and gated recurrent unit (GRU) have been applied to photovoltaic power and load demand prediction. However, the existing models lack optimization in activation functions and network structures, and fail to fully exploit data features, resulting in limited prediction accuracy.
[0005] Reinforcement learning algorithms such as Q-learning algorithm have also been gradually applied in the fields of energy storage arbitrage and control. However, existing research rarely considers the impact of greedy and non-greedy action strategies (such as Softmax) on training efficiency and the scheduling performance differences at different time scales. At the same time, problems such as model dependence and state space discretization also hinder its effective application in microgrid energy storage scheduling.
[0006] In the field of microgrid energy storage system scheduling optimization, traditional rule algorithms have problems of complex modeling, low computational efficiency and being prone to falling into local optima, making it difficult to cope with the dynamic changes of microgrid energy supply and demand; although exact algorithms such as mixed integer linear programming (MILP) can obtain theoretical optimal solutions, they have the problem of too long calculation time. Existing deep learning models are difficult to fully exploit the complex spatio-temporal features and nonlinear relationships in microgrid energy data due to the limitations of activation functions and network structures; while reinforcement learning algorithms face obstacles such as model dependence and state space discretization in applications and cannot achieve efficient scheduling of energy storage systems. Summary of the Invention
[0007] In view of the deficiencies of the prior art, the present invention provides a method and device for determining a microgrid energy storage operation scheduling strategy.
[0008] According to one aspect of the present invention, a method for determining a microgrid energy storage operation scheduling strategy is provided, including: Predicting multi-source data of the microgrid in the current time period based on an improved GRU neural network to obtain predicted multi-source data within a preset future time period; Optimizing the operation scheduling strategies of various operating states of the microgrid by using a decision-making module combining Markov decision-making and Q-learning algorithm to determine the optimal operation scheduling strategies under various operating states; Determining the operating state corresponding to the predicted multi-source data, and taking the optimal operation scheduling strategy in this operating state as the predicted optimal operation scheduling strategy within the future preset time period; Realizing the operation of the microgrid within the preset future time period based on the predicted optimal operation scheduling strategy.
[0009] Optionally, predicting multi-source data of the microgrid in the current time period based on an improved GRU neural network to obtain predicted multi-source data within a future time period, including: Preprocessing the acquired multi-source data in the current time period to obtain preprocessed multi-source data in the current time period; Predicting the preprocessed multi-source data of the microgrid in the current time period based on the improved GRU neural network to obtain predicted multi-source data within a future time period.
[0010] Optionally, the expression of the improved GRU neural network is: Update gate: ; Reset gate: ; Candidate hidden state: ; Hidden state: ; Wherein, x i is the input data at the current moment, h t-1 is the hidden state at the previous moment, α is the Sigmoid activation function; W xu is the weight matrix from the input vector x i to the update gate, W hu is the weight matrix from the hidden state h t-1 at the previous moment to the update gate, W xr is the weight matrix from the input vector x i to the reset gate, W hr is the weight matrix from the hidden state h at the previous momentt-1 to the weight matrix of the reset gate, together with W xr jointly determine the output of the reset gate, W xh is the weight matrix of the input vector x i to the candidate hidden state, W hh is the weight matrix of the previous hidden state h t-1 to the candidate hidden state; b u 、b ur 、b h are used to adjust the reference values of the update gate, reset gate, and candidate hidden state output to avoid gate control failure when the input is all zero; ⊙ represents element-wise multiplication; f() is the Swish activation function.
[0011] Optionally, a decision-making module combining Markov decision and Q-learning algorithm is used to optimize the operation scheduling strategy of the multiple operating states of the microgrid, and determine the optimal operation scheduling strategy under multiple operating states, including: Initialize the initial state of the microgrid and create a Q-table, where the Q-table is used to store the Q-values of the state-action pairs of the microgrid; Based on the initial state of the microgrid, during the training period, an agent continuously tries the actions of multiple operation scheduling strategies to interact with the microgrid to obtain the optimal operation scheduling strategy under multiple operating states.
[0012] Optionally, the action of the agent is selected according to the current state using or the Softmax strategy.
[0013] Optionally, based on the initial state of the microgrid, during the training period, an agent continuously tries the actions of multiple operation scheduling strategies to interact with the microgrid to obtain the optimal operation scheduling strategy under multiple operating states, including: According to the interaction between the action selected by the agent and the microgrid, obtain the new state S t+1 of the microgrid, and obtain the reward r; Update the Q-value of the previous state of the microgrid according to the preset Q-value update formula; Judge whether the current state is a termination state according to the specific operation rules of the microgrid. If it is not a termination state, the agent continues to select actions to interact with the microgrid; If the current state is a termination state, judge whether the current Q-value converges to the optimal result; If the Q-value converges to the final result, obtain the optimal operation scheduling strategy under multiple operating states of the microgrid, otherwise re-initialize the initial state and Q-table of the microgrid and re-optimize.
[0014] Optionally, the Q-value update formula is: Wherein, is the Q value of the next state; is the Q value of the current state; is the optimal Q value among all actions of the next state, and here the optimal Q value is the maximum Q value; is the learning rate; is the discount factor; r k is the reward of the current state, is the current state and action, is the next state and action.
[0015] According to another aspect of the present invention, there is provided a device for determining a microgrid energy storage operation scheduling strategy, including: A prediction module, configured to predict the multi-source data of the microgrid in the current time period according to the improved GRU neural network, and obtain the predicted multi-source data within a preset future time period; An optimization module, configured to optimize the operation scheduling strategies of various operating states of the microgrid by using a decision module combining Markov decision and Q-learning algorithm, and determine the optimal operation scheduling strategies under various operating states; A determination module, configured to determine the operating state corresponding to the predicted multi-source data, and use the optimal operation scheduling strategy in this operating state as the predicted optimal operation scheduling strategy within a preset future time period; An operation module, which realizes the operation of the microgrid within a preset future time period based on the predicted optimal operation scheduling strategy.
[0016] According to still another aspect of the present invention, there is provided a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is used to execute the method described in any one of the above aspects of the present invention.
[0017] According to still another aspect of the present invention, there is provided an electronic device, and the electronic device includes: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of the above aspects of the present invention.
[0018] Therefore, the present invention proposes the DGRU-QL algorithm, an intelligent scheduling optimization method for microgrid energy storage systems that integrates deep learning and reinforcement learning, aiming to solve problems such as low efficiency of traditional algorithms, insufficient prediction accuracy of deep learning, and imperfect optimization of reinforcement learning strategies in microgrid energy storage scheduling. By constructing an improved deep learning network to improve the prediction accuracy of photovoltaic power and power load demand, modeling the microgrid energy storage scheduling problem as a Markov decision process, combining a specific action exploration strategy to optimize the Q-learning algorithm, and performing scheduling at different time scales, the efficient and intelligent scheduling of the energy storage system is achieved, reducing operating costs and improving energy utilization efficiency and the operating stability of the microgrid. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The exemplary embodiments of the present invention can be more fully understood by referring to the following drawings: Figure 1 is a flowchart showing the method for determining the microgrid energy storage operation scheduling strategy provided by an exemplary embodiment of the present invention; Figure 2 is a flowchart showing the DGRU-QL algorithm for solving the IEM energy storage scheduling strategy provided by an exemplary embodiment of the present invention; Figure 3 is a schematic diagram of the GRU recurrent neural network unit provided by an exemplary embodiment of the present invention; Figure 4 is a schematic diagram of the Markov decision process provided by an exemplary embodiment of the present invention; Figure 5 is a schematic diagram of the neural network structure for prediction provided by an exemplary embodiment of the present invention; Figure 6 is a schematic diagram of the optimization of the battery scheduling strategy based on reinforcement learning provided by an exemplary embodiment of the present invention; Figure 7 is a schematic diagram of the structure of the device for determining the microgrid energy storage operation scheduling strategy provided by an exemplary embodiment of the present invention; Figure 8 is the structure of an electronic device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.
[0021] It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.
[0022] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present invention are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0023] It should also be understood that in the embodiments of the present invention, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0024] It should also be understood that for any component, data or structure mentioned in the embodiments of the present invention, without clear limitation or contrary indication in the context, it is generally understood as one or more.
[0025] In addition, the term "and / or" in the present invention is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.
[0026] It should also be understood that the present invention emphasizes the differences between various embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be described one by one.
[0027] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0028] The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present invention and its application or use.
[0029] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and devices should be regarded as part of the specification.
[0030] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0031] Embodiments of the present invention can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0032] Terminal devices, computer systems, servers, and other electronic devices can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, target programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0033] Exemplary Method Figure 1 is a schematic flowchart of a method for determining a microgrid energy storage operation and scheduling strategy provided by an exemplary embodiment of the present invention. This embodiment can be applied to an electronic device, such as Figure 1 as shown, the microgrid energy storage operation and scheduling strategy determination method 100 includes the following steps: Step 101, predict the multi-source data of the microgrid in the current time period according to the improved GRU neural network to obtain the predicted multi-source data within a preset future time period; Step 102, optimize the operation and scheduling strategies of various operating states of the microgrid using a decision-making module of reinforcement learning to determine the optimal operation and scheduling strategies under various operating states; Among them, the optimal operation and scheduling strategies under various operating states can be stored in the "operating state - optimal operation and scheduling strategy" mapping library; Step 103, determine the operating state corresponding to the predicted multi-source data, and use the optimal operation and scheduling strategy in this operating state as the predicted optimal operation and scheduling strategy within a preset future time period; Among them, the system queries the state in the mapping library with the highest matching degree with the operating state corresponding to the predicted multi-source data, and determines the optimal strategy corresponding to the state with the highest matching degree as the predicted optimal operation and scheduling strategy within a preset future time period; Step 104: Based on the predicted optimal operation scheduling strategy, the microgrid is operated within a preset future time period.
[0034] Specifically, the present invention proposes an intelligent scheduling optimization method for a microgrid energy storage system that combines deep learning and reinforcement learning. An improved deep learning network is constructed to improve the prediction accuracy of photovoltaic power and power load demand. By using the Swish activation function to optimize the GRU network structure, the photovoltaic power and power load demand are predicted. The microgrid energy storage scheduling problem is modeled as a Markov decision process, and a Q-learning algorithm combined with the Softmax action exploration strategy is designed, and optimal scheduling is performed at different time scales. A reward function is designed by comprehensively considering factors such as the cost of purchasing electricity and the cost of battery charging and discharging. It solves problems such as low computational efficiency of traditional algorithms, insufficient prediction accuracy of deep learning, and imperfect optimization of reinforcement learning strategies, realizes the efficient and intelligent scheduling of the microgrid energy storage system, reduces the operation cost, and improves the energy utilization efficiency and the operation stability of the microgrid.
[0035] The DGRU-QL (Dropout-Gated Recurrent Unit-Q-Learning) algorithm proposed by the present invention is used to solve the IEM energy storage scheduling strategy. Its process integrates two major parts: deep learning and reinforcement learning, which is a key step in realizing the intelligent scheduling of the microgrid energy storage system. The specific process is as Figure 2 shown. The specific implementation steps are as follows: 1. Deep learning prediction process Start: As the starting point of the entire process, it marks the start of the microgrid energy storage scheduling strategy solving process.
[0036] Data preprocessing: In the prediction module based on deep learning, the microgrid generates multi-source data including (a current period of time, for example, if it is the 15th today, select the data from the 10th to the 15th, and the prediction accuracy can be hourly, daily, 15 minutes, etc., which is not limited here) photovoltaic power, power load demand, etc. This step cleans the original data, removes the outliers, fills in the missing values in the data, and performs normalization processing. Through these operations, the data meets the requirements of the improved GRU neural network model training, provides high-quality input data for accurate prediction, and is the basic link of the entire deep learning prediction.
[0037] Training and Testing: Using the preprocessed data, train the improved GRU neural network. During the training process, continuously adjust the network parameters with the help of the backpropagation algorithm so that the model can learn the complex spatio-temporal features and non-linear relationships in the data, that is, master the internal laws of photovoltaic power and power load demand data. After training, use the test data to evaluate the performance of the model to test the accuracy of the model in predicting photovoltaic power and power load demand for future time periods (seven days, 15 days, or other time periods, not limited here).
[0038] Hyperparameter Optimization: Adjust and optimize the hyperparameters in the GRU neural network, such as the learning rate, the number of neurons in the hidden layer, etc. These hyperparameters have an important impact on the performance of the model. By reasonably adjusting them, the prediction ability of the model can be improved, making it better adapt to the characteristics of microgrid data.
[0039] Obtaining the Minimum Prediction Error: Check by monitoring the root mean square error and mean absolute error of the model on the training set and test set in real time to determine whether the prediction error reaches the minimum. If the error value continues to decrease during the iteration process, it means that there is still room for optimization of the model and it is necessary to return to continue hyperparameter optimization; if the loss no longer decreases significantly after 100 steps of iteration, it is considered that the minimum prediction error has been reached and the model already has good prediction performance under the current conditions and can enter the next step.
[0040] Result Correction: Correct the prediction results of the model. Although the model that has been trained and tested already has a certain prediction ability, in order to make it more in line with the actual operation of the microgrid, it is necessary to correct the prediction results (the process of adjusting parameters due to the difference between the prediction results and the actual results), so as to provide more reliable data for the subsequent decision-making module based on reinforcement learning.
[0041] In an embodiment of the present invention, the deep learning prediction module constructs a model using GRU recurrent neural network units (as Figure 3 shown), aiming to capture the complex spatio-temporal features and non-linear relationships in the data using update gates and reset gates, accurately predict photovoltaic power and power load demand, provide a reliable data basis for microgrid energy storage scheduling decisions, and thus improve the prediction accuracy.
[0042] The specific formulas are as follows: Update Gate: ; Reset Gate: ; Candidate Hidden State: ; Hidden State: ; In the formula, x i is the input data at the current moment, ht-1 is the hidden state at the previous moment, α is the Sigmoid activation function, which is used to control the information flow ratio; W is the weight matrix, where W xu is the weight matrix from the input vector x i to the update gate, which is responsible for mapping the input information to the update gate decision space, W hu is the weight matrix from the input hidden state h at the previous moment t-1 to the update gate, which is used to capture the influence of the historical state on the update gate, W xr is the weight matrix from the input vector x i to the reset gate, which controls the degree of forgetting of the current input to the historical state, W hr is the weight matrix from the input hidden state h at the previous moment t-1 to the reset gate, and together with W xr determines the output of the reset gate, W xh is the weight matrix from the input vector x i to the candidate hidden state, which is responsible for mapping the input information to the candidate hidden state space, W hh is the hidden state h at the previous moment t-1 to the weight matrix of the candidate hidden state; b is the bias vector, where b u and b ur and b h are used to adjust the reference values of the update gate, reset gate, and candidate hidden state outputs, to avoid the gating failure when the input is all zero; ⊙ represents element-wise multiplication. The Swish activation function is adopted to optimize the GRU network and enhance the complex data feature extraction ability.
[0043] 2. Reinforcement learning decision-making process Initialize the random state and Q-table: In the decision-making module based on reinforcement learning, the microgrid energy storage scheduling problem is modeled as a Markov decision process. This step is to randomly set the initial state of the microgrid and create a Q-table at the same time. The Q-table is used to store the Q-values of each state-action pair, and the Q-value represents the long-term cumulative reward expected to be obtained by executing a certain action in this state. It is the core data structure for the agent (battery) to learn the optimal strategy.
[0044] Start the training cycle: Start a training cycle, during which the agent (battery) will interact with the environment (microgrid). In this interaction process, the agent continuously tries different actions to learn the optimal energy storage scheduling strategy.
[0045] Use or the Softmax strategy to select an action in state S: The agent selects an action according to the current state S it is in, using the strategy (randomly selects an action with a probability of and with a probability of 1 - Select the action with the largest current Q value with a certain probability. This strategy balances the exploitation of known optimal actions and the exploration of new actions) or the Softmax strategy (calculate the selection probability based on the Q value of the action and adjust the proportion of exploration and exploitation through the temperature adjustment coefficient τ) to select an action.
[0046] Execute this action and transfer to the new state S t+1 And obtain the reward r: After the agent executes the selected action, the state of the microgrid will change and transfer to the new state S t+1 . At the same time, obtain the corresponding reward r according to the feedback of the environment. The setting of the reward is closely related to factors such as the operating cost of the microgrid and the matching degree of energy supply and demand. For example, charging the battery during the low electricity price period and discharging it during the high electricity price period to save costs for the microgrid will obtain a positive reward, otherwise it may get a negative reward.
[0047] Update the Q value of the previous state according to the formula: Use the formula to update the Q value of the previous state. Among them, is the learning rate, which controls the step size when updating the Q value each time; γ is the discount factor, which is used to measure the importance of future rewards. By continuously updating the Q value, the agent gradually learns the optimal actions to be taken in different states.
[0048] Whether the state is a terminal state: Judge whether the current state is a terminal state. The setting of the terminal state is based on the specific operation rules of the microgrid, such as reaching the preset time limit, the state of charge of the battery reaching the limit, etc. If the current state is not a terminal state, return to continue selecting actions and continuously interact with the environment; if it is a terminal state, enter the next judgment.
[0049] Whether the Q value converges to the optimal result: Judge whether the Q value has converged to the optimal result. When the Q value converges, it means that the agent has learned the optimal action strategy in various states; if it has not converged, it is necessary to return to re-initialize the random state and the Q table, start a new training segment, and continue to optimize the strategy.
[0050] Output the optimal scheduling strategy: When the Q value converges, at this time the agent has found the optimal actions in different states, and will output the optimal charge and discharge scheduling strategy of the battery in different states, and the process of the entire DGRU-QL algorithm for solving the IEM energy storage scheduling strategy ends.
[0051] In an embodiment of the present invention, the decision-making module based on reinforcement learning models the microgrid energy storage scheduling problem as a Markov decision process (as Figure 4 shown, a 1. a 2. a3 (which are three discrete actions)), the Q-learning algorithm is used to combine with a specific action exploration strategy to achieve the optimal scheduling of the energy storage system, as follows: 1) Markov decision process (MDP) modeling: State space (S): It is defined as , including the state of charge ( : It needs to satisfy , where and are the lower and upper limits of the state of charge of the battery respectively, and the lower limit is set to . When approaching the upper limit, it tends to discharge to avoid capacity waste. When approaching the lower limit, it needs to charge to ensure subsequent power supply), grid power ( : When , is low (off-peak electricity price period) and is low ( ), the system tends to purchase electricity from the grid or charge with photovoltaic power to ensure the storage. Conversely, the microgrid sells electricity to the grid), photovoltaic output power ( : When , charging can be carried out), power load demand ( : When , electricity needs to be purchased), heat load demand ( ), electricity price ( ), natural gas price ( ), comprehensively reflecting the operating state of the microgrid.
[0052] Action space (A): It is defined as , which consists of battery charge and discharge actions. In actual calculation, based on the rated power of the battery (such as 2 kW), it is extended to 2 times the rated power. The negative sign indicates discharge, the positive sign indicates charge, and 0 indicates idle. Discrete values (such as -4, -2, 0, 2, 4, unit: kW) are obtained through equally spaced division.
[0053] Reward function (R): It is set as , where: C is the natural gas purchase cost and the battery charge and discharge cost, , is the natural gas price at time t, is the natural gas purchase volume at time t, is the unit charge and discharge cost of the battery, represents the total number of time steps in the scheduling period. For example, if 1 day is the scheduling period and it is divided into 96 time steps according to 15 minutes, then T = 96, are the battery discharge volume and charge volume at time t respectively; D is the penalty for energy supply and demand mismatch. When , , otherwise , is the generated electricity of the CHP (Combined Heat and Power) unit at time t; E is the penalty for the battery SOC exceeding the limit. When or , E = 100; otherwise, E = 0. are the upper and lower limits of the battery capacity.
[0054] Policy (π): Adopt and the Softmax action exploration policy. The policy randomly selects an action with probability , and selects the action with the maximum current Q value with probability ( ); The Softmax policy calculates the action selection probability through . τ is the temperature adjustment coefficient, which adjusts the exploration and exploitation ratio. a i represents the i-th action in the action space; Q(a i ) represents the Q value of executing action a i in the current state, that is, the reward value. The higher this value, the greater the probability of selecting this action; k is the total number of actions in the action space. For example, if k = 5, there are 5 actions.
[0055] 2) Neural network structure for prediction As Figure 4 shown, the neural network structure for prediction is designed to efficiently process time series data related to photovoltaic power and power load demand. The meanings of each element in the figure are explained in detail below: Input layer: The input layer receives preprocessed multi-source data, which contains various features related to photovoltaic power and power load demand, such as historical power data, ambient temperature, light intensity, time information, etc. The number of neurons in the input layer is determined by the feature dimension of the input data. Its function is to convert the original data into a numerical form suitable for network processing and provide input for the subsequent GRU layer.
[0056] GRU layer: The GRU layer is the core part of the entire network. In this layer, the GRU unit can effectively capture the long-term dependencies in time series data through the gating mechanism, avoiding the problem of gradient vanishing or explosion in traditional recurrent neural networks. Each unit in the GRU layer shown in the figure receives the input data at the current moment from the input layer and the hidden state at the previous moment, and outputs the hidden state at the current moment.
[0057] Random inactivation layer: The random inactivation layer is mainly used to prevent network overfitting. During the training process, the random inactivation layer will randomly "discard" (i.e., not participate in the calculation) some neurons with a certain probability. This can prevent the network from relying too much on certain specific neurons, thereby enhancing the generalization ability of the network. The data processed by the random inactivation layer will continue to be passed to the subsequent layers for further processing.
[0058] Output layer: The output layer outputs the result according to the prediction task of the network. For the photovoltaic power and power load demand prediction tasks, the number of neurons in the output layer is usually 1, that is, a predicted value is output. The output layer will perform a linear transformation on the hidden state passed from the GRU layer or obtain the final prediction result after being processed by an activation function.
[0059] 3) Q-learning algorithm: Through the formula: Update the Q value, where is the learning rate (0 < < 1), which controls the update step size; γ is the discount factor (0 < γ < 1), which is used to measure the importance of future rewards; r is the reward obtained by performing the action; is the previous state and action, is the current state and action. The agent continuously updates the Q value through interaction until the optimal scheduling strategy is converged. As Figure 6 shown in the schematic diagram of the battery scheduling strategy optimization based on reinforcement learning, it shows the interaction relationship between the agent (battery), the environment (microgrid), state, action and reward, and intuitively presents the optimization process of the reinforcement learning algorithm for the battery scheduling strategy.
[0060] Therefore, the present invention proposes a DGRU-QL algorithm, an intelligent scheduling optimization method for a microgrid energy storage system that combines deep learning and reinforcement learning, aiming to solve problems such as low efficiency of traditional algorithms in microgrid energy storage scheduling, insufficient prediction accuracy of deep learning, and imperfect strategy optimization of reinforcement learning. By constructing an improved deep learning network to improve the prediction accuracy of photovoltaic power and power load demand, modeling the microgrid energy storage scheduling problem as a Markov decision process, combining a specific action exploration strategy to optimize the Q-learning algorithm, and performing scheduling at different time scales, efficient and intelligent scheduling of the energy storage system is realized, the operation cost is reduced, and the energy utilization efficiency and the operation stability of the microgrid are improved.
[0061] Exemplary device Figure 7 is a schematic structural diagram of a device for determining a microgrid energy storage operation scheduling strategy provided by an exemplary embodiment of the present invention. As Figure 7 shown, the device 700 includes: A prediction module 710, configured to predict the multi-source data of the microgrid in the current time period according to an improved GRU neural network, and obtain the predicted multi-source data within a preset future time period; An optimization module 720, configured to optimize the operation scheduling strategies of multiple operating states of the microgrid by using a decision-making module combining Markov decision-making and Q-learning algorithms, and determine the optimal operation scheduling strategies under multiple operating states; A determination module 730, configured to determine the operating state corresponding to the predicted multi-source data, and use the optimal operation scheduling strategy in this operating state as the predicted optimal operation scheduling strategy within a preset future time period; An operation module 740, which realizes the operation of the microgrid within a preset future time period based on the predicted optimal operation scheduling strategy.
[0062] Optionally, the prediction module 710 includes: A preprocessing sub-module, configured to preprocess the acquired multi-source data in the current time period to obtain the preprocessed multi-source data in the current time period; A prediction sub-module, configured to predict the preprocessed multi-source data of the microgrid in the current time period according to an improved GRU neural network, and obtain the predicted multi-source data within a future time period.
[0063] Optionally, the expression of the improved GRU neural network is: Update gate: ; Reset gate: ; Candidate hidden state: ; Hidden state: ; In the formula, x i is the input data at the current moment, h t-1 is the hidden state at the previous moment, and α is the Sigmoid activation function; W xu is the weight matrix from the input vector x i to the update gate, W hu is the weight matrix from the hidden state h t-1 at the previous moment to the update gate, W xr is the weight matrix from the input vector x i to the reset gate, W hr is the weight matrix from the hidden state h t-1 at the previous moment to the reset gate, and together with W xr determines the output of the reset gate, W xh is the weight matrix from the input vector x i to the candidate hidden state, W hh is the hidden state h t-1The weight matrix to the candidate hidden state; b u and b ur and b h are used to adjust the reference values for updating the gate, reset gate, and candidate hidden state output to avoid gate control failure when the input is all zeros; ⊙ represents element-wise multiplication; f() is the Swish activation function.
[0064] Optionally, the optimization module 720 includes: An initialization sub-module for initializing the initial state of the microgrid and creating a Q-table, where the Q-table is used to store the Q-values of the state-action pairs of the microgrid; An interaction sub-module for, based on the initial state of the microgrid, during the training period, using an agent to continuously try the actions of multiple operation scheduling strategies to interact with the microgrid to obtain the optimal operation scheduling strategy under various operating states.
[0065] Optionally, the action of the agent is selected according to the current state using or the Softmax strategy.
[0066] Optionally, the interaction sub-module includes: A first obtaining unit for obtaining the new state S of the microgrid according to the interaction between the action selected by the agent and the microgrid t+1 and obtaining a reward r; An updating unit for updating the Q-value of the previous state of the microgrid according to a preset Q-value updating formula; A first judging unit for judging whether the current state is a termination state according to the specific operation rules of the microgrid. If it is not a termination state, the agent continues to select actions to interact with the microgrid; A second judging unit for judging whether the current Q-value converges to the optimal result if the current state is a termination state; A second obtaining unit for obtaining the optimal operation scheduling strategy of the microgrid under various operating states if the Q-value converges to the final result, otherwise re-initializing the initial state of the microgrid and the Q-table and re-optimizing.
[0067] Optionally, the Q-value updating formula is: where is the Q-value of the next state; is the Q-value of the current state; is the optimal Q-value among all actions in the next state, and here the optimal Q-value is the maximum Q-value; is the learning rate; is the discount factor; r k is the reward of the current state, is the current state and action, For the next state and action.
[0068] Exemplary electronic device Figure 8 is the structure of an electronic device provided by an exemplary embodiment of the present invention. As Figure 8 shown, the electronic device 80 includes one or more processors 81 and a memory 82.
[0069] The processor 81 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0070] The memory 82 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 81 may run the program instructions to implement the methods of the software programs of the various embodiments of the present invention described above and / or other desired functions. In one example, the electronic device may further include: an input device 83 and an output device 84, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0071] In addition, the input device 83 may further include, for example, a keyboard, a mouse, etc.
[0072] The output device 84 may output various information to the outside. The output device 84 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0073] Of course, for simplicity, Figure 8 only some of the components related to the present invention in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.
[0074] Exemplary computer program products and computer-readable storage media In addition to the above methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present invention described in the "Exemplary Method" section above of this specification.
[0075] The computer program product may be written in any combination of one or more programming languages for executing the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0076] In addition, an embodiment of the present invention may also be a computer-readable storage medium storing computer program instructions which, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present invention described in the "Exemplary Methods" section above of this specification.
[0077] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0078] The basic principles of the present invention have been described above in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present invention. In addition, the above-disclosed specific details are only for illustrative and facilitating understanding purposes, and not for limitation. The above details do not limit the present invention to necessarily adopt the above specific details for implementation.
[0079] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference may be made to each other. For the system embodiments, since they basically correspond to the method embodiments, the description is relatively simple, and reference may be made to the partial description of the method embodiments for the relevant parts.
[0080] The block diagrams of the devices, systems, equipment, and systems involved in the present invention are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, systems, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the term "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0081] The methods and systems of the present invention can be implemented in many ways. For example, the methods and systems of the present invention can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is for illustrative purposes only, and the steps of the methods of the present invention are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present invention can also be implemented as a program recorded on a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present invention. Therefore, the present invention also covers a recording medium storing a program for executing the methods according to the present invention.
[0082] It should also be noted that in the systems, equipment, and methods of the present invention, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present invention. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present invention. Therefore, the present invention is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0083] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present invention to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A method for determining a microgrid energy storage operation and dispatching strategy, characterized in that, Including: Predict the multi-source data of the microgrid in the current time period according to the improved GRU neural network to obtain the predicted multi-source data within a preset future time period; Optimize the operation and scheduling strategies of various operating states of the microgrid by using a decision-making module that combines Markov decision-making and Q-learning algorithms to determine the optimal operation and scheduling strategies under various operating states; Determine the operating state corresponding to the predicted multi-source data, and use the optimal operation and scheduling strategy in this operating state as the predicted optimal operation and scheduling strategy within a preset future time period; Realize the operation of the microgrid within a preset future time period based on the predicted optimal operation and scheduling strategy.
2. The method according to claim 1, characterized in that, Predict the multi-source data of the microgrid in the current time period according to the improved GRU neural network to obtain the predicted multi-source data within a future time period, including: Preprocess the acquired multi-source data in the current time period to obtain the preprocessed multi-source data in the current time period; Predict the preprocessed multi-source data of the microgrid in the current time period according to the improved GRU neural network to obtain the predicted multi-source data within a future time period.
3. The method according to claim 1, wherein The expression of the improved GRU neural network is: Update door: ; Reset door: ; Candidate hidden state: ; Hidden state: ; where x i is the input data at the current moment, h t-1 is the hidden state at the previous moment, and α is the Sigmoid activation function; W xu is the weight matrix from the input vector x i to the update gate, W hu is the weight matrix from the hidden state h t-1 at the previous moment to the update gate, W xr is the weight matrix from the input vector x i to the reset gate, W hr is the weight matrix from the hidden state h t-1 at the previous moment to the reset gate, and together with W xr determines the output of the reset gate, W xh is the weight matrix from the input vector x i to the candidate hidden state, W hh is the weight matrix from the hidden state h t-1 at the previous moment to the candidate hidden state; b u , b ur , b h are used to adjust the baseline values of the update gate, reset gate, and candidate hidden state output to avoid gate control failure when the input is all zero; ⊙ represents element-wise multiplication; f() is the Swish activation function.
4. The method according to claim 1, wherein Optimize the operation and scheduling strategies of various operating states of the microgrid by using a decision-making module that combines Markov decision-making and Q-learning algorithms to determine the optimal operation and scheduling strategies under various operating states, including: Initialize the initial state of the microgrid and create a Q-table, where the Q-table is used to store the Q-values of the state-action pairs of the microgrid; Based on the initial state of the microgrid, within the training cycle, use an agent to continuously try the actions of multiple operation and scheduling strategies to interact with the microgrid to obtain the optimal operation and scheduling strategies under various operating states.
5. The method according to claim 4, wherein The action of the agent is selected according to the current state using or the Softmax policy.
6. The method according to claim 4, characterized in that Based on the initial state of the microgrid, within the training cycle, use an agent to continuously try the actions of multiple operation and scheduling strategies to interact with the microgrid to obtain the optimal operation and scheduling strategies under various operating states, including: Based on the interaction between the action selected by the agent and the microgrid, the new state S of the microgrid is obtained t+1 , and a reward r is obtained; Update the Q-value of the previous state of the microgrid according to the preset Q-value update formula; Judge whether the current state is a termination state according to the specific operation rules of the microgrid. If it is not a termination state, the agent continues to select actions to interact with the microgrid; If the current state is a termination state, judge whether the current Q-value converges to the optimal result; If the Q-value converges to the final result, obtain the optimal operation and scheduling strategies under various operating states of the microgrid, otherwise re-initialize the initial state and Q-table of the microgrid and re-optimize.
7. The method according to claim 6, wherein The Q-value update formula is: Among them, is the Q-value of the next state; is the Q-value of the current state; is the optimal Q-value among all actions of the next state, where the optimal Q-value here is the maximum Q-value; is the learning rate; γ is the discount factor, is the reward of the current state, is the current state and action, is the next state and action.
8. An optimization device for the operation and dispatching strategy of a microgrid energy storage, characterized in that, Including: A prediction module for predicting the multi-source data of the microgrid in the current time period according to the improved GRU neural network to obtain the predicted multi-source data within a preset future time period; An optimization module for optimizing the operation and scheduling strategies of various operating states of the microgrid by using a decision-making module that combines Markov decision-making and Q-learning algorithms to determine the optimal operation and scheduling strategies under various operating states; A determination module, configured to determine the operating state corresponding to the predicted multi-source data, and use the optimal operation scheduling strategy in this operating state as the predicted optimal operation scheduling strategy within a preset future time period; An operation module, configured to implement the operation of the microgrid within a preset future time period based on the predicted optimal operation scheduling strategy.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1-7 above.
10. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1-7 above.
Citation Information
Patent Citations
Distributed energy storage system optimization scheduling method based on LSTM and DQN algorithms
CN119382213A
Self-adaptive plug-in architecture optimization method and device based on Q-Learning
CN119621126A
Comprehensive energy system dynamic scheduling method based on multi-agent deep reinforcement learning
CN119904057A
Cited By
Multi-source energy storage self-power-supply method and system for iron tower, equipment and storage medium
CN120497922A
Multi-source energy storage self-powering methods, systems, equipment, and storage media for iron towers
CN120497922B