Optical storage complementary intelligent storage control algorithm applying MPC-DRLC
By applying MPC-DRLC algorithm in the field of complementary intelligent storage control of optical storage, combined with model prediction control and deep reinforcement learning control, the limitations of existing algorithms in real-time and optimal strategy formulation are solved, and more efficient energy utilization and lower electricity consumption costs are achieved.
Patent Information
- Application Number
- CN202510345714.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
The existing optical storage complementary intelligent storage control algorithm has limitations in real-time and optimal strategy formulation, and it is difficult to effectively improve energy utilization and reduce electricity consumption costs.
The MPC-DRLC (Model Prediction Control-Deep Reinforcement Learning Control) algorithm is adopted to construct a mathematical model of the complementary optimization problem of optical storage, combined with the SSA-LSTM prediction model and deep reinforcement learning control, and realize short-term and long-term global intelligent control.
It improves energy utilization, protects the life of the energy storage system, and achieves better working performance and lower electricity consumption costs through rolling optimization mechanism and reward function design.
Smart Images

Figure CN120200232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning, reinforcement learning, and intelligent control and scheduling of building energy, and particularly to a photovoltaic-storage complementary intelligent storage control algorithm applying MPC-DRLC. Background Art
[0002] With the intensification of the energy crisis and environmental problems, the application of renewable energy has received global attention. As a green and clean renewable energy, solar energy has been widely used in the energy transformation of urban buildings. Constructing a photovoltaic-storage multi-energy complementary system is an effective means to achieve green energy use in the energy-saving transformation of office buildings. Studying a photovoltaic-storage complementary intelligent storage control algorithm based on this system can improve energy utilization efficiency and reduce electricity costs. Real-time and online are the cores of intelligent management and control. Based on the set objective function, the power distribution can be controlled and scheduled through real-time data processing. Existing algorithms have limitations in terms of real-time performance and optimal strategy formulation.
[0003] The present invention proposes a photovoltaic-storage complementary intelligent storage control algorithm applying MPC-DRLC. This method constructs an MPC-DRLC algorithm jointly to control the designed optimization problem for the energy consumption and energy supply scenarios of a building photovoltaic system - energy storage system, promoting green energy use, energy conservation, and consumption reduction. In the field of photovoltaic-storage complementary intelligent storage control, existing similar patents such as CN104967149A, CN116914856A, CN119024707A, etc.:
[0004] The above specific patent comparison documents are as follows:
[0005] 1), "A Model Predictive Control Method for a Microgrid with Wind, Photovoltaic, and Energy Storage", patent number CN104967149A. This method discloses a model predictive control method for a microgrid with wind, photovoltaic, and energy storage, including the following steps: establishing a prediction model to predict the maximum output of wind turbines and photovoltaic power generation in the microgrid within a set future time period; using the predicted maximum output of wind power and photovoltaic power as constraint conditions to perform online optimization on the output of wind turbines, photovoltaic power generation, and energy storage batteries in the microgrid, and giving the reference output of the three; and performing feedback adjustment on the reference output of wind turbines, photovoltaic power generation, and energy storage batteries according to the real-time adjustable capacity of wind turbines and photovoltaic power generation. The present invention constructs an SSA-LSTM photovoltaic power prediction model that can automatically adjust neural network hyperparameters, with more efficient and accurate prediction, which is different from the predictive control method in the above invention.
[0006] 2) "Optimization Method for Wind-Solar-Storage Integrated System Based on Multi-Objective Grey Wolf Algorithm", Patent No. CN116914856A. This method discloses an optimization method for a wind-solar-storage integrated system based on a multi-objective grey wolf algorithm, including the following steps: taking minimizing economic cost, minimizing grid-connected power difference, and maximizing carbon emission reduction as optimization objectives, improving the multi-objective grey wolf optimization algorithm by introducing Tent chaotic mapping, non-linear convergence factor, and dynamic weight, and using the improved multi-objective grey wolf optimization algorithm to study the scheduling optimization of the wind-solar-storage integrated system. The present invention designs an intelligent control algorithm of model predictive control - deep reinforcement learning control. The performance of the combined algorithm is superior to that of ordinary intelligent optimization algorithms, and the control decision obtained by solving the optimization problem is more practical, which is different from the algorithm principle in the above invention.
[0007] 3) "Optimization Control Method for Residential Integrated Energy System Based on Multi-Agent Reinforcement Learning", Patent No. CN119024707A. This method discloses an optimization control method for a residential integrated energy system based on multi-agent reinforcement learning, including the following steps: obtaining relevant data from the actual operation of the building, and using the collected data as training sample data for data preprocessing; constructing a multi-agent environment model to simulate and constrain the operation of the system; constructing a multi-agent IPPO model, training it through the policy gradient method using the IPPO algorithm, and generating the optimal regulation strategy for the building energy system after multiple iterations; using the trained model to achieve efficient energy allocation and reduction of operating costs. The algorithm design of the present invention is aimed at building scenarios equipped with photovoltaic systems and energy storage systems, where the energy consumption and supply situations are more complex. The reinforcement learning control process combines predictive control and can achieve short-term and long-term global control, which is different from the method proposed in the above invention. Summary of the Invention
[0008] To solve the above technical problems, the object of the present invention is to provide a light-storage complementary intelligent storage control algorithm applying MPC-DRLC.
[0009] The object of the present invention is achieved by the following technical solutions:
[0010] A light-storage complementary intelligent storage control algorithm applying MPC-DRLC, including:
[0011] Step 10: Obtain the operation data set and energy consumption data set of the office building;
[0012] Step 20: Construct a mathematical model of the light-storage complementary optimization problem;
[0013] Step 30: Construct an SSA-LSTM prediction model to predict the photovoltaic power generation;
[0014] Step 40: Design a model predictive control MPC algorithm to achieve short-term intelligent management and control;
[0015] Step 50 designs a deep reinforcement learning control DRLC algorithm, combines it with model predictive control to obtain the MPC-DRLC algorithm, and realizes short-term and long-term global intelligent management and control;
[0016] Step 60 verifies the algorithm performance through evaluation indexes such as electricity consumption cost, charge and discharge action conversion times of the energy storage system, and reward function.
[0017] Compared with the prior art, one or more embodiments of the present invention may have the following advantages:
[0018] All the data used in the present invention are real operation data, and the studied algorithm has strong generalization ability; the photovoltaic power generation power of office buildings is predicted through the SSA-LSTM prediction model, which can effectively improve the prediction accuracy and efficiency, and the accurate data provides a basis for intelligent management and control; according to the prediction results, a model predictive control MPC algorithm is constructed, and the rolling optimization mechanism is adopted to update the control results in real time, with strong anti-interference ability and good short-term control performance; a deep reinforcement learning control DRLC algorithm is constructed, combines the short-term control optimization results of MPC to design a Markov decision process, designs a reward function that can improve energy utilization rate and protect the life of the energy storage system, and uses a new PPO agent to achieve control; application examples show that compared with single algorithms, MPC-DRLC takes into account short-term and long-term global control and has better working performance. Description of the Drawings
[0019] Figure 1 is the flow chart of the intelligent storage control algorithm for complementary photovoltaic and energy storage applying MPC-DRLC;
[0020] Figure 2 is the power control and scheduling model diagram of the office building - photovoltaic system - energy storage system - power grid;
[0021] Figure 3 is the LSTM network structure diagram for photovoltaic power prediction of office buildings;
[0022] Figure 4 is the flow chart of automatic tuning of LSTM hyperparameters by SSA;
[0023] Figure 5 is the design diagram of the MPC control scheme;
[0024] Figure 6 is the design diagram of the DRLC control scheme;
[0025] Figure 7 is the line chart of the cumulative reward function values of three reinforcement learning algorithms. Detailed Embodiments
[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings.
[0027] As Figure 1 shown, it is the flow of the optical storage complementary intelligent storage control algorithm applying MPC-DRLC, including:
[0028] Step 10: Obtain the operation dataset and energy consumption dataset of the office building;
[0029] Step 20: Construct a mathematical model of the optimization problem for the optical storage complementarity;
[0030] Step 30: Construct an SSA-LSTM prediction model to predict the photovoltaic power generation;
[0031] Step 40: Design a model predictive control MPC algorithm to achieve short-term intelligent management and control;
[0032] Step 50: Design a deep reinforcement learning control DRLC algorithm, and combine the model predictive control to obtain the MPC-DRLC algorithm to achieve short-term and long-term global intelligent management and control;
[0033] Step 60: Verify the algorithm performance through evaluation indicators such as electricity consumption cost, number of charge and discharge action conversions of the energy storage system, and reward function.
[0034] The above-mentioned Step 10 specifically includes: obtaining the time series operation datasets such as the atmospheric temperature, photovoltaic panel patch temperature, direct radiation, diffuse radiation, total radiation, and power generation collected by the photovoltaic system of the office building, and dividing them into a training set, a test set, and a validation set according to 8:1:1; obtaining the datasets related to energy consumption and energy storage.
[0035] The above-mentioned Step 20 specifically includes: As Figure 2 shown, at time t, the photovoltaic system, the energy storage system, and the power grid supply energy s(t), b - (t), d(t) to the office building respectively; the photovoltaic system and the power grid store energy b + (t), q(t) in the energy storage system respectively, and set the optimization objective function (1) and constraint conditions (2) - (8):
[0036]
[0037] Among them, the total number of control time points is T, and the time point set The predicted power generation of the photovoltaic system is p(t), the electricity demand of the user is r(t), and the real-time electricity price is k(t); the upper limits of the single charge and discharge amounts of the energy storage system are c + , c - respectively; the stored electricity and the upper limit of the stored electricity of the energy storage system are B(t) and Q respectively; the charging and discharging efficiencies of the energy storage system are η + , η- ; Energy storage system constraint constant b storage and P storage Constrains the energy storage action, and the occurrence probability of event [·] is Equation (2) stipulates that the predicted power generation p(t) of the photovoltaic system only discharges to the office building and the energy storage system; Equation (3) stipulates that the electricity consumption r(t) of the office building is only provided by the photovoltaic system, the energy storage system, and the power grid; Equations (4) and (5) stipulate the charge and discharge limits; Equation (6) stipulates the stored electricity limit; Equation (7) stipulates the change law of the stored electricity; Equation (8) protects the life of the energy storage system.
[0038] The specific steps of step 30 are as follows: Figure 3 As shown, the long short-term memory network LSTM model is used to predict the photovoltaic power, and the working principle is: the input at time t includes the memory cell state c at time (t - 1) t-1 and the hidden layer state h t-1 and the input x at time t t ; The output includes the memory cell state c at time t t and the hidden layer state h t . Let the weights of the forget gate, input gate, and output gate be W f , W i , W o respectively, and the biases of the forget gate, input gate, and output gate be b f , b i , b o respectively. The Sigmoid activation function is σ(x), and h t-1 and x t are concatenated into a vector [h t-1 , x t . Then, the forget gate, input gate, and output gate can calculate and output f t , i t , o t
[0039]
[0040] Let the weights and biases of the candidate cell state be W c , b c respectively. The input gate calculates the candidate memory cell state C x through the function tanh(x) = (e -x - e x ) / (e -x + e t . The forget gate and input gate update the c information by bitwise multiplication t . The output gate calculates h t :
[0041]
[0042] Step 30 above specifically includes: As Figure 4 shown, the sparrow search algorithm (SSA) is applied to automatically tune the hyperparameters of LSTM. The working principle is as follows: Automatic hyperparameter tuning is achieved through initialization and the update of the positions of producers, foragers, and danger-aware sparrows; in the initialization process, the input optimization target prediction window length output dimension of the hidden layer output dimension of the fully connected layer random inactivation rate of the Dropout strategy learning rate and other hyperparameters are input. The sparrow population size and relevant algorithm parameters are initialized, and the individual positions are determined by sorting the fitness values; in the position update process, the positions of producers, foragers, and danger-aware sparrows are updated according to the judgment conditions. When the upper limit of the iteration times is reached, the current global optimal position, global optimal fitness value, and the corresponding combination are obtained.
[0043] Step 40 above specifically includes: As Figure 5 shown, SSA-LSTM is applied to predict the power generation p(t) of an office building. Let the state combination of B(t), the actual power generation g(t) at time t, and g(t - 1) at time (t - 1) be the state vector x(t) = [B(t), g(t - 1)] T , s(t), b + (t), b - (t), q(t) control decision combination be the control vector u(t) = [s(t), b + (t), b - (t), q(t)] T , then the state vector x(t + 1) at time (t + 1) can be represented by the state space equation of the photovoltaic-storage hybrid system:
[0044] x(t + 1) = Ax(t) + Bu(t) (11)
[0045] where the state vector parameter matrix A and the control vector parameter matrix B are respectively:
[0046]
[0047] Through rolling optimization, continuous and rolling control is performed on the short-term state within a finite time domain. The steps include:
[0048] I. Set N (N < T) time points as the optimization window in the time domain, and the window length remains unchanged during the control process;
[0049] II. Apply the SSA-LSTM prediction model + data within the time domain to predict the latest p(t);
[0050] III. Input x(N) and the constraint conditions, parameters, and solve the optimization problem to obtain u(N);
[0051] IV. The window scrolls to the next moment, solve for x(N + 1), and start the next round of rolling optimization.
[0052] Iteratively perform the above steps I - IV until the full - time - domain control action U(t)=[u(N),u(N + 1),...,u(T)] T , that is, there are a total of (T - N + 1) control decisions corresponding to the future moments N, (N + 1),…,T. Then the system state - space equation for the future with a total of (T - N + 1) steps can be expressed as:
[0053] X(t)=Dx(t)+FU(t) (13)
[0054] Among them, the finite - time - domain state matrix X(t), the finite - time - domain state - vector parameter matrix D, and the finite - time - domain control - vector parameter matrix F can be expressed by x(t), A, B, and the identity matrix E as:
[0055]
[0056] The above - mentioned step 50 specifically includes: As Figure 6 shown, the core of the DRLC algorithm is the design of the Markov chain process MDP and the design of the proximal policy optimization PPO agent structure. MDP defines the action space, state space, reward function, discount factor, and state - transition probability of the photovoltaic - energy - storage complementary control process. The MDP tuple Specifically:
[0057] I. State space Composed of environmental state variables, including k(t), r(t), g(t), p(t), B(t), x(t), etc. at time t, and can be expressed as:
[0058]
[0059] II. Action space Composed of control variables, including s(t), b + (t), b - (t), q(t), and the MPC process control vector u(t) at time t, and can be expressed as:
[0060]
[0061] For calculation needs, denote the combination of state elements of the set at time t as s t , the combination of state elements of the set at time t as at ,
[0062] III. Reward Function R: The real-time reward given by the environment to the PPO after it executes an action. On the premise of minimizing the objective function, the higher the utilization rate of the electric energy of the photovoltaic system directly used in the office building and the closer the battery energy storage system's power is to the target power, the higher the reward function value. Let the target energy storage power of the battery energy storage system be Q tar . If the reward coefficients of the photovoltaic system and the battery energy storage system are ρ and ν respectively, then the single-round reward function r t and the total reward function R can be designed as follows:
[0063]
[0064] IV. Discount Factor γ: The discount coefficient used in the reward function, where γ ∈ [0, 1].
[0065] V. State Transition Probability The probability that the current state s transitions to the new state s' after executing the action a, that is:
[0066]
[0067] The above-mentioned step 50 specifically includes: The PPO agent consists of two Actor networks (new policy network + old policy network), a Critic network (value network), and an experience buffer. First, initialize the important parameters θ and θ of the Actor network old and the important parameters of the Critic network Store the MPC processes x(t) and u(t) as the initial empirical state actions in the experience buffer. The experience buffer generates a policy π and runs for multiple rounds, storing the obtained (s , a t , r t , s t , s t+1 ) sequence. When the sequence volume in the experience buffer is sufficient, randomly sample a small batch of samples for training the Actor network and the Critic network. The trained (s t , a t , r t , s t+1 ) sequences are all fed back to the experience buffer, and at the same time, the policy and corresponding parameters are trained and updated to increase the possibility of the action that maximizes R in the next round of execution.
[0068] The parameters θ and θ of the Actor network old The advantage function A π introduced in the update process of (s t , a t ) can be expressed as:
[0069] Aπ (s t ,a t ) = Q π (s t ,a t ) - V π (s t ) (19)
[0070] The PPO aims to bring the expected reward J(θ) and solves it through gradient ascent ▽ θ J(θ): θ J(θ) solution:
[0071]
[0072] The old network parameter θ old is updated to θ according to J(θ) and the learning rate α of the Actor network:
[0073] θ ← θ old + α▽ θ J(θ)(21)
[0074] The Actor network updates the new policy and parameters, and uses the clipped objective J clip (θ) to limit the policy update range:
[0075]
[0076] where the ratio σ t (θ) of the new policy to the old policy and the clipping function clip(·) are respectively:
[0077]
[0078] The hyperparameter in the clipping process is ε, representing the difference between the old and new in the policy update process. J clip (θ) updates the policy within the range of [1 - ε, 1 + ε], and the minimization function min(·) ensures the safety of the policy update process at all times and prevents the training learning from crashing. The Critic network evaluates and updates the state value following the actions of the Actor network, and this process updates the Critic network parameters through the loss function and the learning rate β of the Critic network to
[0079]
[0080] The DRLC control algorithm constructs an MDP framework to simplify the mathematical model of the optimization problem of the complementary control of photovoltaic and energy storage, combines the variable parameters of the MPC control algorithm, and provides empirical input for the PPO. The PPO interacts with the environment and trains and learns based on the Actor-Critic network. The Actor network outputs actions that affect state transitions, and the Critic network evaluates the action value based on rewards to drive policy optimization, ultimately achieving intelligent management and control.
[0081] Specifically, step 60 above includes: As Figure 7 shown, select the typical operation dataset and energy consumption dataset of a certain office building in December 2022 to achieve intelligent management and control. Use MPC-DQN and MPC-Actor-Critic as comparison algorithms. The final cumulative costs of the MPC-DQN, MPC-Actor-Critic, and MPC-DRLC algorithms are yuan, yuan, yuan, respectively. The final cumulative cost of the MPC-DRLC algorithm is approximately 8.8% and 7.2% lower than that of the MPC-DQN and MPC-Actor-Critic. Define the sum of the number of times the control action of the energy storage system changes from charging to discharging or from discharging to charging at adjacent times as the number of charge-discharge action conversions N change of the energy storage system. Calculate the N change of the MPC-DQN, MPC-Actor-Critic, and MPC-DRLC, which are respectively. The MPC-DRLC algorithm is approximately 10.3% and 5.8% lower than the MPC-DQN and MPC-Actor-Critic. Figure 7 In train , each algorithm selects a small batch of samples from the experience buffer for n DQN = 2000 rounds of training and learning. The reward function values are accumulated in each round to obtain the convergence trends of the final cumulative reward function values of the MPC-DQN, MPC-Actor-Critic, and MPC-DRLC, which are R Actor-Critic →420, R PPO →450, and R →500 respectively. The number of training rounds when the final cumulative reward function value reaches convergence is
[0082] The above experimental results verify the effectiveness and scientificity of the algorithm proposed in the present invention.
[0082] Although the embodiments disclosed in the present invention are as described above, the above content is only an embodiment adopted for the convenience of understanding the present invention and is not intended to limit the present invention. Any person skilled in the art within the technical field to which the present invention pertains may make any modifications and changes in the form of implementation and details without departing from the spirit and scope disclosed by the present invention. However, the scope of patent protection of the present invention shall still be subject to the scope defined by the appended claims.
Claims
1. A photovoltaic-storage complementary intelligent storage control algorithm using MPC-DRLC, characterized in that: The following steps are involved: Step 10: Obtain office building operation data set and energy consumption data set; Step 20: construct a mathematical model for the optimization problem of photovoltaic-storage complementarity; Step 30 constructs an SSA-LSTM prediction model to predict photovoltaic power generation; Step 40: Design a model predictive control MPC algorithm to achieve short-term intelligent control; Step 50: Design a deep reinforcement learning control DRLC algorithm, and combine it with the model predictive control to obtain the MPC-DRLC algorithm to achieve short-term and long-term global intelligent control; Step 60 verifies the algorithm performance through electricity costs, the number of charge and discharge action conversions of the energy storage system, and the reward function evaluation index.
2. The optical-storage complementary intelligent storage control algorithm using MPC-DRLC according to claim 1 is characterized in that: The office building operation data set obtained in step 10 includes: a time series operation data set of atmospheric temperature, photovoltaic panel patch temperature, direct radiation, diffuse radiation, total radiation and power generation power collected by the office building photovoltaic system, and the data set is divided into a training set, a test set and a validation set according to an 8:1:1 ratio.
3. The optical-storage complementary intelligent storage control algorithm using MPC-DRLC according to claim 2 is characterized in that: The step 20 comprises: At time t, the photovoltaic system, energy storage system, and power grid supply energy s(t) and b(t) to the office building respectively. - (t), d(t); the photovoltaic system and the power grid store energy to the energy storage system respectively b + (t), q(t), set the optimization objective function (1) and constraints (2) to (8): Where T is the total number of control time points, = is the set of {1, time 2, time..., point T}, The predicted power generation of the photovoltaic system is p(t), the power demand of the user is r(t), and the real-time electricity price is k(t); the upper limits of the single charge and discharge of the energy storage system are c + 、c - ; The storage capacity and upper limit of the energy storage system are B(t) and Q respectively; The charging and discharging efficiencies of the energy storage system are η respectively + , η - ; Energy storage system constraint constant b storage , P storage Constraints are imposed on the energy storage action, and the probability of event [·] occurring is In formula (2), the photovoltaic system predicts that the power generation p(t) is discharged to the office building and the energy storage system; in formula (3), the office building electricity r(t) is provided by the photovoltaic system, the energy storage system and the power grid.
4. The optical-storage complementary intelligent storage control algorithm using MPC-DRLC according to claim 1 is characterized in that: In step 30, the photovoltaic power is predicted by using the long short-term memory network LSTM model, including: the input at time t includes the memory cell state c at time (t-1) t-1 , hidden layer state h t-1 , input x at time t t ; The output includes the memory cell state c at time t t , hidden layer state h t ; Let the weights of the forget gate, input gate, and output gate be W respectively f , W i , W o , the biases of the forget gate, input gate, and output gate are b respectively f 、b i 、b o , the Sigmoid activation function is σ(x), h t-1 With x t Vector concatenation is [h t-1 ,x t ], then the forget gate, input gate, and output gate calculate the output f respectively through σ(x) t 、i t , o t : Let the candidate cell state weight and bias be W c 、b c , the input gate passes through tanh(x)=(e x -e -x ) / (e x +e -x ) function calculates the candidate memory cell state C t , forget gate, input gate use bitwise multiplication Update c t Information, the output gate calculates h t :
5. The optical-storage complementary intelligent storage control algorithm using MPC-DRLC according to claim 4 is characterized in that: Apply the sparrow search algorithm SSA to automatically tune the LSTM hyperparameters, including: automatic hyperparameter tuning through initialization and producer, forager, and danger position update; initialization process input optimization target prediction window length Hidden layer output dimension Fully connected layer output dimension Dropout strategy random loss rate and learning rate Hyperparameters, initialize the size of the sparrow population and related algorithm parameters, sort the fitness values to determine the individual positions; the position update process updates the positions of producers, foragers and dangerous people according to the judgment conditions, and when the upper limit of the number of iterations is reached, the current global optimal position, global optimal fitness value and corresponding combination.
6. The optical-storage complementary intelligent storage control algorithm using MPC-DRLC according to claim 1 is characterized in that: In step 40, SSA-LSTM is used to predict the power generation p(t) of the office building and the actual power generation g(t), and the state combination of B(t) at time t and g(t-1) at time (t-1) is formed into a state vector x(t)=[B(t),g(t-1)] T ,s(t),b + (t), b - The control decision combination of s(t) and q(t) is the control vector u(t) = [s(t), b + (t),b - (t),q(t)] T , then the state vector x(t+1) at time (t+1) can be expressed by the state space equation of the photovoltaic-storage complementary system: x(t+1)=Ax(t)+Bu(t) (11) The state vector parameter matrix A and the control vector parameter matrix B are: Through rolling optimization, continuous and rolling control of short-term states is achieved in the time domain.
7. The optical-storage complementary intelligent storage control algorithm using MPC-DRLC according to claim 1 is characterized in that: In step 50, the core of the DRLC algorithm is the design of the Markov chain process MDP and the design of the proximal strategy optimization PPO agent structure; the MDP defines the action space, state space, reward function, discount factor, and state transition probability of the photovoltaic storage complementary control process, and obtains the MDP tuple in, is the state space, is the action space, R is the reward function, γ is the discount factor, is the state transition probability.
8. The optical-storage complementary intelligent storage control algorithm using MPC-DRLC according to claim 7 is characterized in that: The state space It consists of environmental state variables, including k(t), r(t), g(t), p(t), B(t), and x(t) at time t, expressed as: Action Space It consists of control variables, including s(t) at time t, b + (t), b - (t), q(t) and MPC process control vector u(t), expressed as: Reward function R: The real-time reward after the environment executes the action on PPO. Under the premise of minimizing the objective function, the higher the utilization rate of the photovoltaic system electricity used in the office building and the closer the energy storage system power is to the target power, the higher the reward function value is. Let the target energy storage power of the energy storage system be Q tar , the reward coefficients of the photovoltaic system and the energy storage system are ρ and ν respectively, then the single-round reward function r t , the total reward function R is designed as: Discount factor γ: discount factor for reward function, γ∈[0,1]; State transition probability After executing action a, the probability of the current state s turning into the new state s' is:
9. The optical-storage complementary intelligent storage control algorithm using MPC-DRLC according to claim 7 is characterized in that: The PPO agent consists of two Actor networks, namely the new strategy network and the old strategy network, a Critic network, and an experience buffer, including initializing important parameters of the Actor network θ, θ old Important parameters of the Critic network The MPC process x(t) and u(t) are stored as the initial experience state action in the experience buffer. Generate strategy π and run multiple rounds, store the obtained (s t ,a t ,r t ,s t+1 ) sequence; when the experience buffer sequence is sufficient, randomly select small batch samples for Actor network and Critic network training, and the training obtains (s t ,a t ,r t ,s t+1 ) sequence and feed it back to the experience buffer, while training and updating the strategy and corresponding parameters to increase the possibility of the action that can maximize R in the next round of execution.
10. The optical-storage complementary intelligent storage control algorithm using MPC-DRLC according to claim 1 is characterized in that: The office building operation data set and energy consumption data set are selected to realize intelligent management and control. MPC-DQN and MPC-Actor-Critic are used as comparison algorithms. The cumulative electricity cost, the number of charge and discharge action conversions of the energy storage system, and the reward function are used to evaluate and verify the algorithm.
Citation Information
Patent Citations
Micro power grid wind and solar energy storage model prediction control method
CN104967149A
Wind-solar-storage combined system optimization method based on multi-target grey wolf algorithm
CN116914856A
Residential integrated energy system optimization control method based on multi-agent reinforcement learning
CN119024707A
Cited By
Photovoltaic regulation and control method and system based on artificial intelligence
CN121097680A