Dynamic charging and discharging strategy collaborative optimization method for prolonging service life of energy storage system

By combining deep reinforcement learning with electricity price forecasting and battery health modeling to develop a dynamic charging and discharging strategy, the problem of strategic adaptability of energy storage systems in the power market environment is solved, economic benefits are improved, and battery life is extended.

CN120601494AActive Publication Date: 2025-09-05KUNMING AUTOMATION WHOLE SET OF EQUIP BUSINESS CO LTD

Patent Information

Application Number
CN202511108549.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-05
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

The charging and discharging strategies of existing energy storage systems are unable to adapt to the rapidly changing electricity market price signals, resulting in low profitability and operating efficiency. Traditional optimization methods cannot accurately reflect the degradation cost of batteries in different states of charge and health, and cannot effectively extend battery life.

Method used

A dynamic charging and discharging strategy is adopted that combines deep reinforcement learning with accurate electricity price prediction and dynamic battery health modeling. Through long short-term memory networks and deep reinforcement learning agents, charging and discharging power control instructions are generated to optimize the charging and discharging behavior of the energy storage system.

Benefits of technology

It has achieved coordinated optimization of charging and discharging strategies in a complex market environment, significantly improved the economic benefits of the energy storage system throughout its life cycle, and effectively extended the service life of the battery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120601494A_ABST
    Figure CN120601494A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of energy storage systems, in particular to a dynamic charging and discharging strategy collaborative optimization method for prolonging the service life of an energy storage system, and the method comprises the steps: inputting historical electricity price time sequence data based on a preset long-short-term memory network model, so as to generate an electricity price prediction sequence; constructing a state vector, wherein the state vector comprises an internal state of the energy storage system at the current moment t and an electricity price prediction sequence; the state vector is input into a pre-trained deep reinforcement learning agent, and the deep reinforcement learning agent outputs a selected specific action based on a predefined discrete charging and discharging action set according to the input state vector; and converting the output specific action into a charging and discharging power control instruction, and issuing the charging and discharging power control instruction to the energy storage converter for execution. According to the method, through deep reinforcement learning, electricity price accurate prediction and battery health dynamic modeling are combined, and charge and discharge strategy collaborative optimization of the energy storage system in a complex market environment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of energy storage systems, and in particular to a method for collaboratively optimizing dynamic charging and discharging strategies for extending the life of energy storage systems. Background Art

[0002] As the core asset of energy storage systems, batteries have limited cycle life and health status. Frequent or improper charging and discharging will accelerate their capacity decay and performance degradation. Therefore, how to formulate an intelligent charging and discharging strategy that can maximize economic benefits while taking into account battery health and extending its service life has become a core scientific problem and key technical bottleneck that needs to be urgently solved in the field of energy storage technology. Existing energy storage system charging and discharging strategies have significant limitations. On the one hand, the widely used rule-based control strategies based on fixed thresholds have rigid logic and cannot adapt to the ever-changing electricity market price signals. They often miss arbitrage opportunities or lead to uneconomical charging and discharging behaviors, resulting in low profitability and operating efficiency. On the other hand, although traditional mathematical programming optimization methods can seek theoretically optimal solutions, they not only require extremely high model accuracy and complex calculation processes, making it difficult to meet the needs of real-time control, but also often use overly simplified models when quantifying the complex and nonlinear electrochemical process of long-term battery health loss. These models cannot accurately reflect the true degradation cost of batteries under different states of charge (SOC) and states of health (SOH), resulting in their decisions not being ideal in terms of extending battery life. Summary of the Invention

[0003] The purpose of the present invention is to provide a dynamic charging and discharging strategy collaborative optimization method for extending the life of energy storage systems. Through deep reinforcement learning, combined with accurate electricity price prediction and dynamic battery health modeling, it realizes the collaborative optimization of the charging and discharging strategy of the energy storage system in a complex market environment, significantly improves the economic benefits of the entire life cycle, and effectively extends the service life of the system.

[0004] The present invention is achieved through the following technical solutions: A dynamic charging and discharging strategy collaborative optimization method for extending the life of an energy storage system includes the following steps: Based on the preset long short-term memory network model, historical electricity price time series data is input to generate an electricity price forecast sequence; Constructing a state vector, wherein the state vector includes the internal state of the energy storage system at the current time t and the electricity price prediction sequence; Inputting the state vector into a pre-trained deep reinforcement learning agent, the deep reinforcement learning agent outputting a specific action selected from a predefined discrete charge and discharge action set according to the input state vector; The specific output actions are converted into charge and discharge power control instructions and sent to the energy storage converter for execution.

[0005] Optionally, the preset long short-term memory network model is constructed as follows: Construct a neural network structure consisting of an input layer, an LSTM hidden layer, and a fully connected output layer; Taking N time steps of historical electricity price time series data as input of the input layer; Serializing N historical electricity price data points through the LSTM hidden layer, updating its internal state at each time step to extract feature information encoding the time dependency of the historical electricity price data; The fully connected output layer generates a time series of electricity price prediction values ​​including M time steps based on the extracted feature information, which is represented as an electricity price prediction sequence.

[0006] Optionally, the constructing state vector is specifically: Obtain the current state of charge SOC(t) and health state SOH(t) of the energy storage system as internal state data; Obtaining the electricity price prediction sequence including M electricity price prediction values ​​from the current time t+1 to t+M; The internal state data and the electricity price prediction sequence are sequentially spliced ​​in a preset order to form the state vector S(t).

[0007] Optionally, the state vector S(t) is calculated as follows:

[0008] Where S(t) represents the complete state vector at time t; SOC(t) represents the battery state of charge at time t, a scalar between 0 and 1; SOH(t) represents the battery state of health at time t, a scalar between 0 and 1; Represents the predicted The grid electricity price at the moment; M is the total number of time steps predicted by the long short-term memory network model.

[0009] Optionally, the deep reinforcement learning agent is specifically carried by a deep Q network, and the construction process of the deep reinforcement learning agent is as follows: Construct a multilayer perceptron network structure consisting of an input layer, at least two fully connected hidden layers, and an output layer; Setting the number of neurons in the input layer to match the dimension of the state vector S(t); The number of neurons in the output layer is set to be exactly the same as the total number of actions in the predefined discrete charge and discharge action set.

[0010] Optionally, the specific training process of the deep reinforcement learning agent is as follows: Initialize the evaluation network and target network, and prepare an experience replay pool for storing interaction data; In the battery simulation environment, each interaction in the battery simulation environment is evaluated using the defined immediate reward function R(t) and generates a data tuple containing state, action, reward, and next state; Storing the data tuple in an experience replay pool; Randomly extract data from the experience replay pool to calculate the target Q value and the current Q value; According to the loss between the target Q value and the current Q value, the weight of the deep Q network is iteratively updated until the maximum number of iterations is reached, thereby completing the training of the deep reinforcement learning agent. The action value function Q of the deep reinforcement learning agent can accurately predict the expected value of the long-term cumulative reward defined in claim 7 that can be brought about by executing each candidate action based on any input state.

[0011] Optionally, the instantaneous reward function R(t) is solved as follows: Extracting the predicted electricity prices corresponding to each of the next N time steps based on the state vector S(t); Each of the N predicted electricity prices is associated with the currently selected specific action Multiply the corresponding charge and discharge power values ​​to obtain N single-step expected benefit values; The N single-step expected return values ​​are cumulatively summed, and the result of the cumulative sum is determined as the expected economic return item for the next N steps; The current state of charge SOC(t) and health state SOH(t) are used as inputs, and the health cost coefficient related to the state is output through the preset health cost coefficient function k; Combine the health cost coefficient with the currently selected specific action The corresponding preset action stress factor f is multiplied, and the product is determined as the state-dependent health cost item; Multiplying the expected economic benefit item in the next N steps by a first preset weight α to obtain a first calculation result; Multiplying the state-dependent health cost item by a second preset weight β to obtain a second calculation result; The first calculation result is subtracted from the second calculation result, and the final difference is represented as the immediate reward function R(t).

[0012] Optionally, the preset health cost coefficient function k is further defined to have the following output characteristics: The region between the first preset SOC threshold and the second preset SOC threshold is defined as a normal SOC operating range; when the input state of charge SOC(t) is outside the normal SOC operating range, the output value of the health cost coefficient function k is higher than the output value when it is within the normal SOC operating range; When the input health state SOH(t) is lower than the preset health level threshold, the output value of the health cost coefficient function k increases nonlinearly as SOH(t) decreases.

[0013] Optionally, the specific action selected based on the output from the predefined discrete charge and discharge action set is specifically: The current state vector S(t) is input into the trained deep reinforcement learning agent to perform each preset candidate action. , get the predicted corresponding action value ; Among all the action values ​​obtained, select the candidate action corresponding to the maximum value , and determine it as a specific output action .

[0014] Optionally, the specific action of outputting is converted into a charge and discharge power control instruction, which is specifically: According to the specific output action , determine the corresponding nominal charge and discharge rate; Obtaining the preset nominal energy capacity of the energy storage system; Multiplying the nominal charge and discharge rate by the nominal energy capacity to determine the product as a target charge and discharge power instruction; The target charge and discharge power instructions are sent to the energy storage converter of the energy storage system as the power control setting value of the energy storage converter.

[0015] The technical solution of the present invention has at least the following advantages and beneficial effects: On the one hand, the present invention uses long-short-term memory networks to deeply mine the temporal dependencies in historical electricity price data, achieving accurate long-term predictions of future electricity prices and providing forward-looking market insights for the decision-making of intelligent agents. On the other hand, the present invention innovatively constructs a composite instant reward function, which not only takes into account long-term economic benefits, but also introduces a health cost term that is dynamically related to the battery state of charge and health status, and can accurately quantify the marginal impact of each charging and discharging decision on the battery life. Moreover, the present invention, through deep reinforcement learning intelligent agents, can perform end-to-end learning and decision-making in this complex, multi-dimensional state space. At each decision-making moment, it can proactively find an optimal balance point that takes into account both short-term benefit maximization and long-term life cost minimization, thereby significantly improving the comprehensive economic benefits of the energy storage system throughout its life cycle. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A flow chart of the method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system provided by the present invention; Figure 2 Schematic diagram of the core module principle of the dynamic charge-discharge strategy collaborative optimization method for extending the life of energy storage systems provided by the present invention. DETAILED DESCRIPTION

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0018] Reference Figure 1 As shown, Figure 1 A flow chart of the collaborative optimization method for dynamic charge-discharge strategies for extending the life of energy storage systems provided by the present invention.

[0019] In one embodiment, a method for collaboratively optimizing dynamic charging and discharging strategies for extending the life of an energy storage system includes the following steps: Based on the preset long short-term memory network model, historical electricity price time series data is input to generate an electricity price forecast sequence; Constructing a state vector, wherein the state vector includes the internal state of the energy storage system at the current time t and the electricity price prediction sequence; Inputting the state vector into a pre-trained deep reinforcement learning agent, the deep reinforcement learning agent outputting a specific action selected from a predefined discrete charge and discharge action set according to the input state vector; The specific output actions are converted into charge and discharge power control instructions and sent to the energy storage converter for execution.

[0020] like Figure 2 As shown, the method of this embodiment is deployed on a computing device, such as an industrial computer, server, or embedded controller, which communicates with the energy storage system's battery management system (BMS) and power storage converter (PCS). The core of the entire method is a deep reinforcement learning agent, acting like an intelligent brain. At each decision moment t, it perceives the environmental state and makes optimal charging and discharging decisions. Specifically, this embodiment of the present invention can be divided into the following core modules: The electricity price prediction module, whose core is a pre-trained long short-term memory (LSTM) model. It receives historical electricity price data and, based on this data, predicts future price trends, providing a key basis for the agent's economic decision-making. The state construction module: This module collects real-time internal battery states (such as SOC and SOH) from the energy storage system's BMS at each decision moment t. Combined with the future electricity price sequence output by the electricity price prediction module, it constructs a comprehensive state vector S(t). This vector serves as the agent's sole entry point for perceiving the environment. The decision-making agent module, the heart of this embodiment, is powered by a pre-trained deep Q-network (DQN). It receives the state vector S(t) generated by the state building module and outputs a discrete charge and discharge action that is optimal under the current state. The control execution module is responsible for converting the abstract actions output by the agent into Translate it into specific, executable engineering instructions, namely the target charge and discharge power instructions, and send them to the PCS of the energy storage system, which accurately executes the power instructions.

[0021] Specifically, the construction process of the preset long short-term memory network model is as follows: Construct a neural network structure consisting of an input layer, an LSTM hidden layer, and a fully connected output layer; Taking N time steps of historical electricity price time series data as input of the input layer; Serializing N historical electricity price data points through the LSTM hidden layer, updating its internal state at each time step to extract feature information encoding the time dependency of the historical electricity price data; The fully connected output layer generates a time series of electricity price prediction values ​​including M time steps based on the extracted feature information, which is represented as an electricity price prediction sequence.

[0022] During implementation, the first step of this embodiment is to generate an accurate forecast of future electricity prices, which is a prerequisite for achieving proactive economic optimization. In this embodiment, the electricity price prediction module adopts a long short-term memory network (LSTM) model. The reason for choosing LSTM is that the electricity market price has significant time series characteristics, that is, the current price is highly correlated with the price in the past period of time. As a special recurrent neural network (RNN), LSTM's internal sophisticated gating structure (forget gate, input gate, output gate) can effectively learn and memorize long-term dependencies in time series data, thereby performing better than traditional statistical models or ordinary feedforward neural networks in prediction tasks. The specific construction process of the LSTM model is as follows: Step S201: Construct the network structure. The model mainly consists of three parts: an input layer, one or more LSTM hidden layers, and a fully connected output layer. Step S202: Define the input. The input layer receives the historical electricity price time series data of the past N time steps, expressed as .in, is the actual electricity price at the i-th time step before the current time t. The time step length can be determined based on the actual application scenario. For example, if electricity price data points are received every 15 minutes, N can be set to 96, representing the input of historical data for the past 24 hours. Step S203: Feature Extraction. After input, the data flows through the LSTM hidden layer. In this layer, the LSTM unit uses its internal cell state (CellState) and gating mechanism to serialize the 96 input historical electricity price data points. At each time step, the LSTM unit updates its internal state, gradually extracting high-level feature information that encodes the temporal dependencies of the historical electricity price data. For example, it can learn to recognize typical bimodal patterns throughout the day, or the difference in electricity prices between weekdays and weekends. Step S204: Generate a prediction sequence. The feature information extracted by the LSTM hidden layer is finally passed to the fully connected output layer. This output layer has a standard multilayer perceptron structure with M neurons. It decodes the feature information encoded by the LSTM layer and finally generates a time series containing the electricity price forecast values ​​for the next M time steps, which is represented as the electricity price forecast sequence . Similarly, if the time step is 15 minutes, M can be set to 24, which means that the electricity price for the next 6 hours is predicted. The LSTM model of this embodiment is trained offline by supervised learning. The training data set can be a large amount of historical electricity price data obtained from power grid operators or power trading centers. During training, the historical data is divided into a large number of (input, label) data pairs, where the input is a sequence of N steps, and the label is the actual electricity price sequence of the subsequent M steps. An optimizer such as Adam is used with the goal of minimizing the mean square error (MSE) between the predicted value and the true value, and the network weights are iteratively updated through the backpropagation algorithm until the model converges.

[0023] In this embodiment, the state vector is constructed as follows: Obtain the current state of charge SOC(t) and health state SOH(t) of the energy storage system as internal state data; Obtaining the electricity price prediction sequence including M electricity price prediction values ​​from the current time t+1 to t+M; The internal state data and the electricity price prediction sequence are sequentially spliced ​​in a preset order to form the state vector S(t).

[0024] The state vector S(t) is calculated as follows:

[0025] Where S(t) represents the complete state vector at time t; SOC(t) represents the battery state of charge at time t, a scalar between 0 and 1; SOH(t) represents the battery state of health at time t, a scalar between 0 and 1; Represents the predicted The grid electricity price at the moment; M is the total number of time steps predicted by the long short-term memory network model.

[0026] Specifically, after obtaining a future electricity price forecast, this embodiment constructs a state vector S(t) for the deep reinforcement learning agent to comprehensively describe the current decision-making environment. The state vector is the sole source of information for the agent's decision-making, and its design directly impacts the ultimate control effectiveness. The construction of the state vector S(t) is as follows: Step S301: Obtaining internal state data. At each decision time t, the system communicates with the energy storage system's battery management system (BMS) to obtain two key battery internal state data in real time: the current state of charge (SOC(t), which represents the percentage of remaining battery charge. Its value is normalized between 0 and 1. SOC is the fundamental constraint that determines whether charging or discharging is currently possible. The current state of health (SOH(t),) represents the degree to which the battery's performance has been retained compared to a new battery, typically measured as capacity retention. Its value is also normalized between 0 and 1. SOH is a key metric for assessing long-term battery lifespan loss. Step S302: Obtaining a sequence of electricity price forecasts. Directly call the output of the electricity price prediction module, which is a sequence of M electricity price prediction values ​​from the next moment t+1 to t+M from the current moment . Step S303: Splicing to form a state vector. The internal state data obtained above and the electricity price prediction sequence are sequentially spliced ​​in a pre-set fixed order. The specific fixed order refers to the order in the above state vector S(t) calculation formula to form the final state vector S(t). It can be understood that if M=24, S(t) is a 26-dimensional vector. Among them, the first dimension is SOC, the second dimension is SOH, and the third to twenty-sixth dimensions are the predicted electricity prices for the next 6 hours (one point every 15 minutes). This embodiment applies this design so that the state vector S(t) simultaneously includes the physical state of the energy storage system itself (SOC, SOH) and the economic environment of the external market (future electricity prices), providing complete information for the intelligent agent to make decisions that take both into account.

[0027] In the specific application of this embodiment, the deep reinforcement learning agent is specifically carried by a deep Q network, and the construction process of the deep reinforcement learning agent is as follows: Construct a multilayer perceptron network structure consisting of an input layer, at least two fully connected hidden layers, and an output layer; Setting the number of neurons in the input layer to match the dimension of the state vector S(t); The number of neurons in the output layer is set to be exactly the same as the total number of actions in the predefined discrete charge and discharge action set.

[0028] In this embodiment, the deep reinforcement learning agent is implemented using a deep Q-network (DQN). DQN is a combination of deep learning and Q-learning, making it particularly well-suited for decision-making problems with high-dimensional state spaces and discrete action spaces, making it a perfect fit for the application scenarios of this invention. The DQN network structure is a standard multi-layer perceptron (MLP). Its construction process is as follows: Step S501: Build the network structure. The network consists of an input layer, at least two fully connected hidden layers, and an output layer. Step S502: Configure the input layer. The number of neurons in the input layer must match the dimensionality of the state vector S(t). In our example, the dimensionality is 2 + M, or 26 neurons. Step S503: Configure the hidden layers. To extract the complex nonlinear features in the state vector, at least two fully connected hidden layers are configured. As you can see, the first hidden layer can have 256 neurons, and the second hidden layer can have 128 neurons. Each neuron is followed by a nonlinear activation function, such as a rectified linear unit (ReLU), to enhance the network's expressive power. Step S504: Configure the output layer. The number of neurons in the output layer must be exactly the same as the total number of actions in the predefined discrete charge and discharge action set A, and the predefined discrete charge and discharge action set A must be defined first. To make the output of the agent easier to process, this embodiment discretizes the continuous charge and discharge power. As an example, a set of 5 actions is defined: , where each action corresponds to a specific nominal charge and discharge rate, for example: Corresponding to -1.0C (discharging at 1x maximum power); Corresponding to -0.5C (discharge at 0.5 times rate); Corresponding to 0C (not charging or discharging, idle); Corresponding to +0.5C (charging at 0.5 times rate); Corresponding to +1.0C (charging at 1x maximum power). Therefore, the output layer should have 5 neurons. When a state S(t) is input, these 5 neurons will output a Q value, which represents the execution of arrive The expected value of the long-term cumulative return that can be obtained from these five actions is output as .

[0029] Specifically, the training process of the deep reinforcement learning agent is as follows: Initialize the evaluation network and target network, and prepare an experience replay pool for storing interaction data; In the battery simulation environment, each interaction in the battery simulation environment is evaluated using the defined immediate reward function R(t) and generates a data tuple containing state, action, reward, and next state; Storing the data tuple in an experience replay pool; Randomly extract data from the experience replay pool to calculate the target Q value and the current Q value; According to the loss between the target Q value and the current Q value, the weight of the deep Q network is iteratively updated until the maximum number of iterations is reached, thereby completing the training of the deep reinforcement learning agent. The action value function Q of the deep reinforcement learning agent can accurately predict the expected value of the long-term cumulative reward defined in claim 7 that can be brought about by executing each candidate action based on any input state.

[0030] In this embodiment, the training process of the deep reinforcement learning agent is an iterative process carried out in a battery simulation environment and follows the DQN classic algorithm framework. Step S501: Initialization. Construct two DQN networks with identical structures: Evaluation Network and target network . The evaluation network is used to generate the current Q-value prediction and perform gradient updates during training. The target network is used to generate the target Q-value. Its weights are periodically copied from the evaluation network to increase the stability of training. Initialize the experience replay pool D. This is a first-in-first-out queue used to store data generated by the interaction between the agent and the environment, breaking the temporal correlation between the data. Step S502: Environmental interaction and data acquisition. Training is performed in a loop. At each time step t: the agent observes the current state S(t) of the simulation environment. The agent uses the ε-greedy strategy to select an action That is, the action with the largest Q value considered by the current evaluation network is selected with a probability of 1-ε, and an action is randomly selected with a probability of ε. The simulation environment executes the action , and calculate the next state S(t+1) after the action is executed and the immediate reward value R(t) obtained from this interaction based on the internal battery model and electricity price model. Store in the experience replay pool D. Step S503: Network learning and updating. When the amount of data in the experience replay pool reaches a certain size, network learning begins: randomly extract small batches of data tuples from the experience replay pool D. For each data tuple, two calculations are performed: Calculate the target Q value: Where γ is a discount factor (such as 0.99), which indicates the importance of future rewards. The target network is used to predict the maximum future value that the next state can bring. Calculate the current Q value: . Used to evaluate the network performance under S(t) The current estimate of the value of . Calculate the loss. Based on the difference between the target Q value y and the current Q value, calculate the loss function: . Through the back propagation algorithm, the loss function is calculated to evaluate the network The gradient of each layer weight is updated along the direction of gradient descent using optimizers such as Adam The weight of . Every certain number of steps, the network will be evaluated All weights of are completely copied to the target network Step S504: Repeat steps S502 and S503 until the preset maximum number of iterations or the total number of training steps is reached. After the training is completed, the evaluation network is obtained. This is a trained deep reinforcement learning agent. Its internal action-value function Q is now able to accurately predict the expected value of the long-term cumulative reward of executing each candidate action, determined by the defined immediate reward function, based on any input state.

[0031] The immediate reward function R(t) defined above is solved as follows: Extracting the predicted electricity prices corresponding to each of the next N time steps based on the state vector S(t); Each of the N predicted electricity prices is associated with the currently selected specific action Multiply the corresponding charge and discharge power values ​​to obtain N single-step expected benefit values; Accumulating and summing the N single-step expected benefit values, and determining the result of the accumulated sum as the expected economic benefit item for the next N steps; The current state of charge SOC(t) and health state SOH(t) are used as inputs, and the health cost coefficient related to the state is output through the preset health cost coefficient function k; Combine the health cost coefficient with the currently selected specific action The corresponding preset action stress factor f is multiplied, and the product is determined as the state-dependent health cost item; Multiplying the expected economic benefit item in the next N steps by a first preset weight α to obtain a first calculation result; Multiplying the state-dependent health cost item by a second preset weight β to obtain a second calculation result; The first calculation result is subtracted from the second calculation result, and the final difference is represented as the instant reward value R(t).

[0032] In this embodiment, the immediate reward function R(t) is the baton for agent training, defining what constitutes good behavior and what constitutes bad behavior. One of the core innovations of this embodiment is the design of a composite reward function that can synergistically optimize economic benefits and battery life. The solution steps are as follows: Step S701: Calculate the expected economic benefits in the next N steps, aiming to quantify the action First, from the current state vector S(t), extract the predicted electricity price corresponding to each of the next N time steps Here N can be equal to M or less than M, representing the time window of short-term economic returns that the agent focuses on. Then, the currently selected specific action Convert to corresponding charge and discharge power values , charging is a negative profit, discharging is a positive profit. Each of the N predicted electricity prices is respectively compared with the power value Multiply them together to get N single-step expected profit values. Finally, add up these N single-step expected profit values ​​and determine the result as the expected economic profit item for the next N steps.

[0033] Step S702: Calculate the state-dependent health cost item to quantify the action The cost of damage to battery life. It is obtained by multiplying two parts: the first part: the health cost coefficient function k. This is a preset function Its output value is closely related to the battery's current state, reflecting the battery's vulnerability under different conditions. Regarding SOC characteristics, two SOC thresholds are set, for example, the first preset SOC threshold is 20% and the second preset SOC threshold is 90%. The region between these two thresholds (20%-90%) is defined as the normal SOC operating range. When the input SOC(t) lies outside this range, i.e., below 20% or above 90% in the over-discharge / overcharge region, the output value of function k will be significantly higher than when it is within the normal range. This forms a U-shaped curve, indicating that charging or discharging when the battery's charge level is too low or too high will result in a sharp increase in health cost. Regarding SOH characteristics, a preset health level threshold is set, for example, 95%. When the input SOH(t) falls below this threshold, the output value of function k will increase nonlinearly as SOH(t) decreases. This means that for an older battery that has already experienced a certain degree of degradation, the health cost of an action that causes the same damage to it will be much higher than that of a new battery. In the implementation of this embodiment, the function k can be set as a piecewise function, or as a two-dimensional lookup table obtained by fitting a large amount of experimental data. Part II: Action stress factor f, which is specifically related to the specific action Related preset values It characterizes the magnitude of the physical stress on the battery at different charge and discharge rates. Generally, the greater the charge and discharge rate, the greater the damage to the battery. Therefore, it can be preset: .For example, , , Multiplying these two parts, the health cost item This product term can dynamically and accurately quantify the health damage caused by performing a specific action under the current battery state.

[0034] Step S703: Calculate the final immediate reward value R(t). This embodiment introduces two preset weighting coefficients: the first preset weight α (economic benefit weight) and the second preset weight β (health cost weight). These weights are hyperparameters, allowing energy storage system operators to adjust them based on their operating strategy (aggressive pursuit of profit or conservative pursuit of longevity). Multiply the economic benefit term in step S701 by α to obtain the first calculation result. Multiply the health cost term in step S702 by β to obtain the second calculation result. The final immediate reward value R(t) is obtained by subtracting the second calculation result from the first calculation result. By treating health cost as a negative reward term, the intelligent agent must learn to pursue economic benefits while actively avoiding behaviors that incur high health costs, i.e., damage battery life, in order to maximize long-term cumulative rewards during training.

[0035] More specifically, the specific action selected based on the output from the predefined discrete charge and discharge action set is: The current state vector S(t) is input into the trained deep reinforcement learning agent to perform each preset candidate action. , get the predicted corresponding action value ; Among all the action values ​​obtained, select the candidate action corresponding to the maximum value , and determine it as the specific output action .

[0036] The specific action of converting the output into a charge and discharge power control instruction is specifically: According to the specific output action , determine the corresponding nominal charge and discharge rate; Obtaining the preset nominal energy capacity of the energy storage system; Multiplying the nominal charge and discharge rate by the nominal energy capacity to determine the product as the target charge and discharge power instruction; The target charge and discharge power instructions are sent to the energy storage converter of the energy storage system as the power control setting value of the energy storage converter.

[0037] In this embodiment, after the training of the intelligent agent is completed, it is deployed to the actual energy storage system for online real-time reasoning and control. The workflow is as follows: Step S601: Acquire and construct the state S(t). The system periodically executes a decision. At the decision time t, all the functions of the aforementioned state construction module are first executed to obtain the real-time SOC(t) and SOH(t), and the LSTM model is called to predict the future electricity price, and finally spliced ​​into the current state vector S(t). Step S602: The intelligent agent makes a decision, and takes the current state vector S(t) as input and sends it to the trained deep Q network (i.e., the decision-making intelligent agent module). The network performs a forward propagation calculation, and the output layer will perform a forward propagation calculation for each preset candidate action. (Right now arrive ), giving a predicted action value The system selects the candidate action with the largest Q value from all the action values ​​obtained. , and determine it as the current optimal specific output action . This process is called Greedy Policy. Step S603: Action conversion and instruction issuance, the action output by the agent It is an abstract symbol that needs to be converted into physical instructions that the PCS can understand. This process is completed by the control execution module. , consult a predefined action-rate mapping table to determine a corresponding nominal charge and discharge rate. For example, if yes , the table query yields a nominal charge and discharge rate of +0.5C. The preset nominal energy capacity of the energy storage system is obtained from the system configuration. For example, a 1MWh energy storage system has a nominal energy capacity of 1000kWh. The nominal charge and discharge rate is multiplied by the nominal energy capacity to obtain the target charge and discharge power command. In the above example, the target charge and discharge power command = 0.5 * 1000kWh = 500kW (charging). The calculated target charge and discharge power command (500kW) is sent to the energy storage system's energy storage converter PCS via the communication interface. After receiving this command, the PCS adjusts the switching state of its internal power electronic devices, such as the IGBT, to precisely control the DC-side battery to charge at a power of 500kW. By repeatedly executing steps S601 to S603, this embodiment implements fully automatic, intelligent, closed-loop control of the energy storage system with the dual goals of extending life and improving economic efficiency.

[0038] The above is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A dynamic charging and discharging strategy collaborative optimization method for extending the life of an energy storage system, characterized by: The steps of the method include: Based on the preset long short-term memory network model, historical electricity price time series data is input to generate an electricity price forecast sequence; Constructing a state vector, wherein the state vector includes the internal state of the energy storage system at the current time t and the electricity price prediction sequence; Inputting the state vector into a pre-trained deep reinforcement learning agent, the deep reinforcement learning agent outputting a specific action selected from a predefined discrete charge and discharge action set according to the input state vector; The specific output actions are converted into charge and discharge power control instructions and sent to the energy storage converter for execution.

2. The method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system according to claim 1, characterized in that: The construction process of the preset long short-term memory network model is as follows: Construct a neural network structure consisting of an input layer, an LSTM hidden layer, and a fully connected output layer; Taking N time steps of historical electricity price time series data as input of the input layer; Serializing N historical electricity price data points through the LSTM hidden layer, updating its internal state at each time step to extract feature information encoding the time dependency of the historical electricity price data; The fully connected output layer generates a time series of electricity price prediction values ​​including M time steps based on the extracted feature information, which is represented as an electricity price prediction sequence.

3. The method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system according to claim 2, characterized in that: The state vector is constructed as follows: Obtain the current state of charge SOC(t) and health state SOH(t) of the energy storage system as internal state data; Obtaining the electricity price prediction sequence including M electricity price prediction values ​​from the current time t+1 to t+M; The internal state data and the electricity price prediction sequence are sequentially spliced ​​in a preset order to form the state vector S(t).

4. The method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system according to claim 3, characterized in that: The state vector S(t) is calculated as follows: Where S(t) represents the complete state vector at time t; SOC(t) represents the battery state of charge at time t, a scalar between 0 and 1; SOH(t) represents the battery state of health at time t, a scalar between 0 and 1; Represents the predicted The grid electricity price at the moment; M is the total number of time steps predicted by the long short-term memory network model.

5. The method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system according to claim 4, characterized in that: The deep reinforcement learning agent, whose specific carrier is the deep Q network, is constructed as follows: Construct a multilayer perceptron network structure consisting of an input layer, at least two fully connected hidden layers, and an output layer; Setting the number of neurons in the input layer to match the dimension of the state vector S(t); The number of neurons in the output layer is set to be exactly the same as the total number of actions in the predefined discrete charge and discharge action set.

6. The method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system according to claim 5, characterized in that: The specific training process of the deep reinforcement learning agent is as follows: Initialize the evaluation network and target network, and prepare an experience replay pool for storing interaction data; In the battery simulation environment, each interaction in the battery simulation environment is evaluated using the defined immediate reward function R(t) and generates a data tuple containing state, action, reward, and next state; Storing the data tuple in an experience replay pool; Randomly extract data from the experience replay pool to calculate the target Q value and the current Q value; According to the loss between the target Q value and the current Q value, the weight of the deep Q network is iteratively updated until the maximum number of iterations is reached, thereby completing the training of the deep reinforcement learning agent. The action value function Q of the deep reinforcement learning agent can accurately predict the expected value of the long-term cumulative reward defined in claim 7 that can be brought about by executing each candidate action based on any input state.

7. The method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system according to claim 6, characterized in that: The immediate reward function R(t) defined above is solved as follows: Extracting the predicted electricity prices corresponding to each of the next N time steps based on the state vector S(t); Each of the N predicted electricity prices is associated with the currently selected specific action Multiply the corresponding charge and discharge power values ​​to obtain N single-step expected benefit values; The N single-step expected return values ​​are cumulatively summed, and the result of the cumulative sum is determined as the expected economic return item for the next N steps; The current state of charge SOC(t) and health state SOH(t) are used as inputs, and the health cost coefficient related to the state is output through the preset health cost coefficient function k; Combine the health cost coefficient with the currently selected specific action The corresponding preset action stress factor f is multiplied, and the product is determined as the state-dependent health cost item; Multiplying the expected economic benefit item in the next N steps by a first preset weight α to obtain a first calculation result; Multiplying the state-dependent health cost item by a second preset weight β to obtain a second calculation result; The first calculation result is subtracted from the second calculation result, and the final difference is represented as the immediate reward function R(t).

8. The method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system according to claim 7, characterized in that: The preset health cost coefficient function k is further defined as having the following output characteristics: The region between the first preset SOC threshold and the second preset SOC threshold is defined as a normal SOC operating range; when the input state of charge SOC(t) is outside the normal SOC operating range, the output value of the health cost coefficient function k is higher than the output value when it is within the normal SOC operating range; When the input health state SOH(t) is lower than the preset health level threshold, the output value of the health cost coefficient function k increases nonlinearly as SOH(t) decreases.

9. The method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system according to any one of claims 6 to 8, characterized in that: The specific action selected based on the output from the predefined discrete charge and discharge action set is specifically: The current state vector S(t) is input into the trained deep reinforcement learning agent to perform each preset candidate action. , get the predicted corresponding action value ; Among all the action values ​​obtained, select the candidate action corresponding to the maximum value , and determine it as a specific output action .

10. The method for collaborative optimization of dynamic charge and discharge strategies for extending the life of an energy storage system according to claim 9, characterized in that: The specific action of converting the output into a charge and discharge power control instruction is specifically: According to the specific output action , determine the corresponding nominal charge and discharge rate; Obtaining the preset nominal energy capacity of the energy storage system; Multiplying the nominal charge and discharge rate by the nominal energy capacity to determine the product as a target charge and discharge power instruction; The target charge and discharge power instructions are sent to the energy storage converter of the energy storage system as the power control setting value of the energy storage converter.

Citation Information

Patent Citations

  • Electricity price prediction method and system based on bidirectional long-short-term neural network

    CN110276638A

  • Coordinated charging method and coordinated charging system based on distributed deep reinforcement learning

    CN114619907A

  • Prediction-based energy storage regulation and control method and system

    CN117013580A

  • Equalization method of energy storage battery pack management system based on neural network and medium

    CN117613421A

  • Light storage power station electricity selling optimization method based on predictive research and judgment and dynamic planning

    CN117808623A

Cited By

  • Energy storage operation optimization method and system based on hierarchical multi-agent reinforcement learning

    CN120952277A

  • Energy storage operation optimization method and system based on hierarchical multi-agent reinforcement learning

    CN120952277B

  • Energy storage type charging pile power adaptive control method based on reinforcement learning

    CN121492737A

  • A power self-adaptive control method for energy storage charging pile based on reinforcement learning

    CN121492737B