Mobile energy storage scheduling system and scheduling method based on double-layer deep reinforcement learning, and storage medium

By introducing a double-layer deep reinforcement learning structure into the mobile energy storage system, which is divided into power decision-making layer and mobile decision-making layer, the problem that traditional energy storage system scheduling is difficult to adapt to the rapidly changing power market environment is solved, and efficient and economical energy storage system scheduling is achieved.

CN120073682APending Publication Date: 2025-05-30STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510136170.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional energy storage system scheduling is difficult to adapt to the rapidly changing power market environment and complex power grid operation state. Especially after the introduction of mobile energy storage equipment, the scheduling problem becomes more complicated, and factors such as the geographical location selection of energy storage equipment, mobile costs, and electricity price differences between different regions need to be considered.

Method used

A mobile energy storage scheduling system based on double-layer deep reinforcement learning is adopted, which is divided into power decision-making layer and mobile decision-making layer. The power decision-making layer dynamically adjusts the charge and discharge power of the energy storage system in real time through model-free reinforcement learning method to achieve electricity price arbitrage. The mobile decision-making layer determines whether the energy storage system needs to be moved to other energy storage power stations based on the electricity price difference and the mobile cost, and optimizes the mobile strategy.

Benefits of technology

It realizes intelligent scheduling of energy storage systems in complex power market environments, improves arbitrage benefits, reduces mobile costs and capacity attenuation costs, and extends the service life of energy storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120073682A_ABST
    Figure CN120073682A_ABST
Patent Text Reader

Abstract

The invention provides a mobile energy storage scheduling system and scheduling method based on double-layer deep reinforcement learning and a storage medium, and relates to the field of intelligent power grids and energy management systems.The method comprises a power decision-making layer and a mobile decision-making layer, and the power decision-making layer determines the current electricity price, the charge state of energy storage and residual capacity information according to the current electricity price, the charge state of energy storage and the residual capacity information; the charging and discharging power of the energy storage system is dynamically adjusted in real time, and electricity price arbitrage is achieved; and the mobile decision-making layer determines whether the energy storage system moves to other power stations or not according to the electricity price data of each energy storage power station, the mobile cost and the current state of the energy storage power station so as to further increase the arbitrage income and reduce the mobile cost. According to the mobile energy storage dispatching system, efficient, economical and reliable operation of the mobile energy storage system in a complex power market environment is realized through an intelligent dispatching mode, an advanced algorithm and a comprehensive energy storage system aging model; and the method is of great significance in improving the energy utilization efficiency, reducing the electric power cost and promoting the sustainable development of energy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart grid and energy management systems, and in particular to a mobile energy storage scheduling system, a scheduling method and a storage medium based on double-layer deep reinforcement learning. Background Art

[0002] The scheduling of traditional energy storage systems mainly relies on manually set rules or optimization algorithms based on simple mathematical models. These methods are often difficult to adapt to the rapidly changing power market environment and complex grid operating conditions. Especially after the introduction of mobile energy storage devices, the scheduling problem becomes more complex because in addition to considering the charge and discharge strategies of energy storage, factors such as the geographical location selection of energy storage devices, mobile costs, and electricity price differences between different regions also need to be considered.

[0003] In recent years, deep reinforcement learning (DRL), as a technology that combines the advantages of deep learning and reinforcement learning, has shown great potential in solving complex decision-making problems. By simulating the human learning process, DRL can automatically extract features from a large amount of data and optimize decision-making strategies, and is suitable for dealing with problems in high-dimensional state spaces and continuous action spaces. However, applying DRL to mobile energy storage scheduling still faces many challenges, such as how to design reasonable state spaces, action spaces, and reward functions to accurately reflect the actual operating conditions of energy storage systems, and how to effectively integrate the discrete mobile decisions and continuous power decision actions of energy storage.

[0004] Therefore, there is an urgent need in the market for an efficient intelligent scheduling system that can comprehensively consider the charge and discharge strategies and mobile scheduling of energy storage systems and can adapt to real-time electricity price changes. Based on this demand, the present invention proposes an innovative mobile energy storage scheduling system based on double-layer deep reinforcement learning, aiming to decouple discrete and continuous actions through a hierarchical decision-making architecture, optimize the power decision and mobile decision of energy storage, and thus achieve a globally optimal mobile energy storage scheduling strategy. Summary of the Invention

[0005] The present invention addresses the above-mentioned deficiencies in the prior art and provides a mobile energy storage scheduling system, a scheduling method and a storage medium based on double-layer deep reinforcement learning.

[0006] The present invention relates to a mobile energy storage scheduling system based on double-layer deep reinforcement learning, aiming to optimize the charge and discharge strategies and mobile strategies of energy storage systems through intelligent decision-making to achieve maximum arbitrage revenue, while reducing mobile costs and the capacity attenuation cost of energy storage systems. This system is particularly suitable for various types of energy storage systems, including but not limited to household energy storage, community energy storage, industrial energy storage, and grid-level energy storage systems, and can flexibly adjust strategies according to actual needs and conditions to achieve efficient and economic energy management goals.

[0007] The mobile energy storage scheduling system based on double - layer deep reinforcement learning of the present invention includes a power decision layer and a mobile decision layer, where: The power decision layer is used to adjust the charge - discharge power of the energy storage system in real - time and dynamically according to the electricity price in the current period, the state of charge (SOC) of the energy storage, the remaining capacity of the energy storage, and the normalized charge - discharge power of the energy storage in the previous period. The goal of the power decision layer is to enable the energy storage system to charge during the low - price period of electricity and discharge during the high - price period of electricity through a reasonable charge - discharge strategy, so as to achieve electricity price arbitrage and increase revenue. To achieve this goal, the power decision layer adopts a model - free reinforcement learning method to establish a Markov decision process (MDP).

[0008] The objective function is the difference between the arbitrage revenue and cost of the mobile energy storage:

[0009]

[0010] Among them, R t is the energy storage arbitrage revenue, C t is the energy storage operation cost, and f is the objective function.

[0011] The state space includes the site CS where the current period is located, the current period number t, the SOC of the energy storage, the remaining capacity C of the energy storage, and the real - time electricity price p of the current site:

[0012] state = {CS, t, SOC, C, p};

[0013] The action space is the normalized energy storage discharge power P dc , and the actual charge - discharge power of the energy storage system is adjusted through this action space:

[0014] action = {P dc};

[0015] The reward function includes the energy storage arbitrage revenue Y arb , the capacity attenuation cost Y cost :

[0016] reward = Y arb - Y cost ;

[0017] For the calculation of the state transition function, the state space depends on the current charging station, the current period, and the influence of the agent's action execution on the SOC of the mobile energy storage battery:

[0018]

[0019] Among them, η ch and η dis are the charging and discharging efficiencies of the energy storage respectively, Δt is the charge - discharge power decision time interval, Pch P is the charging power for energy storage, and C is the capacity of the energy storage system;

[0020] The specific calculation method of the above reward function is as follows:

[0021] The arbitrage profit is calculated based on the electricity price in the current period and the charge-discharge power of the energy storage system:

[0022]

[0023] where P dc is the discharge power, and η is the charge or discharge efficiency;

[0024] The capacity attenuation cost is calculated based on the aging model of the energy storage system and the charge-discharge power in the current period:

[0025] Y cost = α cap (C ini - C);

[0026] where α cap is the capacity attenuation cost coefficient, and C ini is the initial capacity of the energy storage system;

[0027] The mobile decision-making layer is used to determine whether the energy storage system needs to be moved to other energy storage power stations according to the electricity price data of each energy storage power station, the moving cost between energy storage power stations, and the current state of the energy storage power station. The goal of the mobile decision-making layer is to enable the energy storage system to move to an energy storage power station with a higher electricity price or a lower moving cost through a reasonable moving strategy, so as to further increase the arbitrage profit and reduce the moving cost. To achieve this goal, the mobile decision-making layer also adopts the reinforcement learning method. Its state space includes the site CS where the current period is located, the current period number t, the SOC of the energy storage, the remaining capacity C of the energy storage, the real-time electricity prices p arr of all current sites, and the time T arr required from the current site to all other sites. The action space is the energy storage power station CS next to be moved to in the next period. The reward function is the arbitrage profit Y arb minus the capacity attenuation cost Y cost and the moving cost Y T :

[0028] state = {CS, t, SOC, C, p arr , T arr}

[0029] action = {CS next}

[0030] reward = Y arb - Y cost - Y T

[0031] Furthermore, the present invention also provides a mobile energy storage scheduling method based on double-layer deep reinforcement learning:

[0032] First, a day is divided into K time periods. Within each time period, first, the mobile decision-making layer determines whether the mobile energy storage system is needed. If the decision result of the mobile decision-making layer is to move, then this time period is used for the movement of the energy storage system without performing charge and discharge operations. If the mobile decision-making layer decides not to move, then this time period is for the power decision-making layer to formulate a charge and discharge strategy based on the current state and historical electricity price prediction to obtain electricity price arbitrage benefits. At the end of each time period, the state of the energy storage system is updated, including SOC, remaining capacity, and location information. Then, the above steps are repeated until the end of the day, and the total revenue is calculated and accumulated.

[0033] To further improve the economic benefits of the energy storage system, the present invention also introduces the deep reinforcement learning TD3 algorithm and the energy storage system aging model.

[0034] The TD3 algorithm has an electricity price prediction function and can make reasonable predictions about future short-term or long-term electricity prices by analyzing information such as historical electricity price data and electricity market behaviors. These prediction information can provide a decision-making basis for the power decision-making layer and the mobile decision-making layer, enabling them to formulate charge and discharge strategies and mobile strategies more accurately, thereby further improving the arbitrage benefits.

[0035] The energy storage system aging model is used to predict the capacity attenuation of the energy storage system according to conditions such as the equivalent charge and discharge cycle times and temperature of the energy storage system. This model can provide real-time information on the health status of the energy storage system for the power decision-making layer and the mobile decision-making layer, enabling them to consider the capacity attenuation cost of the energy storage system when formulating strategies, thereby avoiding excessive capacity attenuation caused by overusing the energy storage system.

[0036] In specific implementation, the power decision-making layer is trained using the TD3 algorithm. The TD3 algorithm is an improved algorithm based on deep deterministic policy gradients. It solves the problem of overestimation of the Q-value function in the original algorithm by introducing methods such as double networks and delayed updates, making the training process more stable. During the training process, the power decision-making layer will continuously learn and optimize the charge and discharge strategy to adapt to changes in different electricity prices and energy storage system states. The training process includes the following key points:

[0037] (1) Double network: Two sets of Critic networks are adopted, and the smaller value of the two is taken when calculating the target value, thereby solving the problem of overestimation of the Q-value function. The calculation method of the target value y is:

[0038]

[0039] Where For the action of the next state, add perturbations for smooth regularization. is the Q-value estimation of the i-th target Critic network for state s' and action .

[0040] (2) Target network smooth regularization: When calculating the target value, add truncated noise to the action of the next state to make the value evaluation more accurate. The formula for action smoothing is:

[0041]

[0042] where:

[0043] π φ' (s') is the action output by the target Actor network in state s'; ∈ is the added perturbation term, which is used to increase the robustness and exploration of the model. Here, ∈ is sampled from this normal distribution.

[0044] (3) Delayed update: After the Critic network is updated multiple times, then update the Actor network to ensure the more stable training of the Actor network.

[0045] (4) Network parameter update: Use the gradient descent method to update the parameters of the Critic network to minimize the error between the predicted value and the target value. Subsequently, update the parameters of the Actor network at a lower frequency through the evaluation results of the Critic network, thereby optimizing the policy.

[0046]

[0047] where, θ is the parameter of the Critic network, α is the learning rate, Q θ (s,a) is the predicted value of the Critic network, and y is the target value. Update the parameters of the Critic network by the gradient descent method to minimize the mean square error between the predicted value and the target value.

[0048] The mobile decision-making layer is trained using the Rainbow algorithm. The Rainbow algorithm is an important variant in deep reinforcement learning, which improves the performance of DQN (Deep Q-Network) by integrating multiple methods. These methods include double network structure and delayed update, prioritized experience replay, noisy network, and reward clipping, etc. During the training process, the mobile decision-making layer will continuously learn and optimize the mobile strategy to adapt to the changes in different electricity prices, mobile costs, and energy storage system states. The training process includes the following key points:

[0049] (1) Double Network Structure and Delayed Update: The Rainbow algorithm adopts a double network structure, that is, a current network is used to select actions, and a target network is used to calculate the target Q-value. The target network is updated more slowly than the current network, usually updated every fixed number of steps. The calculation formula for the target Q-value is:

[0050] y = r + γmax a′ Q′(s′, a′; θ′)

[0051] where r is the reward and γ is the discount factor.

[0052] (2) Prioritized Experience Replay: Sample experiences according to the importance of the samples (i.e., the absolute value of the TD error). The importance sampling weight p i is usually proportional to the TD error |δ i |, that is:

[0053]

[0054] where β is a tuning parameter used to control the degree of priority.

[0055] (3) Comprehensive Application of Noisy Networks and Reward Clipping Methods. A noisy network is a method that introduces learnable noise into the neural network parameters to improve the traditional exploration strategy. By adding noise to the weights of the model, the exploration behavior of the agent becomes more diverse and flexible, and the adjustment process of the noise can be learned through backpropagation.

[0056] Given the standard linear layer of a neural network, its input and output variables are x and y respectively, and the weights and biases are W and b respectively. Its output is expressed as:

[0057] y = Wx + b

[0058] In a noisy network, the weights and biases introduce learnable noise terms:

[0059] W = W μ + W σ ⊙ ∈ w

[0060] b = b μ + b σ ⊙ ∈ b

[0061] where W μ and b μ represent the fixed base part (deterministic parameters), W σ and b σ represent the scale of the noise (learnable parameters), ∈ w and ∈ bDenotes random noise, usually sampled from a normal distribution. During forward propagation, noise is randomly sampled first to obtain noisy weights and biases, and forward propagation is completed using the noisy weights and biases. During backpropagation, the gradients are used to update the deterministic and learnable parameters, enabling the noise to be optimized during training.

[0062] Reward clipping is used to prevent the unstable impact on the training process caused by an overly large range of reward values. By applying a fixed clipping operation to the reward values, they are restricted within a certain range. In the present invention, the range of reward clipping is [-1, 1]. If the reward received by the agent at time step t is r t , then the clipped reward is:

[0063] Compared with using the original reward r t , the clipped reward ensures the stability of the update step size, thereby accelerating convergence.

[0064]

[0065] Furthermore, the present invention also provides a storage medium storing computer instructions for causing a processor to execute the mobile energy storage scheduling system based on double-layer deep reinforcement learning.

[0066] Due to the adoption of the above technical solutions, the present invention has at least one of the following beneficial effects:

[0067] The mobile energy storage scheduling system of the present invention realizes the intelligent scheduling of the energy storage system in a complex and changeable electricity market environment by introducing a double-layer deep reinforcement learning structure, namely a power decision-making layer and a mobile decision-making layer. This structure enables the energy storage system to not only formulate efficient charge and discharge strategies based on current electricity price data to achieve electricity price arbitrage, but also dynamically adjust the mobile strategy according to electricity price differences and mobile costs, thereby further improving the arbitrage income and reducing the overall operating cost. This intelligent scheduling method has higher flexibility and adaptability compared with traditional energy storage scheduling methods based on fixed rules or simple optimization algorithms.

[0068] The present invention introduces advanced algorithms in deep reinforcement learning, including the TD3 algorithm and the Rainbow algorithm, for training the power decision-making layer and the mobile decision-making layer. These algorithms enable the agent to continuously learn and optimize strategies during the interaction with the environment, thereby achieving precise control of the charge and discharge and mobile strategies of the energy storage system. This not only improves the arbitrage income of the energy storage system, but also reduces the capacity attenuation cost caused by overuse, and extends the service life of the energy storage system.

[0069] The present invention fully considers the aging problem of the energy storage system. By introducing an energy storage system aging model, it can predict the capacity attenuation of the energy storage system in real time, enabling the power decision-making layer and the mobile decision-making layer to more comprehensively consider the health status of the energy storage system when formulating strategies, avoiding the problem of excessive capacity attenuation caused by overuse, and effectively improving the economy and reliability of the mobile energy storage system.

[0070] The mobile energy storage scheduling system based on double-layer deep reinforcement learning proposed by the present invention realizes the efficient, economic, and reliable operation of the mobile energy storage system in a complex power market environment through an intelligent scheduling method, advanced algorithms, and a comprehensive energy storage system aging model. Brief Description of the Drawings

[0071] Figure 1 It is the algorithm flow chart of the rain flow counting method for calculating the equivalent cycle number of the energy storage in the present invention; Detailed Embodiment

[0072] The following is a detailed description of the embodiments of the present invention: These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

[0073] The present invention proposes an innovative mobile energy storage scheduling system, which ingeniously integrates a double-layer deep reinforcement learning mechanism and uses the deep reinforcement learning algorithm to solve the mobile energy storage optimal scheduling problem that is difficult to solve by traditional methods. The system is divided into a power decision-making layer and a mobile decision-making layer, which use the TD3 algorithm and the Dueling DQN algorithm for intelligent decision-making respectively. The power decision-making layer can accurately control the charging and discharging power of the energy storage device based on information such as real-time electricity price, state of charge of the energy storage system, remaining capacity, and historical charging and discharging power to maximize the arbitrage profit. The mobile decision-making layer, based on the electricity price difference, mobile cost, and current energy storage state of each energy storage power station, intelligently determines the mobile path of the energy storage device to further reduce costs. Through the integrated electricity price prediction module, energy storage system aging model, load response module, and energy storage system state update module, the system can comprehensively consider market dynamics, energy storage health status, and power system requirements, thereby formulating more efficient scheduling strategies. Experimental results show that compared with traditional scheduling methods, the system not only significantly improves the arbitrage profit of the energy storage system but also effectively reduces the mobile cost, demonstrating strong robustness and adaptability. In the future, the system is expected to be further extended to more types of energy storage scheduling, and by optimizing algorithm parameters and network structures, as well as introducing real-time data feedback and online learning mechanisms, more intelligent and flexible energy storage scheduling can be achieved.

[0074] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be described clearly and completely below. Apparently, the described embodiments are some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention.

[0075] The present invention proposes a mobile energy storage scheduling system based on double-layer deep reinforcement learning, aiming to achieve efficient charging, discharging, and mobile scheduling of in-vehicle mobile energy storage devices in the actual power market environment through intelligent algorithms. The system is divided into a power decision-making layer and a mobile decision-making layer, and the TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm and the Dueling DQN (Dueling Deep Q-Network) algorithm are respectively used for optimization decisions. The specific implementation manners of the present invention will be described in detail below.

[0076] The mobile energy storage scheduling system of the present invention consists of a power decision-making layer, a mobile decision-making layer, energy storage system aging, and an energy storage system state update module. Among them:

[0077] The power decision-making layer is used to adjust the charging and discharging power of the energy storage system in real time and dynamically according to the electricity price at the current time period, the state of charge (SOC) of the energy storage, the remaining capacity of the energy storage, and the normalized charging and discharging power of the energy storage in the previous time period. The goal of the power decision-making layer is to enable the energy storage system to charge during the low electricity price period and discharge during the high electricity price period through a reasonable charging and discharging strategy, so as to achieve electricity price arbitrage and increase revenue. To achieve this goal, the power decision-making layer adopts a model-free reinforcement learning method to establish a Markov decision process (MDP). Among them:

[0078] The objective function is the difference between the arbitrage revenue and cost of the mobile energy storage:

[0079]

[0080] Among them, R t is the energy storage arbitrage revenue, C t is the energy storage operation cost, and f is the objective function.

[0081] The state space includes the current site CS, the current time period number t, the SOC of the energy storage, the remaining capacity C of the energy storage, and the real-time electricity price p of the current site:

[0082] state = {CS, t, SOC, C, p}

[0083] The action space is the normalized energy storage discharge power P dc, adjust the actual charging and discharging power of the energy storage system through this action space:

[0084] action = {P dc}

[0085] The reward function includes the energy storage arbitrage revenue Y arb , the capacity attenuation cost Y cost :

[0086] reward = Y arb -Y cost

[0087] Calculation of the state transition function. The state space depends on the current charging station, the time period, and the influence of the agent's actions on the SOC of the mobile energy storage battery:

[0088]

[0089] Among them, η ch and η dis are the charging and discharging efficiencies of the energy storage respectively, Δt is the decision time interval of the charging and discharging power, P ch is the charging power of the energy storage, and C is the capacity of the energy storage system.

[0090] The specific calculation method of the above reward function is as follows:

[0091] The arbitrage revenue is calculated according to the electricity price in the current time period and the charging and discharging power of the energy storage system:

[0092]

[0093] Among them, P dc is the discharging power, and η is the charging or discharging efficiency;

[0094] The capacity attenuation cost is calculated according to the aging model of the energy storage system and the charging and discharging power in the current time period:

[0095] Y cost = α cap (C ini -C)

[0096] Among them, α cap is the capacity attenuation cost coefficient, and C ini is the initial capacity of the energy storage system.

[0097] The mobile decision-making layer is used to determine whether the energy storage system needs to be moved to other energy storage power stations according to the electricity price data of each energy storage power station, the moving cost between energy storage power stations, and the current state of the energy storage power station. The goal of the mobile decision-making layer is to enable the energy storage system to move to an energy storage power station with a higher electricity price or a lower moving cost through a reasonable moving strategy, so as to further increase the arbitrage profit and reduce the moving cost. To achieve this goal, the mobile decision-making layer also adopts the reinforcement learning method. Its state space includes the current site CS, the current time period number t, the SOC of the energy storage, the remaining capacity C of the energy storage, and the real-time electricity prices p of all current sites arr , and the time T required from the current site to all other sites arr . The action space is the energy storage power station CS to be moved to in the next time period next . The reward function is the arbitrage profit Y arb minus the capacity attenuation cost Y cost and the moving cost Y T :

[0098] state = {CS, t, SOC, C, p arr , T arr}

[0099] action = {CS next}

[0100] reward = Y arb -Y cost -Y T

[0101] A mobile energy storage scheduling method based on double-layer deep reinforcement learning of the present invention operates through the following steps:

[0102] First, a day is divided into K time periods. Within each time period, first, the mobile decision-making layer determines whether the energy storage system needs to be moved. If the decision result of the mobile decision-making layer is to move, then this time period is used for the movement of the energy storage system, and no charge and discharge operations are performed. If the mobile decision-making layer decides not to move, then this time period is used by the power decision-making layer to formulate a charge and discharge strategy based on the current state and historical electricity price prediction to obtain the electricity price arbitrage profit. At the end of each time period, the state of the energy storage system is updated, including the SOC, the remaining capacity, and the location information. Then, the above steps are repeated until the end of the day, and the total profit is calculated and accumulated

[0103] To further improve the economic benefit of the energy storage system, the present invention also introduces the deep reinforcement learning TD3 algorithm and the energy storage system aging model

[0104] The TD3 algorithm has the function of electricity price prediction. It can make reasonable predictions about future short-term or long-term electricity prices by analyzing information such as historical electricity price data and electricity market behaviors. These prediction information can provide a decision-making basis for the power decision-making layer and the mobile decision-making layer, enabling them to formulate charging and discharging strategies and mobile strategies more accurately, thereby further improving the arbitrage revenue.

[0105] The energy storage system aging model is used to predict the capacity attenuation of the energy storage system according to conditions such as the equivalent charge and discharge cycle times and temperature of the energy storage system. This model can provide real-time information on the health status of the energy storage system for the power decision-making layer and the mobile decision-making layer, enabling them to consider the capacity attenuation cost of the energy storage system when formulating strategies, thereby avoiding excessive capacity attenuation caused by overusing the energy storage system.

[0106] In specific implementation, the power decision-making layer uses the TD3 algorithm for training. The TD3 algorithm is an improved algorithm based on deep deterministic policy gradient. It solves the problem of overestimation of the Q-value function in the original algorithm by introducing methods such as double networks and delayed updates, making the training process more stable. During the training process, the power decision-making layer will continuously learn and optimize the charging and discharging strategies to adapt to changes in different electricity prices and energy storage system states. The training process includes the following key points:

[0107] (1) Double network: Two sets of Critic networks are adopted, and the smaller value of the two is taken when calculating the target value, thus solving the problem of overestimation of the Q-value function. The calculation method of the target value y is:

[0108]

[0109] where is the action in the next state, adding perturbations for smooth regularization, is the Q-value estimate of the i-th target Critic network for the state s' and the action ;

[0110] (2) Target network smooth regularization: When calculating the target value, truncated noise is added to the action in the next state to make the value evaluation more accurate. The formula for action smoothing is:

[0111]

[0112] where:

[0113] π φ' (s') is the action output by the target Actor network in the state s'; ∈ is the added perturbation term, which is used to increase the robustness and exploratory ability of the model. Here, ∈ is sampled from this normal distribution.

[0114] (3) Delayed update: After the Critic network is updated multiple times, the Actor network is updated to ensure more stable training of the Actor network.

[0115] (4) Network parameter update: Use the gradient descent method to update the parameters of the Critic network to minimize the error between the predicted value and the target value. Subsequently, update the parameters of the Actor network at a lower frequency based on the evaluation results of the Critic network to optimize the policy.

[0116]

[0117] Among them, θ is the parameter of the Critic network, α is the learning rate, Q θ (s,a) is the predicted value of the Critic network, and y is the target value. Update the parameters of the Critic network by the gradient descent method to minimize the mean square error between the predicted value and the target value.

[0118] The mobile decision-making layer is trained using the Rainbow algorithm. The Rainbow algorithm is an important variant in deep reinforcement learning. It improves the performance of DQN (Deep Q-Network) by integrating multiple methods, including double network structure and delayed update, prioritized experience replay, noisy network, and reward clipping. During the training process, the mobile decision-making layer continuously learns and optimizes the mobile strategy to adapt to changes in different electricity prices, mobile costs, and energy storage system states. The training process includes the following key points:

[0119] (1) Double network structure and delayed update: The Rainbow algorithm adopts a double network structure, namely a current network for selecting actions and a target network for calculating the target Q value. The target network is updated more slowly than the current network and is usually updated every fixed number of steps. The formula for calculating the target Q value is:

[0120] y = r + γmax a′ Q′(s′,a′;θ′)

[0121] Among them, r is the reward and γ is the discount factor.

[0122] (2) Prioritized experience replay: Sample experiences according to the importance of the samples (i.e., the absolute value of the TD error). The importance sampling weight p i is usually proportional to the TD error |δ i |, that is:

[0123]

[0124] Among them, β is the adjustment parameter used to control the degree of priority.

[0125] (3)Comprehensively apply the noise network and reward clipping method. The noise network is a method that introduces learnable noise into the neural network parameters to improve the traditional exploration strategy. By adding noise to the weights of the model, the exploration behavior of the agent becomes more diverse and flexible. At the same time, the adjustment process of the noise can be learned through backpropagation.

[0126] For a standard linear layer of a neural network, its input and output variables are x and y respectively, and the weights and biases are W and b respectively. Its output is expressed as:

[0127] y = Wx + b

[0128] In the noise network, learnable noise terms are introduced into the weights and biases:

[0129] W = W μ + W σ ⊙ ∈ w

[0130] b = b μ + b σ ⊙ ∈ b

[0131] Among them, W μ and b μ represent the fixed basic part (deterministic parameters), and W σ and b σ represent the scale of the noise (learnable parameters), ∈ w and b represent random noise, usually sampled from a normal distribution. During forward propagation, the noise is randomly sampled first to obtain the noisy weights and biases, and the forward propagation is completed using the noisy weights and biases. During backpropagation, the gradients are updated for the deterministic parameters and learnable parameters, so that the noise is optimized during training.

[0132] Reward clipping is used to prevent the unstable impact on the training process caused by the excessive range of reward values. By applying a fixed clipping operation to the reward value, it is restricted within a certain range. In the present invention, the range of reward clipping is [-1, 1]. If the reward received by the agent at time step t is r t , then the clipped reward is:

[0133]

[0134] Compared with using the original reward r t , the clipped reward ensures the stability of the update step size, thus accelerating convergence.

[0135] Further, the present invention also provides a storage medium storing computer instructions for causing a processor to execute the mobile energy storage scheduling system based on double-layer deep reinforcement learning.

[0136] The present invention conducts offline training based on time-of-use day-ahead electricity price data and traffic information. After the training is completed, the offline model conducts online testing according to real-time electricity price information, energy storage state information, and traffic conditions. The mobile energy storage scheduling system based on double-layer deep reinforcement learning of the present invention provides a solution for the continuous and discrete joint optimization problem that is difficult to solve by traditional mobile energy storage scheduling methods through technical means such as intelligent scheduling methods, advanced algorithms, and comprehensive energy storage system aging models, and realizes the efficient, economical, and reliable operation of the mobile energy storage system in a complex power market environment.

Claims

1. A mobile energy storage dispatching system based on double-layer deep reinforcement learning, characterized in that: include: The power decision layer is used to dynamically adjust the charging and discharging power of the energy storage system in real time according to the electricity price of the current period, the state of charge of the energy storage, the remaining capacity of the energy storage, and the standardized charging and discharging power of the energy storage in the previous period; The mobile decision layer is used to decide whether the energy storage system needs to be moved to other energy storage stations based on the electricity price data of each energy storage station, the mobility cost between energy storage stations, and the current status of the energy storage station.

2. The mobile energy storage dispatching system based on double-layer deep reinforcement learning according to claim 1 is characterized in that: The power decision layer adopts a model-free reinforcement learning method to establish a Markov decision process, where: The objective function is the difference between the arbitrage benefit and cost of mobile energy storage: Among them, R t is the energy storage arbitrage income, C t is the energy storage operation cost, f is the objective function; The state space includes the current time period's site CS, the current time period's sequence number t, the energy storage's state of charge SOC, the energy storage's remaining capacity C, and the current site's real-time electricity price p: state = {CS, t, SOC, C, p}; The action space is the standardized energy storage discharge power P dc , the actual charging and discharging power of the energy storage system is adjusted through this action space: action={P dc }; The reward function includes the energy storage arbitrage benefit Y arb , capacity degradation cost Y cost : reward=Y arb -AND cost ; The calculation of the state transfer function, the state space depends on the current charging station, the time period and the state of charge SOC of the mobile energy storage battery affected by the action execution of the intelligent agent: Among them, η ch and η dis are the charging and discharging efficiencies of energy storage, Δt is the charging and discharging power decision time interval, P ch is the energy storage charging power, and C is the capacity of the energy storage system.

3. The mobile energy storage dispatching system based on double-layer deep reinforcement learning according to claim 2 is characterized in that: The specific calculation method of the reward function is as follows: The arbitrage profit is calculated based on the electricity price in the current period and the charging and discharging power of the energy storage system: Where P dc is the discharge power, η is the charging or discharging efficiency; The capacity decay cost is calculated based on the aging model of the energy storage system and the charging and discharging power in the current period: Y cost =a cap (C ini -C); where α cap is the capacity decay cost coefficient, C ini is the initial capacity of the energy storage system.

4. The mobile energy storage dispatching system based on double-layer deep reinforcement learning according to claim 3 is characterized in that: The mobile decision layer adopts the reinforcement learning method, and its state space includes the site CS where the current time period is located, the sequence number of the current time period t, the state of charge SOC of the energy storage, the remaining capacity C of the energy storage, and the real-time electricity price p of all current sites. arr , the time required from the current site to all other sites T arr , the action space is the energy storage station CS to be moved to in the next period next , the reward function is the arbitrage profit Y arb Subtract capacity degradation cost Y cost and the moving cost Y T : state={CS,t,SOC,C,p arr ,T arr } action={CS next } reward=Y arb -AND cost -AND T 。 5. The mobile energy storage dispatching system based on double-layer deep reinforcement learning according to claim 4 is characterized in that: It also includes an energy storage system aging model, which is used to predict the capacity attenuation of the energy storage system based on the number of equivalent charge and discharge cycles and temperature conditions of the energy storage system, and provide real-time information on the health status of the energy storage system to the power decision-making layer and the mobile decision-making layer.

6. The mobile energy storage dispatching system based on double-layer deep reinforcement learning according to claim 5 is characterized in that: The energy storage charge and discharge depth and equivalent cycle number within a specific period of time are calculated based on the rain flow counting method.

7. The mobile energy storage dispatching system based on double-layer deep reinforcement learning according to claim 6 is characterized in that: The power decision layer is trained using the TD3 algorithm, and the training process includes the following steps: (1) Dual network: Two sets of critic networks are used. The smaller value of the two is taken when calculating the target value to solve the problem of overestimation of the Q value function. The target value y is calculated as follows: in For the action of the next state, add disturbance for smooth regularization, is the i-th target Critic network for state s' and action Q value estimation of ; (2) Target network smoothing regularization: When calculating the target value, truncation noise is added to the action of the next state. The formula for action smoothing is: in: π φ' (s') is the action output by the target Actor network in state s'; ∈ is the added disturbance term; ∈ is sampled from this normal distribution; (3) Delayed update: After the Critic network is updated multiple times, the Actor network is updated; (4) Network parameter update: The parameters of the Critic network are updated using the gradient descent method to minimize the error between the predicted value and the target value. Subsequently, the parameters of the Actor network are updated at a lower frequency using the evaluation results of the Critic network to optimize the strategy: Among them, θ is the parameter of the Critic network, α is the learning rate, and Q θ (s, a) is the predicted value of the Critic network, y is the target value, and the parameters of the Critic network are updated by the gradient descent method to minimize the mean square error between the predicted value and the target value.

8. The mobile energy storage dispatching system based on double-layer deep reinforcement learning according to claim 7 is characterized in that: The mobile decision layer is trained using the Rainbow algorithm, and the training process includes the following steps: (1) Dual network structure and delayed update: The Rainbow algorithm adopts a dual network structure, that is, a current network is used to select actions, and a target network is used to calculate the target Q value. The target network is updated every fixed step. The calculation formula of the target Q value is: y=r+γmax a′ Q′(s′,a′;θ′); Where r is the reward and γ is the discount factor; (2) Prioritized experience playback: Experience is sampled based on the importance of the sample, with importance sampling weight p i and TD error |δ i | is proportional, that is: Among them, β is a tuning parameter used to control the degree of priority; (3) Given a standard linear layer of a neural network, with input and output variables x and y, weights W and bias b, the output is: y=Wx+b In a noisy network, the weights and biases introduce learnable noise terms: W=W μ +W σ ⊙∈ w b=b μ +b σ ⊙ b Among them, W μ and b μ represents the fixed base part, W σ and b σ represents the scale of the noise, ∈ w With ∈ b Represents random noise. During forward propagation, the noise is first randomly sampled to obtain the noisy weights and biases, and the noisy weights and biases are used to complete the forward propagation. During backward propagation, the gradient updates the deterministic parameters and learnable parameters, so that the noise is optimized during training; The range of reward clipping is [-1,1]. If the reward received by the agent at time step t is r t , then the reward after clipping is:

9. A mobile energy storage scheduling method based on double-layer deep reinforcement learning, characterized in that: The steps include: a. Divide a day into K time periods. In each time period, the mobile decision layer first determines whether a mobile energy storage system is needed; b. If the decision result of the movement decision layer is to move, then this period is used for the movement of the energy storage system, and no charging or discharging operations are performed; c. If the mobility decision layer decides not to move, the power decision layer will formulate a charging and discharging strategy based on the current status and historical electricity price forecasts to obtain electricity price arbitrage benefits; d. At the end of each period, update the status of the energy storage system, including SOC, remaining capacity and location information; e. Repeat steps a to d until the end of the day, and calculate and accumulate the total profit.

10. A storage medium, characterized in that: The storage medium stores computer instructions, and the computer instructions are used to enable the processor to execute the mobile energy storage scheduling system based on double-layer deep reinforcement learning as described in any one of claims 1-8.