Energy storage scheduling method and system based on digital twinning and reinforcement learning
By employing a digital twin and reinforcement learning-based energy storage scheduling method, the real-time scheduling and battery health management issues of energy storage systems have been resolved. This enables the autonomous evolution and long-term stable operation of energy storage devices, improving the adaptability of grid scheduling and the accuracy of battery management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2025-11-25
- Publication Date
- 2026-04-28
AI Technical Summary
Existing energy storage system scheduling methods are difficult to adapt to the rapid changes in renewable energy generation and grid load. They have high computational complexity, poor real-time performance, lack battery health status assessment and fault early warning capabilities, and cannot achieve prior verification, post-effect assessment, closed-loop feedback and autonomous evolution, resulting in low long-term operating efficiency.
A storage scheduling method based on digital twins and reinforcement learning is adopted. By constructing an energy storage scheduling optimization model based on load forecasting, power forecasting and battery evaluation results, and combining multi-battery collaborative control and grid twin model, online optimization and autonomous evolution of energy storage scheduling are realized. Parameter fine-tuning is carried out using a deep Q-network model and an online experience replay buffer, and a multi-dimensional performance audit system is constructed for strategy evaluation.
It has achieved real-time adaptability and autonomous evolution capability of energy storage scheduling, improved the long-term operating efficiency of energy storage equipment, and ensured the continuous optimization of grid stability and battery health status.
Smart Images

Figure CN121216558B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of energy storage optimization technology, specifically an energy storage scheduling method and system based on digital twins and reinforcement learning. Background Technology
[0002] With the rapid development of renewable energy, energy storage devices are playing an increasingly important role in power systems. However, the scheduling and operation of energy storage devices face numerous challenges, such as the intermittency and uncertainty of renewable energy generation and the volatility of grid load. Traditional energy storage system scheduling methods mainly rely on rules and conventional optimization algorithms, which have many limitations, specifically:
[0003] Rule-based scheduling methods rely on expert experience and pre-defined rules, making it difficult to adapt to rapid changes in renewable energy generation and grid load.
[0004] Traditional optimization algorithms, such as dynamic programming and genetic algorithms, can optimize the operation of energy storage systems to some extent, but they suffer from problems such as high computational complexity and poor real-time performance.
[0005] In addition, existing energy storage dispatch also has shortcomings in battery health management. Specifically, traditional battery management systems lack accurate assessment of battery health status and early warning capabilities for faults, which cannot effectively prevent faults from occurring and affect the long-term stable operation of energy storage equipment.
[0006] Current research on intelligent scheduling focuses on improving prediction and optimization algorithms. These models are typically deployed in a fixed manner after offline training, and belong to "open-loop" optimization systems. Such systems lack the ability to effectively verify and evaluate their own decision-making effects, and cannot make online adjustments and continuous evolution based on actual operational feedback. When faced with uncertainties such as model mismatch, equipment aging, and environmental changes, their long-term performance is difficult to guarantee.
[0007] In summary, there is an urgent need for a new energy storage scheduling technology to achieve intelligent scheduling with prior verification, post-effect evaluation, closed-loop feedback, and autonomous evolution, thereby improving the long-term operating efficiency of energy storage equipment. Summary of the Invention
[0008] The purpose of this application is to provide an energy storage scheduling method and system based on digital twins and reinforcement learning, so as to solve the technical problems in the prior art that it is difficult to achieve prior verification, post-effect evaluation, closed-loop feedback and autonomous evolution, and the long-term operating efficiency of energy storage equipment is low.
[0009] To achieve the above objectives, this application discloses an energy storage scheduling method based on digital twins and reinforcement learning, applicable to the same power grid and its loads, wherein the loads include electrical loads and energy storage loads. The method includes:
[0010] By using reinforcement learning, an energy storage scheduling optimization model is constructed, which takes load forecasting results, power forecasting results, and battery evaluation results as inputs and battery charging and discharging control commands as outputs.
[0011] Based on the battery charging and discharging control commands, multi-battery collaborative control is performed to achieve energy storage scheduling;
[0012] Digital twins are used to construct a power grid twin model with actual operating data as input and macro-level experience tuples as output. The macro-level experience tuples include at least an evolutionary reward signal, which is used to quantify the power grid performance, power grid stability, and power grid health during energy storage dispatch.
[0013] When the energy storage scheduling optimization model is running, after the preset scheduling period ends, the energy storage scheduling optimization model samples the macroscopic empirical tuples to optimize the energy storage scheduling optimization model.
[0014] Based on the battery charging and discharging control commands output by the optimized energy storage scheduling model, multi-battery collaborative control is performed to achieve energy storage scheduling.
[0015] Preferably, the energy storage scheduling optimization model is provided with an online experience replay buffer, which is used to store the macroscopic experience tuples;
[0016] The energy storage scheduling optimization model samples the macro-experience tuples from the online experience replay buffer based on the scheduling period, and performs online fine-tuning of the network parameters in the energy storage scheduling optimization model based on the sampled macro-experience tuples.
[0017] Preferably, the energy storage scheduling optimization model adopts a deep Q-network model architecture, specifically:
[0018] The intelligent agent of the energy storage scheduling optimization model is the energy storage load; the action space of the energy storage scheduling optimization model includes the battery charging and discharging control commands, which include charging and discharging power control commands and state switching commands; the state space of the energy storage scheduling optimization model includes load data, power data, battery data, and electricity price data.
[0019] The reward function of the energy storage dispatch optimization model is constructed based on grid revenue, grid cost, grid stability index, and battery loss index. Specifically: grid revenue includes the revenue gained from discharging during peak grid electricity price periods; grid cost includes the cost of charging during off-peak grid electricity price periods and the cost of battery charging and discharging losses; the grid stability index is quantified by the reciprocal of the square of the grid frequency deviation, with a higher value for the smaller the frequency deviation; and the battery loss index is obtained based on the relationship between the battery's charge / discharge depth, cycle count, and lifespan decay, with a higher value for the smaller the battery loss.
[0020] The energy storage scheduling optimization model adopts a three-layer neural network structure. The number of neurons in the input layer is equal to the dimension of the state space, the hidden layer consists of neurons with ReLU activation function, and the number of neurons in the output layer is equal to the dimension of the action space.
[0021] Preferably, the evolutionary reward signal is obtained by auditing the results of multi-battery collaborative control to achieve energy storage scheduling. The audit is based on the grid performance, grid stability, and grid health to obtain corresponding performance indicators, stability indicators, and health indicators. The evolutionary reward signal is obtained by weighted summation of the performance indicators, stability indicators, and health indicators.
[0022] Preferably, the performance indicators are obtained based on a performance indicator calculation formula, and the performance indicators are used to quantify the total net revenue within one scheduling cycle; wherein, the performance indicator calculation formula is specifically as follows:
[0023]
[0024] in, For the calculated performance indicators, The total time step corresponding to the scheduling period. For a moment The discharge power of the energy storage device, For a moment The electricity price of energy storage devices, For a moment The charging power of energy storage devices, For a moment The electricity purchase price of energy storage devices This represents the real-time time step during quantization.
[0025] Preferably, the stability index is obtained based on the stability index calculation formula, and the stability index is used to quantify the overall effect of the energy storage load in smoothing power fluctuations in the power grid within a scheduling cycle; wherein, the stability index calculation formula is specifically as follows:
[0026]
[0027] in, The calculated stability index, The total time step corresponding to the scheduling period. For a moment The load power, For a moment Power generation capacity, For a moment The power of energy storage devices.
[0028] Preferably, the health index is obtained based on a health index calculation formula, and the health index is used to quantify the average relative current stress of the battery pack of the energy storage device within a scheduling cycle to assess its cumulative loss; wherein, the health index calculation formula is specifically as follows:
[0029]
[0030] in, For the calculated health indicators, The total time step corresponding to the scheduling period. For a moment The absolute value of the actual current. This is the rated current of the battery pack. This represents the real-time time step during quantization.
[0031] Preferably, the macro-experience tuple further includes an initial state, an overall strategy sequence, and an ending state; wherein, the initial state, the overall strategy sequence, and the ending state specifically refer to, during energy storage scheduling, the initial state of the power grid and its loads within a scheduling cycle, the overall strategy sequence adopted for the power grid and its loads within the scheduling cycle, and the ending state of the power grid and its loads within the scheduling cycle.
[0032] Preferably, the method includes the following steps before constructing the energy storage scheduling optimization model:
[0033] Basic data is collected and preprocessed to obtain scheduling data; wherein, the basic data includes at least load data, power data and battery data;
[0034] Based on the scheduling data, load forecasting and power forecasting are performed to obtain the load forecasting results and the power forecasting results.
[0035] A health status assessment is performed based on the battery data to obtain the battery assessment result.
[0036] To achieve the above objectives, this embodiment also discloses an energy storage scheduling system based on digital twins and reinforcement learning, applicable to the energy storage scheduling method based on digital twins and reinforcement learning described above. The system includes:
[0037] The energy storage scheduling optimization model construction module is used to construct an energy storage scheduling optimization model with load forecasting results, power forecasting results, and battery evaluation results as inputs and battery charging and discharging control commands as outputs through reinforcement learning.
[0038] The energy storage scheduling module is used to perform multi-battery collaborative control based on the battery charging and discharging control commands to achieve energy storage scheduling;
[0039] The power grid twin model construction module is used to construct a power grid twin model with actual operating data as input and macro-level experience tuples as output through digital twins. The macro-level experience tuples include at least an evolutionary reward signal, which is used to quantify the power grid performance, power grid stability, and power grid health during energy storage dispatch.
[0040] An energy storage scheduling optimization model optimization module is used to sample the macroeconomic empirical tuples after a preset scheduling period ends when the energy storage scheduling optimization model is running, so as to optimize the energy storage scheduling optimization model.
[0041] The energy storage scheduling optimization module is used to perform multi-battery collaborative control to achieve energy storage scheduling based on the battery charging and discharging control commands output by the optimized energy storage scheduling optimization model.
[0042] Beneficial effects: The energy storage scheduling method and system based on digital twins and reinforcement learning proposed in this application establishes an intelligent closed-loop architecture of "decision-verification-evaluation-learning". It realizes the pre-emptive security verification of the strategy through digital twin technology, conducts quantitative audit of the strategy's aftereffects through a multi-objective performance evaluation system, and finally uses the evaluation results as an evolutionary feedback signal to drive the online optimization of the core decision model. This enables energy storage scheduling to evolve from a statically executed automated tool into an intelligent tool that can continuously learn from actual operation, adapt to changes, and autonomously evolve its performance, fundamentally solving the problems of mismatch and performance degradation in existing energy storage scheduling models. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 A flowchart illustrating the energy storage scheduling method based on digital twins and reinforcement learning provided in this application embodiment;
[0045] Figure 2 A flowchart illustrating the application process of the energy storage scheduling method based on digital twins and reinforcement learning provided in this embodiment of the application.
[0046] Figure 3 This is a structural block diagram of an energy storage scheduling system based on digital twins and reinforcement learning, provided in an embodiment of this application.
[0047] The implementation, functional features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0048] The technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0049] In this document, the term "comprising" is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0050] In the management of energy storage devices, the technical and economic benefits are crucial management indicators, with technical benefits serving economic benefits. Regarding technical benefits, energy storage scheduling and battery health status management are important entry points. Regarding economic benefits, the electricity sales price and purchase price of energy storage devices are important entry points. Based on the above technical approach, to achieve intelligent scheduling with pre-verification, post-effect evaluation, closed-loop feedback, and autonomous evolution, thereby improving the long-term operational efficiency of energy storage devices, this embodiment discloses an energy storage scheduling method based on digital twins and reinforcement learning.
[0051] Reference Figure 1 , Figure 1 This is a flowchart of the energy storage scheduling method based on digital twins and reinforcement learning in this embodiment.
[0052] like Figure 1 As shown, the energy storage scheduling method based on digital twins and reinforcement learning in this embodiment is applied to the same power grid and its loads, including electricity loads and energy storage loads. The method includes:
[0053] By using reinforcement learning, an energy storage scheduling optimization model is constructed, which takes load forecasting results, power forecasting results, and battery evaluation results as inputs and battery charging and discharging control commands as outputs.
[0054] Based on battery charging and discharging control commands, multi-battery collaborative control is performed to achieve energy storage scheduling;
[0055] By using digital twins, a power grid twin model is constructed with actual operating data as input and macro-level experience tuples as output. The macro-level experience tuples include at least an evolutionary reward signal, which is used to quantify the power grid performance, power grid stability, and power grid health during energy storage dispatch.
[0056] When the energy storage scheduling optimization model is running, after the preset scheduling period ends, the energy storage scheduling optimization model samples macroscopic empirical tuples to optimize the energy storage scheduling optimization model.
[0057] Based on the battery charging and discharging control commands output by the optimized energy storage scheduling model, multi-battery collaborative control is performed to achieve energy storage scheduling.
[0058] Reference Figure 2 , Figure 2 This is a flowchart illustrating the application process of the energy storage scheduling method based on digital twins and reinforcement learning in this embodiment.
[0059] like Figure 2 As shown, in a specific application, the energy storage scheduling method based on digital twins and reinforcement learning in this embodiment includes the following steps:
[0060] Step 1: Data Collection and Preprocessing;
[0061] Step 2: Forecasting grid load and renewable energy generation capacity;
[0062] Step 3: Health status assessment of energy storage batteries;
[0063] Step 4: Energy management optimization for energy storage devices based on deep reinforcement learning;
[0064] Step 5: Collaborative control of multiple types of energy storage batteries;
[0065] Step Six: Performance evaluation of scheduling strategy and closed-loop optimization of system.
[0066] Specifically, before constructing the energy storage scheduling optimization model, this method includes:
[0067] Collect and preprocess basic data to obtain scheduling data; the basic data includes at least load data, power data, and battery data.
[0068] Load and power forecasts are performed based on scheduling data to obtain load and power forecast results.
[0069] A health status assessment is performed based on battery data to obtain battery assessment results.
[0070] In the specific application of this embodiment, step one, data collection and preprocessing, specifically includes data collection, data cleaning, and feature engineering.
[0071] Regarding data collection, we mainly collect historical load data of the power grid, power generation data of new energy sources, operation data of energy storage batteries, and auxiliary data.
[0072] Historical load data of the power grid is collected at the minute level for more than two years, including load curves under different scenarios such as weekdays, weekends and holidays, and records the maximum load, minimum load and typical daily load characteristics.
[0073] The new energy power generation data collection system collects second-level output data from photovoltaic power plants and wind farms, including power generation characteristic curves under different weather conditions, and also records supporting meteorological monitoring data, including irradiance, wind speed, and ambient temperature.
[0074] The energy storage battery operation data collection system collects real-time operating parameters for each battery, including voltage, current, temperature, state of charge (SOC), state of health (SOH), and number of charge-discharge cycles, with a sampling frequency of 1 second.
[0075] The auxiliary data collection includes real-time electricity price information of the power grid, such as time-of-use pricing, peak pricing, equipment maintenance records, and fault alarm logs.
[0076] For data cleaning, we mainly focus on anomaly handling, data smoothing, and data alignment.
[0077] Anomaly processing employs a sliding window statistical method to identify abnormal data. Data corresponding to anomalies such as voltage surges (rate of change >20% / s) and current anomalies (mismatch with SOC changes) are marked and corrected. For photovoltaic data, irradiance is forcibly zeroed at night when it is 0, and missing values are supplemented during the day using cubic spline interpolation. Wind power data is validated and filled using wind speed-power characteristic curves.
[0078] For noisy data such as temperature and voltage, the Kalman filter algorithm is used to smooth the data, preserving the true trend of change while eliminating random interference.
[0079] Data alignment is based on the different sampling frequencies of different devices. In this embodiment, the power grid data is 1 minute and the battery data is 1 second. A linear interpolation method is used to unify all data to a 1-minute time granularity.
[0080] Feature engineering mainly involves processing basic features, temporal features, periodic features, and interactive features.
[0081] In basic feature processing, statistical features such as mean, variance, maximum and minimum values are extracted directly from the original data.
[0082] When processing time-series features, sliding window statistics are calculated, including moving average, standard deviation, rate of change, etc. In this embodiment, the time windows selected are 5-minute windows, 15-minute windows, and 1-hour windows.
[0083] In the periodic feature processing, the daily and weekly periodic components are extracted by Fourier transform and encoded into sin / cos form for input into the model.
[0084] During interactive feature processing, correlation coefficients between different variables are calculated, such as the relationship between temperature and internal resistance, and the relationship between SOC and charge / discharge efficiency.
[0085] In the specific application of this embodiment, the prediction of grid load and new energy power generation in step one specifically includes the construction of a deep long short-term memory (LSTM) model, model training, real-time prediction, and post-prediction processing.
[0086] When building the LSTM model, make the following settings for the model.
[0087] Input layer: The number of neurons is equal to the number of extracted features, such as the mean, variance, trend term, and seasonality of historical grid loads, and the moving average of solar irradiance and wind speed variation rate in photovoltaic power prediction. In a simple example, if 20 features are extracted, the number of neurons in the input layer is 20.
[0088] Hidden layers: Two LSTM hidden layers are set up, each containing 64 neurons. The LSTM unit structure adopts the standard forget gate, input gate, and output gate structure, with the sigmoid activation function and the tanh activation function for cell state.
[0089] The specific mathematical model for internal computation within the LSTM unit and its corresponding technical effects are as follows:
[0090] The forget gate determines which information is discarded from the cell state:
[0091]
[0092] The input gate determines which new information will be stored in the cell state:
[0093]
[0094]
[0095] Cell status update:
[0096]
[0097] The output gate determines what is output:
[0098]
[0099]
[0100] in, This is the current input. It is the previous hidden state. It is the previous cell state. For the sigmoid function, , , and This is the corresponding weight matrix. , , and For the corresponding bias vector, This represents vector concatenation. This indicates element-wise multiplication.
[0101] Output layer: In this embodiment, for the load prediction result of the power grid, the number of output layer neurons is 1, which outputs the power grid load value at one future time point (1 minute later); for the power prediction result of new energy power generation, the number of output layer neurons is 2, which outputs the photovoltaic power value and wind power value at one future time point respectively.
[0102] During model training, this embodiment uses mean-squared error (MSE) as the loss function, and its calculation formula is as follows:
[0103]
[0104] in, For the sample size, For the true value, These are the model's predicted values.
[0105] The Adam optimization algorithm was employed with a learning rate of 0.001, a batch size of 64, and 300 iterations. The training data was divided into a training set and a validation set in an 80%:20% ratio. During training, the model's performance was evaluated using the validation set to prevent overfitting.
[0106] In real-time forecasting, while the power grid is running, real-time collected grid load data and renewable energy generation data are input into a trained deep LSTM model to predict the changing trends of grid load and renewable energy generation over the next 15 minutes to 4 hours. In a specific example, for short-term load forecasting, such as the next hour, data from the past 2 hours is selected as a time window, and the load statistical characteristics and changing trend characteristics within this window are extracted and input into the model for forecasting.
[0107] In the post-processing of predictions, a physical constraint-based verification is first performed to check the physical rationality of the predicted renewable energy output values. For example, photovoltaic output must be zero at night, and wind power output must not exceed the rated capacity. Dynamic correction is then performed; when the real-time monitored prediction error exceeds a preset threshold, a model fine-tuning mechanism is automatically triggered. Furthermore, when the grid twin model in this embodiment detects that the prediction error continuously exceeds the threshold, an online calibration process for the deep LSTM model used for load and power prediction is automatically triggered. This process utilizes the latest operating data to perform incremental training on the deep LSTM model for a small number of epochs, allowing it to quickly adapt to changes in data distribution and maintain prediction accuracy. In incremental learning, an epoch refers to the process of completely traversing the entire training dataset once and updating the model parameters. After each epoch, the model adjusts its parameters based on the training results, gradually optimizing performance.
[0108] In the specific application of this embodiment, step three, the assessment of the health status of the energy storage battery, specifically includes battery health data collection and preprocessing, feature extraction and dimensionality reduction, support vector machine (SVM) model training, and real-time assessment and early warning.
[0109] During battery health data collection and preprocessing, detailed data on the energy storage battery during operation are continuously collected, including parameters such as voltage, current, temperature, SOC, and internal resistance, with a sampling frequency of 1 second. The collected data undergoes preprocessing, including data cleaning, missing value handling, and noise filtering. For occasional abnormal voltage or current data caused by sensor malfunctions, corrections or removals are made by comparing and analyzing adjacent data points, combined with the battery's operating status. Noise generated during data acquisition is smoothed using a low-pass filter.
[0110] In feature extraction and dimensionality reduction, features related to the battery's charge-discharge curves (such as voltage plateau change rate, capacity decay rate, and internal resistance growth rate), temperature statistics (such as maximum temperature, temperature change rate, and temperature uniformity), and State of Charge (SOC) related features (such as SOC fluctuation range and SOC estimation error) are extracted. Existing Principal Component Analysis (PCA) algorithms are used for dimensionality reduction of these features.
[0111] First, the standardized data matrix... ( One sample, Calculate the covariance matrix using (features) :
[0112]
[0113] Then consider the covariance matrix Perform eigenvalue decomposition: Solving for the eigenvalues and the corresponding feature vector Then sort the feature values in descending order and select the top ones. The eigenvectors corresponding to the largest eigenvalues constitute the projection matrix. ;
[0114] The original data is reduced to a lower dimension through linear transformation. dimension:
[0115]
[0116] in This is the new feature matrix after dimensionality reduction. In this embodiment, the original 15 features are reduced to 5 principal components, with a cumulative contribution rate of over 90%.
[0117] During SVM model training, an SVM is used to construct a battery health status assessment model. The dimensionality-reduced features are used as input, and the actual battery health status (divided into three levels: good, average, and poor) is used as the output label for training. The kernel function chosen is the Radial Basis Function (RBF). The core of this SVM-based implementation (using an RBF kernel) is to solve the following optimization problem and construct a decision function:
[0118] Optimization goal:
[0119]
[0120] Constraints:
[0121]
[0122] in, It is a penalty parameter. It is a slack variable. It is a function that maps samples to a high-dimensional space, and is actually implemented using kernel function techniques.
[0123] The RBF kernel function is:
[0124]
[0125] in, These are kernel function parameters. It is Euclidean distance.
[0126] The final decision function is:
[0127]
[0128] in, These are Lagrange multipliers, obtained by solving the dual problem. The optimal penalty parameters are determined using cross-validation. and kernel function parameters .
[0129] During real-time assessment and early warning, the collected battery data is preprocessed and features extracted during the operation of the energy storage device, and then input into a trained SVM model to quickly assess the current health status of the battery. Based on the model output, a reasonable early warning threshold is set. When the battery health status assessment result is poor, a battery performance degradation warning is issued to remind maintenance personnel to conduct timely inspections and maintenance. Furthermore, during the grid twin model's health monitoring and early warning process in this embodiment, when the false alarm or missed alarm rate increases, the SVM model retraining process is triggered. Newly accumulated labeled data is used to update the classification model, improving assessment accuracy.
[0130] In the specific application of this embodiment, step four, energy management optimization of the energy storage system based on deep reinforcement learning, specifically includes the construction of a Deep Q-Network (DQN) model, DQN model training, and real-time energy management decision-making.
[0131] In the construction of the DQN model, specifically, the energy storage scheduling optimization model adopts a deep Q-network model architecture, as follows:
[0132] The intelligent agent of the energy storage scheduling optimization model is the energy storage load; the action space of the energy storage scheduling optimization model includes battery charging and discharging control commands, which include charging and discharging power control commands and state switching commands; the state space of the energy storage scheduling optimization model includes load data, power data, battery data, and electricity price data.
[0133] The reward function of the energy storage dispatch optimization model is constructed based on grid revenue, grid cost, grid stability index, and battery loss index. Among them, grid revenue includes the revenue obtained from discharging during peak grid electricity price periods; grid cost includes the cost of charging during off-peak grid electricity price periods and the cost of battery charging and discharging losses; grid stability index is quantified by the reciprocal of the square of the grid frequency deviation, and the smaller the frequency deviation, the higher the value of the grid stability index; battery loss index is obtained based on the relationship between the battery's charge and discharge depth, the number of cycles, and lifespan decay, and the smaller the battery loss, the higher the value of the battery loss index.
[0134] The energy storage scheduling optimization model adopts a three-layer neural network structure. The number of neurons in the input layer is equal to the dimension of the state space, the hidden layer consists of neurons with ReLU as the activation function, and the number of neurons in the output layer is equal to the dimension of the action space.
[0135] In one embodiment, the intelligent agent is the energy storage load, i.e., the energy storage device. The action space includes the charging and discharging power and state switching (charging, discharging, standby, etc.) of different energy storage batteries. For example, for lithium-ion batteries, the charging and discharging power ranges from -5kW (discharging) to 5kW (charging), with a power interval of 0.5kW; state switching includes three states: charging, discharging, and standby. The state space includes information such as grid load, renewable energy generation power, battery SOC, and grid electricity price. For example, the grid load ranges from 0 to 1000kW, the renewable energy generation power (photovoltaic + wind power) ranges from 0 to 800kW, the battery SOC ranges from 0 to 1, and the grid electricity price is divided into three levels: off-peak price (0.3-0.4 yuan / kWh), normal price (0.5-0.6 yuan / kWh), and peak price (0.8-1.0 yuan / kWh).
[0136] As a preferred embodiment of this invention, we define the reward function as follows:
[0137] R = α × (Revenue - Cost) + β × Power Grid Stability Index + γ × Battery Life Loss Index
[0138] The revenue mainly includes the income from discharging during peak electricity price periods, while the cost includes the cost of charging during off-peak electricity price periods and the cost of battery charging and discharging losses. The grid stability index is quantified by the reciprocal of the square of the grid frequency deviation; the smaller the frequency deviation, the higher the index value. The battery life loss index is calculated based on the relationship model between the battery's charge and discharge depth, the number of cycles, and life decay; the smaller the loss, the higher the index value.
[0139] In one embodiment, the reward function The formula is expressed as:
[0140]
[0141] in: , and The preset weighting coefficients, and .
[0142] Further:
[0143]
[0144]
[0145]
[0146] In one embodiment, the network structure adopts a three-layer neural network structure, with the number of neurons in the input layer corresponding to the state space dimension. In a specific example, there are four states: grid load, renewable energy generation power, battery SOC, and grid electricity price. After appropriate quantization, each state has 10 neurons in the input layer, 64 neurons in the hidden layer, and the activation function is ReLU. The number of neurons in the output layer corresponds to the action space dimension. Specifically, the lithium-ion battery has 21 charge / discharge power levels and 3 state transitions, totaling 24 actions, resulting in 24 neurons in the output layer.
[0147] During DQN model training, historical running data is used to train the DQN model, allowing the agent to learn the optimal action policy through continuous trial and error to maximize cumulative reward. The loss function of DQN is:
[0148]
[0149] Where the target value for:
[0150]
[0151] During training, an experience replay technique was used, with the replay buffer size set to 10,000. Each time, 32 experiences were randomly selected from the buffer for training. A target network was used, and the parameters of the target network were updated every 1,000 steps to stabilize the training process of the model.
[0152] The learning rate is 0.0001, the discount factor is 0.9, and the initial exploration rate is 1.0, which gradually decays to 0.1 as training progresses. Specifically, after deploying the DQN model, the energy storage scheduling optimization model is equipped with an online experience replay buffer, which stores macroscopic experience tuples.
[0153] The energy storage scheduling optimization model samples macro-experience tuples from the online experience replay buffer based on the scheduling cycle, and performs online fine-tuning of the network parameters in the energy storage scheduling optimization model based on the sampled macro-experience tuples.
[0154] In one embodiment, after completing the DQN model deployment, a separate online experience replay buffer is established. This buffer is specifically used to receive the evolutionary feedback signal from step six of the system-level macro performance evaluation and its corresponding macro experience tuples. During energy storage scheduling, samples are periodically taken from this buffer to fine-tune the parameters of the DQN model online at a low learning rate. This allows the decision-making strategy to continuously evolve based on actual operational results, thereby addressing long-term challenges such as system aging and environmental changes.
[0155] In real-time energy management decision-making, for the real-time operation of energy storage devices, the current grid load, renewable energy generation power, battery SOC, grid electricity price, and other state information are quantified and processed, and then input into a trained DQN model. The model outputs charging and discharging power and state control commands for the energy storage battery based on the learned optimal strategy. In one embodiment, when the grid is in a low-price period and renewable energy generation power is high, the model controls the energy storage battery to charge at a higher power, and at the same time, it rationally allocates charging power according to the battery's current SOC and health status to avoid overcharging. When the grid is in a high-price period and load demand is high, the model controls the energy storage battery to discharge at an appropriate power, prioritizing the grid load demand, and optimizes the discharge strategy based on the battery's discharge performance and remaining capacity to achieve the best balance between economic benefits and system stability.
[0156] When implementing coordinated control of multiple types of energy storage batteries based on step five, we adopt a coordinated control strategy for multiple types of energy storage batteries based on the charging and discharging power and state control commands output by the DQN model, combined with the characteristics of different energy storage batteries, such as the high energy density and fast power response of lithium-ion batteries, the low cost and long cycle life of sodium-ion batteries, and the mature technology and high reliability of lead-acid batteries. In one embodiment, during charging, priority is given to utilizing off-peak electricity price periods and periods of surplus renewable energy generation, and charging is carried out in the order of lithium-ion batteries, sodium-ion batteries, and lead-acid batteries. At the same time, reasonable charging power is allocated according to the SOC and health status of the batteries. When the SOC of the lithium-ion battery is low and its health status is good, it is charged with a higher power first to quickly reach a higher SOC. When the SOC of the lithium-ion battery is close to the upper limit or its health status is average, its charging power is appropriately reduced, and the charging power of the sodium-ion battery and the lead-acid battery is increased. In another embodiment, based on grid load demand and electricity price, lithium-ion batteries are prioritized for discharge to meet high power output requirements. When the state of charge (SOC) of the lithium-ion battery is low or the depth of discharge is close to the limit, sodium-ion batteries and lead-acid batteries discharge in tandem to ensure the stability and continuity of grid power supply. For example, during peak grid electricity price periods and when there is a sudden increase in load, lithium-ion batteries discharge at rated power, while sodium-ion batteries and lead-acid batteries assist in discharge at a certain proportion of power to jointly meet grid load demand.
[0157] Based on steps one through five above, we have constructed a complete energy storage scheduling method. However, in practical applications, we have found that in the static state of steps one through five, data deviations arise due to the accumulation of data during multiple energy storage scheduling operations, leading to performance degradation in energy storage scheduling. To form a complete intelligent scheduling closed-loop optimization and ensure the safety and economy of the decision-making strategy, this embodiment conducts a comprehensive "audit" and evaluation of the overall effectiveness of the executed strategy after a complete scheduling cycle, for example, after 24 hours, i.e., step six: scheduling strategy performance evaluation and system closed-loop optimization. By constructing a multi-dimensional performance audit system and relying on a digital twin system for pre-simulation and post-root cause analysis, the quantitative audit conclusions are ultimately transformed into feedback signals driving the evolution of the core decision-making model, realizing continuous autonomous optimization of the system strategy and solving the performance degradation problem caused by model mismatch and environmental changes in traditional systems. Among these, the most important aspect is how to obtain the feedback signal driving the evolution of the core decision-making model, i.e., the evolution reward signal.
[0158] Specifically, the evolutionary reward signal is obtained by auditing the results of multi-battery collaborative control to achieve energy storage scheduling. This audit is based on the grid performance, grid stability, and grid health, resulting in corresponding performance indicators, stability indicators, and health indicators. The evolutionary reward signal is obtained by weighted summation of the performance indicators, stability indicators, and health indicators.
[0159] Based on the above, this embodiment abandons a single indicator and adopts a comprehensive weighted audit scoring card system to quantify the overall performance of the strategy from three core dimensions, forming a macro performance audit indicator system.
[0160] Specifically, the performance indicators are derived from the performance indicator calculation formula, and are used to quantify the total net revenue within a scheduling cycle; the specific calculation formula for the performance indicators is as follows:
[0161]
[0162] in, For the calculated performance indicators, The total time step corresponding to the scheduling period. For a moment The discharge power of the energy storage device, For a moment The electricity price of energy storage devices, For a moment The charging power of energy storage devices, For a moment The electricity purchase price of energy storage devices This represents the real-time time step during quantization.
[0163] Specifically, the stability index is derived from the stability index calculation formula, and it is used to quantify the overall effect of energy storage load in mitigating power fluctuations in the power grid within a scheduling cycle; the stability index calculation formula is as follows:
[0164]
[0165] in, The calculated stability index, The total time step corresponding to the scheduling period. For a moment The load power, For a moment Power generation capacity, For a moment The power of energy storage devices.
[0166] Specifically, the health index is derived from a health index calculation formula, and it is used to quantify the average relative current stress of the battery pack of the energy storage device within a scheduling cycle to assess its cumulative loss; the specific calculation formula for the health index is as follows:
[0167]
[0168] in, For the calculated health indicators, The total time step corresponding to the scheduling period. For a moment The absolute value of the actual current. This is the rated current of the battery pack. This represents the real-time time step during quantization.
[0169] In one specific application of this embodiment, the final comprehensive audit score for:
[0170]
[0171] in: , and These are the weighting coefficients, and Its value can be adjusted according to the strategic goals of the power grid operator at different stages.
[0172] To ensure the accuracy of a comprehensive "audit" and evaluation of the overall effectiveness of the implemented strategy, we conduct pre-implementation compliance verification based on a high-fidelity digital twin. Before strategy execution, it can be sent to a high-fidelity digital twin simulation platform for simulation. This platform incorporates accurate electrochemical-thermal-lifetime coupled battery models, power grid equivalent models, and market models. After receiving the strategy, the digital twin executes a complete cycle in the virtual environment and outputs the estimated audit score. In addition, security and performance thresholds are set to verify the compliance of the policy. .like If the strategy is deemed high-risk or inefficient, the system will refuse to execute it and trigger a preset backup strategy generation mechanism.
[0173] Based on the above technical foundation, we will discuss how the energy storage scheduling method based on digital twins and reinforcement learning in this embodiment achieves intelligent scheduling with prior verification, post-effect evaluation, closed-loop feedback and autonomous evolution from three perspectives: periodic aftereffect audit and root cause analysis, evaluation conclusion feedback and system evolution, and long-term digital twin model update, thereby improving the long-term operating efficiency of energy storage equipment.
[0174] For post-cycle auditing and root cause analysis, after a real energy storage dispatch cycle is completed, the actual audit score is calculated based on the actual operating data collected from the Battery Management System (BMS), meters, etc. Then calculate the deviation between the simulation and the actual results. This allows for the establishment of a mapping database between performance deviations and potential causes, enabling automated root cause analysis. In one embodiment, if... The large discrepancy, primarily caused by EPI bias, suggests possible reasons such as inaccurate electricity price forecasting models or market model mismatch. The large deviation, primarily caused by AHDI bias, suggests a possible reason: insufficient accuracy of the battery life model. Larger and Meets the standards If the standard is not met, the root cause is speculated to be an error in the power conversion system (PCS) acting as the actuator, or a communication delay.
[0175] Regarding the feedback of evaluation conclusions and system evolution, specifically, the macro-level empirical tuple also includes the initial state, the overall strategy sequence, and the final state; among which, the initial state, the overall strategy sequence, and the final state are, in a scheduling cycle during energy storage scheduling, the initial state of the power grid and its loads in that scheduling cycle, the overall strategy sequence adopted for the power grid and its loads in that scheduling cycle, and the final state of the power grid and its loads in that scheduling cycle.
[0176] In one specific application of this embodiment, the actual audit score calculated in the previous step is... As an evolutionary reward signal ,Right now The system state at the initial stage of strategy execution will then be recorded. The overall strategy sequence adopted throughout the entire cycle Evolutionary reward signals and the state after the cycle ends. Together they form a macro-level empirical tuple. Then, this macroscopic experience tuple is stored in a dedicated evolutionary experience replay buffer. The Deep Reinforcement Learning (DRL) agent in step four, i.e. the aforementioned DQN model, periodically samples from this buffer and uses it as a reward value to update its network parameters. This allows the agent to learn not only instantaneous rewards, but also long-term overall policy performance, thereby continuously optimizing.
[0177] For long-term digital twin model updates, we use long-term post-audit deviations as calibration information for the digital twin. When specific types of deviations continue to occur, the parameter calibration process of the corresponding sub-models in the digital twin, such as the battery degradation model and the photovoltaic power output model, is automatically triggered to ensure their consistency with the physical entity, thereby maintaining the reliability of pre-verification.
[0178] Reference Figure 3 , Figure 3 This is a block diagram of the digital twin and reinforcement learning-based energy storage scheduling system in this embodiment.
[0179] like Figure 3 As shown, this embodiment also discloses an energy storage scheduling system based on digital twins and reinforcement learning, applicable to the energy storage scheduling method based on digital twins and reinforcement learning described above. The system includes:
[0180] The energy storage scheduling optimization model construction module is used to construct an energy storage scheduling optimization model with load forecasting results, power forecasting results, and battery evaluation results as inputs and battery charging and discharging control commands as outputs through reinforcement learning.
[0181] The energy storage scheduling module is used to perform multi-battery collaborative control based on battery charging and discharging control commands to achieve energy storage scheduling;
[0182] The power grid twin model construction module is used to construct a power grid twin model with actual operating data as input and macro-level experience tuples as output through digital twins. The macro-level experience tuples include at least an evolutionary reward signal, which is used to quantify the power grid performance, power grid stability, and power grid health during energy storage dispatch.
[0183] The energy storage scheduling optimization model optimization module is used to sample macroscopic empirical tuples after a preset scheduling period ends when the energy storage scheduling optimization model is running, so as to optimize the energy storage scheduling optimization model.
[0184] The energy storage scheduling optimization module is used to perform multi-battery collaborative control based on the battery charging and discharging control commands output by the optimized energy storage scheduling optimization model in order to achieve energy storage scheduling.
[0185] It should be noted that the energy storage scheduling system based on digital twins and reinforcement learning in this embodiment corresponds to the aforementioned energy storage scheduling method based on digital twins and reinforcement learning. Therefore, any content not specifically described in the energy storage scheduling system based on digital twins and reinforcement learning in this embodiment, including but not limited to functional definitions, working principles, and technical effects, can be referred to the description in the aforementioned energy storage scheduling method based on digital twins and reinforcement learning, and will not be repeated here.
[0186] In summary, the energy storage scheduling method and system based on digital twins and reinforcement learning in this embodiment differs from existing technologies by mainly making the following optimizations, thereby producing corresponding beneficial effects:
[0187] (1) Accurate prediction capability
[0188] Introduction of Deep LSTM Model: Unlike traditional methods that struggle to accurately capture renewable energy generation and grid load fluctuations, this embodiment introduces a deep long short-term memory network (LSTM), whose unique gating mechanism can effectively handle long-term dependencies in time series data.
[0189] Beneficial effects: Deep LSTM models can more accurately predict the changing trends of grid load and renewable energy generation, providing more reliable data support for the scheduling decisions of energy storage systems and reducing grid fluctuations and energy waste caused by inaccurate predictions.
[0190] (2) Adaptive optimization scheduling
[0191] This embodiment innovatively utilizes the Deep Q-Network (DQN) reinforcement learning algorithm to optimize the charging and discharging strategy of an energy storage system. DQN can autonomously learn and formulate the optimal charging and discharging strategy based on real-time grid and battery conditions, dynamically balancing grid demand and system cost.
[0192] Beneficial effects: By learning and adapting to changes in grid characteristics in real time, the DQN algorithm can realize intelligent scheduling of energy storage systems, effectively reducing system operating costs and improving system economy and flexibility while meeting grid demand.
[0193] (3) Detailed battery health management
[0194] This solution combines principal component analysis (PCA) with support vector machine (SVM) to address the shortcomings of existing energy storage system battery health status assessment. PCA is used to reduce the dimensionality of battery operation data and extract key features, and then SVM is used to build a health status classification model.
[0195] Beneficial effects: It can provide early warning of potential battery failures, extend battery life, reduce system maintenance costs and failure risks, and improve the reliability and security of energy storage dispatch.
[0196] (4) Coordinated control of multiple types of energy storage batteries
[0197] Optimized Control Strategy: Unlike existing technologies that lack effective coordinated control methods for different types of energy storage batteries, this embodiment proposes a coordinated control strategy for multiple types of energy storage batteries. Based on battery characteristics and grid demands, charging and discharging power is rationally allocated to achieve complementary advantages among various battery types.
[0198] Beneficial effects: Fully utilize the energy density, power response speed and cycle life characteristics of different batteries to improve the overall performance and efficiency of energy storage systems, enhance the ability to regulate fluctuations in new energy power generation, and better meet the peak shaving and frequency regulation needs of the power grid.
[0199] (5) System-level closed-loop evolution capability
[0200] Most existing technologies are "open-loop" systems, with models that remain fixed after deployment. This embodiment, for the first time, establishes an intelligent closed-loop architecture of "decision-verification-evaluation-learning." It achieves pre-emptive security verification of strategies through digital twin technology, quantitatively audits the post-strategy effects through a multi-objective performance evaluation system, and finally uses the evaluation results as an evolutionary feedback signal to drive the online optimization of the core decision-making model. This transforms energy storage scheduling from a statically executed automated tool into an intelligent tool that can continuously learn from actual operation, adapt to changes, and autonomously evolve its performance, fundamentally solving the problems of mismatch and performance degradation in existing energy storage scheduling models.
[0201] In the embodiments provided in this application, it should be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code, or any suitable combination thereof. For hardware implementation, the processor may be implemented in one or more of the following: application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to implement the functions described herein, or combinations thereof. For software implementation, some or all of the processes of the embodiments may be performed by a computer program instructing the associated hardware. During implementation, the program may be stored in a computer-readable storage medium or transmitted as one or more instructions or code on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available medium accessible to a computer. Computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code having the form of instructions or data structures and accessible to a computer.
[0202] Finally, it should be noted that the above description is only a preferred embodiment of this application and is not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An energy storage scheduling method based on digital twins and reinforcement learning, applied to the same power grid and its loads, wherein the loads include electrical loads and energy storage loads, characterized in that, The method includes: Through reinforcement learning, an energy storage scheduling optimization model is constructed, which takes load forecasting results, power forecasting results, and battery evaluation results as inputs and battery charge and discharge control commands as outputs; wherein, the battery charge and discharge control commands are charge and discharge power and state control commands. Based on the battery charging and discharging control commands, multi-battery collaborative control is performed to achieve energy storage scheduling; wherein, the multi-battery collaborative control adopts a collaborative control strategy for multiple types of energy storage batteries based on the charging and discharging power and state control commands output by the DQN model and combined with the characteristics of different energy storage batteries. Digital twins are used to construct a power grid twin model with actual operating data as input and macro-level experience tuples as output. The macro-level experience tuples include at least an evolutionary reward signal, which is used to quantify the power grid performance, power grid stability, and power grid health during energy storage scheduling. When the energy storage scheduling optimization model is running, after the preset scheduling period ends, the energy storage scheduling optimization model samples the macroscopic empirical tuples to optimize the energy storage scheduling optimization model. Based on the battery charging and discharging control commands output by the optimized energy storage scheduling model, multi-battery collaborative control is performed to achieve energy storage scheduling. The energy storage scheduling optimization model is equipped with an online experience replay buffer, which is used to store the macroscopic experience tuples. The energy storage scheduling optimization model samples the macro-experience tuples from the online experience replay buffer based on the scheduling period, and performs online fine-tuning of the network parameters in the energy storage scheduling optimization model based on the sampled macro-experience tuples. The evolutionary reward signal is obtained by auditing the results of multi-battery collaborative control to achieve energy storage scheduling. This audit is based on the grid performance, grid stability, and grid health, resulting in corresponding performance indicators, stability indicators, and health indicators. The evolutionary reward signal is obtained by weighted summation of the performance indicators, stability indicators, and health indicators. Specifically: the performance indicators quantify the total net benefit within a scheduling cycle; the stability indicators quantify the overall effect of the energy storage load in mitigating grid power fluctuations within a scheduling cycle; and the health indicators quantify the average relative current stress of the energy storage device's battery pack within a scheduling cycle to assess its cumulative losses. The macro-level experience tuple also includes an initial state, an overall strategy sequence, and an ending state; wherein, the initial state, the overall strategy sequence, and the ending state are specifically, during energy storage scheduling, within a scheduling cycle, the initial state of the power grid and its loads in that scheduling cycle, the overall strategy sequence adopted for the power grid and its loads in that scheduling cycle, and the ending state of the power grid and its loads in that scheduling cycle.
2. The energy storage scheduling method based on digital twin and reinforcement learning according to claim 1, characterized in that, The energy storage scheduling optimization model adopts a deep Q-network model architecture, specifically: The intelligent agent of the energy storage scheduling optimization model is the energy storage load; the action space of the energy storage scheduling optimization model includes the battery charging and discharging control commands, which include charging and discharging power control commands and state switching commands; the state space of the energy storage scheduling optimization model includes load data, power data, battery data, and electricity price data. The reward function of the energy storage dispatch optimization model is constructed based on grid revenue, grid cost, grid stability index, and battery loss index. Specifically: grid revenue includes the revenue gained from discharging during peak grid electricity price periods; grid cost includes the cost of charging during off-peak grid electricity price periods and the cost of battery charging and discharging losses; the grid stability index is quantified by the reciprocal of the square of the grid frequency deviation, with a higher value for the smaller the frequency deviation; and the battery loss index is obtained based on the relationship between the battery's charge / discharge depth, cycle count, and lifespan decay, with a higher value for the smaller the battery loss. The energy storage scheduling optimization model adopts a three-layer neural network structure. The number of neurons in the input layer is equal to the dimension of the state space, the hidden layer consists of neurons with ReLU activation function, and the number of neurons in the output layer is equal to the dimension of the action space.
3. The energy storage scheduling method based on digital twin and reinforcement learning according to claim 1, characterized in that, The performance indicators are obtained based on the performance indicator calculation formula, and the performance indicators are used to quantify the total net revenue within one scheduling cycle; wherein, the specific calculation formula for the performance indicators is: in, For the calculated performance indicators, The total time step corresponding to the scheduling period. For a moment The discharge power of the energy storage device, For a moment The electricity price of energy storage devices, For a moment The charging power of energy storage devices, For a moment The electricity purchase price of energy storage devices This represents the real-time time step during quantization.
4. The energy storage scheduling method based on digital twin and reinforcement learning according to claim 1, characterized in that, The stability index is obtained based on the stability index calculation formula, and the stability index is used to quantify the overall effect of the energy storage load in smoothing power fluctuations in the power grid within a scheduling cycle; wherein, the stability index calculation formula is specifically as follows: in, The calculated stability index, The total time step corresponding to the scheduling period. For a moment The load power, For a moment Power generation capacity, For a moment The power of energy storage devices.
5. The energy storage scheduling method based on digital twin and reinforcement learning according to claim 1, characterized in that, The health index is obtained based on a health index calculation formula, and the health index is used to quantify the average relative current stress of the battery pack of the energy storage device within a scheduling cycle to assess its cumulative loss; wherein, the specific calculation formula of the health index is: in, For the calculated health indicators, The total time step corresponding to the scheduling period. For a moment The absolute value of the actual current. This is the rated current of the battery pack. This represents the real-time time step during quantization.
6. The energy storage scheduling method based on digital twin and reinforcement learning according to claim 1, characterized in that, Before constructing the energy storage scheduling optimization model, the method includes: Basic data is collected and preprocessed to obtain scheduling data; wherein, the basic data includes at least load data, power data and battery data; Based on the scheduling data, load forecasting and power forecasting are performed to obtain the load forecasting results and the power forecasting results. A health status assessment is performed based on the battery data to obtain the battery assessment result.
7. An energy storage scheduling system based on digital twins and reinforcement learning, applicable to the energy storage scheduling method based on digital twins and reinforcement learning as described in any one of claims 1-6, characterized in that, The system includes: The energy storage scheduling optimization model construction module is used to construct an energy storage scheduling optimization model with load forecasting results, power forecasting results, and battery evaluation results as inputs and battery charging and discharging control commands as outputs through reinforcement learning. The energy storage scheduling module is used to perform multi-battery collaborative control based on the battery charging and discharging control commands to achieve energy storage scheduling; The power grid twin model construction module is used to construct a power grid twin model with actual operating data as input and macro-level experience tuples as output through digital twins. The macro-level experience tuples include at least an evolutionary reward signal, which is used to quantify the power grid performance, power grid stability, and power grid health during energy storage dispatch. An energy storage scheduling optimization model optimization module is used to sample the macroeconomic empirical tuples after a preset scheduling period ends when the energy storage scheduling optimization model is running, so as to optimize the energy storage scheduling optimization model. The energy storage scheduling optimization module is used to perform multi-battery collaborative control to achieve energy storage scheduling based on the battery charging and discharging control commands output by the optimized energy storage scheduling optimization model.
Citation Information
Patent Citations
Digital twin energy management method and system for source network load storage cooperative scheduling
CN120749911A