Energy storage reverse flow suppression method and system based on time series convolution prediction and offline reinforcement learning

CN122801241APending Publication Date: 2026-09-22ZHONGYAODA DIGITAL ENERGY ECOLOGICAL TECH (ZHEJIANG) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611265502.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-20
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,储能系统的调度策略直接决定了逆流抑制效果与经济收益的平衡,若调度策略不当,仍可能在特定时段出现并网点逆流,或因过度保守而损失峰谷套利空间

Benefits of technology

[0041]1、适应低频采集场景,降低硬件部署门槛。本发明基于常规分钟级采集数据即可有效运行,无需增设高频采集硬件,可低成本部署于存量工商业光储场站,显著降低系统改造成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122801241A_ABST
    Figure CN122801241A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of distributed photovoltaic and energy storage control, and particularly relates to a kind of energy storage reverse flow suppression method and system based on timing convolution prediction and offline reinforcement learning.The present application includes obtaining historical operation data, constructing training sample set;Construct a double-branch timing convolution network, train using the training sample set, for outputting the net load prediction sequence of future time;Energy storage scheduling decision model based on offline reinforcement learning is constructed, with the optimization goal of minimizing grid-connected point reverse flow and maximizing economic benefit, the energy storage scheduling decision model is trained on offline data set;The state vector collected in real time is input into the trained energy storage scheduling decision model, and the charge and discharge power is obtained;Based on the current energy storage power, the charge and discharge power is hard constrained and clipped, for controlling the operation of energy storage system.The present application realizes high-precision load prediction, and through offline reinforcement learning, a robust scheduling strategy is constructed, taking into account the dual goals of reverse flow suppression and peak-valley arbitrage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of distributed photovoltaic and energy storage control technology, and particularly relates to an energy storage backflow suppression method and system based on temporal convolutional prediction and offline reinforcement learning. Background Technology

[0002] With the large-scale grid-connected application of distributed photovoltaic and energy storage systems, the problem of reverse current at the grid connection point is becoming increasingly prominent. When the sum of the discharge power of the energy storage system and the output of the photovoltaic system exceeds the local load absorption capacity, the excess electricity will be injected back into the grid, forming a reverse current. This can cause uncontrollable power on the grid side, power quality pollution, and other hazards, and in severe cases, can lead to safety accidents.

[0003] In industrial and commercial applications, many factories and commercial complexes choose to build distributed photovoltaic (PV) systems on their premises or building rooftops to reduce electricity costs. During periods of intense sunlight, such as midday on sunny days, PV power generation often far exceeds the user's load demand, easily leading to reverse current. To address this, some industrial and commercial users have further integrated energy storage systems, forming a combined PV-storage system. The energy storage system can, on the one hand, charge and absorb excess electricity and suppress reverse current when PV output is excessive; on the other hand, it can leverage the peak-valley difference in time-of-use pricing to discharge electricity to supply the load during peak hours, generating arbitrage profits. However, the scheduling strategy of the energy storage system directly determines the balance between reverse current suppression and economic benefits. An inappropriate scheduling strategy may still result in reverse current at the grid connection point during specific periods, or excessive conservatism may lead to the loss of peak-valley arbitrage opportunities. Once reverse current occurs, the operator will face grid connection assessments and economic penalties from the power grid company, directly impacting the overall profitability of the PV-storage system.

[0004] Existing energy storage anti-reverse flow control technologies have significant shortcomings at both the forecasting and scheduling levels. At the forecasting level, existing methods struggle to effectively handle the prevalent nonlinear and non-stationary characteristics of photovoltaic power output and load power, especially under complex and variable weather conditions, resulting in large forecasting errors and inaccurate feedforward control commands. At the scheduling level, existing methods lack the ability to proactively address forecast uncertainties and exhibit insufficient robustness.

[0005] For example, patent CN110098628B discloses an anti-backflow method for an energy storage demand control system. This method calculates the low-order harmonic current distortion rate by performing a Fourier transform on load power monitoring data and comparing it with a load impact threshold. If the threshold is exceeded, a power reduction command is issued. However, this method relies on cycle-level sampling analysis of the current waveform, requiring dedicated high-frequency acquisition hardware. This places high demands on the sampling frequency of the acquisition equipment, resulting in significant system deployment costs and making it difficult to implement at low cost in industrial and commercial scenarios.

[0006] Patent CN114362249B discloses a backflow prevention control method for a source-load system. It employs an LSTM-BP deep neural network to perform short-term predictions of power supply and load power, assessing backflow risk in advance and executing anti-backflow operations before it occurs. However, the effectiveness of this method highly depends on the accuracy of the prediction model. Under atypical conditions such as drastic load fluctuations, sudden weather changes, or holidays, prediction errors often increase significantly, leading to delayed or inaccurate feedforward control commands, ultimately failing to effectively prevent backflow. Furthermore, this method directly maps prediction results to control actions, lacking proactive capabilities to address prediction uncertainties, resulting in insufficient system robustness. Therefore, it is necessary to introduce a decision-making method capable of learning optimal scheduling strategies from historical operating data and having a certain tolerance for prediction errors to compensate for the limitations of pure prediction-driven control under complex operating conditions.

[0007] Patent application CN115276100A discloses an anti-reverse current control system for photovoltaic energy storage integrated machines. This system involves setting power sampling points at the output terminals of the photovoltaic unit, the AC output terminal of the inverter, and the grid output terminal. A dedicated anti-reverse current acquisition and control device monitors the power of each node in real time, and controls the switching action based on the power comparison results to achieve grid-connected mode switching. However, this method requires additional dedicated acquisition and control devices and multiple power sensors, resulting in high hardware costs and significant system integration complexity. Summary of the Invention

[0008] The purpose of this invention is to solve the problems in the prior art and to propose a method and system for suppressing backflow in energy storage based on temporal convolutional prediction and offline reinforcement learning. While achieving high-precision load prediction, it constructs a robust scheduling strategy through offline reinforcement learning, thus taking into account the dual objectives of backflow suppression and peak-valley arbitrage.

[0009] To achieve the above objectives, the technical solution provided by this invention is as follows:

[0010] Firstly, a method for suppressing energy storage backflow based on temporal convolutional prediction and offline reinforcement learning is provided, including:

[0011] Obtain grid connection point power data and photovoltaic power data, and calculate the historical net load sequence;

[0012] The historical net load sequences are integrated to obtain the full net load sequence, and the corresponding multidimensional time feature vectors are extracted. A training sample set is constructed based on the full net load sequence.

[0013] A dual-branch temporal convolutional network is constructed and trained using a training sample set to output the net load prediction sequence for future time periods.

[0014] An energy storage scheduling decision model based on offline reinforcement learning is constructed. The normalized full net load sequence, the normalized net load prediction sequence, the current state of charge and the state vector composed of multi-dimensional time feature vectors are used as inputs. The optimization objective is to minimize the backflow at the grid connection point and maximize the economic benefits. The energy storage scheduling decision model is trained on an offline dataset.

[0015] The real-time collected state vector is input into the trained energy storage scheduling decision model, and the normalized action command is output. The normalized action command is scaled to the charging and discharging power.

[0016] Hard constraints are applied to the charging and discharging power based on the current energy storage capacity to control the operation of the energy storage system.

[0017] The energy storage capacity is updated based on the trimmed charge and discharge power, and the state of charge for the next time step is obtained based on the updated energy storage capacity.

[0018] Furthermore, the construction of the training sample set based on the full net load sequence includes:

[0019] Extract the input subsequence for the current time step from the full net load sequence;

[0020] Centered on the current time step, extract the net load values ​​of historical reference points across cycles from the historical net load sequence to construct the context feature vector;

[0021] The context feature vector is concatenated with the multidimensional time feature vector to obtain the auxiliary feature vector;

[0022] Using the normalized net load value within a preset future step size as the prediction target, a training sample set for training a two-branch temporal convolutional network is constructed based on the input subsequence, auxiliary feature vector, and prediction target.

[0023] Furthermore, the dual-branch temporal convolutional network includes a first residual convolutional branch and a second residual convolutional branch. The first residual convolutional branch and the second residual convolutional branch are respectively composed of two layers of causal dilated convolution, GELU activation function and Dropout regularization layer connected in sequence to form the main path. The output of the main path is weighted by the channel attention mechanism and residually connected with the input features aligned with the number of channels. Then, after batch normalization, the outputs of the first residual convolutional branch and the second residual convolutional branch are obtained.

[0024] Among them, the inflation coefficients for constructing the first residual convolution branch and the second residual convolution branch are different;

[0025] The outputs of the first residual convolution branch and the second residual convolution branch are compressed by adaptive average pooling and then concatenated with the auxiliary feature vector. Finally, the net load prediction sequence is output through a three-layer fully connected prediction head.

[0026] Furthermore, the energy storage scheduling decision model is trained using a conservative Q-learning algorithm, which introduces a conservative regularization term into the Q-value estimation. The loss function is expressed by the following formula:

[0027]

[0028] in, The total loss function for conservative Q-learning, For standard Bellman error loss, The conservatism coefficient. To calculate the mathematical expectation of all states in the offline dataset, This represents the logarithmic summation of the Q values ​​over all possible actions using an exponential operation. behavioral strategies The expected Q-value of the covered actions, To generate the empirical distribution of historical scheduling strategies for offline datasets, This is the Q-value of the action.

[0029] Furthermore, the reward function of the energy storage scheduling decision model is expressed by the following formula:

[0030]

[0031] in, For the first The reward function for each time step. This is a truncation function. Punishment for going against the tide Penalty for exceeding demand limits. Penalty for exceeding the charge state limit, Penalty for charging and discharging losses, For peak-valley arbitrage rewards, As a reward scaling factor, To truncate the upper limit.

[0032] Furthermore, the reverse current penalty is calculated based on the grid connection point power, the demand over-limit penalty is the difference between the preset demand threshold and the grid connection point power, the state of charge over-limit penalty is calculated based on the state of charge, the charging and discharging loss penalty is calculated based on the charging and discharging power, and the peak-valley arbitrage reward is 1 if charging is done during off-peak hours and 0.5 if discharging is done during peak hours.

[0033] Furthermore, the hard constraint trimming of charging and discharging power based on the current energy storage capacity includes: trimming the charging and discharging power to a limit range according to the rated capacity of the energy storage system, the current energy storage capacity, and the charging and discharging efficiency.

[0034] Secondly, a system for suppressing energy storage backflow based on temporal convolutional prediction and offline reinforcement learning is provided, including:

[0035] The data preprocessing module is used to acquire grid-connected power data and photovoltaic power data, calculate the historical net load sequence, integrate the historical net load sequence to obtain the full net load sequence, and extract the corresponding multi-dimensional time feature vector; and construct a training sample set based on the full net load sequence.

[0036] The dual-branch temporal convolution prediction module is used to construct a dual-branch temporal convolutional network. It is trained using a training sample set constructed based on the full payload sequence and is used to output the payload prediction sequence for future time periods.

[0037] The energy storage scheduling decision modeling module is used to build an energy storage scheduling decision model based on offline reinforcement learning. It takes the state vector composed of the full net load sequence, the net load prediction sequence, the current state of charge, and the multi-dimensional time feature vector as input, and the optimization objective is to minimize the backflow at the grid connection point and maximize the economic benefits.

[0038] The conservative Q-learning strategy training module is used to train the energy storage scheduling decision model on an offline dataset.

[0039] The online inference and hard constraint execution module is used to input the real-time acquired state vector into the trained energy storage scheduling decision model, output normalized action commands, scale the normalized action commands into charging and discharging power, perform hard constraint pruning on the charging and discharging power based on the current energy storage capacity, and use it to control the operation of the energy storage system; update the energy storage capacity according to the pruned charging and discharging power, and obtain the state of charge of the next time step according to the updated energy storage capacity.

[0040] Compared with the prior art, the significant advantages of this invention are:

[0041] 1. Adaptable to low-frequency data acquisition scenarios, reducing the hardware deployment threshold. This invention can operate effectively based on conventional minute-level data acquisition, without the need for additional high-frequency acquisition hardware. It can be deployed at low cost in existing industrial and commercial optical storage sites, significantly reducing system upgrade costs.

[0042] 2. High prediction accuracy and strong adaptability to complex operating conditions. By constructing a bi-branch temporal convolutional network with different basic expansion coefficients and combining it with the SEBlock attention mechanism, it can effectively capture the multi-scale temporal features of the net load sequence. It maintains high prediction accuracy even under atypical operating conditions such as holidays and seasonal changes, providing reliable forward-looking information for scheduling decisions.

[0043] 3. Robust scheduling strategy, balancing backflow suppression and economic benefits. CQL offline reinforcement learning learns the optimal scheduling strategy from historical data, has a certain tolerance for prediction errors, effectively suppresses backflow at grid connection points, and fully utilizes the peak-valley difference of time-of-use electricity prices to achieve arbitrage profits, thus balancing the dual goals of safety compliance and economic benefits for industrial and commercial users. Attached Figure Description

[0044] Figure 1 This is a flowchart of the energy storage backflow suppression method based on temporal convolutional prediction and offline reinforcement learning according to the present invention;

[0045] Figure 2 This is a daily comparison curve of grid connection point power on a typical load day according to the present invention;

[0046] Figure 3 This is a daily scheduling diagram of the charging and discharging power output of this invention;

[0047] Figure 4 This is a daily variation curve of the state of charge of the energy storage system under the regulation of this invention;

[0048] Figure 5 This is a bar chart showing the distribution of the number of daily backflow reductions during the evaluation period of this invention.

[0049] Figure 6 This is a comparison chart of the number of daily backflows during the evaluation period of this invention;

[0050] Figure 7 This is a bar chart comparing the daily backflow rate trend during the evaluation period of this invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0052] Example 1.

[0053] like Figure 1 As shown, this invention provides a method for suppressing energy storage backflow based on temporal convolutional prediction and offline reinforcement learning. The specific steps are as follows:

[0054] Step S101: Obtain historical power data from industrial and commercial power plants, including grid connection point power data (gridEm2) and photovoltaic output power data (powerEm2), calculate net load data, remove abnormal date data and missing values, retain valid data, and form a historical net load sequence.

[0055] Specifically, the power acquisition data file of the power station is read, with a data granularity of 1 minute. The fields (including timestamp, grid connection point power and photovoltaic output power) are shown in Table 1.

[0056] Table 1 Historical Operation Data Table

[0057] t1 g1 pv1 l1 t2 g2 pv2 l2 t3 g3 pv3 l3 … … … …

[0058] The net load is calculated using the following formula:

[0059]

[0060] The power field of specified abnormal dates (such as dates where data is distorted due to equipment maintenance, communication interruption, etc.) is set to null and removed, while valid data is retained, ultimately forming a complete historical net load sequence.

[0061] Step S102: Integrate the historical net load sequence (including splicing, cleaning and alignment, gap filling and correction and standardization) to obtain the full net load sequence. Use mean-variance normalization (StandardScaler) to fit the full net load sequence and save the scaling parameters for use in the inference stage.

[0062] For each data record's timestamp Extract the 12-dimensional multidimensional time feature vector as shown in Table 2. :

[0063] Table 2 Explanation of Time Feature Vectors

[0064]

[0065] Where tod is the number of minutes of the day (0-1439), dow is the week number (0-6), month is the month number (1-12), and hour is the number of hours (0-23).

[0066] Step S103: Set the historical sequence window length HIST_LEN=96 (i.e., 96 time steps, corresponding to a 96-minute historical window at a 1-minute granularity), and the look-ahead window length PRED_LEN=12 (i.e., predict the net load for the next 12 minutes) required for scheduling decisions.

[0067] Traverse the full net load sequence using a sliding window approach, for the first... time steps ( ), cut The full net load sequence of the interval is used as the input subsequence. Simultaneously, the net load values ​​of historical reference points are extracted to construct the context feature vector. Among them, the historical reference point is the first... The time step is the corresponding time step of the previous natural week (7 days ago), and there are 7 sampling points in total, including 3 time steps before and after that corresponding time step. With multidimensional time feature vectors Concatenate to obtain auxiliary feature vectors ; with the first The normalized net load value at the time step is the forecast target. Using training samples Construct a training sample set.

[0068] Step S104: Construct a dual-branch temporal convolutional network (ImprovedTCN), with the following network structure:

[0069] (1) The basic residual unit (TCNBlock) of the temporal convolutional network: Each TCNBlock consists of two layers of causal dilated convolutions, a GELU activation function, and a regularization layer (Dropout) concatenated to form the main path. Tail pruning removes redundant time steps introduced at the end of the sequence due to causal padding, ensuring that the convolutional output depends only on information from the current and historical moments, and does not include information from future moments. The main path output is weighted by a channel attention mechanism (such as SEBlock, Squeeze-and-Excitation Block) (SEBlock compression ratio...). The intermediate layer dimension is taken The result is added to the residuals that have been aligned to the number of channels by a 1×1 convolution, and then batch-normalized (BatchNorm1d) before being output, as shown in the following formula:

[0070]

[0071] in, Indicates the first The input feature map of the TCNBlock, i.e. the th TCNBlock Output feature map of each TCNBlock This represents a 1×1 convolutional layer. This represents causal dilated convolution. This indicates the channel attention mechanism. This indicates batch normalization processing. Indicates the first Output feature map of each TCNBlock.

[0072] (2) Dual-branch structure: The first residual convolution branch constructs a multi-layer TCNBlock with the first basic expansion coefficient, and the expansion coefficients of each layer are 1, 2, 4 and 8 respectively; the second residual convolution branch constructs a multi-layer TCNBlock with the second basic expansion coefficient, and the expansion coefficients of each layer are 2, 4, 8 and 16 respectively; the number of channels in the two branches are [64, 128, 128, 64] respectively, to capture the short-term fluctuation characteristics and long-term periodic trends of the net load.

[0073] As another optional implementation, the dilation coefficients of each layer in the first residual convolution branch can be set to 1, 3, 9, and 27 sequentially, and the dilation coefficients of each layer in the second residual convolution branch can be set to 2, 6, 18, and 54 sequentially. That is, the first residual convolution branch uses a base dilation ratio of 3, and the second residual convolution branch uses a base dilation ratio of 2. The first residual convolution branch uses a base dilation ratio of 2, and the dilation coefficients of each layer in the second residual convolution branch are... ( , The base expansion ratio of the corresponding branch is set to increase exponentially according to a regular pattern. The second residual convolution branch is set according to the expansion coefficient of each layer. The regularity index is set to increase incrementally, so that the two branches have different receptive field ranges of different scales, which can also achieve the differentiated extraction effect of short-term fluctuation characteristics and long-term cycle trends.

[0074] Therefore, the expansion coefficient configuration of the first residual convolution branch and the second residual convolution branch can be generated exponentially by using different base expansion ratios (such as 2, 3, etc.). The specific numerical combinations are not limited to (1, 2, 4, 8 and 2, 4, 8, 16) and (1, 3, 9, 27 and 2, 6, 18, 54) listed in this embodiment. Those skilled in the art can select different base expansion ratios to differentiate the configuration of the two branches according to the periodic characteristics of the actual net load sequence and the model's requirements for extracting short-term fluctuations and long-term trend characteristics.

[0075] (3) Feature Fusion and Prediction Head: The outputs of the two branches are compressed to 1 dimension by adaptive average pooling and then concatenated to obtain a 128-dimensional feature vector. Auxiliary Feature Vector The features are encoded into 32 dimensions through a two-layer fully connected network (19→64→32, GELU activation). After concatenation, the features are passed through a three-layer fully connected prediction head (160→128→64→1, including a Dropout regularization layer) to output the 32-dimensional feature. Net load forecast at time step .

[0076] It should be noted that Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs) can be used to replace two-branch temporal convolutional networks. Both have temporal modeling capabilities, but they are not as efficient in parallel computing as two-branch temporal convolutional networks and have a slower training speed.

[0077] Step S105: Train the dual-branch temporal convolutional network using the training sample set constructed in step S103. After convergence, perform batch inference on the full historical data (batch size 512) to obtain the net load prediction sequence. ,in, This represents the total number of records in the full dataset.

[0078] Using the scaling parameters saved in step S102, the full net load sequence is... With predicted sequence Normalize them separately to obtain the normalized full net load sequence. With normalized net load forecast sequence For use in step S106 state construction:

[0079]

[0080] in, For the first The original net load value at each time step, For the first Normalized forecast net load at each time step, The mean of the standardized fitted net load sequence, is the standard deviation of the fitted net load sequence.

[0081] Step S106: Model the energy storage scheduling problem as a Markov Decision Process (MDP), with the triples defined as follows:

[0082] (1) State space : No. State vector at time step It consists of the following four parts:

[0083]

[0084] in, , which is the normalized historical net load window extracted from the normalized full net load sequence; , which is the normalized net load forecast window extracted from the normalized net load forecast sequence; For the first State of charge at the time step; This is a multidimensional time feature vector. The sum of the four dimensions is... dimension.

[0085] In scenarios requiring high prediction accuracy, the state space can be simplified by using only the normalized net load prediction sequence, state of charge, and multidimensional time feature vectors to form the state vector, thereby reducing the state dimension and the number of policy network parameters.

[0086] (2) Action space :action For the first The normalized charge / discharge power command at each time step is linearly mapped to obtain the charge / discharge power:

[0087]

[0088] in, For the first The charging and discharging power at each time step This represents the rated power of the energy storage system (102kW). A positive value indicates charging, and a negative value indicates discharging.

[0089] (3) Reward function : Execute action Subsequently, the grid connection power was updated to The reward function is defined as:

[0090]

[0091] The various reward categories are shown in Table 3:

[0092] Table 3. Explanation of each item in the reward function

[0093]

[0094] in, For the first The reward function for each time step. This is a truncation function. As a reward scaling factor, To truncate the upper limit, This is the demand threshold. In a charged state, This is the upper limit of the state of charge. This is the limit of the charged state.

[0095] Step S107: Construct the Actor policy network (energy storage scheduling decision model). The network structure is a fully connected network with an input layer dimension of 121, a hidden layer structure of 256→256→128 (each layer contains Layer Normalized Norm and Exponential Linear Unit ELU activation), and an output layer dimension of 1 (Hyperbolic Tangent Function Tanh activation). The output is a normalized action value. .

[0096] The Conservative Q-Learning (CQL) algorithm was used to train the Actor network on an offline historical dataset. CQL, based on the standard Actor-Critic framework, introduces a conservative regularization term for Q-value estimation:

[0097]

[0098] in, Standard Bellman error loss; This is the conservatism coefficient, which controls the intensity of the conservatism penalty. To calculate the mathematical expectation of all states in the offline dataset, This means taking the log-sum-exp value of Q for all possible actions and applying a push-up penalty to actions outside the dataset; behavioral strategies The expected Q-value of the covered actions applies pull-down protection to actions within the dataset; To generate the empirical distribution of historical scheduling strategies for offline datasets, This is the action Q-value. This regularization term makes the Q-function more conservative in its estimation of actions outside the dataset, thus preventing the policy from selecting unverified and dangerous actions due to overestimation during offline training, and improving the robustness of the policy in actual deployment.

[0099] Alternatively, offline TD3 (Twin Delayed Deep Deterministic Policy Gradient), TD3+BC (Twin Delayed DDPG with Behavior Cloning), or offline SAC (Soft Actor-Critic) algorithms can be used to replace CQL. These algorithms can also train continuous action policies on offline datasets and converge faster in some scenarios compared to CQL, but their conservative constraint mechanisms are different.

[0100] Step S108: During the evaluation phase, for each time step of the target date... From the dataset formed in step S105, the corresponding normalized full net load sequence and normalized net load prediction sequence are extracted, and combined with the current state of charge and multi-dimensional time feature vector, they are concatenated to form a 121-dimensional state vector in the manner defined in step S106. Input the trained Actor network and perform forward inference to obtain the normalized charging and discharging power command. Scaled to charge / discharge power .

[0101] Step S109: For Apply SOC hard constraint clipping, with the following clipping rules:

[0102]

[0103] Then limit the charging and discharging power of the trimmed parts to the rated range. Within the system, execute the command and update the energy storage capacity using the following formula:

[0104]

[0105] in, These are charging efficiency and discharging efficiency, respectively. This is the single-step time interval (corresponding to a 1-minute data collection granularity). This refers to the rated capacity of the energy storage system. and These are the upper and lower limits of the state of charge, respectively; and These represent the maximum and minimum allowable storage capacity of the energy storage system, respectively. For the first The energy stored before the time step executes the charge / discharge command. For the first The energy storage capacity before the time step executes the charge / discharge command.

[0106] Step S110: Repeat steps S108 to S109, completing the 1440-step closed-loop control simulation for the day step by step. Record the power at the grid connection point at each step. Charging and discharging power The SOC value is also output, and the minute-by-minute running curve for the day is summarized and output.

[0107] The evaluation indicators are calculated as follows:

[0108]

[0109]

[0110]

[0111] in, .

[0112] The overall backflow elimination rate is output by summarizing and statistically analyzing all dates within the evaluation period, thus verifying the backflow suppression effect of the method of the present invention.

[0113] To verify the actual control effect of the method of the present invention, simulation results of typical days were selected for analysis. The simulation data came from the historical operation records of real photovoltaic power plants, and the evaluation period covered multiple typical operating days. The simulation process was strictly executed according to the procedures described in steps S101 to S109. The initial SOC of the energy storage system was set to 50%, and the rated capacity was... Rated power The SOC constraint range is The simulation results are as follows: Figures 2 to 4 As shown.

[0114] Figure 2 This is a daily comparison curve of grid-connected power on a typical load day (April 9, 2026). The blue curve represents the original grid-connected power without energy storage regulation, the green curve represents the grid-connected power after regulation using the method of this invention (TCN-CQL), the orange dashed line represents the peak shaving threshold (30 kW), the red filled area represents the original reverse flow area, the orange filled area represents the residual reverse flow area that still exists after regulation, and the light red filled area represents the area that still exceeds the peak after regulation.

[0115] Depend on Figure 2As can be seen, the original grid-connected power exhibited a significant negative value between approximately 10:00 and 13:00, dropping to a minimum of approximately -400kW, indicating that the photovoltaic output far exceeded the local load absorption capacity during this period, resulting in severe reverse current. After regulation using the method of this invention, the energy storage system actively absorbed excess photovoltaic power during the peak reverse current period, effectively suppressing the negative shift of the grid-connected power. The number of reverse current events on this typical day decreased from 211 to 151, achieving a reverse current elimination rate of 28.4%; the number of peak-exceeding events decreased from 669 to 412, demonstrating a significant peak-shaving effect. The above results verify the effective suppression capability of the method of this invention for photovoltaic reverse current.

[0116] Figure 3 This is the daily scheduling curve of energy storage charging and discharging power output by the method of this invention (TCN-CQL). ​​The red bar area represents the charging period (power is positive, the energy storage system absorbs electrical energy), and the purple bar area represents the discharging period (power is negative, the energy storage system releases electrical energy to the load side). The red and blue dotted lines represent the upper limits of the rated charging and discharging power, respectively. ).

[0117] Depend on Figure 3 As can be seen, the method of this invention begins charging operations during the initial ramp-up phase of photovoltaic output from 00:00 to 08:00, with a charging power of approximately 20-30 kW; the charging power increases significantly during the rapid rise in photovoltaic output from 08:00 to 10:00; during the peak photovoltaic output period from 10:00 to 13:00, it switches to high-power charging, with the charging power approaching the rated upper limit of 102 kW, to absorb excess photovoltaic energy and suppress reverse current; after 13:00, it switches to an alternating mode of discharging and low-power charging, flexibly adjusting according to load demand. The total daily charging amount is 412.5 kWh, and the discharging amount is 358.6 kWh. The charging and discharging power does not exceed the rated upper limit throughout the entire process, the dispatching command is physically feasible, and the energy storage system is fully utilized.

[0118] Figure 4 This is a daily variation curve of the state of charge (SOC) of the energy storage system under the regulation of the method (TCN-CQL) of this invention. The purple curve is the minute-by-minute trajectory of SOC change, and the purple filled area intuitively shows the dynamic range of SOC. The red and blue dashed lines are the upper limit constraint (95%) and the lower limit constraint (5%) of SOC, respectively.

[0119] Depend on Figure 4As can be seen, the State of Charge (SOC) continuously increased from its initial value (approximately 5%) during charging operations, reaching a peak (approximately 90%) at approximately 08:30. It then rapidly decreased during the discharging phase, reaching a low point (approximately 25%) at approximately 10:00. After another charging cycle, it rebounded, reaching a second peak (approximately 93%) at approximately 12:30. Subsequently, it continued to decrease with discharging operations, reaching near the lower limit of SOC (5%) after 16:00. The SOC ranged from 5.0% to 95.0% throughout the day, with no instances of exceeding the limit. This indicates that the method of this invention effectively meets the safety operation constraints of the energy storage system while achieving the reverse current suppression target. It possesses good engineering practicality and safety reliability.

[0120] This invention has been verified through simulation experiments. Based on historical operating data of real photovoltaic power plants, the method of this invention (TCN-CQL) was used to perform closed-loop simulation evaluation on minute-by-minute data from January 6, 2026 to April 20, 2026, a total of 103 days, with a cumulative data processing count of 147,302 records.

[0121] Experimental results show that the total number of original backflows during the assessment period was 24,566, which was reduced to 17,472 after adjustment using the method of this invention, eliminating a total of 7,094 backflows, with an overall backflow elimination rate of 28.9%. In the 103 assessment days, there were 97 days of improvement, 3 days of stability (days with zero original backflows), and only 3 days of deterioration (February 26, March 1, and March 2, 2026, with deterioration amounts of 7, 13, and 8 respectively, all of which are low-backflow days with very few original backflows, resulting in extremely small absolute deterioration).

[0122] Furthermore, a comparative experiment was conducted between the method of this invention and a simple numerical trend prediction algorithm for backflow prevention. Using January 6, 2026 as a typical comparison date, the algorithm of this invention eliminated 186 additional backflows compared to the numerical trend prediction method. Over a 103-day overall evaluation period, the method of this invention improved the proportion of backflows eliminated by 37.1% compared to the numerical trend prediction method, verifying the significant superiority of the method of this invention in backflow suppression. Experimental results are as follows... Figures 5 to 7 As shown.

[0123] Figure 5 This is a bar chart showing the distribution of daily countercurrent reductions during the assessment period (January 6 to April 20, 2026, a total of 103 days). Green bars indicate a decrease in the number of countercurrents after the daily adjustment (improvement), and red bars indicate an increase in the number of countercurrents after the daily adjustment (deterioration). The vertical axis represents the number of countercurrent reductions (positive values ​​indicate improvement, and negative values ​​indicate deterioration).

[0124] Depend on Figure 5As can be seen, 97 out of 103 days showed improvement, accounting for approximately 94.2%, while only 3 days showed deterioration with minimal deterioration (the maximum deterioration was 13 items). Dates with the greatest improvement included: February 21 (+181 items), February 12 (+181 items), February 18 (+176 items), April 5 (+239 items), and March 11 (+185 items), indicating that the method of this invention is particularly effective under conditions of strong photovoltaic output and significant countercurrent regularity. The overall elimination rate reached 28.9%, a significant improvement compared to the previous version (5.6%), validating the significant effectiveness of the model optimization.

[0125] Figure 6 A line chart comparing the daily backflow rate (the percentage of backflows out of 1440 time steps) within the evaluation period is used. The red line represents the original backflow rate, the blue line represents the backflow rate after adjustment using the method of this invention (TCN-CQL), the green filled area represents the improved area (lower than the original after adjustment), and the orange filled area represents the deteriorated area (higher than the original after adjustment). Figure 6 As can be seen, the original average backflow ratio was 16.6%, which was reduced to approximately 11.9% after adjustment using the method of this invention, representing an average improvement of about 4.8 percentage points. The blue line was significantly lower than the red line in almost all time periods, and the area of ​​improvement in green was much larger than the area of ​​deterioration in orange, visually verifying the overall positive effect of the method of this invention. During periods with a high backflow ratio (such as mid-February when the backflow ratio was about 40%), the method of this invention could still reduce the backflow ratio to below about 30%, demonstrating the robust control capability of the method under high backflow intensity conditions.

[0126] Figure 7 This is a bar chart showing the daily distribution of backflow reductions during the assessment period (January 6 to April 20, 2026, a total of 103 days). Green bars indicate a decrease in the number of backflows after daily adjustments (improvement), red bars indicate an increase in the number of backflows after daily adjustments (deterioration), and gray bars indicate no change in the number of backflows. The vertical axis represents the number of backflow reductions (positive values ​​indicate improvement, negative values ​​indicate deterioration). Figure 7 As can be seen, the method of this invention (TCN-CQL) eliminated a net 7094 backflows during the evaluation period, with an overall backflow elimination rate of 28.9%. Of the 103 days, 97 days showed improvement, accounting for approximately 94.2%, while only 3 days showed deterioration with minimal deterioration (the maximum deterioration was 13 backflows). Dates with the greatest improvement included April 5th (+239 backflows), indicating that the method of this invention is particularly effective under conditions of strong photovoltaic output and significant backflow regularity.

[0127] Example 2.

[0128] This embodiment provides an energy storage backflow suppression system based on temporal convolutional prediction and offline reinforcement learning, including:

[0129] The data preprocessing module is used to acquire grid-connected power data and photovoltaic power data, calculate the historical net load sequence, normalize the historical net load sequence to obtain the full net load sequence, and extract the corresponding multi-dimensional time feature vector; and construct a training sample set based on the full net load sequence.

[0130] The dual-branch temporal convolution prediction module is used to construct a dual-branch temporal convolutional network. It is trained using a training sample set constructed based on the full payload sequence and is used to output the payload prediction sequence for future time periods.

[0131] The energy storage scheduling decision modeling module is used to build an energy storage scheduling decision model based on offline reinforcement learning. It takes the state vector composed of the full net load sequence, the net load prediction sequence, the current state of charge, and the multi-dimensional time feature vector as input, and the optimization objective is to minimize the backflow at the grid connection point and maximize the economic benefits.

[0132] The conservative Q-learning strategy training module is used to train the energy storage scheduling decision model on an offline dataset.

[0133] The online inference and hard constraint execution module is used to input the real-time acquired state vector into the trained energy storage scheduling decision model, output normalized action commands, scale the normalized action commands into charging and discharging power, perform hard constraint pruning on the charging and discharging power based on the current energy storage capacity, and use it to control the operation of the energy storage system; update the energy storage capacity according to the pruned charging and discharging power, and obtain the state of charge of the next time step according to the updated energy storage capacity.

[0134] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for suppressing energy storage backflow based on temporal convolutional prediction and offline reinforcement learning, characterized in that, The energy storage backflow suppression method based on temporal convolutional prediction and offline reinforcement learning includes: Obtain grid connection point power data and photovoltaic power data, and calculate the historical net load sequence; The historical net load sequences are integrated to obtain the full net load sequence, and the corresponding multidimensional time feature vectors are extracted. A training sample set is constructed based on the full net load sequence. A dual-branch temporal convolutional network is constructed and trained using a training sample set to output the net load prediction sequence for future time periods. An energy storage scheduling decision model based on offline reinforcement learning is constructed. The normalized full net load sequence, the normalized net load prediction sequence, the current state of charge and the state vector composed of multi-dimensional time feature vectors are used as inputs. The optimization objective is to minimize the backflow at the grid connection point and maximize the economic benefits. The energy storage scheduling decision model is trained on an offline dataset. The real-time collected state vector is input into the trained energy storage scheduling decision model, and the normalized action command is output. The normalized action command is scaled to the charging and discharging power. Hard constraints are applied to the charging and discharging power based on the current energy storage capacity to control the operation of the energy storage system. The energy storage capacity is updated based on the trimmed charge and discharge power, and the state of charge for the next time step is obtained based on the updated energy storage capacity.

2. The energy storage backflow suppression method based on temporal convolutional prediction and offline reinforcement learning according to claim 1, characterized in that, The construction of the training sample set based on the full net load sequence includes: Extract the input subsequence for the current time step from the full net load sequence; Centered on the current time step, extract the net load values ​​of historical reference points across cycles from the historical net load sequence to construct the context feature vector; The context feature vector is concatenated with the multidimensional time feature vector to obtain the auxiliary feature vector; Using the normalized net load value within a preset future step size as the prediction target, a training sample set for training a two-branch temporal convolutional network is constructed based on the input subsequence, auxiliary feature vector, and prediction target.

3. The energy storage backflow suppression method based on temporal convolutional prediction and offline reinforcement learning according to claim 2, characterized in that, The dual-branch temporal convolutional network includes a first residual convolutional branch and a second residual convolutional branch. The first and second residual convolutional branches are respectively composed of two layers of causal dilated convolution, GELU activation function and Dropout regularization layer connected in sequence to form the main path. The output of the main path is weighted by the channel attention mechanism and residually connected with the input features aligned with the number of channels. Then, it is batch normalized to obtain the output of the first and second residual convolutional branches. Among them, the inflation coefficients for constructing the first residual convolution branch and the second residual convolution branch are different; The outputs of the first residual convolution branch and the second residual convolution branch are compressed by adaptive average pooling and then concatenated with the auxiliary feature vector. Finally, the net load prediction sequence is output through a three-layer fully connected prediction head.

4. The energy storage backflow suppression method based on temporal convolutional prediction and offline reinforcement learning according to claim 1, characterized in that, The energy storage scheduling decision model is trained using a conservative Q-learning algorithm, which introduces a conservative regularization term into the Q-value estimation. The loss function is expressed by the following formula: in, The total loss function for conservative Q-learning, For standard Bellman error loss, The conservatism coefficient. To calculate the mathematical expectation of all states in the offline dataset, This represents the logarithmic summation of the Q values ​​over all possible actions using an exponential operation. behavioral strategies The expected Q-value of the covered actions, To generate the empirical distribution of historical scheduling strategies for offline datasets, This is the Q-value of the action.

5. The energy storage backflow suppression method based on temporal convolutional prediction and offline reinforcement learning according to claim 1, characterized in that, The reward function of the energy storage dispatch decision model is expressed by the following formula: in, For the first The reward function for each time step. This is a truncation function. Punishment for going against the tide Penalty for exceeding demand limits. Penalty for exceeding the charge state limit, Penalty for charging and discharging losses, For peak-valley arbitrage rewards, As a reward scaling factor, To truncate the upper limit.

6. The energy storage backflow suppression method based on temporal convolutional prediction and offline reinforcement learning according to claim 5, characterized in that, The reverse flow penalty is calculated based on the grid connection point power, the demand over-limit penalty is the difference between the preset demand threshold and the grid connection point power, the state of charge over-limit penalty is calculated based on the state of charge, the charging and discharging loss penalty is calculated based on the charging and discharging power, and the peak-valley arbitrage reward is 1 if charging is done during off-peak hours and 0.5 if discharging is done during peak hours.

7. The energy storage backflow suppression method based on temporal convolutional prediction and offline reinforcement learning according to claim 1, characterized in that, The hard constraint trimming of charging and discharging power based on the current energy storage capacity includes: trimming the charging and discharging power to a limit range according to the rated capacity of the energy storage system, the current energy storage capacity, and the charging and discharging efficiency.

8. An energy storage backflow suppression system based on temporal convolutional prediction and offline reinforcement learning, characterized in that, include: The data preprocessing module is used to acquire grid-connected power data and photovoltaic power data, calculate the historical net load sequence, integrate the historical net load sequence to obtain the full net load sequence, and extract the corresponding multi-dimensional time feature vector; and construct a training sample set based on the full net load sequence. The dual-branch temporal convolution prediction module is used to construct a dual-branch temporal convolutional network. It is trained using a training sample set constructed based on the full payload sequence and is used to output the payload prediction sequence for future time periods. The energy storage scheduling decision modeling module is used to build an energy storage scheduling decision model based on offline reinforcement learning. It takes the state vector composed of the full net load sequence, the net load prediction sequence, the current state of charge, and the multi-dimensional time feature vector as input, and the optimization objective is to minimize the backflow at the grid connection point and maximize the economic benefits. The conservative Q-learning strategy training module is used to train the energy storage scheduling decision model on an offline dataset. The online inference and hard constraint execution module is used to input the real-time acquired state vector into the trained energy storage scheduling decision model, output normalized action commands, scale the normalized action commands into charging and discharging power, perform hard constraint pruning on the charging and discharging power based on the current energy storage capacity, and use it to control the operation of the energy storage system; update the energy storage capacity according to the pruned charging and discharging power, and obtain the state of charge of the next time step according to the updated energy storage capacity.

Citation Information

Patent Citations

  • An energy storage demand control system and its backflow prevention method and device

    CN110098628B

  • A source-load system backflow prevention control method, device and source-load system

    CN114362249B

  • Anti-countercurrent control system and method applied to photovoltaic energy storage all-in-one machine

    CN115276100A