An economic dispatch method for optical storage system fusing uncertainty information
By initializing the storage pool network using temporal convolutional networks and sparse coding methods, and combining it with a photovoltaic power plant simulation environment, a near-end policy optimization algorithm is constructed. This algorithm solves the problems of slow policy convergence and prediction uncertainty in traditional algorithms in photovoltaic energy storage systems, and achieves efficient and reliable generation of economic dispatch policies, thereby improving the intelligence level of photovoltaic energy storage systems.
Patent Information
- Application Number
- CN202511643506.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-11-11
AI Technical Summary
In large-scale photovoltaic energy storage collaborative scheduling scenarios, traditional reinforcement learning algorithms face challenges such as slow policy convergence speed, weak generalization ability, and scheduling policy failure caused by the uncertainty of photovoltaic power prediction, resulting in high computational complexity and difficulty in achieving a balance between economy and security.
A temporal convolutional network is used for photovoltaic power prediction to generate prediction results with multiple confidence intervals. The reserve pool network is initialized by combining compressed sensing and sparse coding methods. A policy network and a value network with embedded near-end policy optimization algorithm are constructed. Through iterative training in a photovoltaic power plant simulation environment, the economically optimal day-ahead scheduling policy is output.
It improves the intelligence level of photovoltaic energy storage system scheduling under uncertain environment, enhances the perception and adaptability to photovoltaic power fluctuations, reduces the computational complexity, and achieves a multi-objective balance of maximizing operating revenue, minimizing energy storage costs and fluctuation over-limit compensation costs, thereby improving the economy and stability of scheduling.
Smart Images

Figure CN121097685B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of optical storage system scheduling, in particular to an optical storage system economic scheduling method fusing uncertainty information. BACKGROUND
[0002] With the continuous improvement of the penetration rate of photovoltaic power generation in the power system, the intermittency and randomness of its output power have become a key factor restricting the safe and stable operation of the power grid. As an important means of adjusting new energy fluctuations, achieving energy time shift and peak load shifting, the reasonable configuration and efficient scheduling of energy storage systems have core values in improving new energy consumption capacity and ensuring system economy and safety.
[0003] In the scenario of large-scale photovoltaic energy storage collaborative scheduling, the scheduling model needs to consider multi-dimensional state variables and continuous decision variables at the same time, resulting in an exponential growth of the system state action space, greatly increasing the computational complexity and optimization difficulty. Traditional reinforcement learning algorithms face challenges such as slow convergence speed of strategies and weak generalization ability in such problems, and due to the natural uncertainty of photovoltaic power prediction, the future state has strong fuzziness and unknowability, and the traditional algorithm model often fails to implement the scheduling strategy. SUMMARY
[0004] The purpose of the present application is to provide an optical storage system economic scheduling method fusing uncertainty information, aiming to improve the intelligent level of photovoltaic energy storage system scheduling in an uncertain environment.
[0005] To achieve the above purpose, the technical scheme adopted by the present application is as follows:
[0006] The present application provides an optical storage system economic scheduling method fusing uncertainty information, comprising: S1: obtaining photovoltaic array power generation power and meteorological data through a data acquisition module, and dividing them into a training set and a test set after preprocessing; wherein the data acquisition module refers to a combination of hardware devices for real-time acquisition of photovoltaic array power generation power and meteorological data, and preprocessing includes data cleaning and normalization processing; S2: training a time series convolution network using the training set, performing photovoltaic power day-ahead point prediction through the time series convolution network, and generating multi-confidence interval prediction results based on the prediction error distribution; S3: using compressed sensing and sparse coding methods, converting the interval prediction results into a sparse connection weight matrix, completing the initialization of the reservoir network and performing spectral radius normalization, and then based on the initialized reservoir network, constructing a strategy network and a value network embedded with a proximal policy optimization algorithm; S4: in a preset photovoltaic power station simulation environment, iteratively training the strategy network and the value network, and outputting a day-ahead scheduling strategy that meets the economic optimization; wherein the photovoltaic power station simulation environment refers to a simulation system containing various income models, cost models and constraint conditions.
[0007] In step S1, data is collected by voltage sensors, current sensors, temperature sensors, irradiance sensors and wind speed sensors; the Z-score standardization method is used for preprocessing to map the data measured by the sensors to the [0, 1] interval.
[0008] In step S2, the upper and lower bounds of the power prediction interval under different confidence levels are constructed by a probability statistical method based on the inverse function of the cumulative distribution function of the prediction error.
[0009] In step S3, the construction of the sparse connection weight matrix includes: based on the compressed sensing theory, an optimization model containing a robust loss function and a regularization constraint is constructed, and a sparse reconstruction matrix is solved by an iterative shrinkage threshold algorithm.
[0010] The robust loss function adopts a robust function based on a scaling factor, and the regularization constraint includes an L1 norm constraint on the sparse reconstruction matrix and a Frobenius norm constraint on the noise-induced matrix.
[0011] In step S3, spectral radius normalization refers to scaling the sparse connection weight matrix obtained by initialization according to the modulus of its maximum eigenvalue, so that the echo state characteristics of the reservoir network are met.
[0012] In step S3, the policy network outputs the action probability distribution under the current state, and the value network outputs the expected return estimate of the current state; the policy network and the value network share the reservoir network structure.
[0013] The photovoltaic power station simulation environment in step S4 includes: a VRB energy storage system operation income model, a photovoltaic power station power grid operation income model, a VRB energy storage system investment cost model, and a photovoltaic grid fluctuation overrun compensation model. The VRB energy storage system operation income model is as follows:
[0014]
[0015] wherein, is the operation income of the VRB energy storage system, is the discharge unit price of the VRB energy storage system at time t, is the charging unit price of the VRB energy storage system at time t, is the discharge power of the VRB energy storage system at time t, is the charging power of the VRB energy storage system at time t;
[0016] The VRB energy storage system investment cost model is as follows:
[0017]
[0018]
[0019]
[0020] wherein, is the average installation cost per time point of the VRB energy storage system. is the average operation and maintenance cost per time point of the VRB energy storage system, is the unit power cost of the VRB energy storage system, is the rated power of the VRB energy storage system, is the unit capacity cost of the VRB energy storage system, is the rated capacity of the VRB energy storage system, is the discount rate, is the life cycle of the VRB energy storage system, is the annual operation and maintenance cost coefficient;
[0021] The compensation model for grid-connected photovoltaic fluctuation overrun is as follows:
[0022]
[0023] wherein, is the compensation cost coefficient of grid-connected photovoltaic fluctuation overrun, is the fluctuation rate upper limit specified according to the installed capacity and the target period, is the actual fluctuation degree of the grid-connected power of the photovoltaic power plant at time t, is the part of the actual fluctuation rate of the grid-connected power of the photovoltaic power plant at time t that exceeds the grid-allowed upper limit
[0024] The constraint conditions of the photovoltaic power plant simulation environment include the charge and discharge power constraint, the capacity upper and lower limit constraint, and the photovoltaic grid-connected power fluctuation rate constraint of the VRB energy storage system. The charge and discharge power constraint of the VRB energy storage system is as follows:
[0025]
[0026] wherein, is a binary set, indicating that the ESS is in a charging or discharging state;
[0027] The capacity upper and lower limit constraint is as follows:
[0028]
[0029] wherein, and are the minimum and maximum values of the capacity, respectively;
[0030] The photovoltaic grid-connected power fluctuation rate constraint is as follows:
[0031]
[0032]
[0033] wherein, is a photovoltaic power station rated power value, and respectively represent the photovoltaic power station grid-connected power value at the target period time and time.
[0034] The optimization target of strategy training includes: maximum operation income, minimum VRB energy storage system investment cost, and minimum photovoltaic grid-connected fluctuation overrun compensation fee.
[0035] Compared with the prior art, the application has the beneficial effects that:
[0036] 1. The photovoltaic energy storage system economic dispatch method fusing uncertainty information provided by the application generates a multi-confidence interval prediction result through a time sequence convolution network, initializes a reserve pool network by combining a compressed sensing and a sparse modeling method, enhances the perception and adaptability of the system to photovoltaic power fluctuation, avoids the failure of the dispatch strategy caused by prediction error, and improves the response capability of the photovoltaic energy storage system to uncertainty. In addition, the spectral radius normalization processing of the reserve pool network ensures its dynamic stability, and the sparse connection structure reduces the calculation complexity, so that the reserve pool network can still converge efficiently in a multi-dimensional state variable and a continuous decision space, and the photovoltaic energy storage system improves the intelligent level of dispatch in an uncertain environment to a great extent, and can more efficiently and reliably obtain a matched dispatch strategy.
[0037] 2. The reinforcement learning framework constructed based on the proximal policy optimization algorithm realizes the multi-objective balance of maximum operation income, minimum energy storage cost, and minimum fluctuation overrun compensation fee, takes into account short-term income and long-term cost control, and optimizes the economy and stability of the dispatch strategy. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 is a flowchart of the photovoltaic energy storage system economic dispatch method fusing uncertainty information provided by the application;
[0039] Figure 2 is a whole framework diagram of the photovoltaic energy storage system economic dispatch method fusing uncertainty information provided by the application;
[0040] Figure 3 is an initial weight matrix construction framework diagram of the reserve pool network provided by the application;
[0041] Figure 4 is a photovoltaic energy storage economic dispatch framework diagram based on the PPO algorithm provided by the application. DETAILED DESCRIPTION
[0042] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.
[0043] With reference to Figure 1 and Figure 2 The present application provides a photovoltaic power generation economic scheduling method fusing uncertainty information, comprising:
[0044] S1: Obtain photovoltaic array power generation power and meteorological data through a data acquisition module, and divide the data into a training set and a test set after preprocessing. The data acquisition module refers to a combination of hardware devices for real-time acquisition of photovoltaic array power generation power and meteorological data, and the preprocessing includes data cleaning and normalization processing.
[0045] In step S1, data is collected through a voltage sensor, a current sensor, a temperature sensor, an irradiance sensor, and a wind speed sensor. After the sensors collect data, the data collected by the sensors is transmitted to a data collector perception network through an RS485 / 422 communication protocol, the data is monitored in real time, and time information is recorded synchronously.
[0046] In preprocessing, data cleaning is mainly used to remove noise and outliers in the original collected data, such as jump data and missing values caused by sensor failure, to avoid interference of false data on model training and ensure the accuracy and reliability of input samples, thereby providing high-quality basic data for subsequent prediction and scheduling models. The normalization processing maps photovoltaic power, meteorological data, etc. of different magnitudes to a unified numerical interval. For example, the normalization processing uses a Z-score standardization method to map the data measured by the sensor to the [0, 1] interval.
[0047] In some embodiments, the preprocessed data is divided into a training set and a test set in a ratio of 8:2. A training data matrix is constructed according to the training set, and each row of the training data matrix contains time stamp, photovoltaic power generation power, meteorological data, etc. When dividing, the time sequence of the data is ensured, and the test set does not overlap.
[0048] S2: Train a time series convolution network using the training set, perform day-ahead point prediction of photovoltaic power through the time series convolution network, and generate multi-confidence interval prediction results based on the prediction error distribution.
[0049] Step S2 aims to achieve photovoltaic power prediction and quantify uncertainty. Training the time series convolution network with the training set and generating multi-confidence interval prediction results can improve the accuracy and reliability of photovoltaic power prediction. The time series convolution network captures the long and short term dependencies of time series data through dilated convolution, which helps to reduce single point prediction error. At the same time, the multi-confidence interval prediction generated based on the error distribution breaks through the limitation of traditional point prediction which only provides a single numerical value, and quantifies the fluctuation range of future power, providing data support for dealing with uncertainty. By quantifying uncertainty, the reserve pool network built subsequently can more accurately capture the dynamic characteristics of the system, reduce the interference of extreme fluctuations on dispatch optimization, lay a data foundation for achieving economic optimal dispatch, and improve the stable operation ability of the photovoltaic storage system under complex working conditions.
[0050] The time series convolution network (TCN) performs photovoltaic power day-ahead point prediction, and the prediction value is obtained. The error is obtained by comparing and calculating the prediction value with the actual value.
[0051] In step S2, the upper and lower bounds of the power prediction interval under different confidence levels are constructed based on the inverse function of the cumulative distribution function of the prediction error through the probability and statistics method.
[0052] S3: Using compressed sensing and sparse coding method, the interval prediction result is converted into sparse connection weight matrix, after completing the initialization of reserve pool network and performing spectral radius normalization, the strategy network and value network embedded with proximal policy optimization algorithm are constructed based on the initialized reserve pool network.
[0053] The process of using compressed sensing and sparse coding method to process interval prediction results and construct network can strengthen the network's perception ability of uncertainty information. Through compressed sensing and sparse coding, the fluctuation range and probability characteristics of photovoltaic power contained in the interval prediction result are converted into a sparse connection weight matrix, so that the uncertainty information is embedded in the reserve pool network during initialization, enhancing its dynamic capture ability of photovoltaic power fluctuation, and avoiding information loss caused by traditional random initialization. Spectral radius normalization ensures that the reserve pool network meets the echo state characteristics, ensuring the stability of its dynamic response and avoiding divergence problems during training; while the sparse connection structure greatly reduces the size of network parameters, reduces the computational complexity, and enables the model to run efficiently when processing multi-dimensional scheduling variables.
[0054] For example, with reference to Figure 3 In step S3, the construction of the sparse connection weight matrix includes: based on the compressed sensing theory, an optimization model containing a robust loss function and a regularization constraint is constructed, and a sparse reconstruction matrix is solved by an iterative shrinkage threshold algorithm.
[0055] The robust loss function adopts a robust function based on a scaling factor, and the regularization constraint includes an L1 norm constraint on the sparse reconstruction matrix and a Frobenius norm constraint on the noise-induced matrix.
[0056] In step S3, the spectral radius normalization refers to scaling the initialized sparse connection weight matrix according to the modulus of the maximum eigenvalue, so that the echo state property of the reservoir network is satisfied.
[0057] In step S3, the policy network outputs an action probability distribution under the current state, and the value network outputs an expected return estimate of the current state; the policy network and the value network share the reservoir network structure.
[0058] S4: Iteratively training the policy network and the value network in a preset photovoltaic power station simulation environment to output a day-ahead scheduling strategy meeting economic optimization. The photovoltaic power station simulation environment refers to a simulation system including various revenue models, cost models and constraint conditions.
[0059] The photovoltaic power station simulation environment in step S4 includes: a VRB energy storage system operation revenue model, a photovoltaic power station power grid connection operation revenue model, a VRB energy storage system investment cost model and a photovoltaic grid connection fluctuation overrun compensation model.
[0060] For example, for constructing the VRB energy storage system operation revenue model, the VRB energy storage system operation revenue is generally related to the charging and discharging power, and can be specifically expressed as:
[0061] (1)
[0062] wherein, represents the discharging unit price of the VRB energy storage system at time t, represents the charging unit price of the VRB energy storage system at time t, represents the discharging power of the VRB energy storage system at time t, represents the charging power of the VRB energy storage system at time t. For constructing the photovoltaic power station power grid connection operation revenue model, the planned operation revenue of the photovoltaic power station is an important indicator for measuring its economic benefit, mainly derived from the electricity sales revenue of the photovoltaic grid-connected power. At time t, the economic benefit of the grid-connected photovoltaic power can be expressed by the following formula:
[0063] (2)
[0064] wherein, represents the electricity sales price to the grid, represents the photovoltaic power generation power.
[0065] For the construction of VRB energy storage system investment cost model, the investment cost of VRB energy storage system mainly includes installation and operation and maintenance cost. In the day-ahead dispatching plan, originally a day is taken as the settlement period, in order to unify the time scale of configuration and operation, and adapt to the optimization demand of each time point, the investment cost is calculated as the average investment cost of each time point according to the service life of VRB energy storage system. Specifically, it can be expressed as the following formula:
[0066] (3)
[0067] (4)
[0068] (5)
[0069] Among them, is the average installation cost of VRB energy storage system per time point. is the average operation and maintenance cost of VRB energy storage system per time point, represents the unit power cost of VRB energy storage system, represents the rated power of VRB energy storage system, represents the unit capacity cost of VRB energy storage system, represents the rated capacity of VRB energy storage system, represents the discount rate, represents the life cycle of VRB energy storage system, represents the annual operation and maintenance cost coefficient. Since VRB energy storage system relies on ion exchange to realize charging and discharging, there is no obvious aging mechanism in the battery. Therefore, the life cycle of VRB energy storage system is a constant.
[0070] For the construction of photovoltaic grid-connected fluctuation overrun compensation model, if the VRB energy storage system cannot suppress the photovoltaic grid-connected fluctuation, resulting in the photovoltaic grid-connected fluctuation exceeding the upper limit, then other units with flexible adjustment capacity must be used to adjust the excess fluctuation. Therefore, the operation cost of the system will inevitably increase, and the calculation formula of this cost can be expressed as:
[0071] (6)
[0072] Among them, represents the photovoltaic grid-connected fluctuation overrun compensation cost coefficient, is the fluctuation rate upper limit specified according to the installed capacity and target period, represents the actual fluctuation degree of photovoltaic power station grid-connected power at the dispatching time step t, that is, at time t, represents the part of photovoltaic power station grid-connected power actual fluctuation rate exceeding the upper limit allowed by the grid at the dispatching time step t.
[0073] The constraint conditions of the photovoltaic power station simulation environment include: the charge and discharge power constraint of the VRB energy storage system, the upper and lower limit constraint of the capacity, and the photovoltaic grid-connected power fluctuation rate constraint.
[0074] In the charge and discharge power constraint of the VRB energy storage system, the VRB energy storage system has limited charging power and discharging power , wherein the charging power is negative and the discharging power is positive. and respectively represent the maximum power that can be charged from the system and discharged from the system.
[0075] (7)
[0076] wherein, is a binary set, indicating that the ESS is in a charging or discharging state. When indicates a charging mode, the capacity change of the VRB energy storage system can be represented as:
[0077] (8)
[0078] When indicates a discharging mode. The capacity change of the VRB energy storage system is shown as follows:
[0079] (9)
[0080] wherein, and are the energy storage values of the energy storage unit at and moments, indicates the power loss rate, indicating the energy lost by the VRB energy storage system during charging and discharging.
[0081] For the upper and lower limit constraint of the capacity, since the VRB energy storage system has a limited capacity, in order to prevent overcharging or overdischarging of the VRB energy storage system, the specific constraint is as follows:
[0082] (10) wherein, and are the minimum and maximum values of the defined capacity.
[0083] For the photovoltaic grid-connected power fluctuation rate constraint, according to the relevant technical regulations of the state on photovoltaic grid connection, the grid-connected power fluctuation should meet the fluctuation index requirements on the premise of not affecting the normal operation of the power grid. The fluctuation rate constraint of the photovoltaic power in the target time period t can be given by the following expression:
[0084] (11)
[0085] (12)
[0086] wherein, is the rated power value of the photovoltaic power station, and respectively represent the photovoltaic power station grid-connected power value at the target period time and time.
[0087] The optimization objective of the strategy training includes: maximum operation income, minimum VRB energy storage system investment cost, and minimum photovoltaic grid-connected fluctuation overrun compensation fee.
[0088] As a possible implementation manner, the specific expression of the optimization objective is as follows:
[0089] (13)
[0090] (14)
[0091] wherein, represents the operation income of each time step, represents the limited scheduling range, are calculated by formulas (1) to (6) respectively.
[0092] The state vector and action vector of the light storage system are defined, wherein the state vector reflects the current state of the photovoltaic power station, including photovoltaic power generation power, photovoltaic grid-connected power, energy storage system capacity, grid electricity selling price, and grid-connected fluctuation rate, and can be represented by the following formula:
[0093] (15)
[0094] A remarkable feature of the photovoltaic power station grid-connected operation is that the VRB energy storage system is the only decision variable, and its charging and discharging behavior directly affects the operation state of the system. Therefore, the operation mode and charging and discharging power of the VRB energy storage system are defined as the action vector, and are specifically represented as: (16)
[0095] Based on the training set and test set divided in step S1, the TCN network is used for photovoltaic power day-ahead prediction, and the specific formula is as follows: (17)
[0096] wherein, in the formula, N is a historical feature matrix, is a TCN neural network, is the obtained photovoltaic power day-ahead prediction value.
[0097] The prediction interval is calculated by analyzing the difference between the prediction value of the TCN and the actual observation value, i.e. the prediction error. The prediction error The calculation formula is as follows:
[0098] Prediction error Quantile Obtained by estimating the inverse function of the cumulative distribution function, and the specific formula is as follows:
[0099] Wherein, is the estimated CDF of the prediction error, is the inverse function thereof.
[0100] In step S2, based on the obtained prediction error distribution, the photovoltaic power interval prediction result under different confidence levels is constructed by a probability statistical method, and the confidence interval The quantile of the error distribution is used to define:
[0101]
[0102] Wherein, ranges from 85% to 95% with a step of 5%. For each time point , the prediction interval of the actual value under different confidence is calculated by the following formula:
[0103]
[0104] Wherein, and respectively represent the upper limit and the lower limit of the prediction interval with a confidence of at time .
[0105] In order to depict the potential structural relationship in the observation data and guide the neural network to form a connection mode with information selectivity, the method provided in the embodiments of the present application introduces a sparse learning mechanism based on the compressive sensing theory. In the compressive sensing theory, the sparse representation model assumes that data can be reconstructed by a group of non-zero element coefficients. This idea can be formalized as the following optimization objective:
[0106]
[0107] Wherein, represents an interval prediction data matrix, is a sparse reconstruction matrix, is a noise-induced matrix; functions and are used to constrain and . The present application adopts a robust function , wherein is a proportional factor to enhance the robustness of the model to noise.
[0108] In the learning process of the reconstruction matrix, the design of the robust loss function is crucial. In order to ensure that the learned reconstruction matrix has sparsity, the present application applies a regularization (L1 norm) constraint to the sparse reconstruction matrix , and at the same time, applies an F norm (Frobenius norm) constraint to the noise-induced matrix , the main purpose of which is to control the global energy of the noise, thereby improving the robustness and stability of the model. The above optimization problem can be equivalently transformed into:
[0109] (23)
[0110] , wherein represents a proportional factor of the robust loss function, represents the F norm, represents the L1 norm, and the parameter is used to balance the sparsity of the representation matrix term and the robustness of the model. The objective function has both robustness and sparsity control ability, and is suitable for information compression modeling in uncertain environments.
[0111] To solve the above non-convex objective function, the present application adopts the ISTA (Iterative Shrinkage-Thresholding Algorithm) algorithm, which realizes sparse solution approximation by alternately performing gradient descent and soft thresholding operation. Finally, the sparse reconstruction matrix is obtained, which is used as the initial weight matrix of the reservoir pool network.
[0112] For step S3, in order to ensure that the reservoir pool network satisfies the echo state property (ESP), the spectral radius of the reservoir pool weight matrix is usually normalized. For example, the target spectral radius is set to , which usually satisfies . By scaling the initial reservoir pool matrix according to its spectral radius, the normalized reservoir pool connection weight matrix is obtained, and the specific form is:
[0113] (24)
[0114] , wherein represents the spectral radius of the matrix , i.e. the modulus of its largest eigenvalue. This normalization operation can effectively ensure the stability of the network dynamics and enhance its response ability to the input sequence.
[0115] For example, with reference to Figure 4 In step S3, a policy network and a value network embedded with a proximal policy optimization algorithm are constructed based on the initialized reservoir network, the policy network receives the current environment state as input. The input is injected into the reservoir network through an input weight matrix The state of the reservoir internal neurons is updated according to the following formula:
[0116] (25)
[0117] where, is a hyperbolic tangent activation function applied element by element, providing nonlinear transformation capabilities. is the previous time step reservoir state, providing necessary temporal memory. The output of the policy network, i.e., the action is obtained by the following formula:
[0118] (26)
[0119] where, is the only trainable parameter of the policy network, denoted as . To solve the optimal policy parameter , the PPO algorithm is adopted in the embodiments of the present application. The core of the algorithm is to construct a target function with a clipping mechanism, which can ensure the improvement of the policy under the premise of stable update of the policy. The target function with the clipping mechanism can be defined as:
[0120] (27)
[0121] (28)
[0122] where, is the policy network parameter before update, which is the parameter state after the last iteration or before the current update step; denotes the step size of policy network parameter update, which is used to control the parameter iteration speed; means the partial derivative vector of the policy parameter, i.e., the gradient; is the loss function of the policy network. The probability ratio denotes the probability ratio of the current policy and the old policy under the same state-action pair, which is shown as follows:
[0123] (29)
[0124] where, denotes the clipping of the probability ratio to ensure that the magnitude of policy update is controlled within the interval , so as to prevent drastic changes in policy and improve the stability of the training process. The parameter is the clipping coefficient, which is set to 0.1 in the embodiments of the present application. denotes the estimated value of the advantage function, which is used to measure the degree of advantage or disadvantage of the action selected in state relative to the average policy behavior.
[0125] For example, for constructing a value network, the value network is also constructed on the reserve pool computing framework, and the goal is to parameterize the approximation of the state value function that follows the policy, and provide a benchmark for policy optimization. The reserve pool state evolution value network receives the same environment state as input. The input is injected into the reserve pool network through an independent input weight matrix . The update mechanism of the reserve pool state of the reserve pool network is the same as that of the policy network, satisfying the following formula:
[0126] (30)
[0127] wherein is the reserve pool weight matrix that is exactly the same as the policy network, is the state of the reserve pool of the value network at the previous moment. The output of the value network is the value scalar estimate of the state , which is obtained by linear transformation on the high-dimensional dynamic features :
[0128] (31)
[0129] wherein is the only trainable parameter set of the value network, denoted as . It maps the reserve pool state to the real-valued state value estimate. The learning goal of the parameter is to minimize the mean square error between the predicted value of the value network and the target return , and the objective function is defined as follows:
[0130] (32)
[0131] (33)
[0132] wherein denotes the parameter state of the value network before the current iteration update step starts, E(⋅) is the loss function of the value network, which represents the mathematical expectation, R(⋅) is the target return, usually defined as the discounted and cumulative reward in the future, to measure the long-term benefits of the action policy in the current state. By minimizing the above loss function, the estimation ability of the value network for the state value can be effectively improved, thereby providing accurate reference for policy optimization.
[0133] The reward function is used to measure the immediate feedback signal obtained by the agent from the environment after taking a specific action, as a key guiding signal in the learning process, by reflecting the positive or negative impact of the current behavior on the task goal, prompting the agent to adjust the policy to make optimal decisions. Its mathematical form is defined as follows
[0134] (34)
[0135] wherein, R(t) represents the reward received by the agent at time step t, is a discount factor that determines the degree of influence of future rewards on current returns.
[0136] Through the policy network and value network constructed in iteration step S4, the network parameters are optimized through the loss function and reward function, and finally the policy that satisfies the economic optimization of day-ahead scheduling is obtained.
[0137] The present application is committed to improving the intelligent scheduling level of photovoltaic energy storage system in uncertain environment, breaking through the limitations of traditional methods in state representation and optimization efficiency. The embodiment of the present application proposes an economic scheduling method of photovoltaic energy storage system integrating uncertainty information. First, a photovoltaic power prediction module based on time series convolution network (TCN) is built, and interval prediction results are generated by analyzing prediction error distribution to quantify the fluctuation range of future power; then the principle of compressed sensing is introduced, a sparse connection weight matrix is constructed, and the reserve pool network is initialized with it, forming a dynamic state expression structure with uncertainty perception ability. At the same time, combined with spectral radius normalization processing, it ensures that the reserve pool network meets the echo state characteristics and improves the response performance to time series input. Further, the structure is embedded into the deep reinforcement learning framework based on proximal policy optimization (PPO), and the policy network and value network are constructed, and the economic optimal scheduling of photovoltaic energy storage system is realized through iterative optimization.
[0138] In the description of the present specification, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0139] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for economic dispatch of optical storage system fusing uncertainty information, characterized in that, The method comprises the following steps: S1: obtaining photovoltaic array power generation and meteorological data through a data acquisition module, and dividing the data into a training set and a test set after preprocessing; S2: training a time series convolution network using the training set, performing day-ahead point prediction of photovoltaic power through the time series convolution network, and generating a multi-confidence interval prediction result based on the prediction error distribution; S3: using compressed sensing and sparse coding methods, converting the interval prediction result into a sparse connection weight matrix, initializing a reservoir network after spectral radius normalization, and constructing a policy network and a value network based on the initialized reservoir network; S4: in a preset photovoltaic power station simulation environment, iteratively training the policy network and the value network to output a day-ahead scheduling strategy that meets the economic optimization; wherein the photovoltaic power station simulation environment refers to a simulation system containing various revenue models, cost models and constraint conditions. 2.The method of claim 1, wherein: In step S1, data is collected by voltage sensors, current sensors, temperature sensors, irradiance sensors and wind speed sensors; the preprocessing uses the Z-score standardization method to map the data measured by the sensors to the [0, 1] interval. 3.The method of claim 1, wherein: In step S2, the upper and lower bounds of the power prediction interval under different confidence levels are constructed based on the inverse function of the cumulative distribution function of the prediction error through a probability statistical method. 4.The method of claim 1, wherein, In step S3, the construction of the sparse connection weight matrix includes: based on the compressed sensing theory, an optimization model containing a robust loss function and a regularization constraint is constructed, and a sparse reconstruction matrix is solved by an iterative shrinkage threshold algorithm.
5. The method of claim 4, wherein: The robust loss function uses a robust function based on a scaling factor, and the regularization constraint includes an L1 norm constraint on the sparse reconstruction matrix and a Frobenius norm constraint on the noise-induced matrix. 6.The method of claim 1, wherein: In step S3, the spectral radius normalization refers to scaling the sparse connection weight matrix obtained by initialization according to the modulus of its maximum eigenvalue, so that the reservoir network satisfies the echo state property.
7. The method of claim 1, wherein: In step S3, the policy network outputs the action probability distribution under the current state, and the value network outputs the expected return estimate of the current state; the policy network and the value network share the reservoir network structure. 8.The method of claim 1, wherein, The photovoltaic power station simulation environment in step S4 includes: a VRB energy storage system operation income model, a photovoltaic power station power grid operation income model, a VRB energy storage system investment cost model, and a photovoltaic grid fluctuation overrun compensation model; the VRB energy storage system operation income model is as follows: wherein, is the income from the use of the VRB energy storage system, is the discharge unit price of the VRB energy storage system at time t, is the charge unit price of the VRB energy storage system at time t, is the discharge power of the VRB energy storage system at time t, is the charge power of the VRB energy storage system at time t; The VRB energy storage system investment cost model is as follows: wherein, is the average installation cost per time point for the VRB energy storage system; is the average operation and maintenance cost per time point for the VRB energy storage system, is the unit power cost for the VRB energy storage system, is the rated power for the VRB energy storage system, is the unit capacity cost for the VRB energy storage system, is the rated capacity for the VRB energy storage system, is the discount rate, is the life cycle of the VRB energy storage system, is the annual operation and maintenance cost coefficient; The photovoltaic grid fluctuation overrun compensation model is as follows: Wherein, is the cost coefficient of photovoltaic grid-connected fluctuation over-limit compensation, is the fluctuation upper limit specified according to the installed capacity and the target period, is the actual fluctuation degree of the photovoltaic power station grid-connected power at time t, is the part of the photovoltaic power station grid-connected power actual fluctuation rate at time t that exceeds the grid allowed upper limit. 9.The method of claim 8, wherein, The constraint conditions of the photovoltaic power station simulation environment include: the charge and discharge power constraint of the VRB energy storage system, the capacity upper and lower limit constraint, and the photovoltaic grid power fluctuation rate constraint; The charge and discharge power constraint of the VRB energy storage system is as follows: wherein is a binary set indicating that the ESS is in a charging or discharging state; The capacity upper and lower limit constraint is as follows: wherein and are the minimum and maximum values of the capacity, respectively; The photovoltaic grid power fluctuation rate constraint is as follows: wherein, is a photovoltaic power station rated power value, and respectively represent the photovoltaic power station grid-connected power value at the time and in the target period, is a photovoltaic power generation power.
10. The method of claim 1, wherein The optimization objectives of the strategy training include: maximizing the operation income, minimizing the VRB energy storage system investment cost, and minimizing the photovoltaic grid fluctuation overrun compensation cost.
Citation Information
Patent Citations
Wireless sensor network-based distributed energy microgrid power control system
CN108365638A
Distributed photovoltaic information fusion and standardized data model construction method
CN117632921A