Optical storage inverter system optimization control method based on RBMO-LSTM photovoltaic prediction method

By introducing RBMO algorithm to optimize hyperparameters in LSTM neural network, the problem of cumbersome hyperparameter adjustment in traditional LSTM in photovoltaic power prediction is solved, and more accurate photovoltaic power prediction and better grid stability are achieved.

CN120109765AInactive Publication Date: 2025-06-06SUZHOU DEVOWER ENERGY EQUIP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411334322.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional LSTM neural networks have cumbersome and inefficient hyperparameter adjustment in photovoltaic power prediction, making it difficult to adapt to dynamically changing data and tasks, resulting in photovoltaic power fluctuations having a great impact on grid voltage and frequency stability.

Method used

The photovoltaic prediction method based on RBMO-LSTM is adopted to optimize the hyperparameters of LSTM through the RBMO algorithm, including learning rate, number of hidden layer neurons and L2 regularization coefficient, to improve the accuracy of photovoltaic power prediction.

Benefits of technology

The hyperparameter configuration effect of the LSTM neural network is improved, making the photovoltaic power prediction more accurate, can better adapt to dynamically changing data and tasks, and reduce the instability impact on the grid voltage and frequency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120109765A_ABST
    Figure CN120109765A_ABST
Patent Text Reader

Abstract

The invention discloses an optical storage inverter system optimization control method based on an RBMO-LSTM photovoltaic prediction method. An RBMO algorithm is adopted to optimize various hyper-parameters of LSTM. In the algorithm, a position variable of an RBMO group represents an LSTM hyper-parameter which needs to be optimized. In each iteration, a prediction value of a new hyper-parameter and RMSE values of a prediction result and an actual value are calculated. The degree of fit is defined as the RMSE value of the prediction result. And performing iterative optimization for multiple times so as to determine an optimal hyper-parameter value. According to the method, the value of the hyper-parameter does not need to be manually adjusted, after the training set is obtained, the range of the hyper-parameter is determined, the hyper-parameter can be optimized and selected, the more the number of iterations is, the higher the selection optimization reliability of the hyper-parameter is, and the selection of the hyper-parameter is more intelligent. The configuration effect of the LSTM neural network is effectively improved, the photovoltaic power fitted by the LSTM neural network is more accurate, and the effect of the method is better when meteorological data is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an optimization control method for a photovoltaic storage inverter system based on a RBMO-LSTM photovoltaic prediction method, belongs to the field of artificial intelligence and power electronics technology, and in particular relates to an optimization control method for a photovoltaic storage inverter system based on a RBMO-LSTM photovoltaic prediction method with more accurate fitted photovoltaic power. Background Art

[0002] Photovoltaic power forecasts are divided into four types according to the forecast time scale: statistical methods, physical or meteorological methods, hybrid methods and other innovative methods. Statistical models do not need to analyze the inherent characteristics of the weather system itself and are more like data-driven methods. These forecasts are more dependent on the amount of input data and the accuracy of the data. Generally, artificial intelligence forecasting methods are classified as statistical methods, which may be different in different classification systems.

[0003] Hyperparameters have an important impact on the performance of photovoltaic power generation prediction in LSTM neural networks. The main hyperparameters are: number of hidden layer units, learning rate, batch size, number of layers, time step, regularization, and parameter activation function. Insufficient traditional hyperparameter adjustment. Traditional LSTM neural networks often use empirical methods to manually adjust hyperparameters during training. Manually adjusting hyperparameters requires experience and a large number of experiments. The process is cumbersome and inefficient, consuming a lot of time and computing resources, and it is difficult to find the optimal hyperparameter combination. Later, some scholars proposed a random search method: randomly selecting parameter combinations within a preset parameter range for search. Although this method is more efficient than grid search, the probability of finding the optimal parameters is still low and requires a large number of experiments. In addition, these adjustment methods lack intelligent adjustment mechanisms: lack of intelligent exploration mechanisms for hyperparameter space, it is difficult to use historical test results for guidance, and it is impossible to fully utilize known information to optimize the search process, and the overall efficiency is low. It is difficult to adapt to dynamic changes. Traditional methods have good tuning effects on fixed data sets and tasks, but have poor adaptability to dynamically changing data and tasks. In practical applications, it is difficult to quickly adjust hyperparameters to cope with new data and task changes.

[0004] The power generation equipment of small and medium-sized power grids is small in scale and relies on small generator sets and distributed power generation, such as wind power and photovoltaics. Most of these distributed power sources do not have the rotating mass of traditional synchronous generators and contribute less to the system inertia. Distributed power sources (such as photovoltaic power generation) lack mechanical inertia and can only be connected to the grid through power electronic devices. At present, grid-connected inverters for distributed power generation rarely use grid-type control, and these devices cannot provide system inertia. There is a direct relationship between voltage and power in the power grid. When the power of photovoltaic power generation fluctuates, it will cause voltage fluctuations at the grid connection point. Especially in small and medium-sized regional power grids, the load and power supply of the power grid are relatively small, and the power fluctuation of photovoltaic power generation has a more significant impact on the overall grid voltage.

[0005] The grid frequency reflects the balance between power generation and load. When the photovoltaic power generation power suddenly increases or decreases, this balance will be broken, causing grid frequency fluctuations. In small and medium-sized regional power grids, the system inertia is small, and power fluctuations are more likely to cause frequency instability.

[0006] The traditional LSTM neural network structure is relatively inflexible, and in actual application scenarios, the data is complex and high-dimensional. In the process of training and predicting complex data, some problems will arise, including slow learning speed and tendency to converge to local extreme values. These problems are solved by adjusting hyperparameters. However, the application of a single hyperparameter may not be sufficient to solve all the problems of LSTM, which requires a multi-iteration, dynamic and long-term hyperparameter optimization process. Determining the optimal hyperparameter configuration remains a major challenge. For the above reasons, current large-scale photovoltaic power stations all require the addition of a certain proportion of energy storage. However, due to the uncertainty of photovoltaic output, when meteorological conditions such as light intensity and temperature change rapidly, energy storage cannot respond quickly to changes in photovoltaic power, which will still have a certain degree of impact on the power grid. Summary of the invention

[0007] The technical problem to be solved by the present invention is to provide a photovoltaic storage inverter system optimization control method based on the RBMO-LSTM photovoltaic prediction method, which has the characteristic of more accurate photovoltaic power fitting by the LSTM neural network.

[0008] In order to solve the above technical problems, the technical solution of the present invention is: a photovoltaic storage inverter system optimization control method based on the RBMO-LSTM photovoltaic prediction method, and its innovation lies in: the photovoltaic storage inverter system optimization control method based on the RBMO-LSTM photovoltaic prediction method comprises the following steps:

[0009] Step S1: RBMO algorithm modeling

[0010] The prediction method matrix for photovoltaics is:

[0011]

[0012] The variable n represents the number selected when the RBMO algorithm is initialized, and dim represents the dimension of the problem variable. From the matrix, we can see that the multidimensional coordinates of each row vector of the matrix represent the position of the variable in the spatial hypercube; the upper and lower limits of the variable position are:

[0013] x i,j =(u b -l b )·Rand 1 +l b (2)

[0014] ub and l b are the upper and lower limits of the problem variables, representing the constraints of the variables in practical problems;

[0015] All of them are random numbers; the algorithm simulates the process of variables looking for the location of variables:

[0016]

[0017] Where t is the number of iterations, X i (t+1) is the position of the i-th new search, p represents the number of small groups participating in the variable, X m (t) is the mth individual randomly selected, X i (t) represents the i-th individual, X rs (t) is a randomly selected individual, q is the number of small groups of another size, the number of p is between 2-5, the number of q is between 10 and n, and formula (4) represents a large population;

[0018] The following expression simulates the process of variable attacking the position of the variable

[0019]

[0020] Where X food (t) is the position of the variable,

[0021] The following formula expresses the process of variable position storage, which is used to simulate the fitness function.

[0022] The objective function of the problem

[0023]

[0024] fitness old i and fitness new i are the old and new fitness functions, and the meaning of the above formula (7) is to update the position of the predicted photovoltaic power generation variable when the effect of the fitness function is better than the old value;

[0025] Step S2: Obtain the optimal hyperparameters of LSTM based on the RBMO algorithm

[0026] The acquisition of the RMSE value of the predicted data and the actual data, that is, the fitness function, the fitness function

[0027] for

[0028]

[0029] Where N is the number of samples, xf is the predicted value, x a is the true value;

[0030] After obtaining the RMSE value, the position of the predicted photovoltaic power generation variable is updated according to formula (7) and the next iteration is performed; before the algorithm runs, the number of iterations is set to m times, and the algorithm stops running after the number of iterations reaches m. The position of the predicted photovoltaic power generation variable obtained at this time is the optimal LSTM hyperparameter.

[0031] After the algorithm iteration is completed, the newly generated LSTM with the optimal hyperparameters will be used to predict the PV power generation data, and the hyperparameter combination that shows the best fit will be retained and subsequently used to predict the PV power generation;

[0032] Preferably, after obtaining the predicted photovoltaic power generation, a graph is drawn for comparison with the actual photovoltaic power, and subsequent optimization control is continued. The subsequent optimization control is LSTM training. When performing LSTM training, the LSTM neural network is a neural network with multi-dimensional input and single-dimensional output. The input is specific weather data, and the weather data of the photovoltaic inverter application location is extracted. The output of the LSTM neural network is the photovoltaic power value.

[0033] Preferably, the weather data includes temperature, humidity, surface wind speed, surface horizontal radiation intensity, surface direct normal radiation intensity and scattered radiation intensity.

[0034] Preferably, the LSTM training includes forward propagation, calculation of loss function, back propagation and gradient calculation, and updating weights according to gradient.

[0035] Preferably, the forward propagation is the process by which the LSTM network processes the input data and generates a predicted output. The input data is processed according to the formula X = [x 1 ,x 2 ,...,x t ] to calculate, where each x i They are all vectors containing multiple meteorological features;

[0036] The loss function is calculated using the mean square error MSE as the loss function, assuming that there are N sample prediction values The true value is y t , then the loss function is defined as:

[0037]

[0038] MSE measures the error between the predicted value and the true value, and the goal is to minimize this error.

[0039] Preferably, the back propagation and gradient calculation are key algorithms for calculating gradients in neural networks. Through back propagation, the network calculates gradients layer by layer from the output layer forward, and passes the error back, thereby updating the weights and biases of each layer.

[0040] Preferably, in the updating of weights according to the gradient, mathematically, the gradient is a partial derivative vector of a multivariate function, which indicates the rate at which the function changes in each direction; in a neural network, the gradient indicates the partial derivative of the loss function with respect to the network parameters; the gradient can be expressed as:

[0041]

[0042] Where w is the weight, which is the parameter connecting different layers in the neural network; the input of each neuron is the weighted sum of the output of the previous layer of neurons and the corresponding weight;

[0043] This gradient value reveals the direction and magnitude of the change in the loss function; in a neural network, the loss function measures the gap between the model's predicted output and the actual label; in the mean square error MSE, the predicted value for N samples is The true value is y i , then the loss function can be defined as:

[0044]

[0045] To find the minimum value of the loss function, it is necessary to calculate its partial derivative with respect to the network parameters, i.e. the gradient;

[0046] Gradient descent repeatedly adjusts parameters to gradually reduce the loss function and find its minimum value;

[0047] In gradient descent, the parameter update formula is:

[0048]

[0049] Among them, ω is the weight that needs to be updated, η is the learning rate, which controls the amplitude of each parameter update. is the gradient of the loss function with respect to ω.

[0050] Preferably, the LSTM neural network is a type of recursive neural network RNN, comprising a storage unit and a gate mechanism, wherein the gate mechanism comprises obtaining a forget gate function, obtaining an input gate function, obtaining an output gate function and obtaining an activation function sigmoid,

[0051] The formula of the forget gate function is: t =σ(W f ·[h t-1 ,x t ]+b f )(9)

[0052] In the formula, f t is the output of the forget gate:

[0053] f t is the output of the forget gate at time step t, which is a vector with the same dimension as the hidden state of the LSTM unit. t The value range is between [0,1], which determines the cell state C at the previous time step. t-1 How much information will be retained and passed to the cell state C at the current time step t t ;

[0054] When f t When an element of is close to 1, it means that the corresponding information will be retained; when it is close to 0, it means that the corresponding information will be forgotten;

[0055] Where σ is the Sigmoid activation function: its specific expression is Map the input value to the range of [0,1]. Since the output of the forget gate needs to be between [0,1], use the Sigmoid function to process the result of the linear combination;

[0056] W f is the weight matrix of the forget gate, which contains the input x of the current time step t and the hidden state h at the previous time step t-1 The relevant weight, W f The dimension is (n hidden ,n hidden +n input ), where n hidden It is LSTM

[0057] The dimension of the hidden state of the unit, n input is the input vector x t The dimension of this matrix determines the input x at the current time step. t and the hidden state h at the previous time step t-1 Impact on the output of the forget gate;

[0058] [h t-1 ,x t ]The concatenation of the hidden state and the input, [h t-1 ,x t ] is the hidden state h of the previous time step t-1 and the input x at the current time step t The concatenation vector, h t-1 is the hidden state vector of the previous time step, with dimension n hidden , which contains the information of the previous time step, x t is the input vector of the current time step, with dimension ninput , which contains the input data for the current time step, b f is the bias vector of the forget gate, dimension is n hidden , the bias vector is used to adjust the baseline value of the forget gate output to help the model fit the data;

[0059] According to the function of the forget gate, when the function outputs 1, the transmitted information is completely retained, and when the function value is 0, the transmitted value is 0 and the information is completely "forgotten";

[0060] The formula for the input gate function is:

[0061] i t =σ(W i ·[h t-1 ,x t ]+b i ) (10)

[0062]

[0063] In the formula, i t : Output vector of the input gate, dimension is n hidden ,i t Determines the input x of the current time step t and the hidden state h at the previous time step t-1 For unit status C t The impact of t The value of is between [0,1]. The closer the value is to 1, the more the information should be written into the unit state, and the closer the value is to 0, the more the information should be ignored.

[0064] Where σ is the Sigmoid activation function: its specific expression is Map the input value to the range of [0,1]. Since the output needs to be between [0,1], use the Sigmoid function to process the result of the linear combination.

[0065] W i : The weight matrix of the input gate, dimension is (n hidden ,n hidden +n input ), which controls the hidden state h of the previous time step t-1 and the input x at the current time step t The weight in the input gate;

[0066] [h t-1 ,x t ]: hidden state h t-1 and the input vector x t The concatenation vector of t-1 is the hidden state vector of the previous time step, with dimension nhidden , x t is the input vector of the current time step, with dimension n input ;

[0067] b i : Bias vector of input gate, dimension is n hidden ; The bias term adjusts the activation value of the input gate to help the network fit the data;

[0068] The input gate normalizes the new input information to a value between -1 and 1 through the hyperbolic tangent function and converts the previous f t and i t Merge into new C t Numeric value;

[0069] The formula of the output gate function is:

[0070] o t =σ(W o [h t-1 ,x t ]+b o ) (12)

[0071] h t =o t tanh(C t ) (13)

[0072] o t : Output vector of the output gate, dimension is n hidden ;o t Determines the cell state C at the current time step t

[0073] How much information needs to be output to the hidden state h t ;o t The value of is between [0,1]. The closer the value is to 1, the more the information should be output, and the closer it is to 0, the more the information should be suppressed.

[0074] W o : The weight matrix of the output gate, dimension is (n hidden ,n hidden +n input ), which controls the hidden state h of the previous time step t-1 and the input x at the current time step t The weight in the output gate;

[0075] [h t-1 ,x t ]: hidden state h t-1 and the input vector x t The concatenation vector of t-1is the hidden state vector of the previous time step, with dimension n hidden , x t is the input vector of the current time step, with dimension n input ;

[0076] b o : Bias vector of the output gate, dimension is n hidden , the bias term adjusts the activation value of the output gate to help the network fit the data;

[0077] The output gate represents the output component of the LSTM neuron;

[0078] The formula for the activation function is:

[0079]

[0080] Where σ is the Sigmoid activation function: its specific expression is Map the input value to the range of [0,1]. Since the output needs to be between [0,1], use the Sigmoid function to process the result of the linear combination.

[0081] The output is obtained through an activation function sigmoid, and then multiplied by the normalized C function tanh. t value.

[0082] Preferably, the LSTM hyperparameters include learning rate, number of hidden layer neurons, number of layers, and L2 regularization coefficient;

[0083] In the gradient descent algorithm, the learning rate is the coefficient used to reduce the gradient and determines the gradient loss rate at each step;

[0084]

[0085] Symbol θ t Represents the current parameters, and the gradient of the loss function relative to the parameters is expressed as Represents, and the learning rate is represented by η;

[0086] In the LSTM neural network, an additional regularization term is added to the loss function, which is L2 regularization. L2 regularization prevents the model from overfitting;

[0087] The number of neurons in the hidden layer determines the capacity and expressiveness of the LSTM network structure.

[0088] The advantages of the present invention are: the present invention adopts the RBMO algorithm to optimize various hyperparameters of LSTM, and the position variables of the RBMO group in the algorithm represent the LSTM hyperparameters that need to be optimized. In each iteration, the predicted value of the new hyperparameter and the RMSE value of the predicted result and the actual value are calculated. The fitness is defined as the RMSE value of the predicted result. In this way, multiple iterative optimizations are performed to determine the optimal hyperparameter value.

[0089] Compared with the existing method for determining hyperparameters in LSTM optimization technology, the advantages of the present invention are:

[0090] 1. There is no need to manually adjust the value of hyperparameters. Traditional methods use empirical methods and trial and error methods, and there is no definite adjustment direction. A large number of solutions with poor fitness will be obtained, and the time and computational costs are huge. The method proposed in the present invention determines the range of hyperparameters after obtaining the training set, and then optimizes the hyperparameters. The more iterations, the higher the reliability of the optimization of the hyperparameter selection, and the more intelligent the hyperparameter selection is.

[0091] 2. The configuration effect of the LSTM neural network is effectively improved, making the photovoltaic power fitted by the LSTM neural network more accurate. The more accurate the meteorological data, the better the effect of this method. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0093] Figure 1 It is a flowchart of the RBMO algorithm in the photovoltaic storage inverter system optimization control method based on the RBMO-LSTM photovoltaic prediction method of the present invention.

[0094] Figure 2 It is a flowchart of the RBMO-LSTM algorithm in a photovoltaic storage inverter system optimization control method based on the RBMO-LSTM photovoltaic prediction method of the present invention.

[0095] Figure 3 It is a schematic diagram of the LSTM neural network structure in a photovoltaic storage inverter system optimization control method based on the RBMO-LSTM photovoltaic prediction method of the present invention.

[0096] Figure 4 It is a prediction effect diagram of a training set in a photovoltaic storage inverter system optimization control method based on the RBMO-LSTM photovoltaic prediction method of the present invention.

[0097] Figure 5 It is a prediction effect diagram of a test set in a photovoltaic storage inverter system optimization control method based on the RBMO-LSTM photovoltaic prediction method of the present invention. DETAILED DESCRIPTION

[0098] The RBMO-LSTM algorithm proposed in the present invention is an improved algorithm based on the traditional LSTM neural network time prediction, and an efficient and effective calculation method needs to be developed. This application proposes a new algorithm RBMO-LSTM based on RBMO. The basic premise of the algorithm is to use RBMO to find the optimal hyperparameters of LSTM.

[0099] 1.1. Brief description of variable algorithm (RBMO)

[0100] Variables are known for their striking appearance and adaptable foraging ability. The basis of the RBMO algorithm is the principle of variable behavior of animals. The algorithm is briefly described below. The prediction method matrix of PV is:

[0101]

[0102] The variable n is the number selected when the RBMO algorithm is initialized. The larger the number, the greater the amount of calculation. Generally, 10-20 is selected. dim represents the dimension of the problem variable. It can be seen from the matrix that the multidimensional coordinates of each row vector of the matrix represent the position of the variable in the spatial hypercube, representing several variables selected by the present invention that will affect the photovoltaic power generation power: the weather data extracted by the present invention include temperature, humidity, surface wind speed, surface horizontal irradiation intensity, surface direct normal irradiation intensity and scattered radiation intensity; the upper and lower limits of the position of the variable are:

[0103] x i,j =(u b -l b )·Rand 1 +l b (2)

[0104] u b and l b are the upper and lower limits of the problem variables, representing the constraints of the variables in practical problems.

[0105] The numbers represented are all random numbers. This algorithm simulates the process of variables looking for the position of variables:

[0106]

[0107] Where t is the number of iterations, X i (t+1) is the position of the i-th new search, p represents the number of small groups participating in the variable, X m (t) is the mth individual randomly selected, X i (t) represents the i-th individual, X rs(t) is a randomly selected individual, q is the number of small groups of another size, the number of p is between 2-5, the number of q is between 10 and n, and formula (4) represents the large population. The following expression simulates the process of the variable attacking the position of the variable

[0108]

[0109] Where X food (t) is the position of the variable,

[0110] The following formula expresses the process of variable position storage, which is used to simulate the fitness function, which is the objective function of the optimization problem:

[0111]

[0112] fitness old i and fitness new i are the old and new fitness functions. The meaning of the above formula (7) is to update the position of the variable when the effect of the fitness function is better than the old value.

[0113] 1.2 RBMO-LSTM prediction method

[0114] The RBMO-LSTM photovoltaic power output prediction method proposed in the present invention optimizes the hyperparameters of the L2 regularization coefficient, the initial learning rate, and the number of hidden layer nodes. These three variables are used as the position variables of the population and change with the variable process in each iteration. The variables move in the hypercube according to formulas (3)-(6). The fitness function of this algorithm is the RMSE value of the predicted data and the actual data, and the calculation formula is:

[0115]

[0116] Where N is the number of samples, x f is the predicted value, x a is the true value;

[0117] After obtaining the RMSE value, the position in the population is updated according to formula (7) and the next iteration is performed. Before the algorithm is run, the number of iterations is set to m. When the number of iterations reaches m, the algorithm stops running. The position of the variable population variable obtained at this time is the optimal LSTM hyperparameter.

[0118] After the algorithm iteration is completed, the newly generated LSTM with the optimal hyperparameters will be used to predict the photovoltaic power generation data. The RMSE (root mean square error) of the data is used as the fitting function. After the iteration is completed, the hyperparameter combination that shows the best fit will be retained and the photovoltaic power generation will be predicted subsequently. After the prediction results are obtained, they are compared with the actual photovoltaic power and plotted, and subsequent optimization control is continued.

[0119] When LSTM training is performed, the network is a neural network with multi-dimensional input and single-dimensional output. The input is specific weather data. The algorithm needs to extract weather data at the location where the photovoltaic inverter is applied. The weather data extracted by the present invention includes temperature, humidity, surface wind speed, surface horizontal irradiation intensity, surface direct normal irradiation intensity and scattered radiation intensity. Historical weather data can be obtained from NASA, the European Meteorological Center, etc., and future weather data can be obtained from meteorological station forecasts. The output of the neural network is the photovoltaic power value.

[0120] 2. LSTM neural network prediction principle

[0121] The prediction principle of the LSTM neural network described in this application corresponds to the use of optimized LSTM in the process

[0122] The principle and steps of the network for photovoltaic power prediction. The algorithm proposed in the present invention is essentially to use the variable algorithm to continuously adjust the hyperparameters of the LSTM neural network, and then iterate the training and prediction after adjusting the hyperparameters. After the iteration is completed, the neural network is configured with the hyperparameters with the best effect, and then the final prediction is performed to obtain the final result.

[0123] 2.1 LSTM topology and neuron propagation principle

[0124] LSTM is a special type of recurrent neural network (RNN) that uses specialized memory cells and gate mechanisms. Its purpose is to solve the gradient vanishing and gradient exploding problems common in RNNs when processing long time series data.

[0125] The formula of the forget gate is:

[0126] f t =σ(W f ·[h t-1 ,x t ]+b f ) (9)

[0127] In the formula, f t is the output of the forget gate:

[0128] f t is the output of the forget gate at time step t, which is a vector with the same dimension as the hidden state of the LSTM unit. tThe value range is between [0,1], which determines the cell state C at the previous time step. t-1 How much information will be retained and passed to the cell state C at the current time step t t ;

[0129] When f t When an element of is close to 1, it means that the corresponding information will be retained; when it is close to 0, it means that the corresponding information will be forgotten;

[0130] Where σ is the Sigmoid activation function: its specific expression is Map the input value to the range of [0,1]. Since the output of the forget gate needs to be between [0,1], use the Sigmoid function to process the result of the linear combination;

[0131] W f is the weight matrix of the forget gate, which contains the input x of the current time step t and the hidden state h at the previous time step t-1 The relevant weight, W f The dimension is (n hidden ,n hidden +n input ), where n hidden It is LSTM

[0132] The dimension of the hidden state of the unit, n input is the input vector x t The dimension of this matrix determines the input x at the current time step. t and the hidden state h at the previous time step t-1 Impact on the output of the forget gate;

[0133] [h t-1 ,x t ]The concatenation of the hidden state and the input, [h t-1 ,x t ] is the hidden state h of the previous time step t-1 and the input x at the current time step t The concatenation vector, h t-1 is the hidden state vector of the previous time step, with dimension n hidden , which contains the information of the previous time step, x t is the input vector of the current time step, with dimension n input , which contains the input data for the current time step, b f is the bias vector of the forget gate, dimension is n hidden , the bias vector is used to adjust the baseline value of the forget gate output to help the model fit the data;

[0134] According to the function of the forget gate, when the function outputs 1, the transmitted information is completely retained, and when the function value is 0, the transmitted value is 0, the information is completely "forgotten", and the value retained in (0,1) is this The mechanism of the forget gate.

[0135] The formula for the input gate is:

[0136] i t =σ(W i ·[h t-1 ,x t ]+b i ) (10)

[0137]

[0138] In the formula, i t : Output vector of the input gate, dimension is n hidden ,i t Determines the input x of the current time step t and the hidden state h at the previous time step t-1 For unit status C t The impact of t The value of is between [0,1]. The closer the value is to 1, the more the information should be written into the unit state, and the closer the value is to 0, the more the information should be ignored.

[0139] Where σ is the Sigmoid activation function: its specific expression is Map the input value to the range of [0,1]. Since the output needs to be between [0,1], use the Sigmoid function to process the result of the linear combination.

[0140] W i : The weight matrix of the input gate, dimension is (n hidden ,n hidden +n input ), which controls the hidden state h of the previous time step t-1 and the input x at the current time step t The weight in the input gate;

[0141] [h t-1 ,x t ]: hidden state h t-1 and the input vector x t The concatenation vector of t-1 is the hidden state vector of the previous time step, with dimension n hidden , x t is the input vector of the current time step, with dimension n input ;

[0142] b i: Bias vector of input gate, dimension is n hidden ; The bias term adjusts the activation value of the input gate to help the network fit the data;

[0143] The input gate normalizes the new input information to a value between -1 and 1 through the hyperbolic tangent function and converts the previous f t and i t Merge into new C t Numeric value.

[0144] The output gate function is:

[0145] o t =σ(W o [h t-1 ,x t ]+b o ) (12)

[0146] h t =o t tanh(C t ) (13)

[0147] o t : Output vector of the output gate, dimension is n hidden ;o t Determines the cell state C at the current time step t

[0148] How much information needs to be output to the hidden state h t ;o t The value of is between [0,1]. The closer the value is to 1, the more the information should be output, and the closer it is to 0, the more the information should be suppressed.

[0149] W o : The weight matrix of the output gate, dimension is (n hidden ,n hidden +n input ), which controls the hidden state h of the previous time step t-1 and the input x at the current time step t The weight in the output gate;

[0150] [h t-1 ,x t ]: hidden state h t-1 and the input vector x t The concatenation vector of t-1 is the hidden state vector of the previous time step, with dimension n hidden , x t is the input vector of the current time step, with dimension n input ;

[0151] b o: Bias vector of the output gate, dimension is n hidden , the bias term adjusts the activation value of the output gate to help the network fit the data;

[0152] The output gate represents the output component of the LSTM neuron. The output is obtained through a sigmoid function and then multiplied by the C normalized by the tanh function. t value.

[0153] The above activation function is:

[0154]

[0155] Where σ is the Sigmoid activation function: its specific expression is Map the input value to the range of [0,1]. Since the output needs to be between [0,1], use the Sigmoid function to process the result of the linear combination.

[0156] 2.2 LSTM Hyperparameter Selection

[0157] After selecting the network, one of the key steps before starting the training process is to determine the hyperparameters of the neural network. The hyperparameters of the LSTM network include a series of factors, including learning rate, number of neurons in the hidden layer, number of layers, choice of activation function, regularization coefficient, etc. In this chapter, the main hyperparameters studied are learning rate, L2 regularization coefficient, and number of neurons in the hidden layer. The rest of this section will clarify the functions of these hyperparameters.

[0158] In the gradient descent algorithm, the learning rate can be defined as the coefficient used to reduce the gradient, thereby determining the gradient loss rate at each step.

[0159]

[0160] Symbol θ t Represents the current parameters, and the gradient of the loss function relative to the parameters is expressed as denoted by η, and the learning rate is η. The learning rate directly affects the training speed of the model. The higher the learning rate, the faster the parameters are updated, thus accelerating the convergence speed. However, a learning rate that is too high may cause the model to oscillate or diverge during training and fail to converge to the global optimal solution or the best solution. On the contrary, a lower learning rate may lead to more stable convergence, but may lead to slow convergence and increase training time.

[0161] In neural networks, an additional regularization term is added to the loss function, called L1 or L2, also known as L1 regularization or L2 regularization. L1 regularization refers to the sum of the absolute values ​​of each element in the weight vector w. L2 regularization can prevent the model from overfitting.

[0162] The number of nodes in the hidden layer determines the capacity and expressiveness of the LSTM network. Adding additional hidden units can capture more complex and subtle feature information, thereby improving the network's ability to represent data. In the case of complex tasks or large data sets, increasing the number of hidden layer nodes often improves the performance of the model.

[0163] Increasing the number of hidden layer nodes significantly increases the computational complexity and storage requirements of the model. Each hidden unit is responsible for calculating the state of the input gate, forget gate, output gate, and storage unit. The amount of these calculations is proportional to the number of hidden units. In addition to the increased demand for computing resources and memory, the increase in the number of hidden units will also lead to an increase in training and inference time.

[0164] 2.3. LSTM neural network prediction steps

[0165] Before this step, it is assumed that the data has been cleaned, normalized, smoothed, and partitioned, and all training data has been divided into training set, validation set, and test set. After data preparation is completed, the structure of the LSTM network needs to be designed. A typical LSTM network consists of the following layers: Input layer:

[0166] Receive meteorological characteristic data. LSTM layer: contains multiple hidden units to capture the characteristics of time series.

[0167] Fully connected layer: used to map the output of the LSTM layer to the target output. Output layer: used to generate the final predicted value (PV power).

[0168] The following uses a neuron as an example to illustrate the training and prediction steps of the LSTM neural network.

[0169] 2.3.1 Forward Propagation

[0170] Forward propagation is the process by which the LSTM network processes input data and generates predicted output.

[0171] The unit contains a forget gate, an input gate, an output gate, and a unit state. The detailed process is shown in 2.1. The input data will be calculated according to the formula. Assume that the input weather data is X = [x 1 ,x 2 ,...,x t ], where each x i are all vectors containing multiple meteorological features. The specific calculation is:

[0172] Forget gate calculation: determines how much information in the cell state of the previous time step needs to be retained. The formula for the forget gate is:

[0173] f t =σ(Wf ·[h t-1 ,x t ]+b f )

[0174] The formula for the input gate is:

[0175] i t =σ(W i ·[h t-1 ,x t ]+b i ) (10)

[0176]

[0177] Update unit status: unit status C t Combined with the state C of the previous time step t-1 and the candidate state of the current time step

[0178]

[0179] Here, f t ·C t-1 is the retained part of the previous state, This is newly added information.

[0180] Calculation of output gate and hidden state:

[0181] The output gate function is:

[0182] o t =σ(W o [h t-1 ,x t ]+b o ) (12)

[0183] The hidden state is calculated as:

[0184] h t =o t tanh(C t ) (13)

[0185] 2.3.2 Calculation of loss function

[0186] For regression tasks, the mean square error (MSE) is used as the loss function. Suppose we have N sample predictions The true value is y t , then the loss function is defined as:

[0187]

[0188] MSE measures the error between the predicted value and the true value, and the goal is to minimize this error.

[0189] 2.3.3 Back Propagation and Gradient Calculation

[0190] After the forward propagation is completed, we calculate the value of the loss function. Next, the gradient of each time step is calculated through back propagation (BPTT). Backpropagation is the key algorithm for calculating gradients in neural networks. Through backpropagation, the network calculates the gradient layer by layer from the output layer, passes the error back, and updates the weights and biases of each layer. The key to backpropagation is to use the chain rule to propagate the gradient layer by layer through the dependencies between time steps.

[0191] Assume that the loss function L is the output layer activation value a (L) function, and a (L) It is the hidden layer activation value a again (L-1) The function of $a^{(L-1)}$, and so on, until the input layer:

[0192] L = f(a (L) (a (L-1) (...a (1) (x))))

[0193] Using the chain rule, the loss function for a certain weight The partial derivative (i.e. gradient) of can be expressed as

[0194]

[0195] For each weight W, its gradient The calculation formula is:

[0196]

[0197] Expanding layer by layer, each item is: The partial derivative of the loss function with respect to the activation value of this layer; The derivative of the activation function; The transfer of activation values ​​from the previous layer.

[0198] This chain-like process allows the gradient to be passed from the output layer to the input layer, which is the essence of backpropagation. Each term of the gradient calculates the partial derivative of the loss function with respect to the output, the output with respect to the hidden state, and the hidden state with respect to the weight. After calculating the gradient, LSTM uses the gradient descent algorithm to update the weights. The specific process is:

[0199] 2.3.4 Update weights based on gradients

[0200] After calculating the gradient, LSTM uses the gradient descent algorithm to update the weights. The specific process is as follows: In neural networks, the gradient is the partial derivative of the loss function (such as mean square error, MSE) with respect to network parameters (such as weights and biases). If the weights are slightly adjusted, the direction and magnitude of the change in the loss function. Through the gradient descent algorithm, we can gradually adjust the parameters of the network so that the loss function gradually decreases. This parameter adjustment process is training. Mathematically, the gradient is the partial derivative vector of a multivariate function, which indicates the rate at which the function changes in each direction. In neural networks, the gradient represents the partial derivative of the loss function (such as mean square error, MSE) with respect to network parameters (such as weights and biases). For a certain weight, the gradient can be expressed as:

[0201]

[0202] Where w is the weight, which is the parameter connecting different layers in the neural network; the input of each neuron is the weighted sum of the output of the previous layer of neurons and the corresponding weight; specifically, suppose we have an input vector x = [x 1 ,x 2 ,...,x n ], weight vector w = [w 1 ,w 2 ,...,w n ], then the weighted sum of neurons in a certain layer can be expressed as:

[0203] z=w 1 ·x 1 +w 2 ·x 2 +…+w n ·x n +b

[0204] Where b is the bias; Loss is the loss function, which is a function used to measure the difference between the predicted output of the neural network and the true label. Common loss functions include Mean Squared Error (MSE) and Cross-Entropy. The present invention uses Mean Squared Error.

[0205] This gradient value reveals the direction and magnitude of the change in the loss function. In a neural network, the loss function measures the difference between the model's predicted output and the actual label. Take the mean square error (MSE) as an example. Suppose we have N samples and the predicted value is The true value is y i , then the loss function can be defined as:

[0206]

[0207] During training, we want to minimize the loss function, which means making the model's predicted output as close to the true value as possible. The process of minimizing the loss function is achieved by adjusting the network parameters. In order to find the minimum value of the loss function, we need to calculate its partial derivative with respect to the network parameters, that is, the gradient. Gradient descent is an optimization algorithm that aims to find its minimum value by repeatedly adjusting the parameters so that the loss function gradually decreases. In gradient descent, the parameter update formula is:

[0208]

[0209] Where ω is the weight to be updated, η is the learning rate, and controls the magnitude of each parameter update. is the gradient of the loss function with respect to ω.

[0210] In LSTM, backpropagation through time (BPTT) calculates the gradient of each time step and updates the parameters of the network based on these gradients. Weight update means that the network learns the mapping from input to output, making the error smaller and smaller, and finally achieving fit to the target data.

[0211] 3. Real-time correction of PV storage fluctuation suppression strategy

[0212] The precondition of this strategy is to use the RBMO-LSTM algorithm proposed in this invention to obtain a trained LSTM neural network, obtain the meteorological data of the location of the photovoltaic storage system in real time, and predict the future photovoltaic power based on the meteorological data. Based on the prediction of future photovoltaic power, a photovoltaic storage real-time correction fluctuation suppression strategy is proposed.

[0213] The photovoltaic storage system involved in the present invention is a photovoltaic grid-connected inverter and energy storage battery system. There is also an AC load in the system. According to the traditional control strategy, when the photovoltaic output is greater than the AC load, the excess power is connected to the grid through the photovoltaic grid-connected inverter for power generation. When the photovoltaic power fluctuates too much, it will have a greater impact on the power grid.

[0214] The real-time photovoltaic power sampling is constructed into a time series, and the RMSE is calculated with the predicted photovoltaic power, and real-time correction is performed according to the predicted RMSE value. The smaller the real-time RMSE, the higher the reliability of the photovoltaic prediction accuracy. When the RMSE decreases, the charging or discharging power of the energy storage is increased.

[0215] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0216] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention, and the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A photovoltaic storage inverter system optimization control method based on RBMO-LSTM photovoltaic prediction method, characterized in that: The photovoltaic storage inverter system optimization control method based on the RBMO-LSTM photovoltaic prediction method comprises the following steps: Step S1: RBMO algorithm modeling The prediction method matrix for photovoltaics is: The variable n represents the number selected when the RBMO algorithm is initialized, and dim represents the dimension of the problem variable. From the matrix, we can see that the multidimensional coordinates of each row vector of the matrix represent the position of the variable in the spatial hypercube; the upper and lower limits of the variable position are: x i,j =(u b -L b )·Rand1+l b (2) u b and l b are the upper and lower limits of the problem variables, representing the constraints of the variables in practical problems; the following Rnad represents random numbers; the algorithm simulates the process of variables finding the location of variables: Where t is the number of iterations, X i (t+1) is the position of the i-th new search, p represents the number of small groups participating in the variable, X m (t) is the mth individual randomly selected, X i (t) represents the i-th individual, X rs (t) is a randomly selected individual, q is the number of small groups of another size, the number of p is between 2-5, the number of q is between 10 and n, and formula (4) represents a large population; The following expression simulates the process of variables attacking the position of variables: Where X food (t) is the position of the variable, The following formula expresses the process of variable position storage, which is used to simulate the fitness function, which is the objective function of the optimization problem: fitness old i and fitness new i are the old and new fitness functions, and the meaning of the above formula (7) is to update the position of the predicted photovoltaic power generation variable when the effect of the fitness function is better than the old value; Step S2: Obtain the optimal hyperparameters of LSTM based on the RBMO algorithm The RMSE value of the predicted data and the actual data is obtained, that is, the fitness function, which is Where N is the number of samples, x f is the predicted value, x a is the true value; After obtaining the RMSE value, the position of the predicted photovoltaic power generation variable is updated according to formula (7) and the next iteration is performed. Before the algorithm is run, the number of iterations is set to m. When the number of iterations reaches m, the algorithm stops running. The position of the predicted photovoltaic power generation variable obtained at this time is the optimal LSTM hyperparameter. After the algorithm iteration is completed, the newly generated LSTM with the optimal hyperparameters will be used to predict the PV power generation data, and the hyperparameter combination that shows the best fit will be retained and subsequently used to predict the PV power generation.

2. The method for optimizing and controlling a photovoltaic storage inverter system based on the RBMO-LSTM photovoltaic prediction method according to claim 1, characterized in that: After obtaining the predicted photovoltaic power generation, it is compared with the actual photovoltaic power and plotted, and subsequent optimization control is continued. The subsequent optimization control is LSTM training. When performing LSTM training, the LSTM neural network is a neural network with multi-dimensional input and single-dimensional output. The input is specific weather data, and the weather data of the photovoltaic inverter application location is extracted. The output of the LSTM neural network is the photovoltaic power value.

3. The method for optimizing and controlling a photovoltaic storage inverter system based on the RBMO-LSTM photovoltaic prediction method according to claim 2, characterized in that: The weather data includes temperature, humidity, surface wind speed, surface horizontal radiation intensity, surface direct normal radiation intensity and scattered radiation intensity.

4. The method for optimizing and controlling a photovoltaic storage inverter system based on the RBMO-LSTM photovoltaic prediction method according to claim 2, characterized in that: The LSTM training includes forward propagation, calculation of loss function, back propagation and gradient calculation, and updating weights according to gradient.

5. The method for optimizing and controlling a photovoltaic storage inverter system based on the RBMO-LSTM photovoltaic prediction method according to claim 4, characterized in that: The forward propagation is the process by which the LSTM network processes the input data and generates the predicted output. The input data is processed according to the formula X = [x1, x2, ..., x t ] to calculate, where each x i They are all vectors containing multiple meteorological features; The loss function is calculated using the mean square error MSE as the loss function, assuming that there are N sample prediction values The true value is y t , then the loss function is defined as: MSE measures the error between the predicted value and the true value, and the goal is to minimize this error.

6. The method for optimizing and controlling a photovoltaic storage inverter system based on the RBMO-LSTM photovoltaic prediction method according to claim 4, characterized in that: The back propagation and gradient calculation are key algorithms for calculating gradients in neural networks. Through back propagation, the network calculates the gradient layer by layer from the output layer forward, passes the error back, and thus updates the weights and biases of each layer.

7. The method for optimizing and controlling a photovoltaic storage inverter system based on the RBMO-LSTM photovoltaic prediction method according to claim 4, characterized in that: In the updating of weights according to the gradient, mathematically, the gradient is the partial derivative vector of a multivariate function, which indicates the rate at which the function changes in each direction; in a neural network, the gradient indicates the partial derivative of the loss function with respect to the network parameters; the gradient can be expressed as: Where w is the weight, which is the parameter connecting different layers in the neural network; the input of each neuron is the weighted sum of the output of the previous layer of neurons and the corresponding weight; This gradient value reveals the direction and magnitude of the change in the loss function; in a neural network, the loss function measures the gap between the model's predicted output and the actual label; in the mean square error MSE, the predicted value for N samples is The true value is yi, then the loss function can be defined as: To find the minimum value of the loss function, it is necessary to calculate its partial derivative with respect to the network parameters, i.e. the gradient; Gradient descent repeatedly adjusts parameters to gradually reduce the loss function and find its minimum value; In gradient descent, the parameter update formula is: Among them, ω is the weight that needs to be updated, η is the learning rate, which controls the amplitude of each parameter update. is the gradient of the loss function with respect to ω.

8. The method for optimizing and controlling a photovoltaic storage inverter system based on the RBMO-LSTM photovoltaic prediction method according to claim 2, characterized in that: The LSTM neural network is a type of recursive neural network RNN, including a storage unit and a gate mechanism, wherein the gate mechanism includes obtaining a forget gate function, obtaining an input gate function, obtaining an output gate function, and obtaining an activation function sigmoid. The formula of the forget gate function is: t =σ(W f ·[h t-1 ,x t ]+b f ) (9) In the formula, f t is the output of the forget gate: f t is the output of the forget gate at time step t, which is a vector with the same dimension as the hidden state of the LSTM unit. t The value range is between [0,1], which determines the cell state C at the previous time step. t-1 How much information will be retained and passed to the cell state C at the current time step t t ; When f t When an element of is close to 1, it means that the corresponding information will be retained; When it is close to 0, it means that the corresponding information will be forgotten; Where σ is the Sigmoid activation function: its specific expression is Map the input value to the range of [0,1]. Since the output of the forget gate needs to be between [0,1], use the Sigmoid function to process the result of the linear combination; W f is the weight matrix of the forget gate, which contains the input x of the current time step t and the hidden state h at the previous time step t-1 The relevant weight, W f The dimension is (n hidden ,n hidden +n input ), where n hidden is the dimension of the hidden state of the LSTM unit, n input is the input vector x t The dimension of this matrix determines the input x at the current time step. t and the hidden state h at the previous time step t-1 Impact on the output of the forget gate; [h t-1 ,x t ]The concatenation of the hidden state and the input, [h t-1 ,x t ] is the hidden state h of the previous time step t-1 and the input x at the current time step t The concatenation vector, h t-1 is the hidden state vector of the previous time step, with dimension n hidden , which contains the information of the previous time step, x t is the input vector of the current time step, with dimension n input , which contains the input data for the current time step, b f is the bias vector of the forget gate, dimension is n hidden , the bias vector is used to adjust the baseline value of the forget gate output to help the model fit the data; According to the function of the forget gate, when the function outputs 1, the transmitted information is completely retained, and when the function value is 0, the transmitted value is 0 and the information is completely "forgotten"; The formula for the input gate function is: i t =σ(W i ·[h t-1 ,x t ]+b i ) (10) In the formula, i t : Output vector of the input gate, dimension is n hidden ,i t Determines the input x of the current time step t and the hidden state h at the previous time step t-1 For unit status C t The impact of t The value of is between [0,1]. The closer the value is to 1, the more the information should be written into the unit state, and the closer the value is to 0, the more the information should be ignored. Where σ is the Sigmoid activation function: its specific expression is Map the input value to the range of [0,1]. Since the output needs to be between [0,1], use the Sigmoid function to process the result of the linear combination. W i : The weight matrix of the input gate, dimension is (n hidden ,n hidden +n input ), which controls the hidden state h of the previous time step t-1 and the input x at the current time step t The weight in the input gate; [h t-1 ,x t ]: hidden state h t-1 and the input vector x t The concatenation vector of t-1 is the hidden state vector of the previous time step, with dimension n hidden , x t is the input vector of the current time step, with dimension n input ; b i : Bias vector of input gate, dimension is n hidden ; The bias term adjusts the activation value of the input gate to help the network fit the data; The input gate normalizes the new input information to a value between -1 and 1 through the hyperbolic tangent function and converts the previous f t and i t Merge into new C t Numeric value; The formula of the output gate function is: the t =σ(W o [h t-1 ,x t ]+b o ) (12) h t =o t ·tanh(C t ) (13) o t : Output vector of the output gate, dimension is n hidden ;o t Determines the cell state C at the current time step t How much information needs to be output to the hidden state h t ;o t The value of is between [0,1]. The closer the value is to 1, the more the information should be output, and the closer it is to 0, the more the information should be suppressed. W o : The weight matrix of the output gate, dimension is (n hidden ,n hidden +n input ), which controls the hidden state h of the previous time step t-1 and the input x at the current time step t The weight in the output gate; [h t-1 ,x t ]: hidden state h t-1 and the input vector x t The concatenation vector of t-1 is the hidden state vector of the previous time step, with dimension n hidden , x t is the input vector of the current time step, with dimension n input ; b o : Bias vector of the output gate, dimension is n hidden , the bias term adjusts the activation value of the output gate to help the network fit the data; The output gate represents the output component of the LSTM neuron; The formula for the activation function is: Where σ is the Sigmoid activation function: its specific expression is Map the input value to the range of [0,1]. Since the output needs to be between [0,1], use the Sigmoid function to process the result of the linear combination. The output is obtained through an activation function sigmoid, and then multiplied by the normalized C function tanh. t value.

9. The method for optimizing and controlling a photovoltaic storage inverter system based on the RBMO-LSTM photovoltaic prediction method according to claim 1, characterized in that: The LSTM hyperparameters include learning rate, number of hidden layer neurons, number of layers, and L2 regularization coefficient; In the gradient descent algorithm, the learning rate is the coefficient used to reduce the gradient and determines the gradient loss rate at each step; Symbol θ t Represents the current parameters, and the gradient of the loss function relative to the parameters is expressed as Represents, and the learning rate is represented by η; In the LSTM neural network, an additional regularization term is added to the loss function, which is L2 regularization. L2 regularization prevents the model from overfitting; The number of neurons in the hidden layer determines the capacity and expressiveness of the LSTM network structure.