Wind power generation power prediction method
By combining the LSTM model with the mRMR and PSO algorithms to screen meteorological features and optimize parameters, the problems of insufficient redundancy and temporal nature of meteorological features in wind power forecasting are solved, achieving higher prediction accuracy and stability.
Patent Information
- Application Number
- CN202510579312.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-09-23
Smart Images

Figure CN120687795A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wind power generation prediction, and in particular relates to a short-term wind power generation power prediction method. Background Art
[0002] With the continuous advancement of renewable energy generation technologies, wind power, as a clean, green energy source, has been widely adopted and promoted worldwide. Within the wind power sector, meteorological conditions significantly influence the output power of wind turbines. Therefore, accurate wind power forecasting is crucial to improving the efficiency and reliability of wind power generation. Existing wind power forecasting methods fall into three categories: physical methods, statistical methods, and artificial intelligence methods.
[0003] When building physical models, it involves describing complex physical phenomena such as topography and geocyclones, which are often difficult to accurately quantify, thus introducing large prediction errors. Furthermore, traditional statistical models are based on the assumption of linear relationships between data. This assumption may not fully capture the nonlinear characteristics of actual data, making it difficult to guarantee predictive performance. Artificial intelligence methods are insensitive to initial error information and have good nonlinear data fitting and parameter learning capabilities. Therefore, the present invention selects artificial intelligence methods for prediction.
[0004] Currently, forecasting methods such as artificial neural networks, hybrid density neural networks, and those combining multivariate time series clustering algorithms with deep learning networks have all improved the accuracy of wind power forecasts to a certain extent. However, because these methods fail to simultaneously account for the redundancy and temporal nature of meteorological characteristics, there is still considerable room for improvement in wind power forecast accuracy. Therefore, the exploration of more advanced forecasting methods has become increasingly necessary. Summary of the Invention
[0005] In order to address the shortcomings of the existing technology, the present invention aims to provide a short-term wind power prediction method, which adopts the LSTM model and takes advantage of its time series modeling to more accurately capture the temporal variation law of wind power, thereby significantly improving the accuracy and robustness of wind power prediction.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A method for predicting short-term wind power generation includes the following steps:
[0008] (1) Obtain historical meteorological data and corresponding power generation data to form an original wind power prediction dataset and preprocess it;
[0009] (2) Using the mRMR method, the 10 most representative features were selected from the original meteorological features while retaining the power generation data;
[0010] (3) Divide the feature set and the corresponding data set in step (2) into a training set and a test set;
[0011] (4) The LSTM algorithm is used to analyze the training set data. During the training process, the model parameters are continuously adjusted according to whether the value of the loss function meets the preset requirements, and the final prediction model is obtained through iterative optimization;
[0012] (5) Input the test set into the trained final model to obtain the prediction results.
[0013] Furthermore, in step (1), historical meteorological data and wind power generation data are obtained and preprocessed, specifically including the following steps:
[0014] (1-1) Obtain a meteorological data set with a sampling interval of 1 hour and 31 features, as well as the corresponding historical power generation data with a sampling interval of 5 minutes;
[0015] (1-2) Taking the hourly historical power generation as the target variable, the outliers that exceed the upper and lower boundaries of the data are first eliminated through the box plot method; then a strategy based on neighboring data is adopted, that is, the missing values are filled with the average values of the upper and lower adjacent values of the vacant data; finally, the data is processed through the maximum and minimum normalization method to eliminate the dimensional differences between different features, and the final data points are used to establish the model.
[0016] Furthermore, step (2) specifically includes:
[0017] (2-1) The data points obtained after preprocessing are divided into training set and test set. Assume that the power generation data of the training set is vector y=(y1,…,y k ,…,y 3563 ) T , where y1, y k and y 3563 Represent the power generation of the first, kth and last data points in the training set respectively; at the same time, the meteorological characteristics of the training set are the matrix C = (c1,…,c i ,…,c j ,…,c 31 ), where c1, c i 、c j and c 31 are the column vectors of the first, i-th, j-th and thirty-first meteorological characteristics respectively, and i≠j;
[0018] (2-2) Find the feature set C5 containing 5 meteorological features, including:
[0019] Make the average value of the mutual information between all features in the feature set and the generated power the largest, that is,
[0020]
[0021] in, Representative meteorological characteristics c i The mutual information between the generated power y, p(c i ,y) represents the joint probability distribution of the i-th meteorological characteristic and the power generation y, p(c i ) and p(y) are their respective marginal probability distributions;
[0022] At the same time, it is also necessary to ensure that the average mutual information value between all meteorological features in the feature set is minimum. The average mutual information value between all meteorological features in the feature set is:
[0023]
[0024] in, Representative meteorological characteristics c i with c j The mutual information between them, p(c i ,c j ) represents meteorological characteristics c i and c j The joint probability distribution of simultaneous occurrence, p(c i ) and p(c j ) represent meteorological characteristics c i and c j marginal probability distribution of individual occurrences;
[0025] (2-3) The mRMR algorithm is used to find the optimal meteorological feature set by combining maximum correlation and minimum redundancy, which is expressed as:
[0026]
[0027] in, represents the maximum value in the mRMR score, D represents the correlation score between the feature and the target variable, and R represents the redundancy score between features.
[0028] Furthermore, step (4) specifically includes:
[0029] (4-1) Using the Sigmoid function to control the forget gate f in the LSTM network model t Constraints are expressed as:
[0030] f t =σ(Wf ·[h t-1 ,x t ]+b f );
[0031] Among them, σ represents the Sigmoid function, W f represents the weight matrix of the forget gate, b f represents the bias term of the forget gate, h t-1 Represents the hidden state of the previous moment, x t Represents the input at the current moment;
[0032] (4-2) Use Sigmoid function to adjust the input gate i in the LSTM network model t Constraints are expressed as:
[0033] i t =σ(W i ·[h t-1 ,x t ]+b i );
[0034] Among them, W i represents the weight matrix of the input gate, b i represents the bias term of the input gate, h t-1 Represents the hidden state of the previous moment, x t Represents the input at the current moment;
[0035] (4-3) Use the hyperbolic tangent function to calculate the candidate cell state Constraints are expressed as:
[0036]
[0037] Among them, tanh represents the hyperbolic tangent activation function, which is used to map the input value to between -1 and 1; W s The weight matrix representing the candidate cell state, b s A bias term representing the candidate cell state;
[0038] At the same time, the new cell state S is obtained t , expressed as:
[0039]
[0040] Among them, f t is the output of the forget gate, which is a value between 0 and 1 and determines the previous cell state s t-1 How much information is retained in the current cell state S t middle,
[0041] is the candidate cell state at the current time step;
[0042] (4-4) Use Sigmoid function to output gate O t Constraint, this gate determines how much of the current cell state is output to the hidden state h t , expressed as:
[0043] O t =tanh(W o ·[h t-1 ,x t ]+b o ),
[0044] h t =O t tanhS t ;
[0045] Among them, W o Represents the weight matrix of the output gate, which is used to perform linear transformation on the input; b o Represents the bias term of the output gate;
[0046] (4-5) Use the PSO algorithm to optimize the weights and parameters of LSTM. The update rule of PSO is:
[0047]
[0048] in, represents the value of the velocity of particle m in dimension n at time step t-1; ω is the inertia weight, which is used to control the influence of the previous velocity of the particle; represents the value of the velocity of particle m in dimension n at time step t; p mn represents the individual optimal position of particle m in dimension n; g n represents the value of the global optimal position in dimension n; k1 and k2 are acceleration constants used to adjust the speed of the particle moving towards the individual optimal position and the global optimal position; r1 and r2 are random numbers in the range of [0,1] used to increase randomness; Represents the value of the position of particle m in dimension n at time step t; Represents the value of the position of particle m in dimension n at time step t+1.
[0049] The beneficial effects of the present invention are:
[0050] The proposed short-term wind power forecasting method not only achieves lower average error but also more accurately captures overall trends. This improved performance is primarily attributed to its effective handling of temporal and redundancy features, which allows it to capture key features and improve model robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is the overall flow chart of the present invention;
[0052] Figure 2 This is a cell structure diagram of the LSTM grid of the present invention;
[0053] Figure 3 The wind power prediction curves of four methods are shown in Figure 2. DETAILED DESCRIPTION
[0054] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of application of the present invention.
[0055] like Figure 1-3 As shown, the present invention proposes a method for predicting short-term wind power generation, comprising the following steps:
[0056] Step (1) obtains historical meteorological data and corresponding power generation data to form an original wind power generation prediction data set and preprocesses it. Specifically, it includes the following steps:
[0057] (1-1) Obtain a meteorological dataset with a sampling interval of 1 hour and 31 features, as well as the corresponding historical power generation data with a sampling interval of 5 minutes, for a total of 39,454 data points;
[0058] (1-2) Taking the hourly historical power generation as the target variable, the outliers that exceed the upper and lower boundaries of the data are first eliminated through the box plot method; then a strategy based on neighboring data is adopted, that is, the missing values are filled with the average values of the upper and lower adjacent values of the vacant data; finally, the data is processed through the maximum and minimum normalization method to eliminate the dimensional differences between different features, and finally 3663 data points are obtained for model establishment.
[0059] Step (2) uses the mRMR method to select the 10 most representative features from the original meteorological features while retaining the power generation data. Specifically, it includes:
[0060] (2-1) The 3663 pre-processed data points are divided into a training set and a test set. Assume that the power generation data of the training set is a vector y = (y1,…,y k ,…,y 3563 ) T , where y1, y k and y 3563 Represent the power generation of the first, kth and last data points in the training set respectively. At the same time, the meteorological characteristics of the training set are the matrix C = (c1,…,c i ,…,c j ,…,c31 ), where c1, c i 、c j and c 31 are the column vectors of the first, i-th, j-th and thirty-first meteorological characteristics respectively, and i≠j;
[0061] (2-2) Find the feature set C5 containing 5 meteorological features, including:
[0062] To maximize the average value of the mutual information between all features in the feature set and the generated power, that is,
[0063]
[0064] in, Representative meteorological characteristics c i The mutual information between the generated power y, p(c i ,y) represents the joint probability distribution of the i-th meteorological characteristic and the power generation y, p(c i ) and p(y) are their respective marginal probability distributions;
[0065] At the same time, it is also necessary to ensure that the average mutual information value between all meteorological features in the feature set is minimum. The average mutual information value between all meteorological features in the feature set is:
[0066]
[0067] in, Representative meteorological characteristics C i with C j The mutual information between them, p(c i ,c j ) represents meteorological characteristics c i and c j The joint probability distribution of simultaneous occurrence, p(c i ) and p(c j ) represent meteorological characteristics c i and c j marginal probability distribution of individual occurrences;
[0068] (2-3) The mRMR algorithm is used to find the optimal meteorological feature set by combining maximum correlation and minimum redundancy, which is expressed as:
[0069]
[0070] in, represents the maximum value in the mRMR score, D represents the correlation score between the feature and the target variable (generated power), and R represents the redundancy score between features.
[0071] Step (3) divides the feature set and the corresponding data set of step (2) into a training set and a test set.
[0072] Step (4) uses the LSTM algorithm to analyze the training set data. During the training process, the model parameters are continuously adjusted according to whether the value of the loss function meets the preset requirements, and the final prediction model is obtained through iterative optimization. Specifically, it includes:
[0073] (4-1) Using the Sigmoid function to control the forget gate f in the LSTM network model t Constraints are expressed as:
[0074] f t =σ(W f ·[h t-1 ,x t ]+b f );
[0075] Among them, σ represents the Sigmoid function, W f represents the weight matrix of the forget gate, b f represents the bias term of the forget gate, h t-1 Represents the hidden state of the previous moment, x t Represents the input at the current moment;
[0076] (4-2) Use Sigmoid function to adjust the input gate i in the LSTM network model t Constraints are expressed as:
[0077] i t =σ(W i ·[h t-1 ,x t ]+b i );
[0078] Among them, W i represents the weight matrix of the input gate, b i represents the bias term of the input gate, h t-1 Represents the hidden state of the previous moment, x t Represents the input at the current moment;
[0079] (4-3) Use the hyperbolic tangent function to calculate the candidate cell state Constraints are expressed as:
[0080]
[0081] Among them, tanh represents the hyperbolic tangent activation function, W s The weight matrix representing the candidate cell state, b s A bias term representing the candidate cell state.
[0082] At the same time, the new cell state S is obtained t , expressed as:
[0083]
[0084] Among them, f t is the output of the forget gate, which is a value between 0 and 1 and determines the previous cell state s t-1 How much information is retained in the current cell state S t in,i t is the input of the input gate, is the candidate cell state at the current time step;
[0085] (4-4) Use Sigmoid function to output gate O t Constraint, this gate determines how much of the current cell state is output to the hidden state h t , expressed as:
[0086] O t =tanh(W o ·[h t-1 ,x t ]+b o ),
[0087] h t =O t tanhS t ;
[0088] Among them, the hyperbolic tangent activation function tanh is used to map the input value to between -1 and 1; W o Represents the weight matrix of the output gate, which is used to perform linear transformation on the input; b o Represents the bias term of the output gate;
[0089] (4-5) Use the PSO algorithm to optimize the weights and parameters of LSTM. The update rule of PSO is:
[0090]
[0091] in, represents the value of the velocity of particle m in dimension n at time step t-1; ω is the inertia weight, which is used to control the influence of the previous velocity of the particle; represents the value of the velocity of particle m in dimension n at time step t; p mn represents the individual optimal position of particle m in dimension n; g nrepresents the value of the global optimal position in dimension n; k1 and k2 are acceleration constants used to adjust the speed of the particle moving towards the individual optimal position and the global optimal position; r1 and r2 are random numbers in the range of [0,1] used to increase randomness; Represents the value of the position of particle m in dimension n at time step t; Represents the value of the position of particle m in dimension n at time step t+1.
[0092] Step (5) inputs the test set into the trained final model to obtain the prediction results.
[0093] Comparative experiment:
[0094] To verify the effectiveness of the mRMR-PSO-LSTM method proposed in this paper in processing the redundancy and temporal nature of meteorological features, a comparative experiment was conducted. The BP neural network and PSO-LSTM algorithms were applied to a data set containing all 31 meteorological features and a data set containing only the five meteorological features selected by the mRMR algorithm. The parameters of the BP neural network, PSO algorithm, and LSTM model after PSO algorithm optimization are shown in Table 1.
[0095] Table 1 Parameter settings of the model
[0096]
[0097] After training, the four models predicted the wind power generation in the next 100 hours. The prediction curves are shown in the attached figure. Figure 3 The prediction curve of the PSO-LSTM model is closer to the actual wind power generation, which shows that LSTM can better capture the time series characteristics of meteorological data and thus more accurately predict future wind power generation. Figure 3 Judging from the overall fluctuation amplitude of the four subgraphs, the model based on the mRMR algorithm exhibits a smaller fluctuation range around the actual wind power generation. This demonstrates that the mRMR algorithm can filter out redundant meteorological features, making the model's prediction performance more robust. Notably, the mRMR-PSO-LSTM model proposed in this paper achieves the most accurate prediction performance of the four models, after comprehensively considering the temporal and redundancy of meteorological features.
[0098] In order to quantify the difference in prediction accuracy among the four models, the present invention evaluated the four models. The evaluation data is shown in Table 2.
[0099] Table 2 Model comparison table
[0100]
[0101] The mRMR-PSO-LSTM model achieved the best results in terms of MAPE, MAE, and RMSE, at 2.1328, 2.9038, and 3.3819, respectively. It also achieved the highest R² value of 0.9785. Specifically, the MAPE of the mRMR-PSO-LSTM model decreased by 2.5994, 2.3647, and 0.9141 compared to the BP, mRMR, and PSO-LSTM, respectively. Its MAE decreased by 6.9851, 3.6986, and 1.2445 compared to the BP, mRMR, and PSO-LSTM, respectively. Furthermore, the RMSE of the mRMR-PSO-LSTM model decreased by 9.6574, 4.2966, and 1.4494 compared to the BP, mRMR, and PSO-LSTM, respectively. Most notably, the R2 of mRMR-PSO-LSTM is improved by 0.29, 0.0892, and 0.0223 compared to BP, mRMR, and PSO-LSTM, respectively.
[0102] These results demonstrate that the mRMR-PSO-LSTM model achieves significant advantages in wind power forecasting, not only achieving lower average error but also more accurately capturing overall trends. This performance improvement is primarily attributed to its effective handling of temporal order and redundancy, which enables it to capture key features and improve model robustness. These results provide strong support for the mRMR-PSO-LSTM model's ability to more reliably forecast wind power in practical applications.
[0103] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A method for predicting short-term wind power generation, characterized by: The following steps are involved: (1) Obtain historical meteorological data and corresponding power generation data to form an original wind power prediction dataset and preprocess it; (2) Using the mRMR method, the 10 most representative features were selected from the original meteorological features while retaining the power generation data; (3) Divide the feature set and the corresponding data set in step (2) into a training set and a test set; (4) The LSTM algorithm is used to analyze the training set data. During the training process, the model parameters are continuously adjusted according to whether the value of the loss function meets the preset requirements, and the final prediction model is obtained through iterative optimization; (5) Input the test set into the trained final model to obtain the prediction results.
2. The method for predicting short-term wind power generation according to claim 1, characterized in that: In step (1), historical meteorological data and wind power generation data are obtained and preprocessed, which specifically includes the following steps: (1-1) Obtain a meteorological data set with a sampling interval of 1 hour and 31 features, as well as the corresponding historical power generation data with a sampling interval of 5 minutes; (1-2) Taking the hourly historical power generation as the target variable, the outliers that exceed the upper and lower boundaries of the data are first eliminated through the box plot method; then a strategy based on neighboring data is adopted, that is, the missing values are filled with the average values of the upper and lower adjacent values of the vacant data; finally, the data is processed through the maximum and minimum normalization method to eliminate the dimensional differences between different features, and the final data points are used to establish the model.
3. The method for predicting short-term wind power generation according to claim 2, characterized in that: Step (2) specifically includes: (2-1) The data points obtained after preprocessing are divided into training set and test set; let the power generation data of the training set be vector y=(y1,…,y k ,…,y 3563 ) T , where y1, y k and y 3563 Represent the power generation of the first, kth and last data points in the training set respectively; at the same time, the meteorological characteristics of the training set are the matrix C = (c1,…,c i ,…,c j ,…,c 31 ), where c1, c i 、c j and c 31 are the column vectors of the first, i-th, j-th and thirty-first meteorological characteristics respectively, and i≠j; (2-2) Find the feature set C5 containing 5 meteorological features, including: Make the average value of the mutual information between all features in the feature set and the generated power the largest, that is, in, Representative meteorological characteristics c i The mutual information between the generated power y, p(c i ,y) represents the joint probability distribution of the i-th meteorological characteristic and the power generation y, p(c i ) and p(y) are their respective marginal probability distributions; At the same time, it is also necessary to ensure that the average mutual information value between all meteorological features in the feature set is minimum. The average mutual information value between all meteorological features in the feature set is: in, Representative meteorological characteristics c i with c j The mutual information between them, p(c i ,c j ) represents meteorological characteristics c i and c j The joint probability distribution of simultaneous occurrence, p(c i ) and p(c j ) represent meteorological characteristics c i and c j marginal probability distribution of individual occurrences; (2-3) The mRMR algorithm is used to find the optimal meteorological feature set by combining maximum correlation and minimum redundancy, which is expressed as: in, represents the maximum value in the mRMR score, D represents the correlation score between the feature and the target variable, and R represents the redundancy score between features.
4. The method for predicting short-term wind power generation according to claim 3, characterized in that: Step (4) specifically includes: (4-1) Using the Sigmoid function to control the forget gate f in the LSTM network model t Constraints are expressed as: f T =σ(W f ·[h T-1 ,x t ]+b f ); Among them, σ represents the Sigmoid function, W f represents the weight matrix of the forget gate, b f represents the bias term of the forget gate, h t-1 represents the hidden state at the previous moment, x t Represents the input at the current moment; (4-2) Use Sigmoid function to adjust the input gate i in the LSTM network model t Constraints are expressed as: i t =σ(W i ·[h t-1 ,x t ]+b i ); Among them, W i represents the weight matrix of the input gate, b i represents the bias term of the input gate, h t-1 Represents the hidden state of the previous moment, x t Represents the input at the current moment; (4-3) Use the hyperbolic tangent function to calculate the candidate cell state Constraints are expressed as: Among them, tanh represents the hyperbolic tangent activation function, which is used to map the input value to between -1 and 1; W s The weight matrix representing the candidate cell state, b s A bias term representing the candidate cell state; At the same time, the new cell state S is obtained t , expressed as: Among them, f t As the output of the forget gate, it is a value between 0 and 1 that determines the previous cell state s t-1 How much information is retained in the current cell state S t middle, is the candidate cell state at the current time step; (4-4) Use Sigmoid function to output gate O t Constraint, this gate determines how much of the current cell state is output to the hidden state h t , expressed as: O t =tanh(W o ·[h t-1 ,x t ]+b o ), h t =O t tanhS t ; Among them, W o Represents the weight matrix of the output gate, which is used to perform linear transformation on the input; b o Represents the bias term of the output gate; (4-5) Use the PSO algorithm to optimize the weights and parameters of LSTM. The update rule of PSO is: in, represents the value of the velocity of particle m in dimension n at time step t-1; ω is the inertia weight, which is used to control the influence of the previous velocity of the particle; represents the value of the velocity of particle m in dimension n at time step t; p mn represents the individual optimal position of particle m in dimension n; g n represents the value of the global optimal position in dimension n; k1 and k2 are acceleration constants used to adjust the speed of the particle moving towards the individual optimal position and the global optimal position; r1 and r2 are random numbers in the range of [0,1] used to increase randomness; Represents the value of the position of particle m in dimension n at time step t; Represents the value of the position of particle m in dimension n at time step t+1.