A Short-Term Wind Power Prediction Method Based on Sequence Correlation Mechanism and Informer
Through a short-term wind power power prediction method based on sequence correlation mechanism and Informer, discrete wavelet decomposition and singular spectrum analysis combined with Fourier self-attention mechanism, the problems of high computational cost and data sensitivity of the existing wind power prediction model are solved, and higher precision wind power prediction is achieved.
Patent Information
- Application Number
- CN202311165559.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-09-11
AI Technical Summary
The existing wind power power prediction model has shortcomings in terms of high computational cost, strong parameter dependence, high data sensitivity and the inability to fully capture the nonlinear and time dependence of wind power data, resulting in low prediction accuracy.
The short-term wind power power prediction method based on sequence correlation mechanism and Informer is adopted to extract the frequency and trend characteristics of wind power sequences through discrete wavelet decomposition and singular spectrum analysis, and feature extraction is used for Fourier self-attention and sequence correlation mechanism, and prediction is carried out in combination with the encoding-decoding structure.
It effectively weakens the signal-to-noise ratio drop caused by too deep wavelet decomposition layers, improves prediction accuracy, can better handle the time-dependent and multi-dimensional data characteristics of wind power, and improves prediction accuracy and adaptability.
Smart Images

Figure CN117251724B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wind power, and in particular to a short-term wind power prediction method based on a sequence correlation mechanism and Informer. Background Art
[0002] With the rapid development of China's social economy and technology, the demand and consumption of energy have gradually increased. People have gradually realized the importance of energy conservation, emission reduction and sustainable development. Therefore, the development of new energy has been increasingly emphasized. Vigorously promoting the construction of new energy such as solar power generation and wind power generation has gradually become the main trend of energy transformation. As a kind of clean energy, wind power generation is pollution-free, renewable, inexhaustible, has a short infrastructure cycle, flexible installed capacity, low operation and maintenance costs, and can also be complementary to other power generation systems. It is a power generation method with large-scale commercial development conditions. However, affected by climatic conditions, wind power generation has dynamics, randomness and uncertainty, and the grid connection operation of wind power has a greater impact on the power balance and the safe operation of the power grid. Therefore, it is necessary to effectively predict the wind power to maintain the stable operation of the power system.
[0003] Wind power prediction models are mainly divided into physical models, statistical models, and artificial intelligence models;
[0004] The construction of a physical prediction model requires a set of mathematical equations and physical laws based on multi-variable physical parameters, and at the same time considers meteorological information such as the current and future climate, temperature and pressure of the wind farm, as well as site-specific conditions, such as NWP. Numerical weather prediction (NWP) technology is the most commonly used physical model. By using microscale meteorology and computational fluid dynamics to convert data into wind speeds at the height of the wind turbine, NWP indirectly predicts the wind power.
[0005] Statistical models make full use of a large amount of historical data of wind farms as much as possible, and use algorithms such as Kalman filtering, autoregressive (AR), autoregressive integrated moving average (ARMA) to extract the linear relationship between wind power input features, establish the mapping relationship between them, and thus replace the physical causal relationship analysis. Statistical methods are often used for short-term wind power prediction, ultra-short-term prediction, and wind farm power prediction lacking meteorological information.
[0006] Artificial intelligence models are based on intelligent algorithms such as artificial neural networks, support vector machines, and fuzzy logic to establish a non-linear mapping relationship model, and then use the correlation between input data and output data to predict wind power. For example, the CNN-Informer model, the model based on Informer, and the convolutional layer that extracts time features of different frequencies are used to predict the average wind power in the next three hours.
[0007] When using the above technologies, the following technical problems are found in the prior art:
[0008] Physical models are applicable to newly built wind farms and do not require historical wind farm power data support. However, they need a large amount of accurate numerical weather prediction data, and the calculation cost is too high. 1. The power curve method predicts only based on the power curve of wind turbines, ignoring the complexity and nonlinear characteristics of the wind farm. At the same time, the prediction effect of this method is poor for wind speeds outside the power curve. 2. Physical fluid dynamics models require complex fluid dynamics calculations, with high calculation costs and long time consumption. In addition, the accuracy of the model highly depends on the detailed modeling and parameter adjustment of the wind farm and wind turbines. If the parameter settings are inaccurate, the prediction results may be inaccurate. 3. The wind speed-power density model predicts based on the relationship between wind speed and power density, ignoring the influence of other factors on power output, such as wind direction, temperature, etc. At the same time, this model needs to appropriately select and fit the power density function, and different regions and wind farms may require different function models. 4. Mechanical models need to accurately model the mechanical structure and transmission system of wind turbines, including factors such as the force on the blades, rotor speed, and generator load. If the parameters in the model are inaccurate or lack real-time updates, the prediction results may deviate.
[0009] For statistical methods, if there is sufficient historical data, it has high prediction accuracy. However, the prediction performance of these methods is limited by the nonlinearity and non-smoothness of the data, cannot better adapt to sudden information, and requires collecting a large amount of historical data, with certain usage limitations. 1. Regression analysis models have limited effects in dealing with complex nonlinear relationships and cannot capture the more complex relationship between wind speed and wind power. In addition, regression models are more sensitive to outliers and noise data. 2. Time series models may have poor prediction effects when the data has nonlinear relationships or the trends are not obvious. In addition, time series models may require more complex model structures for long-term seasonal and periodic changes, and parameter selection is also more difficult. 3. Grey system models have high requirements for data and need a certain amount of historical data for modeling. In the case of insufficient data volume or poor data quality, the prediction results may be inaccurate. 4. Support vector regression (SVR) has high calculation costs when dealing with large-scale data, and the time for training and adjusting the model is long. At the same time, SVR is also more sensitive to hyperparameters such as selecting the appropriate kernel function and adjusting the regularization parameter. 5. Random forest models may have overfitting problems when facing high-dimensional data and sparse data. In addition, since random forest is an ensemble model based on decision trees, it is more sensitive to noise and outliers in the input data.
[0010] Artificial intelligence models have strong non - linear fitting and generalization abilities, but they require appropriate network structures and parameter settings, and it is difficult to explain the internal operation mechanism of the models. For example, CNN - Informer only focuses on extracting features in the time dimension and does not process environmental features such as wind direction, wind speed, and temperature in wind power data. There may be complex relationships between these features and wind power generation, and it may not be possible to fully consider the influence of these domain - specific features, failing to capture the non - linear correlations between features, thus leading to limitations in prediction. Moreover, when extracting time features, it only focuses on information at different time scales in the current time period and does not make full use of historical time information, failing to capture the correlations between different time series;
[0011] To this end, we design a short - term wind power prediction method based on sequence - related mechanism and Informer to provide another technical solution for the above - mentioned technical problems. Summary of the Invention
[0012] Based on this, it is necessary to provide a short - term wind power prediction method based on sequence - related mechanism and Informer to solve the technical problems raised in the above - mentioned background technology.
[0013] To solve the above - mentioned technical problems, the present invention adopts the following technical solutions:
[0014] A short - term wind power prediction method based on sequence - related mechanism and Informer, the steps are as follows:
[0015] S1: Obtain the historical power data and features of the wind turbine to be measured;
[0016] S2: Use discrete wavelet decomposition to learn the time patterns in wind power sequence prediction;
[0017] S3: Perform singular spectrum analysis on the periodic components after wavelet decomposition respectively;
[0018] S4: Input the decomposed frequency sequence into the frequency prediction model;
[0019] S5: Obtain the same feature sequences at different time periods from the trend component and calculate the similarity relationship; The new trend component obtained through the sequence - related mechanism is input into the Informer model;
[0020] S6: Add the predicted values obtained from the frequency prediction model and the predicted values obtained from the trend prediction model to obtain the final predicted value.
[0021] As a preferred embodiment of the short-term wind power prediction method based on sequence correlation mechanism and Informer provided by the present invention, in the step S1, the historical power data of the wind turbine includes wind power generation, wind speed, wind direction, temperature and air pressure data.
[0022] As a preferred embodiment of the short-term wind power prediction method based on sequence correlation mechanism and Informer provided by the present invention, the true value after normalizing the historical power data of the wind turbine is multiplied by the standard deviation of the original data, and then added to the mean value of the original data to obtain the final data. The final data is divided into a training set, a validation set and a test set according to the ratio of 7:1:2.
[0023] As a preferred embodiment of the short-term wind power prediction method based on sequence correlation mechanism and Informer provided by the present invention, in the step S3, the periodic component after singular spectrum analysis is reconstructed back to the original sequence, and the detail component after inverse wavelet transform is used as the frequency;
[0024] The sequence correlation mechanism is applied to the trend component after wavelet decomposition, and three new sequences are combined. The three new sequences are reconstructed through inverse wavelet transform and used as the trend term to input into the model.
[0025] As a preferred embodiment of the short-term wind power prediction method based on sequence correlation mechanism and Informer provided by the present invention, in the step S4, the decomposed frequency sequence is input into the frequency prediction model, and the steps are as follows:
[0026] After entering the embedding layer, the sequence will be converted into a new feature sequence;
[0027] Convolution kernels of different sizes are used to capture features of different time scales, and different-scale branches are used to simulate different underlying patterns in the wind power sequence;
[0028] The results of different branches are combined to complete the utilization of the comprehensive information of the sequence;
[0029] After passing through the embedding layer, the sequence is input into the convolutional-discrete Fourier layer.
[0030] As a preferred embodiment of the short-term wind power prediction method based on sequence correlation mechanism and Informer provided by the present invention, the discrete Fourier layer contains multiple branches, and different scale sizes are used to simulate potentially different time patterns. The steps are as follows:
[0031] Feature correlations are extracted through downsampling, and the sequence with more prominent features is transferred to the DFT attention block;
[0032] Capture the periodic patterns or frequency characteristics in the sequence through Fourier transform within the block, thereby enhancing the prediction results.
[0033] As a preferred embodiment of the short-term wind power prediction method based on sequence correlation mechanism and Informer provided by the present invention, in the step S5, obtain the same feature sequences at different time periods and calculate the acquaintance relationship. The steps are as follows:
[0034] Through the sliding window method, randomly combine these n different time series in pairs, and then calculate the similarity relationship between them. The formula is as follows:
[0035]
[0036] where, x t1 and x t2 are the same feature sequences collected at different times, reflects the similarity relationship between sequences, is x t1 and x t2 's standard deviation;
[0037] Through the subsequence aggregation layer, perform swap recombination according to different sequence lengths β1,...,β k and different positions α1,...,α k ;
[0038] The subsequences are aggregated by softmax-normalized confidence. The expression is as follows:
[0039]
[0040] head i = SC(Q i ,K i ,V i )
[0041] MultiHead(Q,K) = W output *Concat(head1,...,head h )
[0042] where, R Q,K (β) is the sequence correlation between sequences x t1 and x t2 , Roll(α i ,β j ) represents the operation at the position of α j with length β i ;
[0043] As a preferred embodiment of the short-term wind power prediction method based on sequence correlation mechanism and Informer provided by the present invention, the model is evaluated through a test set, and the mean square error and mean absolute error between the predicted power and the actual power are calculated.
[0044] It can be undoubtedly seen that through the above technical solution of the present application, the technical problems to be solved by the present application can surely be solved.
[0045] Meanwhile, through the above technical solution, the present invention has at least the following beneficial effects:
[0046] The short-term wind power prediction method based on sequence correlation mechanism and Informer provided by the present invention can effectively weaken the disadvantages such as the decrease in signal-to-noise ratio when the decomposition layer of wavelet decomposition is too deep, and the principal components in the decomposed sequence can be effectively extracted through singular spectrum analysis to effectively weaken the error brought by prediction; at the same time, the method of using Fourier self-attention and sequence correlation mechanism to extract frequency features and trend features respectively can effectively handle long-term dependencies and improve prediction accuracy. Wind power data often has long-term dependencies in time series, that is, the current power value may be affected by time steps in the relatively distant past.
[0047] It has multiple layers of attention mechanism and feed-forward neural network, which can effectively extract relevant features in wind power data; through the encoding and decoding processes of input data, it can automatically learn the patterns and dependencies in time series data without manually designing complex feature engineering; this makes the model have stronger adaptability and can adapt to the characteristics and changes of different wind farms.
[0048] It can accept the input of multi-dimensional time series data; in addition to the wind power value itself, other relevant variables (such as wind speed, temperature, etc.) can also be used as the input of the model to more comprehensively consider the factors affecting wind power and further improve the prediction performance; by combining the self-attention mechanism and the encoding-decoding structure, it can effectively utilize historical data and time context information for prediction; this ability enables the model to more accurately predict the future trend and volatility of wind power. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a schematic diagram of the frequency prediction model of the present invention;
[0051] Figure 2 Schematic diagram of the Fourier attention module of the present invention;
[0052] Figure 3 Schematic diagram of the subsequence aggregation layer of the present invention;
[0053] Figure 4 Schematic diagram of the trend prediction model of the present invention;
[0054] Figure 5 Overall flowchart of the present invention;
[0055] Figure 6 Schematic diagram of the sequence decomposition of the present invention. Detailed implementation manners
[0056] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0057] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0058] It should be noted that, without conflict, the embodiments in the present invention and the features and technical solutions in the embodiments can be combined with each other.
[0059] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0060] Referring to Figures 1-6 , a short-term wind power prediction method based on a sequence correlation mechanism and Informer.
[0061] The prediction accuracy is achieved by making full use of the correlation information between sequences during the model training and learning process; the steps are as follows:
[0062] 1. Obtain the historical power data and other features of the wind turbine to be measured;
[0063] 2. In addition to wind power generation, this data set also provides multi-dimensional meteorological data such as wind speed, wind direction, temperature, and air pressure. Generally speaking, wind power generation is jointly affected by multiple factors such as wind speed, wind direction, temperature, and air pressure. However, the degree of influence of each meteorological factor on wind power output varies. Therefore, before establishing a prediction model, it is necessary to first conduct necessary correlation analysis on the multi-dimensional data. When analyzing the correlation between wind power and various meteorological data with different dimensions, the Pearson correlation coefficient analysis method (PPC) that excludes the influence of dimensions is often used to describe the degree of correlation between them. The PPC analysis method measures the correlation between two variables through the distance between the two variables. When one variable shows an upward or downward trend, and the change trend of the other variable is the same as it, the two variables are positively correlated, and the corresponding PCC coefficient is positive; when one variable shows an upward or downward trend, and the change trend of the other variable is opposite to it, the two variables are negatively correlated, and the corresponding PCC coefficient is negative; if the PCC is 0, it means that the two variables are non-linearly correlated. The value range of PCC is between -1 and 1, and the absolute value thereof characterizes the strength of the correlation relationship. For the data set adopted by this technology, the correlation between wind power and data based on PCC correlation analysis is shown in Table 1;
[0064] Table 1 Correlation Analysis
[0065]
[0066] For the scenario of "output saturation and wind speed overflow", a wind speed data cleaning method is proposed to enhance the correlation between wind power and wind speed. Specifically, the scenarios of full power generation are screened out from the historical data of the wind farm, the wind speed information is statistically analyzed, and the minimum wind speed that ensures all units to output full power is defined as the "full load wind speed". The values exceeding the full load wind speed are flattened to the full load wind speed, and the values lower than the full load wind speed remain unchanged. Due to the volatility of actual wind power data, a large amount of data will cause numerical problems. To accelerate the speed of gradient descent to obtain the optimal solution, this technology standardizes the original power data before constructing the model, multiplies the normalized true value by the standard deviation of the original data, and then adds the mean value of the original data to obtain the final data. The final data is divided into a training set, a validation set, and a test set according to the ratio of 7:1:2;
[0067] 3. Discrete Wavelet Decomposition (DWD) is adopted to learn the time patterns in wind power sequence prediction. Discrete Wavelet Decomposition (DWD) is a method that decomposes a time series signal into a trend term and periodic components (referred to as the approximation signal and the detail signal respectively) using low-pass and high-pass filters. The input signal is decomposed by convolving it with a low-pass filter and a high-pass filter. As the number of decomposition levels increases, the noise in the high-frequency band increases. The wavelet coefficients in the high-frequency band may be affected by mutation signals, resulting in a decrease in the signal-to-noise ratio and the loss of some detail information, causing distortion of the reconstructed signal. At the same time, as the number of decomposition levels increases, the sensitivity to noise and unimportant features also increases, which may lead to overfitting. This means that the model may overfit the wavelet details of the training data and cannot generalize well to new data. To mitigate the above effects, three-level decomposition is selected to maintain the integrity of the signal as much as possible. To extract the main components of the signal, Singular Spectrum Analysis (SSA) is performed on the periodic components (detail signals), and the sequence correlation mechanism mentioned below is applied to the trend term (approximation signal). The superiority of Singular Spectrum Analysis lies in that it is not affected by any prior information or assumption conditions. It can stably identify and strengthen periodic signals, and concentrate the most predictable components into several time series. Therefore, several meaningful components are selected to reconstruct the sequence to reduce the influence of noise. The sequence correlation mechanism can select several components with high similarity from the trend term and remove the hidden noise and frequency components in the sequence to improve the prediction performance. Singular Spectrum Analysis is performed on the periodic components (D1, D2, D3) after wavelet decomposition respectively. The periodic components after Singular Spectrum Analysis are reconstructed back into the original sequence, and the detail components after inverse wavelet transform are used as frequencies. The sequence correlation mechanism is applied to the trend components (A1, A2, A3) after wavelet decomposition, combined into three new sequences, and the three new sequences are reconstructed through inverse wavelet transform and input into the model as the trend term;
[0068] X = A1 + A2 + A3… + A N + D1 + D2 + … + D N
[0069] X freq = SSA(D1 + D2 + …D N )
[0070] X trend = SC(A1 + A2 + …A N )。
[0071] 4. The decomposed frequency sequence is passed into the designed frequency prediction model. As Figure 1As shown, after entering the embedding layer, the sequence is converted into a new feature sequence. Convolution kernels of different sizes are used to capture features at different time scales, and branches of different scales are used to simulate different underlying patterns in the wind power sequence. Then, the results of different branches are combined to complete the utilization of the comprehensive information of the sequence. After passing through the embedding layer, the sequence is fed into the convolutional-discrete Fourier layer. This layer contains multiple branches, and different scale sizes are used to simulate potential different time patterns. In each branch, first, feature correlations are extracted through downsampling, and then the sequence with more prominent features is transferred to the DFT attention block. The Fourier transform within the block helps capture periodic patterns or frequency characteristics in the sequence, thereby enhancing the prediction results. Within the attention mechanism, the Fourier transforms are performed on Q, K, and V respectively. For the query vector Q and the "to-be-query" vector K, information such as the real part, imaginary part, and frequency will be extracted to better capture the frequency relationships between each input sequence element when calculating the attention weights. For the content vector V, only the real part and the imaginary part are extracted because they already contain sufficient information. Finally, the new Q and K are fed into the MatMui module and the time domain is returned. The whole process is as follows:
[0072]
[0073] Q real ,Q imag ,Q freq =F(Q),K real ,K imag ,K freq =F(K),V real ,V imag =F(V)
[0074] Finally, the weighted sum of different weighted features is calculated through the self-attention mechanism:
[0075] Attention(Q f ,K f ,V f )=σ f F -1 (V f )
[0076] Finally, after passing through the merging module, the frequency prediction values of different scales are merged and then linearly projected to output the final prediction result of the frequency components.
[0077] 5. Concatenate the trend components (A1, A2, A3) after wavelet decomposition and input them into the position encoding layer, which is used to assign a unique encoding to each position in the sequence data to capture position information and provide it to the model for learning and inference. Obtain the position encoding (transform the input sequence into a more representative feature representation) and mask (ignore or mask those positions or invalid elements that are not of interest) regarding feature and time information. For each period, the feature sequences such as power, wind direction, and wind speed collected and the mask obtained through the position encoding layer are used as the input of the encoder, and power is used as the input of the decoder. The encoder-decoder structure encodes the input X t into the hidden state H t , and decodes the output Y t from . The decoder calculates the new hidden state according to the previous state and other necessary outputs at the k-th step, and then predicts the (k + 1)-th sequence . This inference method is called dynamic decoding.
[0078] 6. Input A1, A2, A3 into the sequence correlation layer. The feature takes the data sequence of the prediction target and the previous feature data sequence (i is the time series of the i-th feature) at this time point as the input of the sequence association layer. By means of a sliding window, the same feature sequence in different time periods is obtained, and these n different time series are randomly combined in pairs, and then the similarity relationship between them is calculated. The formula is as follows:
[0079]
[0080] where x t1 and x t2 are the same feature sequences collected at different times, reflects the similarity relationship between the sequences, is the standard deviation of x t1 and x t2 ;
[0081] To refine the information difference of different subsequences of x t1 and x t2 in the sequence, the present invention provides a subsequence aggregation layer as shown in the following figure, which can perform swap recombination according to different sequence lengths β1,..., β k and different positions α1,..., α k , such as Figure 3 ;
[0082] This operation allows the mutual convolution of similar subsequences at the same position, emphasizing the extraction of the difference in subsequence information at different positions in the sequences {xt} and {x(t+1)}. Finally, the subsequences are aggregated through a softmax-normalized confidence process as follows:
[0083]
[0084] head i = SC(Q i ,K i ,V i )
[0085] MultiHead(Q,K) = W output *Concat(head1,...,head h )
[0086] where R Q,K (β) is the sequence correlation between sequences x t1 and x t2 , and Roll(α i ,β j ) represents an operation at position α j with length β i . Through the Roll operation, the query Q, key K, and value V of similar subsequences in different sequences are calculated as differential information, multiple Q, K, and V are concatenated, and the differential information is passed to the decoder part of the informer after passing through the linear layer.
[0087] Perform a convolution operation on the curves with relatively poor correlation among them to extract the differential information between them, then randomly shuffle and swap and recombine. The sequence obtained after re-concatenation is normalized to obtain the query, key, and value with differential information;
[0088] 7. The Q and K in the multi-head attention layer in the decoder are weighted by the encoder and the sequence correlation layer.
[0089] 8. In the encoder part, the output is obtained from the positional encoding layer and passed through three linear transformations to obtain the query, key, and value. Query (Q): A query refers to the representation of the target position used to obtain comparisons with other positions in the attention mechanism. It is similar to the information to be queried or the content to be searched for. Key (K): The key is the reference used to compare the query with the representations of other positions. It provides the information for calculating the attention weights. Value (V): The value is the actual representation associated with the position. It stores the information to be extracted or the output obtained from the model. The sparse attention layer is then entered, where important queries are filtered out to reduce the time complexity and memory in space. After that, it enters the downsampling layer, which uses a max-pooling layer with a stride of 2 to downsample the data to half its length after each layer, reducing the memory usage. To enhance the robustness of the distillation operation, the encoder structures are stacked and combined, and their outputs are concatenated to obtain the output of the complete encoder.
[0090] 9. The decoder part consists of two identical multi-head attention layers. The Q and K obtained through the sequence correlation layer and the Q and K obtained from the encoder part are combined into new Q and K through weight coefficients and then enter the sparse attention layer. The output obtained finally passes through a linear layer to be mapped into an output of the same dimension.
[0091] 10. The predicted values obtained from the frequency prediction model and the trend prediction model are added together to obtain the final predicted value.
[0092] 11. The trained model is evaluated using the test dataset, and error metrics between the predicted power and the actual power are calculated, such as the Mean Squared Error (MSE), Mean Absolute Error (MAE), etc., to evaluate the performance of the model.
[0093] The Mean Squared Error (MSE) is a commonly used metric to evaluate the average degree of difference between the predicted results and the actual results. Calculating the MSE requires squaring the difference between the predicted power and the corresponding actual power for each sample and then taking the average. The smaller the value of the MSE, the closer the predicted results are to the actual results.
[0094] The Mean Absolute Error (MAE) is another commonly used error metric, which is used to measure the average absolute difference between the predicted results and the actual results. Calculating the MAE requires taking the absolute value of the difference between the predicted power and the corresponding actual power for each sample, and then taking the average. The smaller the value of the MAE, the closer the predicted results are to the actual results. By calculating and comparing the error metrics of the trained model on the test dataset, the prediction performance of the model can be intuitively understood and compared with other models or benchmarks. This helps to select the best model or make further improvements and optimizations to improve the accuracy and reliability of wind power prediction.
[0095] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A short-term wind power prediction method based on sequence correlation mechanism and Informer, characterized in that, The steps are as follows: S1: Obtain the historical power data and characteristics of the fan to be measured; S2: Use discrete wavelet decomposition to learn the time patterns in wind power sequence prediction; S3: Perform singular spectrum analysis on the periodic components after wavelet decomposition respectively; S4: Input the decomposed frequency sequence into the frequency prediction model; S5: Obtain the same feature sequences at different time periods from the trend component and calculate the similarity relationship; the new trend component obtained through the sequence correlation mechanism is input into the Informer model; S6: Add the predicted values obtained from the frequency prediction model and the predicted values obtained from the trend prediction model to obtain the final predicted value; In step S3, the periodic components after singular spectrum analysis are reconstructed back to the original sequence, and the detail components after inverse wavelet transform are used as frequencies; Apply the sequence correlation mechanism to the trend components after wavelet decomposition, combine them into three new sequences, and reconstruct the three new sequences through inverse wavelet transform as the trend terms input into the model.
2. The short-term wind power prediction method based on sequence correlation mechanism and Informer according to claim 1, characterized in that In step S1, the historical power data of the fan includes wind power generation, wind speed, wind direction, temperature, and air pressure data.
3. A short-term wind power prediction method based on a sequence correlation mechanism and Informer according to claim 1, characterized in that Multiply the true value after normalizing the historical power data of the fan by the standard deviation of the original data, and then add the mean of the original data to obtain the final data. Divide the final data into a training set, a validation set, and a test set according to the ratio of 7:1:
2.
4. A short-term wind power prediction method based on sequence correlation mechanism and Informer according to claim 1, characterized in that, In step S4, input the decomposed frequency sequence into the frequency prediction model. The steps are as follows: After entering the embedding layer, the sequence will be converted into a new feature sequence; Use convolutional kernels of different sizes to capture features at different time scales, and use branches of different scales to simulate different underlying patterns in the wind power sequence; Combine the results of different branches to complete the utilization of the comprehensive information of the sequence; After passing through the embedding layer, the sequence is input into the convolutional-discrete Fourier layer.
5. A short-term wind power prediction method based on a sequence correlation mechanism and Informer according to claim 4, characterized in that, The discrete Fourier layer contains multiple branches, and different scale sizes are used to simulate potentially different time patterns. The steps are as follows: Extract feature correlations through downsampling and transfer the sequence with more prominent features to the DFT attention block; Capture the periodic patterns or frequency characteristics in the sequence through Fourier transform within the block to enhance the prediction results.
6. A short-term wind power prediction method based on sequence correlation mechanism and Informer according to claim 1, characterized in that In step S5, obtain the same feature sequences at different time periods and calculate the similarity relationship. The steps are as follows: Through the sliding window method, randomly combine these n different time series in pairs, and then calculate the similarity relationship between them. The formula is as follows: where x t1 and x t2 are the same feature sequences collected at different times, reflecting the similarity relationship between the sequences, is the standard deviation of x t1 and x t2 ; Through the subsequence aggregation layer, according to different sequence lengths β1,...,β k and different positions α1,...,α k perform swap recombination; The subsequences are aggregated by softmax-normalized confidence. The expression is as follows: head i = SC(Q i ,K i ,V i ) MultiHead(Q,K) = W output *Concat(head1,...,head h ) Among them, R Q,K (β) is the sequence correlation between sequence x t1 and x t2 , and Roll(α i , β j ) represents the operation at the α j position with a length of β i .
7. A short-term wind power prediction method based on a sequence correlation mechanism and Informer according to claim 2, characterized in that, Evaluate the model through the test set, and calculate the mean square error and mean absolute error between the predicted power and the actual power.
Citation Information
Patent Citations
Medium-term electric quantity prediction method based on sequence component decomposition and neural network
CN110298475A
Regional wind power generation power prediction method and system based on multi-scale double-space-time network
CN116451873A