BiTCN-Informer-based ultra-short-term wind power combined prediction method and model
Through the BiTCN-Informer model, combined with the bidirectional time convolution neural network and the Informer network, the probabilistic sparse self-attention mechanism and self-attention distillation mechanism are used to solve the problem of insufficient prediction of fluctuating data intervals near local peaks and mutation data points in the prior art, and achieve higher precision wind power prediction.
Patent Information
- Application Number
- CN202510463338.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-01
AI Technical Summary
The prediction accuracy of the fluctuation data intervals and mutation data points near local peaks is insufficient, and the mutations and fluctuations of data cannot be effectively identified, affecting the accuracy of wind power prediction.
The ultra-short-term wind power combined prediction method based on BiTCN-Informer is adopted, and the bidirectional time convolutional neural network (BiTCN) is combined with the Informer network. Through the probability sparse self-attention mechanism and self-attention distillation mechanism, the timing and spatial characteristics of the data are identified, the dependencies of the data are captured, and the prediction accuracy is improved.
Effectively identifying fluctuation data intervals and mutation data points near local peaks improves the accuracy of wind power prediction and improves the model's prediction effect on special data.
Smart Images

Figure CN120414477A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wind power prediction, and more specifically, to an ultra-short-term wind power combined prediction method and model based on BiTCN-Informer. Background Art
[0002] With the need to build a "modern energy system" and promote the transformation of the traditional structure for developing new energy, the scale of wind power grid connection has gradually expanded. However, the strong volatility and randomness characteristics of wind power are likely to cause local voltage fluctuations, affecting the stable operation of the power system and posing a severe challenge to the operation of the power grid. Therefore, improving the accuracy of wind power prediction is of great significance to the reliability and safety of power grid operation.
[0003] Traditional medium- and long-term wind power predictions (such as daily / weekly scales) are difficult to meet the real-time balance requirements of the power system. Ultra-short-term prediction can dynamically capture the impacts of complex meteorological factors such as sudden wind speed changes and turbulence on wind farms by integrating numerical weather forecasts, artificial intelligence algorithms, and real-time monitoring data, and perform high-precision predictions on the future short-term wind power output. The power system needs to maintain instantaneous supply-demand balance, and the minute-level power fluctuations of wind power may lead to frequency violations or voltage instability. Ultra-short-term prediction provides a decision-making basis for dispatching agencies to quickly adjust standby power sources, energy storage, or demand-side responses. Therefore, under the trend of renewable energy grid connection, ultra-short-term prediction is a key technical means to support the resilience of the power grid and reduce the risk of wind curtailment. Therefore, how to achieve short-term wind power prediction is of great significance to enhancing the reliability and economy of power grid operation.
[0004] In the method for ultra - short - term power prediction of a wind farm disclosed in CN118213974A, after decomposing multiple sequences into sub - signals, the sub - signals of these sequences are combined into a multi - dimensional feature vector, which is processed by a two - dimensional convolutional neural network, and a self - attention encoding - decoding module is used to dynamically mine the global spatio - temporal attention of the data. Specifically, this patent uses the variational mode decomposition algorithm to decompose the time - series data of wind speed and power, constructs a multi - dimensional feature vector, improves the information acquisition ability of the neural network model, and reduces the learning difficulty of the neural network. A combined model of a two - dimensional convolutional neural network and a self - attention encoding - decoding module is constructed. The two - dimensional convolutional neural network is used to extract the feature maps of wind speed and power in different spatio - temporal dimensions, and the spatio - temporal correlation between wind speed - power and power - power in different spatio - temporal dimensions is mined through the self - attention mechanism, significantly improving the prediction accuracy of wind power. This patent is based on predicting the overall feature trend of the data sequence. The two - dimensional convolutional network used in it multiplies unidirectional elements, ignoring the forward and backward dependence relationships between data and being unable to identify data mutations. At the same time, using the self - attention encoding - decoding module cannot judge the length of the fluctuating data interval, and is not flexible in identifying the fluctuating data interval and mutation data points. This patent can achieve ultra - short - term power prediction of a wind farm from the overall data, but the prediction of special data such as the fluctuating data interval and mutation data points near local peaks is insufficient, which will affect the prediction accuracy of the model. Summary of the Invention
[0005] The main technical problem to be solved by the present invention is to provide a combined ultra - short - term wind power prediction method based on BiTCN - Informer for the deficiencies of the prior art in considering less the prediction of special data such as the fluctuating data interval and mutation data points near local peaks and having poor prediction accuracy for the fluctuating data interval and mutation data points.
[0006] Another technical problem solved by the present invention is to provide a combined ultra - short - term wind power prediction model based on BiTCN - Informer.
[0007] The object of the present invention is achieved through the following technical solutions:
[0008] A combined ultra - short - term wind power prediction method based on BiTCN - Informer, the steps include:
[0009] S1. Perform standardized pre - processing on the data;
[0010] S2. Use a bidirectional temporal convolutional neural network to analyze the information characteristics of past observations and future covariates, and extract the spatial characteristics of the data;
[0011] S21. The forward TCN network encodes the future covariates of the time series;
[0012] Classification information acat After being processed by the embedding layer, it is combined with the historical covariate a cov to obtain the information input x that combines all past and future covariate information cov , x cov is processed through a densely connected layer to obtain h cov , h cov is fed into a forward TCN network with a two-layer structure to predict all future information and obtain o cov , h cov and the o output by the single-layer TCN network Ncov are combined to obtain the hidden state information h cov1 .
[0013] S22. The backward TCN network encodes the past observations and covariates of the time series;
[0014] The lagged target value y lag is combined with the historical covariate a cov and the classification information after being processed by the embedding layer to obtain the input x lag , x lag is processed through a densely connected layer to obtain h lag , h lag is fed into a backward TCN network with a two-layer structure to process the past information and obtain o lag , h lag and the o output by the single-layer TCN network Nlag are combined to obtain the hidden state information h lag1 .
[0015] S23. Use a densely connected layer to combine the future information o cov and the past information o lag to obtain the spatial features of the original sequence;
[0016] S3. Adopt the Informer probabilistic sparse self-attention mechanism and self-attention distillation mechanism to simplify network parameters, reduce computational complexity, identify the dependencies of long sequence data, and capture the temporal features of the sequence;
[0017] S31. After processing the input data through the probabilistic sparse self-attention mechanism, output a "positive" Q matrix with a high correlation with K;
[0018] S32. After processing all the input data through the convolutional layer and the embedding layer, input it into the self-attention distillation mechanism to extract the important features of all the data;
[0019] S33. After processing the partially masked input data through the convolutional layer and the embedding layer, input it into the self-attention distillation mechanism to extract the important features of the partially masked data;
[0020] S34. Combine the important features obtained in S32 and all the data and partial mask data obtained in S32 to obtain a feature map, and obtain the K and V matrices from the feature map;
[0021] S35. Use multi-head attention to process the Q, K, and V matrices to obtain weight allocation and obtain the temporal features of the data;
[0022] S4. Use a fully connected layer to integrate the spatial features and temporal features and output the prediction result o.
[0023] Further, the formula for the normalization preprocessing is:
[0024]
[0025] In the formula, z is the normalization result, x is the input observation value, μ is the total data mean, and σ is the total data standard deviation.
[0026] Further, the bidirectional temporal convolutional neural network includes two TCN networks, and the TCN network includes a dilated convolutional network, a GELU activation function, dropout random inactivation regularization, and a densely connected layer network.
[0027] Further, the expression of the forward dilated convolutional network is:
[0028]
[0029] The expression of the backward dilated convolutional network is:
[0030]
[0031] In the formula, G(t) is the convolution value of the t-th element, k is the number of filters, n is the element ordinal number, g(n) is the filter value of element n, is the input value of the t-th element shifted to the right by d n , is the input value of the t-th element shifted to the left by d n .
[0032] Further, the expression of the GELU activation function is:
[0033]
[0034] In the formula, G(x) is the value after GELU activation processing, P(X≤x), Φ(x) are the cumulative distribution functions of the standard normal distribution.
[0035] Further, the expression of the probabilistic sparse self-attention mechanism is:
[0036]
[0037] In the formula, softmax is a normalization function, is a query matrix with a high degree of correlation; K is a key matrix, V is a value matrix, and d is the vector dimension.
[0038] Furthermore, the self-attention distillation mechanism includes attention blocks, convolutional layers, and pooling layers with a multi-layer structure.
[0039] Furthermore, the expression of the self-attention distillation mechanism is:
[0040]
[0041] In the formula, MaxPool(·) is the max pooling operation, ELU(·) is the exponential linear unit, Convld(·) is the 1D convolution operation, and [·] AB is the attention block.
[0042] Furthermore, the Q matrix with a low degree of correlation is replaced by the mean value of each data V matrix.
[0043] Furthermore, the process of all input data and partially masked input data passing through the convolutional layer and the embedding layer is a generative reasoning decoding process, which synthesizes the known sequence and the predicted sequence to achieve a one-time processing of the predicted sequence, expressed as:
[0044]
[0045] In the formula, is the decoded input sequence, is the known sequence of the start token, is the predicted sequence with a padding value of 0.
[0046] A very short-term wind power combined prediction model based on BiTCN-Informer, including a bidirectional temporal convolutional neural network (BiTCN) and an Informer network. The bidirectional temporal convolutional neural network (BiTCN) and the Informer network are combined through a fully connected layer for output;
[0047] The bidirectional temporal convolutional neural network (BiTCN) includes a forward TCN network and a backward TCN network. The TCN network consists of a dilated convolutional network, a GELU activation function, dropout random inactivation regularization, and a densely connected layer network;
[0048] The Informer network includes a probabilistic sparse self-attention mechanism and a self-attention distillation mechanism. The probabilistic sparse self-attention mechanism and the self-attention distillation mechanism are combined through multi-head attention. Among them, the self-attention distillation mechanism includes multiple attention blocks, and there are convolutional layers and max pooling layers between the attention blocks.
[0049] Compared with the prior art, the beneficial effects are as follows:
[0050] In the BiTCN adopted by the present invention, by analyzing the information characteristics of past observations and future covariates, the forward TCN network encodes the future covariates of the time series, and the backward TCN network encodes the past observations and covariates. The BiTCN composed of the bidirectional TCN network identifies the dependence relationship between data from the front and back spaces, determines the fluctuating data interval near the local peak, and fully excavates the spatial characteristics of the data. In addition, the present invention combines the Informer probability sparse self-attention mechanism and the self-attention distillation mechanism to simplify the network parameters and identify the temporal characteristics of the data. The probability sparse self-attention mechanism screens out the query matrix with a high degree of correlation with the key matrix through random sampling and divergence calculation, and replaces the query matrix with a low degree of correlation with the mean value of each data value matrix, reducing unnecessary operation processes. The self-attention distillation mechanism extracts important features, reduces the input dimension, expands the receptive field to obtain more feature information, and predicts the occurrence of mutation data points. The present invention uses the BiTCN in the BiTCN-Informer structure to identify the fluctuating data interval near the local peak, and the Informer determines the occurrence of mutation data points, which can improve the prediction accuracy of special data on the basis of predicting the overall trend of the data curve, making the prediction effect better. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is the structural flow chart of the BiTCN-Informer prediction combination model;
[0052] Figure 2 is the prediction curve graph of seven models;
[0053] Figure 3 is the comparison graph of the prediction curves of seven models. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The following further explains and clarifies with reference to the embodiments, but the specific embodiments do not limit the present invention in any form.
[0055] Embodiment 1
[0056] This embodiment provides a very short-term wind power combined prediction model based on BiTCN-Informer, as Figure 1 shown, including a bidirectional time convolutional neural network (BiTCN) and an Informer network;
[0057] The bidirectional time convolutional neural network (BiTCN) includes two TCN networks. The TCN network consists of multiple layers, and the time block constitutes the basic structure of each layer of the network. The time block includes a dilated convolutional network, a GELU activation function, dropout random inactivation regularization, and a dense connection layer network.
[0058] The BiTCN is divided into a forward TCN network and a backward TCN network according to the position of the recognized variable features, where the forward TCN network encodes the future covariates of the time series; the backward TCN network encodes the past observations and covariates. The expression of the forward dilated convolutional network is:
[0059]
[0060] The expression of the backward dilated convolutional network is:
[0061]
[0062] In the formula, G(t) is the convolution value of the t-th element, k is the number of filters, n is the element ordinal number, g(n) is the filter value of element n, is the input value of the t-th element shifted to the right by d n ; is the input value of the t-th element shifted to the left by d n ;
[0063] The dilated convolutional network expands the recognition range by intermittently extracting data, reducing the problem of layer stacking caused by continuous extraction.
[0064] GELU is a non-linear activation function based on the Gaussian distribution, and its expression is:
[0065]
[0066] In the formula: G(x) is the value after GELU activation processing; P(X≤x) and Φ(x) are the cumulative distribution functions of the standard normal distribution. It has a non-zero gradient when x < 0, alleviating the problem of neuron death to a certain extent and enhancing the expression ability of the model; it is smoother at x = 0, which helps to improve the convergence speed and performance of the training process; the slope is not constant when x > 0, providing an adaptive non-linear gradient with the input variable and reducing the problem of overfitting.
[0067] Dropout regularization can randomly discard some neurons that do not participate in model training, thereby enhancing the generalization ability of the model and reducing overfitting of the model.
[0068] The dense connection layer network combines the neurons between layers by weight connection, captures and integrates the complex relationships between features, and enhances the expression ability of the model.
[0069] The Informer network includes a probabilistic sparse self-attention mechanism and a self-attention distillation mechanism, and the probabilistic sparse self-attention mechanism and the self-attention distillation mechanism are combined by multi-head attention.
[0070] The probability sparse self-attention mechanism screens out the query matrices with a high degree of correlation with the key matrix through random sampling and divergence calculation, and replaces the query matrices with low correlation degrees with the mean values of each data value matrix, reducing unnecessary operation processes. The probability sparse self-attention expression is as follows:
[0071]
[0072] In the formula: softmax is a normalization function; is the query matrix with a high degree of correlation; K is the key matrix; V is the value matrix; d is the vector dimension.
[0073] The self-attention distillation mechanism includes attention block 1, attention block 2, and attention block 3. Convolutional layers and one-dimensional max pooling layers are provided between attention block 1 and attention block 2, and between attention block 2 and attention block 3. The distillation process from the input of the i-th layer to the (i + 1)-th layer is as follows:
[0074]
[0075] In the formula: MaxPool(·) is the max pooling operation; ELU(·) is the exponential linear unit; Convld(·) is the one-dimensional convolution operation; [·] AB is the attention block.
[0076] The process of all data and partial masked data passing through the convolutional layer and the embedding layer is the generative reasoning decoding process, which is expressed as:
[0077]
[0078] In the formula, is the decoding input sequence; is the known sequence of the start token; is the predicted sequence with a padding value of 0.
[0079] The bidirectional temporal convolutional neural network (BiTCN) and the Informer network are combined through a fully connected layer for output.
[0080] Embodiment 2
[0081] This embodiment provides a very short-term wind power combination prediction method based on BiTCN-Informer, as Figure 1 shown. The steps include:
[0082] S1. Perform standardized preprocessing on the data;
[0083] Using z-score normalization to transform multi-type data features that affect wind power and are not on the same scale into the same scale. The z-score normalization formula is as follows:
[0084]
[0085] In the formula: z is the normalization result, x is the input observation value, μ is the total data mean, and σ is the total data standard deviation.
[0086] S2. Use a bidirectional temporal convolutional neural network to analyze the information features of past observations and future covariates, and extract the spatial features of the data;
[0087] S21. The forward TCN network encodes the future covariates of the time series;
[0088] Classification information a cat After being processed by the embedding layer, it is combined with the historical covariate a cov to obtain the information input x of all past and future covariates cov , x cov is processed through a dense connection layer to obtain h cov , h cov is fed into a forward TCN network with a two-layer structure to predict all future information to obtain o cov , h cov and the o output by the single-layer TCN network Ncov are combined to obtain the hidden state information h cov1 .
[0089] S22. The backward TCN network encodes the past observations and covariates of the time series;
[0090] The lagged target value y lag is combined with the historical covariate a cov and the classification information after being processed by the embedding layer to obtain the input x lag , x lag is processed through a dense connection layer to obtain h lag , h lag is fed into a backward TCN network with a two-layer structure to process the past information to obtain o lag , h lag and the o output by the single-layer TCN network Nlag are combined to obtain the hidden state information h lag1 .
[0091] S23. Use a dense connection layer to combine the future information o cov and the past information o lag to obtain the spatial features of the original sequence;
[0092] S3. Simplify the network parameters by adopting the Informer's probabilistic sparse self-attention mechanism and self-attention distillation mechanism, reduce the computational complexity, identify the dependencies of long-sequence data, and capture the temporal features of the sequence;
[0093] S31. Process the input data through the probabilistic sparse self-attention mechanism and output a "positive" Q matrix with high correlation with K;
[0094] S32. Process all the input data through the convolutional layer and the embedding layer, then input it into the attention block 1 of the self-attention distillation mechanism. Reduce the input dimension through the convolutional layer and the one-dimensional max-pooling layer, then send it to the attention block 2. After passing through the convolutional layer and the one-dimensional max-pooling layer again, input it into the attention block 3 to extract the important features of all the data;
[0095] S33. Process half of the masked input data through the convolutional layer and the embedding layer, then input it into the self-attention distillation mechanism, and repeat the distillation process in S32 to extract the important features of the partially masked data;
[0096] S34. Combine the important features output from all the data or partially masked data obtained in S32 and S32 to obtain a feature map, get the K and V matrices from the feature map, and the "positive" Q matrix with high correlation with K. Replace the Q matrices with low correlation with the mean value of each data's V matrix;
[0097] Specifically, by calculating and querying each query vector Q i with all key vectors K j to obtain the KL divergence between the attention distribution and the uniform distribution, and filter out the "positive" Q matrices with large distribution differences. The KL divergence formula is as follows:
[0098]
[0099] where, Q i is the i-th query vector, K j is the j-th key vector, d is the vector dimension, L K is the number of key vectors, that is, the sequence length.
[0100] The larger the KL divergence value, the more "positive" the Q matrix is. Select the top u Qs with the largest KL divergence, where u = clnL Q , c is a constant, L Q is the length of the input sequence, and the K matrix regions corresponding to these Qs are regarded as high-correlation regions.
[0101] S35. Use multi-head attention to process the Q, K, and V matrices to obtain the weight distribution and get the temporal features of the data;
[0102] S4. Use the fully connected layer to integrate the spatial features and temporal features and output the prediction result o.
[0103] Example 3
[0104] This example provides a simulation experiment of a wind power prediction method combining a multi-scale graph convolutional network and an extended long short-term memory network.
[0105] In this example, the wind power data of a wind farm (wind farm site 1) in China, which has been processed (data_processed), is selected. The sampling period is 15 minutes, and the data from 0:00 on March 1, 2020 to 23:45 on December 31, 2020 is taken, with a total of 29,376 valid sampling points. The z-score normalization method is used to convert each feature parameter into the same magnitude. The original data is divided into a training set and a test set in a ratio of 9:1. The training set has a total of 26,422 sampling points, and the test set has a total of 2,922 sampling points. The sliding window method is used to divide and save the data of the training set and the test set into groups of 16 sampling points each, and predictions are made for 2,922 sampling points (from 13:00 on December 1, 2020 to 23:45 on December 31, 2020).
[0106] The wind power prediction method proposed in this invention is programmed in Python 3.9 language. The computing platform is a Win10 system, with an Intel i7-1165G7 @ 2.80GHz and 16.0GB of RAM. There are 12 input data features in addition to the date, and the prediction target is wind power. Therefore, the input feature dimension is 12, and the output feature dimension is 1. After a large number of debuggings, the relevant parameters of BiTCN are as follows: the number of input channels in both the forward and backward directions is [32, 64], including two layers; the number of filters is 3. The relevant parameters of Informer are as follows: the probability sparse self-attention factor is 5; the number of multi-head attention heads is 3; the number of encoders and decoders is 1; the dimension of the fully connected layer network in the decoder is 16; the dropout regularization is set to 0.1; the processing batch size batchsize is set to 128, and the number of iterations epoch is set to 50.
[0107] Introduce the coefficient of determination The normalized mean absolute error (NMAE) τ NMAE ; The normalized root mean square error (NRMSE) τ NRMSE Three error analysis indicators are used to evaluate the effectiveness of the model. Their calculation methods are as follows:
[0108]
[0109] In the formula: n is the number of samples, yi is the true value of wind power is the predicted value of wind power is the average value of wind power
[0110] To verify the effectiveness of the combined prediction model that fuses spatio-temporal features of BiTCN-Informer, the CNN, TCN, BiTCN, Transformer, Informer, and BiTCN-Transformer prediction models were selected for comparative analysis. After 50 iterations, all seven models achieved good training effects. To test the prediction ability of the seven models for unknown data, the seven trained models were used to predict the test set, and the prediction curves of the seven models are shown in Figure 2 .
[0111] As can be seen from Figure 2 a), in the intervals of sampling points 75 - 700, 1050 - 1400, 1510 - 1900, 2200 - 2400, and 2600 - 2700 of the CNN prediction curve, that is, in the intervals of fluctuating data near local peaks, the trend of the prediction curve is generally above that of the original curve as a whole, and the characteristics of this interval are not predicted, and the prediction effect for such intervals is poor. For the prediction of mutation data points in the intervals of sampling points 700 - 1000, 1900 - 2100, 2400 - 2600, and 2750 - 3000, the difference from the original curve is relatively large, and the effect is poor
[0112] As can be seen from Figure 2 b), in the intervals of fluctuating data near the above-mentioned local peaks of the TCN prediction curve, the prediction curve is slightly higher than the original curve, but the difference is not large. The prediction effect for such intervals is good. However, the prediction of mutation data points in the above intervals is insufficient, and the effect is poor
[0113] As can be seen from Figure 2 c), in the intervals of fluctuating data near the above-mentioned local peaks of the BiTCN prediction curve, the trend of the prediction curve is generally consistent with that of the original curve, and the predicted value is closer to the true value than that of TCN. Only a few points have insufficient prediction degree, and the overall prediction effect for such intervals is good. However, the prediction of mutation data points in the above intervals is insufficient, and it is slightly worse than the prediction effect of TCN
[0114] As can be seen from Figure 2 d), for the Transformer prediction curve, there is insufficient prediction in both the intervals of fluctuating data near the above-mentioned local peaks and the mutation data points, and the overall trend characteristics of the original curve are not learned, and the prediction effect is poor
[0115] As can be seen from Figure 2e) It can be seen that the Informer prediction curve shows insufficient prediction in the fluctuation data intervals near the local peaks mentioned above. Compared with the BiTCN prediction curve, it has a better prediction effect on the above-mentioned mutation data points, but there are still problems of insufficient or excessive prediction for some mutation data points, such as sampling points 720, 1010, 2450, 2750, etc.
[0116] From Figure 2 f) It can be seen that the prediction effect of the BiTCN-Transformer prediction curve on the fluctuation data intervals and mutation data points near the local peaks is not as good as that of the BiTCN and Informer prediction curves. Compared with the prediction degree of BiTCN for mutation point data, the prediction effect of BiTCN-Transformer on such data has been slightly improved, such as sampling points 720, 1010, 2450, 2750, etc.
[0117] From Figure 2 g) It can be seen that the prediction trend of the BiTCN-Informer prediction curve in the fluctuation data intervals near the local peaks at sampling points 75 - 700, 1050 - 1400, 1510 - 1900, 2200 - 2400, 2600 - 2700 is roughly the same as that of the original curve, and the prediction effect is good. The prediction effect at mutation data points such as sampling points 720, 1010, 2450, 2750 is better than that of BiTCN and BiTCN-Transformer.
[0118] In summary, Transformer has not accurately learned the overall trend characteristics of the original curve, and its prediction effect on the fluctuation data intervals and mutation data points near the local peaks is worse than that of other models. For the fluctuation data intervals near the local peaks, the overall prediction curves of CNN and TCN are higher than the original curve, and among them, TCN is closer to the original curve; the overall prediction curves of Informer and BiTCN-Transformer are lower than the original curve, and among them, Informer is closer to the original curve; and compared with BiTCN and BiTCN-Informer, TCN and Informer have a slightly worse overall prediction effect for the above intervals.
[0119] For the mutation data points of the original curve, the prediction effects of TCN, Informer, and BiTCN-Informer are better than those of other models, and among them, TCN has a slightly insufficient prediction degree for mutation points.
[0120] BiTCN-Informer has a better prediction effect on the fluctuation data intervals and mutation data points near the local peaks compared with other models.
[0121] From Figure 3It can be seen that BiTCN has a better prediction effect on the fluctuating data interval near the local peak, but a poorer prediction degree for the mutation data points; Informer has a better prediction degree for the mutation data points, but a poorer prediction effect on the fluctuating data interval near the local peak; BiTCN-Informer combines the prediction advantages of both, effectively improving the prediction accuracy for the two special types of data, and the prediction curve is closer to the original curve.
[0122] According to the data in Table 1, it can be known that the prediction model τ of the present invention R2 、τ NRMSE 、τ NMAE indicators are all optimal compared to the other 6 models, with high prediction accuracy. Compared with the CNN, TCN, BiTCN, Transformer, Informer, and BiTCN-Transformer models, the model adopted in the present invention is the closest to 1; τ NRMSE are reduced by 1.1%, 0.3%, 0.7%, 26.9%, 1.3%, and 3.8% respectively; τ NMAE are reduced by 6.9%, 5.2%, 5.9%, 40.7%, 1.5%, and 13.3% respectively.
[0123] Table 1
[0124]
[0125]
[0126] Obviously, the above-mentioned embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A short-term wind power combined prediction method based on BiTCN-Informer, characterized in that the steps Including: S1. Perform standardized preprocessing on the data; S2. Use a bidirectional temporal convolutional neural network to analyze the information features of past observations and future covariates, and extract the spatial features of the data; S21. The forward TCN network encodes the future covariates of the time series; Classification information a cat After being processed by the embedding layer, it is combined with the historical covariate a cov to obtain the information input x of all past and future covariates cov , x cov is processed through a densely connected layer to obtain h cov , h cov is fed into a forward TCN network with a two-layer structure to predict all future information to obtain o cov , h cov and the o output by the single-layer TCN network Ncov are combined to obtain the hidden state information h cov1 ; S22. The backward TCN network encodes the past observations and covariates of the time series; Lagged target value y lag Combined with the historical covariate a cov and the classification information processed by the embedding layer to obtain the input x lag , x lag Processed by the densely connected layer to obtain h lag , h lag Sent to the backward TCN network with a two-layer structure to process past information to obtain o lag , h lag And the o output by the single-layer TCN network Nlag Combined to obtain the hidden state information h lag1 ; S23. Use the dense connection layer to combine future information o cov and past information o lag to obtain the spatial features of the original sequence; S3. Adopt the Informer probability sparse self-attention mechanism and self-attention distillation mechanism to simplify network parameters, reduce computational complexity, identify the dependency relationships of long sequence data, and capture the temporal features of the sequence; S31. After the input data is processed by the probability sparse self-attention mechanism, output a "positive" Q matrix with high correlation with K; S32. After all input data is processed by the convolutional layer and the embedding layer, input it into the self-attention distillation mechanism to extract the important features of all data; S33. After the partially masked input data is processed by the convolutional layer and the embedding layer, input it into the self-attention distillation mechanism to extract the important features of the partially masked data; S34. Combine the important features output from the all data obtained in S32 and the partially masked data obtained in S32 to obtain a feature map, and obtain the K and V matrices from the feature map; S35. Use multi-head attention to process the Q, K, and V matrices to obtain weight allocation to obtain the temporal features of the data; S4. Use a fully connected layer to integrate the spatial features and temporal features, and output the prediction result o.
2. The ultra-short-term wind power combined prediction method based on BiTCN-Informer according to claim 1, wherein The formula for the standardized preprocessing is: In the formula, z is the standardized result, x is the input observation value, μ is the total data mean, and σ is the total data standard deviation.
3. The ultra-short-term wind power combined prediction method based on BiTCN-Informer according to claim 1, wherein, The bidirectional temporal convolutional neural network includes two TCN networks, and the TCN network includes a dilated convolutional network, a GELU activation function, dropout random inactivation regularization, and a densely connected layer network.
4. The ultra-short-term wind power combined prediction method based on BiTCN-Informer according to claim 1, characterized in that The expression of the forward dilated convolutional network is: The expression of the backward dilated convolutional network is: Where, G(t) is the convolution value of the t-th element, k is the number of filters, n is the element ordinal number, g(n) is the filtering value of element n, is the input value of the t-th element translated d n to the right, is the input value of the t-th element translated d n to the left.
5. The ultra-short-term wind power combined prediction method based on BiTCN-Informer according to claim 1, characterized in that, The expression of the GELU activation function is: In the formula, G(x) is the value after GELU activation processing, and P(X≤x), Φ(x) are the cumulative distribution functions of the standard normal distribution.
6. The ultra-short-term wind power combined prediction method based on BiTCN-Informer according to claim 1, wherein The expression of the probability sparse self-attention mechanism is: where softmax is the normalization function, is the query matrix with a high degree of correlation, K is the key matrix, V is the value matrix, and d is the vector dimension.
7. The ultra-short-term wind power combined prediction method based on BiTCN-Informer according to claim 1, wherein The self-attention distillation mechanism includes attention blocks with a multi-layer structure, a convolutional layer, and a pooling layer.
8. The ultra-short-term wind power combined prediction method based on BiTCN-Informer according to claim 1, wherein The expression of the self-attention distillation mechanism is: Wherein, MaxPool(·) is the max pooling operation, ELU(·) is the exponential linear unit, Convld(·) is the 1D convolution operation, and [·] AB is the attention block.
9. The ultra-short-term wind power combined prediction method based on BiTCN-Informer according to claim 1, wherein The process of all input data and partially masked input data through the convolutional layer and the embedding layer is a generative inference decoding process, expressed as: Wherein, is the decoded input sequence; is the known sequence of the start token; is the predicted sequence with a padding value of 0.
10. A very short-term wind power combined prediction model based on BiTCN-Informer, characterized in that, Including a bidirectional temporal convolutional neural network and an Informer network, the bidirectional temporal convolutional neural network and the Informer network are combined and output through a fully connected layer; The bidirectional temporal convolutional neural network includes a forward TCN network and a backward TCN network, and the TCN network consists of a dilated convolutional network, a GELU activation function, dropout random inactivation regularization, and a densely connected layer network; The Informer network includes a probability sparse self-attention mechanism and a self-attention distillation mechanism, the probability sparse self-attention mechanism and the self-attention distillation mechanism are combined through multi-head attention, and the self-attention distillation mechanism includes multiple attention blocks, and there are convolutional layers and max pooling layers between the attention blocks.
Citation Information
Patent Citations
Ultra-short-term power prediction method for wind power plant
CN118213974A