A wind speed time series prediction method based on algorithm optimization combined data denoising
By combining wavelet threshold filtering and the deep learning model BiTCN-BiGRU, and using the NOA optimization algorithm and phase space reconstruction technology, the problems of improper parameter settings and loss of nonlinear information in the wind speed prediction model are solved, and more accurate and robust wind speed time series prediction is achieved.
Patent Information
- Application Number
- CN202411645935.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Existing wind speed prediction models suffer from oversmoothing of data due to improper parameter settings, loss of nonlinear features, and sensitivity of combined system models to noise and outliers, making them difficult to adapt to complex and irregular wind speed data.
By combining wavelet threshold filtering and the deep learning model BiTCN-BiGRU, and optimizing the adjustment factor and parameters through the Nova Optimization Algorithm (NOA), and combining phase space reconstruction and quantile regression (QR), a complete prediction framework is formed.
It improves the accuracy and robustness of wind speed forecasts, better captures the uncertainty and volatility in the data, provides forecasts with multiple quantiles, reduces redundancy in the forecast interval, and enhances the model's generalization ability and robustness.
Smart Images

Figure CN119578622B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wind speed data processing technology, specifically to a wind speed time series prediction method based on algorithm optimization and joint data denoising. Background Technology
[0002] Scholars have proposed the concept of wavelet thresholding based on wavelet transform. Because this method yields the best estimate in Besov space, while no other linear estimate achieves the same result, thresholding has attracted widespread attention worldwide. Wavelet threshold functions are divided into hard threshold functions and soft threshold functions. It is generally believed that wavelet coefficients below the threshold are generated by noise, while those above the threshold are generated by valid signals. Existing technologies include applying wavelet thresholding to wind speed data and combining it with a multivariate LSTM model for prediction, significantly improving prediction accuracy. Other methods employ wavelet denoising algorithms to process wind power data and utilize an improved slime mold optimization algorithm to determine the parameters of the support vector machine, significantly improving the performance of the prediction model. Some scholars have proposed a short-term wind speed prediction model based on WT and CNN, which, through decomposition, denoising, and prediction reconstruction, achieved the lowest error in experiments, significantly outperforming other models. Still others decompose the data into several subsequences, then use wavelet thresholding to process the highest frequency sequence, and finally use LSTM for prediction; this method achieves even higher prediction accuracy. However, the discontinuity of the hard threshold function, the discontinuity of the derivative of the soft threshold function, and the constant deviation between the estimated wavelet coefficients and the wavelet coefficients of the processed signal are precisely the defects that greatly limit the further application of these two methods.
[0003] To overcome the shortcomings of traditional soft and hard thresholding in signal denoising, a wind speed prediction model integrating wavelet soft thresholding and GRU has been proposed in existing technologies. This model improves accuracy by eliminating redundant information and optimizing GRU parameters. Specifically, the signal is first decomposed, and continuous and differentiable wavelet thresholding is applied to high-frequency components, while low-frequency components below a fixed threshold are removed. Experiments show that this model has higher accuracy. Furthermore, two adjustment factors are introduced into the WT threshold function. Parameter tuning ensures the continuity of the function at the threshold point and also solves the problem of wavelet coefficient deviation, significantly improving the signal-to-noise ratio. However, this also raises a new problem: improper adjustment factor settings can lead to excessive smoothing of the WT data, causing the loss of nonlinear features in the original data. To address this issue, researchers have proposed an improved wavelet threshold function with adjustable parameters, using a particle swarm optimization algorithm to find the optimal value of the improved threshold function in the background noise, achieving good filtering results, as shown in Table 1-1. Therefore, combining algorithm optimization and data processing (phase space reconstruction technology) can maximize the preservation of the original signal information and nonlinear features.
[0004] Table 1-1 WT Model
[0005]
[0006] While point prediction methods provide results, they fail to reflect the random uncertainties of wind power generation, offering limited information and making it difficult to assess the reliability and error range of the prediction results. Filtering effectively removes high-frequency noise and random fluctuations, revealing clearer trends and structures in the data, which is particularly beneficial for interval prediction. Currently, interval prediction generally falls into two categories. The first type first performs point predictions, then constructs prediction intervals based on these results. Some researchers have integrated three models using a random forest model to construct a combined prediction model, extending point prediction to interval prediction through fuzzy information granulation. Others have used multi-objective optimization algorithms and fuzzy information granulation to construct wind speed interval prediction systems, simultaneously improving the accuracy of both point and interval predictions. Additionally, some methods convert historical data into interval data using interval construction methods, applying bivariate empirical mode decomposition techniques and least squares support vector machines to predict the upper and lower limits of the intervals. However, this type of interval prediction method relies excessively on the performance of the point prediction model and is susceptible to noise or outliers. The second type fits the prediction interval in a probabilistic statistical manner. Parameterization methods construct prediction intervals by fitting a certain probability distribution function to generate the probability distribution of trajectory points. Some scholars have used the t-location-scale function to describe the probabilistic characteristics of wind power prediction errors and established error models for probabilistic prediction based on this. Others have used normal exponential smoothing and moving kernel density estimation to estimate the probability distribution of prediction errors, and then used entropy weighting to obtain the final prediction interval. Still others have used empirical distribution models to fit the probability distribution of errors, and then performed Monte Carlo sampling to obtain the corresponding prediction interval. However, parametric methods rely on pre-defined data distributions and lack flexibility when dealing with complex and irregular data, making it difficult to adapt to the multimodal or asymmetric nature of the data. Non-parametric methods can perform interval predictions without making any assumptions about the distribution. Some scholars have used kernel density estimation to expand the prediction interval and used velocity-constrained multi-objective particle swarm optimization to adjust the upper and lower bounds of the interval. Other scholars have proposed a wind power prediction interval construction method based on the TCN conformal quantile regression (CQR) algorithm, which significantly improves prediction accuracy. Although QR quantile regression is a flexible and efficient prediction method, the prediction results depend on the selected regression model and require reasonable model settings and parameter adjustments. Therefore, some scholars have proposed an interval prediction model that combines deep learning and QR quantile regression. By optimizing parameters through algorithms, the coverage and accuracy of interval predictions have been significantly improved, exhibiting higher robustness and narrower average bandwidth. Despite the advantages of combined system model predictions, parameter setting and loss of nonlinear information remain two prominent issues. Summary of the Invention
[0007] The purpose of this invention is to provide a wind speed time series prediction method based on algorithm optimization and joint data denoising, which combines wavelet threshold filtering (WT) and deep learning model (BiTCN-BiGRU) to solve the problems of parameter setting and loss of nonlinear information in the prediction of combined system models proposed in the background art.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] A wind speed time series prediction method based on algorithm-optimized joint data denoising includes the following steps:
[0010] S1: The NOA optimization algorithm is used to optimize the two adjustment factors of the wavelet threshold filter WT, and the NOA-WT model is applied to filter the wind speed data.
[0011] S2: Reconstruct the phase space of the processed data, calculate the Lyapunov exponent, and identify its chaotic characteristics;
[0012] S3: Use the NOA-optimized deep learning model BiTCN-BiGRU for quantile regression QR interval prediction.
[0013] Furthermore, the wavelet threshold function in S1 is divided into a hard threshold function and a soft threshold function. Wavelet coefficients smaller than the threshold are generated by noise, while wavelet coefficients larger than the threshold are generated by valid signals. The calculation formula is shown below:
[0014] Hard threshold function:
[0015]
[0016] Soft threshold function:
[0017]
[0018] Among them, W j,k The wavelet coefficients after thresholding, w j,k These are the wavelet decomposition coefficients.
[0019] Furthermore, the functional expression for optimizing the WT filter model using NOA in S1 is as follows:
[0020]
[0021] Here, α and β are adjustment factors. By adjusting these two parameters, the deviation of the signal in the threshold function is reduced, thereby obtaining a better filtering effect.
[0022] Furthermore, the condition for phase space reconstruction in S2 is: d > 2D + 1, where d is the embedding dimension and D is the system correlation dimension. The time delay and embedding dimension are estimated using the correlation integral. Considering both τ and d, the correlation of the time series is obtained through correlation integral analysis of the embedded time series, leading to the statistic. S cor (τ) and according to S cor (τ) To obtain the optimal delay time τ, we need to consider the relationship between τ and τ. d and embedded window τ w Finally, the embedding dimension d is calculated.
[0023] Furthermore, the Bidirectional Temporal Convolutional Network (BiTCN) in the S3 deep learning model utilizes Convolutional Neural Networks (CNNs) to efficiently process time-series data. Specific methods include:
[0024] 1. Dilated Convolution: By stacking diluted kernels with a dilation factor d, a certain number of input data points are skipped during the convolution operation, allowing the convolution kernel to expand the receptive field without increasing computational cost;
[0025] 2. GELU activation function: Using Gaussian error linear units instead of the traditional ReLU activation function, the model returns some small negative values, thereby improving the model's learning ability;
[0026] 3. Dropout layer: Prevents overfitting of the network by randomly discarding the outputs of some neurons, thereby improving the model's generalization ability.
[0027] Furthermore, the BiGRU (Bidirectional Gated Recurrent Unit) in the deep learning model of S3 is a recurrent neural network architecture based on the GRU (Gated Recurrent Unit) for processing sequential data. The calculation process of the GRU is as follows:
[0028] 1. Update gate z t This determines how much of the current time step's state comes from past information and how much from the current input, allowing the network to selectively retain information about the variables.
[0029] z t =σ(W z ·[h t-1 ,x t ]+b z )
[0030] 2. Reset the door r t This determines how the current input is combined with past information. If the value is close to 0, it means that the network will discard past hidden states and only use the current input for updates.
[0031] r t =σ(W r ·[h t-1 ,x t ]+b r )
[0032] 3. Hidden state Adjust the past state according to the reset door, and update door z. t Under its control, new inputs are combined with past hidden states to generate new hidden states:
[0033]
[0034] In the formula, x t Here, σ represents the input data at the current time step, W represents the weight matrix of the update gate, and h represents the input data at the current time step. t-1 represents the hidden state at the previous time step, and b represents the bias term;
[0035] In GRU, the hidden state h t Considering only information from previous time steps and the current input, BiGRU uses two independent GRU networks to process the time series in the forward and reverse directions, respectively.
[0036] Furthermore, the method of using NOA to optimize the deep learning model BiTCN-BiGRU in S3 includes: using the NOA algorithm to optimize four parameters in the BiTCN model, namely the number of filters, the number of neurons in the BiGRU unit, and the learning rate and regularization parameter in the combined model, to find a suitable combination of hyperparameters in a complex search space.
[0037] Furthermore, in S3, quantile regression (QR) is a statistical method used to estimate the relationship between the conditional quantiles of the dependent variable and the independent variables. The goal of QR is to predict a specific quantile of the dependent variable by minimizing the following asymmetric loss function to estimate the regression coefficients at different quantiles:
[0038]
[0039] The above formula can be equivalent to:
[0040]
[0041] Where, ρ τ (u) = u(τ-I(u<0)), where I(Z) is the indicator function, which means that when u is less than zero, it returns 1, otherwise it returns 0; the two models optimized by NOA are combined with phase space reconstruction technology and QR quantile regression to form a complete prediction framework.
[0042] Compared with the prior art, the beneficial effects of the present invention are:
[0043] 1. The wind speed time series prediction method based on algorithm optimization and joint data denoising of the present invention uses the NOA algorithm to optimize the parameters of the WT filtering model and the BiTCN-BiGRU deep learning model to ensure that the model can obtain the minimum error and better generalization ability in different datasets.
[0044] 2. The wind speed time series prediction method based on algorithm optimization and joint data denoising of this invention combines QR quantile regression with BiTCN-BiGRU to provide predictions for multiple quantiles, avoiding the impact of poor model prediction performance on the prediction interval. Simultaneously, the model can better capture the uncertainty and volatility in the data, further improving its ability to capture complex time-series nonlinear features, resulting in more accurate and robust prediction results.
[0045] 3. The accurate wind speed prediction of this invention plays a crucial role in wind power generation, directly affecting the optimization of power generation efficiency and energy management. Wind farms can rationally plan the operation of wind turbine generators, maximize power generation, reduce the impact of wind speed fluctuations on grid stability, and effectively reduce operating and maintenance costs. Attached Figure Description
[0046] Figure 1 This is a flowchart of the NOA optimization WT parameter flowchart of the present invention;
[0047] Figure 2 This is a diagram of the BiTCN model framework of the present invention;
[0048] Figure 3 This is a structural diagram of the GRU model of the present invention;
[0049] Figure 4 This is a structural diagram of the BiGRU model of the present invention;
[0050] Figure 5 This is a flowchart of the NOA optimization process for BiTCN-BiGRU parameters in this invention.
[0051] Figure 6 This is a framework diagram of the prediction method of the present invention;
[0052] Figure 7 This is the wind speed time series plot for data1 of this invention;
[0053] Figure 8 This is the wind speed time series plot for data2 of this invention;
[0054] Figure 9 This is the NOA optimization curve based on data1 in this invention;
[0055] Figure 10This is a graph showing the NOA-WT filtering results based on data1 in this invention;
[0056] Figure 11 This is a comparison chart of the interval predicted values and actual values based on data1 in this invention;
[0057] Figure 12 This is the NOA optimization curve based on data2 in this invention;
[0058] Figure 13 This is a graph showing the NOA-WT filtering results based on data2 in this invention;
[0059] Figure 14 This is a comparison chart of the predicted and actual values based on data2 in this invention;
[0060] Figure 15 This is the prediction graph of the unidirectional deep learning model based on data1 in this invention;
[0061] Figure 16 This is the prediction graph of the bidirectional deep learning model based on data1 in this invention;
[0062] Figure 17 This is a comparison chart of ablation experiment predictions based on data1 in this invention;
[0063] Figure 18 This is a comparison chart of robust predictions based on data1 in this invention;
[0064] Figure 19 This is a diagram showing the SSA-WT filtering results based on data1 in this invention;
[0065] Figure 20 This is a graph showing the NOA and SSA interval prediction results based on data1 in this invention;
[0066] Figure 21 This is the prediction graph of the unidirectional deep learning model based on data2 in this invention;
[0067] Figure 22 This is the prediction graph of the bidirectional deep learning model based on data2 in this invention;
[0068] Figure 23 This is a comparison chart of ablation experiment predictions based on data2 in this invention;
[0069] Figure 24 This is a comparison chart of robust prediction based on data2 in this invention;
[0070] Figure 25 This is a diagram showing the SSA-WT filtering results based on data2 in this invention;
[0071] Figure 26This is a graph showing the NOA and SSA interval prediction results based on data2 in this invention. Detailed Implementation
[0072] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0073] This invention combines wavelet threshold filtering (WT) and a deep learning model (BiTCN-BiGRU) to propose a wind speed time series prediction method based on algorithmic optimization and joint data denoising, comprising the following steps:
[0074] S1: The Nutcracker optimizer algorithm (NOA) is used to optimize the two adjustment factors of the wavelet threshold filter (WT), and the NOA-WT model is applied to filter the wind speed data. The Nutcracker optimizer algorithm (NOA) is evaluated using 23 classic benchmark functions, the CEC2014 test set, the CEC2017 test set, and the CEC2020 test set, as well as 5 engineering problems. It is compared with three existing optimization algorithms. The experimental results show that NOA has outstanding advantages among all methods and has the best overall performance.
[0075] In wavelet thresholding (WT), the wavelet threshold function is divided into a hard threshold function and a soft threshold function. It is generally believed that wavelet coefficients smaller than the threshold are generated by noise, while wavelet coefficients larger than the threshold are generated by valid signals. The calculation formula is shown below:
[0076] Hard threshold function:
[0077]
[0078] Soft threshold function:
[0079]
[0080] Among them, W j,k The wavelet coefficients after thresholding, w j,k These are the wavelet decomposition coefficients.
[0081] In the NOA-optimized WT filtering model, to accurately reflect the uncertainty of wind speed and evaluate the reliability and error range of the prediction results, this invention selects an improved WT model for data filtering. This effectively removes high-frequency noise and random fluctuations, making the data present a clearer trend and structure, which is particularly beneficial for interval prediction. The functional expression of the improved wavelet threshold function for filtering time series data proposed in this invention is as follows:
[0082]
[0083] Here, α and β are adjustment factors. Adjusting these two parameters can effectively reduce the signal deviation in the threshold function. This function is continuous at the threshold point, overcoming the defect of the hard threshold function being non-differentiable at the threshold, and satisfies the odd function condition, ensuring the symmetry of the positive and negative parts of the processed signal. In actual filtering, α and β can be dynamically adjusted according to the magnitude of the deviation and the actual situation to obtain better filtering results. Based on this, this invention proposes a new model that uses the signal-to-noise ratio as the fitness function and the NOA algorithm to optimize the WT parameters. This model can quickly and accurately process wind speed time series. The optimization process is as follows: Figure 1 As shown.
[0084] S2: The processed data undergoes phase space reconstruction, Lyapunov exponents are calculated, and chaotic characteristics are identified. Phase space reconstruction is a method that transforms a one-dimensional time series into a multi-dimensional phase space trajectory through time-delay embedding. The aim of this method is to recover attractors equivalent to the original system dynamics through embedding in a high-dimensional space and to analyze the system behavior. The reconstructed phase space retains the system's dynamic information, allowing the study of nonlinear characteristics and chaotic behavior through time series data. Takens' theorem provides theoretical support for phase space reconstruction, proving that it is feasible to reconstruct the complete dynamics of a system from a one-dimensional time series and giving the condition for reconstruction: d > 2D + 1 (where d is the embedding dimension and D is the system correlation dimension). Based on this, this invention reconstructs the phase space of time series data, uses correlation integrals to estimate the time delay and embedding dimension, considers τ and d, and uses the correlation integral analysis of the embedded time series to obtain the correlation of the time series, deriving statistical measures. S cor ( τ ) and according to S cor ( τ ) , To obtain the optimal delay time based on the relationship with τ. τ d and embedded window τ wFinally, the embedding dimension is calculated. d .
[0085] S3: Predicting QR intervals in quantile regression using the NOA-optimized deep learning model BiTCN-BiGRU; Bidirectional Temporal Convolutional Networks (BiTCN) utilizes convolutional neural networks (CNNs) to efficiently process time-series data. By simultaneously combining past and future time information, it better captures the long-term dependencies between variables, thereby improving prediction accuracy. The model structure is as follows... Figure 2 As shown, it includes:
[0086] 1. Dilated Convolution: By stacking many diluted kernels with an expansion factor d, a certain number of input data points are skipped during the convolution operation, allowing the convolution kernel to expand the receptive field without increasing computational cost;
[0087] 2. GELU activation function: Using Gaussian error linear units instead of the traditional ReLU activation function allows the model to return some small negative values, avoiding the problem of neuron "death" and thus improving the model's learning ability;
[0088] 3. Dropout layer: Prevents overfitting of the network by randomly discarding the outputs of some neurons, thereby improving the model's generalization ability.
[0089] Additionally, the Gated Recurrent Unit (GRU) is a recurrent neural network architecture specifically designed for processing sequential data. It uses a gating mechanism to control information flow and memory updates. Unlike LSTM, GRU combines the forget gate and input gate into a single update gate, simplifying the network's computation and structure. The model structure is as follows: Figure 3 As shown, the calculation process of the GRU cell is as follows:
[0090] 1. Update gate z t It determines how much of the current time step's state comes from past information and how much comes from the current input, allowing the network to selectively retain information about variables;
[0091] z t =σ(W z ·[h t-1 ,x t ]+b z )
[0092] 2. Reset the door r t This determines how the current input is combined with past information. If the value is close to 0, it means the network will discard past hidden states and only use the current input for updates.
[0093] r t =σ(W r ·[h t-1 ,x t ]+b r )
[0094] 3. Hidden state Adjust the past state according to the reset door, and update door z. t Under its control, new inputs are combined with past hidden states to generate new hidden states;
[0095]
[0096] In the formula, x t The input data at the current time step, where σ represents the sigmoid activation function, W represents the weight matrix of the update gate, and h t-1 'b' represents the hidden state at the previous time step, and 'b' represents the bias term.
[0097] In GRU, the hidden state h t By considering only information from previous time steps and the current input, BiGRU uses two independent GRU networks to process the time series in both forward and backward directions, respectively. The hidden state at each time step is a combination of the two computational results, thus ensuring that the hidden state at each time step includes both historical information and future information. The model structure is as follows: Figure 4 As shown.
[0098] Specifically, NOA optimizes the BiTCN-BiGRU interval prediction model: This embodiment of the invention uses wind speed interval prediction as an example to construct a model. To accurately predict the 95% confidence interval of wind speed, the main improvement is based on the BiTCN-BiGRU model. The NOA algorithm is used to optimize four parameters: the number of filters in the BiTCN model, the number of neurons in the BiGRU unit, and the learning rate and regularization parameters in the combined model. This finds a suitable combination of hyperparameters in a complex search space. BiTCN, through convolutional operations, can effectively handle long-term sequence data, capturing long-distance temporal dependencies with fewer layers. Furthermore, gradients propagate more easily in deep networks, avoiding the gradient vanishing problem found in RNNs, thus making the model easier to train. Compared to LSTM, BiGRU has a simplified gating mechanism, resulting in lower computational cost and faster training speed. Without sacrificing the ability to model long-term dependencies, the structure of LSTM is simplified, reducing computational cost. The prediction process of the NOA-BiTCN-BiGRU model is as follows: Figure 5 As shown.
[0099] In the above embodiments, QR quantile regression (QR) is a statistical method used to estimate the relationship between the conditional quantiles of the dependent variable and the independent variables. Its goal is to predict a specific quantile of the dependent variable, rather than predicting the mean or median. This method can provide more comprehensive information about the data distribution and is more robust to regression prediction of skewed data.
[0100] QR estimates regression coefficients at different quantiles by minimizing the following asymmetric loss function:
[0101]
[0102] The above formula can be equivalent to:
[0103]
[0104] Where, ρ τ (u) = u(τ-I(u<0)), where I(Z) is an indicator function that returns 1 when u is less than zero and 0 otherwise.
[0105] Combining the two models optimized using NOA with phase space reconstruction techniques and QR quantile regression forms a complete system framework, such as... Figure 6 As shown.
[0106] To further illustrate the above embodiments, this invention uses two datasets for nonlinear time series analysis. The first dataset (data1) is from the public dataset of the 2022 Baidu KDD CUP competition, with a time resolution of 10 minutes and a total of 1391 data points selected, as follows: Figure 7 As shown. The second dataset (data2) comes from the Sotavento Galicia wind field in Spain (www.sotaventogalicia.com), with a time resolution of 10 minutes, and contains 3312 data points, as shown. Figure 8 As shown. After testing, the selected data did not contain missing values. The first 70% of the dataset was used as training data, and the last 30% was used as test data. Matlab was used to process and analyze the dataset.
[0107] Instance analysis based on data1:
[0108] Using the data1 dataset, when initializing the NOA algorithm, the population size was set to 30, the maximum number of iterations was set to 10, and the adjustment factors α and β were set to ranges of [1, 10] and [0.1, 100], respectively. When the 5th iteration was performed, the fitness value was 29.57 and the curve tended to flatten out. Figure 9As shown. Substituting the optimal solution adjustment factors α = 10 and β = 0.1 into the WT model, setting the wavelet basis to db5 and the wavelet decomposition level to 1, the wind speed time series was filtered. The NCC value was 0.9994, indicating that the similarity between the two signals is very high. The filtering process retains as much useful information as possible, as shown in the results. Figure 10 As shown: The time delay τ = 32 and the embedding dimension d = 3 of the wind speed time series are obtained by algorithm. The Wolf method is used to determine whether the reconstructed wind speed time series has chaotic characteristics. Its Lyapunov exponent is 0.1671, so the time series is considered to be chaotic. The NOA-QR-BiTCN-BiGRU model is used for interval prediction.
[0109] The NOA algorithm is used to optimize the parameters of the BiTCN-BiGRU model to obtain reasonable parameter values. The optimized parameters are: number of filters, number of neurons, initial learning rate, and regularization parameter. The upper bound of the parameter range is [0.1, 18, 10, 0.01], and the lower bound is [0.001, 3, 1, 0.00001]. The other hyperparameters and network structure settings are shown in Table 2 below.
[0110] Table 2 Hyperparameter and Network Structure Settings
[0111]
[0112]
[0113] During the optimization process, the minimum root mean square error is used as the objective function. When the iteration count is complete, each parameter value corresponding to the minimum fitness value is saved as the optimized parameter value. The model proposed in this invention is used to perform interval prediction on the filtered signal after phase space reconstruction. The predicted values are basically consistent with the actual values, such as... Figure 11 As shown, NOA-QR-BiTCN-BiGRU exhibits the lowest prediction error, with MSE, RMSE, and MAE of only 0.2938, 0.5421, and 0.3799, respectively, demonstrating the highest prediction accuracy. Its PICP (Plan-Interval Prediction) of 0.9607 indicates that the model's predicted interval covers approximately 96.07% of the true values, boasting the highest coverage compared to other models, thus capturing more accurate data and demonstrating higher reliability. Its PIMW (Plan-Interval Prediction) of 64.0593 indicates a narrower prediction interval; more accurate predictions mean less redundancy in the prediction range. Therefore, the NOA-QR-BiTCN-BiGRU model not only possesses strong prediction interval coverage but also achieves high accuracy with a relatively small prediction interval.
[0114] Instance analysis based on data2:
[0115] Using the data2 dataset, the same parameters as in Instance Analysis 1 were set when initializing the NOA algorithm. When the third iteration was performed, the fitness value was -27.38092 and the curve began to flatten out. Figure 12 As shown.
[0116] Substituting the optimal solution adjustment factors α=10 and β=100 into the WT model, and setting the same parameters, the wind speed time series was filtered. The NCC value was 0.9990, indicating that the similarity between the two signals was very high. The filtering process aimed to preserve useful information as much as possible. The results are as follows... Figure 13 As shown.
[0117] The time delay τ = 35 and the embedding dimension d = 4 of the wind speed time series were obtained based on the algorithm. The Wolf method was used to determine whether the reconstructed wind speed time series had chaotic characteristics. Its Lyapunov exponent was 0.1862, indicating that the time series had strong nonlinearity. The NOA-QR-BiTCN-BiGRU model was used for interval prediction.
[0118] By setting the same hyperparameters and network structure, the model proposed in this invention is used to perform interval prediction on the filtered signal after phase space reconstruction. The predicted values are basically consistent with the actual values, such as... Figure 14 As shown.
[0119] As shown above, NOA-QR-BiTCN-BiGRU exhibits the smallest prediction error, with MSE, RMSE, and MAE of only 0.6391, 0.7994, and 0.5665, respectively. This indicates that the model possesses high prediction accuracy and can effectively capture trends and features in the data. The PICP (Plan-Interval Prediction) is 0.9719, indicating that the predicted interval covers approximately 97.19% of the true values, demonstrating high interval reliability and providing robust interval prediction results. The PIMW (Plan-Interval Prediction Width) is 93.0, indicating that the model's prediction interval is the narrowest among all compared models. This demonstrates that while maintaining high coverage, the model effectively narrows the prediction interval, improving both accuracy and efficiency. Therefore, NOA-QR-BiTCN-BiGRU not only possesses excellent point prediction performance but also demonstrates good balance in interval prediction, achieving a good balance between coverage and interval width.
[0120] Result analysis based on data1:
[0121] By setting the same hyperparameters and network structures for all single machine learning models and combined deep learning models, the impact of different model combinations on prediction accuracy was studied. The results are shown in Table 3. Figure 15 and Figure 16 As shown.
[0122] Table 3. Prediction and Evaluation Indicators for Different Combination Methods of the Model
[0123]
[0124]
[0125] As shown above, the combined models (QRTCN-GRU and QRBiTCN-BiGRU) exhibit average reductions in MSE, RMSE, and MAE of 65.09%, 40.21%, and 38.10%, respectively, compared to single machine learning models (QRTCN, QRGRU, QRBiTCN, and QRBiGRU), indicating that the combined models can significantly reduce prediction errors. The PICP interval coverage increased by an average of 16.66%, demonstrating that the combined models are more reliable in predicting uncertainty and better predict the probability of the true value falling within the interval. The PIMW interval width decreased by an average of 47.94%, indicating that the combined models improve coverage while also making the prediction interval more compact and thus improving accuracy. Therefore, the combined models integrate the advantages of multiple models, better capturing the complexity of data and improving prediction performance.
[0126] Compared to unidirectional deep learning models, bidirectional deep learning models show an average reduction of 4.20% in MSE, 2.41% in RMSE, and 2.26% in MAE, indicating that bidirectional models can learn from both forward and backward temporal information while still achieving some error reduction. PICP increases by an average of 22.951%, suggesting that bidirectional models are slightly more reliable in predicting intervals and can better encompass actual values. PIMW decreases by an average of 27.57%, indicating that while increasing coverage, bidirectional deep learning models maintain the compactness of prediction intervals, reducing unnecessary interval width and resulting in more accurate predictions.
[0127] Ablation experiments were conducted based on data1:
[0128] The purpose of ablation experiments is to evaluate the contribution of each component to the overall model performance by observing changes in model performance through the gradual removal of components or modules. In the system model proposed in this invention, the NOA algorithm, phase space reconstruction technique, and WT filter are removed respectively. Three comparative experiments are designed to understand the impact of each part on prediction accuracy, as shown in Table 4. Figure 17 As shown.
[0129] Table 4 Comparison of Evaluation Indicators for Ablation Experiments
[0130]
[0131]
[0132] As shown above, the NOA optimization algorithm has the greatest impact on the model, reducing errors MSE, RMSE, and MAE by 80.57%, 55.91%, and 62.40%, respectively. It also increases the average PICP interval coverage by 1.81% and reduces the average PIMW interval width by 250.78%, indicating a narrower prediction interval and thus improved prediction accuracy. Secondly, the phase space reconstruction technique reduces model errors MSE, RMSE, and MAE by 43.92%, 25.11%, and 31.37%, respectively. It increases the average PICP interval coverage by 5.1% and reduces the average PIMW interval width by 167.89%, demonstrating a significant improvement in coverage. The WT filtering technique contributes the least to the model, reducing MSE, RMSE, and MAE by 27.52%, 14.86%, and 17.09%, respectively. It increases the average PICP interval coverage by 3.8% and reduces the average PIMW interval width by 70.90%, thus improving model performance to some extent.
[0133] To further explore the improvement of model robustness by the NOA algorithm, experiments were conducted on datasets with and without data processing (filtering + phase space reconstruction). The experimental results are shown in Table 5. Figure 18 As shown.
[0134] Table 5 Comparison of Evaluation Indicators
[0135]
[0136] As shown above, the NOA optimization algorithm can reduce the model's errors (MSE, RMSE, and MAE) by an average of 23.89%, 32.40%, and 39.71%, respectively; increase the PICP interval coverage by an average of 3.33%; and reduce the average PIMW interval width by 412.28%. Furthermore, compared to the model without data processing, the NOA optimization algorithm reduces the errors of the model with data processing by 20.97%, 14.90%, and 18.05%, respectively; increases the PICP interval coverage by 10.79%; and reduces the average PIMW interval width by 53.80%. Therefore, the NOA optimization algorithm further improves the robustness of the model with data processing, significantly reducing the model's error, increasing the coverage of the prediction interval, and significantly reducing the interval width.
[0137] Comparison based on the data1 algorithm:
[0138] To demonstrate that the NOA algorithm has higher prediction accuracy in optimizing deep learning models, this invention compares it with the state-of-the-art SSA optimization algorithm. The SSA algorithm is widely used due to its superior performance and is one of the most widely applied optimization algorithms today. For a fair comparison, this invention designs experiments using both the NOA and SSA algorithms to optimize deep learning models, ensuring that the experiments are conducted under the same hyperparameters and network structure conditions to objectively evaluate the optimization effects of the two algorithms. First, the SSA algorithm is used to optimize the adjustment factor [α] of the WT filter model. , [β], obtaining the same optimal parameter combination as NOA-WT [10, 0.1], with a corresponding fitness value of -29.57. The optimal adjustment factors α = 10 and β = 0.1 were substituted into the WT model to filter the wind speed time series, and the filtering results are as follows: Figure 19 As shown.
[0139] Next, phase space reconstruction was performed on the SSA-WT filtered data. Its Lyapunov exponent is 0.1671, indicating that the time series is chaotic. Therefore, after phase space reconstruction, the SSA-QR-BiTCN-BiGRU model was used for interval prediction. The final prediction results are as follows: Figure 20 As shown.
[0140] As shown in Table 6, compared to the SSA-optimized model, the NOA-optimized model reduced the errors MSE, RMSE, and MAE by 60.10%, 36.83%, and 43.16%, respectively, indicating smaller prediction errors. The average coverage of the PICP interval increased by 7.39%, indicating that the predicted intervals covered a higher proportion of the true values and had stronger reliability. The average width of the PIMW interval decreased by 67%, meaning that the predicted intervals were more accurate and concentrated, and closer to the actual situation.
[0141] Table 6 Comparison of NOA and SSA evaluation indicators
[0142]
[0143] Analysis of data2 results:
[0144] By setting the same hyperparameters and network structures for all single machine learning models and combined deep learning models, the impact of different model combinations on prediction accuracy was studied. The results are shown in Table 7. Figure 21 and Figure 22 As shown.
[0145] Table 7 Prediction and Evaluation Indicators for Different Combination Methods of the Model
[0146]
[0147]
[0148] As shown above, the combined models (QRTCN-GRU and QRBiTCN-BiGRU) exhibited average reductions in MSE, RMSE, and MAE of 55.11%, 34.43%, and 35.89%, respectively, compared to single machine learning models (QRTCN, QRGRU, QRBiTCN, and QRBiGRU). This indicates that the combined models significantly improved prediction accuracy, better fitting the data and reducing prediction errors. The PICP interval coverage increased by an average of 9.99%, demonstrating higher reliability in interval prediction and more comprehensive coverage of the true values. The PIMW interval width decreased by an average of 101.831%, indicating that the combined models successfully narrowed the prediction interval while maintaining high coverage, improving the efficiency and accuracy of interval prediction.
[0149] Compared to unidirectional deep learning models, bidirectional deep learning models show average reductions in MSE, RMSE, and MAE of 26.46%, 16.50%, and 13.14%, respectively, indicating a significant improvement in prediction accuracy and a more comprehensive capture of data features and trends. PICP increases by an average of 15.27%, demonstrating greater reliability in interval prediction and coverage of more true values. PIMW increases by an average of 71.68%, indicating a wider interval range and potential redundancy in predictions, suggesting room for improvement in interval width optimization. Therefore, while bidirectional models offer improvements in accuracy and coverage, further optimization of interval width is needed to enhance overall efficiency while simultaneously improving prediction interval precision.
[0150] Ablation experiments were performed on data2:
[0151] In Case Study 2, an ablation experiment was also designed, maintaining the same parameter settings and network structure design. The NOA optimization algorithm, phase space reconstruction technique, and VMD data decomposition were removed respectively. Three comparative experiments were designed to understand the impact of each component on prediction accuracy, as shown in Table 8. Figure 23 As shown.
[0152] Table 8 Comparison of Evaluation Indicators for Ablation Experiments
[0153]
[0154] As shown above, the NOA optimization algorithm has the greatest impact on the model, reducing errors MSE, RMSE, and MAE by 72.14%, 47.21%, and 52.34%, respectively. It also increases the average coverage of the PICP interval by 6.05% and reduces the average width of the PIMW interval by 174.62%, indicating a significant narrowing of the prediction interval and improved coverage, resulting in a substantial improvement in the model's prediction accuracy. Secondly, WT filtering reduces the model's errors MSE, RMSE, and MAE by 40.99%, 23.18%, and 27.52%, respectively. It also increases the average coverage of the PICP interval by 4.21% and reduces the average width of the PIMW interval by 135.02%, demonstrating significant effectiveness in reducing errors and improving interval coverage, thus enhancing the accuracy of the prediction interval. Phase space reconstruction technology contributed the least to the model, reducing MSE, RMSE and MAE by 11.65%, 6.01% and 10.31%, respectively. The average coverage of PICP intervals increased by 2.59%, and the average width of PIMW intervals decreased by 138.31%, indicating limited improvement, but still able to optimize the model's predictive performance to some extent.
[0155] To further explore the improvement of model robustness by the NOA algorithm, experiments were conducted on datasets with and without data processing (filtering + phase space reconstruction). The experimental results are shown in Table 9. Figure 24 As shown.
[0156] Table 9 Comparison of Evaluation Indicators
[0157]
[0158] As shown above, the NOA optimization algorithm can reduce the model's errors MSE, RMSE, and MAE by an average of 63.97%, 40.37%, and 67.33%, respectively; increase the PICP interval coverage by an average of 6.73%; and reduce the average PIMW interval width by 126.50%. Furthermore, compared to the model without data processing, the NOA optimization algorithm reduces the errors MSE, RMSE, and MAE by 62.42%, 39.71%, and 68.79%, respectively; increases the PICP interval coverage by 6.73%; and reduces the average PIMW interval width by 126.5067%. Therefore, the NOA optimization algorithm can significantly improve prediction performance, reduce errors, and increase interval coverage under different model conditions, while effectively narrowing the prediction interval width, making the model more accurate and robust under processed data conditions.
[0159] Algorithm comparison based on data2:
[0160] First, the adjustment factor [α] of the WT filter model is optimized using the SSA algorithm. ,[β], obtaining the same optimal parameter combination [10,100] as NOA-WT, with a corresponding fitness value of -27.38092. The optimal adjustment factors α=10 and β=100 were substituted into the WT model to filter the wind speed time series, and the filtering results are as follows: Figure 25 As shown.
[0161] Next, phase space reconstruction was performed on the SSA-WT filtered data. Its Lyapunov exponent was also 0.1862, indicating that the time series was chaotic. Therefore, after phase space reconstruction, the SSA-QR-BiTCN-BiGRU model was used for interval prediction. The final prediction results are as follows: Figure 26 As shown.
[0162] Table 10 shows that, compared to the SSA-optimized model, the NOA-optimized model reduces the errors MSE, RMSE, and MAE by 10.39%, 5.34%, and 7.07%, respectively, indicating smaller prediction errors. The PICP interval coverage increases by 0.78%, demonstrating that the NOA-optimized model has better prediction interval coverage and more accurately covers the range of true values. The average PIMW interval width decreases by 81.78%, meaning that the NOA optimization algorithm not only improves prediction accuracy but also narrows the prediction interval, making the prediction results more concentrated, thereby improving the model's credibility and reliability.
[0163] Table 10 Comparison of NOA and SSA evaluation indicators
[0164]
[0165] In summary, the wind speed time series prediction method based on algorithm optimization and joint data denoising provided by this invention employs a bidirectional deep learning model that achieves better data fitting results compared to a unidirectional machine learning model. This model effectively fits the data, significantly reduces prediction errors (MSE, RMSE, MAE), and improves interval coverage (PICP), demonstrating that the bidirectional model can more comprehensively capture the dependencies in the time series. Secondly, the NOA-WT filtering model of this invention effectively reduces the errors of the deep learning model, effectively filtering out noise while preserving signal characteristics, thereby improving the model's prediction accuracy. Compared to the model without WT filtering, the NOA-WT model significantly reduces prediction errors by approximately 50%-80%, increases PICP interval coverage by approximately 4.21%, and reduces the average width of PIMW intervals by approximately 212%. Thirdly, the phase space reconstruction technology of this invention, by reconstructing the embedding dimension of the time series, can better capture the chaotic characteristics in the time series, significantly reducing prediction errors by approximately 25%-30%, making the predictions of the deep learning model more accurate and further improving the model's prediction performance. Furthermore, the NOA optimization algorithm of this invention can enhance the robustness of deep learning models, significantly reduce prediction errors when processing data, and maintain high robustness under different data processing conditions. In particular, it can improve the model's accuracy and prediction interval coverage under various model conditions. In addition, the NOA-QR-BiTCN-BiGRU model proposed in this invention exhibits lower errors, higher interval coverage, and narrower prediction intervals in various comparisons, significantly outperforming traditional optimization algorithms and other combined models, demonstrating its better comprehensive prediction ability and application value in wind speed interval prediction.
[0166] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A wind speed time series prediction method based on algorithm optimization and joint data denoising, characterized in that, Includes the following steps: S1: The NOA optimization algorithm is used to optimize the two adjustment factors of the wavelet threshold filter WT, and the NOA-WT model is applied to filter the wind speed data. S2: Reconstruct the phase space of the processed data, calculate the Lyapunov exponent, and identify its chaotic characteristics; wherein, the conditions for phase space reconstruction are: ,in, For the embedding dimension, The system correlation dimension is determined; the time delay and embedding dimension are estimated using the correlation integral, while also considering... and The correlation of time series is obtained by using the correlation integral analysis of embedded time series, and the statistical value is derived. , and ,according to , , and To obtain the optimal delay time based on the relationship and embedded window Finally, the embedding dimension is calculated. ; S3: Using the NOA-optimized deep learning model BiTCN-BiGRU for quantile regression QR interval prediction. The method of using NOA to optimize the deep learning model BiTCN-BiGRU includes: using the NOA algorithm to optimize four parameters in the BiTCN model, namely the number of filters, the number of neurons in the BiGRU unit, and the learning rate and regularization parameter in the combined model, to find a suitable combination of hyperparameters in a complex search space.
2. The wind speed time series prediction method based on algorithm optimization and joint data denoising as described in claim 1, characterized in that: In S1, the wavelet threshold function is divided into a hard threshold function and a soft threshold function. Wavelet coefficients smaller than the threshold are generated by noise, while wavelet coefficients larger than the threshold are generated by valid signals. The calculation formula is shown below: Hard threshold function: Soft threshold function: in, These are the wavelet decomposition coefficients. .
3. The wind speed time series prediction method based on algorithm optimization and joint data denoising as described in claim 2, characterized in that: The function expression for optimizing the WT filter model using NOA in S1 is as follows: in, These are the wavelet coefficients after thresholding. and As a regulating factor, by adjusting and Two parameters are used to reduce the deviation of the signal in the threshold function, thereby obtaining a better filtering effect.
4. The wind speed time series prediction method based on algorithm optimization and joint data denoising as described in claim 1, characterized in that: Bidirectional Temporal Convolutional Network (BiTCN) in the S3 deep learning model utilizes Convolutional Neural Networks (CNNs) to efficiently process time series data. Specific methods include: (1). Dilated convolution: by dilating the radix... By stacking diluted kernels, a certain number of input data points are skipped during the convolution operation, allowing the convolution kernel to expand the receptive field without increasing computational cost. (2). GELU activation function: The Gaussian error linear unit is used instead of the traditional ReLU activation function, so that the model returns some small negative values, thereby improving the model's learning ability; (3). Dropout layer: By randomly dropping the output of some neurons, the network is prevented from overfitting, thus improving the generalization ability of the model.
5. The wind speed time series prediction method based on algorithm optimization and joint data denoising as described in claim 4, characterized in that: In the S3 deep learning model, the BiGRU (Bidirectional Gated Recurrent Unit) is a recurrent neural network architecture based on the GRU (Gated Recurrent Unit) for processing sequential data. The calculation process of the GRU is as follows: (1). Update the gate This determines how much of the current time step's state comes from past information and how much from the current input, allowing the network to selectively retain information about the variables. (2). Reset the door This determines how the current input is combined with past information. If the value is close to 0, it means that the network will discard past hidden states and only use the current input for updates. (3) Hidden state Adjust the past state based on the reset door, and update the door. Under its control, new inputs are combined with past hidden states to generate new hidden states: In the formula, Input data for the current time step. This represents the sigmoid activation function. This represents the weight matrix of the updated gate. This indicates the hidden state at the previous moment. Indicates the bias term; In GRU, hidden state Considering only information from previous time steps and the current input, BiGRU uses two independent GRU networks to process the time series in the forward and reverse directions, respectively.
6. The wind speed time series prediction method based on algorithm optimization and joint data denoising as described in claim 5, characterized in that: In S3, quantile regression (QR) is a statistical method used to estimate the relationship between the conditional quantiles of the dependent variable and the independent variables. The goal of QR is to predict a specific quantile of the dependent variable by minimizing the following asymmetric loss function to estimate the regression coefficients at different quantiles: The above formula can be equivalent to: in, , Let be an indicator function, representing when Return when less than zero Otherwise return The two models optimized using NOA are combined with phase space reconstruction technology and QR quantile regression to form a complete prediction framework.