An ultra-short-term probabilistic wind power prediction method based on space-time multi-scale and K-SDW

By using spatiotemporal multi-scale and K-SDW methods to decompose wind speed sequences and combine them with quantile regression models and kernel density estimation, the problem of wind speed uncertainty in wind power forecasting is solved, high-precision probabilistic forecasting is achieved, and the management efficiency of wind power systems is improved.

CN115600500BActive Publication Date: 2026-02-06NANCHANG INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211346533.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2026-02-06
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

Existing wind power forecasting methods are unable to accurately describe the uncertainty of wind speed, leading to serious risks to the power system. Traditional deterministic forecasting models cannot effectively provide probabilistic information on wind speed.

Method used

An ultra-short-term probabilistic wind power prediction method based on spatiotemporal multi-scale and K-SDW is adopted. The wind speed is decomposed into multiple subsequences through variational mode decomposition, spatiotemporal candidate features are reconstructed, and prediction is performed by combining quantile regression model. Furthermore, a combined model is constructed by using the K forward nearest neighbor quantile dynamic sparse weighted combination algorithm and kernel density estimation.

Benefits of technology

It achieves high-precision probabilistic prediction of wind power, provides detailed uncertainty information on future wind speed, and improves the accuracy and reliability of wind power prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115600500B_ABST
    Figure CN115600500B_ABST
Patent Text Reader

Abstract

The application provides an ultra-short-term probability wind power prediction method based on space-time multi-scale and K-SDW, and comprises the following steps: S1, decomposing normalized target wind speed into multiple sub-sequences through variational mode decomposition; S2, reconstructing the sub-sequences and adjacent space wind speed sequences into space-time candidate features, and performing space-time multi-scale feature selection; and S3, predicting each sub-sequence by using a quantile regression model. The application can realize wind power prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wind power technology, and in particular to an ultra-short-term probabilistic wind power prediction method based on spatiotemporal multi-scale and K-SDW. Background Technology

[0002] China will increase its nationally determined contributions and adopt more effective policies and measures. Reducing coal consumption in the power sector is indeed an effective means of reducing carbon dioxide emissions; therefore, confidence stems not only from the reality we see but also from the fact that China's installed wind power capacity has ranked first in the world for ten consecutive years. However, the inherent stochastic fluctuations and uncertainties involved have consistently posed serious risks to the power system. Accurate wind power forecasting is a powerful tool to aid in wind power grid connection management. Therefore, much research has been dedicated to improving the accuracy of wind power forecasting. Summary of the Invention

[0003] This invention aims to at least solve the technical problems existing in the prior art, and in particular, it innovatively proposes an ultra-short-term probabilistic wind power prediction method based on spatiotemporal multi-scale and K-SDW.

[0004] To achieve the above-mentioned objectives of this invention, this invention provides an ultra-short-term probabilistic wind power prediction method based on spatiotemporal multi-scale and K-SDW, comprising the following steps:

[0005] S1, normalized target wind speed is decomposed into multiple subsequences through variational mode decomposition;

[0006] S2, reconstruct the subsequence and adjacent spatial wind speed sequences into spatiotemporal candidate features, and perform spatiotemporal multi-scale feature selection;

[0007] S3 uses a quantile regression model to predict each subsequence.

[0008] In a preferred embodiment of the present invention, step S4 is further included, which combines the prediction results of each model based on the K-forward nearest neighbor quantile dynamic sparse weighted combination algorithm, and then reorganizes the prediction results of the subsequence.

[0009] In a preferred embodiment of the present invention, step S5 is further included, which involves error correction using the RF model and obtaining the probability density function through kernel density estimation.

[0010] In summary, by adopting the above technical solutions, the present invention can achieve wind power prediction.

[0011] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0012] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0013] Figure 1 This is a schematic diagram showing the characteristics of the target wind farm of the present invention at time t.

[0014] Figure 2 This is a schematic diagram of the entire process of the model proposed in this invention.

[0015] Figure 3 This is a schematic diagram of the original wind speed sequence for four seasons in this invention.

[0016] Figure 4 This is a schematic diagram showing partial wind speed data of the target wind turbine and adjacent wind turbines in this invention.

[0017] Figure 5 This is a schematic diagram of the partial probability density prediction results obtained by the method proposed in this invention.

[0018] Figure 6 This is a schematic diagram of the one-hour deterministic prediction results in the test set of this invention.

[0019] Figure 7 This is a schematic diagram illustrating the interval coverage and comparison of the method on four test sets according to the present invention.

[0020] Figure 8 The present invention provides schematic diagrams of partial probability density function prediction results for the proposed method and the comparison method on four test sets. Detailed Implementation

[0021] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0022] 1. Introduction

[0023] 1.1. Background

[0024] China will increase its nationally determined contributions and adopt more effective policies and measures. Reducing coal consumption in the power sector is indeed an effective means of reducing carbon dioxide emissions; therefore, confidence stems not only from the reality we see but also from the fact that China's installed wind power capacity has ranked first in the world for ten consecutive years. However, the inherent stochastic fluctuations and uncertainties involved have consistently posed serious risks to the power system. Accurate wind power forecasting is a powerful tool to aid in wind power grid connection management. Therefore, much research has been dedicated to improving the accuracy of wind power forecasting.

[0025] 1.2. Literature Review

[0026] Currently, wind speed prediction can be divided into physical methods utilizing real-time weather data and statistical methods analyzing historical wind speed data, depending on the research object. Traditional statistical methods, including the Autoregressive Integrated Moving Average (ARIMA) model and Markov chains, can well describe the linear relationship of data, but are clearly insufficient for fluctuating wind speed data. With the development of artificial intelligence technology, it has been found to have excellent processing capabilities for complex nonlinear data and is widely used in wind power prediction. For example, Kong et al. proposed a simplified support vector regression wind speed prediction method, using principal component analysis for feature selection and particle swarm optimization for parameter optimization. Zhang et al. used VMD to uncover the intrinsic patterns of wind speed and used a genetic algorithm to optimize the artificial neural network. Experiments show that this model can effectively detect periodic fluctuations and improve prediction accuracy.

[0027] Wind speed variations are not only related to historical data but also interact with neighboring wind turbines. Studies by Browell et al. have found that considering the correlation between adjacent turbines can effectively improve prediction accuracy. Therefore, many researchers use data from neighboring turbines to simulate the wind speed of a target turbine. Ju et al. constructed a new feature set using time series data from neighboring and target wind farms and extracted information from it using a convolutional neural network (CNN). Hu et al. considered the geographical location and topography of the wind farm, supplementing the input variables with wind speed data from neighboring areas, and established a Gaussian process model combining different kernel functions. The results show that this model has high accuracy and efficiency.

[0028] The above prediction models are all deterministic predictions, only predicting the expected value at a given time, and cannot adequately describe the uncertainty of wind speed at that time. To obtain more information about future wind speed, proposing accurate wind speed probability prediction models is particularly important. Probabilistic predictions can be categorized by their representation into interval predictions, quantile predictions, and probability density functions. Interval predictions provide an approximate range of wind speed at a future point, quantile predictions provide the wind speed value at a certain quantile, and probability density predictions provide the probability density function of future wind speed. Among these, probability density prediction is the quantitative mean with the most information. Based on different modeling methods, probabilistic prediction models can be divided into parametric and non-parametric methods. Parametric methods assume that the predicted probability density function is a known distribution model, such as a Gaussian distribution; the model parameters only need to be estimated based on the error of the deterministic prediction, making the calculation simple. Sun et al. obtained wind power probability prediction results for multiple spatially correlated wind farms using relatively low computational costs; therefore, they only needed to assume that the wind power distribution was a Gaussian mixture model and estimate the model parameters. However, when the assumed density function is inconsistent with the error of the deterministic prediction, the prediction effect will be unsatisfactory.

[0029] Nonparametric methods, which assume no prior assumptions and directly calculate distribution functions through data-driven approaches, have attracted considerable interest from scholars in recent years. Quantile prediction models and kernel density estimation (KDE) are the most commonly used nonparametric estimation tools. Quantile prediction models are mostly derived by improving point prediction models in machine learning, such as Quantile Random Forest (QRF), Quantile Regression Neural Network (QRNN), and Support Vector Quantile Regression (SVQR). Zhang et al. proposed a parallel improved quantile model that not only improved accuracy but also reduced training time. However, in quantile regression, quantiles often overlap. To address this, Zhang et al. proposed a probability density prediction model for electricity load based on a monotonic composite quantile regression neural network and a Gaussian kernel function. Kernel density estimation is another popular nonparametric estimation method that provides more information by converting quantile predictions into density predictions. He et al. transformed quantiles obtained from support vector regression into density functions using KDE with different kernel functions, which was confirmed in three cases in Singapore.

[0030] In recent years, the combined ensemble model that combines multiple point prediction methods has attracted the interest of many scholars. However, the combined model needs to pay attention to the weight setting of each basic model. Xiao et al. optimized the weight coefficient of the combined model based on the Cuckoo Search algorithm and proved that the combined model can absorb the advantages of various models and achieve better prediction results than the single model. In

[21] , a multi-objective optimization combined prediction framework was established to predict wind speed, and the weight of multiple single models was determined by the multi-objective gray wolf optimization algorithm. In fact, some combined methods can also be used for probabilistic prediction models. Bracale et al. selected Bayesian model, Markov chain model and quantile regression method as basic models, obtained different cumulative density functions (CDF), and determined the weight of each model by minimizing the comprehensive continuous ranking probability score (CRPS). Wang et al. designed two multi-objective optimization functions to make the combined interval prediction results cover a wider range and narrower width, and proposed three improved multi-objective clothing group algorithm (TSA) to determine the weight of the two basic models. In the GEFCom2017 competition, the two top teams used a simple averaging strategy to combine the results of multiple quantiles. As far as I know, the literature on combined probabilistic prediction is still relatively limited. A more distinctive approach is that of Wang et al., who solved the problem of minimizing the quantile loss function to find the optimal weights for each quantile model, transforming the nonlinear, nonconvex optimization problem into a computationally efficient and accurate linear programming problem. However, the weights determined by these combination methods do not change over time, meaning they cannot adaptively weight the components.

[0031] 1.3. About this study

[0032] Comprehensive acquisition of uncertain data information and accurate estimation of quantiles are two crucial aspects of probabilistic modeling. To fully explore the uncertainties of different spatiotemporal data characteristics, we conducted information measurement analysis on spatiotemporal multi-scale features. For accurate modeling analysis of spatiotemporal multi-scale features, we adopted several mainstream quantile models as individual models and proposed a combined model based on K-nearest neighbor quantile weighting. To prevent quantile overlap, all predictions were rearranged. Finally, using kernel density estimation as a bridge, the quantile prediction results were transformed into probability density functions.

[0033] The main contributions of this paper are summarized as follows:

[0034] (1) Based on the autocorrelation of historical sequences, the correlation between adjacent sectors is analyzed, spatially correlated sequences are introduced to construct high-dimensional candidate features, and spatiotemporal multi-scale feature selection is performed on them.

[0035] (2) For different quantiles of different models, a new dynamic sparse quantile weighting algorithm based on k forward nearest neighbors is proposed. This algorithm can fully reflect the performance of each model at each time point, rather than a simple weighted combination, giving full play to the advantages of each algorithm and discarding models with poor prediction performance.

[0036] (3) The probability density function of KDE prediction can provide detailed information on the uncertainty of future wind speed. Through experimental analysis of point prediction, interval prediction and probability prediction results, the effectiveness and superiority of the method are verified.

[0037] The remaining sections of this paper are structured as follows: Section 2 introduces spatiotemporal multi-scale feature selection methods, VMD methods, QR model methods, and quantile-weighted combination methods; Section 3 introduces the framework and steps of the proposed methods, as well as evaluation metrics; Section 4 presents a case study of the proposed methods; Section 5 presents the prediction results of the proposed models and compares the performance of comparative models; Section 6 concludes the paper and outlines future research directions.

[0038] 2. Methods and Materials

[0039] 2.1 Spatial Correlation Analysis

[0040] Wind farms are typically built in areas with complex terrain. When forecasting wind speed, relying solely on historical data from the target turbine is often insufficient to capture sudden changes in wind speed. Furthermore, many researchers have found strong correlations between the target turbine and adjacent turbines. Moran's I and local Moran's I are commonly used to measure spatial autocorrelation. Since this analysis focuses on the local spatial correlation of the target turbine, the experiment uses the local Moran's I index, which reflects the spatial clustering of wind speeds between the target turbine and adjacent turbines. The local Moran's I statistic can be expressed as:

[0041]

[0042] I i This represents the Moran index;

[0043] z i This represents the deviation between the wind speed and the average wind speed in the i-th region;

[0044] s represents the standard deviation of wind speed in the study area;

[0045] n represents the total number of all regions in the study area;

[0046] w i,j This represents the spatial weights of the i-th and j-th wind turbines;

[0047] z jThis represents the deviation between the wind speed and the average wind speed in the j-th region.

[0048] in w i,j Let i be the spatial weight of the i-th wind turbine and j-th wind turbine. Here, the spatial weight is defined as:

[0049]

[0050] w i,j This represents the spatial weights of the i-th wind turbine (the target wind turbine) and the j-th wind turbine;

[0051] z i This indicates the deviation between the wind speed in the target area and the average wind speed.

[0052] y i This represents wind speed data for the target area;

[0053] This represents the average wind speed within the study area;

[0054] z j This represents the deviation between the wind speed and the average wind speed in the j-th region.

[0055] y j This represents the wind speed data for the j-th region;

[0056] This represents the average wind speed within the study area;

[0057] s represents the standard deviation of wind speed in the study area;

[0058] n represents the total number of all regions in the study area;

[0059] When Moran Index I i >0, while z i When I < 0, it indicates that the target wind speed is low and the adjacent wind speeds are also low; when I i >0,z i When I > 0, the target wind speed is high, and the adjacent wind speeds are also high; when I i <0, z i When I < 0, the target wind speed is high, and the adjacent wind speed is low; when I i <0,z i When I > 0, the target wind speed is low, and the adjacent wind speed is high; when I i When the value is 0, it indicates that there is no spatial correlation between the target wind turbine and its adjacent wind turbines.

[0060] 2.1.1. Multi-scale feature selection

[0061] When different wind turbines are spatially correlated, the input features include not only wind speeds at adjacent time points but also wind speeds in adjacent spatial locations. In this paper, we first reconstruct N time series into candidate features X and output variables Y, such as... Figure 1 X and Y are shown at time t, where This represents the wind speed data of the i-th wind turbine at time t. Partial autocorrelation analysis is used to determine the time lags T1 and T2 of the wind speeds of the target and neighboring wind turbines, while the Pearson correlation coefficient is used to determine the spatial correlation between each wind turbine and the target wind turbine. For the wind speed Xj of the j-th neighboring wind turbine and the wind speed of the target wind turbine Xi in n samples, the Pearson correlation coefficient is expressed as follows:

[0062]

[0063] r(X j ,X i ) represents the Pearson correlation coefficient;

[0064] n represents the total number of samples;

[0065] This represents the wind speed of the j-th fan at time t;

[0066] This represents the average wind speed of the j-th fan;

[0067] This represents the wind speed of the target wind turbine at time t;

[0068] This represents the average wind speed of the target fan;

[0069] in, This represents the average wind speed of the j-th fan. This represents the average wind speed of the target wind turbine.

[0070] To select highly relevant and low-redundancy features, the multi-scale feature selection steps in this paper are as follows:

[0071] (1) Initialization: Temporal candidate features Spatial candidate features S′←{s′1,s′2,…,s′ d}, where d is the number of spatial candidate features, d = T2 × N, and an empty set Z.

[0072] Indicates will Assign the value to T′;

[0073] This represents the set of T1 time-based candidate features;

[0074] x1 represents the first candidate feature in time;

[0075] x2 represents the second candidate feature in time;

[0076] This represents the first candidate feature at time T1;

[0077] T1 represents the number of candidate features over time;

[0078] S′←{s′1,s′2,…,s′ d} represents the combination of {s′1,s′2,…,s′ d} is assigned to S′;

[0079] {s′1,s′2,…,s′ d} represents the set of d spatial candidate features;

[0080] s′1 represents the first candidate feature in the space;

[0081] s′2 represents the second candidate feature in the space;

[0082] s′ d Represents the d-th candidate feature in the representation space;

[0083] d represents the number of spatial candidate features;

[0084] T2 represents the time lag number of the wind speed of adjacent fans;

[0085] N represents the number of fans;

[0086] (2) Apply the partial autocorrelation function (PACF) to the candidate time features to obtain strongly correlated time features T, Z←Z∪T. Z←Z∪T means assigning the union of set Z and the strongly correlated time features T to set Z;

[0087] (3) Calculate each spatial candidate feature s′ i The Pearson correlation coefficient r between the target wind speed and the target wind speed i .

[0088] (4) When r i When > γ, the corresponding s′ i Add it to the feature set S, where γ is a pre-set minimum correlation coefficient, and then S is sorted in descending order of correlation coefficient.

[0089] (5) The first feature s in S j Add to set Z, Z←Z∪{s} j}

[0090] (6) Extract all features s from S in sequence. i Calculate the correlation r between this feature and each feature in Z. If r < ri Then Z← Z∪{s i}

[0091] (7) Output the selected feature set Z.

[0092] 2.2. Variational Mode Decomposition (VMD)

[0093] VMD is a fully non-recursive adaptive variational mode decomposition model proposed by Dragomiretskiy et al. This method can decompose highly fluctuating wind speed sequences into K sub-modes uk, effectively reducing the non-smoothness and nonlinearity of the original wind speed, and exhibiting robustness and noise resistance. During the decomposition process, the optimal sub-modes are obtained by constructing and solving a variational problem. The constrained variational problem is as follows.

[0094]

[0095] min{} represents taking the minimum value;

[0096] u k This represents the k-th mode;

[0097] ω k This represents the k-th center frequency;

[0098] K represents the number of sub-modes;

[0099] This represents gradient operation;

[0100] δ(t) represents the Dirac distribution;

[0101] This represents the convolution operation;

[0102] u k (t) represents the k-th modal value at time t;

[0103] st indicates that it is restricted;

[0104] u k This represents the k-th mode;

[0105] ||||2 represents the L2 norm;

[0106] Represents the square of the L2 norm;

[0107] Where u k ,ω kLet denoted by the k-th mode and the center frequency, respectively; δ(t) be the Dirac distribution; and f(t) be the original sequence. To solve for the optimal solution of this constrained variational model, a Lagrange multiplier λ and a quadratic penalty parameter α are introduced, resulting in the unconstrained Lagrange function as follows:

[0108]

[0109] α represents the quadratic penalty parameter;

[0110] λ(t) denotes the Lagrange multiplier;

[0111] Indicates inner product operation;

[0112] We use the Alternating Direction Multiplier Method (ADMM) to solve equation (5) above and continuously update it. Until a certain requirement is met. Its update equation is:

[0113]

[0114] express Fourier transform;

[0115] This represents the k-th component after the nth iteration;

[0116] Represent the Fourier transform of f(t);

[0117] f(t) represents the original sequence;

[0118] Indicate u i Fourier transform of (t);

[0119] u i (t) represents the value of the i-th component at time t;

[0120] Represent the Fourier transform of λ(t);

[0121] λ(t) denotes the Lagrange multiplier;

[0122] α represents the quadratic penalty parameter;

[0123] ω represents the center frequency;

[0124] ω k This represents the k-th center frequency;

[0125]

[0126] ω represents the k-th center frequency after the nth iteration (n iterations); ω represents the center frequency.

[0127] express Fourier transform;

[0128] This represents the k-th component after the nth iteration (n iterations);

[0129] || 2 Represents the square of the absolute value;

[0130]

[0131] λ n+1 Represents the Lagrange multiplier in the (n+1)th iteration;

[0132] λ n Represents the Lagrange multiplier of the nth iteration;

[0133] γ represents the parameter to be updated;

[0134] Represent the Fourier transform of f(t);

[0135] f(t) represents the original sequence;

[0136] express Fourier transform;

[0137] This represents the k-th component after the nth iteration (n iterations);

[0138] Where n is the number of iterations. They are respectively u i Fourier transforms of (t), f(t), and λ(t).

[0139] 2.3. Quantile Regression Model

[0140] Traditional regression models typically estimate parameters by solving for the minimum mean squared error, thus obtaining the conditional expectation relationship between the independent and dependent variables. Quantile regression describes the regression relationship between the independent and dependent variables at different conditional quantile points, and can comprehensively describe the conditional distribution of the dependent variable. When training a quantile regression model with n samples, Xi and yi are the input and output features of the i-th sample, respectively. The parameters of the model fq(Xi,Wq) are estimated by solving for the minimum quantile loss, as shown below:

[0141]

[0142] This means finding W under the condition of minimizing the given value. q ;

[0143] n represents the total number of training samples;

[0144] Represents the quantile loss function;

[0145] y i This represents the output feature of the i-th sample;

[0146] X i This represents the input feature of the i-th sample;

[0147] f q (X i W q ) represents the function model at the q-quantile;

[0148] For each quantile q of the i-th sample, there is a corresponding prediction result. The quantile loss function is shown below.

[0149]

[0150] Represents the quantile loss function;

[0151] q represents the quantile;

[0152] y i This represents the true value of the i-th sample;

[0153] Let q represent the prediction result for each quantile q of the i-th sample;

[0154] 2.3.1. Quantile linear regression (QLR)

[0155] Quantile linear regression combines quantile loss and linear regression. Assuming a linear correlation between the independent and dependent variables, then in the k-dimensional input variables... In Chinese, the simple expression for QLR is:

[0156]

[0157] This represents the weight of the 0th independent variable at the q-quantile;

[0158] This indicates the weight of the first independent variable at the q-quantile;

[0159] This represents the weight of the k-th independent variable at the q-quantile;

[0160] This represents the value of the first independent variable in the i-th sample.

[0161] This represents the value of the k-th independent variable in the i-th sample;

[0162] X i Represents a k-dimensional input variable;

[0163] W q This represents the weight matrix at the q-quantile;

[0164] In the formula, Let be the weight of the k-th independent variable at the q-quantile. Then the above problem of minimizing quantile loss becomes:

[0165]

[0166] This means finding W under the condition of minimizing the given value. q ;

[0167] y i This represents the true value of the i-th sample;

[0168] X i Represents a k-dimensional input variable;

[0169] W q This represents the weight matrix at the q-quantile;

[0170] || represents absolute value;

[0171] i|y i >X i W q Indicates in y i >X i W q Under the conditions;

[0172] i|y i ≤X i W q Indicates in y i ≤X i W q Under the conditions;

[0173] When the quantiles take continuous values ​​between 0 and 1, the conditional distribution function of y can be obtained.

[0174] 2.3.2. Quantile Regression Forests

[0175] Random Forest (RF) is an extended variant of Bagging that integrates a large number of decision trees to form a random forest. Its main idea is to select T subsamples from the training set with replacement and choose some features for training the decision trees. Therefore, RF not only randomly selects samples but also randomly selects features. The final predicted value... It is the average of all decision tree predictions.

[0176]

[0177] This represents the predicted value of the i-th sample;

[0178] T represents the total number of subsamples selected from the training set;

[0179] Let represent the predicted value of the i-th sample in the t-th decision tree;

[0180] in Let be the predicted value of the i-th sample in the t-th decision tree.

[0181] To obtain quantile results using Random Forest (RRF), Meinshausen proposed quantile regression based on RRF in 2006. Instead of a simple average of T predicted values, it uses a weighted average to obtain the conditional distribution of the predicted values, thus yielding quantile predictions. The basic steps of the QRF algorithm are as follows.

[0182] Generate a T-tree based on the idea of ​​RF, and observe the predicted value of each leaf node of each tree instead of taking the average value.

[0183] For a given x, calculate the weight w for each tree observation. i (x,θ t ), i = 1, 2, ..., n, and then calculate the weight of each observation.

[0184] T -1 Represents the reciprocal of T;

[0185] T represents the total number of subsamples selected from the training set;

[0186] w i (x,θ t ) represents the weight of each tree observation;

[0187] The distribution function estimate of all y is calculated using the following formula (14).

[0188]

[0189] F(y|X=x) represents the distribution function estimate for all y;

[0190] n represents the number of samples;

[0191] w i (x) represents the weight of the i-th sample;

[0192] 1 represents the observed value;

[0193] {Yi≤y} represents the range of the dependent variable;

[0194] 2.3.3. Quantile Regression Neural Network (QRNN)

[0195] Quantile regression neural network is a neural network-based quantile regression model proposed by Taylor for data where there is a non-linear relationship between independent and dependent variables. It has the following expression:

[0196]

[0197] g[] is the activation function of the output layer function;

[0198] K represents the number of output nodes;

[0199] This represents the weight parameters between hidden layer nodes and the output layer;

[0200] f() represents the activation function of the hidden layer function;

[0201] J represents the hidden layer number;

[0202] This represents the weight parameters from the input layer to the hidden layer;

[0203] x i This represents the i-th sample;

[0204] in, and These are the weight vectors from the input layer to the hidden layer and from the hidden layer to the output layer, respectively. J is the number of hidden layers. f(·) and g(·) are the activation functions of the hidden layer and the output layer, respectively. To calculate the q quantile... and Transform Eq.(9) into the following question:

[0205]

[0206] In the formula, λ1 and λ2 are penalty parameters to prevent the model from overfitting.

[0207] 2.4. Quantile-weighted combination prediction model based on K-SDW

[0208] Let the quantile prediction results of s individual models be respectively Let the weights assigned to the s individual models be... Then the combined prediction result of q quantiles It can be represented as:

[0209]

[0210] This represents the combined prediction result for the q-quantile;

[0211] This represents the weights of the first single model;

[0212] This represents the weights of the second single model;

[0213] This represents the weight of the s-th individual model;

[0214] This represents the quantile prediction result of the first single model;

[0215] This represents the quantile prediction result of the second single model;

[0216] This represents the quantile prediction result of the s-th single model;

[0217] The prediction accuracy of a combined model is highly dependent on the weights of the individual models. Currently, there are two main types: fixed weights and variable weights. Fixed weights are weights that do not change over time, while variable weights do the opposite.

[0218] 2.4.1. Average weight (AW)

[0219] The average weighting strategy assigns equal weights to each model, which is computationally simple, but its performance is affected by each base model. The weights of the i-th model... The calculation is as follows:

[0220]

[0221] This represents the weight of the i-th model;

[0222] S represents the number of sub-models;

[0223] 2.4.2. Static weighted based on pinball loss (SW)

[0224] The static weighting strategy based on bouncing loss is an improvement on the average weighting, where higher-accurate models have larger weights. This paper uses bouncing loss to evaluate the accuracy of model predictions and provides the weights of the i-th model at the q-quantile.

[0225]

[0226] This represents the weight of the i-th model at the q-quantile;

[0227] Let be the bouncing loss value of the i-th model at the q-quantile;

[0228] express The reciprocal of;

[0229] S represents the number of sub-models;

[0230] in, Let be the bouncing loss value of the i-th model at the q-quantile.

[0231] 2.4.3. Sparse Dynamic Weighting Based on K-Forward Nearest Neighbors (K-SDW)

[0232] This paper proposes a sparse dynamic quantile weighting algorithm based on K-forward nearest neighbors to combine multiple quantile regression models. The algorithm learns some ideas from KNN, where the current weight is related to the weights of the previous k iterations. Based on this, a threshold ξ is set; when the weight of an individual model is less than ξ, the predicted value of that model is removed, i.e., the model's weight is set to zero, thus obtaining the final sparse weight vector. For the weights of the i-th model at time t in the q-quantile of the validation set... The solution steps are as follows:

[0233] (1) Calculate the static weights of all quantiles based on the bouncing loss using the validation set.

[0234] (2) When t=1, the weights are initialized and assigned static weights, i.e.

[0235]

[0236] When t < k, the weight at time point t is:

[0237]

[0238] This represents the static weights of the i-th model based on the bouncing loss at time t and q quantile;

[0239] This represents the dynamic weight of the i-th model at time j and the q-quantile;

[0240] When t≥k:

[0241]

[0242] in

[0243] This represents the bouncing loss value of the i-th model at time t-1 and quantile q.

[0244] S represents the number of all sub-models;

[0245] (3) Given a threshold ξ, if the weights at time t are sparse, then perform the following transformation.

[0246]

[0247] This represents the weights of the i-th model at time t, under the q-quantile. When the value is less than ξ, assign it the value 0.

[0248] (4) Calculate the variable weights at all time points in the validation set, and linearly combine S models to obtain the quantile combination prediction values. And calculate The ball loss value.

[0249] (5) Optimize the parameters using the grid search method to obtain the optimal values ​​of k and ξ.

[0250] (6) Calculate the weights of all times in the test set according to formula (22). Then, linearly combine them with S models to obtain the combined predicted value at time t of the q quantile in the final test set.

[0251] Additionally, the pseudocode for this method is as follows.

[0252]

[0253] 2.5. Random Forest Based on Error Correction

[0254] In this section, we introduce the RF from section 2.3.2 to correct errors in wind speed prediction. If F t Let be the actual wind speed value at time t. Then the true error of the q quantile at time t is: F represents the true error of the q quantile at time t; t This represents the actual wind speed at time t; Let represent the combined predicted value at time t for the q-quantile; based on the RF model, an error prediction model is established, and the expression for the prediction error at time t is:

[0255]

[0256] This represents the prediction error at time t;

[0257] f() represents a combinatorial function;

[0258] This represents the true error of the q quantile at time t-1;

[0259] This represents the true error of the q quantile at time t-2;

[0260] This represents the true error of the q quantile at time th;

[0261] Correcting wind speed forecasts using error forecast values:

[0262]

[0263] This represents the prediction result at time t for the q quantile;

[0264] This represents the combined predicted value at time t for the q-quantile;

[0265] This represents the prediction error at time t;

[0266] 2.6. Kernel density estimation (KDE)

[0267] KDE is a classic nonparametric estimation method. It can transform the different conditional quantiles obtained above into a continuous probability density function using a kernel function, without any prior assumptions. The KDE result for the t-th wind speed can be expressed as:

[0268]

[0269] This represents the KDE result for the t-th wind speed;

[0270] Q represents the Q quantiles;

[0271] h represents the smoothing parameter in the kernel function;

[0272] K() represents the kernel function;

[0273] This represents the prediction result at time t for the q quantile;

[0274] in This represents the prediction result at time t for the q quantile, where K(·) is the kernel function and h is the smoothing parameter in the kernel function, also known as the bandwidth.

[0275] 3. Proposed probability density prediction model and evaluation index

[0276] The detailed structure of the proposed spatiotemporal correlated dynamic hybrid model is as follows: Figure 2 As shown in the figure, the process mainly includes data decomposition, feature selection, individual model training, dynamic weights of K-SDW, error correction model of RF, and probability density prediction of KDE. The blue and red contents represent the training and prediction processes of the proposed model, respectively.

[0277] To more clearly express the framework of the entire model, the main steps of the method proposed in this paper are as follows:

[0278] (1) Select the data of one of the 50 wind turbines as the target data for prediction. At the same time, the wind speed data sample of the wind turbine is divided into a 50% training set, a 40% validation set and a 10% test set, and the wind speed data of the target wind turbine is normalized according to formula (26).

[0279]

[0280] x * This represents the normalized value of the wind speed data for the target wind turbine;

[0281] min represents the minimum value in the wind speed data of the target wind turbine;

[0282] max represents the maximum value in the wind speed data of the target wind turbine;

[0283] (2) In order to reduce the volatility and nonlinearity of the data, VMD is used to decompose the wind speed data of the target wind turbine into K components.

[0284] (3) Transform the time series of each component into candidate features X and output Y. Candidate features X includes time lag features and spatial lag features.

[0285] (4) Select the candidate feature X by using partial autocorrelation and Pearson correlation coefficient to select variables with strong correlation in time and space.

[0286] (5) For each of the K components in the training set, train three separate models (QR, QRF, QRNN) for each component. Then use the three models to predict each component of the validation set and the test set respectively.

[0287] (6) Use K-SDW to obtain the dynamic weights of all components in the validation set, and then calculate the weights of all components in the test set step by step to obtain the dynamic combination prediction of the quantiles of the validation set and the test set.

[0288] (7) After integrating all component results, use RF for correction.

[0289] (8) Finally, the probability density result of wind speed is obtained using KDE.

[0290] 3.2. Evaluation Indicators

[0291] 3.2.1. Evaluation of Point Prediction Results

[0292] This paper uses four statistical indicators to determine the performance of the point prediction model, including root mean square error (RMSE), mean percentage error (MAPE), mean absolute error (MAE), and percentage improvement in RMSE. The formulas for each indicator are as follows:

[0293]

[0294] RMSE represents the root mean square error;

[0295] m represents the number of samples;

[0296] y i This represents the true value of the i-th sample;

[0297] This represents the predicted value of the i-th sample;

[0298]

[0299] MAPE represents the average percentage error;

[0300] m represents the number of samples;

[0301] || indicates taking the absolute value;

[0302] y i This represents the true value of the i-th sample;

[0303] This represents the predicted value of the i-th sample;

[0304]

[0305] MAE represents the mean absolute error;

[0306] || indicates taking the absolute value;

[0307] m represents the number of samples;

[0308] y i This represents the true value of the i-th sample;

[0309] This represents the predicted value of the i-th sample;

[0310]

[0311] P RMSE Indicates the percentage improvement in RMSE;

[0312] RMSE represents the root mean square error before model improvement;

[0313] RMSE1 represents the root mean square error after model improvement;

[0314] 3.3.2. Evaluation of Interval Forecast Results

[0315] Predicted Interval Coverage Probability (PICP) and Predicted Interval Normalized Average Width (PINAW) are commonly used methods to evaluate the accuracy of interval prediction. PICP represents the probability that the actual value falls within the predicted interval. The closer this value is to 1, the better the prediction efficiency. When the upper bound of the prediction at time t is Ut and the lower bound is Lt, PICP is defined as follows:

[0316]

[0317] PICP stands for Predicted Interval Coverage Probability;

[0318] N represents the number of samples;

[0319] η t Indicates when y t ∈[L t U t When η t =1; when At that time, η t =0;

[0320] y t Indicates the actual value;

[0321] L t This represents the lower bound of the prediction at time t;

[0322] U t This represents the upper bound of the prediction at time t;

[0323] PINAW is used to calculate the average width of the prediction interval, and its expression is as follows:

[0324]

[0325] Where I represents the range of the dataset. The smaller the PINAW value, the higher the accuracy.

[0326] The objectives of the prediction interval are: (1) max(PICP), (2) min(PINAW). However, when PICP increases, the corresponding PINAW also increases, making it impossible to comprehensively evaluate the interval prediction results. Corrected Prediction Interval Accuracy (CPIA) combines these two indicators and is defined as follows:

[0327]

[0328] CPIA stands for Corrected Prediction Interval Precision.

[0329] T represents the average width of the prediction interval;

[0330] PINAW represents the normalized average width of the prediction interval;

[0331] C t Indicates the intermediate substitution amount;

[0332] Where C t Defined as:

[0333]

[0334] Where d t =U t -L t CPIA∈[0,1], the closer the value of CPIA is to 1, the higher the precision.

[0335] 3.2.3. Evaluation of Probability Prediction Results

[0336] The Average Pinball Loss Function (APL) is similar to the Mean Absolute Error (MAE), but it assigns asymmetrical weights to the negative and positive errors at each quantile.

[0337]

[0338] APL represents the average bouncing loss function;

[0339] T represents the total number of samples;

[0340] Q represents the Q quantiles;

[0341] This represents the loss function for predicting values ​​at the q-quantile;

[0342] The Continuous Ranking Probability Score (CRPS) is a comprehensive index for evaluating the accuracy of probability density models

[36] , and can be seen as an extension of MSE in the probability domain. The smaller the score, the better the model predicts. Its expression is:

[0343]

[0344] CRPS(F,y t ) represents the probability scoring function for continuous sorting;

[0345] F(y′ t ) represents the combination probability function;

[0346] I(y′ t -y t ) represents the threshold function;

[0347] Where y t It is a truth value if y′ t >y t , then I(y′ t -y t ) = 1, otherwise I(y′) t -y t ) = 0.

[0348] 4. Case Studies

[0349] 4.1 Dataset Description

[0350] The data in this paper comes from 50 wind turbines at a wind farm in Shandong Province, China. Wind speed sequences vary by season. To improve the accuracy of the method, this paper divides the target dataset, with a sampling interval of 10 minutes, into different seasons. Figure 3 We provide target wind speed data for four seasons: spring (March 23 to April 23), summer (July 1 to August 1), autumn (October 27 to November 9), and winter (January 1 to February 1). We divide the four datasets into training, validation, and test sets in a 5:3:2 ratio. The training set is used to train various sub-models and hyperparameters; the validation set is used to train the dynamic weights of the combined model and the error correction model; and the test set is used to validate the efficiency of the entire model.

[0351] 4.2 Spatiotemporal Feature Selection Analysis

[0352] To analyze the spatial correlation of wind farms Figure 4The wind speed sequences of the target turbine and adjacent turbines are presented, revealing a high degree of similarity in wind speed between the target and adjacent turbines. In addition, Local Moran's I is used to reflect the spatial similarity of wind speeds. Table 1 shows the local Moran's I at certain time points. These charts demonstrate that the target wind turbine being predicted is related not only to its own historical data but also to the historical wind speeds of neighboring turbines.

[0353] Table 1 Local Moran's I at certain time points

[0354]

[0355] 4.2.1 Multi-scale feature selection

[0356] (1) Time lag selection

[0357] In this paper, PACF is used to calculate time autocorrelation because PACF eliminates the influence of x(t-1),…,x(t-T+1) on x(t), thus obtaining a pure correlation between x(t) and x(tT). Based on previous experience, this paper sets the minimum partial correlation coefficient to 0.2 and selects corresponding lag times greater than 0.2. The number of time lags with strong correlations for each component in the four seasons is shown in Table 2:

[0358] Table 2. Number of time lags for each component in the four datasets.

[0359]

[0360] (2) Spatial lag selection

[0361] In addition to historical features, it is also necessary to determine the spatial correlation between the wind speeds of adjacent wind turbines and the target wind turbine. Experiments revealed that only the target wind turbine wind speed components IMF1 and IMF2 showed a spatial similarity greater than 0.2 across the four datasets. This indicates that only when predicting IMF1 and IMF2 do we need to include neighboring wind speeds with spatial similarity, while other components do not require this. Table 3 presents the spatial similarity results for each component across the four datasets.

[0362] Table 3 Feature selection results for each component spatial feature

[0363]

[0364]

[0365] 42: Indicates the fourth wind turbine, which lags by two time points.

[0366] 4.3 Sparse Dynamic Weights

[0367] Based on the bouncing loss, we obtained the static weights for each percentile of each component in the validation set. Table 4 shows the static weights for each component at the 50th percentile across the four seasons. The static weights represent the overall allocation of the dataset by a single model and do not change over time.

[0368] Table 4. Static weighted average of each component of the 50th quantile

[0369]

[0370] After initializing the weights at the first time point in the validation set as static weights, an optimal k value needs to be pre-set when using K-SDW to solve for the dynamic weights. The k value is set to a range of 1 to 8, and the optimal value is obtained through iterative calculations of 4. Then, the weights for each time point in both the validation and test sets are calculated layer by layer. Table 5 shows the sparse dynamic weights of the two datasets at certain time points in spring when the percentile is 50.

[0371] Table 5 shows the sparse dynamic weights of the 50th percentile at certain time points in spring.

[0372]

[0373] 5. Results and Comparisons

[0374] In this section, we first obtain the probability density function prediction results for each sample in the test set using the proposed method. Then, we verify the prediction performance of the proposed method and the comparison methods from three aspects: point prediction, interval prediction, and probability density prediction. To obtain the probability density estimate, we first reconstruct multiple time series into high-dimensional candidate features and one-dimensional output variables. Second, we extract m-dimensional input features from the candidate features using autocorrelation analysis and Pearson correlation coefficient. Finally, we input the m features into the proposed prediction model for prediction. Figure 5 To test this method, we generated 9-hour probability density function curves for four time periods: spring, summer, autumn, and winter. The red vertical lines represent the actual values. It can be observed that the actual values ​​for each season are almost always near the points of highest probability on the probability density curves. Next, we will further verify the applicability of this method through comparative experiments.

[0375] To demonstrate the superiority of the proposed probability density prediction method, we compared the performance of three single models: VMD-QRF, VMD-QLR, and VMD-QRNN. Furthermore, we compared the effects of average weight based on bouncing loss and static weighted combinations.

[0376] 5.1. Comparison of point prediction results

[0377] To evaluate the performance of the point prediction results obtained from the probability density distribution, this paper assesses the accuracy of the proposed model and comparative models using RMSE, MAPE, MAE, and PRMAE. Figure 6 The hourly prediction results for the four test stations are presented. The predicted times were 10:00 AM on April 18th (spring), 0:00 AM on July 26th (summer), 6:00 PM on November 8th (autumn), and 12:00 PM on January 29th (winter). This was chosen because the wind speeds in these four time periods are near their peak values, allowing for better differentiation and comparison of the model's prediction performance. Table 6 presents the error results of the comparative experiments for the four datasets (spring, summer, autumn, and winter).

[0378] These results show that, without error correction, the proposed method has the same error as the previous one, indicating that the median quantile accuracy is high and error correction techniques are unnecessary. Among the individual models, QLR has the highest accuracy, while QRF has the lowest, especially evident in the winter data, suggesting that QLR is better representative of this data compared to other models. Compared to other individual and combined models, although the improvement effect of this method varies across different datasets, it still outperforms the combined model of QLR and SW. The most significant impact is seen in the combined model of QRF and AW, mainly because the poor performance of QRF affects the average combination. If the performance of QRF is improved, the accuracy of AW will also improve.

[0379] Table 6. Performance Evaluation of Different Model Point Prediction Across 4 Test Sets

[0380]

[0381] 5.2 Interval Forecast Comparison Results

[0382] One aspect of wind speed uncertainty prediction is interval prediction. To evaluate the results of interval prediction, we used PICP, PINAW, and CPIA metrics to assess the interval prediction performance of each model on different datasets, as shown in Table 7. The results from the four datasets show that our proposed method has the highest interval coverage. Although the width of the prediction interval is wider than that of QRNN, the CPIA metric, which combines PICP and PINAW, is the highest. This indicates that our method not only has good prediction results but also strong generalization ability, making it suitable for most datasets. In the comparison of individual models, it is still evident that QRF has a larger prediction interval and performs the worst, while QLR and QRNN have smaller prediction intervals, making the combined metrics incomparable. Among the combined models, AW is influenced by QRF and is the least efficient model besides QRF.

[0383] Figure 7The results of 12-hour interval predictions for four datasets are presented. It can be seen that the spring data yields the best results, with most actual values ​​falling within the predicted interval and the narrowest predicted interval width. The autumn data shows the worst accuracy, especially the QRF prediction interval, where the actual values ​​at some time points are not even within the interval. This is mainly due to the higher variance of the autumn data; wind speeds can suddenly and drastically increase at certain moments, reaching more than three times the previous maximum value, leading to a larger bias in the weights obtained based on the weighted combination of forward nearest neighbors.

[0384] Table 7 Evaluation of the interval prediction performance of different models on four test sets

[0385]

[0386] 5.3 Comparison of Probability Density Prediction Results

[0387] In this section, we apply several prediction models to four datasets and obtain different probability density prediction results, such as... Figure 8 The table shows the probability density prediction results based on Gaussian kernels at certain time points. APL and CRPS are also cited to evaluate the probability density prediction performance of the comparative methods. As shown in Table 8, QRF's prediction performance is worse than other methods, while the prediction performance of the proposed method is not significantly different from that of the method without error correction. Therefore... Figure 8 The probability density function of QRF is not considered, nor is the error correction of this method taken into account.

[0388] Table 8. Performance Evaluation of Probability Density Prediction for Different Models

[0389]

[0390] As shown in the table, the proposed method delivers the best prediction performance compared to the optimal individual model and other combined models. Among the individual models, the spring QRNN exhibits the lowest APL and CRPS, resulting in the best performance, while the summer QLR shows the lowest APL and CRPS, indicating that the optimal individual model differs across datasets, further highlighting the necessity of combined models. Comparing the combined models, the AW model is significantly affected by QRF, with errors even exceeding those of the individual models QLR and QRNN. While the weights in the SW model do not change over time, its weighted combination based on bouncing ball loss allows the QRF model, which has the worst average performance, to be assigned lower weights, resulting in slightly better predictions.

[0391] The above viewpoints are from Figure 8 As can be seen from this, we also found that the probability density estimation of the comparative model in winter does not differ significantly from other seasonal datasets. However, across all datasets, the probability density function estimated by AW is the worst compared to other methods, with even the true values ​​(represented in red) exceeding the range of the probability density function.

[0392] 6. Conclusion and Future Work

[0393] This paper proposes an ultra-short-term wind power probabilistic prediction method based on spatiotemporal multi-scale features and K-SDW weighting. Multiple adjacent wind speed sequences are constructed as high-dimensional spatiotemporal multi-scale candidate and output features, fully utilizing the information from different spatiotemporal data features. Various mainstream statistical models are used to generate quantile prediction results, which are then combined using the proposed dynamic combination algorithm. In the case study, the point prediction, interval prediction, and probabilistic prediction results are evaluated, demonstrating the superiority of this method in improving prediction accuracy and providing more prediction information. Furthermore, our conclusions are as follows:

[0394] (1) The prediction results of individual models vary greatly, and the accuracy of QRF is less than half that of QRNN.

[0395] (2) Compared with other combination methods, the accuracy of combination methods with simple average weights is more affected by the individual models. When the accuracy of a single model is too low, the prediction results of the combination model will be biased towards the less efficient model.

[0396] (3) Compared with other seasons, the extreme wind speed values ​​in the autumn dataset are too far apart to predict the accurate probability density function.

[0397] Furthermore, this paper only considers the probability density function for predicting the wind speed of the target turbine. Future research could consider the probability density functions for predicting the wind speed of multiple turbines.

[0398] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for ultra-short-term probabilistic wind power forecasting based on spatiotemporal multiscale and K-SDW, characterized in that, The method comprises the following steps: S1, decomposing the normalized target wind speed into multiple sub-sequences through variational modal decomposition; S2, reconstructing the sub-sequences and adjacent spatial wind speed sequences into spatio-temporal candidate features, and performing spatio-temporal multi-scale feature selection; S3, predicting each sub-sequence by using a quantile regression model; S4, combining the prediction results of each model based on a K forward nearest neighbor quantile dynamic sparse weighted combination algorithm, and then recombining the prediction results of the sub-sequences; in step S4, comprising: Let the quantile prediction results of s single models be Let the weights assigned to the s single models be The combined prediction result of the q quantile is is expressed as: a combination prediction result representing a q-quantile; represents the weight of the first single model; represents the weight of the 2nd single model; representing the weight of the s-th single model; quantile prediction result of the 1st single model; Quantile prediction results for the 2nd single model; denotes the quantile prediction of the s-th single model; The quantile dynamic sparse weighted combination algorithm based on K forward nearest neighbor is as follows: a threshold value ξ is set, when the weight of an individual model is less than ξ, the prediction value of the model is removed, that is, the weight of the model is assigned as zero, so as to obtain a final sparse weight vector; for the weight of the i-th model at the q quantile t moment in the verification set The solving steps are as follows: (1) calculating the static weights of all quantiles based on the pinball loss according to the verification set; (2) when t = 1, the weight is initialized and assigned to the static weight, that is: When t < k, the weight of the t-th time point is: represents the static weight of the i-th model at time t, q-quantile based on the pinball loss; represents the dynamic weight of the i-th model at the j-th time, q-quantile; When t ≥ k: wherein represents the pinball loss value of the i-th model at the q-quantile at time t - 1; S represents the number of all sub-models; (3) given a threshold ξ, if the weight value at time t is sparse, the following conversion is performed: represents the weight of the ith model at time t, at the q-quantile is assigned the value 0 if it is less than ξ. (4) Obtain the variable weight of all time points in the validation set, and linearly combine S models to obtain the quantile combination prediction value and the pinball loss value of is calculated. (5) the optimal values of k and ξ are obtained by using the grid search method to optimize the parameters; (6) Calculate the weight of all time points in the test set according to formula (22), and linearly combine it with S models to obtain the combined prediction value of the q-quantile t time point in the final test set 2. The ultra-short term probabilistic wind power forecasting method based on spatiotemporal multiscale and K-SDW according to claim 1, characterized in that, Further comprising step S5, using the RF model to correct errors, and obtaining the probability density function through kernel density estimation. 3.The ultra-short term probabilistic wind power forecasting method based on spatiotemporal multiscale and K-SDW according to claim 1, characterized in that, In step S1, comprising: Constrained variational model: min{} represents taking the minimum value; u k denotes the kth modality; ω k denotes the kth center frequency; represents a gradient operation; δ(t) represents the Dirac distribution; denotes a convolution operation; u k (t) denotes the kth modality value at time t; s.t. represents limited by; u k denotes the kth modality; denotes the L2 norm; In order to solve the optimal solution of the constrained variational model, the Lagrange multiplier λ and the quadratic penalty parameter α are introduced to obtain the unconstrained Lagrange function: α represents the quadratic penalty parameter; λ(t) represents the Lagrange multiplier; denotes an inner product operation; We solve (5) using the Alternating Direction Method of Multipliers (ADMM) and update until some requirement is met, whose update equation is representing the Fourier transform of denotes the kth component after the nth cycle; F(f(t)) represents the Fourier transform of f(t); f(t) represents the original sequence; represents u i the Fourier transform of (t); u i (t) denotes the value of the i-th component at time t; Represent the Fourier transform of λ(t); α represents the quadratic penalty parameter; ω represents the center frequency; ω k denotes the kth center frequency; denotes the kth center frequency after the n-th loop; ω represents the center frequency; representing the Fourier transform of denotes the kth component after the nth cycle; || 2 denotes the square of the absolute value; λ n+1 λ represents the Lagrange multiplier of the n+1th cycle; λ n λ represents the Lagrange multiplier for the nth iteration; γ represents the update parameter; F(f(t)) represents the Fourier transform of f(t); representing the Fourier transform of 4. The ultra-short term probabilistic wind power forecasting method based on spatiotemporal multiscale and K-SDW according to claim 1, characterized in that, In step S2, comprising: For the jth adjacent fan in n samples, the wind speed X j and the wind speed of the target fan X i , the Pearson correlation coefficient expression between them is: r(X j ,X i ) represents the Pearson correlation coefficient; n represents the sample number; Vj(t) represents the wind speed of the jth fan at time t; Vj represents the average wind speed of the jth fan; Vt represents the wind speed of the target wind turbine at time t; Vtarget represents the average wind speed of the target wind turbine; wherein, Vj represents the average wind speed of the jth fan, Vt represents the average wind speed of the target fan; The multi-scale feature selection step is as follows: (1) Initialization: Temporal candidate features S' <- {s1', s'2,..., s' d}, d is the number of spatial candidate features, d = T2 x N, an empty set Z. d}, d is the number of spatial candidate features, d = T2 x N, an empty set Z. represents assigning to T'; S' <- {s1', s'2,..., s'n} d} denotes the assignment of {s1', s'2,..., s'n} to S'; d} denotes the assignment of {s1', s'2,..., s'n} to S'; T2 represents the time lag number of adjacent fan wind speeds; N represents the number of fans; (2) using the partial autocorrelation function PACF to obtain the time features T with strong correlation from the time candidate features; (3) Calculate each spatial candidate feature s i Pearson correlation coefficient r between the target wind speed i ; (4) When r i > γ, the corresponding s i ′ is added to the feature set S, where γ is a pre-set minimum correlation coefficient, and then S is sorted by the correlation coefficient from large to small. (5) The first feature s in S j Add to set Z, Z←Z∪{s} j }; (6) Take out all features s in S in turn i , calculate the correlation r between this feature and each feature in Z, if r < r i , then Z <- Z U {s i} (7) output the selected feature set Z.

5. The ultra-short term probabilistic wind power forecasting method based on spatiotemporal multiscale and K-SDW according to claim 1, characterized in that, In step S3, comprising: When training a quantile regression model with n samples, X i and y i are the input and output features of the ith sample, respectively, and the parameters of the model fq(Xi, Wq) are estimated by solving the minimization of the quantile loss as follows: W is found as the minimum under the condition that q ; n represents the total number of training samples; denotes the quantile loss function; y i represents the output feature of the i-th sample; X i Xi represents the input feature of the i-th sample; f q (X i ,W q ) represents a function model at the q-quantile; For each quantile q of the i-th sample, there is a corresponding prediction result where the quantile loss function is given by denotes the quantile loss function; q represents the quantile; y i yi represents the true value of the i-th sample; represents the predicted result corresponding to each quantile q of the i-th sample; or / and assuming a linear correlation between the independent and dependent variables, then the input variables in k dimensions The simple expression for QLR is: represents the weight of the 0th independent variable at the q-quantile; represents the weight of the 1st independent variable at the q-quantile; represents the weight of the kth independent variable at the qth quantile; Xi represents the value of the first independent variable at the i-th sample; denotes the value of the kth independent variable at the ith sample; X i represents a k-dimensional input variable; W q weight matrix representing the q-quantiles; wherein is the weight of the kth argument at the qth quantile, the quantile loss minimization problem above becomes: W is found as the minimum under the condition that q ; y i yi represents the true value of the i-th sample; X i represents a k-dimensional input variable; W q weight matrix representing the q-quantiles; || represents the absolute value; i|y i >X i W q indicates that y i >X i W q under the condition that; i|y i ≤X i W q indicates that y i ≤X i W q under the condition that; When the quantile takes a continuous value between 0 and 1, the conditional distribution function of y can be obtained; or / and the final prediction value is the average of all decision tree predictions: predi represents the prediction value for the i-th sample; T represents the total number of sub-samples selected from the training set; predt,i represents the predicted value of the i-th sample in the t-th decision tree; wherein is the predicted value for the i-th sample in the t-th decision tree; or / and g[] is the activation function of the output layer function; K represents the output node number; denote the weight parameters between the hidden layer nodes and the output layer; f() represents the activation function of the hidden layer function; J represents the number of hidden layers; representing the weight parameters from the input layer to the hidden layer; x i denotes the i-th sample; In the formula, λ1 and λ2 are penalty parameters.

6. The spatiotemporal multiscale and K-SDW based ultra-short term probabilistic wind power forecasting method according to claim 1, characterized in that, In step S5, comprising: The prediction error expression at time t is: represents the prediction error at time t; f() represents the combination function; true error representing the q-quantile at time t-1; true error representing the q-quantile at time t-2; true error representing the q-quantile at time t-h; The wind speed prediction value is corrected by using the error prediction value: denotes the prediction result at the q-quantile at time t; denotes the combined prediction value at quantile q and time instant t; represents the prediction error at time t.

7. The spatiotemporal multiscale and K-SDW based ultra-short term probabilistic wind power forecasting method according to claim 1, characterized in that, Further comprising step S6, performing index evaluation, and the evaluation indexes include root mean square error, average percentage error, average absolute error, improvement percentage of RMSE, prediction interval coverage probability, prediction interval normalized average width, modified prediction interval accuracy, average pinball loss function, continuous ranking probability score one or any combination; Root mean square error: RMSE represents the root mean square error; m represents the sample number; y i yi represents the true value of the i-th sample; predi represents the prediction value for the i-th sample; Average percentage error: MAPE represents the average percentage error; m represents the sample number; y i yi represents the true value of the i-th sample; predi represents the prediction value for the i-th sample; Average absolute error: MAE represents the average absolute error; || denotes taking the absolute value; m denotes the number of samples; y i yi represents the true value of the i-th sample; predi represents the prediction value for the i-th sample; Percentage improvement in RMSE: P RMSE represents the percentage improvement in RMSE; RMSE denotes the root mean square error before model improvement; RMSE1 denotes the root mean square error after model improvement; Prediction interval coverage probability: PICP denotes the prediction interval coverage probability; η t Indicates when y t ∈[L t U t When η t =1; when At that time, η t =0; y t Indicates the actual value; L t represents the prediction lower bound at time t; U t represents the prediction upper bound at time t; N denotes the number of samples; Normalized average width of prediction intervals: Corrected prediction interval accuracy: CPIA denotes the corrected prediction interval accuracy; T denotes the average width of prediction intervals; C t represents an intermediate replacement amount; wherein C t is defined as: where d t = U t -L t , CPIA∈[0,1]; PINAW denotes the normalized average width of prediction intervals; Average pin loss function: APL denotes the average pin loss function; T denotes the total number of samples; a loss function representing a prediction value on a q-quantile; Q denotes the Q quantiles; Probability of consecutive ordering: CRPS(F,y t ) denotes the continuous ranked probability score function; F(y t ′) denotes the combined probability function; I(y t ′-y t ) represents a threshold function; where y t is a true value if y t ' > y t , then I(y t ' - y t ) = 1, otherwise I(y t ' - y t ) = 0.

Citation Information

Patent Citations

  • Method for predicting short-term wind power probability density based on EWT quantile regression forest

    CN107704953A

  • Quantile probabilistic short-term power load prediction integration method

    CN108846517A