Wind driven generator power prediction method

By combining Bayesian wavelet packet denoising and multi-task Gaussian process model, the problems of data noise and insufficiency in wind power forecasting are solved, and high-precision prediction of wind turbine power in complex environments is achieved.

CN120601402APending Publication Date: 2025-09-05DALIAN LANXUE INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510712703.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies require a large amount of high-quality data in wind power forecasting. Data noise and insufficiency lead to low prediction accuracy, especially in wind turbines with only a small amount of operating data in complex environments. The problem of wind power forecasting has not been effectively solved.

Method used

The Bayesian wavelet packet denoising (BDWPT) method is used to reduce data noise, and the Multi-Task Gaussian Process (MTGP) model is combined with a small amount of reliable wind speed data and historical power data of the target wind turbine for prediction.

Benefits of technology

The accuracy and precision of wind power forecasting are improved, especially showing higher prediction accuracy under small data conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120601402A_ABST
    Figure CN120601402A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wind power prediction, and provides a wind driven generator power prediction method, which comprises the following steps: carrying out data preprocessing on original fan data; carrying out noise reduction processing on the preprocessed fan data by using a wavelet packet Bayesian threshold noise reduction method; performing standardization processing on the fan data after noise reduction by using a mean variance normalization method; establishing an MTGP model by taking the power as a target task and the wind speed as an auxiliary task, and training the MTGP model; and performing power prediction on the test set by using the trained MTGP model. According to the method, uncertain noise in original data is reduced through the BDWPT method, so that the data quality and the wind power prediction accuracy are improved, then the migration characteristics of the MTGP model are utilized, a large amount of reliable wind speed data and a small amount of historical power data of the target fan are fused, and the power of the target fan under the complex working condition is predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wind power prediction, and in particular to a method for predicting wind generator power. Background Art

[0002] According to the Global Wind Energy Council's "Global Wind Energy Report 2023," wind power is one of the fastest-growing sources of electricity generation. Over the next five years, global wind power grid-connected capacity is expected to reach 680 GW, with an average annual increase of 136 GW, achieving a compound annual growth rate of 15%. Accurate wind turbine power forecasting is an effective means of improving the dispatchability of power systems. However, in practical engineering applications, due to the extreme randomness of wind, poor operating conditions, poor communications, and sensor failures, training data available for analytical modeling is relatively scarce. Furthermore, this data is of poor quality and contains a lot of noise, making wind power forecasting a challenging problem. This is particularly true for wind turbines operating in complex environments with only a small amount of operating data, and the wind power forecasting problem remains unresolved.

[0003] There are generally two types of modeling methods for wind turbine power prediction: physical models and statistical methods. Physical models predict wind power using numerical weather prediction (NWP) systems based on wind farm topography and field observations. This method typically builds a model based on physical knowledge and calculates the wind turbine's output power by solving nonlinear equations. However, the model must clearly describe the wind power conversion mechanism and requires extremely high accuracy for data such as meteorological forecasts. It also involves a lot of difficult-to-quantify information, such as topography and landforms. Therefore, this method lacks competitiveness in wind power forecasting.

[0004] Statistical methods use historical data to build statistical models to predict future power. These methods primarily include time series and artificial intelligence (AI) approaches. Time series-based methods include the autoregressive moving average (ARMA) model and the autoregressive differential moving average (ARIMA) model. These methods establish a linear mapping between input and output and have been widely used for wind power forecasting. AI models, such as long short-term memory (LSTM) networks, gated recurrent units (GRU) models, and transformer models, demonstrate superior predictive performance compared to traditional statistical methods. AI models rely on large training sets to establish the mapping between input and output variables. Generally speaking, the quality of the training set data directly impacts the model's predictive accuracy. However, in actual projects, data collected by wind farms often suffers from sensor failures due to factors such as harsh operating environments, improper installation, and sampling errors. This can lead to significant deficiencies in data quality and quantity, making it difficult to guarantee the AI ​​model's predictive accuracy. Summary of the Invention

[0005] The present invention mainly solves the technical problem that the existing technology requires a large amount of high-quality data to accurately predict power. If the data contains noise and the power data is insufficient, the prediction accuracy will be low. A wind turbine power prediction method is proposed, which is based on the combination of Bayesian wavelet packet denoising (BDWPT) and multi-task Gaussian process (MTGP) to predict wind turbine power. First, the BDWPT method is used to reduce the uncertainty noise in the original data, thereby improving the data quality and the accuracy of wind power prediction. Then, the migration characteristics of the MTGP model are used to integrate a large amount of reliable wind speed data and a small amount of historical power data of the target wind turbine to predict the power of the target wind turbine under complex working conditions.

[0006] The present invention provides a method for predicting wind turbine power, comprising the following steps:

[0007] Step 1: Obtain wind turbine data during the operation of the wind turbine, perform data preprocessing on the original wind turbine data, and divide it into a training set and a test set; the wind turbine data includes wind speed and power data during the operation of the wind turbine;

[0008] Step 2: noise reduction is performed on the pre-processed wind turbine data using the wavelet packet Bayesian threshold noise reduction method;

[0009] Step 3: Use the mean-variance normalization method to standardize the noise-reduced fan data;

[0010] Step 4: Using power as the target task and wind speed as the auxiliary task, establish the MTGP model and train the MTGP model. Step 4 includes the following steps 401 to 403:

[0011] Step 401, selecting target task wind turbine power data and auxiliary task wind speed data;

[0012] Step 402, establishing an MTGP model;

[0013] Step 403: train the established MTGP model, determine the MTGP model hyperparameters, and obtain a trained MTGP model;

[0014] Step 5: Use the trained MTGP model to perform power prediction on the test set.

[0015] Furthermore, step 2 includes the following steps 201 to 203:

[0016] Step 201: performing discrete wavelet decomposition on the time series wind turbine data to obtain high-frequency coefficients and low-frequency coefficients;

[0017] After wavelet packet transform, N time series data are decomposed into scaling coefficients s j (k) and wavelet coefficient wj (k); The time series is expressed by inverse wavelet transform and discrete wavelet transform coefficients as follows:

[0018]

[0019] Where, represents the reconstructed time series data, j and k represent the kth coefficient of the jth level decomposition, φ j,k (t i ) and ψ j,k (t i ) represents the scaling function and wavelet function of the k-th coefficient of the j-th layer. The scaling function is obtained by scaling the wavelet function. The two sums represent the summation of the scaling subspace and the wavelet subspace.

[0020] The j-th level scaling factor s j (k) and wavelet coefficient w j (k) is expressed as:

[0021]

[0022] Where m = 2k + n, h(·) and h1(·) are scaling and wavelet filter coefficients, respectively, and their values ​​are obtained by setting the wavelet integral to 0 and the square integral to 1; the original data coordinate value is c j Initialize and scale the filter coefficients and wavelet filter coefficient equations to obtain the required decomposition coefficients through iteration;

[0023] Step 202, performing threshold quantization processing on the decomposition coefficients through Bayesian hypothesis test;

[0024] Assume d jk represents the kth decomposition coefficient of the jth layer. Assuming that the decomposition coefficient is contaminated by Gaussian white noise, then:

[0025]

[0026] In the above formula, represents the noise-free coefficient, ε jk ~N(0,1) represents independent and identically distributed error variables with mean 0; σ j represents the noise variance of the j-layer decomposition;

[0027] for and σ j 2 The conditional distribution of the decomposition coefficient is expressed as:

[0028]

[0029] Assumptions The unbiased prior distribution of is as follows:

[0030]

[0031] where γ jk is an independent Bernoulli distribution π j A binary random variable, that is:

[0032] P(γ jk =1) = 1-P(γ jk =0) =π j (8)

[0033] Formula (8) Determination coefficient is zero (γ jk =0) or non-zero (γ jk =1); variance σ j 2 Indicates that in the jth decomposition layer the scale of change;

[0034] According to Bayesian theory, using non-informative prior distribution, the noise reduction coefficients are obtained The posterior distribution of is as follows:

[0035]

[0036] In the variance When, through a normal distribution and the point mass δ(0) at zero, we get The marginal conditional posterior distribution of :

[0037]

[0038] in is the posterior distribution of formula (7), which is described by the following probability

[0039]

[0040] where η jk Yes jk =0 and γ jk = 1, also known as the Bayesian coefficient, is expressed as follows:

[0041]

[0042] The decomposition coefficients are thresholded according to the following hypothesis test:

[0043] H0:d jk =0vs.H1:d jk ≠0 (13)

[0044] Rightjk Apply the indicator function I(·) to obtain the decomposed coefficients

[0045]

[0046] When I(n jk <1)=1, if η in formula (12) jk <1, the null hypothesis H0 is rejected and the corresponding coefficient is retained;

[0047] Step 203, performing wavelet reconstruction on the signal to obtain the fan data after noise reduction;

[0048] The denoising decomposition coefficient obtained by formula (14) The noise-reduced signal can be obtained by using formula (2) Reconstruction processing is performed to obtain the fan data after noise reduction.

[0049] Furthermore, in step 3, the noise-reduced fan data is standardized using the following formula:

[0050]

[0051] Where X is the normalized data, i.e., the input of the subsequent MTGP model, X0 is the data after denoising in the previous step, μ is the mean of the sample, and S is the standard deviation of the sample.

[0052] Furthermore, the step 402, establishing the MTGP model, includes the following process:

[0053] The MTGP model uses Gaussian prior to model the time series. The Gaussian prior consists of the mean function and the covariance function, which is expressed as:

[0054] y=f(x)~GP(m(x),k(x,x′)) (15)

[0055] Where x and y represent the input and output data in the training set, respectively; m(x) is set to 0, and k(x, x′) is the key factor of the model, usually using the square exponential kernel function, which can be expressed as:

[0056]

[0057] Where τ is the Euclidean distance between input variables, θ f and θ l is the hyperparameter that needs to be optimized;

[0058] The hyperparameter optimization method of minimizing the negative log marginal likelihood function (NLML) is used to optimize θ and θ. l For optimization, the NLML function and the optimization function can be expressed as:

[0059]

[0060] The NLML function is a convex function, and the gradient descent algorithm is selected for effective optimization. After model training, the test set data x * The predicted value can be expressed as:

[0061]

[0062] The covariance kernel function of the MTGP model can be expressed as:

[0063]

[0064] In the formula, the matrix K c Reflects the relationship between tasks, in the form of a task label set Decide, m=2, the target task is wind turbine power, and the auxiliary task is wind speed; K t is the covariance matrix determined by the input data in the training set, represents the Kronecker product, θ c ,θ t Represents K c and K t The set of all hyperparameters in; the covariance matrix K t The elements of are kernel functions of the input data, which have a fixed form and all hyperparameters are shared by all tasks;

[0065] When using the MTGP model for time series forecasting, the target task is defined as wind turbine power, and the auxiliary task is defined as wind speed. The MTGP model can be constructed as follows:

[0066]

[0067] Where y r ∈R r×1 and y p ∈R p×1 Represents outputs of two different dimensions, y r is the target task output, y p For auxiliary task output, focus on the output result of the target task wind turbine power, x r and x p Represents the input datasets of two different tasks, x r is a small amount of historical power data of wind turbines, x p is sufficient wind speed data; θ rr ,θ rp ,θ pr and θ pp Indicates the hyperparameters in the MTGP model that need to be optimized using the NLML method;

[0068] The prediction results of the MTGP model can also be expressed as:

[0069]

[0070] Furthermore, the step 5 also includes: evaluating the MTGP model using root mean square error and determination coefficient evaluation indicators.

[0071] The present invention provides a wind turbine power prediction method, including a BDWPT-MTGP model that effectively addresses the problem of low prediction accuracy due to noise and insufficient power data. This is the first application of the BDWPT noise reduction method to wind turbine power prediction, combined with the MTGP method, and successfully applied to small-data wind turbine power prediction tasks. The BDWPT method has extraordinary potential for data noise reduction, while the MTGP model has excellent performance in small-data time series prediction. By combining the advantages of both methods, the present invention achieves higher prediction accuracy than traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 This is a flow chart of the implementation of the wind turbine power prediction method provided by the present invention.

[0073] Figure 2 It is a schematic diagram of the wavelet packet Bayesian threshold denoising method;

[0074] Figure 3 This is the power prediction timing diagram of the BDWPT-MTGP model under different training data amounts: a, 5%; b, 10%; c, 20%. DETAILED DESCRIPTION

[0075] To make the technical problems solved, the technical solutions adopted, and the technical effects achieved by the present invention more clearly apparent, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, rather than all of the contents.

[0076] like Figure 1 As shown, a wind turbine power prediction method provided by an embodiment of the present invention includes the following steps:

[0077] Step 1: Obtain wind turbine data during wind turbine operation, preprocess the original wind turbine data, and divide it into training set and test set.

[0078] The wind turbine data includes wind speed and power data during the operation of a wind turbine (referred to as wind turbine), which can be collected through data collection devices such as sensors.

[0079] The data preprocessing includes but is not limited to data quality checking, deleting negative power data and filling in blank values; the operation data of the wind turbine may contain non-numeric data and blank values ​​due to sensor failure, communication failure and other reasons.

[0080] This step first performs a quality check on the original data, deletes non-numeric data, deletes negative power data, and uses median interpolation to fill in missing values ​​to ensure data integrity.

[0081] Step 2: The pre-processed wind turbine data is subjected to denoising using the Bayesian wavelet packet threshold denoising method (BDWPT).

[0082] like Figure 2 As shown, step 2 includes the following steps 201 to 203:

[0083] Step 201 : performing discrete wavelet decomposition on the time series wind turbine data to obtain high-frequency coefficients and low-frequency coefficients.

[0084] When performing a three-layer discrete wavelet packet transform on a time series signal, the method first decomposes the original signal into an approximate coefficient A and a detail coefficient D. The approximate coefficient A is then further decomposed into AA and AD, which represent the approximate and detail coefficients of the second-level decomposition of A, respectively, while D is further decomposed into DA and DD, which represent the approximate and detail coefficients of the second-level decomposition of D, respectively. The decomposition is repeated until the preset decomposition level is reached. In this way, when the original signal is decomposed to the jth level, 2 j For example, suppose a sensor obtains N time series data f(t i )(i=1,2,3....,N), after three discrete wavelet packet decompositions, 8 decomposition coefficients are obtained. The three-layer DWPT decomposition formula is as follows:

[0085] f(t i )=AAA+DAA+ADA+DDA+AAD

[0086] +DAD+ADD+DDD(1)

[0087] After wavelet packet transform, N time series data are decomposed into scaling coefficients s j (k) and wavelet coefficient w j (k). The time series is expressed by inverse wavelet transform and discrete wavelet transform coefficients as:

[0088]

[0089] In the formula represents the reconstructed time series data, j and k represent the kth coefficient of the jth level decomposition, φ j,k (t i ) and ψ j,k (t i ) represents the scaling function and wavelet function of the k-th coefficient of the j-th layer. The scaling function is obtained by scaling the wavelet function. The two sums represent the sum of the scaling subspace and the wavelet subspace. The scaling coefficient s of the j-th level j (k) and wavelet coefficient w j (k) is expressed as:

[0090]

[0091] Where m = 2k + n, h(·) and h1(·) are scaling and wavelet filter coefficients respectively, and their values ​​are obtained by setting the wavelet integral to 0 and the square integral to 1. The original data coordinate value pair c j Initialization, scaling of filter coefficients and wavelet filter coefficient equations are iteratively used to obtain the required decomposition coefficients.

[0092] Step 202: Perform threshold quantization processing on the decomposition coefficients through a Bayesian hypothesis test.

[0093] Assume d jk represents the kth decomposition coefficient of the jth layer. Assuming that the decomposition coefficient is contaminated by Gaussian white noise, then:

[0094]

[0095] In the above formula, represents the noise-free coefficient, ε jk ~N(0,1) represents independent and identically distributed error variables with mean 0. σ j Denotes the noise variance of the j-layer decomposition. Therefore, denoising becomes a univariate nonparametric estimation problem. The purpose of the Bayesian threshold is to obtain the original noise coefficient d jk Get a set of denoising coefficients Then reconstruct the denoised signal. and σ j 2 The conditional distribution of the decomposition coefficient is expressed as:

[0096]

[0097] Assumptions The unbiased prior distribution of is as follows:

[0098]

[0099] where γ jk is an independent Bernoulli distribution πj A binary random variable, that is:

[0100] P(γ jk =1) = 1-P(γ jk =0) =π j (8)

[0101] Formula (8) Determination coefficient is zero (γ jk =0) or non-zero (γ jk =1). Variance σ j 2 Indicates that in the jth decomposition layer In practical applications, all coefficients in the jth layer are assigned the same unbiased value, so that π j =0.5 and Then the standard deviation σ is estimated by dividing the median of the wavelet coefficients of the j-th layer DWPT decomposition by a factor j Based on previous experience, this method uses a factor of 0.6745 to estimate the standard deviation.

[0102] According to Bayesian theory, using non-informative prior distribution, the noise reduction coefficients are obtained The posterior distribution of is as follows:

[0103]

[0104] In the variance When, through a normal distribution and the point mass δ(0) at zero, we get The marginal conditional posterior distribution of :

[0105]

[0106] in is the posterior distribution of formula (7), which is described by the following probability

[0107]

[0108] where η jk Yes jk =0 and γ jk = 1, also known as the Bayesian coefficient, is expressed as follows:

[0109]

[0110] The decomposition coefficients are thresholded according to the following hypothesis test:

[0111] H0:d jk=0vs.H1:d jk ≠0 (13)

[0112] Right jk Apply the indicator function I(·) to obtain the decomposed coefficients

[0113]

[0114] When I(n jk <1)=1, if η in formula (12) jk <1, the null hypothesis H0 is rejected and the corresponding coefficient is retained.

[0115] Step 203: Perform wavelet reconstruction on the signal to obtain the fan data after noise reduction.

[0116] The denoising decomposition coefficient obtained by formula (14) The noise-reduced signal can be obtained by using formula (2) Reconstruction processing is performed to obtain noise-reduced fan data for subsequent analysis.

[0117] Step 3: Use the mean-variance normalization method to standardize the noise-reduced fan data. The formula is as follows:

[0118]

[0119] Where X is the normalized data, i.e., the input of the subsequent MTGP model, X0 is the data after denoising in the previous step, μ is the mean of the sample, and S is the standard deviation of the sample.

[0120] This step standardizes the noise-reduced fan data in preparation for subsequent modeling.

[0121] Step 4: Using power as the target task and wind speed as the auxiliary task, establish the MTGP model and train the MTGP model.

[0122] Step 4 includes the following steps 401 to 403:

[0123] Step 401: Select target task wind turbine power data and auxiliary task wind speed data.

[0124] The wind speed data and wind turbine power data are obtained from the raw data (usually the operating data collected by the SCADA system of a wind farm).

[0125] Step 402: Establish an MTGP model.

[0126] The MTGP (Multi-Task Gaussian Process, multi-task Gaussian process regression) model is a multi-task learning model extended from the Gaussian process regression model (Gaussian Process Regression, GPR), and its basic principle is roughly the same as GPR.

[0127] The MTGP model uses Gaussian prior to model the time series. The Gaussian prior consists of the mean function and the covariance function, which is expressed as:

[0128] y=f(x)~GP(m(x),k(x,x′)) (15)

[0129] Where x and y represent the input and output data in the training set, respectively. In most applications, m(x) is set to 0, and k(x, x′) is the key factor of the model, usually using the square exponential kernel function, which can be expressed as:

[0130]

[0131] Where τ is the Euclidean distance between input variables, θ f and θ l is the hyperparameter that needs to be optimized. The present invention uses a hyperparameter optimization method that minimizes the negative logarithmic marginal likelihood function (NLML) to optimize θ f and θ are optimized, the NLML function and the optimization function can be expressed as:

[0132]

[0133] The NLML function is a convex function in the present invention, so the gradient descent algorithm is selected for effective optimization. After model training, the test set data x * The predicted value can be expressed as:

[0134]

[0135] The key to the MTGP model is the need for a covariance kernel function K MTGP , which can simultaneously describe the correlation between different input parameters of a single task and the relationship between different output parameters of multiple tasks. Therefore, the covariance kernel function of the MTGP model can be expressed as:

[0136]

[0137] In the formula, the matrix K c Reflects the relationship between tasks, in the form of a task label set In the present invention, m=2, the target task is the wind turbine power, and the auxiliary task is the wind speed. tis the covariance matrix determined by the input data in the training set, represents the Kronecker product, θ c ,θ t Represents K c and K t The set of all hyperparameters in . Similar to the Gaussian process, the elements of the covariance matrix K are the kernel functions of the input data, which have a fixed form and all hyperparameters are shared by all tasks.

[0138] When using the MTGP model for time series forecasting, to better illustrate the kernel matrix in MTGP, we take two tasks as an example, where the target task is defined as wind turbine power and the auxiliary task is defined as wind speed. The MTGP model can be constructed as follows:

[0139]

[0140] Where y r ∈R r×1 and y p ∈R p×1 Represents outputs of two different dimensions, y r is the target task output, y p For auxiliary task output, in this invention, we only focus on the output result of the target task wind turbine power, x and x p Represents the input datasets of two different tasks, x r is a small amount of historical power data of wind turbines, x p is sufficient wind speed data. rr ,θ rp ,θ pr and θ pp Represents the hyperparameters in the MTGP model that need to be optimized using the NLML method. When building an MTGP model, the input data is usually standardized to avoid difficulties in parameter optimization due to a large data range.

[0141] The prediction results of the MTGP model can also be expressed as:

[0142]

[0143]

[0144] Note that for time series without a trend, the mean function of the MTGP model is typically set to 0. A non-zero mean function is typically used when a clear trend is observed in the time series or when a reasonable assumption about the trend term is made. In this paper, since we do not observe a consistent power trend term within the forecast horizon, the mean function is set to 0.

[0145] Step 403: train the established MTGP model, determine the MTGP model hyperparameters, and obtain a trained MTGP model.

[0146] Use the established MTGP model to complete training and determine the MTGP model hyperparameters.

[0147] Step 5: Use the trained MTGP model to perform power prediction on the test set; and use the root mean square error (RMSE) and determination coefficient (R 2 ) evaluation indicators to evaluate the MTGP model.

[0148] The present invention is described in detail below using actual wind turbine operating data as an example:

[0149] 1. The data used is operational data collected from wind turbines at a wind farm in Jiangsu Province, China, covering the period from June 1 to December 1, 2018. The data is collected by a SCADA system with a collection frequency of 10 minutes, totaling 26,292 sample points.

[0150] 2. Perform a quality check on the original data and delete non-numeric data. Then delete the data with power less than 0. The present invention does not consider the case where the power is less than 0. Finally, take the median of the two adjacent data of the vacant value to fill the vacant value to ensure the integrity of the data.

[0151] 3. Use the BDWPT method to reduce noise on the data.

[0152] 4. Normalize the data.

[0153] 5. Build the MTGP model. The target task is a small amount of power data, and the auxiliary task is a large amount of wind speed data. To better compare the results, the wind power training data volume is set to only 5%, 10%, and 20% of the data samples, respectively.

[0154] 6. Use the established MTGP model to perform power prediction, using the root mean square error (RMSE) and determination coefficient (R 2 ) evaluation indicators to evaluate the model. Table 1 is the statistical table of model accuracy indicators, Figure 3 This is the model's prediction when the amount of power data is 5%, 10%, and 20% of the data sample.

[0155] Table 1 Statistics of model prediction accuracy indicators

[0156]

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications to the technical solutions described in the above embodiments, or equivalent replacement of some or all of the technical features therein, do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting wind turbine power, characterized in that: The following processes are included: Step 1: Obtain wind turbine data during the operation of the wind turbine, perform data preprocessing on the original wind turbine data, and divide it into a training set and a test set; the wind turbine data includes wind speed and power data during the operation of the wind turbine; Step 2: noise reduction is performed on the pre-processed wind turbine data using the wavelet packet Bayesian threshold noise reduction method; Step 3: Use the mean-variance normalization method to standardize the noise-reduced fan data; Step 4: Using power as the target task and wind speed as the auxiliary task, establish the MTGP model and train the MTGP model. Step 4 includes the following steps 401 to 403: Step 401, selecting target task wind turbine power data and auxiliary task wind speed data; Step 402, establishing an MTGP model; Step 403: train the established MTGP model, determine the MTGP model hyperparameters, and obtain a trained MTGP model; Step 5: Use the trained MTGP model to perform power prediction on the test set.

2. The wind turbine power prediction method according to claim 1, characterized in that: Step 2 includes the following steps 201 to 203: Step 201: performing discrete wavelet decomposition on the time series wind turbine data to obtain high-frequency coefficients and low-frequency coefficients; After wavelet packet transform, N time series data are decomposed into scaling coefficients s j (k) and wavelet coefficients w j (k); The time series is expressed by inverse wavelet transform and discrete wavelet transform coefficients as follows: Where, represents the reconstructed time series data, j and k represent the kth coefficient of the jth level decomposition, φ j,k (t i ) and ψ j,k (t i ) represents the scaling function and wavelet function of the k-th coefficient of the j-th layer. The scaling function is obtained by scaling the wavelet function. The two sums represent the summation of the scaling subspace and the wavelet subspace. The j-th level scaling factor s j (k) and wavelet coefficients w j (k) is expressed as: Where m = 2k + n, h(·) and h1(·) are scaling and wavelet filter coefficients, respectively, and their values ​​are obtained by setting the wavelet integral to 0 and the square integral to 1; the original data coordinate value is c j Initialize and scale the filter coefficients and wavelet filter coefficient equations to obtain the required decomposition coefficients through iteration; Step 202, performing threshold quantization processing on the decomposition coefficients through Bayesian hypothesis test; Assume d jk represents the kth decomposition coefficient of the jth layer. Assuming that the decomposition coefficient is contaminated by Gaussian white noise, then: In the above formula, represents the noise-free coefficient, ε jk ~N(0,1) represents independent and identically distributed error variables with mean 0; σ j represents the noise variance of the j-layer decomposition; for and σ j 2 The conditional distribution of the decomposition coefficient is expressed as: Assumptions The unbiased prior distribution of is as follows: where γ jk is an independent Bernoulli distribution π j A binary random variable, that is: P(γ jk =1)=1-P(γ jk =0)=π j (8) Formula (8) Determination coefficient is zero (γ jk =0) or non-zero (γ jk =1); variance σ j 2 Indicates that in the jth decomposition layer the scale of change; According to Bayesian theory, using non-informative prior distribution, the noise reduction coefficients are obtained The posterior distribution of is as follows: In the variance When, through a normal distribution and the point mass δ(0) at zero, we get The marginal conditional posterior distribution of : in is the posterior distribution of formula (7), which is described by the following probability where η jk Yes jk =0 and γ jk = 1, also known as the Bayesian coefficient, is expressed as follows: The decomposition coefficients are thresholded according to the following hypothesis test: H0:d jk =0vs·H1:d jk ≠0 (13) Right jk Apply the indicator function I(·) to obtain the decomposed coefficients When I(n jk <1)=1, if η in formula (12) jk <1, the null hypothesis H0 is rejected and the corresponding coefficient is retained; Step 203, performing wavelet reconstruction on the signal to obtain the fan data after noise reduction; The denoising decomposition coefficient obtained by formula (14) The noise-reduced signal can be obtained by using formula (2) Reconstruction processing is performed to obtain the fan data after noise reduction.

3. The wind turbine power prediction method according to claim 1, characterized in that: In step 3, the noise-reduced fan data is standardized using the following formula: Where X is the normalized data, i.e., the input of the subsequent MTGP model, X0 is the data after denoising in the previous step, μ is the mean of the sample, and S is the standard deviation of the sample.

4. The wind turbine power prediction method according to claim 1, wherein: The step 402, establishing the MTGP model, includes the following process: The MTGP model uses Gaussian prior to model the time series. The Gaussian prior consists of the mean function and the covariance function, which is expressed as: y=f(x)~GP(m(x),k(x,x′)) (15) Where x and y represent the input and output data in the training set, respectively; m(x) is set to 0, and k(x, x′) is the key factor of the model, usually using the square exponential kernel function, which can be expressed as: Where τ is the Euclidean distance between input variables, θ f and θ l is the hyperparameter that needs to be optimized; The hyperparameter optimization method of θ is used to minimize the negative log marginal likelihood function (NLML). f and θ l For optimization, the NLML function and the optimization function can be expressed as: The NLML function is a convex function, and the gradient descent algorithm is selected for effective optimization. After model training, the test set data x * The predicted value can be expressed as: The covariance kernel function of the MTGP model can be expressed as: In the formula, the matrix K c Reflects the relationship between tasks, in the form of a task label set Decide, m=2, the target task is wind turbine power, and the auxiliary task is wind speed; K t is the covariance matrix determined by the input data in the training set, represents the Kronecker product, θ c ,θ t Represents K c and K t The set of all hyperparameters in; the covariance matrix K t The elements of are kernel functions of the input data, which have a fixed form and all hyperparameters are shared by all tasks; When using the MTGP model for time series forecasting, the target task is defined as wind turbine power, and the auxiliary task is defined as wind speed. The MTGP model can be constructed as follows: Where y r ∈R r×1 and y p ∈R p×1 Represents outputs of two different dimensions, y r is the target task output, y p For auxiliary task output, focus on the output result of the target task wind turbine power, x r and x p Represents the input datasets of two different tasks, x r is a small amount of historical power data of wind turbines, x p is sufficient wind speed data; θ rr ,θ rp ,θ pr and θ pp Indicates the hyperparameters in the MTGP model that need to be optimized using the NLML method; The prediction results of the MTGP model can also be expressed as:

5. The wind turbine power prediction method according to claim 4, characterized in that: The step 5 further includes: evaluating the MTGP model using root mean square error and determination coefficient evaluation indicators.