An ultra-short-term power prediction method for offshore wind power based on multi-head self-attention integrated boosting mode

Through the multi-head self-attention and Adaboost integrated learning method, an ultra-short-term power prediction model for offshore wind power is constructed, which solves the problem of uncertainty in offshore wind power prediction and achieves higher accuracy and stable ultra-short-term power prediction.

CN116050621BActive Publication Date: 2025-08-26KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310049281.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-01
Publication Date
2025-08-26
Estimated Expiration
2043-02-01

AI Technical Summary

Technical Problem

The uncertainty of offshore wind power and complex climate conditions make it difficult for the existing technology to achieve accurate ultra-short-term power predictions, affecting the stability of the power grid.

Method used

The ultra-short-term power prediction method of offshore wind power is adopted, combined with the multi-scale time block autocoding mechanism and Adaboost integrated learning method, and the wind power power prediction model is constructed, including data collection, preprocessing, multi-scale time block autocoding and integrated learning improvement.

Benefits of technology

It improves the accuracy and stability of ultra-short-term power prediction of offshore wind power, enhances the generalization and portability of the model, and is better than traditional prediction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116050621B_ABST
    Figure CN116050621B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of wind power prediction, and provides a multi-head self-attention offshore wind power ultra-short-term power prediction method with an integrated boosting mode, which further improves the accuracy of offshore wind power ultra-short-term power prediction; the adopted technical solution includes the following steps: S1, data collection and preprocessing, S2, introducing a multi-scale time block self-encoding mechanism as an embedding layer to construct a wind power prediction model, S3, using the Adaboost integrated learning method to improve the prediction model, and S4, case analysis and verification; the prediction model constructed by the present invention has excellent generalization and portability, and compared with traditional prediction models, has further improved the accuracy of offshore wind power ultra-short-term power prediction; integrated learning can further improve the model prediction performance, and is comprehensively superior to the prediction efficiency of traditional wrapped sparse constraint algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention discloses an ultra-short-time power prediction method for offshore wind power with multiple self-attention in an integrated lifting mode, belonging to the technical field of wind power prediction. Background Art

[0002] Compared to onshore wind power, offshore wind power presents greater uncertainty and instability, which can impact the stable operation of power systems. Furthermore, complex offshore climatic conditions limit the accuracy of offshore wind power forecasts. The strong coupling between sea steam and waves necessitates an urgent need for accurate ultra-short-term offshore wind power forecasts to ensure stable grid operation.

[0003] Generally speaking, wind power forecasting can be categorized into two types: physical models and statistical models. Physical methods primarily rely on predicting wind speed based on environmental conditions such as air pressure and temperature surrounding the wind farm, combined with numerical weather prediction (NWP) models. This approach, however, is costly and requires numerous assumptions, hindering rapid scalability. Statistical models, on the other hand, are data-driven. Their core focus is on mapping independent and dependent variables within collected data, further generalizing to the prediction of unknown data. Based on model construction and solution methods, statistical models can be broadly divided into two categories: traditional statistical analysis and artificial intelligence. The former includes methods such as multivariate linear regression and partial least squares, while the latter includes methods such as support vector machine regression, decision trees, and random forests. While these methods have been applied in wind power forecasting scenarios, further advancements are needed because these models fail to fully consider the temporal and spatial variations in wind power.

[0004] Thanks to the development of sensor, communication and storage technologies, we can now obtain massive amounts of high-temporal-resolution wind power monitoring data. This makes it possible to build deep learning time series models, which is expected to model the time dependence of wind power and further improve the accuracy of ultra-short-term offshore wind power forecasts. Summary of the Invention

[0005] The present invention overcomes the deficiencies of the prior art and aims to solve the following technical problems: providing an integrated boosting mode multi-head self-attention offshore wind power ultra-short-term power prediction method to further improve the accuracy of offshore wind power ultra-short-term power prediction.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: a multi-head self-attention offshore wind power ultra-short-term power prediction method with an integrated boosting mode, comprising the following steps:

[0007] S1, data collection and preprocessing;

[0008] S2, introduces a multi-scale time block autoencoding mechanism as an embedding layer to construct a wind power prediction model;

[0009] S3. Use Adaboosts ensemble learning method to improve the prediction model;

[0010] S4. Case analysis and verification.

[0011] Beneficial effects:

[0012] The prediction model constructed by the present invention has excellent generalization and portability. Compared with traditional prediction models, it has further improved the accuracy of ultra-short-term power prediction of offshore wind power. Integrated learning can further improve the prediction performance of the model, and is generally superior to the prediction efficiency of traditional parcel-type sparse constraint algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The present invention will be further described in detail below with reference to the accompanying drawings;

[0014] Figure 1 It is a schematic diagram of the prediction process of the present invention;

[0015] Figure 2 This is a graph showing the prediction performance of the random sampling point model in the example analysis of the present invention;

[0016] Figure 3 This is a cross-validation score graph of different base learners in the example analysis of the present invention;

[0017] Figure 4 This is the cross-validation score graph of the ensemble learning algorithm in the example analysis of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0019] The present invention provides an ultra-short-term power prediction method for offshore wind power with multiple self-attention in an integrated boosting mode, comprising the following steps:

[0020] S1, data collection and preprocessing;

[0021] S2, introduces a multi-scale time block autoencoding mechanism as an embedding layer to construct a wind power prediction model;

[0022] S3. Use Adaboost ensemble learning method to improve the prediction model;

[0023] S4. Case analysis and verification.

[0024] The content of the data collection and preprocessing in step S1 is: the collected wind farm power data is time-sequenced at a time resolution of 5 seconds, invalid and missing data are removed, and sparse autoencoding semi-supervised learning is performed to optimize the latent space dimension.

[0025] The steps of introducing the multi-scale time block autoencoding mechanism as an embedding layer to construct a wind power prediction model in step S2 are as follows:

[0026] S21, Sparse Temporal Block Autoencoder Network

[0027] Sparse Time Block Autoencoder Network Sparse Time Block Autoencoder Network is an adjustment of the traditional sparse autoencoder network (Sparse Autoencoder, SAE). The core difference is that the input layer needs to use a flattening layer to flatten the unequal length sequence vectors with neighborhood blocks as the basic unit, and fill the number of sequence units to 2048 by filling zeros. The encoder input is X, and X∈R (2048×N) Next, the network adopts the SAE encoder (denoted as A0)-decoder (denoted as A2) structure, and the decoder output is recorded as have:

[0028]

[0029] In the formula, both the decoder and the encoder are fully connected layers, and the activation function is sigmoid, so:

[0030]

[0031] Where A1 is the output of the encoder and the input of the decoder. It is located between the encoder and the decoder and is implicitly expressed in the model architecture. Therefore, it is also called the latent space. Under the constraints proposed by this invention, the latent space is a low-rank approximation of the original data, that is, a low-rank feature expression; W1 and W2 represent the weights of the encoder and decoder respectively; b1 and b2 represent the biases of the encoder and decoder respectively.

[0032] sigmoid(z)=(1+e -z ) -1 (3)

[0033] Where z represents a universal vector in the real field, -z represents the negative value of z; z∈R n represents any vector in the real field;

[0034] Formula (3) serves as formula (2) and is used to explain the operation of sigmoid in formula 2;

[0035] The training goal of the IAE network is to minimize the reconstruction loss while introducing sparse constraints, so the IAE loss function is:

[0036]

[0037] Where J represents the total loss function of the IAE network; M represents the total number of valid time series samples participating in the IAE network training; x (i) represents the true wind power value of the i-th sample, represents the wind power value output by the decoder of the ith sample; λ is the given regularized sparseness, β is the given sparse constraint sparseness; F represents the matrix Frobenius norm; D represents the latent space dimension; p, is an intermediate variable, and its specific calculation method is given in formula (5);

[0038]

[0039] Where, Represents the output value of the jth neuron of the i-th sample in the latent space A1; Represents the input value of the jth neuron of the i-th sample of encoder A0;

[0040] With formula (4) as the optimization target, after the sparse time block autoencoder network is trained using the stochastic gradient descent method, the resulting latent space A1 is the result of the equal-length autoencoder mapping;

[0041] S22, Multi-head Attention Architecture

[0042] The classic multi-head attention architecture consists of a decoder and an encoder, both of which contain multi-head self-attention networks to achieve sequence-to-sequence modeling. The prediction model of the present invention is a time series multi-head attention architecture, and its core task is ultra-short-term wind power prediction. Therefore, the present invention makes an overall adjustment to the decoder in the classic multi-head attention architecture, replacing it with a single linear layer network. The encoding result of the encoder at the prediction head position is converted, and the wind power prediction value is the output.

[0043] In addition, the encoder structure in the architecture has been adjusted. The encoder is composed of L stacked basic units, each of which contains a multi-head self-attention network (MSA), a multilayer perceptron (MLP), and layer normalization (LN). The MSA is composed of multiple self-attention networks (SA):

[0044] MSA(z)=[SA1(z);SA2(z);...;SA k (z)]U msa (6)

[0045] Among them, k is the total number of long positions; U msa is the multi-head mapping parameter, which is the variable to be learned in the model;

[0046] Each SA returns the relevance of each element of the input sequence relative to other elements in the sequence in the form of a value through the query (q), key (k), and value (v) mechanism. The calculation formula is as follows:

[0047]

[0048] [q, k, v] = zU qkv (8)

[0049] Among them, U qkv is the mapping parameter, which is the variable to be learned in the model; q represents the query vector, T represents the matrix transpose operation, and D h Represents the length of the sequence vector and v represents the value vector.

[0050] Ensemble learning (EL) includes boosting and bagging ensembles. Both enhance the model from the aspects of improving the learner and improving the data validity. Boosting ensemble learning is to enhance the potential hypothesis theory based on weak learners that can be stacked into strong learners by voting on weak learners, thereby improving the accuracy and performance of the overall model. Among them, Adaboost (AB) is one of the classic methods of boosting ensemble learning, and it performs well in generalization ability. The present invention will use Adaboosts ensemble learning method to improve the prediction model.

[0051] The step S3 uses the Adaboost ensemble learning method to improve the prediction model, which is to improve the weak learner to a strong learner. Before the ensemble improvement, the number of ensembles T needs to be given. The specific steps are as follows:

[0052] S31, training a benchmark model as a base learner in the original training sample set with normal initialization parameters;

[0053] S32. Calculate the mean square error (MSE) of each sample in the training set for the current learner, and sort the samples from largest to smallest according to the MSE.

[0054] S33, select half (round down) of the samples with the larger MSE, and retrain a base learner on the samples;

[0055] S34, performing a weighted average of the new base learner and the original base learner according to the additive model rule, where the weighting parameter is the percentage weight of the MSE value of the training sample corresponding to the base learner;

[0056] S35. Repeat S32-S34 until the number of iterations reaches the upper limit of ensemble learning times.

[0057] The multi-head self-attention offshore wind power ultra-short-term power prediction model framework of the integrated promotion mode provided by the present invention includes three parts as a whole to complete the model prediction of ultra-short-term wind power, wherein the steps completed by each part are as follows:

[0058] 1) Time coefficient constraint

[0059] The model serializes the collected wind farm power data at a 5-second time resolution, removes invalid and missing data, and performs sparse autoencoding semi-supervised learning by optimizing the latent space dimension.

[0060] 2) Time Block Embedding

[0061] The model uses the obtained latent space tensor as the antecedent and the wind power fully connected layer output as the posterior to perform spatiotemporal position encoding, and then flattens the output and embeds it into the encoder of the multi-head attention architecture.

[0062] 3) Multi-head attention architecture

[0063] Taking multi-head self-attention as the basic unit, the embedding layer vector space is encoded and decoded to finally obtain the prediction power.

[0064] The experimental process includes the data preparation and collection stage and the experimental stage.

[0065] The experimental process is as follows Figure 1 shown.

[0066] The preparation collection stage is the data collection and preprocessing of step S1 of the present invention.

[0067] The specific contents of data preparation and collection include:

[0068] S11. Data recording of offshore wind turbines

[0069] The offshore wind farm units are equipped with standardized power monitoring sensors, which perform parallel data acquisition at a time resolution of 1 minute as specified in the present invention. The digital signals are connected to the database through the Internet of Things, thereby completing the real-time recording and storage of the wind power of the specific units.

[0070] S12. Data preprocessing in the database

[0071] The input data of the preprocessing stage is the original data collected and recorded in the database in step A1, and the output is the time series data samples after data cleaning and serialization. The specific steps include:

[0072] (1) Data cleaning

[0073] During wind power acquisition and digital signal transmission and storage, signal interruptions or loss may occur periodically. This can manifest as missing values ​​or abnormal codes within the database. The data cleaning step removes these missing values ​​and abnormal codes, retaining valid signals.

[0074] (2) Sample serialization

[0075] This invention aims to model the time-dependent pattern of offshore wind power, specifically ultra-short-term forecasting. Sample serialization involves reorganizing independent discrete sample points into a sequence, including five historical monitoring values ​​as model independent variables and the power value of the next node as the prediction target.

[0076] The experimental phases are as follows:

[0077] With the preparatory collection phase complete, the present invention continues to investigate and analyze the changes in model accuracy under the sparse autoencoder scale constraint. Finally, the present invention enhances the model using the Adaboost ensemble algorithm and compares it with a related basis learner and its corresponding ensemble learning to demonstrate the model's stability and reliability, completing the model's ultra-short-term offshore wind power forecast.

[0078] The following is a detailed description of the case analysis and verification based on specific data.

[0079] The content of the example analysis and verification in step S4 includes:

[0080] S41. Determine the evaluation index of the prediction results

[0081] A cluster of offshore wind farms was selected for case analysis. The data samples were derived from the historical power data set of the wind farm. The capacity of a single unit was 1.5 MW, and the time resolution used was 1 minute. After the data was sorted, it was stored in the database according to the wind turbine number.

[0082] The unification of the dimensions and ranges of each neuron node in the deep learning model helps to avoid network weight bias. Based on this, formula (1) is used to normalize the data to the interval (0, 1), that is:

[0083]

[0084] Where: x is the actual optimal original offshore wind power input power value in the sample; x tis the optimal offshore wind power value after normalization in the sample; x min is the minimum original offshore wind power in the sample; x max is the maximum value of raw offshore wind power in the sample;

[0085] The offshore wind power ultra-short-term power prediction results are evaluated using mean absolute error (MAE), mean square error (MSE), and mean absolute percentage error (MAPE). The three expressions are as follows:

[0086]

[0087] And select the coefficient of determination R 2 Evaluate the prediction quality of the prediction model, R 2 The value range of is (0,1). The closer the result is to 1, the higher the accuracy of the model fitting data and the better the model quality. The formula is as follows:

[0088]

[0089] Among them, Z i is the actual value of offshore wind power sample; is the predicted value of offshore wind power; is the average value of the actual value of the offshore wind power samples; m is the total number of offshore wind power samples;

[0090] S42. Benchmark model construction and experimental optimization

[0091] S421. Split the input sequence data into time blocks to match the input form of the multi-head self-attention module. Before splitting the time blocks, the time information and its corresponding variable attribute information need to be cleaned and sorted. The focus is on organizing the original data into sequential sequence pairs at equal time step intervals:

[0092] S={(t1,s1),(t2,s2),...,(t N ,s N )}

[0093] Where N is the total number of valid samples, t k ,1≤k≤N represents t k The attribute information of the moment record value is s k ;

[0094] S422. Specify multi-scale time block size array

[0095] The experimental goal is to complete the ultra-short-term wind power forecast, so the default specified time block size group is A = {1, 2, 3, 4, 5, 6, 7, 8, 9}. Then, the initial sequence S is reorganized according to each element in A as the time neighborhood size to obtain the sequence group:

[0096] {B1,B2,...,B9}

[0097] Among them, B k (1≤k≤9) represents A k Serialize X for the neighborhood size; A k represents the k-th element of A, whose default value is exactly k;

[0098] After time block segmentation, the initial sequence X has been reorganized into 9 sequence groups of unequal length. In order to connect the reorganized sequence groups with the embedding layer, sparse autoencoding is performed on {B1, B2, ..., B9} and mapped to the latent space A1∈R d , where d is the latent space dimension;

[0099] According to the model network structure configuration table shown in Table 1, the model group of the present invention under the multi-scale latent space sparse constraint condition is constructed;

[0100] Table 1 Network structure configuration table of the model of the present invention

[0101] Table 1 models structure configuration table

[0102]

[0103] As the latent space dimension increases, the model's constraints on time block sparsity weaken, and the number of neuron nodes also tends to increase. There is a correlation between the model's storage space overhead, training time, and the dimension of the constraint space. The model performs a masking operation on future power values ​​in the prediction sequence to prevent the introduction of prediction information into autoregression, thereby enhancing model robustness. Under this configuration, the model is trained based on sample data using the Adam optimizer.

[0104] In order to comprehensively analyze the prediction ability of the benchmark model, the original data from the offshore wind farm cluster sample within the range of 9:00-17:00 on a certain day were randomly selected for experimental analysis. The experiment was conducted with a prediction step of 5 minutes. The predicted value of wind power and the actual value are shown in the figure below. Figure 2 As shown, at the sampling point position in the stable period, the prediction model of the present invention has excellent prediction accuracy and ability, which shows that the prediction model has excellent prediction performance in the stable power change period;

[0105] S423, Model Performance Verification

[0106] According to the probabilistically approximately correct learning framework (PAC), weak learnability can be improved to strong learnability; in order to further enhance the model of the present invention, finally, an integrated learning experiment will be conducted; in order to verify the comprehensive ability of the model of the present invention, a group of base learners LASSO regression, LSTM, GRU (K-NearestNeighbor, KNN), classification and regression tree (CART), support vector machine (SVM) are synchronously set to cross-validate with the prediction model of the present invention. Cross-validation can effectively evaluate the pros and cons of the prediction performance of the model of the present invention in the sample data of the offshore wind farm cluster, and reduce overfitting to a certain extent. The present invention uses k-fold cross-validation, the idea of ​​which is to randomly divide the original sample data of wind power into k parts without repeated sampling, and select one of the wind power sample data as the test set each time, and the remaining k-1 sample data are used as the training set for model training until all data complete the experimental test. During the experiment, the sample is set to k=5. Cross-validation can reflect the generalization ability of the model. The evaluation criterion of the experiment setting of the present invention is negative mean square error. The results are as follows Figure 3 shown.

[0107] The validation score trends of the six models are as follows Figure 3 , comprehensive analysis shows that the base learner with the best and most stable prediction performance is the model of the present invention, and the prediction performance of the LASSO regression model is second only to the prediction model of the present invention. LASSO regression is similar to our time block sparse autoencoder, and both use the l1 norm to regularize the learning variable dimension to achieve the effect of feature dimensionality reduction. By comparing the prediction model of the present invention, the LASSO regression model and other models horizontally, it can be found that the performance of the model after sparse constraint is more stable than that of the general model in the ultra-short-term power prediction scenario of wind power. The sparse constraint of the LASSO regression model is a non-autonomous training process, and it is impossible to manually intervene in the latent space dimension and probability distribution. The sparse time block autoencoder in the prediction model of the present invention uses the KL divergence to perform low-rank approximation constraints on the latent space, so that the final performance of the prediction model of the present invention is better than the LASSO regression model.

[0108] AdaBoost is an ensemble learning algorithm that can improve weak learners to strong learners. This group of experiments performed AdaBoost on the basis of the previous stage of base learners and performed cross-validation on them. The results are as follows Figure 4 shown.

[0109] like Figure 4It can be seen that among the six horizontal comparison models, the most outstanding and stable models are still the AB-present invention prediction model and the AB-LASSO model, and the integrated present invention model is still better than the AB-LASSO model after integration in terms of stability and prediction performance. In order to further verify the performance of the prediction model of the present invention in ultra-short-term offshore wind power prediction after ensemble learning, and at the same time prove the generalization ability of the prediction model constructed by the present invention, the present invention finally randomly selects the wind power of four hours on a certain day from the sample data of the offshore wind farm cluster, and uses the prediction model of the present invention, the LASSO regression model, the AB-present invention prediction model, and the AB-LASSO model to predict the ultra-short-term offshore wind power. The final prediction results are shown in Table 2.

[0110] Table 2 Power prediction accuracy of ensemble learning and base learners

[0111] Table 5 Power forecasting accuracy table of ensemble learning and basic learning

[0112]

[0113] As can be seen from the table, the new model after the integration of the present invention is slightly better than the benchmark model of the present invention. And the integrated AB-LASSO model is also better than the LASSO model. In the comprehensive numerical analysis comparison, the integrated model of the present invention is better than the original model of the present invention. MSE The determination coefficient R 2 The MSE of AB-LASSO model increased by 2.59% compared with LASSO model, and the coefficient of determination R 2 The result shows that in the ultra-short-term power forecast of offshore wind power, the ensemble learning AB algorithm can effectively improve the performance of the selected prediction base learner. Based on this, the present invention can further obtain the integrated improvement model of the present invention. On the other hand, the integrated prediction model improves the MSE by 13.26% compared with the AB-LASSO model, and the determination coefficient R 2 The improvement is 3.12%, which further proves that the sparse time block autoencoding in the prediction model of the present invention has better performance than the traditional wrapped sparse constraint in the wind power time series prediction model.

[0114] The conclusions of the case analysis and verification are as follows:

[0115] The model constructed by this invention has excellent generalization and portability. Compared with traditional prediction models, it has further improved the accuracy of ultra-short-term offshore wind power prediction. Ensemble learning can further improve the model's prediction performance, and our model is generally superior to the traditional parcel-type sparse constraint algorithm in prediction efficiency.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-head self-attention offshore wind power ultra-short-term power prediction method with integrated boosting mode, characterized by The steps include: S1, data collection and preprocessing; S2, introduces a multi-scale time block autoencoding mechanism as an embedding layer to construct a wind power prediction model; The steps of introducing the multi-scale time block autoencoding mechanism as an embedding layer to construct a wind power prediction model in step S2 are as follows: S21, Sparse Temporal Block Autoencoder Network The input layer of the sparse time block autoencoder network needs to use a flattening layer to flatten the unequal length sequence vectors with neighborhood blocks as the basic unit, and fill the number of sequence units to 2048 by filling zeros. The encoder input is X, and X∈R (2048×N) Next, the network uses the SAE encoder, denoted as A0, and the decoder, denoted as A2 structure. At the same time, the decoder output is have: In the formula, both the decoder and the encoder are fully connected layers, and the activation function is sigmoid, so: Where A1 is the output of the encoder and the input of the decoder. It is located between the encoder and the decoder and is implicitly expressed in the model architecture. Therefore, it is also called the latent space. Under the proposed constraints, the latent space is a low-rank approximation of the original data, that is, a low-rank feature expression. W1 and W2 represent the weights of the encoder and decoder respectively; b1 and b2 represent the biases of the encoder and decoder respectively; sigmoid(z)=(1+e -z ) -1 (3) Where z represents a universal vector in the real field, -z represents the negative value of z; z∈R n represents any vector in the real field; Formula (3) serves as formula (2) and is used to explain the operation of sigmoid in formula 2; The training goal of the IAE network is to minimize the reconstruction loss while introducing sparse constraints, so the IAE loss function is: Where J represents the total loss function of the IAE network; M represents the total number of valid time series samples participating in the IAE network training; x (i) represents the true wind power value of the i-th sample, Represents the wind power value output by the i-th sample decoder; λ is the given regularized sparsity, and β is the given sparse constraint sparsity; F represents the Frobenius norm of the matrix; D represents the dimension of the latent space; p, is an intermediate variable, and its specific calculation method is given in formula (5); Where, Represents the output value of the jth neuron of the i-th sample in the latent space A1; Represents the input value of the jth neuron of the i-th sample of encoder A0; With formula (4) as the optimization target, after the sparse time block autoencoder network is trained using the stochastic gradient descent method, the resulting latent space A1 is the result of the equal-length autoencoder mapping; S22, Multi-head Attention Architecture The decoder in the classic multi-head attention architecture is comprehensively adjusted and replaced with a single linear layer network. The encoding result of the encoder at the prediction head position is converted, and the wind power prediction value is the output. In addition, the encoder structure in the architecture has also been adjusted. The encoder is composed of L stacked basic units. Each basic unit contains a multi-head self-attention network, a multi-layer perceptron, and layer normalization. The multi-head self-attention network (MSA) is composed of multiple self-attention networks (SAs): MSA(z)=[SA1(z);SA2(z);...;SA k (with)]U msa (6) Among them, k is the total number of long positions; U msa is the multi-head mapping parameter, which is the variable to be learned in the model; Each SA returns the relevance of each element of the input sequence relative to other elements in the sequence in the form of a value through the query (q), key (k), and value (v) mechanism. The calculation formula is as follows: [q,k,v]=zU qkv (8) Among them, U qkv is the mapping parameter, which is the variable to be learned in the model; q represents the query vector, T represents the matrix transpose operation, and D h Represents the length of the sequence vector, v represents the value vector; S3. Use Adaboost ensemble learning method to improve the prediction model; S4. Case analysis and verification.

2. The method for ultra-short-term offshore wind power prediction based on multi-head self-attention in an integrated boosting mode according to claim 1 is characterized in that: The content of the data collection and preprocessing in step S1 is: the collected wind farm power data is time-sequenced at a time resolution of 5 seconds, invalid and missing data are removed, and sparse autoencoding semi-supervised learning is performed to optimize the latent space dimension.

3. The method for ultra-short-term offshore wind power prediction based on multi-head self-attention in an integrated boosting mode according to claim 1 is characterized in that: The step S3 uses the Adaboost ensemble learning method to improve the prediction model, which is to improve the weak learner to a strong learner. Before the ensemble improvement, the number of ensembles T needs to be given. The specific steps are as follows: S31, training a benchmark model as a base learner in the original training sample set with normal initialization parameters; S32. Calculate the mean square error (MSE) of each sample in the training set for the current learner, and sort the samples from largest to smallest according to the MSE. S33, selecting half of the samples with the largest MSE and rounding down to the nearest integer, and retraining a base learner on the samples; S34, performing a weighted average of the new base learner and the original base learner according to the additive model rule, where the weighting parameter is the percentage weight of the MSE value of the training sample corresponding to the base learner; S35. Repeat S32-S34 until the number of iterations reaches the upper limit of ensemble learning times.

4. The method for ultra-short-term offshore wind power prediction based on multi-head self-attention in an integrated boosting mode according to claim 1 is characterized in that: The content of the example analysis and verification in step S4 includes: S41. Determine the evaluation index of the prediction results Select data samples and use formula (1) to normalize the data to the interval (0,1), that is: Where: x is the actual optimal original offshore wind power input power value in the sample; x t is the optimal offshore wind power value after normalization in the sample; x min is the minimum original offshore wind power in the sample; x max is the maximum value of raw offshore wind power in the sample; The offshore wind power ultra-short-term power prediction results are evaluated using mean absolute error (MAE), mean square error (MSE), and mean absolute percentage error (MAPE). The three expressions are as follows: And select the coefficient of determination R 2 Evaluate the prediction quality of the prediction model, R 2 The value range of is (0,1). The closer the result is to 1, the higher the accuracy of the model fitting data and the better the model quality. The formula is as follows: Among them, Z i is the actual value of offshore wind power sample; is the predicted value of offshore wind power; is the average value of the actual value of the offshore wind power samples; m is the total number of offshore wind power samples; S42. Benchmark model construction and experimental optimization S421. Split the input sequence data into time blocks to match the input form of the multi-head self-attention module. Before splitting the time blocks, the time information and its corresponding variable attribute information need to be cleaned and sorted. The focus is on organizing the original data into sequential sequence pairs at equal time step intervals: S={(t1,s1),(t2,s2),...,(t N ,s N )} Where N is the total number of valid samples, t k ,1≤k≤N represents t k The attribute information of the moment record value is s k ; S422. Specify multi-scale time block size array The experimental goal is to complete the ultra-short-term wind power forecast, so the default specified time block size group is A = {1, 2, 3, 4, 5, 6, 7, 8, 9}. Then, the initial sequence S is reorganized according to each element in A as the time neighborhood size to obtain the sequence group: {B1,B2,...,B9} Among them, B k (1≤k≤9) represents A k Serialize X for the neighborhood size; A k represents the k-th element of A, whose default value is exactly k; After time block segmentation, the initial sequence X has been reorganized into 9 sequence groups of unequal length. In order to connect the reorganized sequence groups with the embedding layer, sparse autoencoding is performed on {B1, B2, ..., B9} and mapped to the latent space A1∈R d , where d is the latent space dimension; Build a model group under multi-scale latent space sparse constraints according to the model network structure configuration table; As the latent space dimension increases, the model's constraints on time block sparsity weaken, and the number of neuron nodes also tends to increase. There is a correlation between the model's storage space overhead, training time, and the dimension of the constraint space. The model performs a masking operation on future power values ​​in the prediction sequence to prevent the introduction of prediction information into autoregression, thereby enhancing model robustness. Under this configuration, the model is trained based on sample data using the Adam optimizer. In order to comprehensively analyze the prediction ability of the benchmark model, the original data of the data sample was randomly sampled for experimental analysis. The experiment was conducted with a prediction step of 5 minutes to obtain the predicted value and the actual value of wind power, and the verification results were analyzed and obtained. S423, Model Performance Verification The prediction model is cross-validated with multiple groups of base learner models to evaluate the prediction performance of the prediction model in sample data of offshore wind farm clusters, and the verification results are analyzed and obtained.