Offshore wind plant generation power prediction method based on time sequence convolutional network and Transform

By combining the timing convolution network and the Transformer model, DCT extracts frequency domain features and performs multi-dimensional information complementarity, the problems of limited feature data and insufficient periodic features in the power prediction of offshore wind farm power generation are solved, and more efficient and accurate prediction effects are achieved.

CN120262383APending Publication Date: 2025-07-04HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510387145.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The power generation power prediction of offshore wind farms has problems such as limited feature data acquisition, single feature extraction method, and insufficient consideration of periodic features, resulting in insufficient prediction accuracy and stability.

Method used

The time-sequential convolution network (TCN) and the Transformer model are combined to extract frequency domain features through discrete cosine transformation (DCT), and the long-term dependencies are captured by the TCN model, and deep analysis and dynamic prediction are performed through the Transformer model. The feature information is processed in combination with the encoder embedding layer to achieve multi-dimensional information complementarity.

Benefits of technology

It significantly improves the accuracy and stability of offshore wind power power prediction, can capture the complex laws of wind power power changes more comprehensively, and improves the robustness and efficiency of the prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120262383A_ABST
    Figure CN120262383A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of offshore wind power prediction, and discloses a prediction method fusing a time sequence convolutional network and a Transform model. In order to solve the problems that time sequence frequency features are difficult to fully utilize and prediction precision is limited in the prior art, a discrete cosine transform technology is introduced to carry out frequency domain feature extraction on an original time sequence, a periodic rule is deeply mined, and feature weight distribution is optimized; on this basis, a TCN-Transformer composite model is constructed, the TCN model captures a long-time dependency relationship by using expansion causal convolution, and a processed feature information matrix and a feature information matrix passing through an encoder embedding layer are fused to realize multi-dimensional information complementation; and the Transform module is used for carrying out deep analysis and dynamic prediction on the fused features through a self-attention mechanism. According to the method, the offshore wind power prediction precision is remarkably improved, the method can be applied to the field of power grid dispatching and energy optimization, and reliable support is provided for efficient utilization of offshore wind power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to a method for predicting the power generation of an offshore wind farm based on a Temporal Convolutional Network (TCN) and a Transformer. This method first introduces the Discrete Cosine Transform (DCT) technology, aiming to effectively extract the frequency-domain features of time-series data, fully exploit the frequency-domain characteristics to optimize the allocation of feature weights. On this basis, a TCN-Transformer composite model is constructed. The TCN model uses dilated causal convolution to capture long-term dependencies, and fuses the processed feature information matrix with the feature information matrix passing through the encoder embedding layer to achieve multi-dimensional information complementarity. The Transformer module then conducts in-depth analysis and dynamic prediction on the fused features through the self-attention mechanism, realizing deep learning and accurate prediction of the data. Background Art:

[0002] Compared with onshore wind farms, the operation and dynamic management of offshore wind farms present more complex and variable characteristics. The power conversion power has uncertainty and volatility, which undoubtedly poses great challenges to the safe and stable operation of the power grid. Under this background, how to more effectively utilize offshore wind power resources and ensure the stable operation of the power system, the accurate prediction of offshore wind power output is crucial.

[0003] In the field of wind power prediction, the current mainstream methods can be roughly divided into two categories: physical prediction and statistical prediction. Physical prediction methods are based on the understanding of physical principles and mainly rely on numerical weather prediction (NWP) algorithms customized for long-term prediction paradigms. Although this method has certain advantages in predicting wind power trends, it has limitations in dealing with wind power fluctuations. Statistical prediction methods focus on using historical data and mathematical models, mainly relying on time series prediction models. The most common one is the Autoregressive Integrated Moving Average Model (ARMA), which is widely used in wind power prediction but cannot accurately capture the non-linear relationships, complex dynamic characteristics and external influencing factors in the time series, and the prediction accuracy will be limited when dealing with long-term data. Therefore, offshore wind power prediction still faces many challenges.

[0004] In the field of offshore wind power prediction, the first major challenge is that its prediction accuracy is restricted by various complex factors, including but not limited to unpredictable weather conditions and the operating efficiency of power generation equipment. Particularly crucial is that in actual operation, limited by acquisition means and environmental conditions, the characteristic data that can be obtained and used for prediction is often extremely limited. Therefore, how to efficiently and fully utilize this limited characteristic information has become a key technical problem that urgently needs to be solved, and it is of vital significance for improving the accuracy of offshore wind power prediction.

[0005] Secondly, most of the current mainstream wind power prediction methods rely on the simple extraction and utilization of time-domain characteristic information. This single-dimensional characteristic processing method obviously cannot comprehensively and deeply explore the complex laws of wind power changes, resulting in the accuracy of the prediction results being difficult to reach an ideal level.

[0006] Furthermore, although many data-driven wind power prediction models can recognize the importance of the time dependence of time series data during the construction process, they are unable to handle the periodic laws of each characteristic's change over time. This neglect of the periodic change of characteristics not only limits the performance improvement of the model in short-term prediction but also shows obvious deficiencies in medium- and long-term prediction, resulting in a large deviation between the prediction results and the actual situation.

[0007] To sum up, the main problems and defects of the existing technologies in the field of offshore wind power prediction can be summarized as the following three points:

[0008] (1) Limited collection of characteristic data: Due to the complex influence of various factors such as meteorological conditions and the state of power generation equipment, as well as the limitation of acquisition means, the characteristic data that can be obtained in actual applications is extremely limited, seriously restricting the improvement of prediction accuracy.

[0009] (2) Single characteristic extraction method: The current processing of characteristics mainly stays at the time-domain level, failing to fully utilize the richness and diversity of characteristic information, resulting in the prediction model being unable to comprehensively capture the complex laws of wind power changes, thereby affecting the accuracy of the prediction results.

[0010] (3) Insufficient consideration of periodic characteristics: Although the existing data-driven wind power prediction models can capture time series dependence, their modeling ability for periodic characteristics (such as daily / seasonal fluctuations) is insufficient, resulting in the accumulation of prediction errors (see the cases in Tables 2 and 3), which to a certain extent weakens the performance and stability of the model in medium- and long-term prediction. Summary of the Invention:

[0011] Due to the influence of various complex factors on offshore wind farms, such as wind speed, wind direction, humidity, etc., their power generation presents high nonlinearity and time series characteristics. To solve the problem that traditional prediction methods are difficult to accurately capture the complex characteristics of offshore wind farms, resulting in limited prediction accuracy, the present invention discloses a method for predicting the power generation of offshore wind farms based on a temporal convolutional network and Transformer.

[0012] For this reason, the present invention provides the following technical solutions:

[0013] 1. A method for predicting the power generation of an offshore wind farm based on a temporal convolutional network and Transformer, characterized in that the method follows the following steps:

[0014] Step 1: Original data collection and preprocessing

[0015] Collect the original data of the offshore wind farm in the specified area, and perform preprocessing operations on the original data, including effective processing of missing values, normalization and inverse normalization operations of the data, and the application of the sliding window method, so as to generate a sample data set for model training.

[0016] Step 2: Data division

[0017] Divide the historical feature data sequence of offshore wind power obtained in Step 1 into a training set and a test set to generate a two-dimensional data sequence.

[0018] Step 3: Feature analysis

[0019] Use the Pearson correlation coefficient formula to calculate the data series generated in Step 2 to obtain the correlation coefficients of each feature, and determine the feature weights for predicting the target power generation.

[0020] Step 4: DCT processing

[0021] Adopt DCT technology to effectively transform the temporal feature information into frequency domain feature information, accurately identify and effectively remove high-frequency noise components, and obtain discrete cosine transform coefficients.

[0022] Step 5: Construct a TCN model

[0023] Construct a TCN model specifically for processing time series data, and use the frequency features processed by DCT in Step 4 as the input data of the TCN model. The hidden layer of the TCN model includes: a dilated causal convolutional layer, a Dropout layer, and an activation function layer. Utilize the convolutional structure and long-term dependence characteristics of the TCN model to achieve more efficient feature extraction and time series analysis.

[0024] Step 6: Encoder embedding layer processing

[0025] The input data is processed using the encoder embedding layer. The input data and the corresponding time stamps are converted into an embedded format through the encoder embedding layer, providing the necessary feature information support for subsequent prediction tasks.

[0026] Step 7: Feature matrix fusion

[0027] The feature information matrix processed by the TCN model and the feature information matrix passing through the encoder embedding layer are concatenated to obtain a new feature matrix. This step aims to enhance feature representation, make better use of effective information, and improve prediction accuracy.

[0028] Step 8: Construct the Transformer model

[0029] The feature matrix generated in Step 7 is used for prediction by the Transformer model. The encoder in the Transformer model integrates the self-attention mechanism to capture long-range dependencies between features, a feed-forward neural network for non-linear transformation of features, residual connections to alleviate the vanishing gradient problem, and layer normalization operations to accelerate model convergence; the decoder, on this basis, additionally includes a linear layer for generating sequences and a softmax activation function, and through the synergistic effect of the self-attention mechanism, feed-forward neural network, residual connections, and layer normalization, accurate prediction of subsequent time points is achieved.

[0030] Step 9: Model evaluation and optimization

[0031] For the above-designed prediction model, the data in the test set is input into the model to evaluate and verify the prediction model and optimize the parameters.

[0032] According to the method for predicting the power of offshore wind farms based on the temporal convolutional network and the Transformer model described in claim 1, characterized in that in step 1, the original data of the offshore wind farm in the specified area is collected, including meteorological data such as wind speed, wind direction, air density, humidity, turbulence intensity (I), wind shear above the turbine height (S_a), wind shear below the turbine height (S_b), and power generation data. To eliminate the dimensionality differences between different features, the linear function normalization (Min-Max Scaling) method is used to process the feature data. This method maps the original data to the [0,1] interval through a linear transformation, and its normalization formula is as follows:

[0033]

[0034] Where: X′ is the data after normalization processing, X min is the minimum value in the dataset, X max is the maximum value in the dataset, and record these several data to remap back to the original value range for inverse normalization operation.

[0035] A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that in step 2, the sequence of historical offshore feature data obtained in step 1 is divided into a training set and a test set, and for the training set and the test set, a two-dimensional data sequence is generated.

[0036] In this embodiment, it is selected to divide the obtained historical offshore feature data, with the first 80% of the data as the training set and 20% of the data as the test set. The length of the training set data is 0.8*N, and the length of the test set data is 0.2*N, where N is the total length of the data.

[0037] A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that in step 3, a correlation analysis is performed on the collected data. For the collected features, the Pearson correlation coefficient formula is used to calculate the feature correlation coefficient to determine the relevant features for predicting the target power generation. The Pearson correlation coefficient calculation formula is as follows:

[0038]

[0039] where: cov(X,Y) is the covariance of X and Y, σ X and σ Y are the standard deviations of X and Y respectively.

[0040] A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that in step 4, a discrete cosine transform is performed on the time domain information. For example, the most common 1-D DCT definition formula for a discrete time series x(n) of length N (n = 1, 2... N) is as follows:

[0041]

[0042] where: y(k) represents the transformed frequency domain signal, indicating the amplitude of the signal at frequency n, also known as the DCT coefficient, and w(k) is a normalization coefficient used to ensure that the energy of the transform remains consistent.

[0043] When using the DCT coefficient y(k) (k = 1, 2... N) to reconstruct the original time series x(n), this process is called the inverse discrete cosine transform (IDCT), and the definition formula is as follows:

[0044]

[0045]

[0046] Among many DCT transformation methods, DCT-II uses the cosine function as the basis, and this characteristic results in a relatively low correlation among DCT-II coefficients. In the frequency domain, low correlation means that the transformed coefficients are more independent, so that these independent coefficients can be used more effectively to represent the signal. Therefore, this method uses the DCT-II method to convert the time-domain signal into a frequency-domain signal, and the formula is as follows:

[0047]

[0048] The inverse transformation formula of DCT-II is the process of recovering the original time-domain signal from the frequency-domain coefficients after DCT-II transformation. The inverse transformation formula of DCT-II can be expressed as:

[0049]

[0050] A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, wherein in step 5, a temporal convolutional network (TCN) model is constructed, and the specific structure of the model is as follows:

[0051] The hidden layer of the TCN model includes: a dilated causal convolutional layer, a Dropout layer, and an activation function layer. Using the convolutional structure and long-term dependence characteristics of the TCN model, more efficient feature extraction and temporal analysis are realized.

[0052] Input the data set X and X W into the hidden layer for causal dilation respectively. Assume that there is a filter set F = {f1, f2, f3,..., f K}, then the corresponding convolution operation can be expressed as:

[0053]

[0054] In the formula: s represents the current sliding position of the convolution kernel, Y(s) represents the output sequence after convolution, f i is the filter, d is the dilation factor, which is usually set to increase exponentially with the depth of the network, that is, d = 2 m , m is the number of layers of the network, K is the size of the convolution kernel, and x s-i×d is the historical data.

[0055] A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, wherein in step 6, the encoder embedding layer is used to process the data, and the specific operation is as follows:

[0056] The input data is processed using an encoder embedding layer and a decoder embedding layer. The input data and corresponding time stamps are converted into an embedding format through the encoder embedding layer. This embedding vector is designed to accurately capture and augment the semantic content and feature information of the input data, thus comprehensively reflecting its internal characteristics and dynamic changes. Subsequently, the decoder embedding layer is used to process the data and time stamps of the decoder input, providing the necessary feature information support for subsequent prediction tasks.

[0057] A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, wherein the feature matrix fusion in step 7 is as follows:

[0058] The frequency-domain features F TCN ∈R L×32 output by the TCN are concatenated with the time-domain features F Embed ∈R L×D (D is the embedding dimension) along the feature dimension. The concatenated feature matrix contains both time-domain and frequency-domain information and is input into the Transformer model for global dependency modeling.

[0059] Through the parallel connection of feature matrices, it is possible to more comprehensively capture the complex features of the data, which helps the model to use more information for better weight allocation during decision-making, thereby improving the stability and robustness of the model and effectively preventing information loss. At the same time, it can also effectively reduce the feature dimension without losing key information, which is helpful for the training process of the model and reduces the risk of overfitting.

[0060] A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, wherein the construction of the Transformer model in step 8 has the following specific structure:

[0061] The encoder in the Transformer model includes a self-attention mechanism, a feed-forward neural network, a residual connection, and a layer normalization operation. The decoder in the Transformer model includes a self-attention mechanism, a feed-forward neural network, a residual connection, and a layer normalization. The output of the decoder is converted into a probability distribution through a linear layer and a softmax activation function to obtain the prediction result.

[0062] The calculation process of the self-attention mechanism can be expressed as the dot product operation of three vectors: query (Q), key (K), and value (V), as follows:

[0063] Q = XW Q

[0064] K = XW K

[0065] V = XW V

[0066] Where: Q is the query matrix, K is the key matrix, and V is the value matrix.

[0067] The multi-head self-attention mechanism divides the input sequence into multiple heads, each head independently calculates the self-attention, and then concatenates the results. The formula for multi-head self-attention is as follows:

[0068] MultiHead(Q, K, V) = Concat(head1,..., head h )W0

[0069] Where: and W0 are learnable weight matrices.

[0070] The feed-forward neural network usually contains two linear transformations and an activation function. Given the input x, the output of the feed-forward neural network can be expressed as:

[0071] FFN = max(0, xW1 + b1)W2 + b2

[0072] Where: W1 is the weight matrix of the first linear transformation, W2 is the weight matrix of the second linear transformation, b1 is the bias term of the first linear transformation, and b2 is the bias term of the second linear transformation.

[0073] Specifically, for each position i, its output can be expressed as the weighted sum of the values of all positions j, and the weights are calculated by the dot product of query i and key j through the softmax function. The formula is as follows:

[0074]

[0075] Where: Q is the query matrix, K is the key matrix, V is the value matrix, d k is the dimension of the key vector

[0076] A method for predicting the power of offshore wind farms based on a temporal convolutional network and a Transformer model according to claim 1, wherein the specific steps of the model evaluation and optimization in step 9 are as follows:

[0077] Input the test data in the test set into the trained prediction model to obtain the prediction result. According to the comparison between the predicted value and the true value, three evaluation indicators, R 2 , MAE, and RMSE, are used. The first indicator represents the level of the model's prediction ability. The higher the score of the indicator, the better the model's prediction ability. The latter two evaluation indicators represent the error between the predicted value and the true value. The smaller the value of the indicator, the more accurate the prediction result. The calculation formulas for the evaluation indicators of the prediction model are as follows:

[0078]

[0079]

[0080] In the formula: y i represents the true value, i.e., the observed value, represents the model predicted value, y avg represents the average value of the true values, and n represents the number of data points.

[0081] Beneficial effects:

[0082] 1. In exploring the innovative path of offshore wind power prediction technology, the present invention ingeniously applies the DCT technology to achieve an efficient conversion of features from the time domain to the frequency domain. This conversion strategy brings significant advantages: by utilizing the energy concentration characteristic of DCT, it focuses the main energy of the signal on a few key coefficients in the frequency domain, effectively improving the utilization rate of feature information and the accuracy of prediction; at the same time, the feature coefficients after DCT transformation have strong anti-interference ability, which can resist the interference of external factors such as noise to a certain extent, enhancing the stability and reliability of the prediction results. In addition, the good compatibility of DCT with the human perception system makes the transformed features more in line with the requirements of practical applications, helping to simplify the prediction model and improve the prediction efficiency. To sum up, the DCT technology adopted by the present invention provides an efficient, accurate and stable feature conversion method for offshore wind power prediction, with remarkable technological innovation and practical value.

[0083] 2. In the present invention, we propose an innovative solution, that is, to use the TCN model to process the frequency domain features after DCT processing. This method makes full use of the advantages of the TCN model, such as the ability to extract long-term dependence information, efficient parallel computing, and stable training process. By processing the frequency domain features after DCT transformation with the TCN model, we can more accurately capture the temporal information and frequency domain information in the signal, improving the prediction accuracy. At the same time, the scalability and flexibility of the TCN model also enable us to adjust and optimize the model according to the requirements of different application scenarios to adapt to a wider range of wind power prediction tasks.

[0084] 3. The present invention also uses the encoder embedding layer to process the feature information. Through the encoder embedding layer, we can effectively extract and encode the key features in the input data into low-dimensional representations. This process helps to remove redundant information and retain key features, thereby improving the efficiency and accuracy of subsequent processing and significantly enhancing the ability of this method to process complex feature information.

[0085] 4. This method combines the feature information processed by the encoder embedding layer with the data processed by the DCT and TCN models. By doing so, it can make full use of their respective advantages to achieve more comprehensive feature extraction and data representation. This combination method not only improves the accuracy and efficiency of data processing but also provides a richer and more accurate information basis for subsequent data analysis and applications.

[0086] 5. This method inputs the combined feature matrix into the Transformer model for prediction, making full use of the powerful modeling ability of the Transformer model. Through the self-attention mechanism, it captures the long-range dependencies in the input sequence and calculates the global dependencies, thereby understanding the data features more comprehensively. At the same time, the Transformer model has the ability of parallel computing, enabling the model to be more efficient when processing large-scale data, capable of achieving in-depth mining and efficient utilization of data features, and improving the accuracy and reliability of prediction. Description of the Drawings:

[0087] Figure 1 It is a schematic diagram of causal convolution and dilated convolution in the TCN model used in the present invention.

[0088] Figure 2 It is the structure diagram of the Transformer model in the present invention.

[0089] Figure 3 It is the flowchart of the method for predicting the power generation of an offshore wind farm based on a temporal convolutional network and Transformer provided by the method example.

[0090] Figure 4 It is the Pearson correlation coefficient diagram of the seven features (wind speed, wind direction, air density, humidity, turbulence intensity, wind shear above turbine height, wind shear below turbine height) provided by the method example.

[0091] Figure 5 It is the amplitude-frequency diagram of visualizing the discrete cosine transform coefficients after discrete cosine transform in the present invention.

[0092] Figure 6 It is the comparison diagram of the prediction results of the method provided by the embodiment of the present invention and other models using the WT5 dataset at different prediction steps.

[0093] Figure 7 It is the comparison diagram of the prediction results of the method provided by the embodiment of the present invention and other models using the WT6 dataset at different prediction steps. Detailed Embodiments:

[0094] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0095] The method for predicting the power generation of an offshore wind farm based on a temporal convolutional network and a Transformer according to an embodiment of the present invention includes the following steps.

[0096] Step 1: Acquisition and preprocessing of raw data

[0097] Collect the raw data of the offshore wind farm in the specified area, including meteorological data such as wind speed, wind direction, air density, humidity, turbulence intensity (I), wind shear above the turbine height (S_a), wind shear below the turbine height (S_b), and power generation data. Preprocess the raw data, including effective processing of missing values, normalization and inverse normalization operations of the data (using linear function normalization to convert the raw data to the range [0,1]), and the application of the sliding window method to generate a sample data set for model training. The normalization formula is as follows:

[0098]

[0099] In the formula: X′ is the data after normalization processing, X min is the minimum value in the data set, X max is the maximum value in the data set, and record these several data to remap back to the original value range for inverse normalization operation.

[0100] Step 2: Data division

[0101] Divide the offshore historical feature data sequence obtained in Step 1 into a training set and a test set, and generate two-dimensional data sequences for the training set and the test set.

[0102] In this embodiment, it is selected to divide the obtained offshore historical feature data, with the first 80% of the data as the training set and 20% of the data as the test set. The length of the training set data is 0.8*N, and the length of the test set data is 0.2*N, where N is the total length of the data.

[0103] In the generated two-dimensional data, the first dimension of the training set is the length 0.8*N of the training set historical feature data, and the second dimension is the number of input features 8. The first dimension of the test set is the length 0.2*N of the test set historical feature data, and the second dimension is the number of input features 8.

[0104] The predicted power generation value y in this embodiment t+1 and the offshore historical input data {x t ,x t-1...A mapping relationship is formed between x1 and x0, and its formula is:

[0105] y t+1 = f(x t , x t-1 ...x1, x0)

[0106] Where: f is the mapping function.

[0107] Step 3: Feature analysis

[0108] In order to deeply explore the key factors affecting the prediction of offshore power generation, this study first conducted a correlation analysis on the collected data. Using the Pearson correlation coefficient formula for the collected features, the Pearson correlation coefficient was calculated to determine the relevant features for predicting the target power generation. The Pearson correlation coefficient calculation formula is as follows:

[0109]

[0110] Where: cov(X, Y) is the covariance of X and Y, σ X and σ Y are the standard deviations of X and Y respectively.

[0111] The covariance calculation formula of X and Y is as follows:

[0112]

[0113] The mean calculation formulas of X and Y are as follows respectively:

[0114]

[0115] The Pearson correlation coefficients of the two datasets WT5 and WT6 were calculated respectively, and the Pearson correlation coefficient matrix diagram was obtained, as Figure 4 .

[0116] When discussing the key factors affecting offshore power generation, special attention was also paid to the important physical quantity of wind shear, which specifically describes the change of the wind vector in the horizontal or vertical distance. Wind shear is usually expressed by the change rate of wind speed (or wind direction), and its formula is as follows:

[0117]

[0118] Where: S is the wind shear, representing the wind speed change rate, V h represents the wind speed at height h, V0 is the wind speed at the reference height, h is the observation height, and h0 is the reference height.

[0119] Wind shear can be further divided into horizontal wind shear and vertical wind shear. Horizontal wind shear refers to the difference in wind speed and / or direction between two points at the same altitude. The calculation formula for horizontal wind shear is as follows:

[0120] W = (U2 - U1) 2 + (V2 - V1) 2

[0121] In the formula: U2 and U1 respectively represent the values of the eastward wind speed components at the lower and higher levels (or different altitude levels); V2 and V1 respectively represent the values of the northward wind speed components at the lower and higher levels.

[0122] Step 4: Discrete cosine transform processing

[0123] Using the discrete cosine transform (DCT) technology, the time series feature information is effectively transformed into frequency domain feature information. Through DCT processing, high-frequency noise components can be accurately identified and effectively removed, making the high-frequency noise part of the signal concentrated in the high-frequency range, while the low-frequency coefficients carry the main energy information, thereby significantly improving the signal-to-noise ratio of the signal and providing clearer and more accurate input data for subsequent feature extraction and time series analysis. The most common 1-D DCT definition formula for a discrete time series x(n) of length N (n = 1, 2... N) is as follows:

[0124]

[0125] In the formula: y(k) represents the transformed frequency domain signal, indicating the amplitude of the signal at frequency n, also known as the DCT coefficient, and w(k) is a normalization coefficient used to ensure that the energy of the transform remains consistent.

[0126] When using the DCT coefficients y(k) (k = 1, 2... N) to reconstruct the original time series x(n), this process is called the inverse discrete cosine transform (IDCT), and the definition formula is as follows:

[0127]

[0128] With the help of an extended model based on the discrete cosine transform (DCT), based on limited time observation data, the change trend at subsequent time points is predicted in advance. Although N0 is less than N, we can still try to directly use the original definition formula to estimate the DCT coefficients for prediction modeling. The formula for estimating the DCT coefficients for prediction modeling using the original definition is as follows:

[0129]

[0130] In this way, assuming that we use M discrete cosine transform (DCT) coefficients for prediction modeling, its change is predicted in advance through the inverse discrete cosine transform (IDCT), and the formula is as follows:

[0131]

[0132] Among many DCT transformation methods, DCT-II uses the cosine function as the basis, and this characteristic makes the correlation between DCT-II coefficients relatively low. In the frequency domain, low correlation means that the transformed coefficients are more independent, so that these independent coefficients can be used more effectively to represent the signal. Therefore, this method adopts the DCT-II method to convert the time-domain signal into a frequency-domain signal, and the formula is as follows:

[0133]

[0134] According to the DCT coefficients calculated by DCT-II, the amplitude-frequency diagram is used to visualize the DCT coefficients, as Figure 5 .

[0135] After performing the DCT-II transformation on the time-domain signal with length N, the first 10% of the low-frequency coefficients (i.e., the coefficients from k = 1 to k = 0.1N) are retained, and the remaining high-frequency coefficients are set to zero to filter out noise and retain the dominant frequency-domain mode. The number of retained coefficients is calculated as follows:

[0136] M = [0.1N]

[0137] The inverse transformation formula of DCT-II is the process for recovering the original time-domain signal from the frequency-domain coefficients after the DCT-II transformation. The inverse transformation formula of DCT-II can be expressed as:

[0138]

[0139] The retained low-frequency coefficients are reconstructed into a denoised time-domain signal through the inverse transformation of DCT-II, generating the frequency-domain feature matrix F DCT ∈R L×M , where L is the time step and M is the retained frequency-domain dimension.

[0140] Take F DCT ∈R L×M as the input of the TCN model and process it in parallel with the original time-domain features (as Figure 3 shown).

[0141] Step 5: Construct the TCN model

[0142] Construct a TCN model specifically for processing time-series data, and input the frequency-domain feature matrix F DCT generated in step 4 into the TCN model. The specific structure is as follows:

[0143] Input layer: Receive the frequency-domain feature matrix F DCT ∈R L×M, where L is the time step and M is the dimension of the retained DCT coefficients.

[0144] Dilated causal convolution layer:

[0145] The first layer: dilation factor d = 1, kernel size K = 3, number of output channels 32, followed by a ReLU activation function and Dropout (p = 0.1).

[0146] The second layer: dilation factor d = 2, kernel size K = 3, number of output channels 32, followed by a ReLU activation function and Dropout (p = 0.1).

[0147] Output of the TCN model: After two layers of convolution, the output is the time-frequency fusion feature F TCN ∈R L×32 .

[0148] A typical TCN architecture consists of three core components: causal convolution, dilated convolution, and residual connection. The convolution in TCN is causal. It adopts the structure of one-dimensional full convolution (1-D FCN) to ensure that the input and output of each hidden layer are of equal length and have the same time step in the time dimension. This feature is significantly different from traditional convolutional neural networks. It is a strict time-constrained model, and the output at any moment strictly depends on the current and historical input information and has nothing to do with future information. The causal convolution combines dilated convolution as Figure 2 shown, by increasing the kernel size and dilation coefficient to broaden the receptive field of the data, forming a longer convolutional "memory", so that the model can still maintain efficient computational performance when parallelly inputting multi-dimensional data.

[0149] Input the dataset X and X W respectively into the hidden layers of their respective models for convolution operations through causal convolution and dilated convolution. Assuming there is a filter set F = {f1, f2, f3,..., f K}, then the corresponding convolution operation can be expressed as:

[0150]

[0151] In the formula: Y(x s ) represents the output sequence after convolution, f i is the filter, d is the dilation factor, usually set to increase exponentially with the depth of the network, that is, d = 2 m , m is the number of layers of the network, K is the kernel size, and s-(K-i)d is the historical data.

[0152] In this embodiment, the dilation factor of the first layer is 1 and that of the second layer is 2.

[0153] Meanwhile, the hidden layer of the TCN model also includes a Dropout layer and an activation function layer. The Dropout layer in the TCN model is mainly used to prevent overfitting. By randomly discarding the outputs of some neurons, the model will not overly rely on certain specific neurons during training, thus improving the generalization ability of the model. During the training phase, the Dropout operation can be expressed by the following formula:

[0154]

[0155] In the formula: x is the input vector, m is the random mask (0 means discard, 1 means keep), and p is the dropout rate.

[0156] The activation function layer in the TCN model is used to introduce non-linear factors, enabling the model to learn complex mapping relationships. The definition of the ReLU function is very simple, and its mathematical expression is:

[0157] RELU(x) = MAX(0, x)

[0158] That is, when the input x is greater than 0, the output is x; when the input x is less than or equal to 0, the output is 0. This piecewise linear characteristic enables the ReLU function to introduce non-linear factors in the neural network while maintaining the efficiency of calculation.

[0159] Step 6: Encoder embedding layer processing

[0160] The encoder embedding layer and the decoder embedding layer are used to process the input data. The input data and the corresponding time stamps are converted into an embedding format through the encoder embedding layer. This embedding vector is designed to accurately capture and expand the semantic content and feature information of the input data, thus comprehensively reflecting its internal characteristics and dynamic changes. Then, the decoder embedding layer is used to process the data and time stamps of the decoder input, providing the necessary feature information support for subsequent prediction tasks.

[0161] Step 7: Feature matrix fusion

[0162] The frequency-domain features F TCN ∈R L×32 output by the TCN are concatenated with the time-domain features F Embed ∈R L×D (D is the embedding dimension) output by the encoder embedding layer along the feature dimension. The formula is as follows:

[0163]

[0164] The concatenated feature matrix F fusion contains both time-domain and frequency-domain information and is input into the Transformer model for global dependency modeling.

[0165] By parallelizing the feature matrices, it is possible to capture the complex features of the data more comprehensively, which helps the model to utilize more information for better weight allocation during decision-making, thereby enhancing the stability and robustness of the model and effectively preventing information loss. At the same time, it can also effectively reduce the feature dimension without losing key information, which is beneficial to the training process of the model and reduces the risk of overfitting.

[0166] Step 8: Construct a Transformer model

[0167] The encoder in the Transformer model includes a self-attention mechanism, a feed-forward neural network, a residual connection, and a layer normalization operation. The decoder in the Transformer model includes a self-attention mechanism, a feed-forward neural network, a residual connection, and a layer normalization. The output of the decoder is converted into a probability distribution through a linear layer and a softmax activation function to obtain the prediction result.

[0168] The classic Transformer architecture design consists of two core components: an encoder (Encoder) and a decoder (Decoder). In the encoder, the primary task is to convert the original input sequence (such as text, image, etc.) into a high-dimensional feature representation. Through the collaborative action of the self-attention mechanism (Self-Attention) and the feed-forward neural network (Feed-Forward Neural Network), it aims to capture the complex relationships between elements in the input sequence, construct a context-aware representation, and thereby accelerate the computational and training efficiency of the model.

[0169] This Transformer prediction model contains an encoder and a decoder.

[0170] The fused data features are input into the Transformer encoder for deeper feature abstraction and prediction. In each layer of the encoder, the fused features first pass through the self-attention mechanism to calculate the attention weights of each position in the sequence to other positions.

[0171] The calculation process of the self-attention mechanism can be represented as the dot product operation of three vectors: query (Query), key (Key), and value (Value), as follows:

[0172] Q = XW Q

[0173] K = XW K

[0174] V = XW V

[0175] In the formula: Q is the query matrix, K is the key matrix, and V is the value matrix.

[0176] The multi-head self-attention mechanism divides the input sequence into multiple heads, and each head independently calculates self-attention, and then concatenates the results. The formula for multi-head self-attention is as follows:

[0177] MultiHead(Q, K, V) = Concat(head1,..., head h )W0

[0178] Where: and W0 are learnable weight matrices.

[0179] Then it is passed into the feed-forward neural network. The feed-forward neural network usually contains two linear transformations and an activation function, and it is used to further process the output of multi-head attention in each layer of the encoder and decoder. Given the input x, the output of the feed-forward neural network can be expressed as:

[0180] FFN = max(0, xW1 + b1)W2 + b2

[0181] Where: W1 is the weight matrix of the first linear transformation, W2 is the weight matrix of the second linear transformation, b1 is the bias term of the first linear transformation, and b2 is the bias term of the second linear transformation.

[0182] In each layer, residual connections are used to add the input and output, and layer normalization is performed to finally obtain the output after being processed by multiple layers of the Transformer model.

[0183] The decoder receives the output of the encoder and the target sequence as inputs. First, the target sequence is converted into a high-dimensional feature representation through the embedding layer.

[0184] Secondly, the self-attention mechanism is used to process the target sequence, cross-attention calculation is performed, and the target features are combined with the encoder output.

[0185] Then, the features are non-linearly transformed through the feed-forward neural network. Residual connections and layer normalization are used to improve the stability of the model.

[0186] Finally, the output of the decoder is converted into a probability distribution through the linear layer and the softmax activation function to obtain the predicted result output. Specifically, for each position i, its output can be expressed as the weighted sum of the values of all positions j, and the weights are calculated by the dot product of query i and key j through the softmax function. The formula is as follows:

[0187]

[0188] Where: Q is the query matrix, K is the key matrix, V is the value matrix, d k is the dimension of the key vector.

[0189] Step 9: Model Evaluation and Optimization

[0190] Input the test data in the test set into the trained prediction model to obtain a prediction result. According to the comparison between the predicted value and the true value. When evaluating the performance of the model, the following three metrics are adopted in this paper: coefficient of determination (R 2 ), mean absolute error (MAE), and root mean square error (RMSE).

[0191] R 2 represents the proportion of the variation in the variable predicted by the model to the variation in the true variable. The value of R2 ranges between 0 and 1. R 2 is closer to 1, indicating that the model has a better fitting effect on the data; R 2 is closer to 0, indicating that the model has a worse prediction ability.

[0192] R 2 The calculation formula of is as follows:

[0193]

[0194] MAE is the average of the absolute differences between all predicted values and true values. The smaller the MAE, the closer the model's predicted value is to the true value, and the higher the prediction accuracy.

[0195]

[0196] RMSE is the square root of the average of the squares of the differences between the predicted values and the true values. RMSE is very sensitive to large errors because it calculates the square of the differences. Therefore, even if only one or a few predicted values deviate significantly from the actual values, RMSE will increase significantly. This makes RMSE a good metric to punish large prediction errors.

[0197] The calculation formula of RMSE is as follows:

[0198]

[0199] Where: y i represents the true value, i.e., the observed value, represents the model's predicted value, y avg represents the average value of the true values, and n represents the number of data points.

[0200] Data set

[0201] The dataset used in this study is sourced from the wind power data of a wind farm. The data was generated by six wind turbines and three meteorological towers. The layout of the turbines and meteorological towers is as shown in Figure 6 , with a data collection frequency of once every 10 minutes. It was organized into six independent files, containing a total of 52,560 data records. Each file corresponds specifically to one wind turbine, named WT1 to WT6 in sequence. Among them, WT1 to WT4 are mainly used to collect the operation data of the onshore wind farm. In view of this, this study focuses on in-depth analysis of WT5 and WT6 wind turbines.

[0202] For the WT5 and WT6 datasets, the data was collected from January 1, 2009, to December 31, 2009, with a time resolution of 10 minutes. The WT5 dataset contains 16,444 records, and WT6 contains 19,044 records. Both datasets cover seven characteristic variables: wind speed (V), wind direction (D), air density, humidity, turbulence intensity (I), wind shear above turbine height (S _ a), and wind shear below turbine height (S _ b), and additional electric power is used as the output variable for in-depth research.

[0203] The development and implementation of the experiment rely on the PyCharm integrated development environment and are realized with the help of deep learning frameworks such as PyTorch. The data required for the experiment is sourced from reading CSV files, and then preprocessing operations are performed on the data, including but not limited to normalization, denormalization, and the application of the sliding window technique. At the initial stage of model training, the Adam optimizer is selected to optimize the model parameters. At the same time, to prevent the model from losing its generalization ability due to overfitting the training data, the early stopping strategy is implemented, that is, the training is terminated in a timely manner when the performance on the validation set no longer improves, ensuring the efficiency and robustness of model training. Finally, after strict evaluation and verification, the performance parameters of the prediction model are listed in Table 1 in detail.

[0204] Table 1 Parameter Values

[0205]

[0206] Analysis of Prediction Results

[0207] In this study, three common neural network models were selected for comparative experiments: Convolutional Neural Network (CNN), Long Short-Term Memory Network (LSTM), Transformer, and TCN-Transformer model. Given the extensive application and remarkable effectiveness of these four models in time series prediction tasks, the results of this comparative experiment are highly reliable and valuable for reference. To ensure the practicality and generalization ability of the models, actual data from two wind turbines, WT5 and WT6, were used. Experiments with different prediction steps, namely 8h, 16h, and 24h, were conducted using the above models respectively. This experimental design aims to explore the prediction capabilities of each model at different time scales to obtain more comprehensive and in-depth evaluation results.

[0208] The experimental results are shown in Table 2 and Table 3. It can be seen that DCT-TCN-Transformer demonstrated the best performance in all test scenarios. Taking the prediction result of the WT6 dataset with a prediction step of 8h as an example, compared with CNN, LSTM, TCN, Transformer, and TCN-Transformer, the R 2 increased by 5.86%, 5.50%, 6.08%, 0.37%, and 0.07% respectively, the MAE decreased by 73.39%, 72.63%, 75.28%, 35.63%, and 14.09% respectively, and the RMSE decreased by 78.30%, 77.68%, 78.72%, 36.24%, and 11.35% respectively. The above results indicate that this model performs better than other models that only capture time domain information, showing a significant improvement in prediction accuracy and having potential practical application value.

[0209] Table 2 Error metrics of different prediction models for WT5

[0210]

[0211] Table 3 Error metrics of different prediction models for WT6

[0212]

[0213] When predicting the power generation of WT5 and WT6, the error between the predicted value and the true value of the DCT-TCN-Transformer model is significantly smaller compared to other models. This finding is highly consistent with the data presented in Table 2 and Table 3 previously, further verifying the superior performance of the DCT-TCN-Transformer model in power generation prediction.

[0214] Potential applications and market prospects:

[0215] 1. Adaptability to multi-energy scenarios

[0216] The core technical innovation of this invention - the time-frequency dual-domain feature fusion mechanism can be extended to other renewable energy power generation prediction scenarios, including but not limited to:

[0217] Photovoltaic power generation: In view of the fact that the power output of photovoltaic power stations is affected by temporal factors such as sunshine intensity, cloud cover, and temperature fluctuations, the present invention can adjust the DCT window length (such as matching the daily cycle of 24 hours) to extract the frequency domain periodic characteristics of light intensity, and combine the TCN-Transformer model to capture the impact of sudden weather events (such as short-term thunderstorms) on power, thereby improving the prediction robustness.

[0218] Tidal power generation: The power output of tidal power stations has strong periodicity (daily / lunar cycles). By extracting the frequency domain principal components of tidal height time series data (such as 12-hour / 24-hour fundamental frequencies) through DCT and combining it with the long-term dependency modeling capability of TCN, the impact of tidal cycle changes on power generation can be accurately predicted.

[0219] Wave power generation: In view of the non-stationary characteristics of wave height data, DCT is used to separate low-frequency energy distribution and high-frequency random fluctuations, and the dynamic weight mechanism is combined to adaptively adjust the contribution of time-frequency domain features to optimize prediction accuracy.

[0220] 2. Technical Configurability

[0221] According to the data characteristics of different new energy scenarios, the present invention supports the following flexible configurations:

[0222] DCT parameter adjustment: Tidal energy can use a long window (72 hours) to capture the monthly cycle, and photovoltaics can use a short window (24 hours) to match the daily cycle.

[0223] Feature fusion strategy: For strongly periodic energy (such as tides), the frequency domain branch weight can be increased (α>0.7); for energy with strong randomness (such as wind power), balanced weight (α≈0.5) is used.

[0224] 3. Commercial application expansion

[0225] In addition to offshore wind power, the present invention can be embedded in photovoltaic power station monitoring systems, tidal energy dispatching platforms, etc., to provide unified power prediction services for multi-energy coordinated power grids, specifically including:

[0226] Photovoltaic field: predict the next day's power generation and optimize the charging and discharging strategy of the energy storage system.

[0227] Tidal field: Combined with astronomical tide table data, power generation peak is predicted 72 hours in advance to assist in peak load regulation of the power grid.

[0228] The research and development of offshore wind power prediction technology is of great significance for improving the predictability of wind power, optimizing energy management, promoting the utilization of renewable energy, and reducing environmental impacts. Through accurate offshore wind power prediction, we can better schedule and allocate power resources, achieve efficient utilization of energy, and promote sustainable development. In the future, with the continuous progress and optimization of technology, offshore wind power prediction technology will play an even more important role in the energy field. The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. A method for predicting the power generation of an offshore wind farm based on a temporal convolutional network and Transformer, characterized in that The method follows the steps below: Step 1: Original data collection and preprocessing Collect the original data of the offshore wind farm in the specified area, and perform preprocessing operations on the original data, including effective handling of missing values, normalization and denormalization of the data, and the application of the sliding window method, so as to generate a sample data set for model training. Step 2: Data division Divide the offshore wind power historical feature data sequence obtained in Step 1 into a training set and a test set to generate a two-dimensional data sequence. Step 3: Feature analysis Use the Pearson correlation coefficient formula to calculate the data series generated in Step 2 to obtain the correlation coefficients of each feature, and determine the feature weights used to predict the target power generation. Step 4: DCT processing Adopt DCT technology to effectively transform the time series feature information into frequency domain feature information, accurately identify and effectively remove high-frequency noise components, and obtain discrete cosine transform coefficients. Step 5: Construct a TCN model Construct a TCN model specifically for processing time series data, and use the frequency features processed by DCT in Step 4 as the input data of the TCN model. The hidden layer of the TCN model includes: dilated causal convolutional layer, Dropout layer, and activation function layer. Utilize the convolutional structure and long-term dependence characteristics of the TCN model to achieve more efficient feature extraction and time series analysis. Step 6: Encoder embedding layer processing Use the encoder embedding layer to process the input data. Convert the input data and the corresponding time stamps into an embedded format through the encoder embedding layer to provide necessary feature information support for subsequent prediction tasks. Step 7: Feature matrix fusion Perform a parallel operation on the feature information matrix processed by the TCN model and the feature information matrix passed through the encoder embedding layer to obtain a new feature matrix. This step aims to enhance feature expression, make better use of effective information, and improve prediction accuracy. Step 8: Construct a Transformer model Use the Transformer model to predict the feature matrix generated in Step 7. The encoder in the Transformer model integrates the self-attention mechanism to capture the long-range dependencies between features, a feed-forward neural network for non-linear transformation of features, residual connections to alleviate the problem of gradient disappearance, and layer normalization operations to accelerate model convergence; on this basis, the decoder additionally includes a linear layer for generating sequences and a softmax activation function, and through the synergistic effect of the self-attention mechanism, feed-forward neural network, residual connection, and layer normalization, achieve accurate prediction of subsequent time points. Step 9: Model evaluation and optimization For the above-designed prediction model, input the data of the test set into the model to evaluate and verify the prediction model and optimize the parameters.

2. A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that, In step 1, the original data of an offshore wind farm in a specified area is collected, including meteorological data such as wind speed, wind direction, air density, humidity, turbulence intensity (I), wind shear above turbine height (S_a), wind shear below turbine height (S_b), and power generation data. To eliminate the dimensional differences between different features, the linear function normalization (Min - Max Scaling) method is used to process the feature data. Through linear transformation, this method maps the original data to the interval [0, 1], and its normalization formula is as follows: Where: X′ is the data after normalization, X min is the minimum value in the dataset, X max is the maximum value in the dataset. Record these several pieces of data and remap them back to the original value range for the anti-normalization operation.

3. A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, wherein In step 2, the obtained offshore historical feature data sequence in step 1 is divided into a training set and a test set, and two - dimensional data sequences are generated for the training set and the test set. In this embodiment, it is selected to divide the obtained offshore historical feature data. The first 80% of the data is used as the training set, and 20% of the data is used as the test set. The length of the training set data is 0.8 * N, and the length of the test set data is 0.2 * N, where N is the total data length.

4. A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that, In step 3, correlation analysis is performed on the collected data. For the collected features, the Pearson correlation coefficient formula is used to calculate the feature correlation coefficient to determine the relevant features for predicting the target power generation. The Pearson correlation coefficient calculation formula is as follows: where: cov(X,Y) is the covariance of X and Y, σ X and σ Y are the standard deviations of X and Y respectively.

5. A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that, In step 4, the discrete cosine transform is applied to the time - domain information. For example, the most common 1 - D DCT definition formula for a discrete - time sequence x(n) of length N (n = 1, 2... N) is as follows: Where: y(k) represents the transformed frequency - domain signal, indicating the amplitude of the signal at frequency n, and is also called the DCT coefficient. w(k) is a normalization coefficient used to ensure that the energy of the transformation remains consistent. When using the DCT coefficient y(k) (k = 1, 2... N) to reconstruct the original time sequence x(n), this process is called the inverse discrete cosine transform (IDCT), and the definition formula is as follows: Among many DCT transformation methods, DCT - II uses the cosine function as the basis. This characteristic makes the correlation between DCT - II coefficients relatively low. In the frequency domain, low correlation means that the transformed coefficients are more independent, so these independent coefficients can be used more effectively to represent the signal. Therefore, this method uses the DCT - II method to convert the time - domain signal into the frequency - domain signal, and the formula is as follows: The inverse transform formula of DCT - II is the process of recovering the original time - domain signal from the frequency - domain coefficients after DCT - II transformation. The inverse transform formula of DCT - II can be expressed as:

6. A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that In step 5. A temporal convolutional network (TCN) model is constructed, and the specific structure of the model is as follows: The hidden layer of the TCN model includes: dilated causal convolutional layer, Dropout layer, activation function layer. Utilize the convolutional structure and long - term dependence characteristics of the TCN model to achieve more efficient feature extraction and time - series analysis. Input the dataset X and X W into the hidden layer for causal dilation respectively. Assume there is a filter set F = {f1, f2, f3,..., f K}, then the corresponding convolution operation can be expressed as: Where: s represents the current sliding position of the convolution kernel, Y(s) represents the output sequence after convolution, f i is the filter, d is the dilation factor, which is usually set to increase exponentially with the depth of the network, i.e., d = 2 m , m is the number of layers of the network, K is the size of the convolution kernel, x s-i×d is the historical data.

7. A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that, In step 6, the encoder embedding layer is used to process the data. The specific operation is as follows: The input data is processed using an encoder embedding layer and a decoder embedding layer. The input data and corresponding time stamps are converted into an embedding format through the encoder embedding layer. This embedding vector is designed to accurately capture and augment the semantic content and feature information of the input data, thus comprehensively reflecting its internal characteristics and dynamic changes. Then, the decoder embedding layer is used to process the data and time stamps of the decoder input, providing the necessary feature information support for subsequent prediction tasks.

8. A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that The feature matrix fusion in step 7. The specific steps are as follows: The frequency-domain features F output by TCN TCN ∈R L×32 are concatenated with the time-domain features F output by the encoder embedding layer Embed ∈R L×D (where D is the embedding dimension) along the feature dimension. The concatenated feature matrix contains both time-domain and frequency-domain information and is input into the Transformer model for global dependency modeling. Through the parallel connection of feature matrices, complex features of the data can be captured more comprehensively, which helps the model to use more information for better weight allocation during decision-making, thereby enhancing the stability and robustness of the model and effectively preventing information loss. At the same time, it can also effectively reduce the feature dimension without losing key information, which is helpful for the training process of the model and reduces the risk of overfitting.

9. A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that The construction of the Transformer model in step 8. The specific structure of the model is as follows: The encoder in the Transformer model includes a self-attention mechanism, a feed-forward neural network, a residual connection, and a layer normalization operation. The decoder in the Transformer model includes a self-attention mechanism, a feed-forward neural network, a residual connection, and a layer normalization. The output of the decoder is converted into a probability distribution through a linear layer and a softmax activation function to obtain the prediction result. The calculation process of the self-attention mechanism can be expressed as the dot product operation of three vectors: query (Q), key (K), and value (V), as follows: Q = XW Q K = XW K V = XW V In the formula: Q is the query matrix, K is the key matrix, and V is the value matrix. The multi-head self-attention mechanism divides the input sequence into multiple heads, and each head independently calculates the self-attention and then concatenates the results. The formula for multi-head self-attention is as follows: MultiHead(Q,K,V)=Concat(head1,...,head h )W0 In the formula: and W0 are learnable weight matrices. The feed-forward neural network usually contains two linear transformations and an activation function. Given the input x, the output of the feed-forward neural network can be expressed as: FFN = max(0, xW1 + b1)W2 + b2 In the formula: W1 is the weight matrix of the first linear transformation, W2 is the weight matrix of the second linear transformation, b1 is the bias term of the first linear transformation, and b2 is the bias term of the second linear transformation. Specifically, for each position i, its output can be expressed as the weighted sum of the values of all positions j, and the weights are calculated by the dot product of query i and key j through the softmax function. The formula is as follows: Where: Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key vector.

10. A method for predicting the power of offshore wind turbines based on a temporal convolutional network and a Transformer model according to claim 1, characterized in that The specific steps for model evaluation and optimization in step 9 are as follows: The test data in the test set is input into the trained prediction model to obtain the prediction result. According to the comparison between the predicted value and the true value, three evaluation indicators are adopted, R 2 , MAE, and RMSE. The first indicator represents the level of the model's prediction ability. The higher the score of the indicator, the better the model's prediction ability. The latter two evaluation indicators represent the error between the predicted value and the true value. The smaller the value of the indicator, the more accurate the prediction result. The calculation formulas for the evaluation indicators of the prediction model are as follows: Where: y i represents the true value, i.e., the observed value, represents the model predicted value, y avg represents the average value of the true values, and n represents the number of data points.

Citation Information

Cited By

  • Water resource monitoring method based on big data

    CN120579067A

  • Photovoltaic power short-term multi-step prediction method for enhancing feature extraction through FECAM-SERTCN

    CN120855335A

  • A photovoltaic power short-term multi-step prediction method enhanced by feature extraction of fecam-sertcn

    CN120855335B