Inception-BiLSTM-based offshore wind power prediction method
Through the interpolation method of DBSCAN and KNN, and combined with LSTM automatic encoder and feature engineering, the Inception-BiLSTM model is built, which solves the problem of insufficient prediction accuracy of offshore wind power, achieves higher prediction accuracy and reliability, and supports grid scheduling.
Patent Information
- Application Number
- CN202411713963.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional offshore wind power power prediction models are difficult to cope with the problems of violent wind speed changes and complex meteorological factors, resulting in insufficient prediction accuracy and ineffective support for power grid scheduling.
The DBSCAN clustering algorithm was used to detect and reconstruct the outliers by KNN interpolation, and repair timing abnormalities with the LSTM automatic encoder; feature engineering was carried out, key meteorological characteristics were screened, and a composite model of Inception-BiLSTM and multi-head self-attention mechanism was constructed for prediction.
It significantly improves the accuracy and reliability of offshore wind power power prediction, provides reliable technical support for offshore wind power development, and helps achieve the goals of carbon peak and carbon neutrality.
Smart Images

Figure CN120377222A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of offshore wind power, and specifically to an offshore wind power prediction method, aiming to improve the accuracy and reliability of prediction. Background Art
[0002] As an important part of renewable energy, offshore wind power plays a key role in promoting energy transformation. It not only promotes energy transformation but also provides support for achieving the goals of carbon peak and carbon neutrality. However, the development of offshore wind power faces challenges such as drastic changes in wind speed and complex meteorological factors, resulting in fluctuations and uncertainties in power output. Traditional prediction models are difficult to cope with these complexities, leading to insufficient prediction accuracy and being unable to effectively support power grid dispatching.
[0003] To solve the above problems, the present invention proposes an offshore wind power prediction method based on Inception - BiLSTM. This method improves prediction accuracy and reliability through effective data preprocessing, feature engineering, and model construction. In the data preprocessing stage, the DBSCAN clustering algorithm and LSTM autoencoder are used to detect outliers to ensure the retention of time - series features. In the feature engineering stage, key meteorological features are screened through correlation analysis, and auxiliary features such as wind direction classification and wind direction change rate are extracted. Finally, a composite model combining multi - scale convolution (Inception), bidirectional long - short - term memory neural network (BiLSTM), and multi - head self - attention mechanism is constructed to extract local features, capture bidirectional dependencies, and enhance the attention to key features. Through this series of innovative strategies, the present invention significantly improves the accuracy and robustness of wind power prediction, provides reliable technical support for offshore wind power development, and helps to achieve the goals of carbon peak and carbon neutrality. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to propose an offshore wind power prediction method based on Inception - BiLSTM, which significantly improves the accuracy and reliability of offshore wind power prediction through effective data preprocessing, feature engineering, and model construction, provides more reliable technical support for offshore wind power development, and helps to achieve the goals of carbon peak and carbon neutrality.
[0005] Technical Solution: The specific steps of the present invention are as follows:
[0006] S1: First, use the DBSCAN (Density - Based Spatial Clustering of Applications with Noise) clustering algorithm to detect outliers in the original offshore wind farm dataset, and use the KNN (K - Nearest Neighbors) interpolation method for outlier reconstruction.
[0007] S2: An innovative method based on the LSTM autoencoder is introduced to identify and repair outliers in the original data that do not satisfy temporal consistency. Through the unsupervised learning ability of the LSTM autoencoder, we can efficiently capture the inherent time-dependent patterns in the sequence data, thereby accurately locating the data points that deviate from the normal trend. Subsequently, the LSTM regression model is used to reasonably estimate and correct these outliers, ensuring that the processed dataset not only eliminates noise interference but also maximally maintains the continuity and integrity of the original time series.
[0008] S3: Feature selection is performed on the offshore wind farm data. The Pearson correlation coefficient between each numerical feature and power is calculated, and then based on the correlation analysis, features with high correlation (absolute value greater than 0.5), medium correlation (absolute value greater than 0.2 and less than or equal to 0.5), and important physical significance are selected, while low-correlation or non-correlated features (absolute value less than or equal to 0.2) are excluded.
[0009] S4: A unique scheme combining wind direction pre-classification and K-means clustering techniques is proposed to more finely analyze the potential patterns in the offshore wind farm data. First, according to the sensitivity of power output to specific wind directions in historical records, the dataset is divided into several representative wind direction intervals, and corresponding wind direction category labels are assigned to each sample. Then, the K-means algorithm is further applied to perform a secondary clustering operation on the classified dataset to discover the natural grouping structure hidden behind the large-scale observations. This dual classification strategy not only helps to reveal the variation laws of power generation efficiency under different wind directions but also provides richer and more specific context information support for the subsequent feature engineering and prediction modeling stages.
[0010] S5: In the feature engineering stage, not only traditional time-periodic features are extracted to help the model capture seasonal, daily, and weekly periodic variation laws, but also interaction features of wind speed and wind direction are extracted to help the model better understand the interaction between wind speed and wind direction and improve prediction accuracy. Finally, the innovative index of wind direction change rate is introduced, which helps the model better understand and predict the dynamic changes of wind direction, especially providing more accurate and reliable wind power predictions under unstable meteorological conditions.
[0011] S6: A composite model combining a multi-scale convolutional neural network (Inception), BiLSTM, and multi-head self-attention mechanism is proposed, which is specifically used for wind power prediction. First, the model efficiently extracts local features of the input data through three one-dimensional convolutional layers, and uses the ReLU activation function and Dropout layer to enhance non-linearity and prevent overfitting. Then, the bidirectional LSTM layer is used to capture long-term dependencies in time series data, considering the influence of historical and future data. Next, the multi-head self-attention mechanism is utilized to further enhance the attention to key features, dynamically adjusting weights to improve the representation ability and generalization performance. Finally, the complex feature combination is transformed into specific wind power prediction values through a fully connected layer.
[0012] S7: After the above steps are completed, the experiment begins. The dataset used in the Inception-BiLSTM-MultiHead-Attention hybrid neural network model in this example comes from the offshore wind power research and test base in Jiangyin Industrial Park, Fuqing City, Fujian Province. In this example, 80% of the preprocessed dataset is used as the training set, 10% as the validation set, and 10% as the test set. The wind power values are predicted, and the experimental results show that the model has good results in offshore wind power prediction.
[0013] In step S1, the processing of outliers in the dataset includes two parts. First, the DBSCN clustering method is used to detect outliers, and then the KNN interpolation method is used to reconstruct the outliers.
[0014] 1. Principle of the DBSCAN algorithm
[0015] The core idea of the DBSCAN algorithm is to detect clusters by defining neighborhoods and core points, and mark the points that do not meet these conditions as noise points (i.e., outliers). Suppose the sample set is D = (x1,..., x n ). Specifically, DBSCAN uses two key parameters: ε (neighborhood radius), MinPts (the minimum number of points required to form a cluster)
[0016] (1) Core concepts:
[0017] Neighborhood (ε-neighborhood): For a data point x i , its ε-neighborhood N ∈ (x i ) is defined as the set of all points whose distance from x i is less than or equal to ε;
[0018] Core Point: If there are at least MinPts points (including x i ) in the ε-neighborhood of x iitself), then x i is a core point;
[0019] Border Point: If x i is not a core point but lies within the ε-neighborhood of a core point, then x i is a border point;
[0020] Noise Point: A point that is neither a core point nor a border point is considered a noise point, i.e., an outlier.
[0021] (2) Density relationship:
[0022] Density direct reach: If point x j is within the ε-neighborhood of core point x i , then x j is said to be density directly reachable from x i ;
[0023] Density reachable: If there exists a series of points x1, x2,..., x n such that x1 to x n are successively density directly reachable, then x n is said to be density reachable from x1.
[0024] Density connected: If there exists a point x k such that x i and x j are both density reachable from x k , then x i and x j are density connected.
[0025] (3) Execution steps:
[0026] Step 1: Initialize parameters, set the neighborhood radius ε and the minimum number of points MinPts required to form a cluster;
[0027] Step 2: Randomly select an unvisited data point x i , and calculate the ε-neighborhood N i (x ∈ ); i )
[0028] Step 3: If |N ∈ (x i )| ≥ MinPts, then x i is a core point;
[0029] Step 4: Process border points and noise points. If x i is not a core point but lies within the ε-neighborhood of a core point, then x iis a boundary point and is assigned to the cluster of core points within the neighborhood. If x i is neither a core point nor a boundary point, it is marked as a noise point (i.e., an outlier);
[0030] Step 5: Repeat Steps 2 to 4 until all points are processed. The finally obtained noise points are the outlier data.
[0031] 2. KNN Interpolation Method
[0032] After detecting the outlier points, the KNN interpolation method is used to reconstruct the outlier points. The core idea of the KNN interpolation method is to find the K nearest neighbors of the target point based on distance measurement and use the values of these nearest neighbors to estimate the value of the target point. Specifically, the steps of the KNN interpolation method are as follows:
[0033] Step 1: Select an appropriate distance measurement. Select the Euclidean distance, which is defined as:
[0034]
[0035] where x i and x j are two d-dimensional vectors;
[0036] Step 2: Determine the value of K through cross-validation. The value of K refers to the number of nearest neighbors selected;
[0037] Step 3: For each target point x i with a missing value, calculate its distance from all other non-missing value points, and select the K points with the smallest distance as the nearest neighbors;
[0038] Step 4: Estimate the value of the target point based on the values of the K nearest neighbors. Use the weighted average method, and weight according to the distance of the nearest neighbors. The closer the neighbor, the greater the weight. The calculation formula is as follows:
[0039]
[0040] where is the estimated value of the target point, w j is the weight of the jth nearest neighbor, and usually the reciprocal distance is used as the weight.
[0041] In Step S2, the LSTM (Long Short-Term Memory) autoencoder detects the outlier points that do not conform to the time series characteristics and uses the LSTM regression method for outlier reconstruction to retain the time series characteristics.
[0042] 1. Principle of LSTM Autoencoder
[0043] The LSTM autoencoder is an unsupervised learning model based on the LSTM network, which is used to learn the latent representation of time series data and detect outliers through the reconstruction error. Its core components include an encoder and a decoder:
[0044] (1) Encoder
[0045] The input sequence X = [x1, x2,..., x T , where x t ∈R d is the input vector at time step t.
[0046] The encoder hidden state h t is updated through the LSTM cell:
[0047] h t = LSTM(x t , h t-1 )
[0048] where the update formula of the LSTM cell is as follows:
[0049] Input gate i t :
[0050] i t = σ(W i · [h t-1 , x t + b i )
[0051] Forget gate f t :
[0052] f t = σ(W f · [h t-1 , x t + b f )
[0053] Output gate o t :
[0054] o t = σ(W o · [h t-1 , x t + b o )
[0055] Cell state c t :
[0056] c t = f t ⊙ c t-1 + i t ⊙ tanh(Wc · [h t-1 , x t + b c )
[0057] Hidden state h t :
[0058] h t = o t ⊙ tanh(c t )
[0059] (2) Decoder
[0060] The decoder reconstructs the output sequence from the encoded h The initial hidden state s0 of the decoder is set to the encoded h:
[0061] s0 = h
[0062] The hidden state s of the decoder t is updated through an LSTM cell:
[0063]
[0064] where the update formula of the LSTM cell is similar to that of the encoder. The reconstructed output is obtained through a fully connected layer:
[0065]
[0066] 2. Outlier Detection
[0067] Step 1: Collect and preprocess the historical wind power data X, and divide the dataset into a training set X train and a test set X test ;
[0068] Step 2: Construct an LSTM autoencoder;
[0069] Step 3: Train the LSTM autoencoder with the goal of minimizing the reconstruction error: where θ represents the model parameters and N is the number of samples;
[0070] Step 4: Set the threshold τ = quantile(e, 0.95), where e is the distribution of the reconstruction error of the training set, and select the 95% quantile as the threshold;
[0071] Step 5: Mark the outliers:
[0072]
[0073] When the reconstruction error of a data point is greater than the set threshold, the data point is marked as an outlier (marked as '1'), otherwise it is a normal point (marked as '0').
[0074] 3. Outlier Reconstruction
[0075] Step 1: Construct the LSTM regression model LSTM regressor :
[0076] y = LSTM regressor (X)
[0077] Step 2: Use normal data (excluding outliers) to train the LSTM regression model:
[0078]
[0079] where φ represents the parameters of the regression model, and A is the index set of outliers;
[0080] Step 3: Use the trained LSTM regression model to reconstruct the outliers:
[0081]
[0082] Step 4: Replace the outliers in the original data with the reconstructed values:
[0083]
[0084] Through the above steps, the LSTM autoencoder can effectively detect outliers that do not conform to the time series characteristics and use the LSTM regression method for outlier reconstruction to retain the characteristics of the time series.
[0085] In step S3, the features of the offshore wind power are selected to improve the model performance, reduce the computational complexity, avoid overfitting, and ensure that the model can be better generalized and interpreted, so as to more accurately predict the power output and optimize the resource utilization. Therefore, it is necessary to calculate the Pearson correlation coefficient between each numerical feature and the power for feature screening and correlation analysis.
[0086] 1. Pearson Correlation Coefficient
[0087] The Pearson Correlation Coefficient (PCC) is an index to measure the linear correlation between two variables. Its value ranges from -1 to 1, where 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no correlation. Its calculation formula is as follows:
[0088]
[0089] where: X iand Y i are the X and Y values of the i-th sample respectively. and are the means of X and Y respectively. n is the number of samples.
[0090] 2. Feature Selection Steps
[0091] Step 1: Collect and preprocess the historical data of the offshore wind farm (the dataset processed through Steps S1 and S2), including environmental variables such as wind speed, wind direction, temperature, and power data.
[0092] Step 2: For each numerical feature F i , calculate the Pearson correlation coefficient between it and the power P The specific calculation steps are as follows:
[0093] (1) Calculate the mean i of feature F and the mean
[0094]
[0095] (2) Calculate the covariance Cov(F i , P):
[0096]
[0097] (3) Calculate the standard deviation i of feature F and the standard deviation σ P of power P:
[0098]
[0099] (4) Calculate the Pearson correlation coefficient:
[0100]
[0101] Step 3: Select features based on the correlation analysis. According to the calculated correlation coefficients, they are divided into three categories: highly correlated features (the absolute value of the Pearson correlation coefficient is greater than 0.5), moderately correlated features (the absolute value of the Pearson correlation coefficient is greater than 0.2 and less than or equal to 0.5), and low or uncorrelated features (the absolute value of the Pearson correlation coefficient is less than or equal to 0.2);
[0102] Step 4: Select the highly correlated and moderately correlated features, which have a significant impact on power prediction. Exclude those features with low or no correlation with power to reduce the complexity of the model and improve the prediction performance.
[0103] In step S4, first, classify the data set into different wind direction intervals according to the wind direction and generate wind direction classification labels; then use the K-means clustering method to cluster the data and generate clustering labels. Finally, save the wind direction classification labels and clustering labels in the data set for further analysis and modeling.
[0104] 1. Pre-classification of wind direction
[0105] Step 1: Assume that the data set D contains wind direction θ i and the corresponding power P i , where i = 1, 2,..., n;
[0106] D = {(θ1, P1), (θ2, P2),..., (θ n , P n )}
[0107] Step 2: Sort the data according to the wind direction value θ i to obtain a new data set D':
[0108] D' = {(θ1', P1'), (θ2', P2'),..., (θ n ', P n ')}
[0109] θ i ' ≤ θ i+1 ', i = 1, 2,..., n - 1
[0110] Step 3: Calculate the cumulative power and the cumulative power mean:
[0111]
[0112] where k = 1, 2,..., n, P' i is the sorted power value;
[0113] Step 4: Determine the splitting points by finding the points where the change in the cumulative power mean is the largest. Specifically, we look for the k that makes the largest:
[0114]
[0115] Repeat this process and find n splitting points k1, k2, k3,..., k n ;
[0116] Step 5: Generate wind direction classification labels according to the splitting points k1, k2, k3,..., k n :
[0117]
[0118] By generating wind direction classification labels, the power output patterns under different wind direction conditions can be better reflected.
[0119] 2. K-means Clustering
[0120] Based on the pre-classification of wind directions, the K-means algorithm is used to cluster the data and generate clustering labels. K-means is an iterative, distance-based clustering algorithm, and its goal is to divide the data set D into k clusters: C1, C2,..., C k , such that the data points within each cluster are as similar as possible, while the data points between different clusters are as different as possible. At the same time, minimize the within-cluster sum of squares (WCSS):
[0121]
[0122] where c i is the center of cluster C i . The specific implementation steps of K-means clustering are as follows:
[0123] Step 1: Initialization, select a value of k, that is, the number of clusters to be divided. Randomly select k data points as the initial cluster centers: c1, c2,..., c k ;
[0124] Step 2: Assign data points to the nearest cluster. For each data point x i , calculate its distance d(x i , c j ):
[0125]
[0126] where d is the feature dimension. Then assign the data point x i to the cluster to which the nearest cluster center c j belongs:
[0127]
[0128] Step 3: Update the cluster centers, calculate the new centers of each cluster:
[0129]
[0130] where |C j | is the number of data points in cluster C j ;
[0131] Step 4: Repeat Step 2 and Step 3 until the cluster centers do not change or change very little, and stop the iteration. Otherwise, return to Step 2 to continue the iteration. Finally, generate the clustering label column.
[0132] This K-means clustering method based on wind direction pre-classification can more effectively identify natural groupings and patterns in wind power data, optimize resource allocation, thereby enhancing model performance and interpretability, and providing more reliable data analysis results.
[0133] In step S5, in the analysis of time series, especially in the wind power prediction task containing meteorological information, extracting different types of features can help the model better understand the patterns in the data. Extracting time periodic features helps the model identify and utilize the repeating patterns in the data, and understand the changing rules of seasonality, daily periodicity, and weekly periodicity; extracting the interaction features of wind speed and wind direction helps capture the relationship between wind speed and wind direction, and at the same time, converting the wind direction into sine and cosine components can avoid the problem of angle discontinuity; extracting the wind direction change rate helps the model understand the changing trend of wind direction over time, can better capture the dynamic characteristics of the wind, improve the model's adaptability to unstable wind conditions, and thus improve the prediction accuracy.
[0134] 1. Time periodic features
[0135] (1) Seasonality: There are 365 days in a year (or 366 days in a leap year), and seasonality can be represented by sine and cosine functions
[0136]
[0137] where the unit of t is days, that is, the day of the year;
[0138] (2) Daily periodicity: There are 24 hours in a day, and daily periodicity can be represented similarly
[0139]
[0140] where the unit of t is hours, that is, the hour of the day;
[0141] (3) Weekly periodicity: There are 7 days in a week, and it can also be represented by sine and cosine functions
[0142]
[0143] where the unit of t is days, that is, the day of the week.
[0144] 2. Interaction features of wind speed and wind direction:
[0145] InteractionSin = V·sin(θ)
[0146] InteractionCos = V·cos(θ)
[0147] Where V is the wind speed, θ is the wind direction, and sin(θ) and cos(θ) represent the sine and cosine components of the wind direction respectively.
[0148] 3. Wind direction change rate:
[0149] (1) Calculate the interpolation between adjacent time points:
[0150] Δθ(t) = θ(t) — θ(t — 1)
[0151] (2) Handle the cases of crossing 0 degrees or 360 degrees:
[0152]
[0153] (3) Final wind direction change rate:
[0154] WindDirRate(t) = Δθ(t)
[0155] In step S6, the constructed wind power prediction model combines a multi-scale convolutional neural network (Inception), a bidirectional long short-term memory network (BiLSTM), and a multi-head self-attention mechanism. This model first efficiently extracts local features of the input data through a one-dimensional multi-scale convolutional layer, and uses the ReLU activation function and Dropout layer to enhance non-linearity and prevent overfitting. Then, the bidirectional LSTM layer is used to capture the long-term dependencies in the time series, considering the influence of historical and future data. Subsequently, the multi-head self-attention mechanism further strengthens the attention to key features, dynamically adjusts the weights to improve the representation ability and generalization performance. Finally, the fully connected layer converts the complex feature combination into a specific wind power prediction value.
[0156] 1. Inception module
[0157] The Inception module is a variant of the convolutional neural network (CNN). It captures multi-scale information by parallelly using multiple convolutional kernels of different sizes in one module. The core idea of the Inception module is to increase the diversity of the receptive field of the network while reducing the computational amount. The Inception module usually contains the following branches: 1×1 convolution (used to reduce the dimension and computational amount), 3×3 convolution (capturing local features), 5×5 convolution (capturing larger-range features), and max pooling (providing spatial invariance). The outputs of these branches are concatenated in the channel dimension to form the final output. Assume the input tensor is X ∈ R N×C×H×W , where N is the batch size, C is the number of input channels, and H and W are the height and width of the input respectively. Its basic principle is as follows:
[0158] (1) 1×1 convolution branch:
[0159] Y1 = ReLU(W 1×1 * X + b 1×1 )
[0160] where W 1×1 is the weight matrix of the 1×1 convolutional kernel, and b 1×1 is the bias term, and the output
[0161] (2) 3×3 Convolution Branch:
[0162] Y2 = ReLU(W 3×3 * ReLU(W 1×1 * X + b 1×1 ) + b 3×3 )
[0163] where W 1×1 is the weight matrix of the first 1×1 convolutional kernel, and b 1×1 is the bias term of the first 1×1 convolution, W 3×3 is the weight matrix of the 3×3 convolutional kernel, and b 3×3 is the bias term of the 3×3 convolution, and the output
[0164] (3) 5×5 Convolution Branch:
[0165] Y3 = ReLU(W 5×5 * ReLU(W 1×1 * X + b 1×1 ) + b 5×5 )
[0166] where W 1×1 is the weight matrix of the first 1×1 convolutional kernel, and b 1×1 is the bias term of the first 1×1 convolution, W 3×3 is the weight matrix of the 5×5 convolutional kernel, and b 5×5 is the bias term of the 5×5 convolution, and the output
[0167] (4) Max Pooling Branch:
[0168] Y4 = ReLU(W 1×1 * MaxPool(X) + b 1×1 )
[0169] where MaxPool is the 3×3 max pooling layer, W 1×1 is the weight matrix of the first 1×1 convolutional kernel, and b 1×1 is the bias term of the first 1×1 convolution, and the output
[0170] (5) Concatenated Output:
[0171] Y = Concat(Y1, Y2, Y3, Y4
[0172] where the final output
[0173] 2. BiLSTM Network
[0174] Compared with LSTM, BiLSTM contains two independent LSTM layers, one processes the sequence from front to back (forward LSTM), and the other processes the sequence from back to front (backward LSTM). The hidden state h at each time step t t is the forward LSTM hidden state and the backward LSTM hidden state is the concatenation or combination of. The hidden layer of BiLSTM can be expressed as:
[0175]
[0176] In the formula, is the element-wise sum of the forward output component and the backward output component. The BiLSTM network is as Figure 3 shown. Suppose we have an input sequence {x1, x2,..., x T}, where T is the length of the sequence. The input sequence is input into the forward LSTM layer in order. The forward LSTM layer processes the input step by step in time and generates the hidden state output Each hidden state depends on the current input x t and the hidden state at the previous moment The input sequence is input into the backward LSTM layer in reverse order. The backward LSTM layer processes the input step by step in time and generates the hidden state output Each hidden state depends on the current input x t and the hidden state at the next moment At each time step t, the hidden states of the forward and backward LSTM layers are concatenated to obtain the final hidden state h t . Specifically, where the semicolon represents the concatenation operation of vectors or matrices. In this way, the hidden state h at each time step t t contains both information from the past to the present and information from the future to the present.
[0177] 3. Multi-Head Self-Attention Mechanism
[0178] Fusing the multi-head attention mechanism (Multi-Head Self-Attention) in a prediction model allows the model to obtain information from parts of the sequence and learn different types of dependencies in parallel. This mechanism is very effective for processing sequence data. The basic principle of the multi-head self-attention mechanism is as follows:
[0179] (1) Basic definition
[0180] Suppose the input is a tensor X of shape [batch_size, sequence_length, input_dim]. batch_size represents the number of samples in a batch; sequence_length represents the length of the sequence, that is, how many time steps or elements there are in each sample; input_dim represents the feature dimension of each time step (or each element). For each attention head i (i = 1, 2,..., h, where h is the number of attention heads), we need three linear transformation matrices: the query weight matrix the key weight matrix the value weight matrix where d k is the output dimension of a single attention head, usually These matrices are used to generate query, key, and value vectors from the input X. The specific transformation is as follows:
[0181]
[0182] where Q i , K i , V i are all tensors of shape [batch_size, sequence_length, d k .
[0183] (2) Calculate attention scores
[0184] For each attention head i, calculate its attention scores:
[0185]
[0186] Here represents the transpose of K i , and is the scaling factor used to stabilize the gradient;
[0187] (3) Apply the Softmax function
[0188] Apply the Softmax function to the attention scores to obtain the attention weights:
[0189] AttentionWeights i = softmax(Attention Scores i )
[0190] where the shape of the attention weights is [batch_size, sequence_length, sequence_length], indicating the influence degree of each position in the sequence on all other positions;
[0191] (4) Weighted sum
[0192] Use the attention weights to perform a weighted sum on the value vectors to obtain the output of the i-th attention head:
[0193] Head i = Attention Weights i ·V i
[0194] At this time, the shape of Head i is [batch_size, sequence_length, d k ;
[0195] (5) Multi-head merging
[0196] Concatenate the outputs of all attention heads, and then project them back to the original feature space through an additional linear layer W O :
[0197] Concatenated Heads = [Head1; Head2...; Head h
[0198] Multi-Head Output = Concatenated Heads · W O
[0199] where it is used to keep the concatenated multi-head attention output consistent with the initial input X in the feature dimension. The final Multi-Head Output will be a tensor with the shape of [batch_size, sequence_length, input_dim], which fuses the information from different attention heads. In this way, the multi-head self-attention mechanism can capture the complex dependencies in the sequence data and learn information in parallel on different subspaces, thus improving the model's ability to handle sequence tasks.
[0200] 4. Inception-BiLSTM-MultiHead-Attention Hybrid Model
[0201] The Inception-BiLSTM-MultiHead-Attention hybrid model (abbreviated as IBMHA-Model) proposed by the present invention has the following advantages:
[0202] (1) Multi-scale feature extraction: The Inception module provides multi-scale receptive fields, enabling the model to extract rich local features from convolutional kernels of different scales. This not only enhances the ability to capture subtle patterns in the input data but also improves the model's adaptability to different types of information;
[0203] (2) Long-term dependence modeling and context understanding: The BiLSTM layer allows the model to perform bidirectional information flow in the time dimension, effectively capturing long-term dependence relationships in the sequence and being able to consider context information from both the past and the future, further enhancing the depth of understanding of sequence data.
[0204] (3) Adaptive focus on important information: The multi-head self-attention mechanism enables the model to learn different representations of the input sequence in parallel in different subspaces and to adaptively focus on the most important parts. This mechanism significantly enhances the model's performance and robustness in handling complex tasks.
[0205] (4) Flexibility and scalability: This architecture can be easily adjusted to meet different task requirements, such as optimizing model performance by changing the convolutional kernel size, the number of LSTM layers, or the number of attention heads.
[0206] In step S7, the dataset used by the Inception-BiLSTM-MultiHead-Attention hybrid neural network model in this example comes from the offshore wind power research and test base in Jiangyin Industrial Park, Fuqing City, Fujian Province. In this example, 80% of the preprocessed dataset is used as the training set, 10% as the validation set, and 10% as the test set. The dataset includes meteorological data and actual power data, as shown in Appendix 1.
[0207] After the preprocessing and cleaning of the dataset in steps S1 and S2, the outliers were correctly reconstructed and the missing values were filled. Then, in step S3, the features with greater contribution to the output power were screened, and in step S5, more features were extracted. The finally retained meteorological features are: 10-meter wind speed, 10-meter wind direction, 100-meter wind speed, 100-meter wind direction, temperature, air pressure, wind speed and wind direction interaction features, and wind direction change rate. Since this dataset is sampled every 15 minutes, the input sequence length is set to 96, that is, each sample is one-day data, so as to fully consider the change trend of historical data.
[0208] To verify the accuracy of the present invention, taking 00:00 of a certain day in the test set as the start time, the prediction effects of different time steps were analyzed. The final prediction results are the power values at 24 hours (96 time steps), 48 hours (192 time steps), and 72 hours (288 time steps), as shown in the appendix Figure 5 - Appendix Figure 7 as shown. The mean absolute error (MAE) and root mean square error (RMSE) are shown in Table 2 of the appendix.
[0209] Appendix Chart Explanation
[0210] Figure 1 is the structure diagram of the LSTM unit.
[0211] Figure 2 is the structure diagram of the Inception module.
[0212] Figure 3 is the structure diagram of the BiLSTM network.
[0213] Figure 4 is the multi-head self-attention mechanism
[0214] Figure 5 is the structure diagram of the IBMHA-Model
[0215] Figure 6 is the comparison chart of the 24-hour predicted value and the true value
[0216] Figure 7 is the comparison chart of the 48-hour predicted value and the true value
[0217] Figure 8 is the comparison chart of the 72-hour predicted value and the true value Specific Embodiment
[0220] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0221] A method for predicting the power of offshore wind farms based on Inception - BiLSTM includes the following steps:
[0222] S1: First, use the DBSCAN (Density - Based Spatial Clustering of Applications with Noise) clustering algorithm to detect outliers in the original offshore wind farm dataset, and use the KNN (K - Nearest Neighbors) interpolation method to reconstruct the outliers.
[0223] S2: Introduce an innovative method based on the LSTM auto - encoder to identify and repair outliers that do not meet the temporal consistency in the original data. Through the unsupervised learning ability of the LSTM auto - encoder, we can efficiently capture the inherent time - dependent patterns in the sequence data, thereby accurately locating the data points that deviate from the normal trend. Subsequently, use the LSTM regression model to reasonably estimate and correct these outliers, ensuring that the processed dataset not only eliminates noise interference but also maximally maintains the continuity and integrity of the original time series.
[0224] S3: Perform feature selection on the offshore wind farm data, calculate the Pearson correlation coefficient between each numerical feature and power, and then based on the correlation analysis, select features with high correlation (absolute value greater than 0.5), medium correlation (absolute value greater than 0.2 and less than or equal to 0.5), and important physical significance, and exclude features with low or no correlation (absolute value less than or equal to 0.2).
[0225] S4: A unique solution combining wind direction pre-classification and K-means clustering techniques is proposed to more precisely analyze the potential patterns in the data of offshore wind farms. First, based on the sensitivity of power output to specific wind directions in historical records, we divide the dataset into several representative wind direction intervals and assign corresponding wind direction class labels to each sample. Then, the K-means algorithm is further applied to perform a secondary clustering operation on the classified dataset to discover the natural grouping structure hidden behind the large-scale observations. This dual classification strategy not only helps to reveal the variation laws of power generation efficiency under different wind directions but also provides richer and more specific context information support for the subsequent feature engineering and prediction modeling phases.
[0226] S5: In the feature engineering stage, not only traditional time periodic features are extracted to help the model capture seasonal, daily, and weekly periodic variation laws, but also interaction features of wind speed and wind direction are extracted, which helps the model better understand the interaction between wind speed and wind direction and improve prediction accuracy. Finally, the index of wind direction change rate is innovatively introduced, which helps the model better understand and predict the dynamic changes of wind direction, especially providing more accurate and reliable wind power predictions under unstable meteorological conditions.
[0227] S6: A composite model combining multi-scale convolutional neural network (Inception), bidirectional long short-term memory network (BiLSTM), and multi-head self-attention mechanism is proposed for wind power prediction. The model first efficiently extracts local features of the input data through three one-dimensional convolutional layers and uses the ReLU activation function and Dropout layer to enhance non-linearity and prevent overfitting. Then, the bidirectional LSTM layer is used to capture long-term dependencies in time series data, considering the influence of both historical and future data. Next, the multi-head self-attention mechanism is utilized to further enhance the attention to key features, dynamically adjusting weights to improve the representation ability and generalization performance. Finally, the complex feature combinations are transformed into specific wind power prediction values through the fully connected layer.
[0228] S7: After the above steps are completed, the experiment begins. The dataset used in the Inception-BiLSTM-MultiHead-Attention hybrid neural network model in this example comes from the offshore wind power research and test base in Jiangyin Industrial Park, Fuqing City, Fujian Province. In this example, 80% of the pre-processed dataset is used as the training set, 10% as the validation set, and 10% as the test set. Predictions are made on the wind power values, and the experimental results show that the model has good performance in offshore wind power prediction.
[0229] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for predicting the power of offshore wind farms based on Inception - BiLSTM, characterized in that, It includes the following steps: S1: First, use the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm to detect outliers in the original offshore wind farm dataset, and adopt the KNN (K-Nearest Neighbors) interpolation method for outlier reconstruction. S2: An innovative method based on the LSTM autoencoder is introduced to identify and repair outliers that do not meet the temporal consistency in the original data. Through the unsupervised learning ability of the LSTM autoencoder, we can efficiently capture the inherent time-dependent patterns in the sequence data, thereby accurately locating the data points that deviate from the normal trend. Subsequently, the LSTM regression model is used to reasonably estimate and correct these outliers, ensuring that the processed dataset not only eliminates noise interference but also maximally maintains the continuity and integrity of the original time series. S3: Perform feature selection on the offshore wind farm data, calculate the Pearson correlation coefficient between each numerical feature and power, and then based on the correlation analysis, select features with high correlation (absolute value greater than 0.5), medium correlation (absolute value greater than 0.2 and less than or equal to 0.5), and important physical significance, and exclude features with low or no correlation (absolute value less than or equal to 0.2). S4: A unique scheme combining wind direction pre-classification and K-means clustering technology is proposed to more finely analyze the potential patterns in the offshore wind farm data. First, according to the sensitivity of power output to specific wind directions in historical records, we divide the dataset into several representative wind direction intervals and assign corresponding wind direction category labels to each sample. Then, further apply the K-means algorithm to perform a secondary clustering operation on the classified dataset to discover the natural grouping structure hidden behind the large-scale observations. This dual classification strategy not only helps to reveal the variation laws of power generation efficiency under different wind directions but also provides richer and more specific context information support for the subsequent feature engineering and prediction modeling stages. S5: In the feature engineering stage, not only traditional time periodic features are extracted to help the model capture seasonal, daily, and weekly periodic variation laws, but also interaction features of wind speed and wind direction are extracted to help the model better understand the interaction between wind speed and wind direction and improve the prediction accuracy. Finally, the innovative index of wind direction change rate is introduced, which helps the model better understand and predict the dynamic changes of wind direction, especially providing more accurate and reliable wind power predictions under unstable meteorological conditions. S6: A composite model combining multi-scale convolutional neural network (Inception), bidirectional long short-term memory neural network (BiLSTM), and multi-head self-attention mechanism is proposed, which is specifically used for wind power prediction. First, the model efficiently extracts local features of the input data through three one-dimensional convolutional layers, and uses the ReLU activation function and Dropout layer to enhance non-linearity and prevent overfitting. Then, the bidirectional LSTM layer is used to capture long-term dependencies in time series data, considering the influence of historical and future data. Next, the multi-head self-attention mechanism is used to further enhance the attention to key features, dynamically adjust weights to improve the representation ability and generalization performance. Finally, the complex feature combination is transformed into specific wind power prediction values through a fully connected layer. S7: After the above steps are completed, the experiment starts. The dataset used in the Inception-BiLSTM-MultiHead-Attention hybrid neural network model in this example comes from the offshore wind power research and test base in Jiangyin Industrial Park, Fuqing City, Fujian Province. In this example, 80% of the preprocessed dataset is used as the training set, 10% as the validation set, and 10% as the test set. The wind power values are predicted, and the experimental results show that the model has good performance in offshore wind power prediction.
2. The method for predicting the power of offshore wind turbines based on Inception - BiLSTM according to claim 1, wherein: The processing of outliers in the dataset in step S1 includes two parts. First, the DBSCN clustering method is used to detect outliers, and then the KNN interpolation method is used to reconstruct the outliers. (1) DBSCAN algorithm The core idea of the DBSCAN algorithm is to detect clusters by defining neighborhoods and core points, and mark points that do not meet these conditions as noise points (i.e., outliers). Suppose the sample set is D = (x1, x2,..., x n ), specifically, DBSCAN uses two key parameters: ε (neighborhood radius), MinPts (the minimum number of points required to form a cluster) Core concept: Neighborhood (ε-neighborhood): For a data point x i , its ε-neighborhood N ∈ (x i ) is defined as the set of all points whose distance from x i is less than or equal to ε; Core Point: If there are at least MinPts points (including x itself) in the ε-neighborhood of x i , then x i is a core point; i Border Point: If x i is not a core point but lies within the ε-neighborhood of a certain core point, then x i is a border point; Noise Point: A point that is neither a core point nor a border point is considered a noise point, that is, an outlier. Execution steps: Step 1: Initialize parameters, set the neighborhood radius ε and the minimum number of points MinPts required to form a cluster; Step 2: Randomly select a data point x in an unvisited state i , and calculate the ε-neighborhood N i of x ∈ (x i ); Step 3: If |N ∈ (x i )| ≥ MinPts, then x i is a core point; Step 4: Process boundary points and noise points. If x i is not a core point but lies within the ε-neighborhood of a certain core point, then x i is a boundary point and is assigned to the cluster of the core point within the neighborhood. If x i is neither a core point nor a boundary point, then it is marked as a noise point (i.e., an outlier); Step 5: Repeat steps 2 to 4 until all points are processed. The finally obtained noise points are the outlier data. (2) KNN interpolation method After detecting outliers, the KNN interpolation method is used to reconstruct the outliers. The core idea of the KNN interpolation method is to find the K nearest neighbors of the target point based on distance measurement and use the values of these nearest neighbors to estimate the value of the target point. Specifically, the steps of the KNN interpolation method are as follows: Step 1: Select a suitable distance measurement. The Euclidean distance is selected and defined as: where x i and x j are two d-dimensional vectors; Step 2: Determine the value of K through cross-validation. The value of K refers to the number of selected nearest neighbors; Step 3: For each target point x containing missing values i , calculate the distances between it and all other non-missing value points, and select the K points with the smallest distances as the nearest neighbors; Step 4: Estimate the value of the target point based on the values of the K nearest neighbors. The weighted average method is used, and the weights are assigned according to the distance of the nearest neighbors. The closer the neighbor, the greater the weight. The calculation formula is as follows: Among them is the target point estimate value, and w j is the weight of the j-th nearest neighbor. Usually, the reciprocal distance is used as the weight.
3. A method for predicting the power of offshore wind turbines based on Inception - BiLSTM according to claim 1, characterized in that: In step S2, an LSTM (Long Short-Term Memory) autoencoder is introduced to detect outliers that do not conform to the temporal characteristics, and the LSTM regression method is used for outlier reconstruction to preserve the time series characteristics. The LSTM autoencoder is an unsupervised learning model based on the LSTM network, which is used to learn the latent representation of time series data and detect outliers through the reconstruction error. The basic principle of this method is as follows: (1) Encoder Input sequence X = [x1, x2,..., x T , where x t ∈R d is the input vector at time step t. The encoder hidden state h t is updated through the LSTM cell: h t = LSTM(x t , h t-1 ) Among them, the update formula of the LSTM cell is as follows: Input gate i t : i t = σ(W i · [h t-1 , x t + b i ) Forgotten gate f t : f t = σ(W f · [h t-1 , x t + b f ) Output gate o t : o t = σ(W o · [h t-1 , x t + b o ) Cell state c t : c t = f t ⊙ c t-1 + i t ⊙ tanh(W c · [h t-1 , x t + b c ) Hidden state h t : h t = o t ⊙tanh(c t (2) Decoder The decoder reconstructs the output sequence from the encoding h The initial hidden state s0 of the decoder is set to the encoding h: s0 = h The hidden state s of the decoder t Updated by the LSTM cell: Among them, the update formula of the LSTM cell is similar to that of the encoder. The reconstructed output is obtained through the fully connected layer: (3) Detect outliers Step 1: Collect and preprocess the historical wind power data X, and divide the data set into a training set X train and a test set X test ; Step 2: Construct an LSTM autoencoder; Step 3: Train the LSTM autoencoder with the goal of minimizing the reconstruction error: where θ represents the model parameters and N is the number of samples; Step 4: Set the threshold τ = quantile(e, 0.95), where e is the reconstruction error distribution of the training set, and select the 95% quantile as the threshold; Step 5: Mark outliers: When the reconstruction error of a data point is greater than the set threshold, the data point is marked as an outlier (marked as '1'), otherwise it is a normal point (marked as '0'). (4) Reconstruct outliers Step 1: Construct the LSTM regression model LSTM regressor : y = LSTM regressor (X) Step 2: Train an LSTM regression model using normal data (excluding outliers): Among them, φ represents the parameters of the regression model, and A is the set of indices of outliers; Step 3: Use the trained LSTM regression model to reconstruct the outliers; Step 4: Replace the outliers in the original data with the reconstructed values: Through the above steps, the LSTM autoencoder can effectively detect outliers that do not conform to the temporal characteristics and use the LSTM regression method for outlier reconstruction to preserve the characteristics of the time series.
4. A method for predicting the power of offshore wind turbines based on Inception - BiLSTM according to claim 1, characterized in that: In step S3, the features of the offshore wind power are selected to improve the model performance, reduce the computational complexity, avoid overfitting, and ensure that the model can be better generalized and interpreted, so as to more accurately predict the power output and optimize the resource utilization. Therefore, it is necessary to calculate the Pearson correlation coefficient between each numerical feature and the power to perform feature screening and correlation analysis. (1) Pearson correlation coefficient Its value ranges from -1 to 1, where 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no correlation. In the formula, X i and Y i are the X and Y values of the i-th sample respectively. and are the means of X and Y respectively. n is the number of samples. (2) Feature selection steps Step 1: Collect and preprocess the historical data of the offshore wind farm (the dataset processed by steps S1 and S2), including environmental variables such as wind speed, wind direction, temperature, and power data. Step 2: For each numerical feature F i , calculate the Pearson correlation coefficient ρ between it and the power P Fi,P . Step 3: Select features based on correlation analysis. According to the calculated correlation coefficients, they are divided into three categories: highly correlated features (the absolute value of the Pearson correlation coefficient is greater than 0.5), moderately correlated features (the absolute value of the Pearson correlation coefficient is greater than 0.2 and less than or equal to 0.5), and low or uncorrelated features (the absolute value of the Pearson correlation coefficient is less than or equal to 0.2); Step 4: Select highly correlated and moderately correlated features, which have a significant impact on power prediction. Exclude those features with low or no correlation with power to reduce the complexity of the model and improve the prediction performance.
5. The method for predicting the power of offshore wind turbines based on Inception - BiLSTM according to claim 1, wherein: In step S4, first classify the data set into different wind direction intervals according to the wind direction and generate wind direction classification labels; then use the K-means clustering method to cluster the data and generate clustering labels. Finally, save the wind direction classification labels and clustering labels in the data set for further analysis and modeling. (1) Preliminary wind direction classification Step 1: Assume that the dataset D contains the wind direction θ i and the corresponding power P i , where i = 1, 2, …, n; D = { (θ1, P1), (θ2, P2),..., (θ n , P n )} Step 2: Sort the data according to the wind direction value θ i to obtain a new data set D': D' = { (θ1', P1'), (θ2', P2'),..., (θ n ', P n )} θ i ' ≤ θ i+1 ', i = 1, 2,..., n - 1 Step 3: Calculate the cumulative power and the average cumulative power: where k = 1, 2, …, n, P i ' is the sorted power value; Step 4: Determine the segmentation point by finding the point with the largest change in the cumulative power mean. Specifically, we look for the k that maximizes : Repeat this process and find n splitting points k1, k2, k3, ..., k through the method of cross-validation n ; Step 5: According to the segmentation points k1, k2, k3,..., k n , generate wind direction classification labels: By generating wind direction classification labels, the power output patterns under different wind direction conditions can be better reflected. (2) K-means clustering Based on the pre-classification of wind direction, the K-means algorithm is used to cluster the data and generate cluster labels. K-means is an iterative, distance-based clustering algorithm whose goal is to partition the data set D into k clusters: C1, C2, ..., C k , such that the data points within each cluster are as similar as possible, while the data points between different clusters are as different as possible. At the same time, minimize the within-cluster sum of squares (WCSS): where c i is the center of cluster C i . The specific implementation steps of K-means clustering are as follows: Step 1: Initialization. Select a value of k, i.e., the number of clusters to be divided. Randomly select k data points as the initial cluster centers: c1, c2, ..., c k ; Step 2: Assign data points to the nearest cluster. For each data point x i , calculate its distance d(x i , c j ) from each cluster center: where d is the feature dimension. Then the data point x i is assigned to the cluster to which the nearest cluster center c j belongs: Step 3: Update the cluster centers and calculate the new centers of each cluster: where |C j | is the number of data points in cluster C j ; Step 4: Repeat Step 2 and Step 3 until the cluster centers do not change or change very little, and stop the iteration. Otherwise, return to Step 2 to continue the iteration. Finally, generate the clustering label column. This K-means clustering method based on preliminary wind direction classification can more effectively identify the natural groupings and patterns in wind power data, optimize resource allocation, thereby enhancing the model performance and interpretability, and providing more reliable data analysis results.
6. The method for predicting the power of offshore wind turbines based on Inception - BiLSTM according to claim 1, wherein: In step S5, extracting different types of features for the wind power prediction task can help the model better understand the patterns in the data. Extracting time periodic features helps the model identify and utilize the repeating patterns in the data, and understand the changing rules of seasonality, daily periodicity, and weekly periodicity; extracting the interaction features of wind speed and wind direction helps capture the relationship between wind speed and wind direction, and at the same time, converting the wind direction into sine and cosine components can avoid the problem of angular discontinuity; extracting the wind direction change rate helps the model understand the changing trend of the wind direction over time, can better capture the dynamic characteristics of the wind, improve the model's adaptability to unstable wind conditions, and thus improve the prediction accuracy. (1) Time periodic features Step 1: Extract seasonal features. There are 365 days in a year (or 366 days in a leap year), and the seasonality can be represented by sine and cosine functions: Step 2: Extract daily periodic features. There are 24 hours in a day, and the daily periodicity can be represented similarly, where the unit of t is hours, that is, the number of hours of each day: Step 3: Extract weekly periodic features. There are 7 days in a week, and it can also be represented by sine and cosine functions, where the unit of t is days, that is, the number of days of each week: (2) Interaction features of wind speed and wind direction InteractionSin = V·sin(θ) InteractionCos = V·cos(θ) where V is the wind speed, θ is the wind direction, and sin(θ) and cos(θ) represent the sine and cosine components of the wind direction respectively. (3) Wind direction change rate Step 1: Calculate the interpolation between adjacent time points: Δθ(t) = θ(t) — θ(t — 1) Step 2: Handle the cases where it crosses 0 degrees or 360 degrees: Step 3: Obtain the final wind direction change rate: WindDirRate(t) = Δθ(t).
7. A method for predicting the power of offshore wind turbines based on Inception-BiLSTM according to claim 1, wherein: In step S6, the constructed wind power prediction model combines a multi-scale convolutional neural network (Inception), a bidirectional long short-term memory network (BiLSTM), and a multi-head self-attention mechanism. First, the one-dimensional multi-scale convolutional layer efficiently extracts the local features of the input data, and the ReLU activation function and Dropout layer are used to enhance the non-linearity and prevent overfitting. Then, the bidirectional LSTM layer is used to capture the long-term dependencies in the time series, considering the influence of both historical and future data. Subsequently, the multi-head self-attention mechanism further strengthens the attention to key features, dynamically adjusts the weights to improve the representation ability and generalization performance. Finally, the fully connected layer converts the complex feature combinations into specific wind power prediction values. (1) Inception module The Inception module is a variant of the convolutional neural network (CNN) that captures multi-scale information by using multiple convolutional kernels of different sizes in parallel within a module. The core idea of the Inception module is to increase the diversity of the receptive fields of the network while reducing the computational cost. The Inception module typically consists of the following branches: 1×1 convolution (used to reduce the dimension and computational cost), 3×3 convolution (capturing local features), 5×5 convolution (capturing features in a larger range), and max pooling (providing spatial invariance). The outputs of these branches are concatenated along the channel dimension to form the final output. Suppose the input tensor is X ∈ R N×C×H×W , where N is the batch size, C is the number of input channels, and H and W are the height and width of the input respectively. The basic execution steps are as follows: Step 1: Calculate the output of the 1×1 convolutional branch: Y1 = ReLU(W 1×1 *X + b 1×1 Among them, W 1×1 is the weight matrix of the 1×1 convolutional kernel, and b 1×1 is the bias term, and the output Step 2: Calculate the output of the 3×3 convolutional branch: Y2 = ReLU(W 3×3 * ReLU(W 1×1 * X + b 1×1 ) + b 3×3 ) where W 1×1 is the weight matrix of the first 1×1 convolutional kernel, and b 1×1 is the bias term of the first 1×1 convolution, W 3×3 is the weight matrix of the 3×3 convolutional kernel, and b 3×3 is the bias term of the 3×3 convolution, and the output Step 3: Calculate the output of the 5×5 convolutional branch: Y3 = ReLU(W 5×5 * ReLU(W 1×1 * X + b 1×1 ) + b 5×5 ) where W 1×1 is the weight matrix of the first 1×1 convolutional kernel, and b 1×1 is the bias term of the first 1×1 convolution. W 3×3 is the weight matrix of the 5×5 convolutional kernel, and b 5×5 is the bias term of the 5×5 convolution. The output Step 4: Calculate the output of the max pooling branch: Y4 = ReLU(W 1×1 * MaxPool(X) + b 1×1 ) Among them, MaxPool is a 3×3 max pooling layer, and W 1×1 is the weight matrix of the first 1×1 convolutional kernel, and b 1×1 is the bias term of the first 1×1 convolution, and the output Step 5: Concatenate all the outputs: Y = Concat(Y1, Y2, Y3, Y4 Among them, the final output (2) BiLSTM network Compared with LSTM, BiLSTM contains two independent LSTM layers, one processes the sequence from front to back (forward LSTM), and the other processes the sequence from back to front (backward LSTM). The hidden state h at each time step t t is the forward LSTM hidden state and the backward LSTM hidden state concatenated or combined. The hidden layer of BiLSTM can be expressed as: Wherein, is the element-wise sum of the forward output component and the backward output component. The BiLSTM network is shown in Figure 3. Suppose we have an input sequence {x1, x2, …, x T}, where T is the length of the sequence. The input sequence is input into the forward LSTM layer in order. The forward LSTM layer processes the input at each time step and generates the hidden state output Each hidden state depends on the current input x t and the hidden state at the previous moment The input sequence is input into the backward LSTM layer in reverse order. The backward LSTM layer processes the input at each time step and generates the hidden state output Each hidden state depends on the current input x t and the hidden state at the next moment At each time step t, the hidden states of the forward and backward LSTM layers are concatenated to obtain the final hidden state h t . Specifically, where the semicolon represents the concatenation operation of vectors or matrices. In this way, the hidden state h t at each time step t contains both the information from the past to the current and the information from the future to the current. (3) Multi-head self-attention mechanism Fusing the multi-head attention mechanism (Multi-Head Self-Attention) in the prediction model enables the model to obtain information from parts of the sequence and learn different types of dependencies in parallel. This mechanism is very effective for processing sequence data. The basic principle of the multi-head self-attention mechanism is as follows: Assume the input is a tensor X of shape [batch_size, sequence_length, input_dim]. batch_size represents the number of samples in a batch; sequence_length represents the length of the sequence, i.e., how many time steps or elements there are in each sample; input_dim represents the feature dimension of each time step (or each element). For each attention head i (i = 1, 2,..., h, where h is the number of attention heads), we need three linear transformation matrices: the query weight matrix the key weight matrix the value weight matrix where d k is the output dimension of a single attention head, usually These matrices are used to generate query, key, and value vectors from the input X. The specific execution steps are as follows: Step 1: Linear transformation: Q i = X · W i Q K i = X·W i K V i = X · W i V Among them, Q i , K i , V i are all tensors of shape [batch_size, sequence_length, d k . Step 2: Calculate the attention scores For each attention head i, calculate its attention scores: Here denotes the transpose of K i , while is the scaling factor used to stabilize the gradient; Step 3: Apply the Softmax function Apply the Softmax function to the attention scores to obtain the attention weights: AttentionWeights i = softmax(Attention Scores i ) where the shape of the attention weights is [batch_size, sequence_length, sequence_length], indicating the degree of influence of each position in the sequence on all other positions; Step 4: Weighted sum, use the attention weights to perform a weighted sum on the value vectors to obtain the output of the i-th attention head: Head i = Attention Weights i ·V i At this time, Head i has a shape of [batch_size, sequence_length, d k ; Step 5: Multi-head merging, concatenate the outputs of all attention heads and then pass them through an additional linear layer W O Project them back to the original feature space: Concatenated Heads=[Head1;Head2...;Head h Multi-Head Output=Concatenated Heads·W O Among them It is used to keep the output of the concatenated multi-head attention consistent with the initial input X in the feature dimension. The final Multi-Head Output will be a tensor of shape [batch_size, sequence_length, input_dim], which combines information from different attention heads. In this way, the multi-head self-attention mechanism can capture complex dependencies in the sequence data and learn information in parallel on different subspaces, thereby improving the model's ability to handle sequence tasks.
8. A method for predicting the power of offshore wind turbines based on Inception - BiLSTM according to claim 1, characterized in that: In step S7, the dataset used by the Inception-BiLSTM-MultiHead-Attention hybrid neural network model in this example comes from the offshore wind power research and test base in Jiangyin Industrial Park, Fuqing City, Fujian Province. In this example, 80% of the preprocessed dataset is used as the training set, 10% as the validation set, and 10% as the test set. To verify the accuracy of the present invention, taking 00:00 on a certain day in the test set as the start time, the prediction effects of different time steps are analyzed. The final prediction results are the power values at 24 hours (96 time steps), 48 hours (192 time steps), and 72 hours (288 time steps). As shown in Figures 6 - 8. The mean absolute error (MAE) and root mean square error (RMSE) are shown in Table 2.
Citation Information
Cited By
Distributed photovoltaic output prediction method and device, electronic equipment and storage medium
CN120601529A
Distributed photovoltaic output prediction method and device, electronic equipment and storage medium
CN120601529B
Quantum-classical machine learning hybrid wind power prediction method
CN120744336A
Enteromorpha green tide marine environment assessment and prediction method based on deep learning
CN120951064A