Wave field rapid prediction method based on Encoder-Only Transform model
Through the combination of Encoder-Only Transformer model and sliding window method, the problem of inefficient calculation in traditional methods is solved, and the rapid and high-precision prediction of wave fields is achieved.
Patent Information
- Application Number
- CN202510297450.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional numerical calculation methods are inefficient in wave field prediction and are difficult to meet practical application requirements, especially when processing high-resolution data or long-time series prediction.
The Encoder-Only Transformer model is adopted, and the structure adjustment and parameter optimization are combined with the sliding window method, and the training time is reduced through the Encoder-Only structure, and the sliding window method is introduced to quickly update historical data.
It has achieved significant improvement in computing efficiency while ensuring prediction accuracy, greatly improved training speed, and met the actual engineering application needs.
Smart Images

Figure CN120387048A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - technical field of machine learning and ocean engineering, and specifically relates to a fast wave field prediction method based on the Encoder - Only Transformer model (improved Transformer model). Background Art
[0002] As a high - dynamic load, the ocean wave field exhibits significant non - linear characteristics and spatio - temporal variability in China's offshore and deep - sea areas. The violent fluctuations of the wave field will generate complex hydrodynamic actions on ocean engineering structures. Especially in extreme sea conditions, the non - linear characteristics of wave loads may trigger transient shock responses of structures, resulting in sudden increases in stress at key parts and aggravated fatigue damage, seriously threatening the stable operation of offshore platforms, floating wind turbines, and offshore oil and gas equipment. Traditional numerical calculation methods have many limitations in wave field prediction. They usually rely on complex physical equations and fine mesh generation, with large computational amounts. Especially when dealing with high - resolution data or long - time - series predictions, the computational efficiency is low and it is difficult to meet the actual application requirements.
[0003] The Transformer model is a deep - learning architecture based on the self - attention mechanism, which initially achieved breakthrough progress in the field of natural language processing. Its core is the self - attention mechanism. By transforming the input vector into a triple of Query (Q), Key (K), and Value (V), the association strength between each position and other positions is calculated. Using the scaled dot - product formula:
[0004]
[0005] where d k is the vector dimension, and the scaling factor avoids gradient vanishing. This mechanism enables the model to efficiently process long - sequence data and capture long - distance dependencies, such as identifying the formation positions of extreme wave groups. In addition, to solve the position - sensitivity problem of time - series data, Transformer innovatively introduces sine - cosine position encoding:
[0006]
[0007] In the formula, pos represents the time - step position index of an element in the sequence; i represents the dimension index of the position - encoding vector; d model represents the hidden - layer dimension of the model; this encoding injects absolute position information into the input vector, ensuring that the model can recognize the time order and spatial arrangement of wave - field data, enabling the Transformer model to support parallel computing and significantly improving the training speed. These characteristics provide new ideas for wave - field prediction.
[0008] The sliding window method is a commonly used method for time series prediction. By dividing the time series into windows of a fixed size and gradually moving the window forward, training and test data can be generated. In wave field prediction, the sliding window method can quickly update the historical data used for model training and effectively handle the non-stationarity and strong randomness of wave data. Summary of the Invention
[0009] To make up for the deficiencies of the prior art, the present invention provides a fast wave field prediction method based on the Encoder-Only Transformer model, aiming to overcome the technical limitations of traditional numerical calculation methods in terms of prediction efficiency and accuracy.
[0010] The specific steps of the fast wave field prediction method based on the Encoder-Only Transformer model are as follows:
[0011] Step 1: Collect and prepare the data set and perform data preprocessing;
[0012] Step 2: Build a Transformer model. The innovation of the present invention lies in the structural adjustment of the traditional Transformer model, changing its original Encoder-Decoder structure to an Encoder-Only structure, which significantly reduces the training time of the model. Specifically, the input data undergoes transformations through a positional encoder, encoder layers, and a linear layer to obtain the output result, where the encoder layers are composed of a multi-head self-attention layer, a feed-forward neural network, layer normalization, and residual connections;
[0013] Step 3: Initialize the model parameters, initialize the parameters of the sliding window method, and determine the time length of the input data, the time length of the output data, and the sliding window step size determined by the sliding window method;
[0014] Step 4: Input the preprocessed training set data and the validation set into the built network model through the sliding window method for model training;
[0015] Step 5: Input the preprocessed test set data into the model through the sliding window method and use the test set to evaluate the prediction ability of the model;
[0016] Step 6: Update the parameters. If the prediction ability of the model is insufficient, re-initialize the Transformer model parameters and the parameters of the sliding window method;
[0017] Step 7: Repeat steps 4 to 6 to optimize the prediction accuracy.
[0018] Furthermore, in step 1, historical wave field time series monitoring data is obtained, and the missing data is supplemented by the cubic spline interpolation method. Generally, the complete data set is divided into a training set, a validation set, and a test set in a ratio of 7:2:1. The divided data is normalized using the standard score normalization method. The specific formula is as follows:
[0019]
[0020] In the formula, x represents the original data, μ represents the mean value of the original data, σ represents the standard deviation of the original data, and x represents the standard deviation of the original data. norm Represents the normalized data.
[0021] Furthermore, in step 2, the Encoder-Only Transformer model includes a positional encoder, an encoder layer, and a linear layer. The encoder layer consists of a multi-head self-attention layer, a feedforward neural network, layer normalization, and a residual connection. The input data is transformed through the positional encoder, encoder layer, and linear layer to obtain the output result. The mean squared error (MSE) is used as the loss function, and Adam is used as the optimizer.
[0022] Furthermore, in the Encoder-Only Transformer model,
[0023] First, the input data is mapped to an embedding space of fixed dimension, converting the original input features into a vector form that the model can process;
[0024] Then, the position information in the sequence is introduced through position encoding, and the position information of each time step is added to the embedding vector, so that the model can distinguish the features of different time steps. The position encoder uses sine and cosine functions to represent the position information:
[0025]
[0026] Where pos is the position of the time step, i is the dimension index, and d model is the embedding dimension;
[0027] Then, the multi-head self-attention layer allows the model to learn information in different representation subspaces in parallel. The multi-head self-attention layer consists of multiple attention heads, each of which learns a different part of the input data. That is, the input data is divided into multiple "heads", each of which calculates the query, key, and value vectors respectively, and then calculates the weighted sum through the scaled dot product attention mechanism:
[0028]
[0029] Where Q, K, and V are query, key, and value matrices, respectively, and dk is the dimension of the key vector. The multi-head self-attention layer concatenates the outputs of multiple heads and integrates them through a linear transformation;
[0030] In addition, each encoder layer also contains a feed-forward neural network that transforms the vectors at each position separately. The feed-forward network usually consists of two linear layers with a ReLU activation function in the middle:
[0031] FFN(x) = max(0, xW1 + b1)W2 + b2,
[0032] where W1 and W2 are weight matrices, and b1 and b2 are bias terms;
[0033] Finally, after being processed by multiple encoder layers, the final output is transformed through two linear layers to obtain the desired output dimension.
[0034] Furthermore, in step 3, the sliding window method is used for time series prediction. Assume the input window at the h-th step is:
[0035]
[0036] The value of X can be predicted according to the following formula (h+n) :
[0037] X (h+n) = f model (X (h) ),
[0038] where f model is the trained prediction model. Then, the old time steps of the window are removed, and X (h+n) is appended to the end of the window, and the above process is repeated.
[0039] Furthermore, in step 3, the input data time length is set to 1 to 2 wave periods, and the output data time length is set to 0.2 to 0.5 wave periods.
[0040] Furthermore, in step 4, the corresponding learning rate, number of training epochs are set, and the trained model is saved locally. The mean squared error is used as the loss function, and the specific calculation formula is as follows:
[0041]
[0042] In the formula, Y i is the true value of the i-th data, is the predicted value of the i-th data; Adam is used as the optimizer during model training to update the model parameters, and the specific update rule is as follows:
[0043] For parameter θ, calculate the first moment mean mt and the second moment variance v t :
[0044] m t = β1m t-1 +(1 - β1)g t
[0045]
[0046] Update the parameters after correcting the bias:
[0047]
[0048] where g t is the current gradient, η is the learning rate, β1 = 0.9, β2 = 0.999 are the default hyperparameters.
[0049] After updating the model parameters, use the validation set data to evaluate the model performance, and save the model parameters that perform best on the validation set.
[0050] Furthermore, in step 5, use the root mean square error evaluation index to verify the model performance on the test set, and set the corresponding preset threshold according to the requirements of the actual project for the prediction accuracy; when the prediction accuracy does not reach the preset threshold, return to the model training stage for parameter iterative optimization until the accuracy requirements of the actual engineering application are met; use the root mean square error RMSE evaluation index to measure the prediction effect of the model, and the specific formula is as follows:
[0051]
[0052] In the formula, Y i is the true value of the i-th data, is the predicted value of the i-th data.
[0053] Compared with the prior art, the present invention has the following advantages:
[0054] Through the technical solution of combining model architecture optimization and prediction mechanism improvement, the present invention significantly improves the calculation efficiency on the premise of ensuring the prediction accuracy; adjusts the structure of the traditional Transformer model, changes it from the original Encoder-Decoder structure to the Encoder-Only structure, greatly improves the training speed, and introduces the sliding window method to update historical data in a timely manner, thereby realizing the rapid prediction of the wave field. Brief Description of the Drawings
[0055] Figure 1 is the flowchart of the wave field prediction method of the present invention;
[0056] Figure 2The structural diagram of the Encoder-Only Transformer model built for the present invention;
[0057] Figure 3 The training and verification effect diagram according to the embodiments of the present invention;
[0058] Figure 4 The test effect diagram according to the embodiments of the present invention;
[0059] Figure 5 The comparison diagram of the test effects of different models in the embodiments of the present invention;
[0060] Figure 6 The comparison diagram of the training time of different models in the embodiments of the present invention. Detailed implementation manners
[0061] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0062] As Figure 1 shown, a fast wave field prediction method based on an Encoder-Only Transformer model, the specific steps are as follows:
[0063] Step 1: Collect and prepare a data set, and perform data preprocessing.
[0064] Preparing the data set includes data preprocessing steps such as handling missing values, outliers and duplicate values. Data visualization and descriptive statistical analysis can be used to deeply understand the characteristics and distribution of the data, as well as possible correlations.
[0065] Obtain historical wave field time series monitoring data, apply the cubic spline interpolation method to complete the missing data. Generally, the complete data set is divided into a training set, a validation set and a test set according to a ratio of 7:2:1, and the divided data is normalized to improve numerical stability. The method selected is standard score normalization, and the specific formula is as follows:
[0066]
[0067] In the formula, x represents the original data, μ represents the average value of the original data, σ is the standard deviation of the original data, and x norm represents the normalized data.
[0068] Step 2: Build a Transformer model: The input data is transformed through a positional encoder, encoder layers, and a linear layer to obtain an output result. The encoder layers consist of multi-head self-attention layers, feed-forward neural networks, layer normalization, and residual connections.
[0069] This application has adjusted the structure of the traditional Transformer model, changing it from the original Encoder-Decoder structure to an Encoder-Only structure, significantly reducing the training time of the model. As Figure 2 shown, the Encoder-Only Transformer model includes a positional encoder, encoder layers, and a linear layer. The encoder layers consist of multi-head self-attention layers, feed-forward neural networks, layer normalization, and residual connections. The input data is transformed through the positional encoder, encoder layers, and linear layer to obtain an output result. The mean squared error (MSE) is used as the loss function, and Adam is used as the optimizer.
[0070] In the Encoder-Only Transformer model, first, the input data is mapped to an embedding space of a fixed dimension, converting the original input features into a vector form that the model can process. For example, for time series data, the input embedding layer can map the feature values of each time step to a high-dimensional space.
[0071] Then, the positional information in the sequence is introduced through positional encoding, adding the positional information of each time step to the embedding vector so that the model can distinguish the features of different time steps. The positional encoder uses sine and cosine functions to represent the positional information:
[0072]
[0073] where pos is the position of the time step, i is the dimension index, and d model is the embedding dimension.
[0074] Subsequently, the multi-head self-attention layer allows the model to learn information in parallel in different representation subspaces. The multi-head self-attention mechanism is the core part of the Transformer. The multi-head self-attention layer consists of multiple attention heads. Each attention head learns different parts of the input data, that is, the input data is divided into multiple "heads", and each head calculates the query (Q), key (K), and value (V) vectors respectively, and then calculates the weighted sum through the scaled dot-product attention mechanism:
[0075]
[0076] where Q, K, and V are the query, key, and value matrices respectively, and d kis the dimension of the key vector. The multi-head self-attention layer concatenates the outputs of multiple heads and integrates them through a linear transformation.
[0077] In addition, each encoder layer also contains a feed-forward neural network (FFN) that transforms the vectors at each position separately. The feed-forward network usually consists of two linear layers with a ReLU activation function in the middle:
[0078] FFN(x) = max(0, xW1 + b1)W2 + b2,
[0079] where W1 and W2 are weight matrices, and b1 and b2 are bias terms.
[0080] Finally, after being processed by multiple encoder layers, the final output is transformed through two linear layers to obtain the desired output dimension. For example, in a time series prediction task, the output linear layer can map the output of the encoder to the dimension of the target value.
[0081] To alleviate the vanishing gradient problem in the training of deep networks, the Transformer model uses residual connections after each sub-layer (multi-head self-attention layer and feed-forward network) and applies layer normalization after the residual connections. This structure helps to accelerate training and improve the stability of the model.
[0082] Step 3: Initialize the model parameters, initialize the parameters of the sliding window method, and determine the time length of the input data, the time length of the output data, and the sliding window step size determined by the sliding window method.
[0083] Determine the model hyperparameters, including the number of encoding layers of the encoder, the number of multi-head attention heads, the embedding dimension of the model, the hidden layer dimension of the feed-forward neural network, etc. Determine the time length of the input data, the time length of the output data, and the sliding window step size determined by the sliding window method. Here, the time length of the input data is set to 1 to 2 wave periods, and the time length of the output data is set to 0.2 to 0.5 wave periods.
[0084] Use the sliding window method to implement time series prediction. Assume that the input window at the h-th step is:
[0085]
[0086] According to the following formula, the value of X (h+n) can be predicted:
[0087] X (h+n) = f model (X (h) )
[0088] Among them, f model is the trained prediction model. Then, remove the old time steps of the window, and append X (h+n) to the end of the window, and repeat the above process.
[0089] Step 4, Model Training: Input the preprocessed training set data and the validation set into the constructed network model through the sliding window method for model training.
[0090] Set the corresponding learning rate and number of training epochs, and save the trained model locally. Save the model that performs best on the validation set based on the mean squared error loss function. Use the mean squared error as the loss function, and the specific calculation formula is as follows:
[0091]
[0092] In the formula, Y i is the true value of the i-th data, is the predicted value of the i-th data; Adam is used as the optimizer during model training to update the model parameters, and the specific update rule is as follows:
[0093] For the parameter θ, calculate the first moment (mean) m t and the second moment (variance) v t :
[0094] m t =β1m t-1 +(1 - β1)g t
[0095]
[0096] Update the parameters after correcting the bias:
[0097]
[0098] Among them, g t is the current gradient, η is the learning rate, β1 = 0.9, β2 = 0.999 are the default hyperparameters.
[0099] After the model parameters are updated, use the validation set data to evaluate the model performance, and save the model parameters that perform best on the validation set.
[0100] Step 5, Evaluate Fitness: Input the preprocessed test set data into the model through the sliding window method, and use the test set to evaluate the prediction ability of the model.
[0101] The root mean square error (RMSE) evaluation metric is used on the test set to verify model performance. Based on the actual project's prediction accuracy requirements, a corresponding preset threshold is set (the preset threshold is set based on the actual project needs). If the prediction accuracy does not reach the preset threshold, the model returns to the training phase for iterative parameter optimization until the accuracy requirements of the actual engineering application are met. The root mean square error (RMSE) evaluation metric is used to measure the model's prediction performance. The specific formula is as follows:
[0102]
[0103] Where Y i is the true value of the i-th data, is the predicted value of the i-th data.
[0104] Step 6. Update parameters: If the model's prediction ability is insufficient, consider factors such as the length of the wave cycle and reinitialize the Transformer model parameters and the sliding window method parameters, such as the number of encoder layers and the sliding window time step.
[0105] Step 7. Repeat steps 4 to 6 to optimize the prediction accuracy.
[0106] In one specific embodiment, the dataset used was derived from measured and derived wave data collected by an ocean wave measurement buoy anchored in Mooloolaba. This data, sourced from the Queensland government, primarily covers wave data from January 1, 2017, to June 30, 2019, and includes multiple wave field characteristics, as detailed in Table 1. The data can be downloaded from the Ali Tianchi website. This paper tested the effectiveness of the Encoder-Only Transformer model in predicting significant wave height.
[0107] Table 1. Wave field dataset
[0108]
[0109] For the training set, 70% of the actual data points are used to train the model. For the validation set, 20% of the data points are used to verify the accuracy of the trained model. For the test set, 10% of the data points are used to test the model and evaluate the accuracy. The prediction results are as follows: Figures 3 - 6 As shown in the figure, the results show that the Encoder-Only Transformer model performs well on the test set, with a root mean square error of 0.11206 on the test set, and the training time is less than half of the original Transformer model.
[0110] In terms of the model architecture, the present invention structurally reconstructs the traditional Transformer model, improving the original Encoder-Decoder dual-branch architecture to an Encoder-Only architecture. Specifically, in implementation, the decoder module in the traditional architecture is removed, and an adjustable fully connected layer is added at the output end of the encoder to achieve data dimension mapping. This technical solution reduces the number of model parameters and improves the training convergence speed at the same time. To prevent information leakage problems in time series prediction, a sliding window control mechanism is introduced to ensure that the model makes predictions based only on historical information by adjusting the time span of the input sequence.
[0111] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A fast prediction method for wave fields based on the Encoder-Only Transformer model, characterized in that, The specific steps are as follows: Step 1: Collect and prepare the dataset and perform data preprocessing; Step 2: Build a Transformer model. The input data undergoes transformations through a positional encoder, encoder layers, and a linear layer to obtain the output result. Among them, the encoder layer consists of a multi-head self-attention layer, a feed-forward neural network, layer normalization, and a residual connection; Step 3: Initialize the model parameters, initialize the parameters of the sliding window method, and determine the time length of the input data, the time length of the output data, and the sliding window step size determined by the sliding window method; Step 4: Input the preprocessed training set data and the validation set into the built network model through the sliding window method for model training; Step 5: Input the preprocessed test set data into the model through the sliding window method and use the test set to evaluate the prediction ability of the model; Step 6: Update the parameters. If the prediction accuracy of the model is insufficient, re-initialize the Transformer model parameters and the parameters of the sliding window method; Step 7: Repeat Steps 4 to 6 to optimize the prediction accuracy.
2. The rapid wave field prediction method based on the Encoder-Only Transformer model according to claim 1, characterized in that In the said Step 1, obtain the historical wave field time series monitoring data, apply the cubic spline interpolation method to complement the missing data. Generally, the complete dataset is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1, and the divided data is normalized. The method selected is standard score normalization, and the specific formula is as follows: Where x represents the original data, μ represents the average value of the original data, σ is the standard deviation of the original data, and x norm represents the data after normalization.
3. A fast wave field prediction method based on the Encoder-Only Transformer model according to claim 1, characterized in that In the said Step 2, the Encoder-Only Transformer model includes a positional encoder, encoder layers, and a linear layer. Among them, the encoder layer consists of a multi-head self-attention layer, a feed-forward neural network, layer normalization, and a residual connection. The input data undergoes transformations through the positional encoder, encoder layers, and linear layer to obtain the output result; And use the mean squared error MSE as the loss function and adopt Adam as the optimizer.
4. The fast wave field prediction method based on the Encoder-Only Transformer model according to claim 3, wherein In the said Encoder-Only Transformer model, First, the input data is mapped to an embedding space of a fixed dimension, converting the original input features into a vector form that the model can process; Then, the positional information in the sequence is introduced through positional encoding, and the positional information of each time step is added to the embedding vector, enabling the model to distinguish the features of different time steps. The positional encoder uses sine and cosine functions to represent the positional information: where pos is the position of the time step, i is the dimension index, and d model is the embedding dimension; Subsequently, the multi-head self-attention layer allows the model to learn information in parallel in different representation subspaces. The multi-head self-attention layer consists of multiple attention heads. Each attention head learns different parts of the input data, that is, the input data is divided into multiple "heads", and each head calculates the query Query, key Key, and value Value vectors respectively, and then calculates the weighted sum through the scaled dot-product attention mechanism: where Q, K, and C are the query, key, and value matrices respectively, and d k is the dimension of the key vector. The multi-head self-attention layer concatenates the outputs of multiple heads and integrates them through a linear transformation; In addition, each encoder layer also contains a feed-forward neural network, which transforms the vectors at each position respectively. The feed-forward network usually consists of two linear layers, with the ReLU activation function used in the middle: FFN(x) = max(0, xW1 + b1)W2 + b2, where, W1 and W2 are weight matrices, and b1 and b2 are bias terms; Finally, after being processed by multiple layers of encoders, the final output is transformed through two linear layers to obtain the required output dimension.
5. A fast wave field prediction method based on the Encoder-Only Transformer model according to claim 1, characterized in that In step 3, the sliding window method is used for time series prediction. Assume the input window at the h-th step is: X can be predicted according to the following formula (h+n) value: X (h+n) = f model (X (h) ) where f model is the trained prediction model, and then the old time steps of the window are removed, and X (h+n) is appended to the end of the window, and the above process is repeated.
6. A fast prediction method for wave fields based on an Encoder-Only Transformer model according to claim 5, characterized in that In step 3, the time length of the input data is generally set to 1 to 2 wave periods, and the time length of the output data is generally set to 0.2 to 0.5 wave periods.
7. A fast wave field prediction method based on the Encoder-Only Transformer model according to claim 5, characterized in that In step 4, set the corresponding learning rate, number of training epochs, and save the trained model locally. Use the mean squared error as the loss function, and the specific calculation formula is as follows: Where Y i is the true value of the i-th data, is the predicted value of the i-th data; Adam is used as the optimizer during model training to update the model parameters, and the specific update rules are as follows: For parameter θ, calculate the first moment mean m t and the second moment variance v t : Update the parameters after correcting the bias: where g t is the current gradient, η is the learning rate, and β1 = 0.9, β2 = 0.999 are default hyperparameters; After the model parameters are updated, use the validation set data to evaluate the model performance, and save the model parameters with the best performance on the validation set.
8. A fast wave field prediction method based on the Encoder-Only Transformer model according to claim 7, characterized in that In step 5, use the root mean squared error evaluation index to verify the model performance on the test set. Set the corresponding preset threshold according to the requirements of the actual project for the prediction accuracy; when the prediction accuracy does not reach the preset threshold, return to the model training stage for parameter iterative optimization until the accuracy requirements of the actual engineering application are met; use the root mean squared error RMSE evaluation index to measure the prediction effect of the model, and the specific formula is as follows: where Y i is the true value of the i-th data, is the predicted value of the i-th data.
Citation Information
Cited By
Multi-scale physical enhancement type irregular sea wave prediction method
CN121350592A
A multiscale physically enhanced irregular sea wave prediction method
CN121350592B