ROV trajectory prediction method based on Transform-LSTM
Through a hybrid deep learning model based on Transformer-LSTM, the trajectory prediction problem caused by communication delay in USV and ROV collaborative operations is solved, the accuracy and stability of ROV trajectory prediction is improved, and the calculation complexity is adapted to complex marine environments, the calculation complexity is reduced, and the real-time needs are met. It is suitable for various USV and ROV collaborative operations scenarios.
Patent Information
- Application Number
- CN202510215677.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-18
AI Technical Summary
In the collaborative operation of USV and ROV, especially in communication delays or failures, it is difficult to accurately predict the ROV trajectory, resulting in an increase in tracking error, affecting the stability and security of the job, and the calculation complexity of deep learning models is limited, which limits its application in marine operations with high real-time requirements.
Using a hybrid deep learning model based on Transformer-LSTM, a model structure including input layer, position coding module, Transformer encoder module, LSTM module and fully connected output layer is designed by generating ROV and USV data sets, to capture long-term dependencies and process dynamic changes in timing, optimize model structure and parameters, and ensure that trajectory prediction is based on historical data when USV observation fails.
It significantly improves the accuracy and stability of ROV trajectory prediction, adapts to complex marine environments, has low computing complexity, meets real-time requirements, and is suitable for various USV and ROV collaborative operation scenarios, with good versatility and scalability.
Smart Images

Figure CN120339495A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of trajectory prediction for unmanned underwater vehicles, especially the trajectory prediction technology in the collaborative operation of unmanned underwater vehicles and unmanned surface vehicles. Specifically, it is a method for predicting the trajectory of an ROV based on Transformer-LSTM. Background Art
[0002] With the rapid development of ocean engineering and ocean resource development, unmanned surface vehicles (USVs) and remotely operated vehicles (ROVs) are increasingly widely used in fields such as ocean exploration, construction and maintenance, and scientific research. As an offshore operation platform, the USV can perform tasks such as navigation, communication, and data processing on the water surface; while the ROV is responsible for specific operations underwater, such as equipment inspection, sample collection, and maintenance operations. The collaborative operation of USV and ROV can significantly improve the efficiency and safety of ocean operations.
[0003] In the prior art, the research on the trajectory prediction of ROV by USV mainly focuses on traditional algorithms and machine learning methods. For example, traditional methods such as Kalman Filter, Extended Kalman Filter (EKF), and Particle Filter model the historical trajectory data of the ROV to predict its future position. However, these methods have limited prediction accuracy when dealing with complex non-linear dynamic systems and long-term dependencies, and it is difficult to adapt to the changing ocean environment. With the development of deep learning, some research has begun to apply neural network models, such as Feedforward Neural Network (FNN), Recurrent Neural Network (RNN), and its variant Long Short-Term Memory (LSTM) for ROV trajectory prediction. These methods have improved the flexibility and accuracy of prediction to a certain extent, but there are still performance bottlenecks in capturing complex dependencies in long time series data and processing high-dimensional input features.
[0004] In recent years, due to its excellent performance in the field of natural language processing and other fields, the Transformer model has been introduced into the processing and prediction of time series data. In the existing technology, the Transformer is combined with traditional time series models to improve the effect of trajectory prediction. However, in practical applications, especially in the collaborative operation of USV and ROV, these methods have not formed a systematic solution, and the impact of communication delay is not considered enough. In the case of communication failure or large delay, it is difficult for the USV to accurately obtain the real-time position information of the ROV, resulting in an increase in tracking error and affecting the stability and safety of the collaborative operation. At the same time, deep learning models usually have a high computational complexity, which limits their application in marine operations with high real-time requirements. Summary of the Invention
[0005] In order to overcome the above problems existing in the prior art, the present invention proposes a method for predicting the ROV trajectory based on Transformer-LSTM.
[0006] The technical solution adopted by the present invention to solve its technical problems is: a method for predicting the ROV trajectory based on Transformer-LSTM, including the following steps: Step 1, data generation: Simulate the movement trajectory of the ROV in three-dimensional space to generate an ROV data set, and generate a corresponding USV data set according to the ROV movement trajectory. The USV data set simulates the observation error of the real environment by adding Gaussian noise; Step 2, data preprocessing: Normalize the data of the ROV data set and the USV data set obtained in Step 1, convert the normalized data into PyTorch tensors, and use a custom sequence creation function to convert the time series data into a fixed-length input sequence and the corresponding prediction target; Step 3, design the model structure for trajectory prediction: The model is a hybrid deep learning model that combines a Transformer encoder and a long short-term memory network; Step 4, model training: Divide the data set obtained in Step 2 into a training set, a validation set, and a test set. Train the hybrid deep learning model in Step 3 through the training set, perform forward propagation, loss calculation, backward propagation, and parameter update; Evaluate the model through the validation set, perform forward propagation and loss calculation to obtain a trajectory prediction model; Step 5, trajectory prediction: Set the start and end time steps of the USV observation failure, use the trajectory prediction model obtained in Step 4 to perform trajectory prediction on the test set. During the prediction process, for the USV observation failure interval, the trajectory prediction model relies on historical observations and historical prediction results for trajectory inference, draw a three-dimensional trajectory comparison graph, and obtain the prediction result; The trajectory prediction model includes an input layer, a position encoding module, a Transformer encoder module, an LSTM module, and a fully-connected output layer. The input layer receives time series data of a fixed length and maps it to the Transformer encoder module through a linear projection layer; the position encoding module adds position information to the input sequence; the Transformer encoder module extracts high-level feature representations of the input sequence layer by layer through a multi-head self-attention mechanism and a feed-forward neural network, capturing the key long-range dependencies and complex feature interactions in the sequence; the LSTM module receives the output feature representation of the Transformer encoder module and further processes the temporal dynamic change information through a recursive structure, capturing the short-term and long-term temporal dependencies in the sequence; the fully-connected output layer maps the final hidden state of the LSTM module to the three-dimensional pose prediction space of the ROV to generate a prediction result.
[0007] For the above-mentioned ROV trajectory prediction method based on Transformer-LSTM, step 1 is specifically as follows: Use a function to generate the position changes of the ROV on the X, Y, and Z axes, simulating its movement in a complex marine environment; calculate the velocity, acceleration, and angular velocity of the ROV through numerical gradients, and expand the three-dimensional position features into a 12-dimensional feature vector to obtain an ROV dataset; the corresponding USV data is simulated with Gaussian noise to obtain the observation error in the real environment, resulting in a USV dataset.
[0008] For the above-mentioned ROV trajectory prediction method based on Transformer-LSTM, step 2 is specifically as follows: Define the sequence length, concatenate the USV features and the ROV pose as the input sequence, and the position of the ROV at the next time step as the prediction target; use PyTorch's TensorDataset and DataLoader to construct training, validation, and test data loaders.
[0009] For the above-mentioned ROV trajectory prediction method based on Transformer-LSTM, the position encoding module generates a vector of the same dimension as the model for each time step and adds it to the input features. Specifically: where and are position encoding vectors; represents the position of the time step; represents the model dimension index, is the dimension of the model, and the shape of the position encoding matrix PE is where is the maximum sequence length; the position encoding vector is added element-wise to the input feature to obtain an input sequence containing position information.
[0010] The above ROV trajectory prediction method based on Transformer-LSTM, the Transformer encoder module consists of a multi-head self-attention mechanism and a feed-forward neural network, and residual connections and layer normalization are applied after each sub-layer; the multi-head attention mechanism calculates multiple self-attention heads in parallel, and the specific process is as follows: For the input sequence X, query Q, key K, and value matrix V are generated through linear transformation: , , , where , , are the weight matrices of the query, query, key, and value respectively, with a shape of , and calculate the self-attention score matrix: The output of the multi-head attention is the concatenation result of all heads, and then through linear transformation: where is the number of heads, is the input weight matrix; The feed-forward neural network in each Transformer encoder layer includes two linear transformations and an activation function: where and are weight matrices, and are bias vectors; Residual connections and layer normalization are applied after each self-attention and feed-forward neural network.
[0011] The above ROV trajectory prediction method based on Transformer-LSTM, the LSTM module consists of multiple stacked LSTM units, the output of each layer of LSTM units is used as the input of the next layer of LSTM units, and a Dropout layer is introduced in the LSTM module.
[0012] The beneficial effects of the present invention are that by combining the advantages of the Transformer model in capturing long-term dependence relationships with the ability of LSTM in processing temporal dynamic changes, an efficient and robust deep learning model is constructed, significantly improving the accuracy and stability of ROV trajectory prediction. The specific advantages include: 1. The Transformer encoder can effectively capture long - distance dependencies in the input sequence, and the LSTM module further processes the temporal dynamic change information. The dual mechanism improves the prediction accuracy.
[0013] 2. In the case of USV observation failure or communication delay, the model can still rely on historical data and prediction results for continuous prediction, ensuring the accurate tracking of the ROV by the USV.
[0014] 3. By optimizing the model structure and parameter settings, while ensuring high accuracy, the model has a low computational complexity, meeting the requirements of ocean operations with high real - time requirements.
[0015] 4. This method is applicable to various complex ocean environments and different types of USV - ROV cooperative operation scenarios, with good generality and scalability. Description of the Drawings
[0016] Figure 1 is a schematic diagram of the process of the present invention; Figure 2 is a data pre - processing flowchart of the present invention; Figure 3 is a structural diagram of the Transformer - LSTM model of the present invention; Figure 4 is a loss curve graph of model training and verification of the present invention; Figure 5 is a comparison graph of the ROV trajectory prediction and the real trajectory of the present invention; Figure 6 is a prediction error curve graph of the present invention. Detailed Embodiment
[0017] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below in conjunction with the drawings and specific embodiments.
[0018] This embodiment proposes a method for predicting the trajectory of a remotely operated underwater vehicle (ROV) based on Transformer - LSTM, aiming to effectively predict the future trajectory of the ROV through a deep - learning model, so as to solve the problem of information transmission delay caused by umbilical cable communication during the cooperative operation of an unmanned surface vehicle (USV), and improve the tracking accuracy and operation efficiency of the USV for the ROV. The following will introduce the technical solution of the present invention in detail, including specific implementation steps such as data generation and pre - processing, model structure and training, prediction and evaluation.
[0019] A method for predicting the trajectory of an ROV based on Transformer - LSTM, as Figure 1 shown, includes the following steps: Step 1: Data generation and pre - processing, the specific process is asFigure 2 as shown The data generation part simulates the motion trajectory of the ROV in three-dimensional space. By defining a time series, the true position data of the ROV is generated, which contains 12-dimensional features such as position, velocity, acceleration, and angular velocity. The specific implementation steps are as follows: Use predefined mathematical functions to generate the position changes of the ROV on the X, Y, and Z axes. By setting the time series and motion parameters (such as initial velocity, acceleration, steering angle, etc.), the actual motion trajectory of the ROV in a complex ocean environment is simulated. This process takes into account various uncertain factors in the ocean environment, such as ocean currents and waves, to ensure that the generated trajectory data can truly reflect the motion state of the ROV during actual operations.
[0020] Based on the generated three-dimensional position data, the velocity, acceleration, and angular velocity of the ROV are calculated through numerical gradients, and the three-dimensional position features are extended to a 12-dimensional feature vector. This feature extension not only enriches the data representation ability but also provides comprehensive motion information for subsequent deep learning models.
[0021] Partition the dataset: The generated ROV dataset is partitioned into a training set, a validation set, and a test set, with proportions of 70%, 15%, and 15% respectively. The corresponding USV data simulates the observation error in the real environment by adding Gaussian noise, further enhancing the robustness of the model. Specifically, Gaussian noise with a mean of 0 and a standard deviation of 0.1 is added to each feature dimension of the USV data, making the USV observation data closer to the noise environment in actual operations.
[0022] To improve the stability and efficiency of model training, the USV and ROV data are normalized. First, the mean and standard deviation of the training set are calculated respectively, and then these parameters are used to standardize the data of the training, validation, and test sets, making the data have zero mean and unit variance. This process is achieved through the following formula: where is the original data, is the mean, is the standard deviation. The normalization process not only eliminates the dimensional differences between different features but also accelerates the convergence speed of the model and improves the training efficiency.
[0023] Convert the normalized data into PyTorch tensors, and use a custom sequence creation function to convert the time series data into a fixed-length input sequence and the corresponding prediction target. The specific operations include: Define the sequence length, concatenate the USV features and the ROV pose as the input sequence, and use the position of the ROV at the next time step as the prediction target. By using the sliding window technique, a large number of training samples are generated to make full use of the temporal information in the time series.
[0024] Finally, use PyTorch's TensorDataset and DataLoader to construct the training, validation, and test data loaders, which facilitate the subsequent training and evaluation of the model. The data loader supports the efficient reading and preprocessing of batch data, accelerating the model training process.
[0025] Step 2: Design the model structure for trajectory prediction: This embodiment proposes a hybrid deep learning model that combines a Transformer encoder and a long short-term memory network (LSTM) to process time series data and perform ROV trajectory prediction. The specific structure is as Figure 3 shown and includes the following modules: 1. Input layer: The input shape of the model is a tensor of , where is the input batch size; is the length of the input sequence. is the dimension of the input features. In this embodiment, the value is 15 (12-dimensional features of the USV + 3-dimensional pose of the ROV = 15 dimensions). The input data is mapped to the model dimension (such as 64 dimensions) through a linear projection layer to meet the input requirements of the Transformer encoder. The weight matrix of the linear projection layer has a shape of, and the bias vector has a shape of 64. is the projection weight matrix, with a shape of , is the bias vector, with a shape of .
[0026] 2. Position encoding module: To enable the model to capture the position information of each time step in the sequence, position encoding is introduced. The position encoding generates a vector of the same dimension as the model for each time step and adds it to the input features. The specific implementation is as follows: The position encoding vector and are respectively defined as: where pos represents the position of the time step; i represents the model dimension index, is the dimension of the model, and the shape of the position encoding matrix PE is , where is the maximum sequence length. The position encoding vector is added element-wise to the input features to obtain the input sequence containing position information.
[0027] 3. Transformer Encoder Module: The Transformer encoder module is mainly composed of a multi-head self-attention mechanism and a feed-forward neural network. Residual connections and layer normalization are applied after each sub-layer. This module can effectively capture long-range dependencies and complex feature interactions in the sequence.
[0028] First, the multi-head attention mechanism improves the model's ability to learn features in different subspaces by calculating multiple self-attention heads in parallel. The specific process is as follows: For the input sequence X, query (Q), key (K), and value (V) matrices are generated through linear transformations: , , , where , , are the weight matrices of the query, query, key, and value respectively, with the shape of . Then, the self-attention score matrix is calculated: The output of the multi-head attention is the concatenation result of all heads, and then through a linear transformation: where is the number of heads, is the input weight matrix.
[0029] The feed-forward neural network in each Transformer encoder layer includes two linear transformations and a activation function: where and are the weight matrices, and are the bias vectors.
[0030] In addition, residual connections and layer normalization are applied after each self-attention and feed-forward neural network: This design helps to alleviate the problem of gradient vanishing and is beneficial to the training of the neural network structure.
[0031] The entire Transformer encoder module is composed of multiple encoder layers with the same structure stacked on top of each other. Each layer contains the above-mentioned multi-head self-attention mechanism and feed-forward neural network. This multi-layer stacking enables the model to extract higher-level feature representations layer by layer, enhancing the ability to model complex temporal relationships.
[0032] 4. LSTM Module After the output of the Transformer encoder module is processed by positional encoding and the multi-head self-attention mechanism, a feature representation with a shape of is obtained. This feature representation is then input into the LSTM module to further capture the temporal dynamic change information in the sequence.
[0033] LSTM is a variant of the recurrent neural network that can effectively capture long-term dependencies. The LSTM cell includes the following key parts: Forget Gate: ; Input Gate: , Update Gate: ; Output Gate: , .
[0034] Among them, is the sigmoid activation function; is the input at the current time step; is the hidden state at the current time step; is the cell state at the current time step; , , , is the input weight matrix; , , , is the hidden state weight matrix; , , , is the bias vector; represents element-wise multiplication.
[0035] To enhance the model's ability to model temporal dynamic changes, the LSTM module consists of multiple stacked LSTM cells. The output of each layer of LSTM serves as the input to the next layer of LSTM, enabling the model to extract more complex and abstract temporal feature representations layer by layer.
[0036] In addition, introducing a Dropout layer in the LSTM module helps prevent overfitting and improve the generalization ability of the model. Dropout reduces the model's over-reliance on the training data by randomly discarding the outputs of some neurons.
[0037] 5. Fully Connected Output Layer The output of the LSTM module is a feature vector with a shape of , which represents the comprehensive temporal feature representation of the input sequence. To map these features to the 3D pose prediction space of the ROV, the model introduces a fully connected layer: where is the weight matrix of the fully connected layer, with a shape of ; is the bias vector, with a shape of 3; is the final hidden state of the LSTM module, and the output y of the fully connected layer is the predicted value of the 3D pose of the ROV at the next time step.
[0038] Step 3: Connecting Transformer Feature Extraction and LSTM The model realizes the effective connection between Transformer feature extraction and LSTM through the following steps: 1. The input layer receives time series data of a fixed length and maps it to the dimension of the Transformer model through a linear projection layer.
[0039] 2. The position encoding layer adds position information to the input sequence to ensure that the model can recognize the position relationships of each time step in the sequence.
[0040] 3. The Transformer encoder module extracts high-level feature representations of the input sequence layer by layer through the multi-head self-attention mechanism and the feed-forward neural network, capturing the long-range dependencies and complex feature interactions in the sequence.
[0041] 4. The output feature representation of the Transformer encoder is passed to the LSTM module, and the LSTM further processes the temporal dynamic change information through its recursive structure, capturing the short-term and long-term temporal dependencies in the sequence.
[0042] 5. The final hidden state of the LSTM module is mapped to the 3D pose prediction space of the ROV through a fully connected layer to generate the prediction result.
[0043] This connection method fully utilizes the advantages of Transformer in feature extraction and dependency modeling, as well as the ability of LSTM in capturing temporal dynamic changes, forming an efficient and robust hybrid deep learning model.
[0044] It includes an input projection layer and multiple Transformer encoders for extracting high-level features of the input sequence.
[0045] To ensure good performance of the model during training, various hyperparameters are reasonably set, including the input dimension, hidden layer size, number of attention heads, learning rate, etc. The specific parameter settings are as follows: Among them, is the feature dimension of the input sequence (12-dimensional features of USV plus 3-dimensional pose of ROV), is the model dimension of the Transformer encoder, is the number of heads of the multi-head attention mechanism, is the number of Transformer encoder layers, is the hidden layer size of LSTM, is the number of LSTM layers, is the dimension of the predicted output (3-dimensional pose of ROV), is the learning rate, is the number of training epochs, is the weight decay parameter.
[0046] Step 3: Model training and prediction steps: The model training process includes two stages: training and validation. During the training stage, the model is in the training mode, performing forward propagation, loss calculation, backpropagation, and parameter update; during the validation stage, the model is in the evaluation mode, only performing forward propagation and loss calculation, without parameter update.
[0047] During the training process, the mean squared error (MSE) is used as the loss function to measure the gap between the predicted position and the true position.
[0048] The calculation formula of the mean squared error is as follows: Among them, is the predicted value, is the true value, and N is the number of samples. The reason for choosing MSE as the loss function is that it punishes larger errors more severely, which helps the model to fit the data more accurately.
[0049] To optimize the model parameters, optimizer is used. is a variant based on the Adam optimization algorithm, which further prevents model overfitting by introducing a weight decay mechanism. The configuration of the optimizer is as follows: Learning rate: 0.0005; Weight decay parameter: 1e-4.
[0050] During the training process, to prevent the problem of gradient explosion, a gradient clipping technique is introduced to clip the gradients calculated during the backpropagation process, ensuring that the gradient values are within a reasonable range and guaranteeing the stability of the training process.
[0051] The training loop consists of multiple rounds of iterations. In each training round, the model is in the training mode and performs forward propagation, loss calculation, backpropagation, and parameter update. The specific steps are as follows: 1. Forward propagation: Input the input sequence of the observed data of USV in the training set into the model to calculate the prediction results.
[0052] 2. Loss calculation: Calculate the error between the predicted ROV trajectory results and the true labels through the MSE loss function.
[0053] 3. Backpropagation: Calculate the gradients of the model parameters according to the gradients of the loss function.
[0054] 4. Parameter update: Utilize the optimizer to update the model parameters according to the calculated gradients.
[0055] 5. Gradient clipping: During the backpropagation process, clip the gradients to prevent gradient explosion.
[0056] 6. Loss recording: Record the training losses of each batch and calculate the average training loss of the entire training set.
[0057] After each training round, the model switches to the validation mode and only performs forward propagation and loss calculation without parameter update. By calculating the MSE loss on the validation set, evaluate the performance of the model on unseen data. If the current validation loss is better than the previous best validation loss, save the current model parameters to ensure that the model with the best performance is selected for subsequent prediction and application.
[0058] The entire training process iterates through multiple rounds to gradually optimize the model parameters, enabling the model to exhibit good fitting ability and generalization ability on both the training set and the validation set. During the training process, the current training and validation losses are regularly output to help monitor the training progress of the model and prevent overfitting.
[0059] After training, by plotting the training and validation loss curves, visually display the convergence situation and generalization ability of the model during the training process. As Figure 4 shown, the horizontal axis in the figure represents the training rounds, the vertical axis represents the loss value, the blue solid line represents the training loss, and the orange dashed line represents the validation loss. The loss curve can help users understand the learning effect of the model and judge whether it is necessary to further adjust the model structure or training parameters.
[0060] Step 4: Prediction steps in the case of simulated USV observation failure: To evaluate the prediction ability of the model in the case of USV observation failure, the present invention designs a method for simulating USV observation failure. The specific steps are as follows: First, set the start and end time steps of USV observation failure, for example, from time step 100 to 149. During this time period, the USV data is simulated as a failure state, that is, the corresponding USV eigenvalue is set to zero, simulating the missing observation data caused by communication problems in actual operations. In this way, the model cannot rely on the real-time data of the USV for prediction during this time period and needs to rely on previous observations and prediction results for trajectory inference.
[0061] Use the trained Transformer-LSTM model to predict the ROV trajectory for the test set. During the prediction process, for the USV observation failure interval, the model needs to be able to maintain continuous prediction of the ROV trajectory in the case of missing some USV data. The specific implementation steps include: For each test sequence, input the current USV and ROV data, and the model outputs the predicted position of the ROV. During the USV observation failure, the USV data is set to zero, and the model relies on historical observations and previous prediction results for trajectory inference. By gradually updating the input sequence, the model can continuously perform trajectory prediction during the failure interval to ensure the continuity and stability of the prediction results.
[0062] Step Five: Visualization of Prediction Results: By drawing a three-dimensional trajectory comparison graph, visually display the differences between the real trajectory and the predicted trajectory. As Figure 5 shown, the real trajectory is represented by a blue solid line, and the predicted trajectory is represented by an orange dashed line, which is convenient to observe the degree of coincidence between the two. In addition, within the time steps of observation failure, the specific positions of USV observation failure are marked with red crosses to highlight the prediction performance of the model during the failure. This visualization method can directly reflect the prediction effect of the model in different scenarios, especially the robustness and accuracy in the case of USV observation failure.
[0063] To quantify the prediction performance, calculate the error metrics between the prediction results and the real trajectory, including Mean Norm, Root Mean Square Error (RMSE), and Mean Absolute Error (MAE). These metrics can comprehensively reflect the prediction accuracy and stability of the model and help evaluate the performance of the model at different time steps and under different observation conditions.
[0064] By drawing the prediction error curve for each time step, further analyze the performance of the model at different time steps, especially the changing trend of the prediction error during the USV observation failure. As Figure 6As shown, the error curve represents the overall error level of the model with an orange solid line, and within the time steps of observed failures, the errors at specific positions are marked with red crosses for easy observation of the prediction stability and accuracy of the model at critical moments.
[0065] The above embodiments are only exemplary embodiments of the present invention and are not used to limit the present invention. Those skilled in the art can make various modifications or equivalent replacements to the present invention within the essence and protection scope of the present invention, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present invention.
Claims
1. A method for predicting the trajectory of an ROV based on Transformer-LSTM, characterized in that, It includes the following steps: Step 1, data generation: Simulate the movement trajectory of the ROV in three-dimensional space to generate an ROV data set, generate a corresponding USV data set according to the ROV movement trajectory, and add Gaussian noise to the USV data set to simulate the observation error in the real environment; Step 2, data preprocessing: Normalize the data of the ROV data set and the USV data set obtained in Step 1, convert the normalized data into PyTorch tensors, and use a custom sequence creation function to convert the time series data into a fixed-length input sequence and the corresponding prediction target; Step 3, design the model structure for trajectory prediction: The model is a hybrid deep learning model that combines a Transformer encoder and a long short-term memory network; Step 4, model training: Divide the data set obtained in Step 2 into a training set, a validation set, and a test set. Train the hybrid deep learning model in Step 3 through the training set, perform forward propagation, loss calculation, backward propagation, and parameter update; evaluate the model through the validation set, perform forward propagation and loss calculation to obtain a trajectory prediction model; Step 5, trajectory prediction: Set the start and end time steps of USV observation failure, use the trajectory prediction model obtained in Step 4 to perform trajectory prediction on the test set. During the prediction process, for the USV observation failure interval, the trajectory prediction model relies on historical observations and historical prediction results to infer the trajectory, draw a three-dimensional trajectory comparison graph, and obtain the prediction result; The trajectory prediction model includes an input layer, a position encoding module, a Transformer encoder module, an LSTM module, and a fully connected output layer. The input layer receives fixed-length time series data and maps it to the Transformer encoder module through a linear projection layer; the position encoding module adds position information to the input sequence; the Transformer encoder module extracts high-level feature representations of the input sequence layer by layer through a multi-head self-attention mechanism and a feed-forward neural network, capturing the key long-distance dependencies and complex feature interactions in the sequence; the LSTM module receives the output feature representation of the Transformer encoder module and further processes the temporal dynamic change information through a recursive structure, capturing the short-term and long-term temporal dependencies in the sequence; the fully connected output layer maps the final hidden state of the LSTM module to the three-dimensional pose prediction space of the ROV to generate a prediction result.
2. The ROV trajectory prediction method based on Transformer-LSTM according to claim 1, wherein, Specifically, Step 1 is as follows: Use a function to generate the position changes of the ROV on the X, Y, and Z axes to simulate its movement in a complex marine environment; calculate the velocity, acceleration, and angular velocity of the ROV through numerical gradients, and expand the three-dimensional position features into 12-dimensional feature vectors to obtain the ROV data set; the corresponding USV data is added with Gaussian noise to simulate the observation error in the real environment to obtain the USV data set.
3. The ROV trajectory prediction method based on Transformer-LSTM according to claim 1, wherein The specific steps of step 2 are as follows: Define the sequence length, concatenate the USV features and the ROV pose as the input sequence, and use the position of the ROV at the next time step as the prediction target; Use TensorDataset and DataLoader in PyTorch to construct the training, validation, and test data loaders.
4. A ROV trajectory prediction method based on Transformer-LSTM according to claim 1, characterized in that The position encoding module generates a vector with the same dimension as the model for each time step and adds it to the input features. Specifically: Among them, and are position encoding vectors; pos represents the position of the time step; represents the model dimension index, is the dimension of the model, and the shape of the position encoding matrix PE is , where is the maximum sequence length; the position encoding vector is added element-wise to the input features to obtain an input sequence containing position information.
5. A ROV trajectory prediction method based on Transformer-LSTM according to claim 1, characterized in that The Transformer encoder module consists of a multi-head self-attention mechanism and a feed-forward neural network. Residual connections and layer normalization are applied after each sub-layer; The multi-head attention mechanism calculates multiple self-attention heads in parallel. The specific process is as follows: For the input sequence X, generate the query Q, key K, and value matrix V through linear transformation: , , , where , , are the weight matrices of the query, query, key, and value respectively, with shapes of , and calculate the self-attention score matrix: The output of the multi-head attention is the concatenation result of all heads, and then it goes through a linear transformation: Among them, is the number of heads, is the input weight matrix; The feed-forward neural network in each Transformer encoder layer consists of two linear transformations and an activation function: Among them, and are weight matrices, and are bias vectors; Residual connections and layer normalization are applied after each self-attention and feed-forward neural network.
6. The ROV trajectory prediction method based on Transformer-LSTM according to claim 1, wherein The LSTM module consists of multiple stacked LSTM cells. The output of each layer of LSTM cells is used as the input of the next layer of LSTM cells, and a Dropout layer is introduced in the LSTM module.
Citation Information
Cited By
Continuous learning-guided soft measurement method and system for hydrogen production amount in hydrogen production process of electrolytic cell
CN122241159A