AUV operation motion forecasting method based on HAM-BiTL combination network model

Through the HAM-BiTL combined network model, the multi-step forecasting problem of AUV's multi-position state in complex underwater environments is solved, and flexible and accurate forecasting of AUV manipulation motion is achieved, and the forecasting performance is improved.

CN120257469AActive Publication Date: 2025-07-04JILIN UNIVERSITY

Patent Information

Application Number
CN202510295815.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-04
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The prior art is difficult to make flexible and accurate multi-step motion prediction of the multi-position state of AUV in complex underwater environments, especially under nonlinear and multi-dimensional data conditions, traditional algorithms lack effective forecasting capabilities.

Method used

The HAM-BiTL combined network model is adopted, combining the bidirectional time convolution network, channel attention mechanism, bidirectional long and short-term memory network and multi-head self-attention mechanism, and data processing and model training are carried out through the sliding window method to achieve multi-position state forecast of AUV manipulation motion.

Benefits of technology

Multi-step forecasting of AUV manipulation motion is realized, the accuracy and causality of the forecast are improved, and spatial and temporal characteristics are effectively extracted, which is better than traditional network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257469A_ABST
    Figure CN120257469A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of autonomous underwater vehicles, and provides an AUV operation motion forecasting method based on an HAM-BiTL combined network model. According to the method, bidirectional time convolution, a bidirectional long-short-term memory network and a mixed attention mechanism are combined, and a constructed model is named as an HAM-BiTL combined network model. Through combination of a sliding window method and an improvement form thereof, an AUV manipulation motion multivariate state prediction model is trained and verified. The bidirectional time convolution ensures the causality of network forecasting, and can effectively extract the spatial features of multivariate pose data. The bidirectional long-short-term memory network effectively extracts sequence time features, and front-back association of time sequence data is ensured. The mixed attention mechanism can reasonably distribute weights in sample data, so that multi-step forecasting of the multivariate pose data of the AUV manipulation motion under various working conditions is realized. Experiments of three typical control motion scenes prove that the HAM-BiTL combined network model shows excellent forecasting performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous underwater vehicles, and particularly relates to an AUV operation motion prediction method based on a HAM-BiTL combined network model. Background Technique

[0002] In recent years, the development of AUVs (Autonomous Underwater Vehicles) has become increasingly mature. Various configurations and series of underwater unmanned vehicles have been developed by countries around the world. At the same time, advanced underwater navigation equipment already has the ability of autonomous navigation. AUVs play an important role in underwater operation tasks such as ocean development, underwater monitoring, and environmental exploration. The preset programs of many of these tasks are maneuvering motions. However, due to the difficult-to-survey and complex underwater environment, the situation of AUV loss has occurred during the execution of many underwater tasks.

[0003] The research scope of the motion prediction of marine moving objects covers state prediction, trajectory prediction, etc. In the work of predicting marine objects, classical methods with simple algorithms are often used. They have good applicability for data types with clear distribution laws, simple coupling relationships, low dimensions, and time series that conform to the above characteristics. However, in the face of non-linear data, multi-dimensional data, and data with relatively complex coupling relationships, classical methods often cannot meet the prediction requirements, and more complex models need to be used to solve them.

[0004] Data processing models based on artificial intelligence-related methods are more suitable for the prediction requirements of complex time series. This model is data-driven and can construct a relatively accurate non-linear relationship between the input and output data of an unknown system, and has a certain degree of robustness. The current algorithm has achieved relatively accurate real-time prediction of marine moving objects, but there is less research on the motion prediction of AUVs. There are some problems in existing research. Some of the network models in the research lack the verification of experimental data of various motion types; the predicted pose state dimension is relatively low. However, in practical applications, it is necessary to predict multi-variable pose state data. In addition, many motion predictions are only single-step predictions, lacking the practical application value of advanced prediction.

[0005] Therefore, a combined network model based on deep learning is proposed. This model uses a data-driven method and can more flexibly and accurately utilize the pose state data at the previous moment when the mathematical model of the AUV is unknown, so as to realize the prediction of the pose state data at future times during its maneuvering motion process. Summary of the Invention

[0006] The purpose of the embodiments of the present invention is to provide an AUV operation motion prediction method based on a HAM-BiTL combined network model, aiming to solve the problems proposed in the above background technique.

[0007] The embodiments of the present invention are implemented as follows. An AUV operation motion prediction method based on a HAM-BiTL combined network model includes the following steps:

[0008] Step 1: Model construction;

[0009] Construct a bidirectional temporal convolutional network, a channel attention mechanism, a bidirectional long short-term memory network, and a multi-head self-attention mechanism respectively;

[0010] Step 2: Data preparation and processing;

[0011] Collect typical maneuvering motion simulation experiment data of the AUV under different working conditions to construct a data set for carrying out the training and verification work of the network model, and normalize the data using the maximum normalization method;

[0012] Step 3: Specific implementation of the HAM-BiTL combined network model;

[0013] The proposed HAM-BiTL combined network model is realized by cascading the bidirectional temporal convolutional network, the channel attention mechanism, the bidirectional long short-term memory network, and the multi-head self-attention mechanism in step 1 in sequence; the input multi-dimensional time series consists of samples of a sliding window with a specified sampling width, and the BiTCN is designed to convert the multi-dimensional AUV motion sample block into a one-dimensional feature vector; the temporal-spatial features extracted by the BiTCN are weighted by the channel attention mechanism, and then the BiLSTM is used to capture the forward and backward time features in the features processed by the self-attention mechanism; finally, the multi-head self-attention mechanism adaptively assigns time-varying weights to the spatio-temporal feature data within the current window.

[0014] An AUV operation motion prediction method based on the HAM-BiTL combined network model provided by an embodiment of the present invention. The data set used in this method is composed of the simulation results of the REMUS100 multi-condition maneuvering motions and is independently divided into a training set and a test set. Through the verification of the maneuvering motion prediction experiments carried out under three typical maneuvering motion scenarios: vertical plane trapezoidal motion, horizontal plane Z-shaped motion, and spiral diving motion, it is proved that the HAM-BiTL combined network model has excellent prediction performance. And through ablation experiments, it is verified that the bidirectional temporal convolution in the HAM-BiTL component can effectively ensure the causality of the prediction results and at the same time extract effective spatial features. In addition, the bidirectional long short-term memory network effectively extracts the sequential time features and ensures the front-back relationship of the time series data. The hybrid attention mechanism can effectively allocate the weights in the sample data and retain the key time features and spatial features in the data. Combined with the sliding window method and its improved form, the AUV maneuvering motion multi-state prediction model is effectively trained and verified, and multi-step prediction is realized. The experimental results show that the performance of HAM-BiTL is better than that of the three compared network models, and each component it contains is indispensable and necessary. Description of the Drawings

[0015] Figure 1 Schematic diagram of BiLSTM and basic unit structure of LSTM;

[0016] Figure 2 Schematic diagram of BiTCN and basic structure of TCN;

[0017] Figure 3 Basic principle of sliding window and multi-step prediction;

[0018] Figure 4 Structural diagram of HAM-BiTL combined neural network model;

[0019] Figure 5 Part of the prediction results of the test set for the AUV vertical plane trapezoidal maneuvering motion;

[0020] Figure 6 Another part of the prediction results of the test set for the AUV vertical plane trapezoidal maneuvering motion;

[0021] Figure 7 For Figure 5 Corresponding prediction error;

[0022] Figure 8 For Figure 6 Corresponding prediction error;

[0023] Figure 9 MAE of the prediction results for the vertical plane trapezoidal motion;

[0024] Figure 10Part of the prediction results for the test set of the Z-shaped maneuvering motion of the AUV in the horizontal plane;

[0025] Figure 11 Another part of the prediction results for the test set of the Z-shaped maneuvering motion of the AUV in the horizontal plane;

[0026] Figure 12 For Figure 10 The corresponding prediction error;

[0027] Figure 13 For Figure 11 The corresponding prediction error;

[0028] Figure 14 The MAE of the prediction results for the Z-shaped motion in the horizontal plane;

[0029] Figure 15 The MAE of the prediction results for the spiral diving motion;

[0030] Figure 16 The MAE of the prediction results for the ablation experiment of the Z-shaped motion in the horizontal plane. Specific implementation mode

[0031] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0032] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.

[0033] An AUV operation motion prediction method based on the HAM-BiTL combined network model provided by an embodiment of the present invention includes the following steps:

[0034] Step 1: Model construction;

[0035] Step 1.1: Improvement of the long short-term memory network;

[0036] The long short-term memory network (LSTM, Long Short-Term Memory) has three gate structures, namely the forget gate, the input gate and the output gate;

[0037] The function of the forget gate is to precisely control the information to be forgotten in the cell state, so as to ensure that the cell state C t is properly updated during the information flow process. The forget gate receives h t-1 and x t as input parameters, and calculates the corresponding forget gate parameters through the sigmoid layer. When updating the cell state C tAt this time, the LSTM first generates a candidate value for update through the tanh layer At the same time, an input gate parameter i also needs to be obtained through the sigmoid layer t to determine the information to be updated. Subsequently, multiply i t and to obtain the updated information. At the same time, multiply the previously obtained forget gate f t and the old cell state C t-1 to forget part of the old information. The combination of the two gives the updated cell state C. Finally, the LSTM needs to calculate the final output information, which mainly depends on the cell state C t , but it needs to be filtered by the output gate. Specifically, first normalize the value of the cell state Ct to the interval [-1, 1] through the tanh layer, and then obtain the output gate parameter o through the sigmoid layer t . Finally, multiply o t and the normalized cell state C t to obtain the final filtered result. The detailed calculation processes of the forget gate, input gate, and output gate are shown in Equations (1), (2), and (3).

[0038] f t =σ(W f ·[h t-1 , x t +b f ) (1)

[0039] i t =σ(W i ·[h t-1 , x t +b i )

[0040]

[0041]

[0042] Among them, σ represents the sigmoid function, f t represents the forget gate, W f represents the weight matrix of the forget gate, h t-1 represents the hidden state at the previous time step, x t represents the input at the current time step, b f represents the bias term of the forget gate, i t represents the input gate, W i represents the weight matrix of the input gate, b i represents the bias term of the input gate, represents the candidate cell state, tanh represents the hyperbolic tangent function, Wc denotes the weight matrix for calculating the candidate cell state, b C denotes the bias term of the candidate cell state, C t denotes the cell state, o t denotes the output gate, W o denotes the weight matrix of the output gate, b o denotes the bias term of the output gate, h t denotes the hidden state at the current time step.

[0043] The BiLSTM network adopted combines forward and backward information flows in addition to the basic LSTM structure to extract the temporal features of the sequence. At each time step, the BiLSTM network will calculate and update the forward information flow using equations (1)-(3). The schematic diagram of the BiLSTM and the basic unit structure of the LSTM are as Figure 1 shown.

[0044] Step 1.2: Improvement of the convolutional neural network;

[0045] Compared with the traditional convolutional neural network, the Temporal Convolutional Network (TCN) shows more excellent performance in processing time series data. TCN is a network structure specifically designed for processing time series data. Its unique feature is that it uses convolutional layers to replace recurrent layers to handle sequence dependencies. The core features of TCN include causal convolution and dilated convolution. Causal convolution ensures that when predicting the value at the current moment, only the data at the current moment and before is used, ensuring the causality of the model. Dilated convolution further expands the convolutional layer, enabling the network to capture long-range sequence dependencies without increasing the number of parameters or computational complexity. In order to effectively extract the implicit information in the reverse direction and obtain a larger receptive field, a Bidirectional Temporal Convolutional Network (BiTCN) is constructed by bidirectionally linking TCNs to capture the hidden features in the front and back directions, so as to better obtain the long-term dependence of the sequence, effectively extract the spatial features of multi-dimensional pose data, and ensure the causality of network prediction. The structure of the adopted BiTCN is as Figure 2 shown, where the part within the box line is the basic structure of TCN.

[0046] Step 1.3: Improvement of the attention mechanism;

[0047] The Attention Mechanism has the function of imitating human visual attention, enabling the model to focus on the most important parts of the input data. The Channel Attention Mechanism (CAM) is a mechanism that focuses on the importance assignment of the feature map channels in a convolutional neural network. Its main purpose is to emphasize the channels that contribute most to the task and suppress irrelevant or redundant channels by assigning different weights to each channel, thereby improving the performance of the model. The Hybrid Attention Mechanism effectively assigns weights in the sample data and retains the key spatio-temporal features of the motion data sequence.

[0048] The Multi-Head Self-Attention mechanism is an extended form of the Self-Attention mechanism. It obtains the attention distributions of different subspaces of the input sequence by running multiple independent attention mechanisms in parallel, thereby more comprehensively capturing the associations between various potential features in the sequence. During the operation of the multi-head self-attention mechanism, the input sequence first passes through three different linear transformation layers to obtain Query, Key, and Value respectively; Q represents the query vector related to the task, K represents the key vector, and V represents the value vector, and (K, V) represents the key-value pair; the attention distribution is calculated using the vector K, and the aggregated information is calculated using the vector V. The values of Q, K, and V are calculated by Equation (4); for each attention head, a scaled dot-product attention operation is performed, and the operation process is as shown in Equation (5):

[0049]

[0050] Among them, WQ, WK, W V respectively represent the weight matrices obtained through training for Q, K, and V; X represents the input sequence; d k represents the dimension of the vector, and is used to scale the weights to avoid large dot-product results.

[0051] Step 2: Data preparation and processing;

[0052] Step 2.1: Data preparation;

[0053] Based on the mathematical model of AUV Remus 100, simulation experiments are carried out to obtain a dataset; through vertical plane trapezoidal motion experiments, horizontal plane Z-shaped motion experiments and spiral rotary diving experiments on the AUV under different working conditions, the simulation experiment data of its typical maneuvering motions under these working conditions are collected, and a dataset is constructed with this for the training and verification of subsequent network models; at the same time, for the above three typical maneuvering motions, different sailing speeds are preset and matched with the set rudder commands one by one, and diverse experimental working conditions are formed through combination, so as to obtain different datasets; at the same time, an experimental working condition dataset independent of the training set working conditions is also set as a test set to verify the actual effect of the algorithm. Table 1 is the dataset of the vertical plane trapezoidal maneuvering motion, Table 2 is the dataset of the horizontal plane Z-shaped maneuvering motion, and Table 3 is the dataset of the spiral rotary diving motion.

[0054] Table 1 Dataset of Vertical Plane Trapezoidal Maneuvering Motion

[0055]

[0056]

[0057] Table 2 Dataset of Horizontal Plane Z-shaped Maneuvering Motion

[0058]

[0059] Table 3 Dataset of Spiral Rotary Diving Motion

[0060]

[0061] Step 2.2: Data processing;

[0062] Since the features in the dataset usually have different measurement units and dimensions, if directly used for model training, it may lead to network deviation or distortion. Data preprocessing can perform normalization or standardization on the data, making different features comparable numerically and reducing problems caused by different dimensions. Batch normalization can reduce the dependence of the model on the input data, thereby improving the generalization ability of the model and reducing the risk of overfitting. By observing the characteristics and distribution of the dataset, in the scenarios of horizontal plane Z-shaped motion experiments, vertical plane trapezoidal motion experiments and spiral diving motion experiments, many parameters oscillate around the value of 0, so the maximum value normalization method is adopted here, and its calculation principle is shown in the following formula (6):

[0063]

[0064] where y n,t is the initial data at the t-th time step in the n-th feature; is the data with the largest absolute value in the n-th feature, is the data normalized by the maximum value in the nth feature at the tth time step; through such a normalization method, the data corresponding to all time steps of each feature vector is normalized to the range of [-1, 1], ensuring the original distribution characteristics of the data and eliminating the influence of dimensions at the same time.

[0065] After the preprocessing of the experimental data of the vertical plane trapezoidal motion, the horizontal plane Z-shaped motion and the spiral rotary diving are completed respectively, in accordance with the format specifications of the training multi-step prediction network model, further sorting operations are performed on the data set; with the help of the sliding window technology, the information input into the network model for training or verification is continuously updated according to the established step size; the key parameters of the sliding window are set, including the time step of the training input data and the input step size when updating the input data each time. The sliding window size adopted by this method is the size of the input data dimension and time step for each dimension variable, and the basic principle of the sliding window method is as Figure 3 shown.

[0066] The AUV maneuvering motion prediction of this method uses a multi-input multi-output prediction model. Among them, X in formula (7) t|t-(kim-1) as the input for model training and prediction, is a two-dimensional variable with spatial and time dimensions. u in formula (8) t|t-(kim-1) is an external input variable, which, together with X t|t-(kim-1) , serves as the input data for model training and prediction. The X t|t-(kim-1) and u t|t-(kim-1) data within the window range at time t are used as inputs. After being processed by the trained network model, the state quantity at the next moment can be predicted as shown in the formula.

[0067]

[0068] u t|t-(kim-1) = [u de,ta,t-(kim-1) u delta,t-(kim-2) … u delta,t-(kim-n) … u delta,t T (8)

[0069]

[0070] Among them, X t|t-(kim-1) represents the input data composed of the motion states of a total of kim steps from the current time t to the historical time t-(kim-1). x feature,t-(kim-n) represents the vector of all characteristic motion states at the historical time t-(kim-n). x f,t-(kim-n) represents the value of the fth characteristic motion state at the historical time t-(kim-n). u t|t-(kim-1)A vector representing the external instructions for a total of kim steps from the current time t to the historical time t-(kim-1). u delta,t-(kim-n) Represents the value of the external instruction at the historical time t-(kim-n). Represents the motion data X for a total of kim steps t|t-(kim-1) And the vector u of the external instructions for a total of kim steps t|t-(kim-1) Input into the predictor, and the prediction result at time t+1 is obtained

[0071] According to the steps in equations (7)-(9), the sliding window is successively moved downward according to the current time to realize the training of the network model. The data in the sliding window are all real historical data. According to the calculation steps in equations (7)-(9), the test set data is input into the trained model to realize single-step prediction. Then, in order to realize multi-step prediction, an improved sliding window data iteration method is proposed. The estimated value obtained from equation (9) Is incorporated as new data into the next sliding window range to replace the unknown motion state data at future times, and together with the data from the current time t to the historical time t-(kom-2), it constitutes the input data of the new predictor Thus, the prediction at time t+2 is realized at the current time t. The updated data is shown in equation (10). By alternately executing equations (11) and (12) according to this step, the improved sliding window method is used to realize the dim multi-step prediction of the future t+dim time at time t. In the network training stage, the traditional sliding window method is adopted, while when the predictor is actually put into use, the improved sliding window method is adopted to realize multi-step prediction.

[0072]

[0073] Among them, Is the input state vector for multi-step prediction, Is the m-th estimated value for multi-step prediction, where m∈(1, 2, 3,..., dim).

[0074] Step 3: Specific implementation of the HAM-BiTL combined network model;

[0075] The proposed HAM-BiTL combined network model is cascaded in sequence by the bidirectional temporal convolutional network, channel attention mechanism, bidirectional long short-term memory network, and multi-head self-attention mechanism in step 1, as Figure 4As shown in the figure, the input multi-dimensional time series consists of samples of a sliding window with a specified sampling width. BiTCN is designed to convert a multi-dimensional AUV motion sample block into a one-dimensional feature vector. Since the output at each moment in the TCN network is obtained only from the convolution operation of the input at that moment and its previous moments, it ensures causal constraints when processing sequences. Reasonable weights are assigned to the temporal-spatial features extracted by BiTCN through a channel attention mechanism. Subsequently, BiLSTM is used to capture forward and backward time features from the features processed by the self-attention mechanism. Finally, the multi-head self-attention mechanism adaptively assigns time-varying weights to the spatio-temporal feature data within the current window.

[0076] To systematically and deeply understand the training of the HAM-BiTL combined neural network and the implementation method of multi-step forecasting, refer to the pseudo-code in Algorithm 1 (Table 5), where the key steps and logic are presented in detail. The experimental environment configuration is shown in Table 4. During training, the HAM-BiTL model takes the sampling data of 20 time steps from the current moment to the historical moment as input, and the sampling data of the next moment in the future as output. Since the sampling time is 0.5 s, that is, the sampling data of 10 s from the current moment to the historical moment is used as input, and the data of 0.5 s in the future moment is used as output for network model training.

[0077] Table 4 Experimental environment configuration

[0078]

[0079] Table 5 Algorithm 1

[0080]

[0081] As a preferred embodiment of the present invention, according to the data characteristics of the experiment, the root mean square error (RMSE), mean square error (MSE), and mean absolute error (MAE) are used as evaluation indicators for forecasting performance. The calculation of the RMSE, MSE, and MAE indicators can be achieved through Equations (13)-(15). According to the data characteristics, in order to compare the performance of traditional networks and the HAM-BiTL combined neural network, an evaluation indicator for comparing the forecasting performance of networks, namely the Accuracy Improvement Rate (AIR), is proposed here, and this indicator can be expressed by Equation (16).

[0082]

[0083] Among them, x i represents the true value, Denote the predicted value, and N represents the number of samples in the test set. MAE new represents the mean absolute error of the proposed network, and MAE old represents the mean absolute error of the traditional network. AIR represents the accuracy improvement rate of the proposed network compared to the network used for comparison. When AIR≥0, it means the accuracy of the proposed network has improved; when AIR<0, it means the accuracy has decreased.

[0084] As a preferred embodiment of the present invention, datasets corresponding to the vertical plane trapezoidal maneuvering motion, horizontal plane Z-shaped maneuvering motion, and spiral rotation diving motion in Tables 1 to 3 are used to carry out the network training and verification of HAM-BiTL. At the same time, through the experimental data of the classical method, three typical indicators, RMSE, MSE, and MAE, and the proposed AIR indicator are compared to verify the superiority of the prediction performance of the HAM-BiTL combined network model.

[0085] For the prediction experiment of the AUV's vertical plane trapezoidal maneuvering motion, the horizontal rudder angle δ s , speed U, longitudinal speed u, vertical speed w, pitch rate q, yaw rate r, longitudinal displacement x, vertical displacement z, pitch angle θ, and heading angle ψ are used as input variables. The rudder angle command is an external signal, and usually, the result is calculated in advance according to the AUV's motion control system. Therefore, during the training process of the prediction model, the values of other state variables at future moments except for the horizontal rudder angle δ s are used as the output of the prediction. At the same time, the sliding window method is used to carry out network training, that is, during training, the data from the current moment t to the historical moment t - kim + 1 is used as the input, and the output at the future moment t + 1 is predicted in a single step. Then, during the network model testing process, the improved sliding window method is used to update the single-step prediction result into the input data to achieve multi-step prediction from the future moment t + 1 to t + dim. The specific calculation is realized by equations (7)-(12). Here, kim = 20 and dim = 6 are set, that is, the motion state in the next 3s is predicted in real-time using the data from the current moment to the historical moment for a total of 10s. Figure 5 (where subgraphs a, b, c, and d respectively show the prediction results of speed U, longitudinal speed u, vertical speed w, and pitch rate q) and Figure 6 (where subgraphs a, b, c, d, and e respectively show the prediction results of yaw rate r, longitudinal displacement x, vertical displacement z, pitch angle θ, and heading angle ψ) are the prediction results of the test set for the AUV's vertical plane trapezoidal maneuvering motion. The black curve is the real motion state data at the corresponding moment, and the diamonds of different colors are the first prediction results of the algorithms corresponding to the colored curves. From Figure 5 and Figure 6 it can be seen that the prediction performance of HAM-BiTL is excellent. To more clearly and intuitively observe the prediction performance of the algorithm, Figure 7 andFigure 8 The prediction error curve is given. During the prediction process of the multivariate pose states of the AUV by the HAM-BiTL combined network model, the prediction error is significantly smaller than those of the existing research algorithms such as LSTM, GRU, and CNN-LSTM-Attention. Table 6 shows the evaluation indexes of the prediction results of the trapezoidal motion in the vertical plane.

[0086] Table 6 Evaluation Indexes of the Prediction Results of the Trapezoidal Motion in the Vertical Plane

[0087]

[0088]

[0089] From the results of the RMSE, MSE, MAE, and AIR evaluation indexes in Table 6, it can be seen that the RMSE, MSE, and MAE of HAM-BiTL are all smaller than those of the other three algorithms, and the AIR index relative to each algorithm is greater than 0. According to Equation (16), it can be concluded that the combined network model has a significant improvement in prediction accuracy compared with the traditional algorithms. Figure 9 The MAE values of the prediction results of HAM-BiTL, LSTM, GRU, and CNN-LSTM-Attention for each pose state data are shown. For easy observation, the MAE of the unified pose data is normalized. It can be clearly seen from this radar chart that the MAE of HAM-BiTL is significantly smaller than those of the other three algorithms, which also verifies its excellent prediction performance. The MAE of each dimension is normalized when drawing this chart and the subsequent radar charts for easy observation and comparison.

[0090] For the prediction of the Z-shaped maneuvering motion of the AUV in the horizontal plane, the current and historical vertical rudder angles δ v , speed U, longitudinal speed u, vertical speed w, pitch rate q, yaw rate r, longitudinal displacement x, vertical displacement z, pitch angle θ, and heading angle ψ are used as input quantities. The rudder angle command is an external signal, and usually the result is calculated in advance according to the motion control system of the AUV. Therefore, during the training process of the prediction model, the values of the other state variables at future moments except the vertical rudder angle δ v are used as the output quantities of the prediction. The remaining calculation steps and parameters are the same as those in the previous experiment. Figure 10 (where subfigures a, b, c, and d respectively show the prediction results of the speed U, longitudinal speed u, vertical speed w, and pitch rate q) and Figure 11 (where subfigures a, b, c, d, and e respectively show the prediction results of the yaw rate r, longitudinal displacement x, vertical displacement z, pitch angle θ, and heading angle ψ) are the prediction results of the Z-shaped maneuvering motion of the AUV in the horizontal plane. The black curves in the figure are the real motion state data at the corresponding moments. From Figure 10 and Figure 11It can be seen that the prediction performance of HAM-BiTL is excellent. To more clearly and intuitively observe the prediction performance of the algorithm, Figure 12 and Figure 13 the prediction error curves are given. In the process of predicting the multi-dimensional pose state of the AUV by the HAM-BiTL combined network model, the prediction error is significantly smaller than that of the LSTM, GRU, and CNN-LSTM-Attention algorithms.

[0091] The evaluation indexes of the prediction results of the horizontal plane Z-type motion are shown in Table 7 below.

[0092] Table 7 Evaluation Indexes of the Prediction Results of the Horizontal Plane Z-Type Motion

[0093]

[0094]

[0095] From the results of the RMSE, MSE, MAE, and AIR evaluation indexes in Table 7, it can be seen that the RMSE, MSE, and MAE of HAM-BiTL are all smaller than those of the other three algorithms, and the AIR index relative to each algorithm is greater than 0. According to Equation (16), it can be concluded that the combined network model has a significant improvement in prediction accuracy compared with the traditional algorithms. Figure 14 For the MAE values of the prediction results of HAM-BiTL, LSTM, GRU, and CNN-LSTM-Attention for each pose state data, it can also be clearly seen that the MAE of HAM-BiTL is significantly smaller than that of the other three algorithms, which also confirms its excellent prediction performance.

[0096] For the prediction of the spiral diving motion of the AUV, the current and historical vertical rudder angle δ v , horizontal rudder angle δ s , speed U, longitudinal speed u, lateral speed v, vertical speed w, pitch rate q, yaw rate r, longitudinal displacement x, lateral displacement y, vertical displacement z, pitch angle θ, and heading angle ψ are used as input quantities. Since the rudder angle command is an external signal and its result is usually calculated in advance according to the motion control system of the AUV, the values of the remaining state variables used as inputs except for the vertical rudder angle δ v and the horizontal rudder angle δ s at the future moment are used as the output quantities of the prediction.

[0097] Table 8 shows the evaluation indexes of the prediction results of the spiral diving motion.

[0098] Table 8 Evaluation Indexes of the Prediction Results of the Spiral Diving Motion

[0099]

[0100]

[0101] From the results of the RMSE, MSE, MAE, and AIR evaluation metrics in Table 8, it can be seen that the RMSE, MSE, and MAE of HAM-BiTL are all smaller than those of the other three algorithms, and the AIR metrics for each algorithm are all greater than 0. According to Equation (16), it can be concluded that the combined network model has a significant improvement in prediction accuracy compared to traditional algorithms. Figure 15 The MAE values of the prediction results of HAM-BiTL, LSTM, GRU, and CNN-LSTM-Attention for each pose state data are shown. It can be clearly seen from the radar chart that HAM-BiTL is significantly smaller than the other three algorithms, which also verifies its excellent prediction performance.

[0102] To verify the effectiveness of the HAM-BiTL model and the effectiveness of the relevant parts in the model in improving the network prediction effect, ablation experiments were carried out. Through the experimental results of the horizontal plane Z-shaped experiment under the two components of the BiLSTM and BiTL combined model in HAM-BiTL, the relevant experimental settings are the same as those of the horizontal plane Z-shaped experiment, and four evaluation metrics, namely RMSE, MSE, MAE, and AIR, are used to verify the necessity of the proposed network combination method.

[0103] From Figure 16 it can be seen that the MAE of the prediction results of each pose state of the HAM-BiTL model is smaller than those of the other two algorithms, and the MAE of the prediction results of the BiTL model is smaller than that of the BiLSTM model. The evaluation metrics of the prediction results of the horizontal plane Z-shaped motion ablation experiment are shown in Table 9.

[0104] Table 9 Evaluation Metrics of Prediction Results of Horizontal Plane Z-Shaped Motion Ablation Experiment

[0105]

[0106]

[0107] From the specific values of the four evaluation metrics of RMSE, MSE, MAE, and AIR in Table 9, it can be clearly seen that the prediction performance of the HAM-BiTL model is better than that of BiTL and BiLSTM, and the prediction performance of the BiTL model is better than that of BiLSTM. Therefore, it can be concluded that the bidirectional temporal convolution in the proposed HAM-BiTL can effectively ensure the causality of the prediction results and extract effective spatial features at the same time. The hybrid attention mechanism can effectively allocate the weights in the sample data and retain the key temporal and spatial features in the data.

[0108] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An AUV operation motion prediction method based on the HAM-BiTL combined network model, characterized in that It includes the following steps: Step 1: Model construction; Construct a bidirectional temporal convolutional network, a channel attention mechanism, a bidirectional long short-term memory network, and a multi-head self-attention mechanism respectively; Step 2: Data preparation and processing; Collect the typical maneuvering motion simulation experiment data of the AUV under different working conditions as a data set, and construct the data set with this to carry out the training and verification work of the network model, and use the maximum normalization method to normalize the data; Step 3: Specific implementation of the HAM-BiTL combined network model; The proposed HAM-BiTL combined network model is realized by cascading the bidirectional temporal convolutional network, the channel attention mechanism, the bidirectional long short-term memory network, and the multi-head self-attention mechanism in Step 1 in sequence; the input multi-dimensional time series is composed of samples of a sliding window with a specified sampling width, and the BiTCN is designed to convert the multi-dimensional AUV motion sample block into a one-dimensional feature vector; the temporal-spatial features extracted by the BiTCN are weighted by the channel attention mechanism, and then the BiLSTM is used to capture the forward and backward time features in the features processed by the self-attention mechanism; finally, the multi-head self-attention mechanism adaptively assigns time-varying weights to the spatio-temporal feature data within the current window.

2. The AUV operation motion prediction method based on the HAM-BiTL combined network model according to claim 1, characterized in that, The said Step 1 includes the following specific steps: Step 1.1: Feature extraction of the bidirectional long short-term memory network; The LSTM has three gate structures, namely the forget gate, the input gate, and the output gate. The calculation processes of the forget gate, the input gate, and the output gate are shown in Equations (1), (2), and (3); f t = σ(W f · [h t-1 , x t + b f ) (1) Among them, σ represents the sigmoid function, and f t represents the forget gate, and W f represents the weight matrix of the forget gate, and h t-1 represents the hidden state at the previous time step, and x t represents the input at the current time step, and b f represents the bias term of the forget gate, and i t represents the input gate, and W i represents the weight matrix of the input gate, and b i represents the bias term of the input gate, represents the candidate cell state, tanh represents the hyperbolic tangent function, and W c represents the weight matrix for calculating the candidate cell state, and b C represents the bias term of the candidate cell state, and C t represents the cell state, and o t represents the output gate, and W o represents the weight matrix of the output gate, and b o represents the bias term of the output gate, and h t represents the hidden state at the current time step; The adopted BiLSTM network, in addition to including the basic LSTM structure, combines the forward and backward information flows to extract the time features of the sequence; at each time step, the BiLSTM network will calculate and update the information flow using Equations (1)-(3); Step 1.2: Feature extraction of the convolutional neural network; The BiTCN is obtained through a bidirectionally linked TCN to capture the hidden features in the front and back directions, so as to better obtain the long-term dependence of the sequence, thereby extracting the spatial features of the multi-variable pose data and ensuring the causality of the network prediction; Step 1.3: Feature extraction of the attention mechanism; The channel attention mechanism can effectively capture key information by dynamically weighting multi-variable features, enhance the model's ability to model long-term dependencies and non-linear relationships, and at the same time improve the robustness and prediction accuracy. The multi-head self-attention mechanism is adopted. By running multiple independent attention mechanisms in parallel, the attention distribution of the input sequence in different subspaces is comprehensively captured, so as to fully explore the correlations among various potential features in the sequence. During the operation of the multi-head self-attention mechanism, the input sequence will first pass through three different linear transformation layers to obtain Query, Key, and Value respectively. Q represents the query vector related to the task, K represents the key vector, and V represents the value vector. The key-value pair is represented by (K, V). The attention distribution is calculated using the vector K, and the aggregated information is calculated using the vector V. The values of Q, K, and V are calculated by Equation (4). For each attention head, a scaled dot-product attention operation is performed once, and the operation process is shown in Equation (5): Among them, WQ, W K , W V respectively represent that Q, K, and V are weight matrices obtained through training; X represents the input sequence; d k represents the dimension of the vector, and is used to scale the weights.

3. The AUV operation motion prediction method based on the HAM-BiTL combined network model according to claim 2, wherein Step 2 includes the following steps: Step 2.1: Data preparation; Based on the mathematical model of AUV Remus 100, simulation experiments are carried out to obtain a dataset. By conducting vertical plane trapezoidal motion experiments, horizontal plane Z-shaped motion experiments, and spiral rotary diving experiments on the AUV under different working conditions, the simulation experiment data of its typical maneuvering motions under these working conditions are collected, and a dataset is constructed for the subsequent training and verification of the network model. At the same time, for the above three typical maneuvering motions, different sailing speeds are preset and matched with the set rudder commands one by one. By combining them, diverse experimental working conditions are formed, thus obtaining different datasets. At the same time, an experimental working condition dataset independent of the training set working conditions is also set as a test set to verify the actual effect of the algorithm; Step 2.2: Data processing; After obtaining a dataset that meets the requirements, the experimental data under multiple working conditions are preprocessed, that is, the data is normalized or standardized. Here, the maximum normalization method is adopted, and its calculation principle is shown in the following Equation (6): where y n,t is the initial data at the t-th time step in the n-th feature; is the data with the largest absolute value in the n-th feature, is the data normalized by the maximum value at the t-th time step in the n-th feature; Through such a normalization method, the data corresponding to all time steps of each feature vector are normalized to the range of [-1, 1], ensuring the original distribution characteristics of the data and eliminating the influence of dimensions; After the preprocessing of the experimental data of the vertical plane trapezoidal motion, horizontal plane Z-shaped motion, and spiral rotary diving experiments is completed respectively, the dataset is further processed according to the data format requirements of training the multi-step prediction network model. By using the sliding window method, the information input into the network model for training or verification is continuously updated according to the established step size. The key parameters of the sliding window are set, including the time step of the training input data and the input step size when updating the input data each time.

4. The AUV operation motion prediction method based on the HAM-BiTL combined network model according to claim 3, wherein The AUV maneuvering motion prediction uses a multi-input multi-output prediction model; among them, X in Equation (7) t|t-(kim-1) As the model training and prediction input, it is a two-dimensional variable with spatial and temporal dimensions; u in Equation (8) t|t-(kim-1) is an external input variable, which, together with X t|t-(kim-1) serves as the input data for model training and prediction; X within the time window at time t t|t-(kim-1) and u t|t-(kim-1) data as input, after being processed by the trained network model, can predict the state quantity at the next moment as shown in the equation; Among them, X t|t-(kim-1) represents the input data composed of the motion states in a total of kim steps from the current time t to the historical time t-(kim-1); x feature,t-(kim-n) represents the vector of all characteristic motion states at the historical time t-(kim-n); x f,t-(kim-n) represents the value of the f-th characteristic motion state at the historical time t-(kim-n); u t|t-(kim-1) represents the vector of external instructions in a total of kim steps from the current time t to the historical time t-(kim-1); u delta,t-(kim-n) represents the value of the external instruction at the historical time t-(kim-n); represents the motion data X in a total of kim steps t|t-(kim-1) and the vector u of external instructions in a total of kim steps t|t-(kim-1) input into the predictor to obtain the prediction result at time t+1 n∈(1,2,3,K,kim); Move the sliding window downward according to the steps of equations (7)-(9) at the current moment to implement the training of the network model; then, input the test set data into the trained model according to the calculation steps of equations (7)-(9), and then use the improved sliding window data iteration method for multi-step prediction; first, take the estimated value obtained from equation (9) as new data and incorporate it into the next sliding window range to replace the unknown motion state data at future moments, and jointly form the input data of the new predictor with the data from the current moment t to the historical moment t-(kim-2) so as to achieve the prediction at time t+2 at the current moment t; the updated data is shown in equation (10), and then alternate execution of equations (11) and (12) is carried out according to this step, and the improved sliding window method is used to achieve the dim multi-step prediction at time t for the future t+dim moment; it should be noted that in the network training stage, the traditional sliding window method is used, while when the predictor is actually put into application, the improved sliding window method is used to achieve multi-step prediction; Among them, is the input state vector for multi-step prediction, is the m-th estimated value of multi-step prediction, where m ∈ (1, 2, 3, K, dim).

5. The AUV operation motion prediction method based on the HAM-BiTL combined network model according to claim 4, wherein, The root mean square error RMSE, mean square error MSE, and mean absolute error MAE are used as evaluation indicators for the prediction performance. The calculation of the RMSE, MSE, and MAE indicators is realized by Equations (13)-(15). According to the data characteristics, in order to compare the performance of the traditional network and the HAM-BiTL combined neural network, an evaluation indicator for comparing the prediction performance of the network, that is, the accuracy improvement rate AIR, is proposed as an evaluation indicator, and this indicator is represented by Equation (16); Among them, x i represents the true value, represents the predicted value, N represents the number of samples in the test set; MAE new represents the mean absolute error of the proposed network, MAE old represents the mean absolute error of the traditional network, AIR represents the accuracy improvement rate of the proposed network compared to the network used for comparison. When AIR ≥ 0, it means the accuracy of the proposed network has improved. When AIR < 0, it means the accuracy has decreased.

Citation Information

Patent Citations

  • Manta ray imitating aircraft experimental data prediction method based on adversarial neural network

    CN115081318A

  • Ship control motion forecasting method based on self-attention bidirectional long short-term memory network

    CN117634661A

  • Boundary layer transition identification method and system based on flexible intelligent skin sensing data driving

    CN118428256A

  • Underwater sonar image target detection method

    CN118470513A

  • Underwater glider anomaly detection method based on data sequence

    CN119128376A

Cited By

  • Methods and devices for predicting the water entry motion state of cross-medium watercraft, and cross-medium watercraft.

    CN122571056A