Underwater robot target identification and tracking evaluation method based on deep learning

By using deep learning technology and combining multimodal data processing from vision, acoustic and flow field sensors, the problems of optical interference and flow field influence in underwater robot target recognition and tracking were solved, achieving high-precision and robust target recognition and tracking.

CN120949244APending Publication Date: 2025-11-14GUANGZHOU MARITIME INST
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510964133.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Underwater robot target recognition and tracking faces challenges such as optical image scattering interference, limited sonar resolution, and the influence of flow field dynamics. Traditional methods are difficult to effectively fuse multimodal data and dynamically adapt to changes in the flow field, resulting in insufficient recognition and tracking accuracy and robustness.

Method used

A deep learning-based approach is adopted, which utilizes multimodal data collaborative processing and dynamic optimization mechanisms to collect data using visual, acoustic and flow field sensors, performs spatiotemporal alignment and feature fusion, constructs a trajectory prediction model and embeds fluid dynamic constraints to form a closed-loop adaptive system.

Benefits of technology

It improves the accuracy and robustness of underwater target identification and tracking, and can maintain stable performance in complex flow field environments to meet practical application requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949244A_ABST
    Figure CN120949244A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater robot target identification and tracking evaluation method based on deep learning, and the method comprises the steps: obtaining a target optical image, a sonar echo signal and a flow field parameter, constructing multi-modal data, and carrying out the preprocessing and space-time alignment; target apparent features and flow field features are extracted, dynamic weight fusion is carried out, and a target identification result is obtained through classifier processing; constructing a trajectory prediction model based on the target identification result and the flow field dynamic characteristics, and obtaining an initial trajectory prediction result by combining historical trajectory data optimization; constructing an implicit flow field model, and performing joint training with the trajectory prediction model to obtain an optimized trajectory prediction model; and a trajectory prediction result output by the optimized model is obtained, deviation is calculated in combination with actual trajectory data, the physical constraint weight and the target identification parameters in the trajectory prediction model are dynamically adjusted according to the deviation, and high precision and strong robustness of underwater target identification tracking of the underwater robot can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater robot technology, specifically to a deep learning-based method for underwater robot target recognition, tracking, and evaluation. Background Technology

[0002] In the field of underwater robotics, target recognition and tracking face multiple challenges posed by complex environments. The unique characteristics of underwater water make optical images susceptible to scattering, refraction, and interference from suspended objects, significantly degrading image quality. While sonar echo signals can penetrate water, their resolution is limited, and target feature extraction is difficult. Simultaneously, flow field dynamic parameters significantly influence target trajectories, making it difficult for traditional methods to effectively fuse multimodal data and dynamically adapt to changes in the flow field. Existing technologies suffer from insufficient anti-interference capabilities in single-sensor target recognition schemes, a lack of spatiotemporal alignment mechanisms and dynamic weight allocation strategies for multimodal data fusion, and trajectory prediction models often exhibit prediction biases in complex flow fields due to the absence of embedded fluid physics constraints. Furthermore, the lack of adaptive optimization mechanisms based on real-time biases makes it difficult to meet the accuracy and robustness requirements of underwater target recognition and tracking for practical applications. Summary of the Invention

[0003] In order to solve the above-mentioned technical problems, the present invention provides a deep learning-based method for underwater robot target recognition, tracking and evaluation.

[0004] The technical solution of this invention is implemented as follows: a deep learning-based underwater robot target recognition and tracking evaluation method, comprising:

[0005] S1. Acquire underwater target optical images, sonar echo signals, and flow field parameters. Construct multimodal data based on the acquired underwater target optical images, sonar echo signals, and flow field parameters. Perform preprocessing and spatiotemporal alignment on the multimodal data to obtain preprocessed multimodal data.

[0006] S2. For the preprocessed multimodal data, extract the target appearance features and flow field features respectively, perform dynamic weight fusion on the appearance features and flow field features, and classify the dynamic weight fusion result through a classifier to obtain the target recognition result.

[0007] S3. Based on the target identification results and flow field characteristics, construct a trajectory prediction model, and optimize the trajectory prediction model by combining the target's historical trajectory data to obtain the initial trajectory prediction result;

[0008] S4. Based on the initial trajectory prediction results, construct an implicit flow field model, and jointly train the output results of the implicit flow field model with the trajectory prediction model to obtain an optimized trajectory prediction model.

[0009] S5. Obtain the target motion trajectory prediction result output by the optimized trajectory prediction model, calculate the deviation by combining the actual target motion trajectory data, and dynamically adjust the physical constraint weights and target recognition parameters in the trajectory prediction model according to the deviation.

[0010] Further, in step S1, underwater target optical images, sonar echo signals and flow field parameters are acquired, multimodal data is constructed based on the acquired underwater target optical images, sonar echo signals and flow field parameters, and the multimodal data is preprocessed and spatiotemporally aligned to obtain preprocessed multimodal data;

[0011] Furthermore, in step S1, the specific steps are as follows:

[0012] Visual sensors, acoustic sensors, and flow field sensors are deployed to collect optical image data, sonar echo signals, and flow field physical parameters of the target, forming multimodal raw data;

[0013] For the original multimodal data, water scattering correction and noise filtering are performed on the optical image, range-azimuth features are extracted from the sonar data, and Kalman filtering and cubic spline interpolation are performed on the flow field parameters to obtain the processed multimodal data.

[0014] The processed multimodal data is time-aligned using timestamp synchronization technology to obtain preprocessed multimodal data.

[0015] Further, in step S2, for the preprocessed multimodal data, target appearance features and flow field features are extracted respectively, the appearance features and flow field features are dynamically weighted and fused, and the dynamic weighted fusion result is classified by a classifier to obtain the target recognition result;

[0016] Furthermore, in step S2, the specific steps are as follows:

[0017] For the preprocessed multimodal data, a first network branch including a visual branch and an acoustic branch, and a second network branch including a flow field branch are constructed to extract the appearance features and flow field features of the target, respectively.

[0018] The apparent features and flow field features are correlated through a cross-modal attention mechanism to form fused features;

[0019] The fusion features are dynamically weighted according to the real-time flow field intensity to form adaptive fusion features;

[0020] The adaptive fusion features are processed by a classifier to output the target category probability distribution, thus forming the target recognition result.

[0021] Furthermore, in step S3, based on the target identification result and flow field characteristics, a trajectory prediction model is constructed, and the trajectory prediction model is optimized by combining the target's historical trajectory data to obtain an initial trajectory prediction result;

[0022] Furthermore, in step S3, the specific steps are as follows:

[0023] Based on the target recognition results and the flow field features extracted from the preprocessed multimodal data, a trajectory prediction model is constructed using a long short-term memory network. The model takes historical target trajectory data, the identified target features, and the flow field features as inputs and outputs preliminary target trajectory prediction data.

[0024] The preliminary target trajectory prediction data is compared with the actual target motion trajectory data to calculate the trajectory deviation;

[0025] Based on the trajectory deviation, fluid dynamics equation constraints are embedded into the state update equation of the long short-term memory network to form a physically constrained trajectory prediction model.

[0026] The physical constraint trajectory prediction model is optimized and trained to generate initial trajectory prediction results.

[0027] Further, in step S4, an implicit flow field model is constructed based on the initial trajectory prediction result, and the output result of the implicit flow field model is jointly trained with the trajectory prediction model to obtain an optimized trajectory prediction model.

[0028] Furthermore, in step S4, the specific steps are as follows:

[0029] Based on the initial trajectory prediction results, a physical sensing neural network is constructed. The flow field measurement data and fluid dynamics equation constraints are input, and an implicit flow field model is constructed for the unmeasurable flow field region to generate a flow field distribution characteristic representation.

[0030] The flow field distribution characteristics output by the implicit flow field model are correlated with the trajectory prediction model to establish a mapping relationship between flow field characteristics and target motion trajectory.

[0031] Based on the flow field parameters output by the implicit flow field model and the trajectory prediction results of the trajectory prediction model, a joint training objective function including the flow field equation residuals and trajectory fitting errors is designed.

[0032] The physical perception neural network and the trajectory prediction model are jointly trained end-to-end. By optimizing the joint training objective function, an optimized trajectory prediction model that incorporates the physical constraints of the flow field is obtained.

[0033] Further, in step S5, the target motion trajectory prediction result output by the optimized trajectory prediction model is obtained, the deviation is calculated by combining the actual target motion trajectory data, and the physical constraint weights and target recognition parameters in the trajectory prediction model are dynamically adjusted according to the deviation.

[0034] Furthermore, in step S5, the specific steps are as follows:

[0035] The target motion trajectory prediction result output by the optimized trajectory prediction model is obtained, and the actual motion trajectory data of the target is collected. The target motion trajectory prediction result and the actual motion trajectory data are spatiotemporally aligned. The position deviation, velocity deviation, trajectory curvature deviation and recognition error rate are calculated to form a comprehensive deviation.

[0036] Based on the specific values ​​and distribution characteristics of the comprehensive deviation, the adjustment rules and quantification parameters for the fluid dynamics equation constraint weights and target identification parameters are determined.

[0037] For the trajectory prediction model and target recognition parameters, the physical constraint weights and recognition parameters in the model are adaptively updated according to the adjustment rules.

[0038] The underwater robot target recognition and tracking evaluation method based on deep learning described in this invention has the following advantages:

[0039] This invention discloses a deep learning-based method for underwater robot target recognition and tracking evaluation. This method effectively addresses problems such as optical image scattering interference, limited sonar resolution, and the influence of flow field dynamics in underwater environments through multimodal data collaborative processing and dynamic optimization mechanisms. It improves data consistency by utilizing multi-source data acquisition and spatiotemporal alignment preprocessing from visual, acoustic, and flow field sensors. Through cross-modal attention and dynamic weight fusion, it achieves an adaptive combination of apparent features and flow field features, overcoming the shortcomings of traditional methods in terms of insufficient anti-interference capability. By embedding fluid dynamic constraints into the trajectory prediction model and constructing an implicit flow field model, it solves the problem of trajectory prediction deviation in complex flow fields. Based on a real-time deviation-based physical constraint weight and dynamic adjustment mechanism for recognition model parameters, a closed-loop adaptive system from "prediction-feedback-optimization" is formed. This method can effectively improve the high accuracy and robustness of underwater robots in underwater target recognition and tracking, maintain stable performance in complex flow field environments, and provide reliable technical support for the practical engineering applications of underwater robots. Attached Figure Description

[0040] Figure 1 This is a structural block diagram of a deep learning-based underwater robot target recognition and tracking evaluation method in an embodiment of the present invention;

[0041] Figure 2This is a flowchart illustrating the steps of a deep learning-based underwater robot target recognition and tracking evaluation method in an embodiment of the present invention. Detailed Implementation

[0042] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0043] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0044] Please see Figures 1-2 As shown in the embodiment of the present invention, a deep learning-based underwater robot target recognition and evaluation method includes:

[0045] S1. Acquire underwater target optical images, sonar echo signals, and flow field parameters. Construct multimodal data based on the acquired underwater target optical images, sonar echo signals, and flow field parameters. Perform preprocessing and spatiotemporal alignment on the multimodal data to obtain preprocessed multimodal data.

[0046] S2. For the preprocessed multimodal data, extract the target appearance features and flow field features respectively, perform dynamic weight fusion on the appearance features and flow field features, and classify the dynamic weight fusion result through a classifier to obtain the target recognition result.

[0047] S3. Based on the target identification results and flow field characteristics, construct a trajectory prediction model, and optimize the trajectory prediction model by combining the target's historical trajectory data to obtain the initial trajectory prediction result;

[0048] S4. Based on the initial trajectory prediction results, construct an implicit flow field model, and jointly train the output results of the implicit flow field model with the trajectory prediction model to obtain an optimized trajectory prediction model.

[0049] S5. Obtain the target motion trajectory prediction result output by the optimized trajectory prediction model, calculate the deviation by combining the actual target motion trajectory data, and dynamically adjust the physical constraint weights and target recognition parameters in the trajectory prediction model according to the deviation.

[0050] like Figure 2 As shown, in step S1, underwater target optical images, sonar echo signals, and flow field parameters are acquired. Multimodal data is constructed based on the acquired underwater target optical images, sonar echo signals, and flow field parameters. The multimodal data is preprocessed and spatiotemporally aligned to obtain preprocessed multimodal data.

[0051] Specifically, in this embodiment, a visual sensor, an acoustic sensor, and a flow field sensor are deployed to collect optical image data, sonar echo signals, and flow field physical parameters of the target, forming multimodal raw data.

[0052] Specifically, visual sensors, acoustic sensors, and flow field sensors are deployed for the underwater robot. The visual sensor uses an underwater high-definition camera with a sampling frequency of 30 frames per second and a resolution of 1920×1080 pixels. The acoustic sensor uses a forward-looking sonar with an operating frequency of 900kHz, a scanning range of 120°, and a maximum detection distance of 100 meters. The flow field sensors include an acoustic Doppler current profiler and a pressure sensor array, with sampling frequencies of 5Hz and 10Hz, respectively.

[0053] For the acquired multimodal raw data, water scattering correction and noise filtering are performed on the optical images;

[0054] Specifically, water scattering correction employs an improved dark channel prior algorithm to recover the true color information of the image by estimating the transmittance map of the underwater image; noise filtering uses a bilateral filtering algorithm to remove noise while preserving image edge information; range-azimuth features are extracted from sonar data; the original sonar echo signal is converted into a two-dimensional sonar image in polar coordinates using a beamforming algorithm; target echo features are extracted using an adaptive threshold segmentation algorithm; and Kalman filtering and cubic spline interpolation are applied to the flow field parameters. Kalman filtering is used to remove random noise from the flow field measurement data, while cubic spline interpolation is used to supplement missing values ​​in the flow field data, ensuring the continuity and integrity of the flow field data.

[0055] Based on timestamp synchronization technology, the processed multimodal data is time-series aligned to obtain preprocessed multimodal data;

[0056] Specifically, using the timestamps of sonar data as a benchmark, the optical image data and flow field parameter data are time-aligned. A linear interpolation method is used to calculate the data values ​​at asynchronous moments to ensure the consistency of the three modal data in the time dimension. At the same time, the data collected by different sensors are mapped to a unified spatial coordinate system through coordinate transformation to achieve spatial dimension alignment, and finally, preprocessed multimodal data is obtained.

[0057] like Figure 2 As shown, in step S2, for the preprocessed multimodal data, target appearance features and flow field features are extracted respectively, the appearance features and flow field features are dynamically weighted and fused, and the dynamic weighted fusion result is classified by a classifier to obtain the target recognition result.

[0058] Specifically, in this embodiment, for the preprocessed multimodal data, a first network branch including a visual branch and an acoustic branch, and a second network branch including a flow field branch are constructed to extract the appearance features and flow field features of the target, respectively.

[0059] Specifically, the vision branch adopts a deep residual network (ResNet-50) structure, containing 50 convolutional layers. The input is the preprocessed optical image, and the output is a 2048-dimensional feature vector. The acoustic branch adopts an improved VGGNet structure, containing 13 convolutional layers and 3 fully connected layers. The input is the range-azimuth map of sonar data, and the output is a 1024-dimensional feature vector. The flow field branch adopts a graph convolutional network (GCN) structure, which constructs the flow field parameters into a spatial graph structure. It extracts flow field features through 3 graph convolutional layers and outputs a 512-dimensional feature vector.

[0060] The apparent features and flow field features are correlated through a cross-modal attention mechanism to form fused features;

[0061] Specifically, the cross-modal attention mechanism includes three steps: First, the similarity matrix between the appearance features and the flow field features is calculated, and the similarity is measured by cosine distance; second, the similarity matrix is ​​normalized by softmax to obtain the attention weights; finally, the features are weighted and summed according to the attention weights to generate fused features. The fused features have a dimension of 3584 and include visual, acoustic and flow field dynamic information.

[0062] The fusion features are dynamically weighted according to the real-time flow field intensity to form adaptive fusion features;

[0063] The dynamic weight allocation employs an adaptive gating mechanism, automatically adjusting the weight ratios of visual, acoustic, and flow field features based on flow field intensity indicators (including turbulence intensity and velocity gradient). When the flow field intensity is low (e.g., turbulence intensity < 0.2), the weight of visual features is automatically increased, utilizing the high-resolution texture information of optical images. When the flow field intensity is high (e.g., velocity gradient > 0.5), the weights of acoustic and flow field features are significantly increased, leveraging the penetrating power of sonar signals in complex flow fields. The weight calculation formula is as follows:

[0064] W i =softmax(MLP) i (F))

[0065] Among them, W i Here, represents the weight coefficients of the i-th modal feature, softmax is the normalized exponential function, and MLP is... i Let F be the three-layer perceptron network corresponding to the i-th mode, and F be the flow field intensity index vector.

[0066] By calculating the weights, the weight coefficients of each mode are output, thereby achieving adaptive optimization allocation of feature weights under different flow field environments;

[0067] The adaptive fusion features are processed by a classifier to output the target category probability distribution, thus forming the target recognition result;

[0068] Specifically, the classifier adopts a fully connected neural network structure, containing three fully connected layers with 1024 and 512 neurons in the hidden layers, respectively, using ReLU as the activation function. The output layer uses the softmax activation function to output the probability distribution of the target category. The classifier is trained using the cross-entropy loss function, with the Adam optimization algorithm, a learning rate of 0.001, a batch size of 64, and 100 training rounds. After 100 rounds of iterative training, accurate identification of the target category is achieved.

[0069] like Figure 2 As shown, in step S3, a trajectory prediction model is constructed based on the target identification result and flow field characteristics. The trajectory prediction model is then optimized by combining the target's historical trajectory data to obtain the initial trajectory prediction result.

[0070] Specifically, in this embodiment, based on the target recognition results and the flow field features extracted from the preprocessed multimodal data, a trajectory prediction model is constructed using a long short-term memory network. The model is input with historical target trajectory data, the identified target features, and the flow field features, and outputs preliminary target trajectory prediction data.

[0071] Specifically, the Long Short-Term Memory (LSTM) network consists of two LSTM units, each containing 128 hidden states. The input is a concatenation of the target's historical trajectory coordinate sequence (containing position and velocity information from the past 20 time steps), the target recognition feature vector (512-dimensional), and the flow field feature vector (256-dimensional). After forward computation by the LSTM, the output is the position sequence for the next 10 time steps, forming preliminary target trajectory prediction data. The parameters of the forget gate, input gate, and output gate of the LSTM unit are learned through the backpropagation algorithm, and the weights are initialized using the Xavier method to ensure stable gradient propagation.

[0072] The preliminary target trajectory prediction data is compared with the actual target motion trajectory data to calculate the trajectory deviation;

[0073] Specifically, trajectory deviation includes three indicators: average Euclidean distance error (ADE), final point Euclidean distance error (FDE), and trajectory shape similarity (TSS). The ADE is calculated as the average Euclidean distance between corresponding points of the predicted trajectory and the true trajectory; the FDE is the Euclidean distance between the endpoint of the predicted trajectory and the endpoint of the true trajectory; and the TSS uses the Dynamic Time Warping (DTW) algorithm to calculate the shape similarity between the two trajectories.

[0074] Based on the trajectory deviation, fluid dynamics equation constraints are embedded into the state update equation of the long short-term memory network to form a physically constrained trajectory prediction model.

[0075] Specifically, the fluid dynamics equation constraints include a simplified form of the Navier-Stokes equations, comprehensively considering the fluid resistance, buoyancy, and inertial forces acting on the underwater target, and adding physical constraint terms to the LSTM state update equations:

[0076] h t =LSTM(x t ,h t-1 )+λ·Φ(h t-1 ,f t )

[0077] Among them, h t Let x be the hidden state at time t. t For input features, h t-1 Let f be the hidden state at the previous time step, λ be the physical constraint weight parameter, Φ be the constraint function based on the fluid dynamics equations, and f be the hidden state at the previous time step. t The flow field characteristics at time t;

[0078] The physical constraint function Φ includes the target motion equation, fluid resistance model, and buoyancy calculation, and the parameters are dynamically adjusted according to the target type and flow field characteristics.

[0079] The physical constraint trajectory prediction model is optimized and trained to generate initial trajectory prediction results;

[0080] Specifically, the optimization training employs an iterative optimization of the physical constraint trajectory model using a joint loss function, which is set as follows:

[0081] L = L pred +α·L phys

[0082] Among them, L pred For trajectory prediction loss (using mean squared error), L phys The physical constraint loss (measures the consistency between the predicted trajectory and the physical model) is α, which is the balance coefficient. The initial value is set to 0.5 and is dynamically adjusted during the training process.

[0083] The optimization algorithm used is Adam, with a learning rate of 0.0005, a batch size of 32, and 200 training rounds. An early stopping strategy is adopted to avoid overfitting (the training stops if the validation set loss does not decrease for 10 consecutive rounds).

[0084] After training, the model outputs a trajectory prediction result that integrates physical constraints, i.e., the initial trajectory prediction result, which simultaneously satisfies the historical trajectory feature mapping and fluid dynamics laws, and has higher accuracy and robustness compared to the initial trajectory.

[0085] like Figure 2 As shown, in step S4, an implicit flow field model is constructed based on the initial trajectory prediction result. The output result of the implicit flow field model is then jointly trained with the trajectory prediction model to obtain an optimized trajectory prediction model.

[0086] Specifically, in this embodiment, a physical sensing neural network is constructed based on the initial trajectory prediction result. The flow field measurement data and fluid dynamics equation constraints are input, and an implicit flow field model is constructed for the unmeasurable flow field region to generate a flow field distribution feature representation.

[0087] Specifically, the physical perception neural network adopts the Physical Information Neural Network (PINN) architecture, which contains 8 fully connected layers, each with 256 neurons and the activation function tanh. The input is spatial coordinates (x, y, z) and time t, and the output is the velocity vector (u, v, w) and pressure p at that spatiotemporal point. PINN is trained by minimizing the physical residual loss function. The physical residual includes the continuity equation residual and the momentum equation residual, ensuring that the predicted flow field satisfies the basic laws of fluid mechanics. The training data includes flow field data from sensor measurement points and a large number of unlabeled spatial points. The entire flow field is reconstructed through physical constraints, providing flow field feature support for unmeasurable regions for subsequent trajectory models.

[0088] The flow field distribution characteristics output by the implicit flow field model are correlated with the trajectory prediction model to establish a mapping relationship between flow field characteristics and target motion trajectory.

[0089] Specifically, the data association adopts a spatiotemporal attention mechanism. Based on the target's current position and predicted trajectory, flow field features of relevant regions are extracted from the implicit flow field model. A region of interest (ROI) is constructed around the target's predicted trajectory, and the flow field vector field and scalar field features within the region are extracted. The degree of association between the target trajectory and the flow field features is calculated through attention weights. This association process dynamically binds the flow field features to the trajectory, providing a cross-modal data foundation for joint training.

[0090] Based on the flow field parameters output by the implicit flow field model and the trajectory prediction results of the trajectory prediction model, a joint training objective function including the flow field equation residuals and trajectory fitting errors is designed.

[0091] Specifically, the joint training objective function is defined as:

[0092] L=ω1·L traj +ω2·Lphys +ω3·L consist

[0093] Where L is the total loss of joint training, L traj For trajectory prediction loss, L phys For physical constraint loss, L consist The consistency loss is represented by w1, w2, and w3, which are weighting coefficients.

[0094] The trajectory prediction loss measures the trajectory prediction accuracy through mean square error, the physical constraint loss constrains the flow field equation residuals, and the consistency loss ensures the consistency of the flow field-trajectory model. The weight coefficients w1, w2, and w3 are initially set to 0.5, 0.3, and 0.2, respectively, and are dynamically adjusted during training to form a multi-objective balanced optimization system.

[0095] The physical sensing neural network and the trajectory prediction model are jointly trained end-to-end. By optimizing the joint training objective function, an optimized trajectory prediction model that incorporates the physical constraints of the flow field is obtained.

[0096] Specifically, the joint training adopts an alternating optimization strategy. First, the parameters of the trajectory prediction model are fixed, and the implicit flow field model is optimized. Then, the parameters of the implicit flow field model are fixed, and the trajectory prediction model is optimized. Finally, end-to-end fine-tuning is performed. The optimization algorithm is Adam, with a learning rate of 0.0002, a batch size of 16, and 300 training rounds. A learning rate decay strategy is adopted, with the learning rate reduced to 0.5 times the original value every 100 rounds. After 300 rounds of training, an optimized trajectory prediction model that integrates the physical constraints of the flow field is obtained.

[0097] like Figure 2 As shown, in step S5, the target motion trajectory prediction result output by the optimized trajectory prediction model is obtained, the deviation is calculated by combining the actual target motion trajectory data, and the physical constraint weights and target recognition parameters in the trajectory prediction model are dynamically adjusted according to the deviation.

[0098] Specifically, in this embodiment, the target motion trajectory prediction result output by the optimized trajectory prediction model is obtained, and the actual motion trajectory data of the target is collected. The target motion trajectory prediction result and the actual motion trajectory data are spatiotemporally aligned, and the position deviation, velocity deviation, trajectory curvature deviation and recognition error rate are calculated to form a comprehensive deviation.

[0099] Specifically, positional deviation is quantified using the average Euclidean distance error (ADE) and final point Euclidean distance error (FDE); velocity deviation is calculated using the L2 norm of the velocity vector difference between the predicted and actual trajectories at each time point; trajectory curvature deviation is obtained by comparing the curvature distributions of the predicted and actual trajectories; the recognition error rate is the reciprocal of the accuracy of target category recognition, and the overall deviation is calculated by weighted summation as follows:

[0100] E total =β1·E pos +β2·E vel +β3·E curv +β4·E id

[0101] Among them, E total For the overall deviation, E pos E represents the positional deviation. vel For speed deviation, E curv E represents the trajectory curvature deviation. id For the identification error rate, β1 is the position deviation weighting coefficient, β2 is the velocity deviation weighting coefficient, β3 is the trajectory curvature deviation weighting coefficient, and β4 is the identification error rate weighting coefficient.

[0102] The default values ​​for β1, β2, β3, and β4 are 0.4, 0.3, 0.2, and 0.1, respectively, and can be dynamically adjusted according to the application scenario.

[0103] Based on the specific values ​​and distribution characteristics of the comprehensive deviation, the adjustment rules and quantification parameters for the fluid dynamics equation constraint weights and target identification parameters are determined.

[0104] Specifically, the adjustment rule adopts a gradient-based adaptive adjustment strategy: when the overall deviation increases, the physical constraint weights are increased to strengthen the constraints of fluid dynamics on trajectory prediction; when the overall deviation decreases, the physical constraint weights are decreased to release the flexibility of the data-driven model. The update formula for the physical constraint weights λ is as follows:

[0105] λ new =λ old +η·sign(ΔE total )·|ΔE total |

[0106] Where, λ new For the updated physical constraint weights, λ old The physical constraint weights before the update are η, the learning rate is set to 0.01, and ΔE is the learning rate. total The change in the overall deviation is represented by sign, where sign is the sign function.

[0107] The target recognition parameters are adjusted using the backpropagation algorithm, which calculates the gradient based on the recognition error rate and updates the recognition parameters.

[0108] For the trajectory prediction model and target recognition parameters, the physical constraint weights and recognition parameters in the model are adaptively updated according to the adjustment rules.

[0109] Specifically, the adaptive update process is executed every 10 time steps. The comprehensive deviation is calculated based on the most recently collected trajectory data, and the model parameters are updated using the above adjustment rules. The updated model parameters are immediately applied to subsequent target recognition and trajectory prediction tasks, forming a closed-loop optimization mechanism. The value range of the physical constraint weights is limited to [0.1, 1.0] to avoid over-reliance on or neglect of the physical model. The target recognition parameter update adopts the momentum gradient descent method to maintain the smoothness and stability of the parameter update.

[0110] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0111] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and rules of the present invention should be included within the scope of protection of the present invention.

Claims

1. A deep learning-based method for underwater robot target recognition, tracking, and evaluation, characterized in that, Includes the following steps: S1. Acquire target optical images, sonar echo signals and flow field parameters, construct multimodal data, preprocess and spatiotemporally align the multimodal data to obtain preprocessed multimodal data; S2. For the preprocessed multimodal data, extract the target appearance features and flow field features respectively, perform dynamic weight fusion on the appearance features and flow field features, and classify the dynamic weight fusion result through a classifier to obtain the target recognition result. S3. Based on the target identification results and flow field characteristics, construct a trajectory prediction model, and optimize the trajectory prediction model by combining the target's historical trajectory data to obtain the initial trajectory prediction result; S4. Based on the initial trajectory prediction results, construct an implicit flow field model, and jointly train the output results of the implicit flow field model with the trajectory prediction model to obtain an optimized trajectory prediction model. S5. Obtain the target motion trajectory prediction result output by the optimized trajectory prediction model, calculate the deviation by combining the actual target motion trajectory data, and dynamically adjust the physical constraint weights and target recognition parameters in the trajectory prediction model according to the deviation.

2. The underwater robot target recognition and tracking evaluation method based on deep learning according to claim 1, characterized in that: The S1 step specifically includes: The optical image data, sonar echo signal, and flow field physical parameters of the target are collected to form multimodal raw data; For the original multimodal data, water scattering correction and noise filtering are performed on the optical image, range-azimuth features are extracted from the sonar data, and Kalman filtering and cubic spline interpolation are performed on the flow field parameters to obtain the processed multimodal data. The processed multimodal data is time-aligned using timestamp synchronization technology to obtain preprocessed multimodal data.

3. The underwater robot target recognition and tracking evaluation method based on deep learning according to claim 1, characterized in that: The S2 step specifically includes: For the preprocessed multimodal data, a first network branch including a visual branch and an acoustic branch, and a second network branch including a flow field branch are constructed to extract the appearance features and flow field features of the target, respectively. The apparent features and flow field features are correlated through a cross-modal attention mechanism to form fused features; The fusion features are dynamically weighted according to the real-time flow field intensity to form adaptive fusion features; The adaptive fusion features are processed by a classifier to form the target recognition result.

4. The underwater robot target recognition and tracking evaluation method based on deep learning according to claim 1, characterized in that: The S3 step specifically includes: Based on the target recognition results and the flow field features extracted from the preprocessed multimodal data, a trajectory prediction model is constructed using a long short-term memory network. The model takes historical target trajectory data, the identified target features, and the flow field features as inputs and outputs preliminary target trajectory prediction data. The preliminary target trajectory prediction data is compared with the actual target motion trajectory data to calculate the trajectory deviation; Based on the trajectory deviation, fluid dynamics equation constraints are embedded into the state update equation of the long short-term memory network to form a physically constrained trajectory prediction model. The physical constraint trajectory prediction model is optimized and trained to generate initial trajectory prediction results.

5. The underwater robot target recognition and tracking evaluation method based on deep learning according to claim 1, characterized in that: The S4 step specifically includes: Based on the initial trajectory prediction results, a physical sensing neural network is constructed. The flow field measurement data and fluid dynamics equation constraints are input, and an implicit flow field model is constructed for the unmeasurable flow field region to generate a flow field distribution characteristic representation. The flow field distribution characteristics output by the implicit flow field model are correlated with the trajectory prediction model to establish a mapping relationship between flow field characteristics and target motion trajectory. Based on the flow field parameters output by the implicit flow field model and the trajectory prediction results of the trajectory prediction model, a joint training objective function is designed that includes the flow field equation residuals and the trajectory fitting error. The physical perception neural network and the trajectory prediction model are jointly trained end-to-end. By optimizing the joint training objective function, an optimized trajectory prediction model that incorporates the physical constraints of the flow field is obtained.

6. The underwater robot target recognition and tracking evaluation method based on deep learning according to claim 1, characterized in that: The S5 step specifically includes: The target motion trajectory prediction result output by the optimized trajectory prediction model is obtained, and the actual motion trajectory data of the target is collected. The target motion trajectory prediction result and the actual motion trajectory data are spatiotemporally aligned. The position deviation, velocity deviation, trajectory curvature deviation and recognition error rate are calculated to form a comprehensive deviation. Based on the specific values ​​and distribution characteristics of the comprehensive deviation, the adjustment rules and quantification parameters for the fluid dynamics equation constraint weights and target identification parameters are determined. For the trajectory prediction model and target recognition parameters, the physical constraint weights and recognition parameters in the model are adaptively updated according to the adjustment rules.

Citation Information

Cited By

  • Video target identification method and system based on deep learning

    CN121545105A

  • A deep learning-based video target recognition method and system

    CN121545105B

  • Positioning method and system for underwater environment

    CN121958897A