A method and system for dynamic tracking of underwater terrain

By combining the Vision Mamba sensing backbone network and the improved Kalman filter, the reliability and robustness issues of observation fusion in underwater topographic surveying are solved, and high-precision underwater topographic dynamic tracking measurement is achieved.

CN121007539BActive Publication Date: 2026-02-17TIANJIN BORUI TESTING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511179903.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2026-02-17
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing underwater topographic surveying methods struggle to obtain highly reliable spatiotemporal observation information in complex environments. Traditional observation fusion methods are susceptible to low visibility and noise interference, leading to cumulative drift and decreased estimation reliability.

Method used

The Vision Mamba sensing backbone network is used for encoding and feature representation of multi-source heterogeneous sensing data. An improved Kalman filter is combined for adaptive covariance adjustment and robustness testing to dynamically remove outlier observations, thereby achieving geometric consistency registration of multimodal features and high-confidence observations.

Benefits of technology

It improves the spatiotemporal continuity and quality controllability of observation data, significantly enhances the noise and drift resistance of dynamic state estimation, and realizes high-precision underwater topographic dynamic tracking measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121007539B_ABST
    Figure CN121007539B_ABST
Patent Text Reader

Abstract

The application discloses a kind of underwater topography dynamic tracking measurement method and system, comprising: outputting calibrated and preprocessed underwater multi-source heterogeneous perception data sequence;Obtain multi-scale multi-modal feature representation carrying visibility score and matching confidence;Unified observation information frame is constructed;Form time-consistent alignment observation sequence;Output prior prediction result;Adaptive adjustment of measurement covariance and process covariance is carried out and consistency test and robust loss estimation are executed, and topography state estimation and its topography state estimation covariance are obtained;Complete underwater topography dynamic tracking measurement.The application can improve the utilization rate of high-confidence observation points in real time in complex scenarios, dynamically suppress the cumulative error caused by abnormal points, realize continuous and high-precision estimation of topography state, effectively improve the automation level of map updating and loop correction, and better complete underwater topography dynamic tracking measurement.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of tracking measurement, in particular to an underwater terrain dynamic tracking measurement method and system. BACKGROUND

[0002] With the rapid development of ocean engineering, ocean resource development and underwater intelligent equipment, underwater terrain dynamic tracking measurement has become a basic ability for practical applications such as underwater inspection, channel survey and target search. Currently, the mainstream underwater terrain measurement method mainly relies on optical cameras and sonar sensors, and three-dimensional reconstruction and dynamic mapping of underwater environment are achieved through feature extraction and multi-sensor fusion. Due to the low contrast, strong scattering, multipath echo interference of optical images in complex underwater environment, and the inconsistency of different sensor frame rates, fields of view and coordinate systems, the existing method often fails to obtain high-credible spatio-temporal observation information in continuous terrain tracking scenarios.

[0003] The observation fusion method based on traditional local feature points or template matching is easily affected by low visibility and sonar false images, and the stability and robustness of the observation results are insufficient. At the same time, the conventional Kalman filter or extended Kalman filter is mainly based on Gaussian noise assumption, and the measurement noise and process noise are generally fixed, which is difficult to adapt to the heavy-tailed distribution, model mismatch and sudden abnormal noise existing in underwater observation. In the case of multi-rate, asynchronous fusion and complex sea conditions, the cumulative drift and estimation credibility decline problems are particularly prominent. SUMMARY

[0004] One object of the present application is to provide an underwater terrain dynamic tracking measurement method and system. The present application can improve the utilization rate of high-confidence observation points in complex scenarios in real time and adaptively, dynamically suppress the cumulative error caused by abnormal points, realize continuous and high-precision estimation of terrain state, effectively improve the automation level of map updating and loop correction, and better complete underwater terrain dynamic tracking measurement.

[0005] An underwater terrain dynamic tracking measurement method according to an embodiment of the present application comprises:

[0006] Collecting underwater multi-source heterogeneous perception data, and calibrating and preprocessing the data to output a calibrated and preprocessed underwater multi-source heterogeneous perception data sequence;

[0007] Encoding the calibrated and preprocessed underwater multi-source heterogeneous perception data sequence into a spatio-temporal Token sequence and inputting it into a Vision Mamba perception backbone network to obtain a multi-scale multi-modal feature representation carrying visibility scores and matching confidence;

[0008] Based on the obtained multi-scale multi-modal feature representation carrying the visibility score and the matching confidence, geometric consistency registration is completed in the cross-modal alignment head of the VisionMamba perception backbone network, and a unified observation information frame is constructed;

[0009] The unified observation information frame is subjected to observation threshold and dynamic weighting processing, and the unified observation information frame subjected to observation threshold and dynamic weighting processing is subjected to multi-rate time alignment with inertial measurement unit observation data, to form a time-aligned observation sequence with consistent timing;

[0010] An interactive multi-model motion prediction set is constructed on the aligned observation sequence, and a sea current disturbance term is introduced to calculate state prior and covariance prior, and output a prior prediction result;

[0011] The unified observation information frame subjected to observation threshold and dynamic weighting processing is fused with the prior prediction result in an improved Kalman filter, and the measurement covariance and the process covariance are adaptively adjusted, and consistency checking and robust loss estimation are performed, to obtain terrain state estimation and terrain state estimation covariance;

[0012] The terrain state estimation and terrain state estimation covariance are projected to a map expression, to complete dynamic tracking measurement of underwater terrain.

[0013] Optionally, the construction of the calibrated and preprocessed underwater multi-source heterogeneous perception data sequence comprises:

[0014] Underwater multi-source heterogeneous perception data is collected, which includes optical camera observation data, sonar observation data, structured light observation data, laser stripe observation data, and inertial measurement unit observation data;

[0015] The optical camera observation data is subjected to dewarping, de-coloring, and light equalization processing;

[0016] The sonar-related observation data is subjected to polar-to-Cartesian domain transformation and intensity normalization processing;

[0017] The structured light observation data and the laser stripe observation data are subjected to stripe extraction and geometric correction processing;

[0018] The inertial measurement unit observation data is subjected to denoising and interpolation padding processing;

[0019] The calibrated and preprocessed underwater multi-source heterogeneous perception data sequence is output.

[0020] Optionally, the multi-scale multi-modal feature representation carrying the visibility score and the matching confidence comprises:

[0021] The calibrated and preprocessed underwater multi-source heterogeneous perception data sequence is block encoded according to a unified space-time structure, each block is a space-time Token, and all Tokens form a complete space-time Token sequence;

[0022] For each Token at a specified time and space position in the space-time Token sequence, the space-time Token features of the current time and all previous times at the corresponding space position are fused to generate long-range time sequence state space features;

[0023] The long-range time sequence state space features at each time are respectively pooled and convoluted at different spatial scales using feature decomposition to obtain multi-scale feature representations, and each scale has corresponding multi-modal feature representations;

[0024] The multi-modal feature representations at each time and each spatial scale are respectively channel spliced and attention enhanced in the Vision Mamba perception backbone network to obtain fused and enhanced multi-scale multi-modal feature representations;

[0025] Based on the fused and enhanced multi-scale multi-modal feature representations, for each time, each spatial scale, and each spatial position, a feature-to-score mapping structure is used to generate visibility scores and matching confidence scores respectively;

[0026] For each time, each spatial scale, and each spatial position, the fused and enhanced multi-scale multi-modal feature representations, visibility scores, and matching confidence scores are combined to form multi-scale multi-modal feature representations carrying visibility scores and matching confidence scores.

[0027] Optionally, the construction of the unified observation information frame includes:

[0028] The multi-scale multi-modal feature representations carrying visibility scores and matching confidence scores are input into the cross-modal alignment head in the Vision Mamba perception backbone network, and for each time, each spatial scale, and each spatial position, the feature representations of the optical camera observation channel, the sonar observation channel, and the inertial measurement unit observation channel are extracted respectively, and the feature representations are mapped to a unified geometric space;

[0029] In the unified geometric space, for each time, each spatial scale, and each spatial position, based on the spatial position consistency constraint, the distance between the feature representations of the optical camera observation channel, the sonar observation channel, and the inertial measurement unit observation channel at the same time and the same spatial position is taken as the registration residual;

[0030] Based on the registration residual of each time, each spatial scale and each spatial position, the feature representation of the optical camera observation channel, the sonar observation channel and the inertial measurement unit observation channel at the same spatial position is dynamically adjusted to obtain a dynamically adjusted unified feature representation;

[0031] The dynamically adjusted unified feature representation, the time index, the spatial index, the visibility score, the matching confidence and the registration residual are taken as observation attributes to construct a unified observation information frame.

[0032] Optionally, the forming of the time-sequentially consistent aligned observation sequence comprises:

[0033] The visibility score, the matching confidence and the registration residual of each observation position at a specified time, spatial scale and spatial position in the unified observation information frame are extracted;

[0034] The visibility score and the matching confidence of each observation position are compared, and the smaller one of the two is taken as an observation confidence score;

[0035] The observation confidence score and the registration residual of each observation position are combined to calculate a final measurement weight of the spatial position;

[0036] An observation weight threshold is set, when the final measurement weight of each observation position is lower than the observation weight threshold, it is determined as an abnormal observation and is removed, all low measurement weight observations are marked as invalid or do not participate in subsequent processing from the unified observation information frame, only the valid observation positions with a weight greater than or equal to the threshold are reserved;

[0037] The unified observation information frame after removing the abnormal observation is time-aligned with the inertial measurement unit observation data to generate observation pairs, and all observation pairs constitute a time-sequentially consistent aligned observation sequence.

[0038] Optionally, the output of the prior prediction result comprises:

[0039] Based on the time-sequentially consistent aligned observation sequence, an interacting multiple model motion prediction set is constructed, and the interacting multiple model motion prediction set comprises a constant speed motion model, a constant acceleration motion model and a coordinated turning motion model;

[0040] The constant speed motion model, the constant acceleration motion model and the coordinated turning motion model all take the state estimation value of the previous time as input, respectively perform linear transformation through the corresponding state transition matrix, superimpose the respective corresponding sea current disturbance term to reflect the influence of the sea current on the target state in the actual underwater environment, and introduce process noise subject to zero mean Gaussian distribution to obtain the state prior of each motion model at the current time;

[0041] For all motion models in the interactive multi-model motion prediction set, the model probability of each motion model is used as a weighting coefficient to weight and fuse the state prior and process noise of each model at the current time, to obtain the fused state prior and covariance prior at the current time, and to use them as the output prior prediction result.

[0042] Optionally, the terrain state estimation including the terrain elevation, terrain curvature, terrain feature point and carrier pose, and the terrain state estimation covariance thereof, comprises:

[0043] The observed threshold and the unified observation information frame after dynamic weighting processing are input into the improved Kalman filter together with the prior prediction result;

[0044] For each spatial position, different adjustment coefficients are respectively given according to the visibility score, matching confidence and registration residual, and weighted superposition is performed to obtain the measurement covariance of the current spatial position;

[0045] The measurement covariance of each spatial position and the prior prediction result are input into the improved Kalman filter update together, and the mapping relationship of the state space and the observation space by the observation matrix is used to calculate the filtering gain at the current time;

[0046] The difference between the observation vector in the unified observation information frame and the state prediction is calculated as a quantity residual;

[0047] The Huber loss function is applied to the quantity residual for nonlinear weighting, the Huber loss function is used to suppress the influence of extreme observation errors on the filtering result, the weighted quantity residual is updated by the filtering gain to obtain the terrain state estimation at the current time, and the covariance is also updated synchronously;

[0048] The output terrain state estimation at the current time is calculated as the geometric variables and attitude variables of the terrain elevation, terrain curvature, terrain feature point and carrier pose, and the terrain state estimation covariance at the current time is retained.

[0049] An underwater terrain dynamic tracking measurement system for performing an underwater terrain dynamic tracking measurement method, comprising:

[0050] A data acquisition and synchronization module is configured to acquire underwater multi-source heterogeneous sensing data, and to complete data calibration and preprocessing under a unified time reference and a unified coordinate system, and to output a calibrated and preprocessed underwater multi-source heterogeneous sensing data sequence;

[0051] The perception coding and feature modeling module comprises a Vision Mamba perception backbone network, and is used for coding a calibrated and preprocessed underwater multi-source heterogeneous perception data sequence into a spatio-temporal Token sequence, performing multi-scale multi-modal feature modeling based on a selective structured scanning mechanism, and outputting a multi-scale multi-modal feature representation carrying a visibility score and a matching confidence;

[0052] The cross-modal registration and unified observation information frame construction module is used for realizing geometric consistency registration of an optical camera observation channel, a sonar observation channel and an inertial measurement unit observation channel in a cross-modal alignment head based on the multi-scale multi-modal feature representation carrying the visibility score and the matching confidence, and generating a unified observation information frame;

[0053] The dynamic weighting and multi-rate alignment module is used for performing observation threshold and dynamic weighting processing on the unified observation information frame, generating an observation credibility score and a measurement weight according to the visibility score, the matching confidence and a registration residual, eliminating abnormal observation, and performing multi-rate time alignment with inertial measurement unit observation data, and outputting an aligned observation sequence with consistent time sequence;

[0054] The motion prediction and state estimation module comprises an interactive multi-model motion prediction structure and an improved Kalman filter, is used for constructing an interactive multi-model motion prediction set based on the aligned observation sequence, generating a prior prediction result, and adaptively adjusting a measurement covariance, and outputs a terrain state estimation and a terrain state estimation covariance;

[0055] The map expression and incremental update module is used for projecting the terrain state estimation and the terrain state estimation covariance to a terrain map subgraph of a signed distance field map and an elevation grid map.

[0056] The present application has the following advantages:

[0057] 1、The present application effectively improves the spatio-temporal continuity and quality controllability of observation data based on long-time sequence multi-modal feature modeling and observation credibility quantization of Vision Mamba, encodes a calibrated and preprocessed underwater multi-source heterogeneous perception data sequence into a unified spatio-temporal Token sequence, uses a selective structured scanning mechanism of Vision Mamba to deeply model long-time sequence information, automatically extracts multi-scale multi-modal features and generates a visibility score and a matching confidence for each spatial position, overcomes the deficiency that a traditional local feature method is easy to fail under strong scattering, low visibility and multi-path echo, and realizes global consistency and high confidence output of observation data in the spatio-temporal dimension.

[0058] 2、The improved Kalman filter realizes observation-driven adaptive covariance adjustment and robustness test, significantly improves the anti-noise and anti-drift ability of dynamic state estimation, dynamically adjusts the measurement covariance of each spatial position based on the visibility score, matching confidence and registration residual of the observation frame output in the Kalman filter update stage, and adaptively adjusts the process covariance combined with the prior prediction result, effectively solves the error amplification and drift problem of the traditional filter under non-Gaussian heavy-tailed noise, model mismatch and sea state mutation, and introduces the robust loss function of the residual to perform consistency test and nonlinear weighting on abnormal observation, and enhances the inhibition ability to sudden abnormality and outliers. BRIEF DESCRIPTION OF DRAWINGS

[0059] The accompanying drawings are included to provide a further understanding of the application, and constitute a part of the specification, together with the embodiments of the application, to explain the application, and do not constitute a limitation on the application. In the drawings:

[0060] Figure 1 A flowchart of an underwater terrain dynamic tracking measurement method and system is provided. DETAILED DESCRIPTION

[0061] Example 1:

[0062] Reference Figure 1 An underwater terrain dynamic tracking measurement method, comprising:

[0063] Collecting underwater multi-source heterogeneous perception data and performing calibration and preprocessing to output calibrated and preprocessed underwater multi-source heterogeneous perception data sequence;

[0064] Encoding the calibrated and preprocessed underwater multi-source heterogeneous perception data sequence into a spatio-temporal Token sequence and inputting it into a Vision Mamba perception backbone network to obtain a multi-scale multi-modal feature representation carrying visibility scores and matching confidence;

[0065] Based on the obtained multi-scale multi-modal feature representation carrying visibility scores and matching confidence, geometric consistency registration is completed in the cross-modal alignment head of the Vision Mamba perception backbone network to construct a unified observation information frame;

[0066] Performing observation threshold and dynamic weighting processing on the unified observation information frame, and performing multi-rate time alignment between the unified observation information frame subjected to observation threshold and dynamic weighting processing and inertial measurement unit observation data to form a time-sequential aligned observation sequence;

[0067] Constructing an interactive multi-model motion prediction set on the aligned observation sequence, and introducing a sea current disturbance term to perform state prior and covariance prior calculation, and outputting a prior prediction result;

[0068] The observed threshold is unified with the dynamic weighted processed unified observation information frame and the prior prediction result in the improved Kalman filter, the measurement covariance and the process covariance are adaptively adjusted according to the visibility score, the matching confidence and the registration residual, and the consistency test and the robust loss estimation are performed, and the terrain state estimation and the terrain state estimation covariance are obtained;

[0069] The terrain state estimation and the terrain state estimation covariance are projected to the map expression, and the underwater terrain dynamic tracking measurement is completed.

[0070] In the embodiment, the construction of the calibrated and preprocessed underwater multi-source heterogeneous perception data sequence includes:

[0071] The underwater multi-source heterogeneous perception data is collected, and the underwater multi-source heterogeneous perception data includes optical camera observation data, sonar observation data, structured light observation data, laser stripe observation data and inertial measurement unit observation data;

[0072] The optical camera observation data is subjected to defogging, color correction and illumination equalization processing;

[0073] The sonar related observation data is subjected to polar to Cartesian domain transformation and intensity normalization processing;

[0074] The structured light observation data and the laser stripe observation data are subjected to stripe extraction and geometric correction processing;

[0075] The inertial measurement unit observation data is subjected to denoising and interpolation padding processing;

[0076] The calibrated and preprocessed underwater multi-source heterogeneous perception data sequence is output.

[0077] In the embodiment, the multi-scale multi-modal feature representation carrying the visibility score and the matching confidence includes:

[0078] The calibrated and preprocessed underwater multi-source heterogeneous perception data sequence is block encoded according to a unified spatio-temporal structure, each block is a spatio-temporal Token, and all Tokens form a complete spatio-temporal Token sequence;

[0079] The spatio-temporal Token is organized in a three-dimensional structure, wherein the spatial scale is identified by two coordinate parameters of row and column, and the time dimension is identified by a frame sequence number parameter.

[0080] For each Token of a specified time and space position in the spatio-temporal Token sequence, the spatio-temporal Token features of the current time and all previous times corresponding to the space position are fused to generate a long-range time series state space feature;

[0081] The long-range time series state space feature represents the time series characteristics of the position fusing historical information at the current time.

[0082] The long-range time sequence state space features of each time are respectively pooled and convoluted at different spatial scales by using feature decomposition, to obtain multi-scale feature representations, and there are corresponding multi-modal feature representations at each scale;

[0083] The multi-scale feature representations enhance the spatial hierarchical information and modal expression capability of the features.

[0084] The multi-modal feature representations at each time, each spatial scale are respectively subjected to channel splicing and attention enhancement in the Vision Mamba perception backbone network, to obtain multi-scale multi-modal feature representations after fusion and enhancement.

[0085] In embodiment 1, the feature channels of all multi-modal feature representations are spliced in the channel dimension, the channel splicing fuses the optical observation feature channel, the sonar observation feature channel, the structured light observation feature channel, the laser stripe observation feature channel and the inertial measurement unit observation feature channel into a high-dimensional feature representation, after the channel splicing, the high-dimensional feature representation after splicing is subjected to attention enhancement by using the feature attention mechanism built in the Vision Mamba perception backbone network, the attention enhancement adaptively highlights the feature area most critical to the current underwater terrain dynamic tracking measurement in the multi-modal high-dimensional feature, and outputs the multi-scale multi-modal feature representations after fusion and enhancement.

[0086] Based on the multi-scale multi-modal feature representations after fusion and enhancement, for each time, each spatial scale and each spatial position, a feature-to-score mapping structure is used to respectively generate a visibility score and a matching confidence;

[0087] In embodiment 1, the multi-scale multi-modal feature representations after fusion and enhancement of the position are input into a specific multi-layer perception structure, which is respectively used to generate a visibility score and a matching confidence, the multi-layer perception structure performs nonlinear feature mapping processing on the input features, and outputs the original visibility score and the original matching confidence, and the original visibility score and the original matching confidence are respectively applied to a normalization function to be limited in the interval [0, 1], to obtain the visibility score reflecting the observation reliability of the spatial position and the matching confidence reflecting the multi-modal registration accuracy of the spatial position.

[0088] For each time, each spatial scale and each spatial position, the multi-scale multi-modal feature representations after fusion and enhancement, the visibility score and the matching confidence are combined to form multi-scale multi-modal feature representations carrying the visibility score and the matching confidence.

[0089] In the embodiment, the construction of the unified observation information frame includes:

[0090] The multi-scale multi-modal feature representation carrying the visibility score and the matching confidence is input into a cross-modal alignment head in the Vision Mamba perception backbone network, and for each time, each spatial scale, and each spatial position, feature representations of the optical camera observation channel, the sonar observation channel, and the inertial measurement unit observation channel are extracted respectively, and the feature representations are mapped to a unified geometric space.

[0091] In the unified geometric space, for each time, each spatial scale, and each spatial position, based on the spatial position consistency constraint, distances between the feature representations of the optical camera observation channel, the sonar observation channel, and the inertial measurement unit observation channel at the same time and the same spatial position are taken as registration residuals.

[0092] The registration residual is used to measure the geometric consistency deviation of the spatial position between different observation channels, and the smaller the registration residual value is, the better the registration effect of the spatial position in the unified geometric space.

[0093] Based on the registration residual of each time, each spatial scale, and each spatial position, the feature representations of the optical camera observation channel, the sonar observation channel, and the inertial measurement unit observation channel at the same spatial position are dynamically adjusted to obtain a dynamically adjusted unified feature representation.

[0094] The dynamic adjustment realizes the feature consistency of different observation channels in the unified geometric space by correcting the difference part of the feature, and the dynamically adjusted unified feature representation can represent the geometric consistency characteristics of the fusion of multi-observation channel information at the same spatial position.

[0095] The dynamically adjusted unified feature representation, the time index, the spatial index, the visibility score, the matching confidence, and the registration residual are taken as observation attributes to construct a unified observation information frame.

[0096] In the unified observation information frame, the time index is used to record the time information corresponding to the observation, the spatial index is used to record the spatial position information corresponding to the observation, the visibility score is used to reflect the reliability of the observation, the matching confidence is used to reflect the accuracy of the multi-modal registration, and the registration residual is used to represent the geometric consistency deviation of the observation. All observation attributes are one-to-one corresponding and are stored together in the unified observation information frame.

[0097] The time index in the unified observation information frame is one-to-one corresponding to the unified time reference, and the spatial index in the unified observation information frame is one-to-one corresponding to the unified coordinate system, forming a complete association of the observation information frame and the time and spatial reference. The associated observation information sequence is a time-sequential and spatial index-unified observation information sequence.

[0098] In the embodiment, the time-sequential alignment observation sequence is formed, including:

[0099] extracting the visibility score, the matching confidence and the registration residual of each observation location at the specified time, spatial scale and spatial position in the unified observation information frame;

[0100] comparing the visibility score and the matching confidence of each observation location, and taking the smaller value as the observation confidence score;

[0101] The observation confidence score measures the basic confidence level of the spatial position observation information, and the value range of the observation confidence score is limited to [0, 1].

[0102] combining the observation confidence score and the registration residual of each observation location, calculating the final measurement weight of the spatial position;

[0103] The final measurement weight is the product of the observation confidence score and the registration residual exponential decay function, which is used to reduce the weight of the observation point with large registration error.

[0104] Setting an observation weight threshold, when the final measurement weight of each observation location is lower than the observation weight threshold, it is determined as an abnormal observation and is rejected, all low measurement weight observations are marked as invalid or do not participate in subsequent processing in the unified observation information frame, only the valid observation positions with weight greater than or equal to the threshold are reserved;

[0105] After rejecting the abnormal observation, the unified observation information frame and the inertial measurement unit observation data are time-aligned to generate observation pairs, and all observation pairs constitute a time-consistent aligned observation sequence.

[0106] The time alignment processing is based on the unified time reference to build an aligned time index, and at each aligned time index, the high-frequency observation data of the inertial measurement unit is mapped to the time position consistent with the unified observation information frame using resampling to generate observation pairs.

[0107] In this embodiment, the output prior prediction result includes:

[0108] Based on the time-consistent aligned observation sequence, an interacting multiple model motion prediction set is constructed, which includes a constant speed motion model, a constant acceleration motion model and a coordinated turning motion model.

[0109] Each motion model is modeled by an independent state transition equation and a process noise equation.

[0110] The constant speed motion model, the constant acceleration motion model and the coordinated turning motion model all take the state estimation value of the previous time as input, respectively perform linear transformation through the corresponding state transition matrix, superimpose the corresponding sea current disturbance term to reflect the influence of the sea current on the target state in the actual underwater environment, and introduce process noise subject to zero mean Gaussian distribution to obtain the state prior of each motion model at the current time;

[0111] The state prior is used to describe the predicted state of the underwater terrain dynamic tracking target at the current time based on the motion hypothesis.

[0112] For all motion models in the interactive multi-model motion prediction set, the model probability of each motion model is used as a weighting coefficient to weight and fuse the state prior and process noise of each model at the current time to obtain the fused state prior and covariance prior at the current time, which are used as the prior prediction result of the output;

[0113] The fused state reflects the predicted state of the underwater terrain dynamic tracking target at the current time, and the fused covariance prior indicates the uncertainty of the predicted state.

[0114] In the embodiment, the terrain state estimation including terrain elevation, terrain curvature, terrain feature points and carrier pose and the terrain state estimation covariance are obtained, including:

[0115] The unified observation information frame after the observation threshold and dynamic weighting processing is input into the improved Kalman filter together with the prior prediction result;

[0116] For each spatial position, different adjustment coefficients are assigned according to the visibility score, matching confidence and registration residual, and weighted superposition is performed to obtain the measurement covariance of the current spatial position;

[0117] The measurement covariance reflects the observation error of the spatial position at the current time, and the numerical value of the measurement covariance dynamically adjusts the filtering gain of the spatial position.

[0118] The measurement covariance of each spatial position and the prior prediction result are input into the improved Kalman filter update, and the mapping relationship of the state space and the observation space is calculated through the observation matrix to calculate the filtering gain at the current time;

[0119] The difference between the observation vector in the unified observation information frame and the state prediction is calculated as the residual error;

[0120] The residual error is used to measure the actual deviation between the current observation and the prior state prediction.

[0121] The Huber loss function is applied to the quantity residual for nonlinear weighting, the Huber loss function is used to suppress the influence of extreme observation errors on the filtering result, the state is updated by the filtering gain, the terrain state estimation at the current time is obtained, and the covariance is also updated synchronously;

[0122] The terrain state estimation at the current time is calculated as terrain elevation, terrain curvature, terrain feature points and geometric variables and attitude variables of the carrier pose, and the terrain state estimation covariance at the current time is retained.

[0123] In the embodiment, the underwater terrain dynamic tracking measurement is completed, including:

[0124] The terrain state estimation and the terrain state estimation covariance are projected into the terrain map subgraph of the signed distance field map and the elevation grid map, respectively, the terrain state estimation includes terrain elevation, terrain curvature, terrain feature points and carrier pose, and all projection operations are based on unified coordinate system for spatial consistency mapping;

[0125] After each terrain state estimation projection, an incremental update operation is performed on the terrain map subgraph of the signed distance field map and the elevation grid map, and the corresponding region state and uncertainty information in the terrain map subgraph are corrected in real time by fusing the latest terrain state estimation and terrain state estimation covariance;

[0126] For all terrain map subgraphs, according to the spatial continuity, boundary consistency and feature point overlap relationship of the terrain state estimation and the terrain state estimation covariance, a subgraph splicing operation is performed, the subgraph splicing operation includes state fusion and uncertainty propagation of the spatial overlap region, and the spliced map subgraph forms a consistent global map expression;

[0127] According to the spatial index and feature description between the historical terrain state estimation and the current terrain state estimation, a loop consistency correction operation is performed, the loop consistency correction aligns the spatial position and state by identifying the visited region, comparing the terrain state estimation and its covariance of the current map subgraph and the historical map subgraph in the overlap region, correcting the cumulative drift, and improving the overall accuracy and stability of the map;

[0128] The terrain map subgraph of the signed distance field map and the elevation grid map after incremental update, subgraph splicing and loop consistency correction is output, forming the final underwater terrain dynamic tracking measurement result, the measurement result includes spatially continuous terrain state estimation, terrain state estimation covariance and spatial index correspondence, meeting the high-precision and repeatable underwater terrain dynamic tracking measurement requirements.

[0129] An underwater terrain dynamic tracking measurement system for performing an underwater terrain dynamic tracking measurement method, comprising:

[0130] a data acquisition and synchronization module configured to acquire underwater multi-source heterogeneous perception data, and to complete data calibration and preprocessing under a unified time reference and a unified coordinate system, and to output calibrated and preprocessed underwater multi-source heterogeneous perception data sequences;

[0131] a perception encoding and feature modeling module including a Vision Mamba perception backbone network, configured to encode the calibrated and preprocessed underwater multi-source heterogeneous perception data sequences into spatio-temporal Token sequences, to model multi-scale multi-modal features based on a selective structured scanning mechanism, and to output multi-scale multi-modal feature representations carrying visibility scores and matching confidence;

[0132] a cross-modal registration and unified observation information frame construction module configured to implement geometric consistency registration of optical camera observation channels, sonar observation channels and inertial measurement unit observation channels in a cross-modal alignment head based on the multi-scale multi-modal feature representations carrying visibility scores and matching confidence, and to generate unified observation information frames;

[0133] a dynamic weighting and multi-rate alignment module configured to perform observation thresholding and dynamic weighting processing on the unified observation information frames, to generate observation credibility scores and measurement weights according to visibility scores, matching confidence and registration residuals, to eliminate abnormal observations, and to perform multi-rate time alignment with inertial measurement unit observation data, and to output time-sequentially consistent aligned observation sequences;

[0134] a motion prediction and state estimation module including an interacting multiple model motion prediction structure and an improved Kalman filter, configured to construct an interacting multiple model motion prediction set based on the aligned observation sequences, to generate prior prediction results, and to adaptively adjust measurement covariance, and to output terrain state estimates and terrain state estimation covariances;

[0135] a map representation and incremental update module configured to project the terrain state estimates and the terrain state estimation covariances to terrain map subgraphs of a signed distance field map and an elevation raster map. Embodiments

[0136] In a pipeline inspection task, an engineering team operates an AUV to measure dynamic terrain in a complex underwater environment. The AUV is equipped with multi-modal sensors such as optical cameras, multi-beam sonars, inertial measurement units, structured light modules, and depth gauges, and collects a large amount of optical images and sonar point cloud data. The environment background is characterized by low visibility, strong scattering, multi-path sonar echoes, and frequent water flow disturbances.

[0137] After the AUV enters the first section, the optical camera collects 10 frames of images per second. In the initial 30 minutes, the traditional method extracts an average of 180 SURF feature points per frame. Due to the reduction in image contrast caused by suspended sediment, the feature stable tracking rate is only 53%, and the inter-frame matching error rate in some areas reaches 21%. In contrast, the Vision Mamba perception backbone network synchronously collects image frames, automatically extracts an average of 305 multi-scale fusion feature points per frame, and the feature stable tracking rate is increased to 89%. The average system visibility score is 0.74, which is 0.18 higher than the traditional method. At this time, the AUV passes through a sand dune obstacle, and the local multipath echo in the sonar point cloud is enhanced. Under the traditional scheme, the corresponding point cloud frame anomaly score is as high as 0.29. The observation points with a matching confidence lower than 0.55 output by Vision Mamba are automatically marked and removed by the system. The observation points with a confidence higher than 0.8 are dynamically given a high weight to participate in state estimation.

[0138] When the AUV passes through the rock slope area, due to the attitude mutation and asynchronous sensors, the traditional method's Kalman filter accumulates drift to 17 cm in 15 minutes, and 193 abnormal observation points are removed. The average error of abnormal observations is 11 cm. After using the improved Kalman filter of the application, the measurement covariance is dynamically adjusted according to the visibility score, matching confidence, and registration residual output by the current multi-modal perception. In the same period, the cumulative drift is limited to within 5 cm, there are only 58 abnormal observation points, and the average error of abnormal observations is 3 cm.

[0139] As the AUV enters the second section of the pipeline covered area, a short-term current mutation occurs on site, and the IMU data abnormally fluctuates with a peak value of 1.7 degrees per second. In a loop closure identification process, the traditional method's map subgraph stitching position deviates, with a maximum height difference of 12 cm; using the map loop correction of the application, the height difference in the overlapping area of the historical and current subgraphs is stable within 2 cm.

[0140] In the third section, the AUV needs to conduct fine-grained mapping on the interface between the exposed pipeline surface and the sand pile. Under the traditional method, due to image blur and sonar multipath, the minimum elevation point cloud density between subgraphs is only 102 points per square meter, and the seamless rate of map stitching is 72%. Using the application, all available multi-modal observation data are used to generate high-confidence feature points through the Vision Mamba network, and the elevation point cloud density is increased to 158 points per square meter after dynamic filtering and updating the map subgraph. The seamless rate of map stitching is increased to 95%. The total area of the cumulative measurement area is 285 square meters. Under the traditional scheme, there are 14 spatial misplacement correction events, and the application method only needs to be adjusted 3 times. The maximum spatial misplacement distance is 3 cm, and the traditional method is 13 cm.

[0141] During the entire task period, the system records the performance comparison data of the two methods respectively:

[0142] The cumulative effective observation frame number of the present application is 67038, and the cumulative effective observation frame number of the traditional method is 48907.

[0143] The effective registration residual mean of the present application is 1.8 cm, and the effective registration residual mean of the traditional method is 6.2 cm.

[0144] The inter-subgraph cumulative stitching error standard deviation of the present application is 2.6 cm, and the inter-subgraph cumulative stitching error standard deviation of the traditional method is 8.1 cm.

[0145] The loop detection trigger number of the present application is 8, all of which are automatically closed and corrected, and the loop detection trigger number of the traditional method is 5, of which 2 are closed failure.

[0146] The terrain elevation standard deviation (taking the IMU and sonar joint calibration as the true value reference) of the present application is ±5 cm, and the terrain elevation standard deviation of the traditional method is ±17 cm.

[0147] 6000 groups of historical scene samples are used in the algorithm training and model convergence stage, wherein the present application method converges at an average of 1800 steps, and the RMSE is reduced to below 6 cm; the traditional method converges at 4200 steps, and the final RMSE is 14 cm, the feature abnormality rejection rate of the present application is 5.1%, and the feature abnormality rejection rate of the traditional method is 12.3%.

[0148] As can be seen from the whole process of embodiment 2, the present application can improve the utilization rate of high confidence observation points in complex scenes in real time, dynamically suppress the cumulative error caused by abnormal points, realize continuous and high-precision estimation of terrain state, and effectively improve the automation level of map updating and loop correction, better complete underwater terrain dynamic tracking measurement.

[0149] The above describes only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can make equivalent replacement or change according to the technical solution and the inventive concept of the present application within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A method of dynamically tracking a water bottom topography, comprising: The method comprises the following steps: Collecting underwater multi-source heterogeneous perception data, and calibrating and preprocessing the data to output calibrated and preprocessed underwater multi-source heterogeneous perception data sequences; Encoding the calibrated and preprocessed underwater multi-source heterogeneous perception data sequences into spatio-temporal Token sequences and inputting them into a Vision Mamba perception backbone network to obtain multi-scale multi-modal feature representations carrying visibility scores and matching confidence; Based on the obtained multi-scale multi-modal feature representations carrying visibility scores and matching confidence, geometric consistency registration is completed in the cross-modal alignment head of the Vision Mamba perception backbone network, and a unified observation information frame is constructed; Performing observation threshold and dynamic weighting processing on the unified observation information frame, and performing multi-rate time alignment between the observation threshold and dynamic weighting processed unified observation information frame and inertial measurement unit observation data to form an aligned observation sequence with consistent timing; Constructing an interactive multi-model motion prediction set on the aligned observation sequence, and introducing a sea current disturbance term to calculate state priors and covariance priors, and outputting prior prediction results; Fusing the observation threshold and dynamic weighting processed unified observation information frame and the prior prediction results in an improved Kalman filter, and adaptively adjusting the measurement covariance and the process covariance and performing consistency checking and robust loss estimation to obtain terrain state estimation and terrain state estimation covariance; Projecting the terrain state estimation and terrain state estimation covariance to a map representation to complete underwater terrain dynamic tracking measurement.

2. The method of claim 1, wherein, The construction of the calibrated and preprocessed underwater multi-source heterogeneous perception data sequence comprises: Collecting underwater multi-source heterogeneous perception data, which includes optical camera observation data, sonar observation data, structured light observation data, laser stripe observation data, and inertial measurement unit observation data; Performing defogging, de-coloring, and light equalization processing on the optical camera observation data; Performing polar to Cartesian domain transformation and intensity normalization processing on the sonar related observation data; Performing stripe extraction and geometric correction processing on the structured light observation data and laser stripe observation data; Performing denoising and interpolation padding processing on the inertial measurement unit observation data; Outputting the calibrated and preprocessed underwater multi-source heterogeneous perception data sequence.

3. The method of claim 1, wherein, The multi-scale multi-modal feature representation carrying visibility scores and matching confidence comprises: The calibrated and preprocessed underwater multi-source heterogeneous perception data sequence is block encoded according to a unified spatio-temporal structure, and each block is a spatio-temporal Token. All Tokens form a complete spatio-temporal Token sequence; For each Token at a specified time and spatial position in the spatio-temporal Token sequence, the spatio-temporal Token features at the current time and all previous times corresponding to the spatial position are fused to generate long-range temporal state space features; Each long-range temporal state space feature at a time is respectively pooled and convolved at different spatial scales using feature decomposition to obtain multi-scale feature representations, and each scale has a corresponding multi-modal feature representation; The multi-modal feature representation at each time, each spatial scale, and each spatial position is subjected to channel splicing and attention enhancement in the Vision Mamba perception backbone network to obtain a fused and enhanced multi-scale multi-modal feature representation. Based on the fused and enhanced multi-scale multi-modal feature representation, the visibility score and the matching confidence are respectively generated for each time, each spatial scale, and each spatial position by using a feature-to-score mapping structure. The fused and enhanced multi-scale multi-modal feature representation, the visibility score, and the matching confidence are combined to form a multi-scale multi-modal feature representation carrying the visibility score and the matching confidence.

4. The method of claim 1, wherein, The construction of the unified observation information frame includes: The multi-scale multi-modal feature representation carrying the visibility score and the matching confidence is input into the cross-modal alignment head in the Vision Mamba perception backbone network, and the feature representations of the optical camera observation channel, the sonar observation channel, and the inertial measurement unit observation channel are respectively extracted for each time, each spatial scale, and each spatial position, and the feature representations are mapped to a unified geometric space. In the unified geometric space, the distance between the feature representations of the optical camera observation channel, the sonar observation channel, and the inertial measurement unit observation channel at the same time and the same spatial position is taken as a registration residual based on a spatial position consistency constraint. Based on the registration residual at each time, each spatial scale, and each spatial position, the feature representations of the optical camera observation channel, the sonar observation channel, and the inertial measurement unit observation channel at the same spatial position are dynamically adjusted to obtain a dynamically adjusted unified feature representation. The dynamically adjusted unified feature representation, the time index, the spatial index, the visibility score, the matching confidence, and the registration residual are taken as observation attributes to construct a unified observation information frame.

5. The method of claim 1, wherein, The formation of the time-consistent aligned observation sequence includes: The visibility score, the matching confidence, and the registration residual of each observation position at a specified time, spatial scale, and spatial position in the unified observation information frame are extracted. The visibility score and the matching confidence of each observation position are compared, and the smaller value of the two is taken as an observation confidence score. The observation confidence score and the registration residual of each observation position are combined to calculate the final measurement weight of the spatial position. An observation weight threshold is set, and when the final measurement weight of each observation position is lower than the observation weight threshold, the observation is determined to be abnormal and is removed. All low-measurement-weight observations are marked as invalid or do not participate in subsequent processing in the unified observation information frame, and only valid observation positions with a weight greater than or equal to the threshold are retained. The unified observation information frame after removing the abnormal observation and the inertial measurement unit observation data are subjected to time alignment processing to generate observation pairs, and all observation pairs constitute a time-consistent aligned observation sequence.

6. The method of claim 1, wherein, The output of the prior prediction result includes: Based on the time sequence consistent alignment observation sequence, an interacting multiple model motion prediction set is constructed, and the interacting multiple model motion prediction set includes a constant speed motion model, a constant acceleration motion model and a coordinated turning motion model; The constant speed motion model, the constant acceleration motion model and the coordinated turning motion model all take the state estimation value of the previous moment as the input, respectively perform linear transformation through the corresponding state transition matrix, superimpose the corresponding sea current disturbance term to reflect the influence of the sea current on the target state in the actual underwater environment, and introduce process noise subject to zero mean Gaussian distribution to obtain the state prior of each motion model at the current moment; For all motion models in the interacting multiple model motion prediction set, the model probability of each motion model is used as a weighting coefficient to weight and fuse the state prior and process noise of each model at the current moment to obtain the fused state prior and covariance prior at the current moment as the prior prediction result of the output.

7. The method of claim 1, wherein, The terrain state estimation including terrain elevation, terrain curvature, terrain feature point and carrier pose and the terrain state estimation covariance thereof are obtained, including: The unified observation information frame after the observation threshold and the dynamic weighting processing and the prior prediction result are jointly input into the improved Kalman filter; For each spatial position, different adjustment coefficients are respectively given according to the visibility score, the matching confidence and the registration residual, and are weighted and superimposed to obtain the measurement covariance of the current spatial position; The measurement covariance of each spatial position and the prior prediction result are jointly input into the improved Kalman filter update, and the mapping relationship of the state space and the observation space is combined with the observation matrix to calculate the filtering gain at the current moment; The difference between the observation vector in the unified observation information frame and the state prediction is calculated as the quantity residual; The Huber loss function is applied to the quantity residual for nonlinear weighting, the Huber loss function is used to suppress the influence of extreme observation errors on the filtering result, the weighted quantity residual is updated through the filtering gain to obtain the terrain state estimation at the current moment, and the covariance is also updated synchronously; The terrain state estimation output at the current moment is calculated as the geometric variables and attitude variables of the terrain elevation, terrain curvature, terrain feature point and carrier pose, and the terrain state estimation covariance at the current moment is retained.

8. A dynamic tracking measurement system for underwater terrain, for performing a dynamic tracking measurement method for underwater terrain according to any one of claims 1 to 7, characterized in that It includes: A data acquisition and synchronization module is used for acquiring underwater multi-source heterogeneous sensing data, and completing data calibration and preprocessing under a unified time reference and a unified coordinate system, and outputting calibrated and preprocessed underwater multi-source heterogeneous sensing data sequences; A perception coding and feature modeling module includes a Vision Mamba perception backbone network, which is used for coding the calibrated and preprocessed underwater multi-source heterogeneous sensing data sequences into spatio-temporal Token sequences, performing multi-scale multi-modal feature modeling based on a selective structured scanning mechanism, and outputting multi-scale multi-modal feature representations carrying visibility scores and matching confidences; The cross-modal registration and unified observation information frame construction module is configured to implement geometric consistency registration of the optical camera observation channel, the sonar observation channel and the inertial measurement unit observation channel in the cross-modal alignment head based on multi-scale multi-modal feature representations carrying visibility scores and matching reliabilities, and generate a unified observation information frame; The dynamic weighting and multi-rate alignment module is configured to perform observation thresholding and dynamic weighting processing on the unified observation information frame, generate observation credibility scores and measurement weights according to the visibility scores, the matching reliabilities and the registration residuals, eliminate abnormal observations, and perform multi-rate time alignment with the inertial measurement unit observation data, and output an aligned observation sequence with consistent timing; The motion prediction and state estimation module includes an interacting multiple model motion prediction structure and an improved Kalman filter, and is configured to construct an interacting multiple model motion prediction set based on the aligned observation sequence, generate a prior prediction result, and adaptively adjust a measurement covariance, and output a terrain state estimation and a terrain state estimation covariance; The map expression and incremental update module is configured to project the terrain state estimation and the terrain state estimation covariance to a terrain map subgraph of a signed distance field map and an elevation raster map.

Citation Information

Patent Citations

  • Underwater terrain matching and positioning method based on Gaussian process regression learning

    CN111220146A

  • Underwater multi-source detection data normalization method

    CN114355474A