Underwater autonomous mapping method based on target detection and deep semantic key point extraction
By combining the underwater SLAM system of trinocular vision sensors and active sonar, using adaptive corner extraction and YOLACT dynamic object culling, and optimizing feature point matching, the problems of map construction accuracy and dynamic object interference in underwater environments are solved, achieving high-precision three-dimensional map construction and autonomous navigation.
Patent Information
- Application Number
- CN202510454346.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-09-09
AI Technical Summary
Traditional SLAM technology has difficulty in effectively building accurate maps in underwater environments. It is limited by light attenuation, interference from complex underwater environments, and misjudgment of dynamic objects. It consumes a lot of computing resources and suffers from severe positioning drift.
The trinocular vision sensor and active sonar are combined, and the extended Kalman filter algorithm is integrated to extract environmental features. The adaptive corner point extraction and YOLACT algorithm are combined to remove the interference of dynamic objects. The DISK algorithm is used to optimize feature point matching. The multimodal feature robust descriptor is used to perform cross-modal comparative learning and elastic feature aggregation.
It significantly improves the accuracy of underwater positioning and map construction, reduces interference from dynamic objects, enhances the robustness of the system in complex underwater environments, and constructs high-precision three-dimensional maps to provide support for autonomous navigation and path planning.
Smart Images

Figure CN120612439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a map construction method for an autonomous system in a complex environment, and in particular to an underwater autonomous mapping method based on target detection and depth semantic key point extraction. Background Art
[0002] Underwater operations and maintenance tasks in marine engineering are more complex than those on land. The emergence of autonomous underwater vehicles (AUVs) has greatly improved the efficiency of exploration, development, and utilization of marine resources. AUVs can replace humans in certain complex tasks, providing technical support to those on land. However, in underwater environments, the performance of some effective land-based sensors, such as lidar and optical cameras, is significantly reduced due to limitations in their principles, such as underwater light attenuation. Complex currents, waves, and terrain often exist underwater, placing higher demands on the motion control of AUVs. In particular, the presence of complex terrain, such as shipwrecks and reefs, complicates the research and application of underwater systems. Traditional monocular vision sensors cannot obtain sufficient information due to factors such as light attenuation, underwater light spots, shadows, and varying water quality. Furthermore, due to the lack of depth information, traditional SLAM technology is difficult to apply underwater.
[0003] Traditional SLAM algorithms typically estimate the system's posture and environmental map through filtering or optimization methods. Over long periods of time, errors can easily accumulate, leading to a decrease in posture and map accuracy. Furthermore, traditional SLAM processes are computationally intensive, such as data matching, filtering, nonlinear optimization, data association, and optimization. These processes take a long time to compute, consume significant computing resources, and pose challenges in ensuring real-time performance. Traditional SLAM methods typically assume the system's environment is static—simply assuming the surrounding environment remains unchanged as the system moves. Without an effective mechanism for filtering out dynamic features, dynamic objects can be easily mistaken for static features of the environment, leading to poor map construction and positioning accuracy. This can cause the system to produce inconsistent maps and even cause positioning drift after long periods of operation, significantly reducing task execution efficiency and reliability. Traditional methods often consume significant computing resources to address these issues. Therefore, a new approach that can effectively address these issues is needed. Summary of the Invention
[0004] Purpose of the invention: In order to solve the problems of high difficulty and complex computation load in constructing maps by autonomous systems in underwater environments in the prior art, the present invention proposes an underwater autonomous mapping method based on target detection and depth semantic key point extraction.
[0005] Technical solution: The present invention comprises the following steps:
[0006] (1) The autonomous system collects environmental information through trinocular vision sensors and active sonar;
[0007] (2) The three-dimensional image is obtained by trinocular vision, the depth data is obtained by sonar, and the fusion extended Kalman filter algorithm MEK-SLAM is used to fuse the data, extract environmental features, and estimate the system position and attitude;
[0008] (3) Extract ORB feature points through adaptive corner point extraction threshold, and update the map through feature point extraction and matching;
[0009] (4) During the map update process, the YOLACT algorithm is used to detect and segment dynamic objects in the image to remove the interference of dynamic objects on map construction;
[0010] (5) Based on the DISK algorithm, the matching and descriptor generation of feature points are optimized. To address the problem of difficult stable matching between sonar point clouds and visual features, multimodal feature robust descriptors are used to perform cross-modal contrast learning and elastic feature aggregation to improve the accuracy of image registration.
[0011] Furthermore, the trinocular vision sensor in step (1) includes three cameras and a controller, and the three cameras form an equilateral triangle arrangement to simultaneously collect two-dimensional image data from three different angles; the active sonar provides depth information of the surrounding environment to the autonomous system by emitting sound waves and receiving echo information.
[0012] Furthermore, the implementation process of step (2) is as follows:
[0013] A two-branch network is constructed using a combined physics-deep learning framework: transmittance maps and background light estimation are calculated based on underwater optical imaging equations; a conditional GAN is used to generate detail enhancement results; and a differentiable rendering layer is used to achieve joint optimization of the two branches, restoring details while maintaining physical plausibility.
[0014] Design a spectrally sensitive convolution kernel and dynamically adjust the color correction matrix according to the water type to solve the problem of cross-depth color shift;
[0015] A filter is used to remove noise; a camera distortion is corrected using a camera intrinsic parameter matrix K and a camera distortion coefficient D; and a histogram equalization method and other methods are used to improve the contrast and visibility of the image; the camera intrinsic parameter matrix K includes information about the focal length and the optical center position; and the distortion coefficient D corrects the radial and tangential distortion of the image.
[0016] The system's position and attitude are updated in real time. The pose estimation is used to ensure the data accuracy of subsequent feature extraction and environment mapping by MEK-SLAM, eliminating the impact of positioning errors on feature extraction.
[0017] Furthermore, the process of extracting ORB feature points by using the adaptive corner point extraction threshold in step (3) is as follows:
[0018] Dynamically adjust the corner point extraction threshold based on the local information of the image, and use the local gradient and brightness difference of the image to adjust the corner point extraction threshold:
[0019] T(x,y)=mean(I(x,y))+k·std(I(x,y))(1)
[0020] After detecting the corner points, use the BRIEF descriptor to describe each corner point and generate a binary vector to represent the local area of the feature point;
[0021] The FAST algorithm is used to assign a direction to each corner point according to the gradient direction; the BRIEF descriptor is used for encoding, and the descriptor is invariant to rotation.
[0022] Furthermore, the map updating process by extracting and matching feature points in step (3) is as follows:
[0023] Use descriptors for matching, use brute force matching algorithm for feature matching, and calculate the similarity between feature points:
[0024]
[0025] The matched feature points are compared with the points in the existing map, and the location and feature information in the map are updated; if the new feature points have a good match with some points in the existing map, the map point cloud is updated.
[0026] Furthermore, the implementation process of step (4) is as follows:
[0027] The YOLACT algorithm performs instance-level segmentation through CNN to generate a segmentation mask for each object. It first extracts image features through the backbone network, then predicts the bounding box of each object and the segmentation mask of the object.
[0028] Motion consistency verification is performed using time-series-aware dynamic object filtering: a three-stage filtering mechanism is constructed to calculate the 3D position of objects through multi-view parallax, eliminating objects that do not conform to the rigid body motion assumption. The RAFT optical flow network is used to detect the difference between object motion and background flow fields. A Transformer-based trajectory prediction module is used to establish a cross-frame motion logic chain.
[0029] Bayesian filtering is used to continuously update the static probability value of each area, guiding the SLAM system to dynamically adjust the feature point selection strategy;
[0030] Combined with self-supervised learning, the segmentation accuracy of dynamic objects is optimized; self-supervised learning automatically adjusts the network weights through unsupervised or weakly supervised methods to adapt to feature point and object detection in dynamic scenes;
[0031] The mask generated by YOLACT is used to eliminate the area of dynamic objects and retain only the feature points in the static environment, ensuring that the SLAM system only performs positioning and map construction in the static environment.
[0032] Furthermore, the implementation process of step (5) is as follows:
[0033] DISK uses CNN to detect key points in images and optimizes feature point detection by training the network. CNN captures structural information in images from local to global perspectives through multi-level feature maps.
[0034] For each key point, the image information of the local area is used to generate a descriptor; for each feature point P i =(x i ,y i ), descriptor d generated by DISK i Expressed as:
[0035] d i =DISK(LocalRegion(I,P i )) (3)
[0036] Match feature points in an image using high-quality descriptors:
[0037]
[0038] Among them, d i and d j is the descriptor of two feature points || d i || and ||d j || is their norm;
[0039] Then, cross-modal contrastive learning is performed and a two-stream network is designed using the improved InfoNCE loss function, where the InfoNCE loss function is expressed as:
[0040]
[0041] where v i 、s j Enforces modal invariance for learning visual and sonar features at the same location;
[0042] Elastic feature aggregation: Introducing a deformable convolutional network to dynamically adjust the receptive field based on local feature density, automatically expanding the aggregation range in areas with sparse sonar data;
[0043] Preliminary feature point pairs are obtained by matching feature points between two images. RANSAC is used to filter out incorrect matches and retain high-quality matching feature point pairs. The relative pose between the images is estimated using the camera pose estimation algorithm based on the matched feature point pairs.
[0044] The matching point pair of the two images is (P1, P2), and the pose estimation formula is:
[0045] P2=RP1+T(6)
[0046] Where P1 and P2 are points in the camera coordinate system;
[0047] Based on the estimated preliminary pose, the estimated pose is optimized using bundle adjustment:
[0048]
[0049] Reduce registration error and reprojection error, and improve image registration accuracy.
[0050] Furthermore, the dual-stream network is a visual stream network and a sonar stream network.
[0051] Beneficial effects: Compared with the existing technology, the beneficial effects of the present invention are: compared with the traditional SLAM system, this method gives full play to the advantages of multi-sensor data fusion in the underwater environment, and significantly improves the accuracy of underwater positioning and map construction; the present invention combines trinocular vision with sonar, and can obtain more accurate environmental information in underwater low-visibility environments. MEK-SLAM can obtain high-precision pose estimation and data fusion, thereby ensuring more accurate positioning, which has obvious advantages when facing complex underwater terrain; in response to the problems of color distortion, atomization effect and low-contrast interference in underwater images, a visual enhancement algorithm guided by physical models is proposed, and a dual-branch network is constructed using a joint physics-deep learning framework; adaptive wavelength compensation: a spectrally sensitive convolution kernel is designed, and the color correction matrix is dynamically adjusted according to the water type to solve the problem of cross-depth color shift; the high-low fusion method combining trinocular vision and sonar can effectively reduce the interference of underwater dynamic objects and improve the robustness of the system in complex underwater dynamic environments; in response to the problem that traditional instance segmentation is prone to misjudgment between fast-moving objects and semi-static objects, motion consistency is proposed Verification, construct a three-stage filtering mechanism, dynamic-static probability map, continuously update the static probability value of each area through Bayesian filtering, and guide the SLAM system to dynamically adjust the feature point selection strategy; use adaptive corner point extraction and ORB feature point extraction combined with YOLACT dynamic object detection and segmentation to effectively remove the interference of underwater dynamic objects and improve the robustness of the system in underwater dynamic environments; to address the problem of difficult stable matching of sonar point clouds and visual features, cross-modal contrast learning is proposed, and a dual-stream network (visual stream + sonar stream) is designed; elastic feature aggregation, introduces a deformable convolutional network, dynamically adjusts the receptive field according to the local feature density, and automatically expands the aggregation range in sparse sonar data areas; uses the DISK algorithm to optimize feature point matching and descriptor calculation and generation, improves the image registration accuracy in underwater environments, and reduces the accumulation of registration errors caused by water flow, lighting or noise; through the above optimization, this method can overcome the interference of complex underwater environments and construct a more accurate three-dimensional map, which can provide strong support for underwater autonomous navigation and path planning, especially in unknown underwater environments and when dynamic changes are large. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a flow chart of the present invention;
[0053] Figure 2 This is a schematic diagram of the installation position of the trinocular vision sensor;
[0054] Figure 3 This is a schematic diagram of the workflow of the YOLACT instance segmentation algorithm in underwater autonomous systems;
[0055] Figure 4 This is a schematic diagram of the workflow of the DISK algorithm in the SLAM system. DETAILED DESCRIPTION
[0056] The present invention will be described in further detail below with reference to the accompanying drawings.
[0057] like Figure 1 As shown in the figure, this paper proposes an underwater autonomous mapping method based on target detection and deep semantic key point extraction. It uses adaptive corner point extraction, YOLACT dynamic object removal, and DISK optimized feature point matching to enhance the system's robustness to dynamic environments and the accuracy of 3D map construction, thereby improving the accurate guidance of autonomous navigation and path planning. The specific steps are:
[0058] Step 1: The autonomous system collects environmental information through trinocular vision sensors and active sonar.
[0059] like Figure 2 As shown in the figure, the trinocular vision sensor consists of three cameras and a controller. The three cameras form an equilateral triangle arrangement, with one camera located in the center, as shown by O3 in the figure. Object P appears as P3 in its image. The other two cameras are located on either side, as shown by O1 and O2 in the figure. Object P appears as P1 and P2 in their respective images. It can provide multiple two-dimensional images collected simultaneously from three different angles, which can not only capture environmental texture but also provide visual information for subsequent feature point selection. Active sonar provides depth information of the surrounding environment to the autonomous system by emitting sound waves and receiving echo information. It is more effective in environments with poor lighting or insufficient visual information.
[0060] Step 2: After obtaining the sensor data, the Mixed Extended Kalman Filter Simultaneous Localization and Mapping (MEK-SLAM) algorithm is used to fuse these sensor data. However, the sensor data needs to be preprocessed before data fusion to ensure the subsequent processing process.
[0061] Since the data obtained by the sensors are affected by factors such as noise and distortion during the process of obtaining sensor data, certain measures need to be taken to correct these sensor data before performing sensor data fusion.
[0062] First, a two-branch network is constructed using a joint physics-deep learning framework. The network calculates transmittance maps and background light estimates based on underwater optical imaging equations. A conditional GAN (generative adversarial network) is then used to generate detail enhancement results. A differentiable rendering layer is used to achieve joint optimization of the two branches, restoring detail while maintaining physical plausibility.
[0063] The spectrally sensitive convolution kernel is redesigned to dynamically adjust the color correction matrix based on the water type (blue / green light dominance) to address cross-depth color shift. Before sensor data fusion, the MEK-SLAM algorithm also requires image denoising. Image noise can affect subsequent feature extraction and pose estimation. Therefore, a filter is used to smooth the image and remove the noise. During the data fusion process, image data and sonar data are calibrated based on timestamps to ensure spatiotemporal consistency between the two data.
[0064] There is a certain spatial deviation between the two-dimensional image data provided by the trinocular vision sensor and the depth data provided by the sonar. However, MEK-SLAM accurately fuses vision and sonar to extract key environmental features and estimate the system's position and attitude. The core purpose of this fusion is to extract environmental features and update the system's attitude in real time to ensure the stability of subsequent feature extraction and environmental mapping. Camera distortion correction is then performed. Due to the physical effects of the camera lens, images captured by a camera will have a certain degree of distortion. Wide-angle lenses are particularly prone to distortion at the edges of the image. To ensure the accuracy of the captured image data, correction is required using the camera's intrinsic parameter matrix K and distortion coefficient D. The camera's intrinsic parameter matrix K includes information such as focal length and optical center position. The distortion coefficient D is used to correct radial and tangential image distortion. These two parameters can be used to correct image distortion and produce an ideal, distortion-free image.
[0065] Image enhancement is another preprocessing step. Image contrast, brightness, and other information often vary depending on the environment, affecting image visibility and information content. To improve image quality, MEK-SLAM uses histogram equalization to enhance image contrast. Histogram equalization redistributes the image's grayscale values, making the image brightness distribution more uniform. This enhances image detail and increases visibility, facilitating subsequent feature extraction.
[0066] After the data preprocessing is completed, the system continues to use the MEK-SLAM algorithm for pose estimation. Pose estimation refers to the position and orientation of the system in three-dimensional space relative to a reference coordinate system. MEK-SLAM uses a nonlinear optimization algorithm to estimate and update the fused three-eye vision data and sonar data in real time, wherein MEK-SLAM uses the improved MEK technology of the present invention to update the pose in real time. MEK-SLAM corrects the system pose in real time by jointly estimating visual data and sonar data. In a more complex dynamic environment, it can quickly and reduce the pose error of the system, ensuring that the system can always correctly perceive its current position in the environment in each process. At the same time, pose estimation has very high requirements on the quality of sensor data and the coordination between sensors. Therefore, the MEK-SLAM algorithm can minimize the inconsistency between sensor data during the real-time data fusion process, making the system positioning more robust and accurate. The MEK-SLAM algorithm provides good support for subsequent feature extraction and environment mapping through the accuracy of pose estimation. When using feature extraction, pose estimation can provide the system's location information. In this sub-process, the error points caused by the deviation of the system positioning will also have a certain impact on the extraction work. Therefore, pose error is a common problem in SLAM systems. In a dynamic environment using laser positioning, it is easy to cause system pose estimation errors due to the influence of factors such as sensor noise and object movement. To avoid this situation, MEK-SLAM updates the system's pose in real time, providing high-precision positioning data for subsequent feature extraction during the environment mapping process; on the other hand, MEK-SLAM updates the pose in real time to avoid the accumulation of errors in the transmission of environmental information during the map construction process; traditional SLAM systems are based on the repeated accumulation of positioning errors, which makes the map construction process prone to deviations. MEK-SLAM avoids error accumulation through real-time pose estimation and data fusion, further constructs a high-precision three-dimensional map, and thus improves the system's navigation and path planning capabilities.
[0067] Step 3: ORB (Oriented FAST and Rotated BRIEF) features, which are composed of FAST (Features from Accelerated Segment Test) corner points and BRIEF (Binary Robust Independent Elementary Features) descriptors, are processed using the MEK-SLAM method, that is, ORB is extracted and matched using an adaptive corner extraction threshold, and the map is updated.
[0068] First, the pre-processed image data is used to extract corner points using an adaptive corner point extraction threshold. The corner point extraction threshold is dynamically adjusted based on the local information of the image. Generally, the local gradient and brightness difference of the image are used to adjust the corner point extraction threshold:
[0069] T(x,y)=mean(I(x,y))+k·std(I(x,y))(1)
[0070] After detecting the corner points, the BRIEF descriptor is used to describe each corner point and a binary vector is generated to represent the local area of the feature point.
[0071] Combining FAST and BRIEF, ORB provides feature points. Using the FAST algorithm, each corner point is assigned a direction based on the gradient direction. BRIEF descriptors are used for encoding, and the descriptors are invariant to rotation.
[0072] Use descriptors for matching, use brute force matching algorithm for feature matching, and calculate the similarity between feature points:
[0073]
[0074] The matched feature points are compared with the existing points in the map, and the positions and features on the map are updated. If the new feature points match some points in the existing map well, the map point cloud is updated.
[0075] Step 4: If Figure 3 As shown in the figure, during map updates, the YOLACT instance segmentation algorithm is used to detect and segment dynamic objects in the image. YOLACT (You Only LookAt Coefficients Ts, a real-time instance segmentation method) uses self-supervised learning to eliminate the interference caused by dynamic objects to MEK-SLAM, allowing MEK-SLAM to use only feature points in the static environment data for positioning and map construction.
[0076] YOLACT uses CNN (Convolutional Neural Networks) to achieve instance-level segmentation of each object and generate a segmentation mask for each object. The convolutional neural network backbone first obtains image features, and then predicts the bounding box and segmentation mask for each object.
[0077] Motion consistency verification is performed using time-series-aware dynamic object filtering: a three-stage filtering mechanism is constructed to calculate the three-dimensional position of objects through multi-view parallax, exclude targets that do not conform to the rigid body motion assumption, use the RAFT optical flow network to detect the difference between object motion and background flow field, and establish a cross-frame motion logic chain based on the Transformer trajectory prediction module; and continuously update the static probability value of each area through Bayesian filtering to guide the SLAM system to dynamically adjust the feature point selection strategy.
[0078] And through self-supervised learning, the object segmentation accuracy is optimized. Self-supervised learning is unsupervised or weakly supervised learning, which automatically updates the network weights without supervision or weak supervision to adapt to the extraction of feature points and objects in dynamic scenes.
[0079] The mask generated by YOLACT is used to remove dynamic objects in the area, retaining only the feature points in the static environment, ensuring that SLAM is only positioned and mapped in the static environment.
[0080] Step 5: Figure 4 As shown in the figure, DISK (Deep Image Structure Keypoints, a feature detection and description method based on deep learning) is used to optimize feature point matching and descriptor generation in the SLAM system, providing more accurate feature points and descriptors and improving the accuracy of image registration.
[0081] MEK-SLAM uses optimized feature points for pose estimation and image registration, enhancing the accuracy of the constructed 3D map. DISK provides high-quality feature point matching, helping MEK-SLAM reduce cumulative errors in complex environments, ultimately obtaining a high-precision 3D underwater map for autonomous navigation and path planning. The specific implementation steps are as follows:
[0082] DISK uses a deep convolutional neural network to locate keypoints. Its main idea is to train and optimize feature point detection. CNNs capture structural information by transforming local features of an image into global features. Given an image keypoint, DISK optimizes the image end-to-end based on the trained network to obtain the confidence level of each point.
[0083] For each key point, the DISK algorithm uses the local area image information to generate a descriptor. The DISK descriptor is more robust and effective than the previous BRIEF or OROB. i =(x i ,y i ), descriptor d generated by DISK i It can be expressed as:
[0084] d i=DISK(LocalRegion(I,P i ))(3)
[0085] Using high-quality descriptors d i To match feature points in an image:
[0086]
[0087] Among them, d i and d j is the descriptor of two feature points || d i || and ‖d j ‖ is their norm.
[0088] We then conduct cross-modal contrastive learning and design a two-stream network (visual stream + sonar stream) using the improved InfoNCE loss function. The InfoNCE loss function can be expressed as:
[0089]
[0090] where v i 、s j Enforces modality invariance for learning visual and sonar features at the same location.
[0091] Elastic feature aggregation: Introducing a deformable convolutional network to dynamically adjust the receptive field based on local feature density, automatically expanding the aggregation range in areas with sparse sonar data;
[0092] The high-quality feature points and descriptors provided by DISK improve the accuracy of the next step of image registration. The image registration steps are as follows:
[0093] For two images (e.g., I1 and I2), preliminary feature point pairs (P1, P2) are obtained through feature point matching. RANSAC (Random Sample Consensus) is used to filter out false matches and retain high-quality matching feature point pairs.
[0094] The relative pose between the images is estimated using the camera pose estimation algorithm through the matched feature point pairs. Assuming that the matching point pairs of the two images are (P1, P2), where P1 and P2 are points in the camera coordinate system, the pose estimation formula is:
[0095] P2=RP1+T(6)
[0096] Use bundle adjustment to optimize the initially estimated pose, reduce registration error and reprojection error, and further improve the accuracy of image registration:
[0097]
[0098] By dynamically updating the 3D point cloud in the map, the MEK-SLAM system generates a complete 3D point cloud of the environment, continuously improving the map information and accumulating new environmental point clouds, so that the MEK-SLAM system can form a complete underwater 3D map. Each matching and estimation is based on optimizing the current map to reduce error accumulation.
[0099] Since the high-quality feature points and descriptors provided by the DISK algorithm are more accurate, they can effectively reduce the error accumulation in the SLAM system.
[0100] With the continuous dynamic updating of the 3D point cloud in the map, the occurrence of mismatches during the 3D map point cloud generation and matching process is reduced, reducing the errors caused by mismatches. Through precise pose estimation and random sampling consensus algorithms, errors caused by dynamic objects, noise, or inaccurate matching are reduced. Utilizing accurately estimated high-quality feature points and poses, the DISK algorithm helps MEK-SLAM reduce error accumulation and generate more accurate maps. By effectively generating feature point matching and descriptors, DISK's optimized feature point and descriptor matching and descriptor generation help the MEK-SLAM system generate high-precision 3D underwater maps, providing a reliable basis for autonomous navigation and path planning.
[0101] The present invention is an autonomous system map construction method based on underwater target detection, combined with real-time target separation and depth semantic key points. It adopts adaptive corner point extraction and ORB feature point extraction as well as YOLACT dynamic object detection and segmentation to eliminate dynamic objects and improve the robustness of the system in dynamic environments. It uses DISK to optimize feature point matching algorithm and descriptor generation, reduces the cumulative error of image registration, and improves the image registration accuracy, thereby improving the system's three-dimensional map construction accuracy in complex and dynamic environments, providing effective guarantees for autonomous navigation and path planning.
[0102] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An underwater autonomous mapping method based on target detection and deep semantic key point extraction, characterized in that: The following steps are involved: (1) The autonomous system collects environmental information through trinocular vision sensors and active sonar; (2) The three-dimensional image is obtained by trinocular vision, the depth data is obtained by sonar, and the fusion extended Kalman filter algorithm MEK-SLAM is used to fuse the data, extract environmental features, and estimate the system position and attitude; (3) Extract ORB feature points through adaptive corner point extraction threshold, and update the map through feature point extraction and matching; (4) During the map update process, the YOLACT algorithm is used to detect and segment dynamic objects in the image to remove the interference of dynamic objects on map construction; (5) Based on the DISK algorithm, the matching and descriptor generation of feature points are optimized. To address the problem of difficult stable matching between sonar point clouds and visual features, multimodal feature robust descriptors are used to perform cross-modal contrast learning and elastic feature aggregation to improve the accuracy of image registration.
2. The underwater autonomous mapping method based on target detection and deep semantic key point extraction according to claim 1 is characterized in that: The three-eye vision sensor in step (1) includes three cameras and a controller. The three cameras form an equilateral triangle arrangement and simultaneously collect two-dimensional image data from three different angles.
3. The underwater autonomous mapping method based on target detection and deep semantic key point extraction according to claim 1 is characterized in that: The active sonar in step (1) transmits sound waves and receives echo information to provide the autonomous system with depth information of the surrounding environment.
4. The underwater autonomous mapping method based on target detection and deep semantic key point extraction according to claim 1 is characterized in that: The implementation process of step (2) is as follows: A two-branch network is constructed using a combined physics-deep learning framework: transmittance maps and background light estimation are calculated based on underwater optical imaging equations; a conditional GAN is used to generate detail enhancement results; and a differentiable rendering layer is used to achieve joint optimization of the two branches, restoring details while maintaining physical plausibility. Design a spectrally sensitive convolution kernel and dynamically adjust the color correction matrix according to the water type to solve the problem of cross-depth color shift; Use filters to remove noise; use the camera intrinsic parameter matrix K and the camera distortion coefficient D to correct camera distortion; use histogram equalization to improve image contrast and visibility; The system's position and attitude are updated in real time. The pose estimation is used to ensure the data accuracy of subsequent feature extraction and environment mapping by MEK-SLAM, eliminating the impact of positioning errors on feature extraction.
5. The underwater autonomous mapping method based on target detection and deep semantic key point extraction according to claim 4 is characterized in that: The camera intrinsic parameter matrix K includes focal length and optical center position information; the distortion coefficient D corrects the radial and tangential distortion of the image.
6. The underwater autonomous mapping method based on target detection and deep semantic key point extraction according to claim 1 is characterized in that: The process of extracting ORB feature points by adaptive corner point extraction threshold in step (3) is as follows: Dynamically adjust the corner point extraction threshold based on the local information of the image, and use the local gradient and brightness difference of the image to adjust the corner point extraction threshold: T(x,y)=mean(I(x,y))+k·std(I(x,y))(1) After detecting the corner points, use the BRIEF descriptor to describe each corner point and generate a binary vector to represent the local area of the feature point; The FAST algorithm is used to assign a direction to each corner point according to the gradient direction; the BRIEF descriptor is used for encoding, and the descriptor is invariant to rotation.
7. The underwater autonomous mapping method based on target detection and deep semantic key point extraction according to claim 1 is characterized in that: The process of updating the map by extracting and matching feature points in step (3) is as follows: Use descriptors for matching, use brute force matching algorithm for feature matching, and calculate the similarity between feature points: Compare the matched feature points with the points in the existing map and update the location and feature information in the map; If the new feature point has a good match with some points in the existing map, the map point cloud is updated.
8. The underwater autonomous mapping method based on target detection and deep semantic key point extraction according to claim 1 is characterized in that: The implementation process of step (4) is as follows: The YOLACT algorithm performs instance-level segmentation through CNN to generate a segmentation mask for each object. It first extracts image features through the backbone network, then predicts the bounding box of each object and the segmentation mask of the object. Motion consistency verification is performed using time-series-aware dynamic object filtering: a three-stage filtering mechanism is constructed to calculate the 3D position of objects through multi-view parallax, eliminating objects that do not conform to the rigid body motion assumption. The RAFT optical flow network is used to detect the difference between object motion and background flow fields. A Transformer-based trajectory prediction module is used to establish a cross-frame motion logic chain. Bayesian filtering is used to continuously update the static probability value of each area, guiding the SLAM system to dynamically adjust the feature point selection strategy; Combined with self-supervised learning, the segmentation accuracy of dynamic objects is optimized; self-supervised learning automatically adjusts the network weights through unsupervised or weakly supervised methods to adapt to feature point and object detection in dynamic scenes; The mask generated by YOLACT is used to eliminate the area of dynamic objects and retain only the feature points in the static environment, ensuring that the SLAM system only performs positioning and map construction in the static environment.
9. The underwater autonomous mapping method based on target detection and deep semantic key point extraction according to claim 1 is characterized in that: The implementation process of step (5) is as follows: DISK uses CNN to detect key points in images and optimizes feature point detection by training the network. CNN captures structural information in images from local to global perspectives through multi-level feature maps. For each key point, a descriptor is generated using the image information of the local area; For each feature point P i =(x i ,y i ), descriptor d generated by DISK i Expressed as: d i =DISK(LocalRegion(I,P i )) (3) Match feature points in an image using high-quality descriptors: Among them, d i and d j is the descriptor of two feature points || d i || and ||d j || is their norm; Then, cross-modal contrastive learning is performed and a two-stream network is designed using the improved InfoNCE loss function, where the InfoNCE loss function is expressed as: where v i 、s j Enforces modal invariance for learning visual and sonar features at the same location; Elastic feature aggregation: Introducing a deformable convolutional network to dynamically adjust the receptive field based on local feature density, automatically expanding the aggregation range in areas with sparse sonar data; Preliminary feature point pairs are obtained by matching feature points between two images. RANSAC is used to filter out incorrect matches and retain high-quality matching feature point pairs. The relative pose between the images is estimated using the camera pose estimation algorithm based on the matched feature point pairs. The matching point pair of the two images is (P1, P2), and the pose estimation formula is: P2=RP1+T(6)where P1 and P2 are points in the camera coordinate system; Based on the estimated preliminary pose, the estimated pose is optimized using bundle adjustment: Reduce registration error and reprojection error, and improve image registration accuracy.
10. The underwater autonomous mapping method based on target detection and deep semantic key point extraction according to claim 9 is characterized in that: The dual-stream network is a visual stream network and a sonar stream network.
Citation Information
Cited By
Feature extraction and matching method and system, electronic equipment and medium
CN121147875A