Underwater SLAM Method Based on Dynamic Detection and IMU DVL Fusion
By introducing dynamic detection and IMUDVL fusion methods in SLAM technology, we can identify and eliminate dynamic underwater targets, and through multi-sensor optimization, the problem of large SLAM positioning error in far-reaching marine environments is solved, achieving higher positioning accuracy and robustness.
Patent Information
- Application Number
- CN202510513158.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing SLAM technology is susceptible to the influence of dynamic environments in complex dynamic environments in deep seas, resulting in large positioning errors, making it difficult to achieve reliable navigation and positioning.
The underwater SLAM method based on dynamic detection and IMUDVL is adopted, real-time object detection is carried out through the YOLOv11 model, dynamic objects are identified, and feature points are classified and eliminated in combination with the optical flow method to construct a factor graph for tight coupling optimization of multiple sensors.
It significantly improves the robustness and positioning accuracy of the system, solves the problem of traditional SLAM algorithms missing tracking or map drifting in dynamic environments, and provides reliable technical support for the stable navigation and positioning of AUVs in complex underwater environments in deep seas.
Smart Images

Figure CN120071115B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for simultaneous localization and mapping of an autonomous underwater vehicle (AUV), and particularly to an underwater SLAM method based on dynamic detection and IMU-DVL fusion. Background Art
[0002] Since the 21st century, with the increasing depletion of land resources, humans have shifted the focus of energy development and exploration to the ocean, especially the deep and far-reaching sea areas. Compared with the coastal waters, the underwater environment in the deep and far-reaching sea has significant differences: its water body is relatively clear due to being far from land-based pollution, but at the same time, it also faces unique challenges such as high pressure, low light, limited long-distance communication, frequent occurrence of dynamic targets (such as fish schools, diving operation equipment), and complex ocean current disturbances. As a new generation of underwater robots, AUVs have advantages such as safety and intelligence, and play a crucial role in exploring and developing deep-sea resources. Reliable navigation and positioning technology is a prerequisite for AUVs to complete tasks safely and efficiently.
[0003] As an advanced algorithm for simultaneous localization and mapping, ORB-SLAM3 has demonstrated superior real-time performance and robustness in many applications. However, its application in the dynamic scenarios of the deep and far-reaching sea faces multiple bottlenecks. First, the poor underwater lighting conditions, turbid water quality, and the presence of suspended matter make it difficult for visual sensors to obtain clear images. Second, although ORB-SLAM3 can perform well in static environments, its adaptability to dynamic environments is poor, and problems such as lost tracking or map drift often occur. Therefore, the current urgent problem to be solved is to optimize the ability of the SLAM system to handle dynamic environments by modifying the front-end and back-end links of ORB-SLAM3 for the complex underwater dynamic environment. Summary of the Invention
[0004] Object of the Invention: To solve the problem that existing SLAM technologies are vulnerable to dynamic environments and cause large positioning errors in the complex underwater environment of the deep and far-reaching sea, the present invention proposes an underwater SLAM method based on dynamic detection and IMU-DVL fusion.
[0005] Technical Solution: The underwater SLAM method based on dynamic detection and IMU-DVL fusion of the present invention includes the following steps:
[0006] An underwater SLAM method based on dynamic detection and IMU-DVL fusion, characterized in that the method includes the following steps:
[0007] S1. Train the YOLOv11 model to obtain a trained YOLOv11 model;
[0008] S2. Collect underwater images through an IDS industrial camera at a frequency of 10 Hz, input the current frame into the YOLOv11 model trained in step S1, and obtain the bounding box information of dynamic targets;
[0009] S3. For the current frame image input in step S2, use the superpoint model to extract feature points, and classify the feature points in the current frame image according to the bounding box information of the dynamic targets obtained in step S2: Feature points outside the bounding box are regarded as preliminary static feature points, and feature points inside the bounding box are regarded as dynamic feature points and excluded;
[0010] S4. For the preliminary static feature points of adjacent frames, calculate the optical flow of the feature points by introducing the Lucas-Kanade optical flow method improved by the camera model, set the residual norm threshold, exclude the optical flow of feature points with a residual norm greater than the threshold, and further save the optical flow of feature points less than the threshold as the input of the Tracking thread. The system makes a preliminary estimate through bag-of-words model matching and outputs a preliminary estimated trajectory;
[0011] S5. Construct a factor graph for the tight coupling optimization of multi-sensors. The error factors of the factor graph include: image reprojection error factor, IMU pre-integrated attitude error factor, and DVL observed velocity error factor;
[0012] S6. Construct the objective function of the factor graph through the error factors of the factor graph obtained in step S5, that is, minimize the error function. Solve this minimized error function through the Levenberg-Marquardt algorithm. In each iteration, the algorithm linearizes the residual terms of each sensor factor, constructs and solves the normal equation, and ensures the numerical stability of the optimization process by adaptively adjusting the damping factor. Finally, output the optimal trajectory estimation result after multi-sensor fusion calibration.
[0013] Furthermore, the specific method of step S2 is to uniformly scale the input current frame image to 968×608×3. In the YOLOv11 model, extract multi-level features through the multi-scale feature pyramid extraction module, including: 2×2 pixel blocks of the image are used to capture small target details, 4×4 pixel blocks of the image are used to balance semantic and spatial information, and 8×8 pixel blocks of the image are used to integrate global context to handle target occlusion; Embed a dynamic perception enhancement module in the Backbone module, combine spatio-temporal attention mechanism and local-global feature interaction to suppress the interference of suspended particles and ocean currents; Use a lightweight CSP-DenseNet structure in the feature extraction stage; Introduce an adaptive feature selection mechanism in the Neck part, use learnable gating to dynamically allocate multi-scale feature weights, and finally output the prior recognition box of dynamic targets , including type, confidence, and coordinates, expressed as:
[0014] ,
[0015] where represents the target category, indicating that the target is a diver, indicating that the target is a fish; is the confidence score, the closer it is to 1, the higher the confidence of the model in the detection result; is the center point coordinate; represents the width of the prior recognition box, represents the height of the prior recognition box.
[0016] Furthermore, the complete network structure of the superpoint model described in step S3 includes three parts: the shared encoding network Shared Encoder, the interest point decoding network Interest Point Decoder, and the descriptor decoding network Descriptor Decoder;
[0017] The method of extracting feature points using the superpoint model in step S3 and classifying the feature points in the current frame image according to the bounding box information of the dynamic target obtained in step S2 is as follows: First, input an image with height H and width W into the shared encoding network Shared Encoder, and obtain a compressed feature map with height H / 8, width W / 8, and descriptor of 128 bits after passing through four block blocks; input the output of the shared encoding network Shared Encoder into the interest point decoding network Interest Point Decoder, and obtain a feature map with height H / 8, width W / 8, and descriptor of 65 through two layers of convolution. After softmax normalization, retain the top N non-maximum suppression points to obtain the key point coordinates; at the same time, input the output of the shared encoding network Shared Encoder into the descriptor decoding network Descriptor Decoder, and obtain a 256-dimensional descriptor through a convolutional decoder, bilinear interpolation, and L2 normalization.
[0018] Furthermore, the specific method of step S4 includes:
[0019] The internal parameter matrix K of the camera is obtained in advance through the calibration method. The rotation matrix R and translation vector t of the camera movement are calculated through consecutive frames. The projection matrix P is expressed as [ R | t ] , then the optical flow of the static feature points is expressed as:
[0020] ,
[0021] Calculate the optical flow residual: the actual optical flow and the optical flow of static feature points The residual between , then the residual norm is expressed as:
[0022] ,
[0023] In the formula, represents the horizontal coordinate of the actual optical flow point, represents the horizontal coordinate of the static optical flow point;
[0024] Classify the optical flow, preset a threshold , by comparing the residual norm with the threshold , classify the optical flow, expressed as:
[0025] ,
[0026] In the formula, indicates that the optical flow is classified as static optical flow, indicates that the optical flow is classified as dynamic optical flow.
[0027] Furthermore, in step (5), the IMU pre-integration factor is expressed as:
[0028] ,
[0029] where, represents the position estimate of the IMU at time, represents the position estimate of the IMU at -1 time, represents the velocity of the IMU at time, represents the IMU at time acceleration, represents the time step;
[0030] The DVL provides the velocity observation of the underwater object, which plays a constraint role in the velocity estimation of the IMU. Therefore, the DVL velocity factor is expressed as:
[0031] ,
[0032] where, represents the velocity observed by the DVL, represents the velocity calculated by the IMU;
[0033] Construct the visual reprojection factor: For the feature points extracted by the superpoint model in step S3, the weights of the static feature points and the dynamic feature points in the factor graph are expressed as:
[0034] ,
[0035] Then the visual reprojection factor is expressed as:
[0036] f visual (x)=ω(x) ⋅ [x - π(X, T )] ,
[0037] wherein, represents a certain pixel point on the image, represents the pixel point 's position in three-dimensional space, represents the external parameter matrix of the camera, represents the function that projects a point in three-dimensional space onto a two-dimensional plane.
[0038] Furthermore, the error factor of the factor graph obtained in step S6 through step S5 constructs the objective function of the factor graph , which is expressed as:
[0039] ,
[0040] Then the optimal state estimation is obtained by solving through the Levenberg-Marquardt algorithm.
[0041] Beneficial effects: The present invention provides an IMU / DVL-optimized simultaneous localization and mapping algorithm for integrating dynamic object detection in the deep and far sea. Aiming at the problem that in the complex underwater environment of the deep and far sea, the existing SLAM technology is vulnerable to the influence of the dynamic environment, resulting in large positioning errors, an innovative solution is proposed. Specifically, the present invention introduces the YOLOv11 deep learning algorithm to perform real-time target detection on underwater images, which can efficiently identify dynamic objects such as fish schools and divers, and combines the optical flow method to accurately eliminate dynamic feature points. The YOLOv11 algorithm can maintain a high detection accuracy in the complex underwater lighting and turbid environment through the pre-trained underwater fish and diver datasets, ensuring the accurate identification of dynamic objects. On this basis, the optical flow method calculates the optical flow vectors of feature points to distinguish static and dynamic feature points, effectively eliminates dynamic feature points, and inputs the filtered feature points into the Tracking link, thereby significantly improving the robustness and positioning accuracy of the system. In the backend, the state is optimized by constructing a factor graph that combines the DVL observation velocity error, the IMU pre-integration error, and the image feature point reprojection error, providing more accurate position and attitude information. This method not only solves the problem that traditional SLAM algorithms are prone to losing tracking or generating map drift in dynamic environments, but also provides reliable technical support for the stable navigation and positioning of AUVs in the complex underwater environment of the deep and far sea. Description of the Drawings
[0042] Figure 1 System block diagram of YOLOv11 and the optical flow method;
[0043] Figure 2 Overall block diagram of the present invention;
[0044] Figure 3 Visualization results of dynamic target detection under the RUOD, DUO, and Archive datasets. Among them, (a) and (b) are the detection results of the YOLOv11 model on the test set, (c) is the diver detected in the underwater image sequence of the Archive dataset, and (d) is the fish school detected in the underwater image sequence of the Archive dataset;
[0045] Figure 4 Index diagram of YOLOv11 model training. Among them, (a) represents the index mAP50, (b) represents the index precision, and (c) represents the index recall;
[0046] Figure 5 Visualization results of eliminating target dynamic feature points under the Archive dataset. The objects framed by the blue boxes represent the dynamic objects detected by the YOLOv11 model, the green dots represent the effective feature points, and the red dots represent the eliminated invalid feature points. Among them, (a) shows the elimination of the diver, and (b) shows the elimination of the fish school target. Specific implementation mode
[0047] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0048] As Figure 1 shown, the underwater SLAM method based on dynamic detection and IMU DVL fusion in this embodiment solves the problem that visual odometry is difficult to handle dynamic environments in complex underwater environments. The YOLOv11 algorithm is used to detect images and identify dynamic objects, which are framed by bounding boxes. ORB feature points are extracted from the images. For the feature points within the bounding box, the sparse optical flow algorithm is used to classify static and dynamic feature points. All feature points outside the bounding box and static feature points within the bounding box are retained for subsequent visual simultaneous localization and mapping. The advantages of the method are as follows:
[0049] (a) Dynamically recognize and frame dynamic objects in real time, avoiding interference of dynamic objects on positioning and mapping;
[0050] (b) Combine the optical flow method to screen the feature points within the bounding box, which can more accurately distinguish static and dynamic feature points, significantly improving the accuracy of feature point screening;
[0051] (c) The IMU can provide high-frequency dynamic motion information, and the low-frequency velocity information provided by the DVL can effectively correct the drift of the IMU, avoiding error accumulation. Through the fusion and optimization of the IMU, DVL, and visual images, the accuracy of pose estimation can be improved in the case of relatively fast motion.
[0052] The specific steps include:
[0053] S1. Use 2000 datasets such as RUOD and DUO to train the YOLOv11 model to obtain a trained YOLOv11 model;
[0054] S2. Collect underwater images through an IDS industrial camera at a frequency of 10 Hz, and input the current frame into the YOLOv11 model obtained by training in step S1 to obtain the bounding box information of dynamic targets;
[0055] S3. For the current frame image input in step S2, use the superpoint model to extract feature points, and classify the feature points in the current frame image according to the bounding box information of the dynamic targets obtained in step S2: Feature points located outside the bounding box are regarded as preliminary static feature points, and feature points located within the bounding box are regarded as dynamic feature points and excluded;
[0056] S4. For the preliminary static feature points of adjacent frames, by introducing the Lucas-Kanade optical flow method improved with a camera model, calculate the optical flow of the feature points, set the residual norm threshold, eliminate the optical flow of feature points with a residual norm greater than the threshold, and further save the optical flow of feature points less than the threshold as the input of the Tracking thread. The system makes a preliminary estimate through bag-of-words model matching and outputs a preliminary estimated trajectory;
[0057] S5. Construct a factor graph for the tight coupling optimization of multi-sensors. The error factors of the factor graph include: image reprojection error factor, IMU pre-integrated attitude error factor, and DVL observed velocity error factor. In the factor graph, the image reprojection error factor ensures the visual positioning accuracy, the IMU pre-integrated attitude error factor provides continuous motion constraints to handle fast motion or short-term occlusion, and the DVL observed velocity error factor effectively suppresses the cumulative error of inertial navigation. This not only enhances the robustness of the system in complex scenarios such as texture loss and dynamic interference, but also significantly improves the smoothness and consistency of trajectory estimation.
[0058] S6. Construct the objective function of the factor graph through the error factors of the factor graph obtained in step S5, that is, minimize the error function. Solve this minimized error function through the Levenberg-Marquardt algorithm. In each iteration, the algorithm linearizes the residual terms of each sensor factor, constructs and solves the normal equation, and ensures the numerical stability of the optimization process by adaptively adjusting the damping factor. Finally, output the optimal trajectory estimation result after multi-sensor fusion calibration.
[0059] In step S1, to improve the detection accuracy of the model in the dynamic environment of the deep and far sea, the present invention constructs a targeted annotation dataset, focusing on two types of typical dynamic targets: divers and fish schools. The RUOD and DUO datasets are used and reclassified and annotated. Bounding boxes are generated through a semi-automated annotation tool and verified manually. There are a total of 1552 training sets, 191 test sets, and 172 validation sets.
[0060] The specific method of step S2 is to use the YOLOv11 model obtained in step S1 to perform object detection on the current input frame through the real seabed image sequence with a frequency of 10Hz collected by the IDS camera in the Archive dataset publicly available at the University of Haifa, Israel. The input images of the current frame are uniformly scaled to 968×608×3. In the YOLOv11 model, multi-scale feature pyramid extraction modules are used to extract multi-level features, including: 2×2 pixel blocks of the image for capturing details of small objects such as fish schools, 4×4 pixel blocks of the image for balancing semantic and spatial information, and 8×8 pixel blocks of the image for integrating global context to handle object occlusion; a dynamic perception enhancement module is embedded in the Backbone module, which combines spatio-temporal attention mechanism and local-global feature interaction to effectively suppress the interference of suspended particles and ocean currents; the lightweight CSP-DenseNet structure is adopted in the feature extraction stage: the CSP-DenseBlock in the Backbone multiplexes features through dense connections, reducing the number of parameters by 30% while enhancing gradient propagation; an adaptive feature selection mechanism is introduced in the Neck part, which dynamically assigns multi-scale feature weights using learnable gates to enhance the response intensity of the dynamic target area. Finally, the prior recognition boxes of the dynamic targets are output , including type, confidence, and coordinates, represented as:
[0061] ,
[0062] where represents the target category, indicating that the target is a diver, indicating that the target is a fish; is the confidence score, the closer it is to 1, the higher the confidence of the model in the detection result; is the center point coordinate; represents the width of the prior recognition box, represents the height of the prior recognition box.
[0063] Furthermore, the complete network structure of the superpoint model described in step S3 includes three parts: the shared encoding network Shared Encoder, the interest point decoding network Interest Point Decoder, and the descriptor decoding network Descriptor Decoder;
[0064] In step S3, the method of extracting feature points using the superpoint model and classifying the feature points in the current frame image according to the bounding box information of the dynamic target obtained in step S2 is as follows: First, an image with height H and width W is input into the shared encoding network Shared Encoder, and after passing through four block blocks, a compressed feature map with height H / 8, width W / 8, and a descriptor of 128 bits 128 is obtained; the output of the shared encoding network Shared Encoder is input into the interest point decoding network Interest Point Decoder, and through two layers of convolution, a feature map with height H / 8, width W / 8, and a descriptor of 65 is obtained. After softmax normalization, the first N non-maximum suppression points are retained to obtain the key point coordinates; at the same time, the output of the shared encoding network Shared Encoder is input into the feature point description network Descriptor Decoder, and a 256-dimensional descriptor is obtained through a convolutional decoder, bilinear interpolation, and L2 normalization.
[0065] Further, in step S4, in the deep-sea underwater environment, due to the turbidity of the water body, uneven illumination, and frequent interference from dynamic targets (such as fish schools, divers), the traditional optical flow method is prone to misjudgment due to background noise or the movement of dynamic objects. To solve this problem, the present invention improves the traditional optical flow model by combining the camera motion model. The internal parameter matrix K of the camera is obtained in advance through the calibration method, and the rotation matrix R and translation vector t of the camera motion are calculated through consecutive frames. The projection matrix P is expressed as [ R | t ] , then the static feature point optical flow is expressed as:
[0066] ,
[0067] Calculate the optical flow residual: The actual optical flow and the static feature point optical flow The residual between them, then the residual norm is expressed as:
[0068] ,
[0069] In the formula, represents the horizontal coordinate of the actual optical flow point, represents the horizontal coordinate of the static optical flow point;
[0070] Classify the optical flow, preset a threshold , by comparing the residual norm with the threshold , classify the optical flow, expressed as:
[0071] ,
[0072] In the formula, indicates that the optical flow is classified as static optical flow, indicates that the optical flow is classified as dynamic optical flow.
[0073] Furthermore, in step S5, the IMU pre-integration factor is expressed as:
[0074] ,
[0075] where represents the position estimation of the IMU at time, represents the position estimation of the IMU at -1 time, represents the velocity of the IMU at time, represents the acceleration of the IMU at time, represents the time step;
[0076] The DVL provides the velocity observation of the underwater object, which plays a constraint role in the velocity estimation of the IMU. Therefore, the DVL velocity factor is expressed as:
[0077] ,
[0078] where represents the velocity observed by the DVL, represents the velocity calculated by the IMU;
[0079] Construct the visual reprojection factor: For the feature points extracted by the superpoint model in step S3, the weights of the static feature points and dynamic feature points in the factor graph are expressed as:
[0080] ,
[0081] Then the visual reprojection factor is expressed as:
[0082] f visual (x)=ω(x) ⋅ [x - π(X, T )] ,
[0083] Among them, represents a certain pixel point on the image, represents the pixel point 's position in three-dimensional space, represents the external parameter matrix of the camera, represents the function that projects a point in three-dimensional space onto a two-dimensional plane.
[0084] Furthermore, the error factor of the factor graph obtained through step S5 in step S6 constructs the objective function of the factor graph , expressed as:
[0085] ,
[0086] Then, the optimal state estimation is obtained by solving through the Levenberg-Marquardt algorithm.
[0087] Simulation experiment:
[0088] The test results on the RUDO and DUO datasets show that the specially trained YOLOv11
[0089] model demonstrates excellent detection performance, where the test effects on the RUDO and DUO datasets are as Figure 3 shown in (a) and (b) therein, and the detection effect on the underwater image sequence of Archive is as Figure 3 shown in (c) and (d) therein. The model metrics mAP50 reaches 0.779, precision reaches 0.837, and recall reaches 0.723, as Figure 4 shown in (a), (b), and (c) therein.
[0090] The extraction effect of the superpoint feature points after double screening by YOLOv11 object detection and the improved LK optical flow method based on camera motion is as Figure 5 shown in (a) and (b) therein. Among them Figure 5 in (a), 217 effective feature points are extracted, 31 dynamic feature points are removed, and the feature point removal rate is 12.5%; Figure 5 in (b), 342 effective feature points are extracted, 32 dynamic feature points are removed, and the feature point removal rate is 8.56%. The experimental results show that through the collaborative optimization of YOLOv11 dynamic detection and multi-sensor fusion in the present invention, the SLAM problem in the underwater dynamic environment is effectively solved. In particular, the test on the Archive real dataset verifies the practicability of the algorithm and provides a reliable solution for the navigation and positioning of AUVs in complex underwater environments.
Claims
1. An underwater SLAM method based on dynamic detection and IMUDVL fusion, characterized in that: The method comprises the following steps: S1. Train the YOLOv11 model to obtain a trained YOLOv11 model; S2. Collect underwater images through an IDS industrial camera at a frequency of 10 Hz, input the current frame into the YOLOv11 model trained in step S1, and obtain the bounding box information of the dynamic target; S3. For the current frame image input in step S2, the superpoint model is used to extract feature points, and according to the bounding box information of the dynamic target obtained in step S2, the feature points in the current frame image are classified: the feature points outside the bounding box are regarded as preliminary static feature points, and the feature points inside the bounding box are regarded as dynamic feature points for elimination; S4. For the preliminary static feature points of adjacent frames, the Lucas-Kanade optical flow method improved by introducing the camera model is used to calculate the optical flow of the feature points, and the residual norm threshold is set. The optical flow of the feature points with a residual norm greater than the threshold is eliminated, and the optical flow of the feature points with a residual norm less than the threshold is further saved as the input of the Tracking thread. The system matches the preliminary estimate through the bag-of-words model and outputs the preliminary estimated trajectory; S5. Construct a factor graph for tight coupling optimization of multi-sensors. The error factors of the factor graph include: image reprojection error factor, IMU pre-integration attitude error factor, and DVL observation velocity error factor. S6. The objective function of the factor graph, i.e., the minimization error function, is constructed through the error factors of the factor graph obtained in step S5. The minimization error function is solved by the Levenberg-Marquardt algorithm. In each iteration, the algorithm linearizes the residual terms of each sensor factor, constructs and solves the normal equations, and ensures the numerical stability of the optimization process by adaptively adjusting the damping factor. Finally, the optimal trajectory estimation result after multi-sensor fusion calibration is output.
2. The underwater SLAM method based on dynamic detection and IMUDVL fusion according to claim 1, characterized in that: The specific method of step S2 is to uniformly scale the current frame input image to 968×608×3, and in the YOLOv11 model, extract multi-level features through a multi-scale feature pyramid extraction module, including: image 2×2 pixel blocks are used to capture small target details, image 4×4 pixel blocks are used to balance semantic and spatial information, and image 8×8 pixel blocks are used to integrate global context to deal with target occlusion; embed a dynamic perception enhancement module in the Backbone module, combine the spatiotemporal attention mechanism with local-global feature interaction, and suppress the interference of suspended particles and ocean currents; adopt a lightweight CSP-DenseNet structure in the feature extraction stage; introduce an adaptive feature selection mechanism in the Neck part, use learnable gating to dynamically allocate multi-scale feature weights, and finally output the prior recognition box of the dynamic target , including type, confidence and coordinates, expressed as: , in represents the target category, Indicates that the target is a diver. Indicates that the target is fish; is the confidence score, The closer it is to 1, the higher the confidence of the model in the detection result; is the center point coordinate; Represents the width of the prior recognition box, Indicates the height of the prior recognition box.
3. The underwater SLAM method based on dynamic detection and IMUDVL fusion according to claim 1 is characterized in that, The complete network structure of the superpoint model in step S3 includes three parts: a shared encoding network SharedEncoder, a feature point decoding network Interest Point Decoder, and a feature point description network DescriptorDecoder; The method of extracting feature points using the superpoint model described in step S3, and classifying the feature points in the current frame image according to the bounding box information of the dynamic target obtained in step S2 is as follows: first, input an image with a height of H and a width of W to the shared encoding network SharedEncoder, and obtain a compressed feature map with an output height of H / 8, a width of W / 8, and a descriptor of 128 bits after four blocks; input the output of the shared encoding network Shared Encoder to the feature point decoding network Interest Point Decoder, obtain a feature map with a height of H / 8, a width of W / 8, and a descriptor of 65 after two layers of convolution, retain the first N non-maximum suppressed points after softmax normalization, and obtain the key point coordinates; at the same time, input the output of the shared encoding network Shared Encoder to the feature point description network Descriptor Decoder, and obtain a 256-dimensional descriptor after convolution decoder, bilinear interpolation and L2 normalization.
4. The underwater SLAM method based on dynamic detection and IMUDVL fusion according to claim 1 is characterized in that, The specific method of step S4 includes: The camera's intrinsic parameter matrix K is known in advance through the calibration method, and the rotation matrix R and translation vector t of the camera motion are calculated through continuous frames. The projection matrix P is expressed as , then the static feature point optical flow It is expressed as: , Calculating optical flow residual: actual optical flow Optical flow with static feature points The residual between , then the residual norm It is expressed as: , In the formula, Represents the horizontal coordinate of the actual optical flow point, Represents the horizontal coordinate of the static optical flow point; Classify the optical flow and preset the threshold , by comparing the residual norm With threshold , classify the optical flow, expressed as: , In the formula, Indicates that the optical flow is summarized as static optical flow, Indicates that the optical flow is summarized as dynamic optical flow.
5. The underwater SLAM method based on dynamic detection and IMUDVL fusion according to claim 4 is characterized in that, In step (5), the IMU pre-integration factor It is expressed as: , in, Indicates that the IMU is Estimated position at the moment, Indicates that the IMU is -1 moment position estimate, Indicates that the IMU is The speed of time, Indicates that the IMU is The acceleration of time, represents the time step; DVL provides the velocity observation of underwater objects, which constrains the velocity estimation of IMU. Therefore, the DVL velocity factor It is expressed as: , in, represents the velocity observed by DVL, Indicates the speed calculated by IMU; Construct visual reprojection factors: For the feature points extracted by the superpoint model in step S3, the weights of static feature points and dynamic feature points in the factor graph It is expressed as: , The visual reprojection factor It is expressed as: , in, Represents a pixel point on the image. Represents pixel The position in three-dimensional space, represents the camera's extrinsic matrix, Represents a function that projects a point in three-dimensional space onto a two-dimensional plane.
6. The underwater SLAM method based on dynamic detection and IMUDVL fusion according to claim 1 is characterized in that: Step S6 constructs the objective function of the factor graph using the error factor of the factor graph obtained in step S5. , expressed as: , Then the optimal state estimation is obtained by using the Levenberg-Marquardt algorithm.
Citation Information
Patent Citations
Bionic polarization semantic SLAM method based on neural radiation field
CN119444857A
Stackable unmanned aerial vehicle (UAV) system and portable hangar system therefor
US20190077519A1