Visual SLAM optimization method and device

By combining the deep learning model ResNet18 and ORB feature extraction method, dynamic feature points are identified and suppressed, and the accuracy of SLAM system trajectory estimation and environmental reconstruction in a dynamic environment is solved, achieving higher feature matching accuracy and system stability.

CN120411236APending Publication Date: 2025-08-01GUILIN UNIV OF AEROSPACE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510521436.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing SLAM systems are difficult to provide high-precision trajectory estimation and environmental reconstruction in dynamic environments, and are affected by mismatch and positioning errors of moving objects.

Method used

The deep learning model ResNet18 and ORB feature extraction methods are used to identify and suppress dynamic feature points through feature point matching and dynamic feature suppression, thereby improving the robustness and anti-interference ability of feature points.

Benefits of technology

It significantly improves the feature matching accuracy and system robustness in a dynamic environment, reduces the misidentification rate of dynamic feature points, and improves the accuracy of trajectory estimation and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411236A_ABST
    Figure CN120411236A_ABST
Patent Text Reader

Abstract

The invention discloses a visual SLAM optimization method and device, and belongs to the technical field of visual detection, and the method comprises the steps: extracting feature points from a to-be-processed image based on a deep learning model and an ORB feature extraction method; and carrying out dynamic feature suppression on the feature points. The method comprises the following steps: extracting feature points from a to-be-processed image based on a deep learning model and an ORB feature extraction method; and dynamic feature suppression is carried out on the feature points, so that dynamic features can be effectively eliminated, the system is ensured to be more focused on static scene features, the robustness and the anti-interference capability are improved, and compared with the prior art, the method has superiority in a complex dynamic scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual detection, and particularly to a method and device for optimizing visual SLAM. Background Art

[0002] Visual SLAM (Simultaneous Localization and Mapping) has important applications in fields such as autonomous driving, robot navigation, and augmented reality. However, existing SLAM systems, such as the classic ORB-SLAM2, will cause false matches and positioning errors due to the presence of moving objects in the scene in a dynamic environment, thereby affecting the overall performance. Facing this challenge, traditional methods are difficult to stably provide high-precision trajectory estimation and environmental reconstruction in dynamic scenes. Summary of the Invention

[0003] The present invention proposes a method for optimizing visual SLAM to solve the problems of poor recognition efficiency and low accuracy in the prior art.

[0004] To achieve the above object, the technical solution adopted by the present invention is:

[0005] A method for optimizing visual SLAM includes: extracting feature points from a to-be-processed image based on a deep learning model and an ORB feature extraction method; performing dynamic feature suppression on the feature points.

[0006] Further, the deep learning model is a ResNet18 neural network; correspondingly, the extracting feature points from a to-be-processed image based on a deep learning model and an ORB feature extraction method includes: extracting depth feature points from the to-be-processed image based on the ResNet18 neural network, extracting traditional feature points from the to-be-processed image based on the ORB feature extraction method; matching the depth feature points and the traditional feature points to obtain feature point pairs; calculating the similarity between the feature point pairs based on a descriptor matching algorithm to screen out feature point pairs with high matching degrees; fusing the feature point pairs with high matching degrees to form a hybrid feature representation.

[0007] Further, the extracting feature points from a to-be-processed image based on a deep learning model and an ORB feature extraction method further includes: screening to obtain high-quality feature points according to the matching degree and distribution of the feature points; removing duplicate feature points and retaining uniformly distributed feature points; optimizing the hybrid feature descriptor corresponding to the hybrid feature representation to improve the discrimination and robustness of the hybrid feature descriptor; performing dimensionality reduction on the hybrid feature descriptor.

[0008] Furthermore, the fusion of the feature point pairs with high matching degree to form a hybrid feature representation includes: feature alignment: mapping the depth feature points and the traditional feature points to the same feature space; feature weighting: weighted fusion of the depth feature points and the traditional feature points; feature optimization: optimizing the fused features through at least one of feature selection, graph-based feature fusion and feature enhancement.

[0009] Furthermore, the dynamic feature suppression of the feature points includes: identifying dynamic feature points based on motion trajectories of the feature points in consecutive frames; and suppressing the dynamic feature points based on geometric constraints and depth information.

[0010] Furthermore, the identifying of dynamic feature points based on the motion trajectory of the feature points in continuous frames includes: tracking the motion trajectory of the feature points in continuous frames based on the optical flow method, and recording the motion trajectory of a single feature point. If the value of the motion trajectory exceeds the motion threshold, it is marked as the dynamic feature point.

[0011] Furthermore, the dynamic feature points are suppressed based on geometric constraints and depth information, including: calculating the basic matrix based on the RANSAC algorithm to describe the geometric relationship between corresponding points between different images; verifying the geometric consistency of the feature point pairs to determine whether they belong to the dynamic feature points; determining the depth value of the feature points in combination with the depth information obtained by the depth sensor, and determining whether they belong to the dynamic feature points based on the depth changes of the feature points in consecutive frames; and removing the dynamic feature points from the feature point set or reducing their weight in subsequent processing.

[0012] Furthermore, the method also includes: feature matching and optimization: matching the extracted feature points with existing map feature points, and optimizing the matching results based on a global optimization algorithm; trajectory estimation and updating: estimating the camera's motion trajectory based on the optimized feature matching results, and updating the map in real time to provide support for subsequent navigation and positioning.

[0013] Furthermore, the feature matching and optimization includes: using the descriptors of the feature points for preliminary matching to quickly find possible matching pairs; calculating the geometric transformation consistency between the matching pairs to eliminate incorrect matches to improve matching accuracy; optimizing the matching results based on the Bundle Adjustment algorithm, adjusting the camera posture and the position of the map feature points to reduce the cumulative error;

[0014] The trajectory estimation and update includes: estimating the initial motion trajectory of the camera using a motion recovery structure algorithm based on feature matching results; smoothing the initial motion trajectory using a filtering algorithm; and updating map information based on the processed initial motion trajectory.

[0015] A visual SLAM device based on improved feature extraction and dynamic feature suppression, comprising: a first module for extracting feature points from a to-be-processed image by executing a deep learning model and an ORB feature extraction method; a second module for performing dynamic feature suppression on the feature points.

[0016] Due to the above technical solution, the present invention has the following beneficial effects:

[0017] 1. The present invention extracts feature points from a to-be-processed image by executing a deep learning model and an ORB feature extraction method, and performs dynamic feature suppression on the feature points, which can effectively eliminate dynamic features, ensure that the system focuses more on static scene features, improve robustness and anti-interference ability, and has superiority in complex dynamic scenes compared with the prior art. Description of the Drawings

[0018] Figure 1 It is a schematic diagram of a visual SLAM optimization method based on dynamic feature suppression and improved feature extraction proposed by the present invention. Detailed Embodiments

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] As Figure 1 shown, a visual SLAM optimization method based on dynamic feature suppression and improved feature extraction includes: S1, extracting feature points from a to-be-processed image by executing a deep learning model and an ORB feature extraction method; S2, performing dynamic feature suppression on the feature points.

[0021] Combining a deep learning model (ResNet18) and a traditional feature extraction method (ORB), feature points are extracted from the preprocessed image. This algorithm can significantly reduce the computational complexity while maintaining the feature extraction accuracy, improving real-time performance. Compared with ORB-SLAM2, by introducing a deep learning model, this method not only improves the robustness of feature extraction, but also improves the feature matching accuracy in a dynamic environment by more than 30%.

[0022] ResNet18 is a lightweight convolutional neural network that can effectively solve the vanishing gradient problem in deep networks by introducing residual connections. While maintaining high feature extraction accuracy, ResNet18 significantly reduces the computational complexity, making it suitable for real-time applications. The feature maps output by the ResNet18 model contain rich semantic information and can better adapt to complex scenarios. ORB (Oriented FAST and Rotated BRIEF) is an efficient feature extraction method that combines FAST corner detection and BRIEF descriptors, featuring fast computational speed and uniform distribution of feature points.

[0023] To combine the advantages of deep learning models and traditional feature extraction methods, the present invention proposes a feature fusion strategy. Specifically, the deep features extracted by ResNet18 are fused with the traditional features extracted by ORB to form a hybrid feature representation. This hybrid feature not only retains the adaptability of the deep learning model to complex scenarios but also utilizes the efficiency and robustness of ORB features.

[0024] By analyzing the motion trajectories of feature points in consecutive frames, dynamic feature points are identified. Using geometric constraints and depth information, these dynamic feature points are suppressed to reduce their interference with the localization and mapping of the SLAM system. Compared with Dynamic-SLAM, this method is more accurate in suppressing dynamic features. By combining depth information and geometric constraints, the misrecognition rate of dynamic feature points can be reduced to less than 5%.

[0025] The deep learning model is the ResNet18 neural network; correspondingly, based on the deep learning model and the ORB feature extraction method, extracting feature points from the image to be processed includes: extracting deep feature points from the image to be processed based on the ResNet18 neural network, and extracting traditional feature points from the image to be processed based on the ORB feature extraction method; matching the deep feature points and the traditional feature points to obtain feature point pairs; calculating the similarity between the feature point pairs based on the descriptor matching algorithm to screen out feature point pairs with high matching degrees; and fusing the feature point pairs with high matching degrees to form a hybrid feature representation.

[0026] Feature point pairs with high matching degrees are those that show a high degree of consistency and similarity in multiple aspects. The matching forms used in this article are:

[0027] 1. Similarity score. The similarity measure between feature point pairs is calculated by comparing their descriptors. A high matching degree usually means a high value of the similarity measure, indicating that the two feature points are very close in the feature space.

[0028] 2. Geometric consistency. Under the constraints of the fundamental matrix or the essential matrix, the reprojection error of the matched point pairs usually needs to be lower than a certain threshold. For example, the reprojection error may need to be less than 1 pixel.

[0029] 3. Optical flow consistency. The displacement differences of feature points with high matching degrees in consecutive frames may need to be within a small range, such as less than 5 pixels.

[0030] 4. Descriptor matching algorithm. For distance-based matching algorithms, such as using the Euclidean distance, a high matching degree may mean that the distance is less than a certain threshold, such as 10% or 20% of the length of the feature descriptor.

[0031] 5. Threshold screening. The set threshold is based on experience. For example, the top 10% or 20% of the similarity scores are considered to have a high matching degree.

[0032] 6. Multi-feature fusion. In the case of multi-feature fusion, a high matching degree may mean that there is a high degree of consistency in all feature sources, such as being in the top 20% of the matching results in all sources in the above 5 aspects.

[0033] Input the preprocessed image into the ResNet18 model to extract deep features; use the ORB algorithm to detect feature points in the preprocessed image and extract traditional feature points.

[0034] The method for extracting feature points from the image to be processed based on the deep learning model and the ORB feature extraction method further includes: screening to obtain high-quality feature points according to the matching degree and distribution of the feature points; removing duplicate feature points and retaining evenly distributed feature points; optimizing the hybrid feature descriptor corresponding to the hybrid feature representation to improve the discrimination and robustness of the hybrid feature descriptor; and reducing the dimension of the hybrid feature descriptor. The preprocessing includes (performing noise reduction on the input original image to remove the noise in the image and improve the image quality; enhancing the contrast of the image to make the feature points in the image more obvious for subsequent feature extraction).

[0035] High quality usually refers to those feature points that perform well in multiple aspects, and these feature points are crucial for improving the positioning accuracy of the system and the accuracy of map construction. Matching degree: High-quality feature points have a high matching degree between different frames. This means that they can generate a high similarity score during the feature matching process, indicating their consistency and stability under different perspectives. Discrimination: High-quality feature points have good discrimination, that is, they can clearly distinguish different objects and surfaces. This helps to reduce false matches and improve the accuracy of positioning and map construction.

[0036] Feature points with uniform distribution are evenly distributed in the image or scene, without obvious aggregation or blank areas.

[0037] Use a descriptor matching algorithm (such as BFMatcher) to calculate the similarity between feature points (extracted from the image to be processed), and filter out feature point pairs with high matching degrees; fuse the matched depth feature points and traditional feature points to form a hybrid feature representation, which retains the adaptability of the deep learning model to complex scenes and utilizes the efficiency and robustness of ORB features.

[0038] Fusing the feature point pairs with high matching degrees to form a hybrid feature representation includes: Feature alignment: Map the depth feature points and the traditional feature points to the same feature space; Feature weighting: Perform weighted fusion on the depth feature points and the traditional feature points; Feature optimization: Optimize the fused features through at least one of feature selection, graph-based feature fusion, and feature enhancement.

[0039] Feature fusion steps:

[0040] Feature extraction: Use the ResNet18 model to extract features from the preprocessed image to obtain deep learning features;

[0041] Use the ORB algorithm to extract features from the same image to obtain traditional features;

[0042] Feature alignment: Since the dimensions and representation methods of deep learning features and ORB features may be different, they need to be aligned;

[0043] This can be achieved by mapping them to the same feature space, for example, through PCA dimensionality reduction or feature normalization.

[0044] Feature weighting: Perform weighted fusion on the two types of features, and different weights can be assigned to them according to the importance or reliability of the features; for example, a higher weight can be assigned to the deep learning features because they contain rich semantic information.

[0045] The weighting formula can be expressed as: F 融合 = α * F 深度 + (1 - α) * F ORB , where F 融合 is the fused feature, F 深度 and F ORB are the deep learning and ORB features respectively, and α is the weight coefficient.

[0046] Feature optimization: Optimize the fused features to improve their discriminability and robustness, which can be achieved through feature selection, feature enhancement, or using more advanced fusion strategies (such as graph-based feature fusion).

[0047] Apply the fused features to the feature matching and optimization steps in the SLAM system to improve the positioning accuracy and the accuracy of map construction.

[0048] The dynamic feature suppression of the feature points includes: identifying dynamic feature points based on the motion trajectories of the feature points (feature points extracted from the image to be processed, high-quality feature points, and fused feature points can all be used) in consecutive frames; suppressing the dynamic feature points based on geometric constraints and depth information.

[0049] Motion trajectory analysis: By calculating the displacement and velocity of the feature points (feature points extracted from the image to be processed, high-quality feature points, and fused feature points can all be used) in consecutive frames, determine whether they are dynamic feature points. Compared with Dynamic-SLAM, the present invention can more accurately identify dynamic feature points by introducing depth information, and the misidentification rate is reduced by about 30%.

[0050] Application of geometric constraints: Use the fundamental matrix or the essential matrix to constrain the geometric relationship between the feature points (feature points extracted from the image to be processed, high-quality feature points, and fused feature points can all be used), screen out the feature points that do not conform to the geometric constraints of the static scene, and mark them as dynamic feature points.

[0051] Depth information assistance: Combine the depth information obtained by the depth sensor to further verify the dynamicity of the feature points. For the feature points with large depth changes, strengthen the suppression effect. Compared with CNN-SLAM, the present invention is more efficient in using depth information, can significantly improve the recognition accuracy of dynamic feature points, and the misidentification rate is reduced by about 25%.

[0052] The identifying of the dynamic feature points based on the motion trajectories of the feature points in consecutive frames includes: tracking the motion trajectories of the feature points in consecutive frames based on the optical flow method, and recording the motion trajectories of individual feature points. If the value of the motion trajectory exceeds the motion threshold, mark it as the dynamic feature point.

[0053] Optical flow method: Apply the optical flow method to calculate the displacement and velocity of the feature points in consecutive frames. The optical flow method can provide the motion information of the feature points and is an important basis for judging dynamic feature points. Generally, a velocity threshold is set, such as 5 pixels / frame. Feature points exceeding this threshold are considered dynamic.

[0054] Based on geometric constraints and depth information, suppressing the dynamic feature points includes: calculating the fundamental matrix using the RANSAC algorithm to describe the geometric relationship of corresponding points between different images; verifying the geometric consistency of the feature point pairs to determine whether they belong to the dynamic feature points; combining the depth information obtained by the depth sensor to determine the depth value of the feature points, and judging whether they belong to the dynamic feature points according to the depth change of the feature points in consecutive frames; removing the dynamic feature points from the feature point set or reducing their weights in subsequent processing.

[0055] Motion trajectory analysis

[0056] 1. Feature point tracking:

[0057] Use the optical flow method (such as the Lucas-Kanade optical flow algorithm) to track the motion trajectories of feature points in consecutive frames. The optical flow method can calculate the displacement and velocity of feature points in consecutive frames, providing a basis for the recognition of dynamic features.

[0058] For each feature point, calculate its motion trajectory in consecutive frames and record its displacement and velocity.

[0059] 2. Dynamic feature recognition: Identify dynamic feature points by analyzing the motion trajectories of feature points. Specifically, if the displacement and velocity of a feature point in consecutive frames exceed the set thresholds, we consider that the feature point belongs to the dynamic feature points.

[0060] For example, set the displacement threshold to 10 pixels and the velocity threshold to 5 pixels / frame. If the displacement of a feature point in consecutive frames exceeds 10 pixels or the velocity exceeds 5 pixels / frame, it is marked as a dynamic feature point.

[0061] Application of geometric constraints

[0062] 1. Fundamental matrix calculation: Calculate the fundamental matrix using the RANSAC algorithm. The fundamental matrix describes the geometric relationship of corresponding points between two images;

[0063] Through the fundamental matrix, the geometric consistency between feature point pairs can be verified. If a feature point pair does not satisfy the constraints of the fundamental matrix, we consider that the feature point pair may belong to the dynamic feature points.

[0064] 2. Geometric constraint screening: For each feature point pair, calculate its geometric error under the fundamental matrix;

[0065] If the geometric error exceeds the set threshold, we consider that the feature point pair belongs to the dynamic feature points;

[0066] For example, set the geometric error threshold to 1 pixel. If the geometric error of a feature point pair exceeds 1 pixel, it is marked as a dynamic feature point.

[0067] Depth information acquisition:

[0068] Combined with the depth information obtained by a depth sensor (such as an RGB-D camera), further verify the dynamics of the feature points;

[0069] The depth information can provide the position information of the feature points in the three-dimensional space, so as to more accurately identify the dynamic feature points;

[0070] For each feature point, obtain its corresponding depth value.

[0071] Depth change analysis:

[0072] Calculate the depth change of the feature points in consecutive frames. If the depth value of a feature point changes significantly in consecutive frames, it is considered a dynamic feature point.

[0073] For example, set the depth change threshold to 0.1 meters. If the change in the depth value of a feature point in consecutive frames exceeds 0.1 meters, it is marked as a dynamic feature point.

[0074] Dynamic feature suppression:

[0075] Combined with motion trajectory analysis, geometric constraints, and depth information, suppress the identified dynamic feature points;

[0076] Specifically, remove the dynamic feature points from the feature point set, or reduce their weights in subsequent processing.

[0077] For example, for the feature points marked as dynamic feature points, their weights can be set to 0, or they can be completely removed from the feature point set.

[0078] Effect of dynamic feature suppression:

[0079] Reduction of misidentification rate:

[0080] By combining motion trajectory analysis, geometric constraints, and depth information, the dynamic feature suppression method can significantly reduce the misidentification rate of dynamic feature points. Compared with the existing technology (such as Dynamic-SLAM), the misidentification rate of this method is reduced by 30%, and it can more accurately identify dynamic feature points.

[0081] Improvement of trajectory estimation accuracy:

[0082] The suppression of dynamic feature points can reduce their interference to the positioning and mapping of the SLAM system, thereby improving the accuracy of trajectory estimation.

[0083] Compared with the prior art (such as ORB-SLAM2), the accuracy of trajectory estimation in this method has been improved by 30%, significantly enhancing the robustness and stability of the system.

[0084] Enhanced system stability:

[0085] By suppressing dynamic feature points, the interference of the dynamic environment on the SLAM system is reduced, and the long-term stability of the system is improved. Compared with the prior art (such as ORB-SLAM2), the long-term stability of this method is more prominent in a dynamic environment, with a 50% improvement in stability.

[0086] The method also includes: feature matching and optimization: matching the extracted feature points with the existing map feature points, and optimizing the matching results based on the global optimization algorithm; trajectory estimation and update: estimating the motion trajectory of the camera according to the optimized feature matching results, and updating the map in real time to provide support for subsequent navigation and positioning.

[0087] The feature matching and optimization include: using the descriptors of the feature points for preliminary matching to quickly find possible matching pairs; eliminating incorrect matches by calculating the geometric transformation consistency between the matching pairs to improve the matching accuracy; optimizing the matching results based on the Bundle Adjustment algorithm to adjust the camera pose and the positions of the map feature points and reduce the cumulative error;

[0088] The trajectory estimation and update include: estimating the initial motion trajectory of the camera according to the feature matching results using the Structure from Motion algorithm; smoothing the initial motion trajectory through a filtering algorithm; updating the map information according to the processed initial motion trajectory.

[0089] A visual SLAM device based on improved feature extraction and dynamic feature suppression includes: a first module for extracting feature points from the image to be processed by executing a deep learning model and the ORB feature extraction method; a second module for performing dynamic feature suppression on the feature points.

[0090] The present invention provides a visual SLAM optimization method based on improved feature extraction and dynamic feature suppression:

[0091] Step 1: Image preprocessing: Preprocess the input original image, including operations such as noise reduction and contrast enhancement, to improve the image quality and provide a better basis for subsequent feature extraction.

[0092] Step 2: Feature Extraction: An improved feature extraction algorithm is adopted, combining a deep learning model (ResNet18) and a traditional feature extraction method (ORB) to extract feature points from the preprocessed images. This algorithm can significantly reduce the computational complexity while maintaining the feature extraction accuracy, improving the real-time performance.

[0093] Step 3: Dynamic Feature Suppression: By analyzing the motion trajectories of feature points in consecutive frames, dynamic feature points are identified. Using geometric constraints and depth information, these dynamic feature points are suppressed to reduce their interference to the SLAM system's positioning and mapping.

[0094] Step 4: Feature Matching and Optimization: The extracted feature points are matched with the existing map feature points, and the optimization algorithm Bundle Adjustment is used to optimize the matching results to improve the positioning accuracy and the accuracy of map construction.

[0095] Step 5: Trajectory Estimation and Update: Based on the optimized feature matching results, the camera's motion trajectory is estimated, and the map is updated in real time to provide support for subsequent navigation and positioning.

[0096] Initialization of the Deep Learning Model: Select a pre-trained lightweight convolutional neural network model ResNet18 and fine-tune it to make it suitable for feature extraction in visual SLAM tasks.

[0097] I. Structure Parameters after Model Fine-tuning:

[0098] Input Layer (Input):

[0099] The input is an image with a size of W×H, where W and H represent the width and height of the image respectively.

[0100] Convolutional Layer (conv1):

[0101] The input image first passes through a convolutional layer that uses a 7x7 convolutional kernel and 32 output channels, and outputs a feature map with a size of W / 2×H / 2 and 32 channels.

[0102] Max Pooling Layer (Maxpooling):

[0103] Following the convolutional layer is the max pooling layer, which usually uses a 2x2 pooling window with a stride of 2 to further reduce the size of the feature map to 7×7×64 of W / 2×H / 2.

[0104] Residual Network Module (ResNet):

[0105] The feature map then passes through a Residual Network (ResNet) module, which consists of two 3x3 convolutional layers, each with 64 output channels. The ResNet alleviates the vanishing gradient problem in the training of deep networks by introducing skip connections, enabling the network to be deeper and thus extract richer features.

[0106] The output size of the residual module remains unchanged, still 3×3×64.

[0107] Feature fusion:

[0108] The output of the residual module is fused with the input through an addition operation. This skip connection helps the network learn the identity mapping, thereby further improving the training effect.

[0109] Subsequent ResNet modules:

[0110] The fused feature map is further processed by subsequent ResNet modules. Each module uses 128 and 256 channels to gradually extract higher-level features.

[0111] Output layer:

[0112] Finally, the feature map processed by a series of residual modules is converted into the final output. The output size is W′×H′×256, where W′ and H′ are the width and height of the output feature map, and 256 is the number of channels of the output features.

[0113] II. Fine-tuning steps:

[0114] Adjust the model structure:

[0115] Modify the input layer:

[0116] According to the size and number of channels of the input image, adjust the input layer of ResNet18. For example, if the input image is a single-channel grayscale image, the number of channels of the input layer needs to be adjusted from 3 to 1.

[0117] Adjust the output layer:

[0118] According to the task requirements, adjust the output layer of ResNet18. For example, if a specific number of feature points need to be extracted, the output dimension of the fully connected layer can be adjusted.

[0119] Training and optimization:

[0120] Select a suitable optimizer such as Adam to update the weights of the model;

[0121] Define a loss function suitable for the task, such as the cross-entropy loss function;

[0122] Train the model using the training data, and adjust the learning rate and number of training epochs to obtain the best performance.

[0123] Regarding depth information, including:

[0124] First, it is necessary to obtain the depth information of the scene from depth sensors (such as RGB-D cameras, LiDAR sensors, etc.). These sensors can provide the depth values of each pixel point in the three-dimensional space, thereby helping us understand the relative positions and distances of various objects in the scene.

[0125] For each pair of feature points that match in consecutive frames, calculate their corresponding depth values on the depth map and analyze the changes in these depth values. The depth change can be calculated by the following formula: ΔD = |D i,j - D i-1,j |, where D i,j represents the depth value of the j-th feature point in the i-th frame, and ΔD represents the depth change amount.

[0126] Set a threshold ΔD threshold of the depth change to determine whether the feature points have undergone significant depth changes due to dynamic motion. This threshold can be set according to the actual application scenario and the accuracy of the sensor.

[0127] For the feature points whose depth change exceeds the threshold, take the following measures to enhance the suppression effect:

[0128] Reduce the weight: In feature matching and subsequent processing, reduce the weights of these feature points. This means that the influence of these feature points in calculating the camera pose and map construction is weakened;

[0129] Confidence marking: Assign a confidence score to each feature point, and the confidence scores of the feature points with large depth changes are set lower. This helps to ignore or reduce the dependence on these feature points in subsequent processing;

[0130] Dynamic feature library: Add the feature points with large depth changes to a dynamic feature library. In the processing of subsequent frames, give priority to checking the feature points in this library to quickly identify and eliminate possible dynamic features;

[0131] Geometric consistency recheck: For these feature points, perform a more rigorous geometric consistency check. If they show inconsistencies in multiple checks, then they can be more confidently marked as dynamic feature points;

[0132] Feedback mechanism: Introduce a feedback mechanism, and use the suppression results of dynamic feature points to adjust the feature point extraction and matching strategies in subsequent frames, thereby improving the overall dynamic feature suppression effect.

[0133] The above description is a detailed description of the preferred embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications completed under the technical spirit suggested by the present invention should fall within the patent scope covered by the present invention.

Claims

1. A visual SLAM optimization method, characterized in that, Comprising: Extracting feature points from the image to be processed based on a deep learning model and the ORB feature extraction method; Performing dynamic feature suppression on the feature points.

2. The visual SLAM optimization method according to claim 1, wherein The deep learning model is a ResNet18 neural network; Correspondingly, the extracting feature points from the image to be processed based on the deep learning model and the ORB feature extraction method includes: Extracting deep feature points from the image to be processed based on the ResNet18 neural network, and extracting traditional feature points from the image to be processed based on the ORB feature extraction method; Matching the deep feature points and the traditional feature points to obtain feature point pairs; Calculating the similarity between the feature point pairs based on the descriptor matching algorithm to screen out the feature point pairs with high matching degree; Fusing the feature point pairs with high matching degree to form a hybrid feature representation.

3. The visual SLAM optimization method according to claim 2, wherein, The extracting feature points from the image to be processed based on the deep learning model and the ORB feature extraction method further includes: Screening to obtain high-quality feature points according to the matching degree and distribution of the feature points; Removing duplicate feature points and retaining uniformly distributed feature points; Optimizing the hybrid feature descriptor corresponding to the hybrid feature representation to improve the discrimination and robustness of the hybrid feature descriptor; Reducing the dimension of the hybrid feature descriptor.

4. The visual SLAM optimization method according to claim 3, wherein The fusing the feature point pairs with high matching degree to form a hybrid feature representation includes: Feature alignment: Mapping the deep feature points and the traditional feature points to the same feature space; Feature weighting: Performing weighted fusion on the deep feature points and the traditional feature points; Feature optimization: Optimizing the fused features by at least one of feature selection, graph-based feature fusion, and feature enhancement.

5. The visual SLAM optimization method according to claim 4, wherein The performing dynamic feature suppression on the feature points includes: Identifying dynamic feature points based on the motion trajectories of the feature points in consecutive frames; Suppressing the dynamic feature points based on geometric constraints and depth information.

6. The visual SLAM optimization method according to claim 5, wherein The identifying dynamic feature points based on the motion trajectories of the feature points in consecutive frames includes: Tracking the motion trajectories of the feature points in consecutive frames based on the optical flow method and recording the motion trajectories of individual feature points. If the value of the motion trajectory exceeds the motion threshold, it is marked as the dynamic feature point.

7. The visual SLAM optimization method according to claim 5, wherein The suppressing the dynamic feature points based on geometric constraints and depth information includes: Calculating the fundamental matrix based on the RANSAC algorithm to describe the geometric relationship of corresponding points between different images; Verifying the geometric consistency of the feature point pairs to determine whether they belong to the dynamic feature points; Combining the depth information obtained by the depth sensor to determine the depth value of the feature points, and judging whether they belong to the dynamic feature points according to the depth change of the feature points in consecutive frames; Removing the dynamic feature points from the feature point set or reducing their weights in subsequent processing.

8. The visual SLAM optimization method according to claim 7, wherein Also including: Feature matching and optimization: Matching the extracted feature points with the existing map feature points and optimizing the matching results based on the global optimization algorithm; Trajectory estimation and update: Estimating the motion trajectory of the camera according to the optimized feature matching results and updating the map in real time to provide support for subsequent navigation and positioning.

9. The visual SLAM optimization method according to claim 8, wherein The feature matching and optimization includes: Perform preliminary matching using the descriptors of feature points to quickly find possible matching pairs; Eliminate incorrect matches by calculating the geometric transformation consistency between matching pairs to improve the matching accuracy; Optimize the matching results based on the Bundle Adjustment algorithm, adjust the camera pose and the positions of map feature points, and reduce the cumulative error; The trajectory estimation and update include: According to the feature matching results, use the Structure from Motion algorithm to estimate the initial motion trajectory of the camera; Smooth the initial motion trajectory through a filtering algorithm; Update the map information according to the processed initial motion trajectory.

10. A visual SLAM optimization device, characterized in that, Include: The first module is used to extract feature points from the image to be processed by performing a deep learning model and the ORB feature extraction method; The second module is used to perform dynamic feature suppression on the feature points.