Method and system for improving precision of visual odometer based on deep learning
Through the lightweight deep learning model XFeat and the adaptive feature matching algorithm, the problem of insufficient feature points and estimation error of visual odometers in complex environments is solved, the accuracy and robustness of visual odometers are improved, and the dynamic changes in different scenarios are adapted.
Patent Information
- Application Number
- CN202510137349.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art under complex conditions such as dynamic environments, low-texture areas and lighting changes, the feature point extraction of visual odometers is insufficient and estimation errors are large. Traditional methods rely on local geometric information to perform poorly. Deep learning methods are complex in calculations and have high resource requirements, making it difficult to adapt to changes in different environments.
The lightweight deep learning model XFeat is used to extract feature points, combine PIDNet and Light Glue models to remove dynamic feature points, improve feature matching accuracy through adaptive mechanisms and polar line constraints, and enhance robustness by using multi-scale fusion features and context information.
Improve the accuracy and robustness of visual odometers in complex environments, reduce the demand for computing resources, realize real-time feature extraction and matching, and adapt to texture sparseness, lighting changes and dynamic object interference in different scenarios.
Smart Images

Figure CN120070973A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and visual odometry of mobile robots, and particularly relates to a method and system for improving the accuracy of visual odometry based on deep learning. Background Art
[0002] The statements in this section merely provide background art related to the present invention and do not necessarily constitute prior art.
[0003] Visual SLAM (Simultaneous Localization And Mapping) technology plays a crucial role in robot navigation, and visual odometry is a key part of it because it directly affects the effects of subsequent local mapping and loop detection. Compared with laser sensors, cameras can capture richer visual information and have higher cost performance, and have received more and more attention and research in recent years, which provides robots with more powerful pose estimation capabilities, thereby achieving more stable navigation and positioning in various application scenarios. Visual odometry relies on extracting and tracking stable feature points from images to estimate the motion of the camera, but this process faces significant challenges under complex conditions such as dynamic environments, low-texture areas, and illumination changes. In a dynamic environment, moving objects may introduce artifacts, interfering with feature matching and pose estimation; low-texture areas lack sufficient visual information, resulting in difficult feature extraction, which in turn affects the positioning accuracy; illumination changes, such as shadows or uneven illumination, will also cause unstable or unmatched features, thus affecting the stability and accuracy of visual odometry.
[0004] To address these problems, traditional geometric-method-based visual odometry has adopted various strategies. In a dynamic environment, dynamic object detection and removal are usually carried out through geometric methods to exclude the interference of moving objects and ensure the stability of the static scene; in low-texture areas, visual odometry is fused with an IMU for multi-sensor fusion to make up for the shortage of feature points and improve the positioning accuracy; for illumination changes, traditional methods use illumination-invariant feature descriptors and improved image matching algorithms to enhance the stability of features, and reduce the impact of illumination differences through local illumination adjustment. Although these strategies have improved the robustness and accuracy of visual odometry in complex environments to a certain extent, due to still relying on traditional feature point methods such as ORB, SIFT, and SURF, mainly relying on local geometric information such as corners and edges of images and not combining the context information in the images, there are still certain limitations in the camera pose estimation effect in some extreme scenarios and it does not have better universality.
[0005] With the development of deep learning, methods based on deep learning have gradually become a new trend for improving the accuracy of visual odometry. Although the end-to-end visual odometry algorithm based on deep learning can better adapt to different environmental changes, it also has some deficiencies. First, deep learning models usually require a large amount of labeled data for training, which makes it difficult to implement them in applications with scarce data or high annotation costs. Second, the computational complexity of deep neural networks is relatively high, especially on resource-limited embedded devices or mobile platforms, where real-time requirements may not be met. Third, deep learning methods have high requirements for the diversity and quality of training data. An insufficiently trained model may lead to poor generalization ability and difficulty in adapting to new environments or tasks. Finally, although deep learning can automate the feature extraction and optimization process, the training, tuning, and deployment processes of the entire end-to-end system are still relatively complex and require a large amount of computational resources and engineering support.
[0006] The prior art has the following defects: (1) Traditional feature point methods mainly rely on local geometric information of images, such as corner points and edges, and perform poorly in complex environments, easily leading to feature matching failures or camera pose estimation biases. (2) The sparse feature points extracted by traditional methods cannot effectively handle large-scale or detail-rich scenes, and their descriptors have limited adaptability to scale, rotation, and illumination changes. (3) End-to-end visual odometry algorithms require a large amount of labeled data, with high training costs, high requirements for computational resources, and poor integration with the removal of dynamic points. Summary of the Invention
[0007] To solve the deficiencies of the prior art, the present invention provides a method and system for improving the accuracy of visual odometry based on deep learning, which solves the problems of insufficient feature point extraction and large estimation errors when performing camera pose estimation based on visual sensors under complex conditions such as dynamic environments, low-texture regions, and illumination changes, and improves the accuracy of visual odometry in complex environments.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] In the first aspect, the present invention provides a method for improving the accuracy of visual odometry based on deep learning.
[0010] A method for improving the accuracy of visual odometry based on deep learning includes the following processes:
[0011] Convert the acquired image data into grayscale image data, and use a pre-trained first deep learning model to extract feature points and descriptors corresponding to the feature points in the grayscale image data;
[0012] Use a pre-trained second deep learning model to extract static feature points and prior dynamic feature points from the image data, map the feature points extracted by the first deep learning model one by one with the prior dynamic feature points, and remove the prior dynamic feature points from the feature points extracted by the first deep learning model;
[0013] Process the descriptors corresponding to the remaining feature points after removing the prior dynamic feature points, input them into the third deep learning model for feature point matching, perform epipolar constraint determination on the successfully matched feature points, remove the feature points that do not meet the requirements to obtain the final remaining feature points, and perform camera pose estimation based on the final remaining feature points.
[0014] As a further limitation of the first aspect of the present invention, the first deep learning model uses the XFeat model, including: a lightweight backbone network module, a feature point extraction module, and a descriptor extraction module;
[0015] The lightweight backbone network module includes six sequentially connected basic blocks. Each basic block halves the resolution of the input image. The spatial resolution of the input image of the first basic block is H×W, and the spatial resolution of the output image of the last basic block is H / 32×W / 32. The output of the fourth basic block, the upsampling result of the output of the fifth basic block, and the convolution and upsampling results of the output of the sixth basic block are fused to obtain a multi-scale fusion feature;
[0016] The spatial resolution of the multi-scale fusion feature is H / 8×W / 8. The feature point extraction module extracts feature points based on the feature image output by the fourth basic block, and the descriptor extraction module extracts descriptors based on the multi-scale fusion feature.
[0017] As a further limitation of the first aspect of the present invention, in the feature point extraction module, the multi-scale fusion feature is represented as a two-dimensional grid of 8×8 pixels. Each grid cell is reshaped into a 64-dimensional feature, and the feature point coordinates are regressed through four 1×1 convolutions to obtain the feature points Classify the feature points into one of 64 possible positions or the case of no feature points.
[0018] As a further limitation of the first aspect of the present invention, in the descriptor extraction module, the output of the fourth basic block, the upsampling result of the output of the fifth basic block, and the convolution and upsampling results of the output of the sixth basic block are summed element by element, and then a convolution fusion block is used to obtain the descriptor
[0019] As a further limitation of the first aspect of the present invention, the second deep learning model is deployed using TensorRT. The PIDNet model adopted by the first deep learning model extracts static feature points and prior dynamic feature points from the image data, including: normalizing and preprocessing the pixels of the image data, extracting features based on the preprocessed image, performing feature fusion and enhancement, and then performing semantic prediction and classification to obtain a semantic segmentation result, and further obtaining static feature points and prior dynamic feature points.
[0020] As a further limitation of the first aspect of the present invention, the third deep learning model adopts the Light Glue model, and the feature point matching is performed by inputting into the Light Glue model, including:
[0021] The remaining feature points after removing the prior dynamic feature points and the corresponding descriptors are input into the Light Glue model. Each local feature contains a 2D point position and a visual descriptor. The state of each local feature is initialized to the corresponding visual descriptor and updated through multiple layers. Each layer consists of a self-attention unit and a cross-attention unit;
[0022] In the self-attention unit, each point focuses on other points in the same image, and the representation is enhanced through relative position encoding, enabling the model to capture the geometric relationships between points. In the cross-attention unit, each point focuses on points in another image to obtain cross-image context information;
[0023] Through a lightweight head, each layer calculates a matching score matrix between all point pairs in the two images, representing the similarity between point pairs in the two images; at the same time, calculates the matchability score of each point, indicating whether there is a corresponding point for this point in the other image; combines the similarity and matchability scores to generate a final assignment matrix for predicting matching point pairs, and according to the final assignment matrix, selects the point pairs with matching scores higher than the threshold as the matching results;
[0024] During the feature matching process, at the end of each layer, an adaptive depth mechanism is used to infer the confidence of the predicted assignment of each point. If the confidence of a point is high, indicating that its prediction result is reliable and is the final state, the LightGlue model stops reasoning in the early layer. If some points are predicted to be non-matchable, an adaptive width mechanism is used to exclude them from the input of the subsequent layer.
[0025] In the second aspect, the present invention provides a system for improving the accuracy of visual odometry based on deep learning.
[0026] A system for improving the accuracy of visual odometry based on deep learning includes the following processes:
[0027] A feature extraction unit, configured to: convert the acquired image data into grayscale image data, and extract feature points and descriptors corresponding to the feature points in the grayscale image data by using a pre-trained first deep learning model;
[0028] A prior dynamic feature point removal unit, configured to: extract static feature points and prior dynamic feature points in the image data by using a pre-trained second deep learning model, perform one-to-one mapping between the feature points extracted by the first deep learning model and the prior dynamic feature points, and remove the prior dynamic feature points from the feature points extracted by the first deep learning model;
[0029] A feature point matching unit, configured to: process the descriptors corresponding to the feature points remaining after removing the prior dynamic feature points, input them into a third deep learning model for feature point matching, perform epipolar constraint determination on the successfully matched feature points, remove the feature points that do not meet the requirements to obtain the final remaining feature points, and estimate the camera pose according to the final remaining feature points.
[0030] In a third aspect, the present invention provides a computer device, including: a processor and a computer-readable storage medium;
[0031] A processor, adapted to execute a computer program;
[0032] A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, the method for improving the accuracy of visual odometry based on deep learning as described in the first aspect of the present invention is implemented.
[0033] In a fourth aspect, the present invention provides a computer-readable storage medium, in which a computer program is stored, and the computer program is adapted to be loaded and executed by a processor to implement the method for improving the accuracy of visual odometry based on deep learning as described in the first aspect of the present invention.
[0034] In a fifth aspect, the present invention provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the method for improving the accuracy of visual odometry based on deep learning as described in the first aspect of the present invention is implemented.
[0035] Compared with the prior art, the beneficial effects of the present invention are:
[0036] 1. The present invention innovatively proposes a method for improving the accuracy of visual odometry based on deep learning. The feature points extracted by the first deep learning model are mapped one by one with the prior dynamic feature points. The prior dynamic feature points are removed from the feature points extracted by the first deep learning model, and the descriptors corresponding to the remaining feature points after removing the prior dynamic feature points are processed and input into the third deep learning model for feature point matching. This solves the problems of insufficient feature point extraction and large estimation errors in camera pose estimation based on visual sensors under complex conditions such as dynamic environments, low-texture areas, and illumination changes in the prior art, and improves the accuracy of visual odometry in complex environments.
[0037] 2. The XFeat model pays more attention to the global and semantic relevance of features. In images with noise interference, illumination changes, and dynamic scenes, it can still accurately capture key features, generate rich and highly discriminative features, thereby enhancing the richness and robustness of feature representation. Compared with the currently very popular image feature extraction SuperPoint model, the XFeat model adopts a lightweight CNN architecture and a carefully designed strategy. Its training process is more efficient, it performs well in data utilization efficiency, can achieve an ideal convergence effect with a relatively small amount of data, greatly saves training resources and time costs, and can perform real-time feature extraction on resource-constrained devices without relying on specific hardware optimization.
[0038] 3. Deploying the lightweight semantic segmentation PIDNet model using TensorRT can significantly improve the inference speed. It can combine operations such as continuous convolution and activation functions, reduce computational redundancy. When the model is converted into a TensorRT engine, through technical means such as quantization, the data precision can be accurately adjusted, and the resource occupancy can be reduced, so that the semantic segmentation result can be obtained more quickly, and the prior dynamic feature points can be removed. This can not only speed up the feature matching speed but also improve the accuracy of visual odometry and enhance the robustness of the system.
[0039] 4. The Light Glue model is extremely flexible and can work in cooperation with the feature extraction XFeat model. The Light Glue model adopts an adaptive mechanism. Facing image pairs of different complexities, it automatically adjusts the computational complexity. When the prediction result has reached a sufficiently high confidence level, it will choose to terminate the calculation process in advance. When encountering obviously mismatched points, it will discard them early, avoiding resource waste while improving efficiency. Moreover, it uses a multi-layer attention mechanism to capture complex relationships and is paired with an optimized descriptor to make feature matching more accurate.
[0040] 5. The method of using deep learning for feature point extraction and matching, combined with geometric methods to remove feature points that do not meet the requirements, can not only extract more discriminative and robust features, but also improve the matching accuracy by using context information and global structure, can adapt to texture sparsity, illumination changes and dynamic object interference in different scenarios, and enhances the ability of visual odometry to handle extreme or dynamic illumination conditions.
[0041] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0043] Figure 1 Schematic flow chart of the method for improving the accuracy of visual odometry based on deep learning provided in Embodiment 1 of the present invention;
[0044] Figure 2 Schematic diagram of the lightweight backbone network module provided in Embodiment 1 of the present invention;
[0045] Figure 3 Schematic diagram of the feature point extraction module and descriptor extraction module provided in Embodiment 1 of the present invention;
[0046] Figure 4 Flow chart of removing prior dynamic feature points provided in Embodiment 1 of the present invention;
[0047] Figure 5 Schematic diagram of the Light Glue model provided in Embodiment 1 of the present invention;
[0048] Figure 6 Schematic diagram of the epipolar geometry constraint provided in Embodiment 1 of the present invention;
[0049] Figure 7 Schematic diagram of the system for improving the accuracy of visual odometry based on deep learning provided in Embodiment 2 of the present invention;
[0050] Figure 8 Schematic diagram of a kind of computer device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0052] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0053] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0054] Embodiment 1:
[0055] Visual Odometry (VO) is a process of analyzing a relevant sequence of camera images to determine the position and orientation of an object (usually a robot or a vehicle) in three-dimensional space. It mainly relies on the image data collected by a camera and combines computer vision algorithms to infer the displacement and pose changes of the object in the environment. This process is similar to humans perceiving their own movement by observing the continuous changes in the surrounding environment. Visual Odometry focuses on estimating the motion trajectory of a camera in three-dimensional space by analyzing a sequence of consecutive images. Visual Odometry has broad application prospects in multiple fields, including: unmanned driving (helping autonomous vehicles achieve precise positioning and path planning), augmented reality (enabling precise positioning and tracking of virtual objects in real scenes), robot navigation (as an essential part of robot navigation, helping robots perceive the environment and plan paths), etc.
[0056] Currently, the application of deep learning in improving the accuracy of visual odometry has been widely involved in fields such as robot autonomous navigation, intelligent driving, unmanned aerial vehicles, and unmanned ships. However, there are still some challenges, such as the interpretability problem of deep learning models, robustness in complex scenarios, and real-time requirements, etc. In view of this, as Figure 1 shown, this implementation method proposes a method for improving the accuracy of visual odometry based on deep learning, including the following processes:
[0057] S1: Obtain image data through a camera, use the OpenCV library to convert the image into a grayscale image, and input each frame of the converted grayscale image into a pre-trained XFeat model to extract the feature points and corresponding descriptors of the image;
[0058] S2: Deploy the PIDNet model using TensorRT. Input the acquired image data into the model. According to prior knowledge, classify the feature points in the image into static feature points and prior dynamic feature points. Map the feature points extracted by the XFeat model one by one with the prior dynamic feature points, remove the prior dynamic feature points in the image, and only retain the remaining feature points extracted by the XFeat model. In this way, the obtained feature points retain both the static feature points and the feature points that are not detected by the PIDNet model but are detected by the XFeat model, maximizing the retention of more feature points;
[0059] S3: Perform processing operations such as normalization on the descriptors corresponding to the remaining feature points to ensure compliance with the input requirements of the LightGlue model. Input the feature points and the processed descriptors into the Light Glue model for feature point matching to complete the matching operation between images;
[0060] S4: Perform epipolar constraint on the successfully matched feature points to further determine the feature points in the image that do not meet the requirements; if there are dynamic feature points or feature points that change significantly under light changes in the image, perform removal operations, retain the remaining feature points, and use the remaining feature points to complete the estimation of the camera pose.
[0061] In step S1 of this implementation method, the XFeat model is used to extract the feature points and corresponding descriptors of the image. The XFeat model is a lightweight and efficient convolutional neural network architecture designed for fast image matching on resource-constrained devices; it reduces the computational overhead by optimizing the basic design of the convolutional network while ensuring efficient and accurate local feature extraction and matching; it mainly consists of a lightweight backbone network module, a feature point extraction module, and a descriptor extraction module.
[0062] Lightweight backbone network module: As Figure 2 shown, with a grayscale image is the input, in pixels, where H is the height, W is the width, C represents the number of channels, and the network adopts a lightweight design. Starting from the initial convolutional layer, the number of channels is minimized as much as possible. Starting from C = 4, as the spatial resolution decreases, the number of channels is increased at a three-fold rate instead of the traditional two-fold rate. This method ensures the minimum number of channels in the early layers while compensating for the reduction in the number of parameters in the entire network backbone. It not only significantly reduces the computational load in the early stage, especially for high-resolution images, but also optimizes the overall capacity of the network by more effectively managing the number of convolutional channels. The network consists of six basic blocks (the first basic block, the second basic block, the third basic block, the fourth basic block, the fifth basic block, and the sixth basic block in sequence). Each block contains a different number of channels, successively reducing the resolution by half and increasing the number of channels, finally reaching a spatial resolution of H / 32 × W / 32. Each basic block consists of a 2D convolution with a kernel size ranging from 1 to 3, ReLU activation, Batch Normalization, and a stride of 2 for reducing the spatial resolution. In addition, it includes a fusion block for multi-resolution features.
[0063] Specifically, Figure 2 in it, the input image is H, W, 1, the output of the first basic block is H, W, 4, the output of the second basic block is H / 2, W / 2, 8, the output of the third basic block is H / 4, W4 / , 24, the output of the fourth basic block is H / 8, W / 8, 64, the output of the fifth basic block is H / 16, W / 16, 4, and the output of the sixth basic block is H / 32, W / 32, 128.
[0064] The output H / 8, W / 8, 64 of the fourth basic block is upsampled, the output H / 16, W / 16, 4 of the fifth basic block is upsampled, and the output H / 32, W / 32, 128 of the sixth basic block is processed by a 1×1 convolution and upsampled. The above upsampled content is input into the fusion block to obtain an output of H / 8, W / 8, 64.
[0065] The feature point extraction module, as Figure 3 shown, uses the final encoder features with 1 / 8 of the original image resolution, but introduces a dedicated parallel branch for feature point detection, focusing on low-level image structures. To maintain the spatial resolution without sacrificing speed, the input image is represented as a two-dimensional grid of 8×8 pixels, each grid cell is reshaped into a 64-dimensional feature, and the feature point coordinates are regressed through 4 1×1 convolutions to obtain the feature points The feature points are classified as one of 64 possible positions or the case of no feature points.
[0066] The descriptor extraction module, as Figure 3 shown, extracts the descriptor by merging multi-scale features from the encoder (that isFigure 2 (the output of the fusion block in the middle). Specifically, the feature pyramid strategy is used, and consecutive convolutional blocks are applied to reduce the image resolution to 1 / 32 to increase the receptive field of the network. Then, bilinear upsampling and element-wise summation are used to merge the intermediate representations at three different scales (1 / 8, 1 / 16, 1 / 32). Finally, a convolutional fusion block is used to combine them into the final descriptor representation F.
[0067] In step S2 of this implementation method, removing the prior dynamic feature points can not only speed up the feature matching speed but also improve the accuracy of the visual odometer. The flowchart for removing the prior dynamic feature points is as Figure 4 shown. First, the pixels of the input image are pre-normalized to eliminate the numerical differences caused by different image acquisition devices and lighting conditions, ensuring the unity of the data scale. At the same time, according to the input size requirements preset by the PIDNet model, the length and width of the image are adjusted to an integer multiple of 32 to facilitate the efficient operation of subsequent network layers.
[0068] Immediately afterwards, it enters the feature extraction stage. The preprocessed image is fed into the input layer of the model. Here, the convolutional layer is paired with an activation function to quickly capture the initial image features and activate the latent visual information. Subsequently, it enters the backbone network composed of a series of cascaded residual blocks, continuously and deeply mining features at different scales, like a progressive filter, screening out key features at different scales. The unique three-branch structure of the PIDNet model plays a core role in this process: the proportion branch (P) focuses on maintaining high resolution and does not miss any details. Micro-elements such as the texture of leaves and the fine direction of hair can be accurately captured, preserving first-hand detailed information for subsequent precise segmentation; the integral branch (I) starts from a macroscopic perspective and constructs long-range dependence relationships by cleverly aggregating local and global context information. For example, judging whether a continuous area is a forest or a grassland, endowing the image with high-level semantic cognition; the derivative branch (D) is extremely sensitive to high-frequency changes and specifically locks in the mutations at the boundaries of objects, clearly outlining the contours of different objects and assisting in distinguishing the boundaries of objects.
[0069] Subsequently, it enters the feature fusion and enhancement stage. The ratio branch (P) utilizes the attention mechanism of the Pixel-attention-guided fusion module (Pag) to selectively obtain semantic information from the integral branch (I), integrating semantic knowledge into its own detailed features to make the detailed features more class-directed. The Boundary-attention-guided fusion module (Bag) focuses on the boundaries, relying more on the fine delineation of the ratio branch in the boundary areas, filling the remaining areas with the semantic information of the integral branch (I), and combining the boundary clues of the derivative branch (D) to make the boundaries of the fused feature map sharper and the semantic coherence stronger.
[0070] Finally, semantic prediction and classification are carried out. The fused feature map is fed into the classifier, a component composed of fully connected layers or convolutional layers, which carefully calculates the probability of each pixel belonging to each semantic category. Then, by means of threshold segmentation, the categories above the established threshold are preliminarily determined. Further optimization is carried out through connected component analysis to remove isolated misjudged pixels and regularize the same-class regions. Thus, the semantic segmentation of the entire image is completed. After obtaining the semantic segmentation result based on the PIDNet model, the prior dynamic feature points marked as categories such as people, vehicles, and animals are mutually mapped with the feature points extracted by the XFeat model, and the successfully mapped feature points are removed.
[0071] In step S3 of this implementation manner, Light Glue is a deep neural network for image feature matching, which is improved based on SuperGlue. By reexamining multiple design decisions, such as adopting relative position encoding, decoupling similarity and matchability prediction, and introducing an adaptive depth and width mechanism, this network can adaptively adjust the computational amount according to the difficulty of the image pair, with a faster inference speed for easily matchable image pairs, making it excellent in terms of accuracy, efficiency, and training simplicity.
[0072] The model diagram of Light Glue is as Figure 5 shown, consisting of L identical layers that jointly process two feature sets. Each layer consists of self-attention and cross-attention units for updating the representation of each point. Then, a classifier decides whether to stop inference at each level to avoid unnecessary calculations. In addition, Light Glue also introduces a lightweight head for calculating partial matching results from the updated representation. This structure enables Light Glue to reduce the computational complexity while maintaining high accuracy.
[0073] First, input the features after removing the prior dynamic feature points and the corresponding processed descriptors into the LightGlue model. Each local feature contains a 2D point position and a visual descriptor. Then, initialize the state of each local feature to its corresponding visual descriptor and update it through a series of layers, each layer consisting of a self-attention unit and a cross-attention unit. In the self-attention unit, each point attends to other points in the same image, enhancing the representation through relative position encoding, enabling the model to capture the geometric relationships between points. In the cross-attention unit, each point attends to points in another image to obtain cross-image context information. In this way, the model can learn the matching relationships between images. Then, through a lightweight head, a matching score matrix between all point pairs in the two images is calculated for each layer. represents the similarity between point pairs in images A and B. At the same time, calculate the matchability score σ i , indicating whether the point has a corresponding point in the other image. Combine the similarity and matchability scores to generate a soft partial assignment matrix P for predicting matching point pairs. Finally, according to the final assignment matrix P, select the point pairs with matching scores higher than the threshold as the matching results.
[0074] During the feature matching process, at the end of each layer, the Light Glue model uses an adaptive depth mechanism to infer the confidence of the predicted assignment for each point. If a point has a high confidence, indicating that its prediction result is reliable and is in the final state, the model can stop reasoning at the early layer. If some points are predicted to be non-matchable, use the adaptive width mechanism to exclude them from the input of subsequent layers, thereby reducing the computational amount.
[0075] In step S4 of this implementation method, first calculate the fundamental matrix F based on the already matched feature points in the image. Then, further identify the feature points in the image that do not meet the requirements with the help of the epipolar constraint theory. The epipolar constraint means that there is a specific geometric relationship between corresponding points on the imaging planes of two cameras, that is, for a point on one camera imaging plane, its corresponding point on the other camera imaging plane must be on the epipolar line determined by the fundamental matrix or the essential matrix. If there are dynamic feature points in the image and feature points that change significantly under illumination changes, remove them and only retain the remaining feature points. In this way, the accuracy and effectiveness of camera tracking in subsequent visual SLAM can be improved, providing a more reliable data basis for related visual SLAM tasks.
[0076] The epipolar geometry constraint is as Figure 6 shown. For two consecutive frames of images I 1 and I 2 , there are two cameras with centers O 1 and O 2。There is a point P in space, and its corresponding pixel in the image I 1 is p 1 , and its corresponding pixel in the image I 2 is p 2 . Since the epipolar geometry constraint is also called the five-point coplanarity constraint, it means that O 1 , O 2 , p 1 , p 2 , and P all lie on the same plane. Now let's explain the geometric relationship between them. First is the epipolar plane. The line connecting O 1 p 1 and the line connecting O 2 p 2 will intersect at point P in three-dimensional space. The epipolar plane is the plane determined by the three points O 1 , O 2 , and P at this time. Then are the epipoles and the baseline. The intersection points of the line connecting O 1 O 2 with the image planes I 1 , I 2 are respectively the epipoles e 1 , e 2 , and O 1 O 2 is called the baseline. Finally is the epipolar line. The lines l 1 , l 2 where the epipolar plane intersects the two image planes I 1 , l 2 are the epipolar lines.
[0077] With the above geometric relationship, the epipolar geometry constraint can be described. Ideally, the projection point p 1 of the spatial point P in the first frame I 1 can be transformed to the corresponding epipolar line l 2 in the second frame I 2 using the fundamental matrix F. Let the projection matching points of the spatial point P in the first and second frames be p 1 , p 2 respectively. Denote p 1 = [a 1 , b 1 , 1], p 2 = [a 2 , b 2 , 1] as the normalized coordinates in the camera coordinate system corresponding to the projection matching points p 1 , p 2 respectively. Calculate the epipolar line l 2 = Fp 1 through the fundamental matrix F. The formula is:
[0078]
[0079] Among them, the parameters of the epipolar line l 2 are A, B, and C respectively, and the fundamental matrix between two frames of images I 1 and I 2 is F.
[0080] The projected matching point is p 2 =[a 2 , b 2 , 1], and the equation of the epipolar line l 2 is: Aa + Bb + C = 0. Using the formula, the distance between the projected matching point p 2 and the epipolar line l 2 is:
[0081]
[0082] Since it is necessary to remove the dynamic feature points not screened by the lightweight semantic segmentation network PIDNet and the feature points with obvious changes under light changes, after obtaining the distance, a threshold is set to determine whether the matching feature points are the feature points that do not meet the requirements. The formula is:
[0083] d > th (3);
[0084] Among them, th is the set threshold, and the threshold size is 3 pixel points; if the distance is greater than the set threshold, then the matching point is the feature point that does not meet the requirements and needs to be removed.
[0085] Embodiment 2:
[0086] As Figure 7 shown, this implementation provides a system for improving the accuracy of visual odometry based on deep learning, including the following processes:
[0087] Feature extraction unit, configured to: convert the acquired image data into grayscale image data, and extract feature points and descriptors corresponding to the feature points in the grayscale image data using a pre-trained first deep learning model;
[0088] Prior dynamic feature point removal unit, configured to: extract static feature points and prior dynamic feature points in the image data using a pre-trained second deep learning model, map the feature points extracted by the first deep learning model one by one with the prior dynamic feature points, and remove the prior dynamic feature points from the feature points extracted by the first deep learning model;
[0089] The feature point matching unit is configured to: process the descriptors corresponding to the feature points remaining after removing the prior dynamic feature points, input them into the third deep learning model for feature point matching, perform epipolar constraint determination on the successfully matched feature points, remove the feature points that do not meet the requirements to obtain the final remaining feature points, and estimate the camera pose based on the final remaining feature points.
[0090] For the working methods of the specific units, please refer to the introduction in Embodiment 1 and will not be elaborated here.
[0091] It can be understood that the above-mentioned units can be separately or all combined into one or several other units to form, or some of them can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the system may also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.
[0092] According to another embodiment of this application, the system described in this embodiment can be constructed and the method of Embodiment 1 of this application can be realized by running a computer program (including program code) that can execute the respective steps involved in the corresponding method described in Embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0093] Embodiment 3:
[0094] As Figure 8 shown, this implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer-readable storage medium 1003. Among them, the processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 can be connected through a bus or other means.
[0095] Among them, the communication interface 1002 is used to receive and send data. The computer-readable storage medium 1003 can be stored in the memory of the electronic device. The computer-readable storage medium 1003 is used to store a computer program. The computer program includes program instructions. The processor 1001 is used to execute the program instructions stored in the computer-readable storage medium 1003.
[0096] The processor 1001 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device. It is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.
[0097] The processor 1001 is configured to execute the following process:
[0098] Obtain image data through the camera, convert the image into a grayscale image using the OpenCV library, input each frame of the converted grayscale image into the pre-trained XFeat model, and extract the feature points and corresponding descriptors of the image;
[0099] Deploy the PIDNet model using TensorRT, input the obtained image data into the model, classify the feature points in the image into static feature points and prior dynamic feature points according to prior knowledge, map the feature points extracted by the XFeat model one by one with the prior dynamic feature points, remove the prior dynamic feature points in the image, and only retain the remaining feature points extracted by the XFeat model. In this way, the obtained feature points retain both static feature points and feature points that are not detected by the PIDNet model but are detected by the XFeat model, maximizing the retention of more feature points;
[0100] Perform processing operations such as normalization on the descriptors corresponding to the remaining feature points to ensure compliance with the input requirements of the Light Glue model. Input the feature points and the processed descriptors into the Light Glue model for feature point matching to complete the matching operation between images;
[0101] Perform epipolar constraint on the successfully matched feature points to further determine the feature points that do not meet the requirements in the image. If there are dynamic feature points or feature points that change significantly under light changes in the image, perform removal operations, retain the remaining feature points, and use the remaining feature points to complete the estimation of the camera pose.
[0102] For the specific working method, see the introduction in Embodiment 1, which will not be elaborated here.
[0103] Embodiment 4:
[0104] This implementation provides a computer-readable storage medium (Memory). A computer-readable storage medium is a memory device in an electronic device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the electronic device.
[0105] Moreover, one or more instructions suitable for being loaded and executed by a processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0106] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor loads and executes one or more instructions stored in the computer-readable storage medium to implement the following process:
[0107] Obtain image data through a camera, use the OpenCV library to convert the image into a grayscale image, input each frame of the converted grayscale image into a pre-trained XFeat model, and extract the feature points and corresponding descriptors of the image;
[0108] Deploy the PIDNet model using TensorRT, input the obtained image data into the model, classify the feature points in the image into static feature points and prior dynamic feature points according to prior knowledge, map the feature points extracted by the XFeat model one by one with the prior dynamic feature points, remove the prior dynamic feature points in the image, and only retain the remaining feature points extracted by the XFeat model. In this way, the obtained feature points retain both static feature points and feature points that are not detected by the PIDNet model but are detected by the XFeat model, maximizing the retention of more feature points;
[0109] Perform processing operations such as normalization on the descriptors corresponding to the remaining feature points to ensure compliance with the input requirements of the Light Glue model, input the feature points and the processed descriptors into the Light Glue model for feature point matching, and complete the matching operation between images;
[0110] Perform epipolar constraint on the successfully matched feature points to further determine the feature points in the image that do not meet the requirements. If there are dynamic feature points or feature points that change significantly under light changes in the image, perform removal operations, retain the remaining feature points, and use the remaining feature points to complete the estimation of the camera pose.
[0111] For the specific working method, please refer to the introduction in Embodiment 1 and will not be elaborated here.
[0112] Embodiment 5:
[0113] This implementation provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the following process:
[0114] Obtain image data through a camera, use the OpenCV library to convert the image into a grayscale image, input each frame of the converted grayscale image into a pre-trained XFeat model, and extract feature points and corresponding descriptors of the image;
[0115] Deploy the PIDNet model using TensorRT, input the obtained image data into the model, classify the feature points in the image into static feature points and prior dynamic feature points according to prior knowledge, map the feature points extracted by the XFeat model one by one with the prior dynamic feature points, remove the prior dynamic feature points in the image, and only retain the remaining feature points extracted by the XFeat model. In this way, the obtained feature points retain both static feature points and feature points that are not detected by the PIDNet model but are detected by the XFeat model, maximizing the retention of more feature points;
[0116] Perform processing operations such as normalization on the descriptors corresponding to the remaining feature points to ensure compliance with the input requirements of the Light Glue model, input the feature points and the processed descriptors into the Light Glue model for feature point matching, and complete the matching operation between images;
[0117] Perform epipolar constraint on the successfully matched feature points to further determine the feature points in the image that do not meet the requirements. If there are dynamic feature points or feature points that change significantly under light changes in the image, perform removal operations, retain the remaining feature points, and use the remaining feature points to complete the estimation of the camera pose.
[0118] For the specific working method, please refer to the introduction in Embodiment 1 and will not be elaborated here.
[0119] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0120] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data processing device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state disk (SSD)), etc.
[0121] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for improving the accuracy of visual odometer based on deep learning, characterized in that: The process includes: Convert the acquired image data into grayscale image data, and use a pre-trained first deep learning model to extract feature points in the grayscale image data and descriptors corresponding to the feature points; Using a pre-trained second deep learning model to extract static feature points and a priori dynamic feature points in the image data, mapping the feature points extracted by the first deep learning model with the a priori dynamic feature points one-to-one, and removing the a priori dynamic feature points from the feature points extracted by the first deep learning model; The descriptors corresponding to the feature points remaining after removing the prior dynamic feature points are processed and input into the third deep learning model for feature point matching. The successfully matched feature points are subjected to polar constraint judgment and the feature points that do not meet the requirements are removed to obtain the final remaining feature points. The camera pose is estimated based on the final remaining feature points.
2. The method for improving the accuracy of visual odometer based on deep learning as claimed in claim 1, characterized in that: The first deep learning model adopts the XFeat model, including: a lightweight backbone network module, a feature point extraction module and a descriptor extraction module; The lightweight backbone network module includes six basic blocks connected in sequence, each of which halves the resolution of an input image, the spatial resolution of the input image of the first basic block is H×W, and the spatial resolution of the output image of the last basic block is H / 32×W / 32. The output of the fourth basic block, the up-sampling result of the output of the fifth basic block, and the convolution and up-sampling result of the output of the sixth basic block are fused to obtain a multi-scale fusion feature; The spatial resolution of the multi-scale fusion feature is H / 8×W / 8, the feature point extraction module extracts feature points based on the feature image output by the fourth basic block, and the descriptor extraction module extracts descriptors based on the multi-scale fusion feature.
3. The method for improving the accuracy of visual odometer based on deep learning as claimed in claim 2, characterized in that: In the feature point extraction module, the multi-scale fusion features are represented as a two-dimensional grid of 8×8 pixels. Each grid unit is reshaped into a 64-dimensional feature. The feature point coordinates are regressed through four 1×1 convolutions to obtain the feature point Classify feature points into one of 64 possible locations or into the absence of feature points.
4. The method for improving the accuracy of visual odometer based on deep learning according to claim 2 or 3, characterized in that: In the descriptor extraction module, the output of the fourth basic block, the upsampled result of the output of the fifth basic block, and the convolution and upsampled result of the output of the sixth basic block are summed element by element, and then the convolution fusion block is used to obtain the descriptor 5. The method for improving the accuracy of visual odometer based on deep learning as claimed in claim 1, characterized in that: A second deep learning model is deployed using TensorRT. The second deep learning model adopts a PIDNet model. The pre-trained PIDNet model is used to extract static feature points and prior dynamic feature points in the image data, including: normalizing and preprocessing the pixels of the image data, extracting features based on the preprocessed image, fusing and enhancing the features, performing semantic prediction and classification, obtaining semantic segmentation results, and then obtaining static feature points and prior dynamic feature points.
6. The method for improving the accuracy of visual odometer based on deep learning as claimed in claim 1, characterized in that: The third deep learning model adopts the Light Glue model and is input into the Light Glue model for feature point matching, including: Remove the feature points and corresponding descriptors remaining after the prior dynamic feature points are added to the Light Glue model. Each local feature contains a 2D point position and a visual descriptor. The state of each local feature is initialized to the corresponding visual descriptor and updated through multiple layers. Each layer consists of a self-attention unit and a cross-attention unit. In the self-attention unit, each point pays attention to other points in the same image, and the representation is enhanced by relative position encoding, so that the model can capture the geometric relationship between points. In the cross-attention unit, each point pays attention to points in another image to obtain contextual information across images; Through a lightweight head, each layer calculates a matching score matrix between all point pairs in the two images, indicating the similarity between the point pairs in the two images; at the same time, it calculates the matchability score of each point, indicating whether the point has a corresponding point in the other image; the similarity and matchability scores are combined to generate a final assignment matrix for predicting matching point pairs. According to the final assignment matrix, point pairs with matching scores higher than the threshold are selected as matching results; During the feature matching process, at the end of each layer, the adaptive depth mechanism is used to infer the confidence assigned to each point prediction. If a point has a high confidence, indicating that its prediction result is reliable and is the final state, the Light Glue model stops reasoning at the early layer. If some points are predicted to be unmatched, they are removed from the input of subsequent layers using the adaptive width mechanism.
7. A system for improving the accuracy of visual odometer based on deep learning, characterized in that: The process includes: A feature extraction unit is configured to: convert the acquired image data into grayscale image data, and extract feature points in the grayscale image data and descriptors corresponding to the feature points using a pre-trained first deep learning model; The prior dynamic feature point removal unit is configured to: extract static feature points and prior dynamic feature points from the image data using a pre-trained second deep learning model, perform one-to-one mapping between the feature points extracted by the first deep learning model and the prior dynamic feature points, and remove the prior dynamic feature points from the feature points extracted by the first deep learning model; The feature point matching unit is configured to: process the descriptors corresponding to the feature points remaining after removing the prior dynamic feature points, input them into the third deep learning model for feature point matching, perform polar constraint judgment on the successfully matched feature points and remove the feature points that do not meet the requirements to obtain the final remaining feature points, and estimate the camera pose based on the final remaining feature points.
8. A computer device, characterized in that: include: A processor and a computer readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the method for improving the accuracy of a visual odometer based on deep learning as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is suitable for being loaded by a processor and executing the method for improving the accuracy of visual odometer based on deep learning as described in any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method for improving the accuracy of visual odometer based on deep learning as described in any one of claims 1 to 6.
Citation Information
Cited By
Camera pose regression estimation system and method based on long-time-sequence arbitrary point tracking
CN121259094A