VSLAM method and device for indoor low-texture environment

By combining the VSLAM method of Shi-Tomasi corner point and LSD line segment extractor, the problem of insufficient VSLAM robustness in indoor low-texture environments is solved, and higher positioning accuracy and trajectory accuracy are achieved.

CN120070788AActive Publication Date: 2025-05-30HAINAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510145260.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The existing VSLAM method has problems such as insufficient robustness and failed positioning and mapping in indoor low texture environments.

Method used

Using a VSLAM method combining point features and line features, the relative rotation and translation optimization process based on the Harmanton world hypothesis is constructed to optimize the pose matrix to achieve high-quality positioning and graph building through Shi-Tomasi corner point feature extraction and the near-line merger and short-line culling algorithm of the LSD line segment extractor.

Benefits of technology

The robustness and positioning accuracy of the VSLAM system are significantly improved in indoor low-texture environments. Compared with the benchmark method ORB-SLAM3, the track accuracy is improved by an average of 29.89%, and a 59.85% increase in specific environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070788A_ABST
    Figure CN120070788A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine vision and image processing, in particular to a VSLAM method and device for an indoor low-texture environment, and the method comprises the steps: obtaining an original color image, converting the original color image into a grayscale image only containing single-channel information, carrying out the Shii-Tomasi corner feature extraction to obtain point features, and carrying out the line feature extraction to obtain a VSLAM image; the method comprises the following steps: extracting point features and line features, carrying out near line merging and short line elimination on the extracted line features, and constructing a new pose matrix through an optimized rotation matrix and an optimal translation vector based on the extracted point features and line features and a relative rotation and translation optimization process of a Karman world hypothesis; the obtained pose matrix is used for re-projection error optimization of the point features and the line features, positioning and mapping in the indoor low-texture environment are achieved, and high-quality extraction and matching of the feature points and the feature lines can be ensured in the indoor low-texture environment; and the precision and robustness of the system can be improved by more efficiently utilizing the structural information in an indoor artificial regular environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision and image processing, and particularly to a VSLAM method and device for an indoor low-texture environment. Background Art

[0002] In recent years, with the rapid development of the SLAM (Simultaneous Localization and Mapping) technology, the vision SLAM (VSLAM) system based on camera sensors has attracted more and more extensive attention from researchers because it has lower hardware costs compared with Lidar SLAM systems and laser SLAM systems, and camera sensors can obtain richer information. The VSLAM system mainly relies on the multi-view geometric information in the scene, combines computer vision algorithms, and accurately estimates the position of the robot and simultaneously generates a 3D map of the environment. Therefore, it is widely used in the fields of 3D vision reconstruction, mobile robots, 3D holographic technology, visual positioning, AR / VR applications, and autonomous driving technologies such as micro aerial vehicles (MAVs).

[0003] Currently, researchers have developed VSLAM methods based on various vision sensors such as monocular cameras, stereo binocular cameras, RGB-D cameras, and event cameras. The existing VSLAM methods are mainly divided into feature-based methods and direct methods. Among them, feature-based methods have always been the focus of research because they rely on classical computer vision algorithms to extract geometric features and have stronger robustness than direct methods that rely on image pixel values when dealing with light changes. In addition, the demand for computing resources of feature-based methods is also lower than that of direct methods. Since the assumptions required by direct methods are often difficult to meet in general camera sensors and environments, feature-based methods are more favored by researchers. However, in some low-texture scenes of artificial indoor environments, feature-based methods often have uneven or insufficient point feature distributions, resulting in abnormalities in the tracking of the VSLAM system and ultimately may lead to failures in positioning and mapping.

[0004] To address this challenge, many studies have introduced other geometric elements in multi-view geometry, such as line segments, planes, or vanishing points, into the VSLAM system, which can significantly enhance the robustness of the system. These geometric elements can provide additional information support for scenes with sparse features, thereby improving the performance of the system in low-texture environments. Usually, many indoor environments are artificial regular structures where three mutually orthogonal line segments can be extracted, enabling the establishment of a general Manhattan world hypothesis. However, most existing structured VSLAM methods rely on the global Manhattan world hypothesis and need to estimate the vanishing point for each frame. Nevertheless, this hypothesis is not always applicable in actual indoor environments, which may lead to time-consuming operations and poor optimization effects. In addition, global Manhattan world optimization usually relies on the initial frame estimation and is easily affected by the initial error in subsequent optimization. For point features in low-texture environments, these methods do not optimize or fully utilize them, but still use traditional processing methods or directly discard them. In most scenarios, the number and matching success rate of line features are usually much lower than those of point features. Therefore, line features are more often used as an auxiliary means to supplement the deficiencies of point features. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to propose a VSLAM method and device for indoor low-texture environments to solve the limitations and insufficient robustness of traditional feature visual SLAM in some indoor low-texture environments.

[0006] Based on the above purpose, the present invention provides a VSLAM method for indoor low-texture environments, including the following steps:

[0007] S1. Obtain the original color image and convert it into a grayscale image containing only single-channel information;

[0008] S2. Perform Shi-Tomasi corner feature extraction on the grayscale image to obtain point features;

[0009] S3. Perform line feature extraction on the grayscale image, and perform near-line merging and short-line elimination on the extracted line features;

[0010] S4. Based on the extracted point features and line features, through the relative rotation and translation optimization process based on the Manhattan world hypothesis, construct a new pose matrix using the optimized rotation matrix and the optimal translation vector;

[0011] S5. Use the obtained pose matrix for reprojection error optimization of point features and line features to achieve positioning and mapping in indoor low-texture environments.

[0012] Preferably, performing Shi-Tomasi corner feature extraction on the grayscale image to obtain point features includes:

[0013] Calculate the image gradient I of the grayscale image in the horizontal direction x and the image gradient I in the vertical direction y , calculate the autocorrelation matrix Q of the grayscale image through the image gradient, and calculate the minimum eigenvalue of the autocorrelation matrix Q;

[0014]

[0015] window is the local window size, and the response function R of the Shi-Tomasi corner detection algorithm Shi-Tomasi is used to determine whether the corner at a certain position in the grayscale image is a qualified Shi-Tomasi corner based on the calculated minimum eigenvalue. The response function R Shi-Tomasi is usually a certain threshold, λ 1 and λ 2 are the eigenvalues calculated through the autocorrelation matrix Q;

[0016] R Shi-Tomasi = min(λ 1 , λ 2 ).

[0017] Preferably, the Shi-Tomasi corner feature extraction of the grayscale image adopts a grid-based and pyramid hierarchical extraction strategy. Among them, the extraction strategy increases the grid size, halves the number of pyramid layers, and halves the maximum number of extracted corners compared with the benchmark version ORB-SLAM3

[0018] Preferably, the near-line merging and short-line elimination of the extracted line features include:

[0019] Adopt the near-line merging and short-line elimination algorithm based on the LSD line extractor to perform near-line merging and short-line elimination.

[0020] Preferably, before the relative rotation and translation optimization process based on the Manhattan world hypothesis, the method also includes the execution logic of the relative rotation optimization pre-step, specifically including:

[0021] First, judge whether the current frame ID is a multiple of an integer to realize the discontinuous Manhattan vanishing point estimation. If the current frame performs the Manhattan vanishing point estimation, judge whether the line clusters clustered according to the Manhattan vanishing point in the current frame meet the condition that all three line clusters exist and each line cluster contains at least three line segments. If this condition is met, calculate the main direction curdk of the line clusters in the current frame and judge whether the main direction curdk of the line clusters in the current frame and the main direction lastdk of the previous line clusters both exist, and the frame ID interval between the two frames where the main directions of the line clusters are located cannot be greater than 15 frames. When all the above conditions are met, perform the relative rotation optimization.

[0022] Preferably, in step S4, the process of relative rotation includes:

[0023] For each line cluster clustered by the Manhattan vanishing point, first calculate the normal vector of each line segment in the line cluster in the current camera coordinate system. Calculate the direction vector l of each line segment through the cross product operation of the starting point and the ending point in homogeneous coordinate form, and normalize the direction vector l. Then, transform the processed line segment direction vector l through the transpose K T of the camera intrinsic matrix K to obtain the normal vector s of the line segment in the current camera coordinate system i , and normalize it to construct a matrix S containing all the normal vectors of the line segments in the three line clusters;

[0024] Transpose the matrix S and perform SVD singular value decomposition to obtain three directions orthogonal to all the normal vectors of the line segments in the three line clusters, that is, the main directions of the three line clusters, and obtain the main direction curdk of the line clusters in the current frame;

[0025] After calculating curdk each time, the system saves and updates it to construct a relative rotation optimization method for the pose between two frames.

[0026] Preferably, the relative rotation optimization method for the pose between two frames includes:

[0027] For the previous main direction lastdk of the line clusters, calculate its transformation direction in the current frame:

[0028]

[0029] where R is the rotation matrix of the current frame, k represents the serial numbers of the main directions of the three different line clusters, is the inverse of the rotation matrix of the frame where the previous main direction of the line clusters is located, is the previous main direction lastdk of the line clusters. By calculating the angular change with the current Manhattan vanishing point direction δ to define the relative rotation angle error amount, and adopt a prior method based on angular weight and distance weight to adjust the internal arrangement order of curdk;

[0030] According to the cost function of minimizing the overall relative rotation error, list the corresponding Jacobian matrix calculation formula:

[0031]

[0032] Since the partial derivative of the translation part is 0, the final Jacobian matrix for each error term θ i is a 1×6 matrix, where the first 3 columns correspond to the translation part and the last three columns correspond to the rotation part. The final Jacobian matrix is expressed as:

[0033]

[0034] The complete Jacobian matrix representing the rotational error terms in each of the different pairwise orthogonal directions is as follows:

[0035]

[0036] Define the above relative rotational error as an edge in the Levenberg - Marquardt optimization algorithm within G2O and add it to the vertex defined by the current frame pose, indicating that this relative rotational error will impose an optimization constraint on the pose of the current frame. The three main directions of the line clusters used in the optimization have the same weight, indicating that they have the same importance during the optimization process, thereby obtaining the optimized rotation matrix R. init .

[0037] Preferably, in step S4, the translation optimization process includes:

[0038] Define a random number generator to randomly select a feature point index in each iteration, and set an appropriate maximum number of iterations and inlier threshold. In each iteration, query whether the feature point with the randomly selected index is a feature point associated with a valid map point.

[0039] After selecting a feature point associated with a valid map point, construct a linear least - squares problem. First, obtain the image coordinates k p of the randomly selected feature point and the world coordinates X w of its associated map point. Then, according to the principal point coordinates (c x , c y ) in the camera intrinsic matrix, calculate the coordinates (u, v) of the random feature point relative to the camera optical center as shown in the following formula:

[0040] u = k p .x - c x , v = k p .y - c y

[0041] Next, construct the matrices M and vector n required for solving the least - squares problem based on the projection model:

[0042]

[0043] where M T is the transpose of matrix M, f x and f y are the focal lengths of the camera, σ -2 is the weight of the random feature point corresponding to the image pyramid, and X w , Y w and Z w are the world coordinates of the map point associated with the random feature point. Solve the following equation according to the least - squares method to obtain the optimal translation vector t.init ;

[0044] t init =(M T M) -1 M T n.

[0045] Preferably, step S5 further includes:

[0046] The present invention adopts the bundle adjustment technology to globally optimize the poses of key frames, map points, and map lines, calculates the three-dimensional coordinates of map points and map lines through triangulation, and uses the global bundle adjustment technology to optimize the poses of key frames, map points, and map lines to improve the accuracy of the map.

[0047] The present invention also provides a VSLAM device for indoor low-texture environments, which is used to execute the above-mentioned VSLAM method for indoor low-texture environments.

[0048] Advantages of the present invention:

[0049] Compared with the benchmark method ORB-SLAM3, the method of the present invention has an average improvement of 29.89% in the APE (RMSE) trajectory accuracy in several EuRoC sequences tested. In addition, for indoor low-texture environments such as V201, the method of the present invention also achieves an average improvement of 59.85% in trajectory accuracy compared with the current state-of-the-art algorithms. Compared with other VSLAM methods, the method of the present invention can not only ensure the high-quality extraction and matching of feature points and feature lines in indoor low-texture environments, but also more efficiently utilize the structural information to improve the accuracy and robustness of the system in indoor artificial regular environments. The method of the present invention demonstrates excellent performance in indoor low-texture environments and makes great contributions to the development of VSLAM systems applied to this scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0051] Figure 1 It is a schematic diagram of the overall framework of the embodiment of the present invention;

[0052] Figure 2 It is a comparison diagram of improving the FAST corner extraction to Shi-Tomasi corner extraction in the embodiment of the present invention;

[0053] Figure 3Schematic diagram of the near-line merging operation logic according to an embodiment of the present invention;

[0054] Figure 4 Schematic diagram of the Manhattan world vanishing point for optimizing the relative rotation of the camera frame pose according to an embodiment of the present invention;

[0055] Figure 5 Flowchart of the pre-step for relative rotation optimization according to an embodiment of the present invention;

[0056] Figure 6 Visual comparison diagram of relative rotation optimization when conditions are met according to an embodiment of the present invention;

[0057] Figure 7 Partial indoor low-texture environment map according to an embodiment of the present invention;

[0058] Figure 8 APE accuracy error line chart and trajectory chart for comparison with the true trajectory according to an embodiment of the present invention;

[0059] Figure 9 Comparison result diagram of all point feature matches between two frames according to an embodiment of the present invention;

[0060] Figure 10 Comparison result diagram of high-quality point feature matches between two frames according to an embodiment of the present invention;

[0061] Figure 11 Effect comparison diagram after improving the line feature extraction according to an embodiment of the present invention. Detailed implementation manners

[0062] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments.

[0063] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs. The "first", "second", and similar terms used in the present invention do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0064] An embodiment of this specification provides a VSLAM method for an indoor low-texture environment. The overall process is as Figure 1 shown, including:

[0065] 1. Input image

[0066] First, the VSLAM system captures an original color image from a binocular camera and converts it into a grayscale image containing only single-channel information.

[0067] 2. Perform Shi-Tomasi point feature extraction on the input image in the VSLAM system

[0068] To avoid the situation that the inlier feature extraction module in the improved VSLAM system cannot balance robustness and real-time performance when performing point feature extraction on an indoor low-texture environment. The present invention proposes a grid-based multi-layer pyramid front-end point feature extraction method based on the Shi-Tomasi corner detection algorithm.

[0069] First, for the input grayscale image, in the front-end point feature extraction part of the VSLAM system of the present invention, the image gradient I of the grayscale image in the horizontal direction will be calculated x and the image gradient I of the grayscale image in the vertical direction y . The autocorrelation matrix Q of the grayscale image is calculated through the image gradient, and the minimum eigenvalue of the autocorrelation matrix Q is calculated. window is the local window size (usually 3×3 pixels).

[0070]

[0071] λ 1 and λ 2 are the eigenvalues calculated through the autocorrelation matrix Q, and the response function R of the Shi-Tomasi corner detection algorithm Shi-Tomasi is used to judge whether a corner at a certain position in the grayscale image is a qualified Shi-Tomasi corner together with the calculated minimum eigenvalue. The response function R Shi-Tomasi is usually a determined threshold.

[0072] R Shi-Tomasi = min(λ 1 , λ 2 ) (2)

[0073] To ensure the uniform distribution and real-time performance of the improved front-end point feature extraction corner points in the VSLAM system of the present invention, the present invention introduces a grid-based and pyramid hierarchical extraction strategy to ensure that sufficient Shi-Tomasi corner points can be uniformly extracted within the grayscale image, thereby reducing the matching failures caused by the drastic movement of the camera. In ORB-SLAM3, the extraction of ORB feature points relies on the detection of FAST corner points in the image. Although the FAST corner point detection is famous for its fast detection speed and can extract a large number of corner points in a short time, and thus is very suitable for real-time application scenarios, its performance in low-texture environments is often not ideal. In low-texture environments, due to the scarcity and uneven distribution of distinguishable feature points in the image, the FAST corner points are easily interfered by noise, resulting in unstable extracted feature points and even a large number of incorrect matches, which may significantly affect the robustness and accuracy of the system. Although the FAST corner points can play an important role in applications with high real-time requirements, in low-texture environments, its dependence on and limitations of feature points are undoubtedly revealed.

[0074] In contrast, the Shi-Tomasi corner point detection algorithm is an improvement of the classical Harris corner point detection algorithm [], mainly by weighting the gradient information of the local area of the image to more accurately detect the stable corner points in the image. In low-texture environments, although there are fewer detectable feature points in the image, the Shi-Tomasi algorithm has stronger feature stability, which means that it can effectively avoid the extraction of incorrect and noisy point features by the FAST corner points in areas with less texture. More importantly, the stability of the Shi-Tomasi corner points enables the system to maintain relatively stable tracking and positioning even in an environment with sparse feature points, thereby reducing the matching failures and trajectory drift problems caused by the low quality of feature points in low-texture environments.

[0075] To ensure the real-time performance of the VSLAM system while improving the stability of feature points, the present invention optimizes the Shi-Tomasi corner point extraction algorithm and simplifies the grid-based and pyramid hierarchical strategies in ORB-SLAM3. Specifically, the present invention appropriately increases the size of the grid to reduce the computational burden by reducing the number of grid cells, thereby improving the speed of feature point extraction; at the same time, the present invention reduces the number of pyramid layers to reduce the overhead of multi-scale image processing. The purpose of these simplification operations is to ensure the stable operation of the system in low-texture environments by reducing the computational complexity. Although the Shi-Tomasi corner point extraction consumes more computational resources than the traditional FAST corner point method, its high-quality feature point characteristics enable the present invention to reduce the maximum number of feature points extracted, further improve the quality of the extracted points and maintain a uniform distribution, while meeting the requirements of real-time performance.

[0076] Through these optimizations, the present invention can effectively improve the positioning accuracy in a low-texture environment and keep the computational cost close to the baseline version. These adjustments not only enhance the robustness of the VSLAM system, enabling the system to cope with the challenges in a low-texture environment, but also ensure that the improved system maintains good real-time performance and reliability while providing higher positioning accuracy. Therefore, the introduction and related optimizations of the Shi-Tomasi corner points successfully address the challenges of point feature extraction in a low-texture environment and effectively improve the overall performance of the VSLAM system.

[0077] As Figure 2 shown, the number of point features extracted by the feature extraction pyramid of the present invention and the number of pyramid layers are half of those of the original ORB-SLAM3 version.

[0078] 3. Perform near-line merging and short-line elimination on the line features extracted in the VSLAM system

[0079] The VSLAM method of the present invention adds the extraction and matching process of line features to the baseline VSLAM method. Aiming at the problem that the redundant line segments of the traditional LSD detector are prone to cause false matching of line features and increase the computational cost when dealing with scenes with rich line textures, the present invention proposes a near-line merging and short-line elimination algorithm based on the LSD line segment extractor, aiming to filter out redundant line features through post-processing steps, thereby reducing the false matching of line features in the VSLAM system and reducing the time required for the VSLAM system to calculate the LBD descriptor of line features. This algorithm strategy realizes the near-line merging and short-line elimination operations through several steps in Table 1.

[0080] Table 1 Algorithm: Line Feature Processing

[0081]

[0082]

[0083] The operation logic of the present invention for near-line merging of the line features extracted by the front end can be represented by Figure 3 to indicate.

[0084] In Figure 3 , line segment A and line segment B represent the line segments that the present invention will judge whether to merge. L1, L2, L3, and L4 are the endpoint distances between the two line segments A and B. d is the distance between the midpoint of line segment A and line segment B, and θ is the included angle of the line segments calculated through the direction vectors.

[0085] The front-end feature extraction process of the VSLAM system is completed through step 2 and step 3.

[0086] The present invention selects the LSD line segment extractor as the method for extracting line segment features. LSD (Line Segment Detector) is a parameter-free high-precision line segment detection method, which not only has strong real-time performance but also can well preserve the details and geometric features of line segments in images. Therefore, it has become the most widely used line segment extractor in current VSLAM methods. The present invention selects the LBD line segment descriptor as the descriptor for line segment features. The LBD descriptor is designed specifically for line segment features and has characteristics such as rotational invariance, high computational efficiency, and strong robustness, and is suitable for line segment matching and feature description in VSLAM systems.

[0087] However, in a low-texture environment, the VSLAM system usually faces problems of sparse features and inaccurate matching, especially during the extraction and matching processes of line features. Although LSD (line segment extractor) can effectively extract line segment features from images, in a low-texture environment, there is often a lack of sufficient distinct features in the image, resulting in fewer and lower-quality line segments being extracted. In addition, line segments in a low-texture environment often exhibit irregular distributions, and the LSD detector may detect redundant line segments or there may be cases of false matching, thus affecting the performance of the system.

[0088] To address the above problems, the near-line merging and short-line elimination algorithm proposed by the present invention plays a crucial role. The core of this algorithm is to screen the detected line segments through a post-processing step to eliminate redundant features and improve the matching quality of line features. Through the near-line merging operation, the present invention can merge adjacent line segments with similar directions into a longer and more stable line segment, thereby reducing the matching errors caused by short and discontinuous line segments. At the same time, the short-line elimination operation can effectively remove unrepresentative short-line features, which is particularly important in a low-texture environment because short-line features are usually more susceptible to noise and edge effects. These processing steps can significantly improve the robustness and positioning accuracy of the system by reducing useless redundant features and false matches. Specifically, the near-line merging operation reduces the possibility of false matching, enabling line segment features to be more stably matched in a low-texture environment, while the short-line elimination reduces the time required to calculate the LBD descriptor, thereby improving the real-time performance of the system.

[0089] 4. Construct relative rotation and translation optimization based on the Manhattan world

[0090] After the front-end feature extraction process of the VSLAM system is completed, the relative rotation and translation optimization process based on the Manhattan world hypothesis will be carried out.

[0091] To ensure the real-time performance and effectiveness of the VSLAM system, the present invention adopts a fast Manhattan vanishing point estimation method based on two line segments. This method can quickly and accurately estimate the Manhattan vanishing point in space from two randomly selected line segments extracted from the input image, and obtain the globally optimal Manhattan vanishing point group through detailed search.

[0092] Based on the above-mentioned vanishing point estimation method under the Manhattan world assumption, the present invention proposes the following relative rotation and translation optimization methods based on the Manhattan world.

[0093] Figure 4 Schematic diagram for optimizing the relative rotation of the camera frame through the Manhattan world vanishing point.

[0094] The present invention integrates the Manhattan vanishing point estimation method into the frame construction process of the front-end feature extraction part of the proposed VSLAM system, that is, the subsequent process of the feature extraction part. At the same time, to ensure the real-time performance of the system and avoid forcibly using the relative rotation and relative translation optimization methods based on the Manhattan vanishing point in a non-Manhattan structure environment, the present invention designs a pre-execution logic for relative rotation optimization, and its execution logic flowchart is as Figure 5 shown.

[0095] In the pre-execution logic of relative rotation optimization, first, it is judged whether the current frame ID is a multiple of an integer to achieve intermittent Manhattan vanishing point estimation. If the Manhattan vanishing point is estimated for the current frame, it is judged whether the line clusters clustered according to the Manhattan vanishing point in the current frame meet the condition that all three types of line clusters exist and each type of line cluster contains at least three line segments. If this condition is met, the main direction curdk of the line cluster of the current frame is calculated, and it is judged whether both the main direction curdk of the line cluster of the current frame and the main direction lastdk of the previous line cluster exist, and the frame ID interval between the two frames where the main directions of the line clusters are located cannot be greater than 15 frames. Finally, when all the above conditions are met, the relative rotation optimization within the present invention is confirmed to be executed.

[0096] When the relative rotation optimization of the system of the present invention is executed, three line segments of different colors will appear in the visualization window interface, representing the line segments within different line clusters respectively. Figure 6 Shows the visualization window without relative rotation optimization and with relative rotation optimization in progress.

[0097] The following part will detail the process of calculating the main direction curdk of the line cluster for each frame. For each line cluster clustered through the Manhattan vanishing point, the present invention first calculates the normal vector of each line segment in the current camera coordinate system. Specifically, for the starting point p s and the ending point p e of each line segment in the pixel coordinates of the image, the present invention converts them into homogeneous coordinate form.

[0098]

[0099] Among them, x s and y s are the two-dimensional pixel coordinates of the starting point of the line segment, and x e and y e are the two-dimensional pixel coordinates of the ending point of the line segment. Then, the direction vector l of each line segment is calculated by performing a cross product operation on the starting point and the ending point in homogeneous coordinate form, and the direction vector l is normalized.

[0100]

[0101] The processed line segment direction vector l is transformed through the transpose K T of the camera internal parameter matrix K to obtain the normal vector s i of the line segment in the current camera coordinate system, and it is normalized.

[0102]

[0103] The present invention constructs a matrix S containing the normal vectors of all line segments in three line clusters. Assuming there are n line segments in each line cluster, the matrix S can be represented as a 3×n matrix.

[0104]

[0105] Among them, s nn is the normal vector of the nth line segment in the nth type of line cluster. Then, the matrix S is transposed and subjected to SVD singular value decomposition, and three directions orthogonal to the normal vectors of all line segments in the three line clusters can be obtained, that is, the main directions of the three line clusters. The present invention can obtain the main direction curdk of the line cluster in the current frame.

[0106] After calculating curdk each time, the system will save and update it, and at most only retain the previous main direction lastdk of the line cluster. When the present invention obtains the current Manhattan vanishing point direction vps, the current main direction curdk of the line cluster, the previous main direction lastdk of the line cluster, and the inverse of the camera pose rotation matrix of the frame where the previous main direction of the line cluster is located, the present invention can construct a relative pose rotation optimization method between two frames.

[0107] Specifically, the present invention defines the error of the relative rotation optimization method as a non-linear least squares problem based on the Levenberg-Marquardt optimization framework. For the previous main direction lastdk of the line cluster, the present invention first calculates its transformation direction in the current frame:

[0108]

[0109] Wherein, R is the rotation matrix of the current frame, k represents the serial numbers of the main directions of three different line clusters, is the inverse of the rotation matrix of the frame where the main direction of the line cluster was located last time, is the main direction lastdk of the line cluster last time. The present invention can calculate the angular change with the current Manhattan vanishing point direction δ to define the relative rotation angle error amount. Since the Manhattan vanishing point estimation of the present invention is random, it is easy to have a random arrangement within the Manhattan vanishing point direction δ, resulting in a random arrangement within curdk when calculating the main direction of the line cluster, which is not conducive to error calculation. Therefore, a prior method based on angle weight and distance weight is adopted to adjust the internal arrangement order of curdk. This prior method will calculate the angular difference of the Manhattan vanishing point vector and the Euclidean distance of the Manhattan vanishing point between the currently estimated Manhattan vanishing point direction vps and the previous Manhattan vanishing point direction lastvps, and obtain the combination with the smallest change by continuously sorting the vectors within the currently estimated Manhattan vanishing point direction vps to become the new Manhattan vanishing point direction vps. Formula (9) is introduced to avoid large angular errors caused by incorrect adjustment of the prior method to a certain extent, and formula (10) reflects the overall angular error minimization cost function.

[0110]

[0111] Wherein, δ is the current Manhattan vanishing point direction, represents the transpose of the ith Manhattan vanishing point direction vector, and i represents the serial numbers of three different Manhattan vanishing point directions, represents the main direction of the line cluster with the same serial number as the Manhattan vanishing point direction, θ i represents the relative rotation angle error of each different pair of mutually orthogonal directions, represents calculate its skew-symmetric matrix. According to the overall relative rotation error minimization cost function, the present invention can list the corresponding Jacobian matrix calculation formula.

[0112]

[0113] Since the partial derivative of the translation part is 0, the final Jacobian matrix for each error term θ i is a 1×6 matrix, where the first 3 columns correspond to the translation part (zero vector), and the last three columns correspond to the rotation part. The following is the final Jacobian matrix.

[0114]

[0115] The complete Jacobian matrix representing the rotation error terms in each different pair of mutually orthogonal directions can be expressed as:

[0116]

[0117] Define the above relative rotation error as an edge in the Levenberg - Marquardt optimization algorithm in G2O and add it to the vertex defined by the current frame pose, indicating that this relative rotation error will impose an optimization constraint on the pose of the current frame. The three main directions of the line clusters used in the optimization have the same weight, indicating that they have the same importance in the optimization process. According to the above method, the optimized rotation matrix R can be obtained. init 。

[0118] After obtaining the optimized rotation matrix R init , the present invention will also define a strategy for optimizing the translation vector through a framework based on the Random Sample Consensus (RANSAC) algorithm. First, define a random number generator for randomly selecting feature point indices in each iteration, and set appropriate maximum iteration times and inlier thresholds. In each iteration, query whether the feature point selected randomly is a feature point associated with a valid map point through the randomly selected feature point index.

[0119] When a feature point associated with a valid map point is selected, the present invention will construct a linear least - squares problem. First, obtain the image coordinates k p of the randomly selected feature point and the world coordinates X w of its associated map point. Then, according to the principal point coordinates (c x , c y ) in the camera intrinsic matrix, the present invention can calculate the coordinates (u, v) of the random feature point relative to the camera optical center as shown in the following formula:

[0120] u = k p .x - c x , v = k p .x - c y (14)

[0121] Next, the present invention will construct the matrix M and the vector n required for solving the least - squares problem based on the projection model.

[0122]

[0123] Among them, M T is the transpose of the matrix M, f x and f y are the focal lengths of the camera, σ -2 is the weight of the image pyramid corresponding to the random feature point, X w , Y w and Z wThe world coordinates of the map points associated with the random feature points. The present invention will solve the following equation according to the least squares method to obtain the optimal translation vector t init .

[0124] t init =(M T M) -1 M T n (17)

[0125] For the valid feature points selected in the current frame, i represents the index of the current valid feature point, is the world coordinates of the map point associated with the feature point with index i. The present invention will calculate the coordinates of the map points associated with these feature points in the current camera coordinate system based on the rotation matrix R init and the translation vector t init obtained by the above method:

[0126]

[0127] The present invention will also calculate the projection error and count the number of inliers u init under the current rotation matrix R init and the translation vector t i and v i represent the coordinates of the feature point with index i with respect to the camera optical center. The projection error formula is calculated as follows:

[0128]

[0129] The sum of the squares of the overall projection error is:

[0130]

[0131] If the error is less than the set threshold, the feature point is considered an inlier. If the total number of inliers in the current iteration exceeds the previous optimal result, the translation vector t init calculated by the least squares method this time is updated as the optimal translation vector.

[0132] The optimized rotation matrix R init and the optimal translation vector t init will construct a new pose matrix T init , and this pose matrix T init will be used as the initial pose for subsequent reprojection error optimization in the VSLAM system of the present invention, so as to reduce the adverse effects brought by the Manhattan world mandatory constraint and improve the accuracy and real-time performance of the reprojection error.

[0133] 5. Motion Estimation and Map Optimization

[0134] Pose matrix T of the above steps init For the reprojection error optimization of point features and line features, the VSLAM system can quickly obtain the accurate current camera pose and determine key frames. The present invention adopts the Bundle Adjustment technology to globally optimize the poses of key frames, map points, and map lines. In addition, the system also calculates the three-dimensional coordinates of map points and map lines through Triangulation and uses the global Bundle Adjustment technology to optimize the poses of key frames, map points, and map lines to improve the accuracy of the map. When the system detects a loop closure, it uses the previous key frames and inactive maps to optimize the poses of the current key frames and active maps, thereby reducing the cumulative error. Even if a key frame is lost during the tracking process, the system can quickly recover and continue the tracking process through relocalization and map fusion technologies. The comprehensive application of these technologies enables the VSLAM system of the present invention to achieve efficient and stable positioning and mapping in an indoor low-texture environment.

[0135] In order to test the performance of the proposed VSLAM system in terms of accuracy and real-time performance, the present invention uses the EuRoC dataset (European Robotics Challenge Dataset) released by the Micro Aerial Vehicle Laboratory (ASL) of ETH Zurich to compare the system of the present invention with the ORB-SLAM3 system. The EuRoC dataset is a dataset widely used to evaluate VSLAM systems and Visual Inertial Odometry (VIO) algorithms, including high-quality datasets collected during the flight of micro aerial vehicles (MAVs) equipped with cameras and IMU sensors in indoor and industrial environments, including indoor factories and indoor rooms, which have the complexity of real scenes, such as changing lighting, low-texture areas, and fast movements.

[0136] The experimental environment of this embodiment is a laptop equipped with an Intel(R) Core(TM) i9-10980HK@2.4GHz 16-core CPU. All experiments are carried out in an 8-core Ubuntu18.04 system within a VMware virtual machine. In the evaluation of trajectory accuracy, the present invention uses the EVO tool to conduct an index evaluation of the accuracy of ORB-SLAM3 and the proposed method. The present invention selects the root mean square error RMSE of the Absolute Trajectory Error APE as the standard for judging the trajectory accuracy of the present invention. The Absolute Trajectory Error APE can be expressed by the following formula.

[0137]

[0138] Wherein represents the estimated camera pose at the i-th moment, T i represents the true camera pose at the i-th moment, and the calculated error e i represents the difference between the estimated camera pose and the true camera pose. In order to obtain a measure of the overall error, the present invention represents it by calculating the root mean square value RMSE of APE, and the specific method is shown in the following formula.

[0139]

[0140] where n is the total number of poses, and RMSE evaluates the overall error size by squaring and averaging the errors at each moment and then taking the square root. The smaller the value, the more accurate the estimated trajectory.

[0141] 1. Comparison of trajectory accuracy

[0142] The present invention will select some sequences in the famous EuRoC dataset for experimental testing and comparison. The selected sequences include left and right monocular grayscale images and IMU data. The size of each image is (752, 480), the image frame rate is 20Hz, the IMU frame rate is 200Hz, and the true trajectory is included. Figure 7 Shows the challenging low-texture part in the experimental scene within the selected sequence.

[0143] Table 2 statistics the APE (RMSE) trajectory accuracy of various methods under the selected sequences. Among them, the method of this embodiment, the OpenVINS method, and the ORB-SLAM3 method run in the binocular IMU mode, and the UV-SLAM and UL-SLAM run in the monocular IMU mode. The best experimental numerical results are highlighted by bold display in the table.

[0144] Table 2 APE (RMSE) accuracy comparison

[0145]

[0146]

[0147] The data in Table 2 show that, compared with ORB-SLAM3, the method of the present invention has an average improvement of 29.89% in overall accuracy. Even when evaluating different improved modules separately, the trajectory accuracy of each module has also been improved compared with ORB-SLAM3. In a dark environment where it is difficult to fully extract point features, such as the MH05 dataset, even though the trajectory accuracy of the method of the present invention is not as good as that of the state-of-the-art UL-SLAM, it can still achieve an accuracy improvement of 48.53% compared with ORB-SLAM3. In an indoor environment, due to factors such as white walls and mirror reflections, low-texture scenes often appear. By combining local Manhattan vanishing point information, the method of the present invention can also provide a more accurate initial pose estimate for reprojection error optimization in complex or non-atypical Manhattan-structured indoor environments. Therefore, the trajectory accuracy has also been improved by 18.91% and 27.29% respectively in the indoor V201 and V203 datasets compared with the sub-optimal ORB-SLAM3 algorithm. These results verify the adaptability and accuracy improvement of the method of the present invention in different scenarios.

[0148] In addition, the present invention also uses the EVO tool to compare the estimated trajectories of ORB-SLAM3 and the method of this embodiment with the actual trajectory on the MH05 dataset, as Figure 8 shown. The figure includes an APE accuracy error line graph and a trajectory graph, and the dashed line represents the true value.

[0149] As can be seen from Figure 8 (a) and (b), the absolute trajectory error of ORB-SLAM3 mainly concentrates between 0.1 and 0.15, while the method of the present invention further reduces the error to between 0.06 and 0.08. In addition, Figure 8 (c) and (d) further verify the accuracy advantage of the method of the present invention in a low-light scenario. Compared with the ground truth, there is an obvious error in the trajectory of ORB-SLAM3, and the error value may reach about 0.2, while the maximum trajectory error of the method of the present invention remains at about 0.1, and the error of some paths is reduced by about 50%. These results fully illustrate the accuracy improvement of the method of the present invention under low-light illumination.

[0150] 2. Comparison of Point Feature Matching Effect and Line Feature Extraction Effect

[0151] The present invention selects two adjacent frames of images on the EuRoC dataset MH01 for experiments on improved point feature matching. All point matching images and high-quality point matching images under the point feature matching used in ORB-SLAM3 and the improved point feature matching of this embodiment are shown in Figures 9 - 10 . The maximum number of extracted feature points is set to 600, and the threshold part follows the settings in ORB-SLAM3 and the system of this embodiment.

[0152] InFigures 9 - 10 Among them, the improvement of the feature matching effect is very obvious. Specifically, among all the point matching results, the number of matching points using the FAST corner extraction method is 11, while the number of matching points using the Shi-Tomasi corner extraction method has increased significantly to 130, showing the significant advantage of the Shi-Tomasi corner in a low-texture environment, which is improved by about 12 times. For high-quality point matching, the number of matching points using the FAST corner extraction method is 6, while the number of matching points using the Shi-Tomasi corner extraction method has increased to 97, which is improved by about 16 times, showing a significant improvement in the quality of feature point matching. We judge which feature point matches belong to high-quality point matches by calculating the Hamming distance between the descriptors of the matched point features and comparing it with the minimum distance. More high-quality point feature matches help to improve the stability of the VSLAM system, and thus improve the tracking accuracy and the accuracy of landmark estimation.

[0153] Figure 11 It reflects the effect of the improved line feature extraction. The method of the present invention filters out many redundant short line segments and near line segments, retains the long line segments that are more conducive to inter-frame matching, reduces the computational overhead of the line feature part, and improves the stability of line feature matching.

[0154] 3. Real-time analysis

[0155] The present invention also conducts real-time experiments and analyzes on ORB-SLAM3 and the method of the present invention. The present invention runs in the EuRoC MH01 dataset and counts the time consumed by ORB-SLAM3 and the method of the present invention in each module. The real-time performance of the method proposed by the present invention is reflected by comparing the overall and time consumption between each module with ORB-SLAM3.

[0156] Table 3 Comparison of time consumption per frame on average

[0157]

[0158]

[0159] Among them, the unit of time consumption in Table 3 is ms / f, which means the time consumed per frame on average. The addition of line features and Manhattan optimization has a certain impact on the time consumption of the method of the present invention, but compared with ORB-SLAM3, the overall system of the present invention can still meet the real-time requirements.

[0160] The method of the present invention consumes 71.40 ms per frame on average in the tracking module, while the time consumption of ORB-SLAM3 is 31.99 ms, showing a certain increase. This is because the present invention introduces line feature matching operations and vanishing point estimation during the tracking process. However, the time consumption of the tracking module is still within an acceptable range and can meet the requirements of a real-time vSLAM system. In the local mapping module, the method of the present invention consumes 377.56 ms, which is an increase compared to 278.74 ms of ORB-SLAM3. The main reason for the time increase is that the present invention adopts more complex line feature processing and Manhattan optimization, increasing the accuracy and complexity of the local map. The performance improvement brought by these enhancement strategies is sufficient to compensate for the additional time overhead. The time consumption of the loop detection module is relatively low. The method of the present invention only consumes 1.19 ms per frame on average, while it is 4.35 ms for ORB-SLAM3. The method of the present invention shows higher efficiency in loop detection.

[0161] Overall, the average time consumption per frame of ORB-SLAM3 is 315.08 ms, while that of the method of the present invention is 450.15 ms. Although the computing time has increased by nearly 135 ms, after introducing more geometric information and optimization strategies, the positioning accuracy and robustness of the system of the present invention have been effectively improved, and it can still maintain relatively good real-time performance.

[0162] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present invention is limited to these examples; under the idea of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above, which are not provided in detail for the sake of brevity. Any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A VSLAM method for indoor low-texture environments, characterized in that: The following steps are involved: S1, obtain the original color image and convert it into a grayscale image containing only single channel information; S2, extract Shi-Tomasi corner features from the grayscale image to obtain point features; S3, extracting line features from the grayscale image, and merging near lines and removing short lines from the extracted line features; S4, based on the extracted point features and line features, a relative rotation and translation optimization process based on the Harmanton world hypothesis is performed to construct a new pose matrix through the optimized rotation matrix and the optimal translation vector; S5. The obtained pose matrix is ​​used to optimize the reprojection error of point features and line features to achieve positioning and mapping in indoor low-texture environments.

2. The VSLAM method for indoor low-texture environment according to claim 1, characterized in that: The Shi-Tomasi corner point feature extraction is performed on the grayscale image to obtain point features including: Calculate the image gradient I of the grayscale image in the horizontal direction x and the vertical image gradient I y , calculate the autocorrelation matrix Q of the grayscale image through the image gradient, and calculate the minimum eigenvalue of the autocorrelation matrix Q; Window is the local window size, and the response function R of the Shi-Tomasi corner detection algorithm is used. Shi-Tomasi The calculated minimum eigenvalue is used to determine whether a corner point at a certain position in the grayscale image is a qualified Shi-Tomasi corner point. The response function R Shi-Tomasi It is usually a certain threshold, λ1 and λ2 are the eigenvalues ​​calculated by the autocorrelation matrix Q; R Shi-Tomasi =min(λ1,λ2).

3. The VSLAM method for indoor low-texture environment according to claim 1, characterized in that: The Shi-Tomasi corner feature extraction for the grayscale image adopts a gridding and pyramid layered extraction strategy, wherein the extraction strategy increases the grid size, reduces the number of pyramid layers by half, and reduces the maximum number of extracted corner points by half compared to the benchmark version ORB-SLAM3.

4. The VSLAM method for indoor low-texture environment according to claim 1, characterized in that: The step of merging near lines and removing short lines from the extracted line features includes: The near-line merging and short-line eliminating algorithm based on LSD line segment extractor is used to merge near-line and eliminate short-line.

5. The VSLAM method for indoor low-texture environment according to claim 1, characterized in that: Before the relative rotation and translation optimization process based on the Harmanton world hypothesis, the method also includes relative rotation optimization pre-step execution logic, specifically including: First, determine whether the current frame ID is a multiple of an integer to achieve discontinuous Manhattan vanishing point estimation. If the current frame performs Manhattan vanishing point estimation, determine whether the line clusters clustered according to the Manhattan vanishing point of the current frame meet the requirements of the existence of all three line clusters and each line cluster contains at least three line segments. If this condition is met, calculate the main direction curdk of the current frame line cluster and determine whether the main direction curdk of the current frame line cluster and the main direction lastdk of the previous line cluster both exist, and the frame ID interval of the frames where the two main directions of the line clusters are located cannot be greater than 15 frames. When all the above conditions are met, perform relative rotation optimization.

6. The VSLAM method for indoor low-texture environment according to claim 1, characterized in that: In step S4, the relative rotation process includes: For each line cluster that passes through the Manhattan vanishing point cluster, first calculate the normal vector of each line segment in the line cluster in the current camera coordinate system, calculate the direction vector l of each line segment by aligning the start and end points in the secondary coordinate form, and normalize the direction vector l. The processed line segment direction vector l is then transformed through the transpose K of the camera intrinsic parameter matrix K. T Convert and get the normal vector s of the line segment in the current camera coordinate system i , and normalize it to construct a matrix S containing all the line segment normal vectors in the three line clusters; Transpose the matrix S and perform SVD singular value decomposition to obtain three directions orthogonal to the normal vectors of all line segments in the three line clusters, namely the main directions of the three line clusters, and obtain the main direction curdk of the line cluster of the current frame; Each time curdk is calculated, the system saves and updates it to build a relative rotation optimization method between the two frames.

7. The VSLAM method for indoor low-texture environment according to claim 6, characterized in that: The relative rotation optimization method between the two frames includes: For the last main direction of the line cluster lastdk, ​​calculate its transformation direction in the current frame: Among them, R is the rotation matrix of the current frame, k represents the main direction numbers of three different line clusters, is the inverse of the rotation matrix of the frame where the main direction of the last line cluster is located, is the main direction of the last line cluster lastdk. By calculating The relative rotation angle error is defined by the angle change of the current Manhattan vanishing point direction δ, and a priori method based on angle weight and distance weight is adopted to adjust the internal arrangement order of Curdk; According to the cost function of minimizing the overall relative rotation error, the corresponding Jacobian matrix calculation formula is listed: Since the translation part derivative is 0, the final Jacobian matrix is i is a 1×6 matrix, where the first three columns correspond to the translation part and the last three columns correspond to the rotation part. The final Jacobian matrix is ​​expressed as: The complete Jacobian matrix representing the rotation error terms in each different pairwise orthogonal direction is given by: The above relative rotation error is defined as an edge in the Levenberg-Marquardt optimization algorithm in G2O and added to the vertex defined by the current frame pose, indicating that the relative rotation error will impose optimization constraints on the current frame pose. The three line cluster main directions used in the optimization have the same weight, indicating that they have the same importance in the optimization process, thereby obtaining the optimized rotation matrix R init .

8. The VSLAM method for indoor low-texture environment according to claim 1, characterized in that: A strategy for optimizing the translation vector is defined by a random sampling consensus algorithm framework. In step S4, the translation optimization process includes: Define a random number generator to randomly select feature point indexes in each iteration, and set appropriate maximum number of iterations and internal point thresholds. In each iteration, query whether the randomly selected feature point index is a feature point associated with a valid map point. After selecting the feature points associated with the valid map points, a linear least squares problem is constructed. First, the image coordinates k of the randomly selected feature points are obtained. p and the world coordinate X of its associated map point w Then, according to the principal point coordinates in the camera intrinsic parameter matrix (c x ,c y ), calculate the coordinates (u, v) of the random feature point relative to the camera optical center, as shown in the following formula: u=k p .x-c x ,v=k p .x-c y Next, the matrix M and vector n required to solve the least squares problem are constructed based on the projection model: Among them, M T is the transpose of matrix M, f x and f y is the focal length of the camera, σ -2 is the weight of the random feature point corresponding to the image pyramid, X w , Y w and Z w The world coordinates of the map point associated with the random feature point are solved by the least squares method to obtain the optimal translation vector t init ; t init =(M T M) -1 M T n。 9. The VSLAM method for indoor low-texture environment according to claim 1, characterized in that: Step S5 also includes: The present invention adopts bundle adjustment technology to globally optimize the position and posture of key frames, map points and map lines, calculates the three-dimensional coordinates of map points and map lines by triangulation, and uses global bundle adjustment technology to optimize the position and posture of key frames, map points and map lines to improve the accuracy of the map.

10. A VSLAM device for indoor low-texture environment, characterized in that: The device is used to execute the VSLAM method for indoor low-texture environment as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Indoor binocular visual odometer method based on point, line and surface features

    CN114004900A

  • Visual SLAM (Simultaneous Localization and Mapping) method based on point-line features in low-texture environment

    CN114627309A

  • Indoor VSLAM illumination adaptive adjustment method and device under restricted resources

    CN118864332A