A High-Precision Visual SLAM Method Based on Improved Quadtree Feature Point Extraction
Through the improved quadtree filtering algorithm and BEBLID feature point description algorithm, the problem of insufficient feature point extraction and description accuracy in visual SLAM is solved, and higher image matching accuracy and more accurate pose estimation are achieved, which improves the positioning and mapping accuracy of visual SLAM.
Patent Information
- Application Number
- CN202310190463.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-02
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-03-02
AI Technical Summary
The existing visual SLAM methods have problems of insufficient accuracy and low efficiency in the process of feature point extraction and description, resulting in low positioning and mapping accuracy, which is prone to loss and system crashes.
The improved quadtree filtering algorithm is used to filter feature points, retain high-response feature points and eliminate isolated weak-response feature points. At the same time, the BEBLID algorithm is introduced for feature point description, and feature extraction is performed using AdaBoost algorithm and sampled image blocks of different sizes.
It improves the accuracy of image matching, achieves more accurate pose estimation and positioning, and improves the trajectory accuracy of visual SLAM and the optimal trajectory and map construction capabilities over a long period of time.
Smart Images

Figure CN116245949B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and relates to a high-precision visual SLAM method based on improved quadtree feature point extraction. Background Art
[0002] In recent years, thanks to the rapid development of computer technology, communication technology and artificial intelligence technology, significant breakthroughs have occurred in related technologies of computer vision, including image matching, face recognition, augmented reality, autonomous driving and 3D reconstruction, etc. And there is an important requirement in augmented reality, autonomous driving and 3D reconstruction technologies: precise positioning. At present, for the outdoor positioning problem, with the help of 5G communication technology and navigation satellite systems such as the Global Positioning System (GPS), it has reached a quite precise level, and basically solved the positioning problems of most outdoor applications. However, in indoor environments, such as indoor parking lots, large warehouses, shopping malls, restaurants, etc., in these scenarios, many current indoor positioning methods, such as infrared positioning, wireless network communication technology positioning and ultra-wideband positioning, etc., cannot achieve a satisfactory effect. Similarly, for applications like augmented reality that need to determine their own position and map construction in real time in an unknown environment, the above solutions cannot achieve ideal effects.
[0003] Visual SLAM has received extensive attention from the academic and industrial communities because of its simple hardware structure, rich collected information that can be further processed using deep learning, and the ability to perform positioning and mapping simultaneously. Visual SLAM mainly realizes the task of simultaneous localization and mapping in an unknown environment through the environmental information collected by a camera. Currently, many excellent results have been achieved in the fields of intelligent robots, augmented reality and autonomous driving. Visual SLAM includes a front-end visual odometer and a back-end loop detection, optimization and mapping. Among them, the visual odometer part mainly estimates the change of the camera pose by extracting and processing environmental information, so as to realize the positioning function. Therefore, how to extract and process environmental information more efficiently and accurately is a research focus in the field of visual SLAM.
[0004] At present, ORB-SLAM, which is widely used, uses ORB for feature point extraction and description in the image processing stage. However, in order to ensure the uniformity of feature point distribution during quadtree screening, it retains a large number of isolated weak response feature points while eliminating many feature points with higher response values, resulting in a significant decrease in image matching accuracy. At the same time, due to the problems of slow description speed and low accuracy in the rBRIEF feature point description algorithm in the ORB algorithm, visual SLAM is prone to losing track or even crashing during camera tracking. These existing problems greatly affect the positioning and mapping accuracy of visual SLAM. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a high-precision visual SLAM method based on improved quadtree feature point extraction to achieve higher trajectory accuracy and positioning accuracy.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A high-precision visual SLAM method based on improved quadtree feature point extraction, characterized in that: the method includes the following steps:
[0008] S1. Collect image information in the environment through a camera;
[0009] S2. Convert the RGB image collected by the camera into a grayscale image, and at the same time construct an image pyramid for each image and divide each layer of the image into grids;
[0010] S3. Determine the number of feature points to be extracted from each layer of the pyramid image according to the area of each layer of the pyramid image and the set number of feature points;
[0011] S4. Extract an excessive number of feature points within the grids divided in each layer of the pyramid image, and then use the improved quadtree to screen the feature points extracted from each layer of the image;
[0012] S5. Use the BEBLID algorithm to describe the feature points screened in step S4;
[0013] S6. Match the images according to the feature points extracted between two adjacent frames of images, then estimate the camera pose through the PnP algorithm and the matching relationship of the feature points, and finally adjust and optimize the estimated camera pose by the method of minimum reprojection error;
[0014] S7. Construct all the motion information and camera observation information into an optimization problem with a larger scale and size, and use bundle adjustment to optimize and solve it to obtain the optimal trajectory and map in a long time.
[0015] Further, step S3 is specifically:
[0016] First, calculate the total area of the layers of the entire image pyramid:
[0017]
[0018] In the formula, H and W respectively represent the height and width of the bottom-layer image, s represents the scaling factor of the image pyramid, and m represents the number of layers of the pyramid;
[0019] Then, calculate the number of feature points per unit area according to the number of feature points to be extracted from each image:
[0020]
[0021] In the formula, Num represents the number of feature points to be extracted from each image.
[0022] Finally, determine the number of feature points to be extracted from each layer according to the area size of each layer of the pyramid. The number of feature points that should be allocated to the i-th layer is:
[0023]
[0024] Furthermore, in step S4, the improved quadtree is used to screen the feature points extracted from each layer of the image. Specifically: First, divide the quadtree according to the number of feature points required for this layer of the image; Second, count the response values of the feature points in each node of the quadtree; Third, determine the adaptive response threshold of this layer of the image according to the median and average values of the response values of the feature points in this layer of the image; Finally, use the calculated adaptive response threshold to screen the feature points in each node, and on the premise of ensuring the uniformity of the distribution of the feature points, retain as many high-response feature points as possible and eliminate isolated weak-response feature points to achieve more accurate pose estimation.
[0025] Judge the number of feature points in each node of the quadtree. If there is only one feature point in a certain node, then judge it according to the adaptive response threshold. If the response value of the feature point is less than the adaptive response threshold, then eliminate this feature point to reduce the impact on the image matching accuracy;
[0026] If there are multiple feature points in the node and the response values of all feature points are less than the adaptive response threshold, then retain the feature point with the highest response value.
[0027] Furthermore, step S5 is specifically as follows:
[0028] First, extract sampled image patches of different sizes around the feature points, and then use the sampled image patch feature extraction function f(x) and threshold T corresponding to each weak classifier in the AdaBoost algorithm to obtain h(x);
[0029] The extraction function of the sampled image patches around the feature points is as follows:
[0030]
[0031] In the formula, p1 and p2 respectively represent the centers of the image patches extracted by each weak classifier, s represents the side length of the image patch, I(p) and I(q) respectively represent the gray values of each pixel point;
[0032] The value of h(x) represents the similarity of the sampled image patch structures selected by each weak classifier in AdaBoost. If the average gray difference between two sampled image patches is less than the threshold T, it is +1, otherwise it is -1, as follows:
[0033]
[0034] Secondly, in order to obtain the binary feature descriptor, it is necessary to judge the value of h(f,T). If h(f,T) is greater than 0, the corresponding binary descriptor bit is taken as 1, otherwise it is 0;
[0035] Finally, by training the descriptors of all feature points in the dataset, optimizing the loss function, and obtaining the pixel positions, image patch sizes, and thresholds of the best descriptor sampling point pairs, the best BEBLID binary descriptor pattern is obtained; the loss function is:
[0036]
[0037] In the formula, N represents the sampled image patches corresponding to N pairs of feature points in the training dataset; x i and y i respectively represent the image patches corresponding to two certain feature points in the training dataset; γ represents the learning rate; k represents the k-th weak classifier, that is, the k-th bit corresponding to the final 256-bit descriptor; h k (x i ) and h k (y i ) respectively represent the similarity of the structures of two sampled image patches selected by the k-th weak classifier; l i represents the label, l i ∈{-1,1}, when l i =1, it means that the image patches corresponding to the two feature points have the same image structure, and when l i =-1, it means that the corresponding image structures are different.
[0038] The beneficial effects of the present invention are as follows: By adopting an improved quadtree screening algorithm to screen feature points, on the premise of ensuring the uniform distribution of feature points, as many high-response feature points as possible are retained, improving the accuracy of image matching and achieving more accurate pose estimation. At the same time, a feature point description algorithm BEBLID based on the AdaBoost algorithm is introduced, which can describe feature points faster. During the description process, sampling image blocks of different sizes are selected, realizing a calculation method similar to the gradient. Therefore, a more accurate feature point description is obtained, achieving higher trajectory accuracy on most dataset sequences and more accurate positioning.
[0039] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0041] Figure 1 Schematic diagram for the construction of the image pyramid and feature point extraction;
[0042] Figure 2 Flowchart for calculating the adaptive response threshold;
[0043] Figure 3 Schematic diagram for the calculation of the BEBLID descriptor;
[0044] Figure 4 Overall flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically, and the following embodiments and the features in the embodiments can be combined with each other without conflict.
[0046] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams rather than actual physical diagrams, and should not be construed as a limitation to the present invention; for better illustration of the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.
[0047] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the attached drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0048] As Figure 4 shown, it is a high-precision visual SLAM method based on improved quadtree feature point extraction. The specific implementation process of this method is as follows:
[0049] S1. Use the camera on the robot or vehicle to obtain the image information in the environment.
[0050] S2. Convert the RGB image obtained by the camera into a grayscale image, and construct an 8-layer image pyramid for each image, and perform grid division on each layer of the pyramid image.
[0051] S3. According to the area of each layer of the pyramid image and the set number of feature points, determine the number of feature points extracted from each layer of the pyramid image, and use the quadtree to divide the nodes on each layer of the image.
[0052] The steps to determine the number of feature points extracted from each layer of the image are as follows:
[0053] S31. First, calculate the total area of the layers of the entire image pyramid:
[0054]
[0055] In the formula, H and W respectively represent the height and width of the bottom layer image, s represents the scaling factor of the image pyramid, and m represents the number of layers of the pyramid;
[0056] S32. Then, calculate the number of feature points per unit area according to the number of feature points to be extracted from each image:
[0057]
[0058] S33. Finally, determine the number of feature points to be extracted for each layer according to the area size of the images at each layer of the pyramid. The number of feature points that should be allocated to the i-th layer is:
[0059]
[0060] S4. Within the grids divided in the images at each layer of the pyramid, use the FAST corner extraction algorithm to set double thresholds for over-extracting feature points, so as to extract as many features as possible, which is convenient for subsequent improvement of the quadtree to screen and eliminate feature points, thereby realizing the uniform distribution of feature points, as Figure 1 shown.
[0061] Subsequently, use the improved quadtree to screen the feature points extracted from the images at each layer:
[0062] First, divide the quadtree according to the number of feature points required for the image at this layer; secondly, count the response values of the feature points in each node of the quadtree; thirdly, determine the adaptive response threshold for the image at this layer according to the median and average values of the response values of the feature points in this layer of the image (take the smaller value of the median and the average value), as Figure 2 shown; finally, use the calculated adaptive response threshold to screen the feature points in each node, and on the premise of ensuring the uniform distribution of feature points, retain as many high-response feature points as possible and eliminate isolated weak-response feature points to achieve more accurate pose estimation.
[0063] In the calculation process of the adaptive response threshold of the present invention, different thresholds are set for different layers of the image pyramid. The purpose of statistically calculating the average value and median value of the response values of the feature points on each layer of the image is to retain as many high-response feature points as possible. If the average value of the distribution of the response values of the feature points on this layer of the image is greater than the median value, it means that there are more high-response feature points in this layer of the image, and vice versa, the proportion of low-response feature points is larger. By selecting the smaller one of the average value and the median value as the threshold for this layer of the image, more feature points can be retained, the repeatability of feature extraction can be improved, and a better image matching effect can be achieved. In image feature detection, the time consumed by feature point extraction accounts for a larger proportion, but the present invention does not increase the number of feature points extracted, so the quadtree screening algorithm based on the adaptive response threshold will not have a great impact on the efficiency of feature detection.
[0064] After the node division is completed, the number of feature points of each node is judged. If there is only one feature point in the node, it means that the feature point is an isolated point, and it is judged according to the adaptive response threshold: if the response value of the feature point is less than the set threshold, it means that the feature point is not only isolated but also an unobvious weak response feature point. Such feature points can often be detected in the previous frame, but may not be detected in the next frame, so it will lead to the occurrence of false matching. Therefore, such feature points need to be removed to reduce the impact on the image matching accuracy. When there are multiple feature points in the node, it is necessary to further judge the response values of the feature points. If there are feature points higher than the set threshold in the node, then the feature points are screened according to the normal adaptive threshold screening algorithm, that is, the feature points with response values higher than the adaptive response threshold are retained, and the feature points with response values lower than the adaptive response threshold are removed; but if the response values of all feature points in the node are less than the adaptive response threshold, then the feature point with the highest response value is retained, as Figure 4 shown. Because of the aggregation of the image feature distribution, when there are multiple feature points in a single node, it means that there are indeed effective features in this area. Although the response values of all feature points in this node of the frame image do not reach the adaptive response threshold, in order to retain as many effective features as possible and ensure the uniformity of the feature point distribution, so still choose to retain the feature point with the highest response value, that is, the most obvious feature point in the node.
[0065] S5. Use the BEBLID algorithm to describe the feature points screened in step S4. Since the BEBLID feature point description algorithm can achieve more accurate description of feature points, the present invention applies it to the front-end visual odometer part of visual SLAM. For the feature points retained by the improved feature extraction and screening algorithm, use the best BEBLID descriptor mode to describe them to obtain higher image matching accuracy, so as to achieve more accurate motion estimation.
[0066] The BEBLID descriptor uses the AdaBoost algorithm to extract sampling image blocks of different sizes in the neighborhood of the feature point, and then compares the average gray difference of these sampling image blocks with the selected threshold to obtain a binary descriptor, as Figure 3 shown.
[0067] First, extract sampling image blocks of different sizes around the feature point, and then use the sampling image block feature extraction function f(x) and threshold T corresponding to each weak classifier in the AdaBoost algorithm to obtain h(x);
[0068] The extraction function of the sampling image block around the feature point is as follows:
[0069]
[0070] Wherein, p1 and p2 respectively represent the centers of the image patches extracted by each weak classifier, s represents the side length of the image patch, I(p) and I(q) respectively represent the gray values of each pixel point;
[0071] The value of h(x) represents the similarity of the sampled image patch structures selected by each weak classifier in AdaBoost. If the average gray difference between two sampled image patches is less than the threshold T, it is +1, otherwise it is -1, as follows:
[0072]
[0073] Secondly, in order to obtain the binary feature descriptor, it is necessary to determine the value of h(f,T). If h(f,T) is greater than 0, the corresponding binary descriptor bit is taken as 1, otherwise it is 0;
[0074] Finally, by training the descriptors of all feature points in the dataset, optimizing the loss function, and obtaining the pixel positions, image patch sizes, and thresholds of the best descriptor sampling point pairs, the best BEBLID binary descriptor pattern is obtained.
[0075] The loss function is:
[0076]
[0077] Wherein, N represents the sampled image patches corresponding to N pairs of feature points in the training dataset; x i and y i respectively represent the image patches corresponding to two certain feature points in the training dataset; where k represents the k-th weak classifier, that is, the k-th bit corresponding to the final 256-bit descriptor; h k (x i ) and h k (y i ) respectively represent the similarity of the structures of the two sampled image patches selected by the k-th weak classifier, that is, when extracting their respective descriptors, the structural similarity of the two sampled image patches in the k-th bit corresponding to the current image (x i or y i ); γ represents the learning rate; l i represents the label, l i ∈{-1,1}, when l i =1, it means that the image patches corresponding to the two feature points have the same image structure, and when l i =-1, it means that the corresponding image structures are different.
[0078] S6. Match the images based on the feature points extracted between two adjacent frames, then estimate the camera pose through the PnP algorithm and the matching relationship of the feature points, and finally adjust and optimize the estimated camera pose by the method of minimum reprojection error;
[0079] S7. Construct all the motion information and camera observation information into an optimization problem with a larger scale and size, and use bundle adjustment to optimize and solve it to obtain the optimal trajectory and map over a long period of time.
[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A high-precision visual SLAM method based on improved quadtree feature point extraction, characterized in that: The method includes the following steps: S1. Collect image information in the environment through a camera; S2. Convert the RGB image collected by the camera into a grayscale image, and at the same time build an image pyramid for each image and divide each layer of the image into grids; S3. Determine the number of feature points to be extracted from each layer of the image pyramid according to the area of each layer of the pyramid and the set number of feature points; S4. Extract an excessive number of feature points within the grids divided in each layer of the pyramid image, and then use an improved quadtree to screen the feature points extracted from each layer of the image. Specifically: First, divide the quadtree according to the number of feature points required for this layer of the image; second, count the response values of the feature points in each node of the quadtree; third, determine the adaptive response threshold for this layer of the image according to the median and average values of the response values of the feature points in this layer of the image; finally, use the calculated adaptive response threshold to screen the feature points in each node, and on the premise of ensuring the uniformity of the feature point distribution, retain as many high-response feature points as possible and eliminate isolated weak-response feature points to achieve more accurate pose estimation; Judge the number of feature points in each node of the quadtree. If there is only one feature point in a certain node, judge it according to the adaptive response threshold. If the response value of the feature point is less than the adaptive response threshold, then eliminate this feature point to reduce the impact on the image matching accuracy; if there are multiple feature points in the node and the response values of all feature points are less than the adaptive response threshold, then retain the feature point with the highest response value; S7. Use the BEBLID algorithm to describe the feature points screened in step S4; S8. Match the images according to the feature points extracted between two adjacent frames of images, then estimate the camera pose through the PnP algorithm and the matching relationship of the feature points, and finally adjust and optimize the estimated camera pose by the method of minimum reprojection error; S9. Construct all the motion information and camera observation information into an optimization problem with a larger scale and size, and use bundle adjustment to optimize and solve it to obtain the optimal trajectory and map in a long time; 2. The high-precision visual SLAM method according to claim 1, characterized in that: Step S3 is specifically: First, calculate the total area of the layers of the entire image pyramid: In the formula, H and W respectively represent the height and width of the bottom layer image, C = HW, s represents the scaling factor of the image pyramid, and m represents the number of layers of the pyramid; Then calculate the number of feature points per unit area according to the number of feature points to be extracted from each image: In the formula, Num represents the number of feature points to be extracted from each image; Finally, determine the number of feature points to be extracted from each layer according to the area size of each layer of the image pyramid. The number of feature points that should be allocated to the i-th layer is:
3. The high-precision visual SLAM method according to claim 1, characterized in that: Step S5 is specifically: First, extract sampling image blocks of different sizes around the feature points, and then use the sampling image block feature extraction function f(x) and threshold T corresponding to each weak classifier in the AdaBoost algorithm to obtain h(x); The extraction function of the sampling image block around the feature point is as follows: Wherein, p1 and p2 respectively represent the centers of the image patches extracted by each weak classifier, s represents the side length of the image patch, I(p) and I(q) respectively represent the gray values of each pixel point; The value of h(x) represents the similarity of the sampled image patch structures selected by each weak classifier in AdaBoost. If the average gray difference between two sampled image patches is less than the threshold T, it is +1, otherwise it is -1, as shown below: Secondly, in order to obtain a binary feature descriptor, it is necessary to determine the value of h(f,T). If h(f,T) is greater than 0, the corresponding binary descriptor bit is taken as 1, otherwise it is 0; Finally, by training the descriptors of all feature points in the dataset, optimizing the loss function, and obtaining the pixel positions, image patch sizes, and thresholds of the best descriptor sampling point pairs, the best BEBLID binary descriptor pattern can be obtained; the loss function is: where N represents the sampled image patches corresponding to N pairs of feature points in the training dataset, x i and y i respectively represent the image patches corresponding to two certain feature points in the training dataset, γ represents the learning rate, h k (x i ) and h k (y i ) respectively represent the similarity of the structures of two sampled image patches selected by the k-th weak classifier, l i represents the label, l i ∈{-1, 1}, when l i = 1, it means that the image patches corresponding to the two feature points have the same image structure, and when l i = -1, it means that the corresponding image structures are different.
Citation Information
Patent Citations
ORB-SLAM2 improved algorithm based on information entropy and sharpening adjustment
CN111709893A
Capsule endoscope image feature point extraction method
CN114926448A