Monocular Camera Visual Odometry Initialization Method Based on Top-View Point Line Feature Fusion

By using dot-line feature fusion method to match features in indoor top view scenes, the problems of feature repetition and sparseness are solved, and stable and fast monocular visual odometer initialization is achieved, and positioning accuracy and real-timeness are improved.

CN114444586BActive Publication Date: 2025-07-01TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210037374.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-13
Publication Date
2025-07-01
Estimated Expiration
2042-01-13

AI Technical Summary

Technical Problem

In indoor top view scenes, feature duplication and point features sparse make it difficult to achieve stable and accurate feature matching, and thus difficult to complete the initialization of the visual odometer.

Method used

The monocular camera visual odometer initialization method based on top view point line feature fusion is adopted. Through large-scale line feature extraction and multi-scale point feature extraction, multi-level point fusion feature matching is performed to obtain stable underlying feature point matching results, and then monocular initialization is completed.

Benefits of technology

It effectively avoids the problems of feature duplication and sparseness in top-view scenes, realizes stable and fast monocular visual odometer initialization, and improves positioning accuracy and real-timeness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114444586B_ABST
    Figure CN114444586B_ABST
Patent Text Reader

Abstract

The present invention relates to a monocular camera visual odometry initialization method based on top-viewpoint line feature fusion. The method comprises the following steps: Step 1: Use a monocular camera to acquire a sequence of images with a set overlap degree in a grid ceiling, called left and right images. Downsample the left and right images to obtain downsampled images and form an image pyramid; Step 2: Parallelly extract large-scale line features and multi-scale point features from the left and right images on the downsampled images; Step 3: Perform multi-level point-line fusion feature matching to obtain stable bottom-layer feature point matching results; Step 4: Perform monocular initialization according to the bottom-layer feature point matching results. Compared with the prior art, the present invention has the advantages of significantly reducing the influence of the problem of feature repetition on matching, effectively reducing the search space for bottom-layer feature matching, and simultaneously taking into account stability and real-time performance, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of positioning information, and in particular to a monocular camera visual odometry initialization method based on the fusion of top-view line features. Background Art

[0002] The positioning technology is a core technology for the operation of indoor service robots. Accurate positioning information is the basis for robots to perform various tasks. It is of great significance and application value to study stable positioning technologies for indoor scenarios such as restaurants, hospitals, and stations. Currently, the SLAM algorithm is usually used to process data from various internal and external sensors to obtain positioning information. Commonly used sensors include radar and cameras. Although methods based on radar and cameras have been widely and continuously studied, there is still a certain distance from large-scale deployment and application, and there are many challenges in actual complex and dynamic environments. For example, there are a large number of moving targets in the above scenarios, which will seriously interfere with the normal operation of the positioning algorithm. For such scenarios, observing the top-view ceiling can well avoid the interference of moving targets. Comparing the characteristics of radar and visual sensors, vision is a low-cost and promising solution because it can obtain stable top-view textures and structured features such as grid-like decorations and smoke alarms.

[0003] The SLAM algorithm using a camera is divided into parts such as visual odometry, backend optimization, and loop detection. Visual odometry is to continuously calculate the pose of images by using the sequence of images obtained by the camera, which is an important part of the visual SLAM algorithm. Among them, the SLAM front-end odometry using a single camera is called monocular visual odometry. Monocular visual odometry extracts features and matches the sequence of images obtained by the camera. At the initial run, based on the matching results and taking the camera coordinate system of the first image as the reference, a world coordinate system is established, the relative pose of the first two images is restored, and the three-dimensional coordinates of the matching points are solved. This step is called monocular camera initialization. After initialization, the three-dimensional coordinates of some spatial points at the current position are obtained. For subsequent images, feature extraction and matching are also used. According to the 2D-3D point pair correspondence, the pose of the corresponding camera of the current image is solved by using PnP, and this process is continuously carried out to continuously and real-time determine the current robot pose.

[0004] In the top - view scenario, the obtained image texture is sparser than that obtained horizontally. There are fewer available features, more repeated feature extractions, and it is difficult to achieve stable and accurate feature matching. Algorithms that work well in the frontal view often struggle to complete the localization task and even have difficulty completing the initialization step. The problem first stems from feature repetition. In indoor top - views, there are often many repeated components, such as alarms and grid intersections. In addition, there are some line segments. In this case, the feature - point descriptors extracted are often mostly similar, especially for rotation - invariant descriptors. This type of descriptor first calculates the main direction of the feature point, rotates the pixels in the neighborhood to the main - direction position, and then calculates the descriptor. Feature points of the same type with different angular distributions become exactly the same after rotation, and without normalizing the angle, correct matching cannot be achieved. Some methods calculate the main direction of line segments for grid ceilings, align the images, and then extract larger circles for SIFT matching. However, this method is rather limited and cannot meet different top - view situations, such as scenarios with height - difference structures or pipeline structures with different directions. Moreover, there are fewer point features in the top - view. Relying solely on point features is not easy to meet the conditions for stable initialization and is difficult to initialize quickly.

[0005] Considering that there are usually point features and line features on the top - view ceiling, by simultaneously using point and line features, deeply fusing feature extraction and feature matching, and using point - line features at a large scale for pre - matching to constrain the underlying feature matching, a stable and fast monocular top - view visual odometry initialization method can be obtained. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the above - mentioned existing technologies and provide a monocular camera visual odometry initialization method based on the fusion of top - view point - line features.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] A monocular camera visual odometry initialization method based on the fusion of top - view point - line features, the method comprising the following steps:

[0009] Step 1: Use a monocular camera to obtain a sequence of images with a set overlap degree in the grid ceiling, called left - right images. Downsample the left - right images to obtain downsampled images and form an image pyramid.

[0010] Step 2: Simultaneously perform large - scale line - feature extraction and multi - scale point - feature extraction on the left - right images on the downsampled images.

[0011] Step 3: Perform multi - level point - line fusion feature matching to obtain a stable underlying feature - point matching result.

[0012] Step 4: Perform monocular initialization based on the underlying feature - point matching result.

[0013] In the said step 2, the process of parallelly extracting large-scale line features and multi-scale point features from the left and right images on the downsampled image specifically includes the following steps:

[0014] Step 201: Use the LSD algorithm to extract line segments on the downsampled image, and perform region growing on the line segment support region and the corresponding region of the line segment endpoints of each line segment respectively to obtain the mask images of the significant corner prediction region, the edge point prediction region, and the general point prediction region;

[0015] Step 202: For the line segments extracted from a single image, calculate the length-weighted direction histogram at the set angular resolution interval b and normalize it;

[0016] Step 203: At the same time, use the FAST operator to extract the underlying features on the original image. Divide the image into uniformly sized grids, and the grids are described by the upper left corner coordinates and the width and length: (X, Y, W, H). Use the FAST operator to extract feature points inside each grid, that is, FAST corners, and assign them to the grid where they are located. If no feature points are extracted, lower the FAST threshold and extract again;

[0017] Step 204: Count the number of feature points extracted using the first FAST operator in the grid. Take out the grid with the largest number of FAST feature points within a certain grid range (X1, Y1, X2, Y2). At the same time, combine the mask image and consider that there are large-scale feature points in the area of this grid and the line segment endpoints. Use the multi-scale detection method to extract the significant feature points in the large-scale space on the corresponding area of the downsampled image, that is, large-scale corners;

[0018] Step 205: Calculate the descriptors of the point-line features.

[0019] In the said step 205, calculating the descriptors of the point-line features specifically is:

[0020] For the line segment features, calculate the line segment descriptors;

[0021] For the large-scale corner features, search for the line segments around the large-scale corners, calculate the line segment directions, and form a direction histogram. Select the local main direction, rotate the pixel blocks within the set neighborhood around the large-scale corners according to the local main direction, and calculate the feature point descriptors.

[0022] In the said step 3, the process of performing multi-level point-line fusion feature matching specifically includes the following steps:

[0023] Step 301: Obtain the rotation angles of the left and right images

[0024] Step 302: Match the large-scale point-line features to obtain the approximate translational and rotational transformation relationships of the left and right images in different regions;

[0025] Step 303: On the basis of the pre-matching of large-scale point-line features, match the underlying features of the left and right images.

[0026] In the above-mentioned step 301, the process of obtaining the rotation angles of the left and right images specifically includes the following steps:

[0027] Step 301A: First, calculate the approximate rotation angles of the left and right images;

[0028] Step 301B: Apply an angular offset a (a = b) to the right image direction histogram, calculate the matching score y with the left image direction histogram, and obtain a matching sequence ((a1, y1), (a2, y2),...);

[0029] Step 301C: Perform curve fitting on the matching sequence obtained in step 301B to obtain an approximate functional relationship y = f(a) between the matching score y and the angular offset a;

[0030] Step 301D: Calculate the matching score corresponding to each angular offset a and obtain the maximum matching score;

[0031] Step 301E: Search for the angular offset corresponding to the maximum matching score, and use an optimization method to find the curve extreme value near it. The offset corresponding to the extreme point is the more accurate rotation angle of the left and right images.

[0032] In the above-mentioned step 302, the process of matching the large-scale point-line features specifically includes the following steps:

[0033] Step 302A: Add large-scale point features on the basis of line feature matching to provide a translation constraint, that is, match the large-scale point-line features according to the rotation angles of the left and right images, the descriptors, and the spatial relationship between the large-scale point-line features. Add large-scale point features and distance constraints to the existing matching method, and respectively screen out candidate matches for the extracted large-scale point-line features according to the descriptors to obtain a candidate match set:

[0034] Step 302B: Arbitrarily select two groups of matching pairs in the candidate match set, that is, large-scale point-line feature combinations, calculate the midpoint positions of each line segment, calculate the consistency scores of the matching pairs according to various constraints, and construct an adjacency matrix for any matching pair according to the consistency scores;

[0035] Step 302C: Use a spectral method to obtain a set of matching pairs with the highest consistency of the adjacency matrix, complete the matching of the large-scale point-line features, and obtain an approximate translation and rotation transformation relationship between the left and right images in different regions.

[0036] In step 302A, the candidate matching set includes large-scale point matching and line segment matching. Specifically, for large-scale point matching, the large-scale points in the left image correspond to those in the right image, and for line segment matching, the line segments in the left image correspond to those in the right image.

[0037] In step 302B, the matching pairs include point-point matching pairs, line-line matching pairs, and point-line matching pairs. Various constraints include the rotation angles of the left and right images, the projection ratio between matching pairs, the descriptor distance, the rotation angles of the left and right images, and the relative offset vector of the centers.

[0038] In step 303, the process of matching the underlying features of the left and right images specifically includes the following steps:

[0039] Step 303A: Match the underlying features and determine whether the underlying feature points are located within the neighborhood set by the large-scale feature points. If so, execute step 303B; if not, execute step 303C.

[0040] Step 303B: Calculate the local principal directions of the large-scale feature points that are successfully matched in the left and right images. Align the pixel blocks around the original matching points of the left and right images based on the local principal directions, perform template matching on the pixel blocks of the left and right images, calculate the local plane transformation relationship, and then perform sub-pixel refinement on the underlying features to obtain the accurate matching of the underlying feature points.

[0041] Step 303C: Search for the first N line segments around the underlying features, calculate the line segment directions to obtain the local rotation angle and the local plane transformation relationship, calculate the descriptor according to the local rotation angle, and search for the feature points matching those in the left image in the right image by windowing according to the local plane transformation relationship.

[0042] In step 4, the process of performing monocular initialization based on the matching results of the underlying feature points is specifically as follows:

[0043] Use the RANSAC algorithm to calculate the homography matrix and the intrinsic matrix respectively according to the matching results of the underlying feature points. Decompose the rotation matrix parameter R and the translation parameter T according to the homography matrix or the intrinsic matrix to restore the relative pose between the two images, and complete the initialization of the monocular visual odometer.

[0044] Compared with the prior art, the present invention has the following advantages:

[0045] 1. By using a monocular camera to obtain indoor top-view images and taking advantage of the grid-like ceiling, smoke detectors, and some decorative objects existing in the indoor top-view environment, it effectively avoids the problem of matching failure caused by a large number of dynamic obstacles in the horizontal images in high-dynamic scenarios. Without the need to pre-lay control points, the initialization work of the visual odometer can be quickly and accurately completed.

[0046] 2. Deep fusion and correlation are carried out in the point-line feature extraction and matching method, effectively utilizing spatial correlations such as the distance between features, local and global main directions, and significantly reducing the impact of feature repetition on matching.

[0047] 3. The multi-level matching strategy effectively reduces the search space for low-level feature matching by using large-scale point-line feature pre-matching, taking into account both stability and real-time performance. Brief Description of the Drawings

[0048] Figure 1 It is a flowchart of feature extraction for the odometer initialization method of the present invention.

[0049] Figure 2 It is a flowchart of feature matching for the odometer initialization method of the present invention. Detailed Embodiments

[0050] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] Embodiment

[0052] As Figure 1 and Figure 2 shown, the present invention provides a monocular camera visual odometer initialization method based on top-view point-line feature fusion. Through point-line feature fusion extraction and matching, stable matching of similar low-level point features is achieved, and then the relative pose of the camera is estimated to complete the monocular initialization work. The method includes the following steps:

[0053] Step 1: Use a monocular camera to capture top-view sequence images of the indoor environment. Select two images with a set overlap degree according to the shooting time, called left and right images. Perform Gaussian downsampling on the left and right images to obtain downsampled images and form an image pyramid, and then execute Step 2;

[0054] Step 2: Simultaneously perform point-line feature extraction on the images, that is, simultaneously perform large-scale line features and multi-scale point features (multi-scale point features include large-scale point features and low-level point features) on the images;

[0055] Step 3: After the left and right images respectively complete point-line feature extraction, perform multi-level point-line fusion feature matching;

[0056] The process of performing LSD line segment feature extraction specifically includes the following steps:

[0057] Step 201: To reduce the time overhead of the LSD algorithm, perform LSD line segment feature extraction on the image downsampled by 2 times or 4 times;

[0058] Step 202: For the line segments extracted in Step 201, in the new image, perform region growing on the line segment support region and the region corresponding to the line segment endpoints of each line segment respectively to obtain the mask images of the significant corner prediction region, the edge point prediction region, and the general point prediction region;

[0059] Step 203: Calculate the angles and lengths of the line segments in Step 201, calculate the length-weighted angle histogram at 20° angle resolution intervals, and normalize it;

[0060] Step 204: Calculate the LBD descriptor of the line segment features extracted in Step 201;

[0061] The process of performing FAST feature point extraction specifically includes the following steps:

[0062] Step 205: Use the FAST operator to extract the underlying features on the original image. To reduce the influence of illumination, divide the image into grids of uniform size, and extract evenly distributed feature points, that is, FAST corners. If no feature points can be extracted within the grid, reduce the FAST threshold and extract feature points again. Assign all the extracted feature points to the grids they belong to. The grids are described by the upper left coordinate and the width and length: (X, Y, W, H);

[0063] Step 206: According to the mask image obtained in Step 202 and the grids divided in Step 205, extract the regions with the possibility of containing large-scale feature points. The regions with the possibility of containing large-scale feature points include the line segment endpoints and the regions where the number of the first FAST corners is significantly more than that of the neighboring grids. Detect large-scale corners on these regions using a multi-scale method;

[0064] Step 207: For the extracted large-scale corners, to solve the problems of feature repetition and the lack of a significant direction for circular features, search for the first N line segments closest to the large-scale corners, calculate the direction histogram, and select the main direction;

[0065] Step 208: Calculate the descriptors of the large-scale corners based on the main direction obtained in Step 207. The descriptors of the large-scale corners include the SIFT descriptor and the ORB descriptor. Only calculate one descriptor;

[0066] The process of performing multi-level point-line fusion feature matching specifically includes the following steps:

[0067] Step 301: Calculate the approximate rotation angles of the left and right images;

[0068] Step 302: Maximize the consistent matching pairs. Calculate K candidate matches based on the descriptors of large-scale corner points, construct an adjacency graph among all candidate matches, calculate the consistency scores among candidate matches under the constraint of the approximate rotation angle, and extract the set of matching pairs with the highest consistency score from the candidate matches through spectral methods to achieve repeated low-level feature matching;

[0069] Step 303: Perform low-level feature matching.

[0070] In Step 302, the process of maximizing the consistent matching pairs specifically includes the following steps:

[0071] Step 302A: Continuously apply an angular offset a = 20°, 40°,... equal to the histogram angular interval to the right image direction histogram, calculate the matching scores y1, y2,... of the left and right image direction histograms, and obtain a matching sequence ((a1, y1), (a2, y2),...);

[0072] Step 302B: Perform curve fitting on the matching sequence obtained in Step 302A to obtain an approximate functional relationship between the matching score y and the angular offset a: y = f(a);

[0073] Step 302C: First, search for the angular offset corresponding to the maximum matching score, and use an optimization method to find the curve extremum near it to obtain a more accurate rotation angle of the left and right images. While calculating the rotation angle of the left and right images, perform Step 302D:

[0074] Step 302D: Add large-scale point features on the basis of online feature matching to provide a translation constraint, that is, add large-scale point features and distance constraints to the existing matching method to solve problems such as periodic mis-matching of ceiling grid patterns. For the large-scale point-line features obtained by feature extraction, use descriptors to screen out candidate matches. Assume that K pairs of point-line matches are screened out and form a candidate match set, including large-scale point matches and line segment matches:

[0075] Step 302E: Arbitrarily select two pairs of matching pairs, that is, large-scale point-line feature combinations. The matching pairs include point-point matching pairs, line-line matching pairs, and point-line matching pairs. Calculate the consistency scores of the matching pairs according to the rotation angles of the left and right images, the projection ratio between the matching pairs, the descriptor distance, and the central relative offset vector, etc., screen out reasonable matching pairs, and construct an adjacency matrix of all matching pairs;

[0076] Step 302F: Use spectral methods to obtain the set of matching pairs with the highest consistency of the matching adjacency matrix, complete the large-scale point-line feature matching, and obtain the approximate translation and rotation transformation relationship between the left and right images in different regions.

[0077] In Step 303, the process of performing low-level feature matching specifically includes the following steps:

[0078] Step 303A: Match the underlying features to determine whether the underlying feature points are located within the neighborhood set by the large-scale feature points. If so, execute Step 303B; if not, execute Step 303C.

[0079] Step 303B: Extract the pixel blocks of a set size around the large-scale feature points corresponding to the underlying point features, rotate them to the local principal direction that is the same on the left and right, and directly perform template matching. The template matching implicitly constrains the spatial relationship between the underlying point features. Under the constraint of the template matching result, search for the feature points that match in the left and right images for the underlying point feature matching, and perform sub-pixel refinement.

[0080] Step 303C: Search for the first N line segments around the underlying features, calculate the line segment directions to obtain the local rotation angle and the local planar transformation relationship, calculate the descriptor according to the local rotation angle, and search for the feature points that match the left image on the right image by windowing according to the local planar transformation relationship.

[0081] In Step 303C, calculating the descriptor according to the local rotation angle specifically means: finding the large-scale points and line segment features that match around the remaining underlying point features, calculating the principal direction of the current underlying point feature, rotating the underlying point feature to this direction, and then calculating the descriptor.

[0082] The present invention includes two parts: feature extraction and feature matching.

[0083] In an indoor horizontal high-dynamic scene, a monocular camera is used to continuously capture indoor top-view images to obtain adjacent image pairs, namely the left and right images. The line segment features of the images are extracted by the LSD (Line Segment Detection) algorithm. After the LSD algorithm extracts candidate line segments, region growing is performed on the line segment support regions and the line segment endpoints respectively. The mask images of edge points, significant corner points, and non-edge points are obtained by using the gradient statistical information of the LSD algorithm. Calculate the LBD (Line Binary Descriptor) descriptor for the extracted line segments, and then statistically calculate the direction histogram of the extracted line segments at a certain angular resolution. At the same time, multi-scale feature points are extracted. The large-scale feature point regions are screened by using the underlying feature aggregation degree and the mask image. Extract large-scale corner points from these large-scale feature point regions, search for the first N line segments closest within a certain radius of them, statistically calculate their direction information, and further select the feature points with a significant principal direction. Calculate the descriptor of the large-scale corner points based on this principal direction.

[0084] After completing the multi-scale point and line feature extraction and association, perform multi-level point and line fusion feature matching. Improve the method for determining the approximate rotation method by matching the existing line direction histogram, calculate the corresponding matching scores for each angle offset, perform curve fitting on the matching sequence, and use the optimization method to find the extreme points of the curve, so as to obtain more accurate rotation angles for the left and right images. After that, add large-scale point features on the basis of line feature matching to provide translational constraints and solve the problem of periodic mis-matching of repeated patterns such as ceiling grids. First, select candidate matches according to the descriptor distance. Then, take any two pairs of large-scale point and line feature combinations to calculate indicators such as their central offset vectors and projection ratios, construct the adjacency matrix of all large-scale point and line feature combinations, and use the spectral method to select the large-scale point and line feature combination with the smallest indicator to obtain the large-scale point and line matching result and complete the large-scale point and line feature matching. Perform bottom-layer feature matching under the large-scale constraint. Around the large-scale feature points, directly select the appropriate main direction for template matching, and obtain the matching point pairs according to the local transformation relationship. For general points, calculate the main direction and approximate translation parameters using the nearest N line segments, and perform window search on the right image of the original image.

[0085] This method extracts large-scale point and line features and bottom-layer point features, and obtains stable bottom-layer point feature matching pairs through multi-level matching, which are used to estimate the homography matrix and fundamental matrix between the left and right images, decompose the rotation matrix parameter R and translation parameter T to restore the relative pose between the two images, and complete monocular initialization. Monocular initialization is to calculate the relative spatial relationship of the cameras when the two images are taken, and the relative spatial relationship is the rotation matrix parameter R and translation parameter T in the three-dimensional space.

[0086] This method includes the following steps:

[0087] Step 1: Use a monocular camera to obtain a sequence of images with a set overlap in the grid ceiling, called left and right images. Perform downsampling on the left and right images to obtain downsampled images and form an image pyramid.

[0088] Step 2: At the same time, perform point and line feature extraction on the left and right images on the downsampled images:

[0089] Use the LSD algorithm to extract line segments on the downsampled images, and perform region growing on the line segment support region and the corresponding region of the line segment endpoints of each line segment respectively to obtain the mask images of the significant corner prediction region, edge point prediction region, and general point prediction region.

[0090] For the line segments extracted from a single image, calculate the length-weighted direction histogram at intervals of the set angular resolution b (Bin) and normalize it.

[0091] Meanwhile, the original image is divided into grids of uniform size. The coordinates of the grids are denoted as (X, Y) according to the image coordinate system. The FAST operator is used to extract feature points, i.e., FAST corner points, inside each grid. If no feature points are extracted, the FAST threshold is decreased and extraction is performed again;

[0092] Count the number of feature points extracted using the FAST operator for the first time inside the grids. Select the grid with the largest number of FAST corner points within a certain grid range (X1, Y1, X2, Y2). Meanwhile, combined with the mask image, it is considered that there are large-scale feature points in the area of this grid and the line segment endpoints. The multi-scale detection method is used to extract significant feature points in the large-scale space, i.e., large-scale corner points, in the corresponding area of the downsampled image;

[0093] Calculate the descriptors of the large-scale point-line features. For the line segment features, calculate their LBD descriptors; for the large-scale point features, search for the line segments around the large-scale corner points, calculate the line segment directions, and form a direction histogram. Select the main direction, rotate the pixel blocks within the set neighborhood around the large-scale corner points according to the main direction, and calculate their descriptors;

[0094] Step 3: Obtain the approximate rotation angles of the left and right images and make improvements. Apply an angular offset a (a = b) to the direction histogram of the right image, calculate the matching score y with the direction histogram of the left image, calculate the matching score corresponding to each angular offset a, find the maximum matching score, perform curve fitting on the matching sequence formed by the angular offset and the matching score, and use the optimization method to calculate the extreme point near the maximum matching score. The offset corresponding to the extreme point is the more accurate rotation angle of the left and right images

[0095] Match the large-scale point-line features according to the rotation angles of the left and right images, the descriptors, and the spatial relationship between the large-scale point-line features. Add large-scale point features and distance constraints to the existing matching method. First, screen out the candidate matches through the descriptors, calculate the midpoint positions of each line segment, and calculate the consistency scores for the point-point matching, line-line matching, and point-line matching in any two pairs of candidate matches. The constraints of the consistency scores include the projection ratio between the matching pairs, the central relative offset vector, the rotation angles of the left and right images, and the descriptor distance, and normalize them to the same magnitude. Use the consistency scores to construct the adjacency matrix of any matching pair, and use the spectral method to find the candidate matching set with the highest consistency score to complete the matching of the large-scale point-line features;

[0096] Based on the pre - matching of large - scale point - line features, match the underlying features of the left and right images. Near the large - scale feature points, calculate the main direction of the large - scale feature points. Align the pixel blocks around the matching points of the original left and right images with the main direction as the reference. Perform template matching on the pixel blocks of the left and right images to calculate a relatively accurate local planar transformation relationship, and then refine the underlying features at the sub - pixel level to obtain the accurate matching of the underlying feature points. For the remaining underlying features, calculate the main direction and descriptors of the nearest line segments, and search for the matching feature points by windowing according to the point - line feature matching results.

[0097] Step 4: Perform monocular initialization through the matched underlying features. Use the RANSAC algorithm to calculate the homography matrix and the essential matrix respectively, and select a suitable model to recover the relative pose of adjacent images from them.

[0098] Calculate the homography matrix and the essential matrix of the two images based on the matched underlying points. Select the homography matrix or the essential matrix according to the scene to calculate the rotation matrix parameter R and the translation parameter T.

[0099] The method proposed by the present invention uses a monocular camera to obtain the top - view images in an indoor scene. Through large - scale point - line feature extraction, descriptor calculation, image rotation angle estimation, and large - scale point - line feature fusion matching to guide the underlying feature matching, stable matching results of underlying feature points can be obtained in the case of sparse and repeated features. On this basis, calculate the essential matrix or the homography matrix between the matching points, thereby completing the initialization of the monocular visual odometer. This method makes full use of the point - line features existing in the environment, and conducts different degrees of fusion on both extraction and matching, effectively overcoming the problems of sparse and repeated features in this scene, realizing fast and reliable monocular initialization. Extracting line segment features can also provide more constraints for initialization and odometer tracking and mapping, resulting in more stable positioning results.

[0100] The above - mentioned are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any staff familiar with the technical field of the present invention can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A monocular camera visual odometry initialization method based on top - view point line feature fusion, characterized in that The method includes the following steps: Step 1: Use a monocular camera to obtain a sequence of images with a set overlap in the grid ceiling, called left and right images. Downsample the left and right images to obtain downsampled images and form an image pyramid; Step 2: Simultaneously perform large-scale line feature extraction and multi-scale point feature extraction on the left and right images on the downsampled images; Step 3: Perform multi-level point-line fusion feature matching to obtain a stable matching result of underlying feature points; Step 4: Perform monocular initialization based on the matching result of the underlying feature points; In the said Step 3, the process of performing multi-level point-line fusion feature matching specifically includes the following steps: Step 301: Obtain the left and right image rotation angles Step 302: Match the large-scale point-line features to obtain an approximate translational and rotational transformation relationship between the left and right images in different regions; Step 303: On the basis of the pre-matching of the large-scale point-line features, match the underlying features of the left and right images; In the said Step 302, the process of matching the large-scale point-line features specifically includes the following steps: Step 302A: Add large-scale point features on the basis of online feature matching to provide translational constraints, that is, according to the rotation angles of the left and right images Descriptors and the spatial relationship between large-scale point and line features to match the large-scale point and line features, add large-scale point features and distance constraints to the existing matching method, and respectively screen out candidate matches for the extracted large-scale point and line features according to the descriptors to obtain a candidate match set: Step 302B: Arbitrarily select two sets of matching pairs, that is, large-scale point-line feature combinations, in the candidate matching set. Calculate the midpoint position of each line segment, calculate the consistency score of the matching pairs according to various constraints, and construct an adjacency matrix of any matching pair according to the consistency score; Step 302C: Use the spectral method to obtain a set of matching pairs with the highest consistency of the adjacency matrix, complete the large-scale point-line feature matching, and obtain an approximate translational and rotational transformation relationship between the left and right images in different regions.

2. A monocular camera visual odometry initialization method based on top view point line feature fusion according to claim 1, characterized in that In the said Step 2, the process of simultaneously performing large-scale line feature extraction and multi-scale point feature extraction on the left and right images on the downsampled images specifically includes the following steps: Step 201: Use the LSD algorithm to extract line segments on the downsampled images. Perform region growing on the line segment support region and the corresponding region of the line segment endpoints of each line segment respectively to obtain mask images of the significant corner prediction region, edge point prediction region, and general point prediction region; Step 202: For the line segments extracted from a single image, calculate the length-weighted direction histogram at the set angular resolution interval b and normalize it; Step 203: At the same time, use the FAST operator to extract the underlying features on the original image. Divide the image into grids of uniform size, and the grids are described by the upper left corner coordinates and width and length: (X, Y, W, H). Use the FAST operator to extract feature points, that is, FAST corners, inside each grid and assign them to the grid where they are located. If no feature points are extracted, lower the FAST threshold and extract again; Step 204: Count the number of feature points extracted using the FAST operator for the first time in the grid. Take out the grid with the largest number of FAST feature points within a certain grid range (X1, Y1, X2, Y2). At the same time, combine the mask image and consider that there are large-scale feature points in the region of this grid and the line segment endpoints. Use the multi-scale detection method to extract significant feature points in the large-scale space, that is, large-scale corners, in the corresponding region on the downsampled image; Step 205: Calculate the descriptors of the point-line features.

3. A monocular camera visual odometry initialization method based on top view point line feature fusion according to claim 2, characterized in that In the said Step 205, calculating the descriptors of the point-line features specifically is: For line segment features, calculate line segment descriptors; For large-scale corner features, search for line segments around the large-scale corners, calculate the line segment directions, and form a direction histogram. Select the local dominant direction, rotate the pixel blocks within a set neighborhood around the large-scale corners according to the local dominant direction, and calculate the feature point descriptors.

4. A monocular camera visual odometry initialization method based on top view line feature fusion according to claim 2, characterized in that In the said step 301, the process of obtaining the left and right image rotation angles specifically includes the following steps: Step 301A: First, calculate the approximate rotation angles of the left and right images. Step 301B: Apply an angular offset a, where a = b, to the right image direction histogram, calculate the matching score y with the left image direction histogram, and obtain a matching sequence ((a1, y1), (a2, y2)…). Step 301C: Perform curve fitting on the matching sequence obtained in Step 301B to obtain an approximate functional relationship y = f(a) between the matching score y and the angular offset a. Step 301D: Calculate the matching score corresponding to each angular offset a and obtain the maximum matching score. Step 301E: Search for the angular offset corresponding to the maximum matching score, and use an optimization method to find the extreme value of the curve near it. The offset corresponding to the extreme point is the more accurate left and right image rotation angles.

5. A monocular camera visual odometry initialization method based on top-viewpoint line feature fusion according to claim 1, characterized in that, In the said Step 302A, the candidate matching set includes large-scale point matching and line segment matching. The large-scale point matching specifically means that the large-scale points of the left image correspond to those of the right image, and the line segment matching specifically means that the line segments of the left image correspond to those of the right image.

6. A monocular camera visual odometry initialization method based on top view point line feature fusion according to claim 1, characterized in that, In the said Step 302B, the matching pairs include point-point matching pairs, line-line matching pairs, and point-line matching pairs. Various constraints include the rotation angles of the left and right images, the projection ratio between the matching pairs, the descriptor distance, the rotation angles of the left and right images, and the relative offset vector of the centers.

7. A monocular camera visual odometry initialization method based on top view point line feature fusion according to claim 4, characterized in that In the said Step 303, the process of matching the underlying features of the left and right images specifically includes the following steps: Step 303A: Match the underlying features and determine whether the underlying feature points are located within the neighborhood set by the large-scale feature points. If so, execute Step 303B; if not, execute Step 303C. Step 303B: Calculate the local dominant directions of the large-scale feature points that are matched on the left and right images. Align the pixel blocks around the original image matching points of the left and right images based on the local dominant direction, perform template matching on the pixel blocks of the left and right images, calculate the local plane transformation relationship, and then perform sub-pixel refinement on the underlying features to obtain the accurate matching of the underlying feature points. Step 303C: Search for the first N line segments around the underlying features, calculate the line segment directions to obtain the local rotation angle and the local plane transformation relationship, calculate the descriptor according to the local rotation angle, and search for the feature points matching the left image on the right image by windowing according to the local plane transformation relationship.

8. A monocular camera visual odometry initialization method based on top-viewpoint line feature fusion according to claim 1, characterized in that, In the said Step 4, the process of performing monocular initialization based on the matching results of the underlying feature points is specifically as follows: Use the RANSAC algorithm to calculate the homography matrix and the essential matrix respectively according to the matching results of the underlying feature points, decompose the rotation matrix parameter R and the translation parameter T according to the homography matrix or the essential matrix to restore the relative pose between the two images, and complete the initialization of the monocular visual odometer.