Robot localization method based on improved point-line feature and IMU fusion

Through the improved dot-line feature fusion method and combined with IMU data, the positioning accuracy and real-time problems of traditional single vision sensors in multi-texture environments are solved, achieving more efficient and robust robot positioning.

WO2025108059A1PCT designated stage expired Publication Date: 2025-05-30CHINA MCC17 GRP CO LTD

Patent Information

Application Number
PCT/CN2024/129584
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2024-11-04
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional single vision sensors have influences on light and velocity, insufficient point features or uneven distribution, and excessive extraction of line features in multi-texture environments, resulting in problems with positioning accuracy and real-time.

Method used

The improved point-line feature fusion method is used to extract feature points through the Shi-Tomasi algorithm, and the LK optical flow is tracked. The improved LSD line segment detection method extracts line features, and the unstable short line segments are reduced through the line segment merging algorithm. The LBD descriptor is used for matching, and pre-integration and nonlinear optimization are combined with IMU data.

Benefits of technology

It improves the real-time and accuracy of robot positioning, reduces computing costs, enhances the robustness and stability of the system, and performs better in textured environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129584_30052025_PF_FP_ABST
    Figure CN2024129584_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of visual SLAM localization. Disclosed is a robot localization method based on improved point-line feature and IMU fusion. The method in the present invention comprises: 1, extracting feature points from an image, performing tracking and matching, using an improved line segment detection (LSD) method to extract lines from the image, using a line segment merging algorithm to merge short line segments belonging to the same straight line, and then using an LBD to perform matching; 2, performing pre-integration on collected IMU data; 3, using the same initialization policy as VINS-Mono; and 4, using information added to a sliding window, so as optimize the pose of a camera. In the present invention, filtering processing is added, a large number of line segments are no longer extracted in a dense region, a line segment merging algorithm is used to remove some short line segments, and long line segments with more stable characteristics are obtained by means of merging, thereby reducing the matching complexity; and in an environment where there are more texture lines, a trajectory generated in the present invention can also effectively improve the precision of localization.
Need to check novelty before this filing date? Find Prior Art

Description

A robot positioning method based on improved point-line features and IMU fusion Technical Field

[0001] The present invention relates to the field of visual SLAM positioning and mapping technology, and more specifically to a robot positioning method based on improved point-line features and IMU fusion. Background Art

[0002] Simultaneous Localization and Mapping (SLAM) is a key technology for robot navigation. Robots equipped with sensors can achieve autonomous localization and mapping navigation without prior environmental information. Cameras are inexpensive and have strong information perception capabilities. Visual SLAM using cameras as sensors is a hot research topic in robotics. However, visual SLAM systems are susceptible to factors such as lighting and movement speed, leading to the trend towards the integration of multiple sensors. Using visual information as a constraint for the IMU (Inertial Measurement Unit) can effectively reduce IMU drift and accumulated error. The IMU can restore the scale information of the monocular camera and provide fast-response positioning for vision. Therefore, SLAM systems that integrate vision and IMUs have become a hot technical topic.

[0003] In vision-integrated IMU (Infrared Measurement Unit) SLAM systems, the main methods for associating visual information with data are direct methods and feature point methods. Feature point methods offer greater robustness and accuracy. Typical feature point-based algorithms include SIFT, PTAM, and SUPF. However, these algorithms suffer from issues such as poor real-time performance and prone to tracking loss. ORB-SLAM, proposed by Mur-Artal et al., brought feature point SLAM technology to its peak. While the use of ORB feature points enables the entire system to achieve excellent tracking and mapping results, it suffers from computational overhead and the sparse feature points extracted are only sufficient for localization but not navigation. Consequently, line features have gradually attracted considerable attention. Line features possess a more complete information structure and are better able to represent environmental information. JEngel et al. proposed LSD-SLAM, which achieves semi-dense tracking through pixel gradients and employs ingenious techniques to ensure real-time and stable tracking. Pumarola et al. proposed the monocular point-line SLAM algorithm, and Gomez-Ojeda et al. proposed the binocular point-line SLAM algorithm, achieving real-time localization and mapping, successfully improving the accuracy and robustness of SLAM. However, when there are a large number of identical line textures in the environment, extracting straight lines will not only extract a large number of identical line segments, but also increase the computational cost and cause mismatching, thus affecting the accuracy.

[0004] Summary of the Invention

[0005] 1. Technical problem to be solved by the invention

[0006] In response to the problems existing in traditional single vision sensors, as well as the problems of insufficient or uneven distribution of point features and over-extraction of line features in multi-texture environments, the present invention provides a robot positioning method based on the fusion of improved point and line features and IMU; the present invention adopts line features to make up for the shortcomings of point features, and uses an improved line feature extraction method to make up for the shortcomings of over-extraction. In environments with a large number of texture lines, the present invention has the advantages of strong real-time performance and high precision.

[0007] 2. Technical solution

[0008] In order to achieve the above object, the technical solution provided by the present invention is:

[0009] The present invention provides a robot positioning method based on improved point-line features and IMU fusion, the steps of which are as follows:

[0010] Step 1: Use the Shi-Tomasi algorithm to extract feature points in the captured image, and use LK optical flow for tracking and matching. At the same time, use the improved LSD line segment detection method to extract line features in the image, and use the line segment merging algorithm to merge short line segments belonging to the same straight line, and then use the LBD descriptor for matching;

[0011] Step 2: Pre-integrate the collected IMU data;

[0012] Step 3: Using the same initialization strategy as VINS-Mono, the visual information obtained in steps 1 and 2 is tightly coupled with the IMU information. The visual SFM is used to estimate the pose between consecutive frames within the sliding window and the inverse depth of the 3D points. Finally, the pose is aligned with the IMU pre-integration result and the initialization parameters are solved.

[0013] Step 4: Using the visual parameters obtained in steps 1 and 2 and the IMU pre-integration results, the point-line reprojection residuals and the IMU data measurement pre-integration residuals are calculated. Nonlinear optimization is performed through sliding window visual constraints, IMU constraints, and landmark points. The three form a close constraint relationship. The point-line reprojection residuals and the IMU pre-integration residuals are used as optimization targets to optimize the camera pose, IMU pose, and the position of the landmark points, thereby obtaining the optimized camera pose graph. Through the camera pose graph, the actual motion trajectory of the robot is obtained to achieve the positioning purpose.

[0014] Furthermore, in step 1, before using the improved LSD line segment detection method to extract straight line features in the image, a filter is first set to screen the line-dense areas in the image. The filter detects areas with high pixel density in the grayscale image based on pixel gradients and changes the pixels in the area to a uniform pixel value. Then, LSD is called to perform line detection on the processed grayscale image.

[0015] Furthermore, the filtering process of the filter is:

[0016] First, traverse a grayscale image pixel and calculate the pixel gradient value. For pixel K, the gradient value is Kgra, and the gradient threshold G is set;

[0017] Secondly, define the regional gradient density number T. When Kgra>G, take the pixel K coordinate as the center and traverse the gradient values ​​of all pixels Kn in the t×t area. The number of gradient values ​​of all pixels Kn in the area where Kngra>G is the regional gradient density number T.

[0018] Then set the regional gradient density threshold When T is greater than When , all pixel values ​​within the 3t×3t area are changed.

[0019] Furthermore, the line segment merging algorithm first groups the line segments based on the line segment angle difference and the distance from the endpoint to the straight line, then implements different sorting schemes according to the line segment angle to achieve an orderly arrangement of the line segments based on relative positions, and finally compares the distance between the line segment endpoints and the sum of the line segment lengths to determine whether they belong to the same straight line, and screens and eliminates the merged lines based on the LBD descriptor.

[0020] Furthermore, the specific process of the line segment merging algorithm is as follows:

[0021] (1) Sort the line segments extracted by LSD in descending order according to their lengths to obtain L = {L1, L2, L3, ..., Ln}; for each line segment L i , filter based on angle to get the candidate line segment group L θ , the longest line segment in each group is called the main line segment l m ;

[0022] (2) Calculate line segment L j Starting point to L i The distance d from the straight line s , line segment L j End at L i The distance d from the straight line e , and then get the line segment L j End points to L i The distance d of the straight line, d=[d s ,d e ];

[0023] (3) For the candidate line segment group L θ After distance screening again, the final candidate line segment group L is obtained g ;

[0024] (4) For the grouped line segments, the line segments belonging to the same straight line are merged through the line segment merging strategy, and the line segments that are incorrectly grouped are eliminated.

[0025] Furthermore, in step (1), the longest line segment is grouped and the length is less than the set threshold L. min The line segments are removed without grouping.

[0026] Furthermore, the angle screening strategy in step (1) is:

[0027] Where: L j Represents the length ratio L in the line segment group L i Short line segment, θ i and θ j L i and L j Angle with the x-axis, θ th Indicates L i and L j The angle difference threshold between them.

[0028] Furthermore, the distance screening strategy in step (3) is:

[0029] Where: d th Indicates the distance threshold from the endpoint to the straight line.

[0030] Furthermore, when merging line segments in step (4), according to the main line segment l m Angle θ with the x-axis m Execute different sorting strategies, align the segment groups according to the direction of the main segment, divide the sorted segments into two groups according to the position of the main segment, and perform prefix sum calculation on each group. When calculating the prefix sum, each group of segments must include the main segment, starting from the segment farthest from the main segment, calculate the main segment l m Starting endpoints m , and the starting endpoint s of the segment farthest from the main segment j The distance d and the total length l of the corresponding line segment in the prefix sum calculated previously sum For comparison, if the total length of the line segment is l sum The proportion of distance d is greater than the threshold r th , then it is considered that from the main line segment l m All line segments between the farthest line segment belong to the same long line segment, connecting the main line segment l m The starting point of the farthest segment and the end point of the farthest segment are used as the merged segment, otherwise the farthest segment is released back to the segment group; at the same time, the second farthest segment is continued to be judged.

[0031] Furthermore, in the matching process using LBD descriptors, the descriptors of the merged line segment and the main line segment are calculated, and the descriptor distance is calculated. When the descriptor distance is greater than the threshold D th , abandon the current merge and keep only the main line segment as the result of the current merge.

[0032] 3. Beneficial effects

[0033] Compared with the existing known technologies, the technical solution provided by the present invention has the following significant effects:

[0034] (1) The present invention provides a robot positioning method based on improved point-line features and IMU fusion. Based on the principle of LSD to extract straight lines, before performing LSD detection, the line feature dense area is screened by a filter. The filter detects the area with high pixel density in the grayscale image and changes the pixels in the area to a uniform pixel value. Then, the grayscale image is subjected to straight line detection. After LSD detection, the unstable short line segments on the same straight line are merged into stable long line segments by sorting and merging the collected short line segments. While reducing the feature lines, it also has certain benefits on the matching efficiency of the LBD descriptor. After screening and line merging, the area will no longer be subjected to dense line collection, but the outline of the area will be extracted. This can make up for the problem of over-extraction of line features in textured environments.

[0035] (2) The robot positioning method based on the improved point-line feature and IMU fusion of the present invention filters and merges the line segments in the dense area, eliminating some unstable line segments, thereby reducing the complexity of system matching, improving the matching efficiency, and further improving the real-time performance of the system.

[0036] (3) The present invention is compared with the VINS_mono algorithm under the MH_04_difficuil (MH_04 for short) sequence. The experiment shows that the trajectories of both are relatively close to the true trajectory, but the maximum error of the present invention is smaller in comparison, and the APE change is more stable. It can be seen that the trajectory of the present invention can effectively improve the positioning accuracy of the robot. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] FIG1 is a processing framework diagram of the present invention;

[0038] FIG2 is a filtering flow chart of the improved LSD line extraction method;

[0039] FIG3 is a schematic diagram of partitioning based on line segment direction;

[0040] FIG4 is a schematic diagram of line segments after grouping;

[0041] Figure 5 (a) and (b) are the line feature extraction results before and after improvement;

[0042] Figure 6 (a) and (b) are the line matching results before and after improvement;

[0043] Figure 7 (a) and (b) are the effect diagrams of the point and line features extracted from the MH_04 sequence by the present invention;

[0044] Figure 8 (a) and (b) are the APE error heatmaps of the trajectories of the VIN_Mono and IPLI algorithms and the true trajectory;

[0045] Figure 9 shows the trajectory changes over time for the VIN_Mono and IPLI algorithms;

[0046] Figure 10 (a) and (b) are the indoor line feature extraction effect diagrams of the present invention;

[0047] Figure 11 (a) and (b) are comparison diagrams of indoor positioning trajectories under VIN_Mono and IPLI algorithms. DETAILED DESCRIPTION

[0048] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0049] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0050] Example 1

[0051] With reference to FIG1 , a robot positioning method based on improved point-line features and IMU fusion in this embodiment is mainly divided into three parts: front-end feature extraction and IMU data preprocessing, initialization, and back-end sliding window joint optimization. The specific steps are as follows:

[0052] Step 1: Use the Shi-Tomasi algorithm to extract feature points from the image and use LK (Lucas-Kanade) optical flow for tracking and matching. Line features are extracted from the image using an improved LSD line segment detection method. Before extracting lines, a filter is applied to filter out dense areas and reduce over-extraction in textured environments. A line segment merging algorithm is used to merge short line segments belonging to the same line, and then matching is performed using the LBD descriptor.

[0053] The principle of LSD line extraction is that areas with high pixel gradient change rates are likely to form straight lines. This embodiment sets a filter to screen dense areas before performing LSD detection. The filter is designed based on pixel gradients, detecting areas of high pixel density in the grayscale image and changing the pixels in that area to a uniform pixel value. LSD is then called to perform line detection on this processed grayscale image. For the modified area, only the modified edge box is detected, and the dense line segments in that area are no longer detected.

[0054] Figure 2 shows the filter flow chart, the specific process is as follows:

[0055] First, we traverse a grayscale image pixel and calculate its gradient value. For pixel K, the gradient value is Kgra, and the gradient threshold G is set.

[0056] Next, define the regional gradient density number T. When Kgra>G, traverse the gradient values ​​of all pixels Kn in the t×t region centered on the pixel K coordinate. The number of gradient values ​​of all pixels Kn in the region where Kngra>G is the regional gradient density number T.

[0057] Then set the regional gradient density threshold When T is greater than Change the values ​​of all pixels in a 3t×3t area.

[0058] t and The setting of the value directly affects the filtering effect. If the threshold is too large, the filtering effect is not obvious. If the threshold is too small, excessive filtering will result in a reduction in the number of line segments detected in subsequent tests. Considering the comprehensive experiments, the t value is set to 3. The value is set to 18.

[0059] After the filter has filtered the dense areas, the LSD algorithm is used to extract the characteristic straight lines in the image. Since the extracted characteristic straight lines contain unstable short line segments, and considering that the long line segment features are more stable, this embodiment uses a line segment merging algorithm to merge the short line segments on the same line to obtain more stable long line segments.

[0060] The line segment merging algorithm first groups line segments based on angle differences and the distances from endpoints to the line. It then implements different sorting schemes based on the angles of the line segments to achieve an orderly arrangement of the line segments based on relative positions. Finally, it compares the distances between the line segment endpoints and the sum of the line segment lengths to determine whether they belong to the same line. The merged lines are then screened and eliminated based on the LBD descriptor.

[0061] The specific process of the line segment merging algorithm is as follows:

[0062] First, sort the line segments extracted by LSD in descending order according to their length, and obtain L = {L1, L2, L3, ..., Ln}. Group the line segments starting from the longest one, whose length is less than the set threshold L. min The line segments are not grouped, but removed. For each line segment L i First, we filter the candidate line segment group L based on the angle θ , the longest line segment in each group is called the main line segment lm. This step can quickly filter out a large number of unmatched line segments and reduce the subsequent calculation amount:

[0063] Where: L j Represents the length ratio L in the line segment group L i Short line segment, θ i and θ j L i and L j Angle with the x-axis, θ th Indicates L i and L j The angle difference threshold between them.

[0064] Then, calculate L j End points to L i The distance of the straight line, let the line segment L j The starting point coordinates are p s , the end point coordinate is p e , line segment L i The straight line is represented as l=[l1,l2,l3] T , that is, the equation of the line is l1x+l2y+l3=0, then the distance from the two end points to L i The distance d of the straight line is:

[0065] Where: d s Indicates L j Starting point to L i The distance of the line, d e Indicates L j End at L i The distance along the straight line.

[0066] Calculate the distance d and filter the previous angle to get the candidate line segment group L θ After distance screening again, the final candidate line segment group L is obtained g :

[0067] Where: d th Indicates the distance threshold from the endpoint to the straight line.

[0068] By filtering under these conditions, we can identify line segments in the image that belong to the same line. These line segments may originally belong to the same line, or they may be line segments from different lines in space projected onto the same line. Therefore, for the grouped line segments, we use a segment merging strategy to merge as many line segments as possible that belong to the same line while removing misgrouped line segments and storing them in the original segment group L extracted by LSD for subsequent processing.

[0069] Since the distance between the endpoints of the line segments is not restricted during the line segment grouping process, the line segment group L g Some line segments may appear to be on the same line in the image but differ significantly in real space. These incorrectly grouped segments must be removed when merging. The lines extracted by the LSD algorithm are directional. If the colors of a binary image are inverted, the lines extracted by LSD remain unchanged, but their directions are completely opposite. However, when calculating the LBD descriptor, the direction of the line segments is distinguished. Therefore, when merging line segments, the relative positions of the line segments must be determined. The starting and ending points must be consistent to ensure accurate subsequent line segment matching.

[0070] First, to prevent sorting ambiguity, the rectangular coordinate system is divided into four parts based on angle, as shown in Figure 3, where S represents the starting point of a segment and E represents the end point of a segment. Taking interval 1 as an example, with the segments in this interval as the main segments, the segment groups obtained after the above segment grouping are all in the form of S at the bottom and E at the top. Therefore, it is only necessary to sort according to the vertical coordinate of the starting point S to obtain a geometrically ordered sequence of segments.

[0071] The long line segment previously grouped for reference is called the main line segment l m , according to the angle θ between the main line segment and the x-axis m Execute different sorting strategies. The purpose of this is to align the line segment groups according to the direction of the main line segment, which is convenient for the subsequent calculation of the endpoint distance. Suppose two line segments l i and l j , the starting point and the end point are (s i ,e i ), (s j ,e j ), the coordinates of each point are (x, y), then the sorting is done according to the following strategy:

[0072] In the formula: sort(*) means sorting based on the size of the variables in the brackets.

[0073] The line segments sorted by this method are sorted according to the endpoint positions according to {s i ,ei ,s j ,e j}, and then divide the sorted segments into two groups according to the position of the main segment, as shown in Figure 4, and perform prefix sum calculation on each group. When calculating the prefix sum, each group of segments must include the main segment. Taking the sorted segment group 2 as an example, let the prefix sum calculation result be s pre =[l sum0 ,l sum4 ,l sum5 ,l sum6 ], then l sum0 That is equal to l m The length of l sum4 Then it is equal to l m The sum of the length of l4, and so on.

[0074] Starting from the line segment farthest from the main line segment, i.e. l6 in Figure 4, calculate the main line segment l m Starting endpoints m and the starting endpoint s of segment l6 j The distance d and the total length l of the corresponding line segment in the prefix sum calculated previously sum6 For comparison, see the following formula:

[0075] Where: r th is the ratio threshold.

[0076] If the ratio of the total length of the line segment to the distance is greater than the threshold r th , then it is considered that from l m All line segments between l and l6 belong to the same long line segment, connecting l m The starting point of l and the end point of l6 are taken as the merged line segment, otherwise the line segment l6 is released back to the line segment group L for subsequent processing. This is achieved by marking the line segment as used or not. At the same time, the next farthest line segment l5 is continued to be judged. Since long line segments often have sufficient information and stability, and the incorrect merging of long line segments that do not belong to the same straight line often causes greater errors in the system, the proportional coefficient r is used in the calculation process. th A linear calculation is performed to limit the merging of long segments. th =α×dis+β

[0077] Where: α and β are parameter factors, and dis is the minimum distance.

[0078] Finally, in order to reduce the error of merging, we filter based on LBD descriptors, calculate the descriptors of the merged line segment and the main line segment, and calculate the descriptor distance. If the main line segment and the line segment it merges belong to the same straight line, the Hamming distance of the descriptors before and after the merger will not be very different. Therefore, when the descriptor distance is greater than a certain threshold Dth , abandon the current merge and keep only the main line segment as the result of the current merge.

[0079] Figure 5 shows the improved line extraction results. (a) in Figure 5 shows the unfiltered line segment extraction result, and (b) shows the filtered line segment extraction result. The Venetian blinds in Figure 5 represent a complex environment with dense line features. After filtering and line merging, the LSD algorithm no longer extracts a large number of line segments. This demonstrates that the filtering and line merging methods proposed in this invention can effectively improve the efficiency and accuracy of subsequent line segment detection and matching, reduce computational costs, and enhance system flexibility.

[0080] To verify the effectiveness of this invention, we conducted experiments using the open-source EuRoC dataset and a real-world multi-texture environment to test positioning accuracy and robustness. The EuRoC dataset was collected by ETH Zurich using an indoor drone and features complex environments with varying lighting and textures. The dataset was tested on a 64-bit Ubuntu 18.04 operating system, while the actual environment was tested on a 64-bit Ubuntu 20.04 operating system.

[0081] Line feature matching uses the LBD descriptor. Figure 6(a) shows the original unprocessed LSD extraction and matching results, while Figure 6(b) shows the extraction and matching results after filtering and line segment merging. It can be seen that the matching results are dense in areas with dense blinds, resulting in complex matching results, which affects the matching quality and real-time performance of the system.

[0082] Table 1 Comparison of matching time before and after improvement

[0083] Complex matching processes increase processing time, but filtering eliminates the need to extract numerous line segments in densely populated areas, reducing matching complexity and improving the system's real-time performance. Table 1 shows a comparison of extraction and matching times before and after the improvement. This comparison demonstrates that the LSD algorithm, after filtering, effectively improves its real-time performance. The improved extraction and matching time is approximately 8.28% less than the unmodified version, validating the algorithm's feasibility.

[0084] Step 2: To avoid re-propagating the IMU measurements, pre-integrate the acquired IMU data.

[0085] Step 3: The initialization process uses the same method as VINS-mono. The visual information obtained in steps 1 and 2 is tightly coupled with the IMU information. The visual SFM is used to estimate the pose between consecutive frames in the sliding window and the inverse depth of the 3D points. Finally, it is aligned with the IMU pre-integration result and the initialization parameters are solved.

[0086] Step 4: Using the visual parameters obtained in steps 1 and 2 and the IMU pre-integration results, the point-line reprojection residuals and the IMU data measurement pre-integration residuals are calculated and sent to the localized sliding window optimization part using the Robot Operating System (ROS) topic publish / subscribe mechanism. Nonlinear optimization is performed using the sliding window visual constraints, IMU constraints, and landmark points. The three form a close constraint relationship. The point-line reprojection residuals and the IMU pre-integration residuals are used as optimization targets to optimize the camera pose, IMU attitude, and the positions of the landmark points, thereby obtaining the optimized camera pose graph.

[0087] Through the camera pose graph, the actual motion trajectory of the robot can be obtained, thereby achieving the purpose of positioning.

[0088] The state variables in the sliding window are defined as follows:

[0089] x k Indicates the state of the IMU at the Kth frame of the camera, including the bias of the accelerometer and gyroscope and the pose information estimated by the pre-integration of the IMU. n represents the number of key frames, m is the inverse depth of all point features, and the orthogonal coordinate parameters of o line features. Indicates the external parameters from the camera to the IMU.

[0090] Using BA (Bundles Adjustment) nonlinear optimization, the sum of the prior and Mahalanobis distances of all measurement residuals is minimized to obtain the maximum a posteriori estimate, thereby obtaining the nonlinear optimization objective function, that is,

[0091] where {r p , H p} represents the prior data of sliding window marginalization, r point represents the point reprojection error, r line represents the line reprojection error, r imu Represents the IMU measurement error. Finally, the nonlinear solver cere-solver is used to perform operations to solve the nonlinear problem.

[0092] To verify the positioning performance of this invention and evaluate its accuracy, we compared it with the current classic algorithms, VINS_mono and PL_vins. In the experiment, we used EVO to align timestamps with the true values ​​and calculate the absolute trajectory error (APE). The root mean square error (RMSE / m) was used to evaluate the system.

[0093] To demonstrate the advantages of the present invention, the VINS_mono algorithm was compared with the present invention using the MH_04_difficuil (MH_04) sequence. The MH_04 sequence is a complex environment with a large number of line textures, which better reflects the system's accuracy. Figure 7 shows the effect of point and line features extracted by the present invention when running on the MH_04 sequence. The MH_04 sequence suffers from complex lighting conditions and sparse environmental textures. The present invention adds an improved line feature extraction thread to compensate for the lack of point features.

[0094] Figure 8 shows the APE error heatmaps between the test trajectories and the true trajectories of the present invention and the VINS_mono algorithm under MH_04, respectively. It can be seen that both trajectories are very close to the true trajectories. The maximum error value of VINS_Mono is 1.316m, and the minimum is 0.021m, while the maximum error value of the present invention is 0.395m, and the minimum is 0.035m. Furthermore, the APE of the present invention varies steadily and evenly, while the APE of VINS_Mono varies dramatically, indicating that the present invention is more stable and more robust. Calculated root mean square error of VINS_Mono is 0.398084m, while the root mean square error of the present invention is 0.189329m. This shows that the trajectory proposed by the present invention can effectively improve the positioning accuracy of the system, especially in environments with a lot of line textures.

[0095] Figure 9 shows the time-varying X, Y, and Z-axis curves of the experimental and actual trajectories for the VINS_mono algorithm and the present invention (IPLI) under MH_04. It can be seen that both algorithms are close to the true curve on the X-axis, but the VINS_mono algorithm's curve is jagged and nearly jagged on the Y and X-axes, while the present invention's curve is smooth and generally close to the true curve.

[0096] Table 2 RMSE / m of different algorithms

[0097] As can be seen from Table 2, the present invention is superior to VINS_mono and PL_vins. Compared with the VINS_mono algorithm, the present invention significantly reduces the RMSE, with an average reduction of 50.7%; compared with the PL_vins algorithm, the present invention partially reduces the RMSE, with an average reduction of 13.2%. In the VI sequence, the RMSE is reduced by 55.6% compared with the VIMS-mono algorithm and by 25.1% compared with the PL-vins algorithm; in the MH sequence, the RMSE is reduced by 47.8% compared with the VIMS-mono algorithm and by 6.1% compared with the PL-vins algorithm. The VI sequence has different problems such as complex scenes, fast motion, blurred scenes and sparse textures, but the RMSE value of the present invention is still significantly reduced in the VI sequence, which shows that the present invention is suitable for complex scenes and illustrates the robustness and stability of the present invention.

[0098] To verify the accuracy of this invention, an experiment was conducted in a library setting. An experimenter walked around the room holding a laptop computer equipped with a D435i camera. The D435i camera features a visual processor, an RGB sensor, a color image signal processing module, and a built-in 6-axis IMU. After calibration, the D435i camera met the required error, thus meeting the experimental requirements. However, due to the randomness and errors of human operation, as well as errors in hardware such as the D435i calibration, only qualitative comparisons were performed.

[0099] Figure 10 shows the results of point and line feature extraction in an experimental library. The library's floor and ceiling have sparse, monotonous textures, so extracting single point features can lead to insufficient feature points. The library is home to numerous bookshelves, resulting in numerous areas with dense lines, which can easily lead to mismatches and affect accuracy. During testing, the experimenter held a handheld device and moved at a constant speed throughout the library. Starting from the starting point, they passed through an aisle between the bookshelves and then turned into a shelf. The light intensity at these turns increased, indicating a dramatic change in lighting. Finally, they returned to the starting point along the same route. The accuracy of the algorithm was measured by comparing the degree of overlap between the starting and ending points.

[0100] As shown in Figure 11, which is a planar display of the experimental running trajectory on the xy axis, the distance between the starting point and the end point of the positioning trajectory of the present invention is only 0.128m different. The positioning trajectory of the original VINS_mono algorithm not only deviates from the true trajectory, but also has a difference of 0.371m between the starting point and the departure point. This fully demonstrates the effectiveness of the present invention and improves the positioning accuracy and robustness of the system.

[0101] Table 3 Distance difference between different algorithms (unit: m)

[0102] Table 3 shows the distances between the starting points and the end points of the present invention and the VINS_Mono algorithm under two different running paths in the library. The results show that the distances between the starting points and the end points of the present invention are smaller than those of the VINS_Mono algorithm in both tests, indicating that the present invention improves positioning accuracy and has higher robustness.

[0103] The above is a schematic description of the present invention and its embodiments, which is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. Therefore, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs a structure and embodiment similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.

Claims

1. A robot positioning method based on improved point-line features and IMU fusion, characterized in that: The steps are: Step 1: Use the Shi-Tomasi algorithm to extract feature points in the captured image, and use LK optical flow for tracking and matching. At the same time, use the improved LSD line segment detection method to extract line features in the image, and use the line segment merging algorithm to merge short line segments on the same straight line, and then use the LBD descriptor for matching; Step 2: Pre-integrate the collected IMU data; Step 3: Adopt the same initialization strategy as VINS-Mono, tightly couple the visual information and IMU information obtained in steps 1 and 2, use the visual SFM to estimate the pose between consecutive frames in the sliding window and the inverse depth of the 3D point, and finally align it with the IMU pre-integration result and solve the initialization parameters; Step 4: Using the visual parameters obtained in steps 1 and 2 and the IMU pre-integration results, the point-line reprojection residuals and the IMU data measurement pre-integration residuals are calculated. Nonlinear optimization is performed through sliding window visual constraints, IMU constraints, and landmark points. The three form a close constraint relationship, among which the point-line reprojection residuals and the IMU pre-integration residuals are used as optimization targets to optimize the camera pose, IMU posture, and the position of landmark points, and then the optimized camera pose graph is obtained; through the camera pose graph, the actual motion trajectory of the robot is obtained to achieve the positioning purpose.

2. The robot positioning method based on improved point-line features and IMU fusion according to claim 1, characterized in that: In step 1, before using the improved LSD line segment detection method to extract straight line features in the image, a filter is first set to screen the line-dense areas in the image. The filter detects areas with high pixel density in the grayscale image based on pixel gradients, and changes the pixels in the area to a uniform pixel value. Then, LSD is called to perform straight line detection on the processed grayscale image.

3. The robot positioning method based on improved point-line features and IMU fusion according to claim 2, characterized in that: The filtering process of the filter is: First, traverse a grayscale image pixel and calculate the pixel gradient value. For pixel K, the gradient value is Kgra, and the gradient threshold G is set; Secondly, define the regional gradient density number T. When Kgra>G, take the pixel K coordinate as the center and traverse the gradient values ​​of all pixels Kn in the t×t area. The number of gradient values ​​of all pixels Kn in the area where Kngra>G is the regional gradient density number T. Then set the regional gradient density threshold When T is greater than When , all pixel values ​​in the 3t×3t area are changed.

4. The robot positioning method based on improved point-line features and IMU fusion according to claim 3 is characterized in that: The line segment merging algorithm first groups the line segments based on the line segment angle difference and the distance from the endpoint to the straight line, then implements different sorting schemes according to the line segment angle to achieve orderly arrangement of the line segments based on relative positions, and finally compares the distance between the line segment endpoints and the sum of the line segment lengths to determine whether they belong to the same straight line, and screens and eliminates the merged straight lines based on the LBD descriptor.

5. The robot positioning method based on improved point-line features and IMU fusion according to claim 3 is characterized in that: The specific process of the line segment merging algorithm is as follows: (1) Sort the line segments extracted by LSD in descending order according to their lengths to obtain L = {L1, L2, L3, …, Ln}; for each line segment L i , based on the angle, we filter the candidate line segment group L θ , the longest line segment in each group is called the main line segment l m ; (2) Calculate line segment L j Starting point to L i The distance d from the straight line s , line segment L j End at L i The distance d from the straight line e , and then get the line segment L j From both ends to L i The distance d of the straight line, d=[d s ,d e ]; (3) For the candidate line segment group L θ The final candidate line segment group L is obtained by distance screening again g ; (4) For the grouped line segments, the line segments belonging to the same straight line are merged through the line segment merging strategy, and the line segments that are incorrectly grouped are eliminated.

6. The robot positioning method based on improved point-line features and IMU fusion according to claim 5, characterized in that: In step (1), grouping starts from the longest line segment whose length is less than the set threshold L min The line segments are removed and not grouped.

7. The robot positioning method based on improved point-line features and IMU fusion according to claim 6, characterized in that: The angle screening strategy in step (1) is: Where: L j Indicates the length ratio L in the line segment group L i Short line segment, θ i and θ j L i and L j Angle with the x-axis, θ th Indicates L i and L j The angle difference threshold between 8. The robot positioning method based on improved point-line features and IMU fusion according to claim 7, characterized in that: The distance screening strategy in step (3) is: Where: d th Indicates the distance threshold from the endpoint to the straight line.

9. The robot positioning method based on improved point-line features and IMU fusion according to claim 8, characterized in that: Step (4) When merging line segments, according to the main line segment l m Angle θ with the x-axis m Execute different sorting strategies, align the line segment groups according to the direction of the main line segment, divide the sorted line segments into two groups according to the position of the main line segment, and perform prefix sum calculation on each group. When calculating the prefix sum, each group of line segments must include the main line segment. Start from the line segment farthest from the main line segment and calculate the main line segment l m Starting endpoints m , and the starting endpoint s of the line segment farthest from the main line segment j The distance d is calculated together with the total length l of the corresponding line segment in the prefix sum calculated previously. sum For comparison, if the total length of the line segment is l sum The proportion of distance d is greater than the threshold r th , then it is considered that from the main line segment l m All line segments between the farthest line segment belong to the same long line segment, connecting the main line segment l m The starting point of the farthest line segment and the end point of the farthest line segment are used as the merged line segment, otherwise the farthest line segment is released back to the line segment group; at the same time, the second farthest line segment is continued to be judged.

10. The robot positioning method based on improved point-line features and IMU fusion according to claim 9, characterized in that: In the process of matching using LBD descriptors, the descriptors of the merged line segment and the main line segment are calculated, and the descriptor distance is calculated. When the descriptor distance is greater than the threshold D th , abandon the current merge and keep only the main line segment as the result of the current merge.

Citation Information

Patent Citations

  • Line feature matching method based on point-line-surface fusion

    CN115512138A

  • Robot positioning method based on improved point-line feature and IMU fusion

    CN117576409A

  • System and method for automated merging

    US20230091276A1

Cited By

  • Intelligent electric energy metering box state identification method and system

    CN121214414A

  • Interaction method based on earphone camera, earphone and storage medium

    CN122387409A