A Visual Localization and Mapping Method, System, Device and Medium Based on Geometry and Semantics

By combining geometric and semantic methods, using neural networks and binomial probability models and other technologies, we eliminate dynamic objects in visual SLAM systems, solving the problem of insufficient positioning and mapping accuracy in high dynamic scenarios, achieving higher robustness and accuracy.

CN117036484BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311080521.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-25
Publication Date
2025-07-01
Estimated Expiration
2043-08-25

AI Technical Summary

Technical Problem

Existing visual SLAM technology is difficult to effectively eliminate dynamic objects in high dynamic scenarios, resulting in insufficient accuracy and robustness of positioning and map construction.

Method used

Geometry and semantics are used to perform semantic segmentation and feature point extraction through neural networks, and dynamic objects are further eliminated using binomial probability model and geometric correlation constraint algorithm to improve the accuracy of positioning and graph construction.

Benefits of technology

It significantly improves the accuracy and robustness of positioning and mapping of visual SLAM systems in high dynamic scenarios, and reduces poor and unstable data associations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117036484B_ABST
    Figure CN117036484B_ABST
Patent Text Reader

Abstract

A method, system, device and medium for visual localization and mapping based on geometry and semantics. The method includes: inputting an original image into a Mask R-CNN network model to segment instance labels and obtain all dynamic object masks of the image; inputting the original image into a GCN network model to extract all feature points; calculating the semantic segmentation dynamic probability of each key point for the image binary mask and all feature points; removing dynamic objects from the original image scene and tracking the image; discriminating the current image frame and screening dynamic features that cannot be removed by the semantic method; inputting the completely processed image frame into a tracking and mapping thread to estimate the camera pose through frame-by-frame association; optimizing loop detection of the global map and global Bundle Adjustment; the system, device and medium are used to implement this method; the present invention improves the accuracy and robustness of the visual SLAM system in real high-dynamic scenes, reduces bad and unstable data associations, and improves the success rate of intelligent robots in performing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent robots, and particularly relates to a vision positioning and mapping method, system, device and medium based on geometry and semantics. Background Technique

[0002] Visual SLAM (Simultaneous Localization and Mapping) is an advanced computer vision technology used for simultaneously and real-time localizing and constructing a map of an unknown environment, and is widely applied in fields such as autonomous navigation, augmented reality, robotics, and virtual reality. Traditionally, navigation and map construction are carried out separately, where navigation relies on a pre-constructed map, and the map requires known position information. However, for many practical applications, such as unmanned aerial vehicles, unmanned vehicles, and robots, the environment may be unknown or dynamically changing, which requires real-time localization and map construction, that is, the emergence of SLAM technology. Visual SLAM utilizes computer vision technology to simultaneously estimate its own position (localization) and construct an environmental map in an unknown environment by processing visual input data from a camera or cameras. It is based on key technologies such as visual feature extraction, feature matching, pose estimation, and three-dimensional reconstruction, and has the following key advantages: 1. Real-time: A visual SLAM system can process sensor data in a real-time environment and quickly update the localization information and the map. 2. No prior information required: Compared with traditional methods, visual SLAM does not require prior maps or prior information such as GPS, and is applicable to unknown or dynamic environments. 3. Accuracy: With the continuous development of computer vision and sensor technologies, the localization and map construction accuracy of visual SLAM is continuously improving. 4. Independence: Compared with other SLAM technologies, such as laser SLAM, visual SLAM only relies on a camera or cameras and does not require additional sensors, reducing the system complexity and cost. However, visual SLAM technology also faces some challenges, such as situations like occlusion, fast movement, and large-scale environments that may affect the accuracy of localization and mapping.

[0003] The patent application with the publication number CN114627184A provides a binocular vision positioning and mapping method for a mobile robot based on the direct method, which segments an image and selects feature points with regional thresholds, designs an initialization algorithm based on bidirectional static stereo matching, and uses a backend graph optimization model based on radiometric error constraints to solve the problem of insufficient real-time performance in a visual SLAM system. However, its system relies on a static environment, making the system unable to robustly operate in an actual high-dynamic scenario.

[0004] The patent application with the publication number CN110838145A provides a method for visual localization and mapping of indoor dynamic scenes, which extracts and matches feature points from query images and target images, accelerates the matching of feature points through the bag-of-words model, adds precision constraints and geometric constraints to the matched feature points, and uses a two-way scoring rule to double-constrain the feature points to obtain the finally filtered feature points, and adds them to the ORB-SLAM visual odometry framework, improving the accuracy of localization and mapping. However, its filtering of dynamic objects only uses the two-way scoring rule and cannot completely eliminate the dynamic objects in the scene. At the same time, it does not optimize the feature points misjudged by the two-way scoring rule. Summary of the Invention

[0005] In order to overcome the above-mentioned deficiencies of the prior art, the purpose of the present invention is to provide a method, system, device and medium for visual localization and mapping based on geometry and semantics. By using a neural network for semantic segmentation, feature point extraction and using a binomial probability model to improve the edges of segmentation information, dynamic objects in the scene are eliminated. By using geometric association constraints, dynamic objects that cannot be processed by semantic segmentation are further supplemented, and it is possible to better eliminate bad dynamic information in the scene. The present invention can effectively improve the accuracy and robustness of the visual SLAM system in real high-dynamic scenes, and greatly reduce bad and unstable data associations.

[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0007] A method for visual localization and mapping based on geometry and semantics, comprising the following steps:

[0008] Step S1: Input the original image into the Mask R-CNN network model, perform per-pixel semantic segmentation on the original image to obtain instance labels and a binary mask of the image. The segmented instance labels are used to track different objects, and all dynamic objects appearing in the original image scene are obtained according to the binary mask of the image.

[0009] Step S2: Input the original image into the GCN network model, extract the features in the original image to obtain all the feature points on the image.

[0010] Step S3: For the binary mask of the image obtained by the Mask R-CNN network model in Step S1 and all the feature points obtained by the GCN network model in Step S2, use the binomial logistic regression model to calculate the semantic segmentation dynamic probability of each key point. Feature points with a probability less than 0.75 in the binary mask of the image are considered static feature points, and the feature points in the processed image are obtained.

[0011] Step S4: After removing the dynamic objects in the original image scene using the binary mask of the image obtained by the Mask R-CNN network model in Step S1, use the feature points in the processed image obtained in Step S3 for positioning to obtain the key frame sequence of the current frame image;

[0012] Step S5: Compare the angle and distance of the features of the current frame image with those of the highest overlapping key frame selected from the corresponding key frame sequence obtained in Step S4. If the feature angles of the two frame images exceed the preset threshold, the features of the current frame image are considered dynamic features. In the case where the feature angles do not exceed the threshold, then judge whether the depth difference exceeds the preset threshold. If the depth difference exceeds a certain threshold, it is considered a dynamic feature, otherwise the feature is considered a static feature. After the comparison, obtain the static feature points in the completely processed image;

[0013] Step S6: Input the static feature points in the image completely processed in Step S5 into the tracking and mapping thread, estimate the camera pose through frame-by-frame association, complete the mapping work, and obtain the global map;

[0014] Step S7: Perform loop detection and global Bundle Adjustment optimization on the global map obtained in Step S6 to optimize the accuracy of global map positioning.

[0015] The binary activation layer of the GCN network model in Step S2 is as follows:

[0016] Forward:

[0017] Backward:

[0018] where b is the binarized version of the feature f, 1 |f|≤1 cancels the gradients representing the feature responses of f with absolute values greater than 1, 1 |f|≤1 is the straight-through estimator of the so-called hard sign function for backpropagation gradients.

[0019] The binomial logistic regression model in Step S3 includes: semantic segmentation result label, set of semantic segmentation result labels of feature points, set of geometric labels of pixel points, distance between feature points and boundary pixel points, and definition of the regression model;

[0020] The definition of the semantic segmentation result label is as follows:

[0021]

[0022] where, represents the semantic segmentation result label of the feature point p i at time t;

[0023] The set of semantic segmentation result labels of the feature points is defined as follows:

[0024]

[0025] Among them, s t represents the set of semantic segmentation result labels of the feature points, where n is the number of feature points;

[0026] The set of geometric labels of the pixel points is defined as follows:

[0027]

[0028] Among them, b t is the set of boundary pixel points at time t , where m is the number of boundary pixel points, and all boundary points are included in b t ;

[0029] The distance between the feature points and the boundary pixel points is defined as follows:

[0030]

[0031] Among them, dist(p i , b t ) represents the distance between the feature point p i and the boundary pixel points;

[0032] The definition of the regression model is as follows:

[0033]

[0034] Among them, represents the dynamic probability of the semantic segmentation feature point p i , where α is the influence factor for balancing the detection result curve.

[0035] In step S5, the angles and distances between the features of the current frame image and the features of the corresponding key frame with the highest overlap are compared, and the dynamic feature points are eliminated through the geometric association constraint algorithm. The process of the geometric association constraint algorithm is as follows:

[0036]

[0037]

[0038] Among them, if the return value is SF, it is determined that the current feature on the current frame image is a static feature; if the return value is DF, it is determined that the current feature on the current frame image is a dynamic feature.

[0039] In the process of performing pose estimation and constructing a global map frame by frame in step S6, the following steps are included:

[0040] Step S6.1: Check whether there are key frames in the buffer queue;

[0041] Step S6.2: Take out the first key frame in the buffer queue in step S6.1 for processing;

[0042] Step S6.3: Remove the bad points that appear in the map;

[0043] Step S6.4: Use the epipolar geometry or triangulation method to create new map points to supplement the map after removing bad points in step S6.3;

[0044] Step S6.5: Fuse the map points formed by the first key frame taken out in step S6.2 and its co-visible key frames;

[0045] Step S6.6: Perform local BA optimization on the map in step S6.3;

[0046] Step S6.7: Remove the redundant key frames in the buffer queue in step S6.1;

[0047] Step S6.8: Add the current key frame taken out in step S6.2 to the loop detection thread.

[0048] The process of loop detection in step S7 includes the following steps:

[0049] Step S7.1.1: Take out the key frame at the head of the buffer queue as the current loop detection key frame;

[0050] Step S7.1.2: If the time since the last loop detection is less than 5 seconds, do not perform loop detection;

[0051] Step S7.1.3: Calculate the maximum similarity between the current loop detection key frame in step S7.1.1 and its co-visible key frames;

[0052] Step S7.1.4: Find the loop candidate key frame of the current loop detection key frame according to the maximum similarity obtained in step S7.1.3;

[0053] Step S7.1.5: Find a match between the loop candidate key frame group formed by the loop candidate key frames in step S7.1.4 and the loop candidate key frame group existing in the loop variable;

[0054] Step S7.1.6: Maintain the loop variable so that the loop candidate key frame group formed by the loop candidate key frames in step S7.1.4 is used as the loop candidate key frame group before the next frame;

[0055] The global Bundle Adjustment optimization of the global map obtained in step S6 in step S7 refers to using the minimization of reprojection error to optimize the solution of moving points, obtaining the rotation matrix R and the translation matrix t, and using the obtained rotation matrix R and translation matrix t to optimize the poses of points on the map;

[0056] Among them, the global Bundle Adjustment optimization of the global map is performed using the observation model:

[0057] Step S7.2.1, World coordinate system to camera coordinate system:

[0058] Let the world coordinates of point P be (x, y, z), and the coordinates transferred to the camera are:

[0059] P′ = RP + t = (x′, y′, z′)

[0060] Step S7.2.2, Normalization:

[0061] P′ C = [u c , v c , 1] T = [x′ / z′, y′ / z′, 1] T

[0062] Step S7.2.3, Distortion removal:

[0063] P′ C (u c , v c ) → (u′ c , v′ c )

[0064] Step S7.2.4, Camera coordinate system to pixel coordinate system:

[0065] u s = f x ·u′ c + c x

[0066] v s = f y ·v′ c + c y

[0067] Step S7.2.5, Camera observation model:

[0068] g(u s , v s ) = h(R, t, x, y, z)

[0069] Step S7.2.6, Reprojection error model:

[0070] e = z - h(R, t, x, y, z)

[0071] Where z is the observed value and h is the value obtained from the estimation model.

[0072] A vision positioning and mapping system based on geometry and semantics, comprising:

[0073] Vision odometer module: used to input the original image into the Mask R-CNN network model, perform per-pixel semantic segmentation on the original image to obtain instance labels and the binary mask of the image, and the segmented instance labels are used to track different objects; input the original image into the GCN network model, extract the features in the original image to obtain all the feature points on the image; calculate the semantic segmentation dynamic probability of each key point through the binomial logistic regression model; perform positioning by using the feature points in the processed image to obtain key frames of several frames of images in a small range; eliminate dynamic feature points through the geometric association constraint algorithm;

[0074] Mapping module: used to input the static feature points in the fully processed image into the tracking and mapping thread, estimate the camera pose through frame-by-frame association, complete the mapping work, and obtain the global map;

[0075] Nonlinear optimization module: used to perform loop detection and global Bundle Adjustment optimization on the global map to optimize the accuracy of global map positioning.

[0076] The present invention also provides a vision positioning and mapping device based on geometry and semantics, comprising:

[0077] Memory: storing the computer program of the above-mentioned vision positioning and mapping method based on geometry and semantics, which is a computer-readable device;

[0078] Processor: used to implement the above-mentioned vision positioning and mapping method based on geometry and semantics when executing the computer program.

[0079] The present invention also provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the above-mentioned vision positioning and mapping method based on geometry and semantics.

[0080] Compared with the prior art, the beneficial effects of the present invention are:

[0081] 1. The present invention provides a vision localization and mapping method based on geometry and semantics. First, semantic segmentation is performed using a neural network, and the edges of the segmentation information are refined using a binomial probability model to eliminate dynamic objects in the scene. The GCN network model is used to extract features in the image, and the extracted features are made more repeatable and widely distributed, enabling more accurate subsequent localization and mapping work. The geometric association constraint method is used to further supplement dynamic objects that cannot be processed by semantic segmentation, which can better eliminate bad dynamic information in the scene and effectively reduce unstable and bad data associations. The processed images are used to estimate the camera pose through frame-by-frame association, and a series of tasks such as backend optimization and mapping are completed.

[0082] 2. The present invention enhances the environmental perception ability of intelligent robots by combining semantic and geometric methods. The GCN network model is used to extract features in the image, enhancing the wide distribution of the features. At the same time, the present invention adopts a method combining semantics and geometry for dynamic objects in large scenes to eliminate dynamic targets in the scene, enabling more accurate localization.

[0083] 3. By adopting the Mask R-CNN network model, the present invention can perform per-pixel semantic segmentation on the original image to obtain instance labels and a binary mask of the image. The instance labels can be used to track different objects, and the binary mask can obtain all dynamic objects in the image, thus enhancing the ability to eliminate dynamic semantic information in the scene.

[0084] 4. Due to the adoption of the GCN network model, the present invention can extract all feature points widely distributed on the image, making the extracted feature points more representative and trackable.

[0085] 5. Due to the adoption of the binomial logistic regression model, the present invention can calculate the semantic segmentation dynamic probability of each key point. Feature points with a probability lower than 0.75 can be reprocessed as static feature points, making the semantic information of the feature points on the mask edge obtained by segmentation using the Mask R-CNN network model more accurate.

[0086] 6. Due to the adoption of the geometric association constraint method, the present invention can use two constraint conditions of angle and distance to further effectively discriminate features, and can also effectively segment dynamic objects that cannot be processed by the Mask R-CNN network model and the binomial logistic regression model, thus enhancing the adaptability of the system in high-dynamic scenes.

[0087] In summary, the present invention solves the problem of environmental perception of intelligent robots in unknown dynamic environments, improves the success rate of intelligent robots in performing tasks, and promotes the development of the intelligent robot industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 It is a flowchart of a vision localization and mapping method based on geometry and semantics.

[0089] Figure 2 It is a demonstration diagram of the results of a binomial logistic regression model.

[0090] Figure 3 It is a schematic diagram of a geometric association constraint dynamic feature point screening algorithm; among them, Figure 3 (a) is a diagram of the present invention determining a feature point as a static feature point through the angle constraint in the geometric association constraint algorithm, Figure 3 (b) is a diagram of the present invention determining a feature point as a static feature point through the angle and distance constraints in the geometric association constraint algorithm, Figure 3 (c) is a diagram of the present invention determining a feature point as a dynamic feature point through the angle and distance constraints in the geometric association constraint algorithm.

[0091] Figure 4 It is a schematic diagram of the feature point extraction result of the present invention.

[0092] Figure 5 It is a schematic diagram of the estimated trajectory of the actual scene of the present invention; among them, Figure 5 (a) is a trajectory diagram of the present invention running in the fr3_walking_xyz sequence in the TUM dataset, Figure 5 (b) is a trajectory diagram of the present invention running in the fr3_walking_static sequence in the TUM dataset, Figure 5 (c) is a trajectory diagram of the present invention running in the fr3_walking_rpy sequence in the TUM dataset, Figure 5 (d) is a trajectory diagram of the present invention running in the fr3_walking_half sequence in the TUM dataset. Detailed implementation manner

[0093] The present invention will be further described in detail below with reference to the accompanying drawings.

[0094] The present invention designs a method for eliminating dynamic objects in high-dynamic scenes based on geometry and semantics, reducing the influence of the system by bad and unstable data associations, and at the same time optimizing the problem of misdiagnosis in the semantic method at the masked edge of dynamic objects, enabling the entire system to achieve better results in actual high-dynamic scenes.

[0095] As Figures 1-5 shown, the present invention discloses a specific embodiment, that is, a vision localization and mapping method based on geometry and semantics, including the following steps:

[0096] As Figure 1As shown in the figure, in step S1, the original image is input into the Mask R-CNN network model. After processing, this model can perform per-pixel semantic segmentation on the image. The instance labels obtained by segmentation can track different objects. By uniquely labeling each object, Mask R-CNN enables us to accurately track the movement and changes of these objects in the image sequence. The output binary mask provides us with an effective means to obtain all the dynamic objects in the image scene. By applying these masks, the dynamic objects can be extracted from the image.

[0097] Among them, the Mask R-CNN network model is trained through the COCO dataset and has shown excellent performance in the instance segmentation and object detection tasks of the COCO dataset;

[0098] When an original image is processed by the Mask R-CNN network model, it includes the following steps:

[0099] Step S1.1: Preprocess the data of the input original image, that is, normalize and scale the size;

[0100] Step S1.2: Input the image preprocessed in step S1.1 into the pre-trained backbone feature extraction network to obtain the feature map of the image;

[0101] Step S1.3: Set the region of interest (ROI) through each point in the feature map obtained in step S1.2 to obtain multiple ROI candidate boxes;

[0102] Step S1.4: Perform binary classification and bounding box regression (BB regression) on the ROI candidate boxes obtained in step S1.3 to filter out some candidate ROI candidate boxes;

[0103] Step S1.5: Perform region of interest alignment (ROIAlign) operation on the remaining ROI candidate boxes in step S1.4;

[0104] Step S1.6: Perform classification and BB regression on the ROI candidate boxes in step S1.5 to generate a mask.

[0105] Step S2: Input the original image into the GCN network model. This model extracts the key features in the image through a complex calculation process. These features are widely distributed in the image and have good repeatability. They can appear stably in different images and can express the content of the image more comprehensively.

[0106] These features extracted by the GCN network model can play an important role in subsequent localization and mapping tasks. They can be used for image or scene localization to help identify location information in the image. In addition, these features can also be used for mapping, that is, constructing a structured representation of the image or environment, enabling objects, entities, or locations in the image to be more precisely labeled and understood.

[0107] The binary activation layer of the GCN network model is as follows:

[0108] Forward:

[0109] Backward:

[0110] where b is the binary version of the feature f.1 |f|≤1 The gradients representing the responses of each feature of f with absolute values greater than 1 are cancelled. It is a straight-through estimator of the so-called hard sign function used for backpropagating gradients.

[0111] As Figure 2 shown, in step S3, for the binary mask of the image obtained by the Mask R-CNN network model in step S1 and all the feature points obtained by the GCN network model in step S2, a binomial logistic regression model is used to calculate the semantic segmentation dynamic probability of each key point, and then to determine whether the features at the mask edge are static features. Feature points with probabilities less than 0.75 in the binary mask of the image are considered static feature points;

[0112] The binomial logistic regression model includes: the semantic segmentation result label, the set of feature point semantic segmentation result labels, the set of pixel point geometric labels, the distance between the feature point and the boundary pixel point, and the definition of the regression model;

[0113] The semantic segmentation result label is defined as follows:

[0114]

[0115] where represents the semantic segmentation result label of the feature point p i at time t.

[0116] The set of feature point semantic segmentation result labels is defined as follows:

[0117]

[0118] where s t represents the set of feature point semantic segmentation result labels with n being the number of feature points.

[0119] The set of pixel geometric tags is defined as follows:

[0120]

[0121] Among them, b t is the set of boundary pixels at time t , where m is the number of boundary pixels, and all boundary points are included in b t .

[0122] The distance between the feature point and the boundary pixel is defined as follows:

[0123]

[0124] Among them, dist(p i , b t ) represents the distance between the feature point p i and the boundary pixel.

[0125] The regression model is defined as follows:

[0126]

[0127] Among them, represents the dynamic probability of the semantic segmentation feature point p i , where α is the influence factor for balancing the detection result curve, which is set to 0.1 in the embodiment.

[0128] Step S4: After removing the dynamic objects in the original image scene using the binary mask of the image obtained by the Mask R-CNN network model in step S1, use the feature points in the processed image obtained in step S3 for positioning to obtain the key frame sequence of the current frame image, preparing for the subsequent geometric association constraint method and enabling the subsequent geometric association constraint method to be effectively carried out; Using feature points for positioning has a very low computational cost and searches for the corresponding relationship of feature points in the static area of the image.

[0129] Step S4 is to perform lightweight version tracking on the image frame after dynamic feature removal.

[0130] Using feature points for positioning mainly uses the reference key frame for positioning, and the positioning process includes the following steps:

[0131] Step S4.1: Calculate the bag of words (BoW) of the current frame;

[0132] Step S4.2: Accelerate the feature point matching between the current frame and the reference frame through the BoW obtained in step S4.1;

[0133] Step S4.3: Use the pose of the previous frame as the initial value of the current frame's pose, and obtain an accurate pose estimate by optimizing the reprojection error of 3D-2D.

[0134] As Figure 3 shown, in step S5, a method based on angle and distance is used to determine whether there are dynamic features between image frames. First, compare the features of the current frame image with the features of the key frame with the highest overlap selected from the corresponding key frame sequence obtained in step S4. If the angular difference between the features of the two frame images exceeds a preset threshold, the features of the current frame image are considered dynamic features. If the angle does not exceed the threshold, then determine whether the depth difference exceeds the preset threshold. If it exceeds, it is identified as a dynamic feature.

[0135] Eliminate dynamic feature points through a geometric correlation constraint algorithm. The process of the geometric correlation constraint algorithm is as follows:

[0136]

[0137] Among them, if the return value is SF, it is determined that the current feature on the current frame image is a static feature. If the return value is DF, it is determined that the current feature on the current frame image is a dynamic feature.

[0138] Step S6: Input the image frame into a geometric and semantic-based visual localization and mapping system, and realize camera pose estimation and mapping work through frame-by-frame association to obtain a global map; at the same time, use the estimation result of the camera pose to further optimize the camera's trajectory and attitude information to improve the accuracy and stability of positioning; if a failure occurs during the process of estimating the camera pose through frame-by-frame association, repositioning estimation can also be performed to avoid incorrect output of the entire pose estimation sequence.

[0139] The above step S6 performs frame-by-frame association on the image frames after dynamic object elimination, so as to realize camera pose estimation and mapping work.

[0140] The process of pose estimation and global map construction through frame-by-frame association in the above step S6 includes the following steps:

[0141] Step S6.1: Check whether there are key frames in the buffer queue;

[0142] Step S6.2: Take out the first key frame in the buffer queue in step S6.1 for processing;

[0143] Step S6.3: Eliminate the bad points that appear in the map;

[0144] Step S6.4: Use the epipolar geometry or triangulation method to create new map points to supplement the map after eliminating bad points in step S6.3;

[0145] Step S6.5: Fuse the first key frame taken out in step S6.2 with the map points formed by its co-visible key frames;

[0146] Step S6.6: Perform local BA optimization on the map in step S6.3;

[0147] Step S6.7: Remove redundant key frames from the buffer queue in step S6.1;

[0148] Step S6.8: Add the current key frame taken out in step S6.2 to the loop closure detection thread.

[0149] Step S7: Through loop closure detection, identify and correct position deviations caused by sensor errors or cumulative errors, thereby improving the accuracy and consistency of positioning. Through global Bundle Adjustment, optimize the poses, depths, and reprojection errors of all image frames and feature points to further improve the accuracy and stability of global map positioning.

[0150] The said step S7 is to further optimize the global positioning and mapping through loop closure detection and global Bundle Adjustment after completing the positioning and simple mapping work.

[0151] Loop closure detection can perform consistent matching on the same locations identified during the positioning process. Global Bundle Adjustment optimizes the results of pose estimation and the landmarks in the map. Loop closure detection and global Bundle Adjustment can greatly improve the accuracy of positioning and mapping;

[0152] The process of loop closure detection in the said step S7 includes the following steps:

[0153] Step S7.1.1: Take out the key frame at the head of the buffer queue as the current loop closure detection key frame;

[0154] Step S7.1.2: If the time since the last loop closure detection is less than 5 seconds, do not perform loop closure detection;

[0155] Step S7.1.3: Calculate the maximum similarity between the current loop closure detection key frame in step S7.1.1 and its co-visible key frames;

[0156] Step S7.1.4: Find the loop closure candidate key frame of the current loop closure detection key frame according to the maximum similarity obtained in step S7.1.3;

[0157] Step S7.1.5: Find a match between the loop closure candidate key frame group formed by the loop closure candidate key frames in step S7.1.4 and the loop closure candidate key frame group existing in the loop variable;

[0158] Step S7.1.6, maintain the loop variable so that the loop candidate keyframe group formed by the loop candidate keyframes in Step S7.1.4 is used as the loop candidate keyframe group before the next frame;

[0159] In Step S7, the global map obtained in Step S6 is optimized by global Bundle Adjustment, which means using the minimization of reprojection error to optimize the solution of moving points to obtain the rotation matrix R and the translation matrix t, and using the obtained rotation matrix R and translation matrix t to optimize the poses of points on the map;

[0160] Among them, the global map is optimized by global Bundle Adjustment using the observation model:

[0161] Step S7.2.1, convert from world coordinate system to camera coordinate system

[0162] Assume that the world coordinates of point P are (x, y, z), and the conversion to the camera coordinate system is:

[0163] P′ = RP + t = (x′, y′, z′)

[0164] Step S7.2.2, normalization:

[0165] P′ C = [u c , v c , 1] T = [x′ / z′, y′ / z′, 1] T

[0166] Step S7.2.3, distortion removal:

[0167] P′ C (u c , v c ) → (u′ c , v′ c )

[0168] Step S7.2.4, convert from camera coordinate system to pixel coordinate system:

[0169] u s = f x ·u′ c + c x

[0170] v s = f y ·v′ c + c y

[0171] Step S7.2.5, camera observation model:

[0172] g(us , v s ) = h(R, t, x, y, z)

[0173] Step S7.2.6, reprojection error model:

[0174] e = z - h(R, t, x, y, z)

[0175] where z is the observed value and h is the value obtained from the estimation model.

[0176] As Figure 4 shown, the effectiveness of the present invention in running in the real world is verified. The result of the extraction of the first sequence of feature points is shown in the first row. It can be clearly seen that there are no distributed feature points on the dynamic objects of the moving person's body and the carried box, and static feature points are widely distributed in the static background. The elimination of dynamic feature points on the carried box shows the effectiveness of the geometric constraint association method. The result of the extraction of the second sequence of feature points is shown in the second row. There are no distributed feature points on the bodies of the two moving people, and the static feature points on the static background are well retained. It can be seen that the present invention can effectively eliminate dynamic feature points in the real scene.

[0177] As Figure 5 shown, the estimated trajectories of the present invention in four high-dynamic environment sequences of the TUM dataset are presented. The black represents the ground truth, the blue represents the estimated trajectory, and the red represents the difference between the ground truth and the estimated trajectory, fully showing that the present invention can robustly calculate results with a small difference from the actual trajectory.

[0178] Compared with the existing visual localization and mapping technologies, the present invention can use a variety of technologies and methods to eliminate bad dynamic information in dynamic scenes, and at the same time use the GCN network model to extract features in images, obtaining more representative and trackable features, which can effectively improve the accuracy and robustness of the visual SLAM system in real high-dynamic scenes, and greatly reduce bad and unstable data associations.

[0179] A vision localization and mapping system based on geometry and semantics, comprising:

[0180] Visual odometry module: It is used to implement the operation of inputting the original image into the Mask R-CNN network model in step S1, performing per-pixel semantic segmentation on the original image to obtain instance labels and the binary mask of the image, and the segmented instance labels are used to track different objects; implementing the operation of inputting the original image into the GCN network model in step S2 to extract features in the original image to obtain all feature points on the image; implementing the operation of calculating the semantic segmentation dynamic probability of each key point through the binomial logistic regression model in step S3; implementing the operation of positioning by using the feature points in the processed image in step S4 to obtain the key frames of several frames of images in a small range; implementing the operation of eliminating dynamic feature points through the geometric correlation constraint algorithm in step S5;

[0181] Mapping module: It is used to implement the operation of inputting the static feature points in the completely processed image into the tracking and mapping thread in step S6, estimating the camera pose through frame-by-frame association, completing the mapping work, and obtaining the global map;

[0182] Nonlinear optimization module: It is used to implement the optimization work of loop detection and global Bundle Adjustment on the global map in step S7 to optimize the accuracy of map positioning.

[0183] The present invention also provides a visual positioning and mapping device based on geometry and semantics, including:

[0184] Memory: It stores the computer program of the above-mentioned visual positioning and mapping method based on geometry and semantics, and is a computer-readable device;

[0185] Processor: It is used to implement the above-mentioned visual positioning and mapping method based on geometry and semantics when executing the computer program.

[0186] The present invention also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the above-mentioned visual positioning and mapping method based on geometry and semantics.

Claims

1. A visual positioning and mapping method based on geometry and semantics, characterized in that: Including the following steps: Step S1: Input the original image into the Mask R-CNN network model, perform per-pixel semantic segmentation on the original image to obtain instance labels and a binary mask of the image. The segmented instance labels are used to track different objects, and all dynamic objects appearing in the original image scene are obtained according to the binary mask of the image; Step S2: Input the original image into the GCN network model to extract features in the original image and obtain all feature points on the image; Step S3: For the binary mask of the image obtained by the Mask R-CNN network model in Step S1 and all feature points obtained by the GCN network model in Step S2, use the binomial logistic regression model to calculate the semantic segmentation dynamic probability of each key point. Feature points with a probability less than 0.75 in the binary mask of the image are considered static feature points, and the feature points in the processed image are obtained; Step S4: After removing the dynamic objects in the original image scene through the binary mask of the image obtained by the Mask R-CNN network model in Step S1, use the feature points in the processed image obtained in Step S3 for positioning to obtain the key frame sequence of the current frame image; Step S5: Compare the angle and distance of the features of the current frame image with those of the highest overlapping key frame selected from the corresponding key frame sequence obtained in Step S4. If the feature angles of the two frame images exceed the preset threshold, the features of the current frame image are considered dynamic features. In the case where the feature angle does not exceed the threshold, it is judged whether the depth difference exceeds the preset threshold. If the depth difference exceeds a certain threshold, it is considered a dynamic feature, otherwise the feature is considered a static feature. After the comparison, the static feature points in the completely processed image are obtained; Step S6: Input the static feature points in the image completely processed in Step S5 into the tracking and mapping thread, estimate the camera pose through frame-by-frame association, complete the mapping work, and obtain the global map; Step S7: Perform loop detection and global Bundle Adjustment optimization on the global map obtained in Step S6 to optimize the accuracy of global map positioning.

2. The method for visual positioning and mapping based on geometry and semantics according to claim 1, characterized in that: The binary activation layer of the GCN network model in Step S2 is as follows: where b is the binarized version of feature f, 1 |f|≤1 canceled the gradients representing the respective feature responses of f with absolute values greater than 1, 1 |f|≤1 is a straight-through estimator of the so-called hard sign function used for backpropagating gradients.

3. A geometric and semantic-based visual positioning and mapping method according to claim 1, characterized in that: The binomial logistic regression model in Step S3 includes: semantic segmentation result label, set of feature point semantic segmentation result labels, set of pixel point geometric labels, distance between feature points and boundary pixel points, and definition of the regression model; The definition of the semantic segmentation result label is as follows: Among them, represents the feature point p i The semantic segmentation result label at time t; The definition of the set of feature point semantic segmentation result labels is as follows: where s t represents the set of feature point semantic segmentation result labels , where n is the number of feature points; The definition of the set of pixel point geometric labels is as follows: where b t is the set of boundary pixels at time t , where m is the number of boundary pixels and all boundary points are included in b t ; The definition of the distance between feature points and boundary pixel points is as follows: Among them, dist(p i , b t ) represents the distance between the feature point p i and the boundary pixel point; The definition of the regression model is as follows: Among them, represents the dynamic probability of the semantic segmentation feature point p i , where α is the influence factor for balancing the detection result curve.

4. A geometric- and semantic-based visual positioning and mapping method according to claim 1, characterized in that: In Step S5, the features of the current frame image are compared with the angles and distances of the corresponding highest overlapping key frame features, and dynamic feature points are removed through the geometric association constraint algorithm.

5. A geometric- and semantic-based visual positioning and mapping method according to claim 1, characterized in that: The process of pose estimation and global map construction through frame-by-frame association in Step S6 includes the following steps: Step S6.1: Check whether there are key frames in the buffer queue; Step S6.2: Take out the first key frame in the buffer queue of step S6.1 for processing; Step S6.3: Eliminate the bad points that appear in the map; Step S6.4: Use the epipolar geometry or triangulation method to create new map points to supplement the map from which bad points are eliminated in step S6.3; Step S6.5: Fuse the map points formed by the first key frame taken out in step S6.2 and its co-visible key frames; Step S6.6: Perform local BA optimization on the map in step S6.3; Step S6.7: Eliminate the redundant key frames in the buffer queue of step S6.1; Step S6.8: Add the current key frame taken out in step S6.2 to the loop closure detection thread.

6. A geometric and semantic-based visual localization and mapping method according to claim 1, characterized in that: The process of loop closure detection in step S7 includes the following steps: Step S7.1.1: Take out the key frame at the head of the buffer queue as the current key frame for loop closure detection; Step S7.1.2: If the time since the last loop closure detection is less than 5 seconds, no loop closure detection is performed; Step S7.1.3: Calculate the maximum similarity between the current key frame for loop closure detection in step S7.1.1 and its co-visible key frames; Step S7.1.4: Find the loop closure candidate key frame of the current key frame for loop closure detection according to the maximum similarity obtained in step S7.1.3; Step S7.1.5: Find a match between the loop closure candidate key frame group formed by the loop closure candidate key frames in step S7.1.4 and the loop closure candidate key frame group existing in the loop variable; Step S7.1.6: Maintain the loop variable so that the loop closure candidate key frame group formed by the loop closure candidate key frames in step S7.1.4 is used as the loop closure candidate key frame group before the next frame; The optimization of the global map obtained in step S7 by performing global Bundle Adjustment in step S6 means using the minimization of the reprojection error to optimize the solution of the moving points, obtaining the rotation matrix R and the translation matrix t, and using the obtained rotation matrix R and translation matrix t to optimize the poses of the points on the map; Among them, the global map is optimized by global Bundle Adjustment using the observation model: Step S7.2.1: World coordinate system to camera coordinate system: Let the world coordinates of point P be (x, y, z), and the coordinates after transformation to the camera coordinate system are: P′ = RP + t = (x′, y′, z′) Step S7.2.2: Normalization: P′ C = [u c , v c , 1] T = [x′ / z′, y′ / z′, 1] T Step S7.2.3: Distortion removal: P′ C (u c ,v c ) → (u′ c ,v′ c ) Step S7.2.4: Camera coordinate system to pixel coordinate system: u s = f x · u c ′ + c x v s = f y ·v c ′ + c y Step S7.2.5: Camera observation model: g(u s ,v s ) = h(R, t, x, y, z) Step S7.2.6: Reprojection error model: e = z - h(R, t, x, y, z) where z is the observed value and h is the value obtained from the estimation model.

7. A geometric and semantic-based visual positioning and mapping system based on the method according to any one of claims 1 to 6, characterized in that: Including: Visual Odometry Module: It is used to input the original image into the Mask R-CNN network model, perform per-pixel semantic segmentation on the original image to obtain instance labels and the binary mask of the image. The segmented instance labels are used to track different objects; input the original image into the GCN network model to extract the features in the original image and obtain all the feature points on the image; calculate the semantic segmentation dynamic probability of each key point through the binomial logistic regression model; perform localization by using the feature points in the processed image to obtain the key frame sequence of the current frame image; eliminate dynamic feature points through the geometric association constraint algorithm. Mapping Module: It is used to input the static feature points in the fully processed image into the tracking and mapping thread, estimate the camera pose through frame-by-frame association, complete the mapping work, and obtain the global map. Nonlinear Optimization Module: It is used to perform loop detection and global Bundle Adjustment optimization on the global map to optimize the accuracy of map localization.

8. A vision positioning and mapping device based on geometry and semantics, characterized in that: Comprising: Memory: It stores the computer program of the method for visual positioning and mapping based on geometry and semantics according to any one of claims 1-6, and is a computer-readable device. Processor: It is used to implement the method for visual positioning and mapping based on geometry and semantics according to any one of claims 1-6 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it can implement the method for visual positioning and mapping based on geometry and semantics according to any one of claims 1-6.

Citation Information

Patent Citations

  • Visual positioning and mapping method for indoor dynamic scene

    CN110838145A

  • Mobile robot binocular vision positioning and mapping method based on direct method

    CN114627184A

  • Visual SLAM method based on semantic segmentation dynamic points

    CN113516664A

  • Pedestrian indoor positioning and AR navigation method based on computer vision and PDR

    CN114739410A