SLAM Closed-Loop Detection and Pose Graph Optimization Method Based on Motion Constraints

By introducing closed-loop detection and pose map optimization methods based on motion constraints in the SLAM system, the problems of slow operation speed, low recall and poor robustness in the prior art are solved, and more efficient and robust real-time positioning and mapping are achieved.

CN115482252BActive Publication Date: 2025-06-13INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110599038.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-31
Publication Date
2025-06-13
Estimated Expiration
2041-05-31

AI Technical Summary

Technical Problem

The existing SLAM closed-loop detection and position map optimization technologies have slow operation speed, low recall rate and insufficient integration of kinematic knowledge, resulting in poor real-time positioning and mapping robustness.

Method used

A closed-loop detection and pose map optimization method based on motion constraints is proposed. By acquiring the historical keyframe sequence and the current frame image, and calculating the relative poses in combination with the visual-inertial odometer, the pose map is constructed. Global binary features were extracted using a pre-trained deep learning network, and closed-loop detection and pose map optimization were performed through a grid-based image feature matching algorithm.

Benefits of technology

By integrating kinematics knowledge, the operation speed, recall rate, and the robustness of SLAM closed-loop detection and position map optimization are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115482252B_ABST
    Figure CN115482252B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer vision, and particularly relates to a SLAM loop closure detection and pose graph optimization method based on motion constraints, aiming to solve the problems that the SLAM loop closure detection and pose graph optimization technologies have a slow running speed, a low recall rate, and do not fully integrate kinematic knowledge, resulting in poor robustness of SLAM. The method of the present invention includes: determining whether the current frame image is a key frame, and if so, calculating the relative poses between key frames and constructing a pose graph; taking the N historical key frames with the smallest global binary feature distance between the current frame image and each historical key frame as loop closure candidate frames; determining whether the distances between each loop closure candidate frame and the current frame image are all greater than a set distance threshold, and if not, optimizing the pose graph, otherwise extracting the local features of each loop closure candidate frame for matching and loop closure detection, and if the loop closure detection is successful, optimizing the pose graph, otherwise re-acquiring the frame image. The present invention improves the robustness of simultaneous localization and mapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a method, system, and device for SLAM loop detection and pose graph optimization based on motion constraints. Background Art

[0002] Simultaneous Localization and Mapping (SLAM) can be described as a robot constructing an environmental map while positioning itself in an unknown environment. This technology has received increasing attention due to augmented reality and autonomous driving. Loop detection is an important module of SLAM, used to correct the cumulative errors generated during its long-term operation.

[0003] Today, a large number of SLAM odometry algorithms have been proposed and achieved amazing performance. However, after long-term exploration of unknown environments by SLAM systems, trajectory prediction errors and mapping errors inevitably occur. Loop detection is a recognized solution to this problem, which can be understood as an online retrieval problem that requires real-time and robust matching of the current location with previously visited locations. Manually designed global features are calculated relatively quickly but are vulnerable to illumination and perspective changes. Manually designed local features are robust and can solve the perspective problem, but the calculation is time-consuming. Clustering techniques for local features have been proposed, and the bag-of-words model based on unsupervised training is widely used in loop detection. With the development of deep learning, convolutional neural networks have achieved amazing performance in image representation and are gradually being tried in location recognition and loop detection. However, the latest proposed CNN-based loop detection methods neither consider the real-time operation performance on mobile platforms nor fully integrate kinematic knowledge. Therefore, the present invention proposes a method for SLAM loop detection and pose graph optimization based on motion constraints. Summary of the Invention

[0004] In order to solve the above problems in the prior art, that is, to solve the problems that the existing SLAM loop detection and pose graph optimization technologies have a slow running speed, a low recall rate, and do not fully integrate kinematic knowledge, resulting in poor robustness of simultaneous localization and mapping, the present invention proposes a method for SLAM loop detection and pose graph optimization based on motion constraints, and the method includes:

[0005] S10. Obtain a historical key frame sequence and a current frame image, and determine whether the current frame image is a key frame through a preset key frame selection method. If so, calculate the relative pose between each key frame in combination with the rotation matrix and translation matrix corresponding to the visual-inertial odometer, and construct a pose graph;

[0006] S20. Extract the global binary feature of the current frame image through a pre-trained deep learning network as the first feature; calculate the distance between the first feature and the global binary features corresponding to each historical key frame, and use the top N historical key frames with the smallest distance as loop candidate frames;

[0007] In S30, it is determined whether the Hamming distances between each closed-loop candidate frame and the current frame image are all greater than a set distance threshold. If not, the closed-loop candidate frame with the smallest Hamming distance is used as the closed-loop frame, and the process jumps to step S40. Otherwise, the local features of each closed-loop candidate frame are extracted as the second features. Through an image feature matching algorithm based on grid-based motion statistics, the second features are matched with the corresponding local features of the current frame image for closed-loop detection. If the closed-loop detection is successful, the closed-loop candidate frame with the largest matching similarity is used as the closed-loop frame, and the process jumps to step S40. Otherwise, the process jumps to step S10.

[0008] In S40, the pyramid LK optical flow method is used to predict the image coordinates of the three-dimensional points observable in the closed-loop frame on the current frame image, and a 3d-2d match is established. Based on the matched 3d-2d points, the pose of the current frame image in the world coordinate system is calculated through the RANSAC algorithm and the PnP algorithm, and the generated pose graph is optimized.

[0009] In some preferred embodiments, the preset key frame selection method is as follows:

[0010] If the number of three-dimensional points observable in the current image frame is greater than N, the parallax between the current frame image and the previous key frame image is greater than M, and the time interval between the current frame image and the previous key frame image is greater than a set interval threshold, then the current frame image is a key frame; N and M are positive integers.

[0011] In some preferred embodiments, the training method of the deep learning network is as follows:

[0012] A10, collect continuous video data of unidirectional motion without closed-loop as input data;

[0013] A20, use the t-th frame image in the input data as the query image, the images in the range of [t - d, t + d] as the similar images, and the images other than the query image and the similar images as the dissimilar images;

[0014] A30, extract the global binary features of the query image, the similar images, and the dissimilar images through a pre-constructed deep learning network as the first global feature, the second global feature, and the third global feature respectively;

[0015] A40, calculate the distance between the first global feature and the second global feature as the first distance; calculate the distance between the first global feature and the third global feature as the second distance; calculate the distance between the second global feature and the third global feature as the third distance;

[0016] A50. Input the first distance, the second distance, and the third distance into a pre-constructed loss function to obtain a loss value; and in combination with the loss value, update the model parameters of the deep learning network through backpropagation.

[0017] A60. Loop and execute steps A30 - A50 until a trained deep learning network is obtained.

[0018] In some preferred embodiments, the pre-constructed loss function Loss is:

[0019]

[0020] where represents the distance between the i-th similar image p and the dissimilar image n i in the Hamming space. represents the distance between the i-th query image q i and the similar image p i in the Hamming space, and the subscript 1 represents the similarity grading parameter. represents the distance between the i-th query image q i and the dissimilar image n i in the Hamming space. represents the hash code corresponding to the continuous video data, p(.) represents the conditional probability, M represents the number of the triple {q i , p i , n i}, λ represents the set weight, N represents the length of the continuous video data, L is a positive integer, representing an L-dimensional vector.

[0021] In some preferred embodiments, for the SLAM loop closure detection and pose graph optimization method based on motion constraints, it is characterized in that "matching each second feature with the local feature corresponding to the current frame image through an image feature matching algorithm based on grid-based motion statistics and performing loop closure detection", and the method is as follows:

[0022] Calculate the similarity between each second feature and the local feature corresponding to the current frame image. If the maximum similarity is greater than the set similarity threshold, then use the loop closure candidate frame corresponding to the maximum similarity as the pending loop closure frame.

[0023] Judge whether there is a pending loop closure frame in the next frame of the current frame. If so, then use the pending loop closure frame corresponding to the current frame as the correct loop closure frame, and the loop closure detection is successful.

[0024] In some preferred embodiments, the optimization objective function corresponding to the pose graph is:

[0025]

[0026]

[0026] Among them, R i and t i respectively represent the rotation matrix and translation vector of the i-th frame relative to the world coordinate system, R ij and t ij respectively represent the relative rotation and translation between the i-th frame and the j-th frame, ε represents the set of edges in the pose graph, (i, j) represents the edge connecting the i-th frame and the j-th frame, T represents the transpose, SO(3) represents the special orthogonal group, and R 3 represents a 3D vector space.

[0027] In some preferred embodiments, the optimization solution process corresponding to the objective function of pose graph optimization is as follows:

[0028] Solve the second error term of the optimization objective function to obtain the initial rotation matrix of the i-th frame relative to the world coordinate system

[0029] For perform singular value decomposition to obtain the final rotation matrix R of the i-th frame relative to the world coordinate system i ; Take R i as the initial value of pose graph optimization, substitute it into the optimization objective function and solve to obtain the camera pose in the visual-inertial sensor after pose graph optimization.

[0030] In the second aspect of the present invention, a SLAM loop closure detection and pose graph optimization system based on motion constraints is proposed. The system includes: a pose graph construction module, a global feature matching module, a local feature matching module, and a pose graph optimization module;

[0031] The pose graph construction module is configured to obtain a historical key frame sequence and the current frame image, and determine whether the current frame image is a key frame through a preset key frame selection method. If so, calculate the relative poses between key frames by combining the rotation matrix and translation matrix corresponding to the visual-inertial odometer, and construct a pose graph;

[0032] The global feature matching module is configured to extract the global binary feature of the current frame image through a pre-trained deep learning network as the first feature; calculate the distances between the first feature and the global binary features corresponding to each historical key frame, and take the first N historical key frames with the smallest distances as loop closure candidate frames;

[0033] The local feature matching module is configured to determine whether the Hamming distances between each closed-loop candidate frame and the current frame image are all greater than a set distance threshold. If not, the closed-loop candidate frame with the smallest Hamming distance is taken as the closed-loop frame, and the pose graph optimization module is jumped to. Otherwise, the local features of each closed-loop candidate frame are extracted as the second features. The second features and the local features corresponding to the current frame image are matched through an image feature matching algorithm based on grid-based motion statistics for closed-loop detection. If the closed-loop detection is successful, the closed-loop candidate frame with the largest matching similarity is taken as the closed-loop frame, and the pose graph optimization module is jumped to. Otherwise, the pose graph construction module is jumped to;

[0034] The pose graph optimization module is configured to predict the image coordinates of the three-dimensional points observable by the closed-loop frame on the current frame image by using the pyramid LK optical flow method and establish 3d-2d matches. Based on the matched 3d-2d points, the pose of the current frame image in the world coordinate system is calculated through the RANSAC algorithm and the PnP algorithm, and the generated pose graph is optimized.

[0035] In a third aspect of the present invention, a device is proposed, including at least one processor; and a memory communicatively connected to at least one of the processors; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned SLAM closed-loop detection and pose graph optimization method based on motion constraints.

[0036] In a fourth aspect of the present invention, a computer-readable storage medium is proposed, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned SLAM closed-loop detection and pose graph optimization method based on motion constraints.

[0037] Advantages of the present invention:

[0038] By integrating kinematic knowledge, the present invention improves the running speed, recall rate, and the robustness of simultaneous localization and mapping of the existing SLAM closed-loop detection and pose graph optimization technologies.

[0039] 1) In the training stage of the present invention: The t-th frame image in the continuous video data is used as the query image, the [t - d, t + d]-th frame images are used as the similar images, and the images other than the query image and the similar images are used as the dissimilar images. The global binary features of the query image, the similar images, and the dissimilar images are extracted and the feature distances are calculated to train the deep learning network, improving the accuracy of feature extraction of the network;

[0040] 2) Detection stage: Calculate the Hamming distance between the current key frame and historical key frames, and take the top N historical key frames with the smallest distance as the loop closure candidate frames; Flexible and efficient retrieval of loop closure frames is achieved according to the Hamming distance between each loop closure candidate frame and the current frame image and the inlier rate corresponding to the local feature matching of each loop closure candidate frame.

[0041] 3) Optimization stage: First, optimize, correct, and perform singular value decomposition on the relative rotation and translation between the i-th frame and the j-th frame to obtain the R i as the initial value of pose graph optimization; Based on the initial value of pose graph optimization, solve the objective function of pose graph optimization to achieve fast and accurate optimization of poses, improving the running speed, recall rate, and robustness of instant localization and mapping of existing SLAM loop closure detection and pose graph optimization technologies. Brief Description of the Drawings

[0042] Other features, purposes, and advantages of this application will become more obvious by reading the detailed description of the non-restrictive embodiments with reference to the following drawings.

[0043] Figure 1 is a schematic flowchart of a method for SLAM loop closure detection and pose graph optimization based on motion constraints according to an embodiment of the present invention;

[0044] Figure 2 is a schematic framework diagram of a system for SLAM loop closure detection and pose graph optimization based on motion constraints according to an embodiment of the present invention;

[0045] Figure 3 is a schematic block diagram of a simplified system for SLAM loop closure detection and pose graph optimization based on motion constraints according to an embodiment of the present invention;

[0046] Figure 4 is a schematic diagram of the principle of an image feature matching algorithm based on grid-based motion statistics according to an embodiment of the present invention;

[0047] Figure 5 is a schematic diagram of a motion state according to an embodiment of the present invention;

[0048] Figure 6 is a schematic diagram of the visualization of system loop closure detection according to an embodiment of the present invention;

[0049] Figure 7 is a schematic diagram of the influence of the Hamming distance threshold on system accuracy and time according to an embodiment of the present invention;

[0050] Figure 8 is a schematic diagram of the trajectory comparison after loop closure detection and pose optimization according to an embodiment of the present invention;

[0051] Figure 9It is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application according to an embodiment of the present invention. Detailed implementation manners

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the related invention and are not intended to limit the invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0054] A method for SLAM closed-loop detection and pose graph optimization based on motion constraints according to the first embodiment of the present invention, as Figure 1 shown, the method includes the following steps:

[0055] S10. Obtain a historical key frame sequence and a current frame image, and determine whether the current frame image is a key frame through a preset key frame selection method. If so, calculate the relative poses between the key frames by combining the rotation matrix and translation matrix corresponding to the visual-inertial odometer, and construct a pose graph;

[0056] S20. Extract the global binary feature of the current frame image through a pre-trained deep learning network as the first feature; calculate the distances between the first feature and the global binary features corresponding to each historical key frame, and use the first N historical key frames with the smallest distances as the closed-loop candidate frames;

[0057] S30. Determine whether the Hamming distances between each closed-loop candidate frame and the current frame image are all greater than a set distance threshold. If not, use the closed-loop candidate frame with the smallest Hamming distance as the closed-loop frame and jump to step S40. Otherwise, extract the local features of each closed-loop candidate frame as the second feature; perform image feature matching on the second features and the local features corresponding to the current frame image through an image feature matching algorithm based on grid-based motion statistics and perform closed-loop detection. If the closed-loop detection is successful, use the closed-loop candidate frame with the largest matching similarity as the closed-loop frame and jump to step S40. Otherwise, jump to step S10;

[0058] S40. Use the pyramid LK optical flow method to predict the image coordinates of the three-dimensional points observable in the closed-loop frame on the current frame image, and establish 3D-2D matching; based on the matched 3D-2D points, calculate the pose of the current frame image in the world coordinate system through the RANSAC algorithm and the PnP algorithm, and optimize the generated pose graph.

[0059] To more clearly illustrate the SLAM closed-loop detection and pose graph optimization method based on motion constraints of the present invention, the following details each step in an embodiment of the method of the present invention.

[0060] The present invention is mainly divided into three parts: self-supervised training of a deep learning network, closed-loop retrieval integrating global and local features, and pose graph optimization. The block diagram is as Figure 3 shown. In the following embodiments, first, the training process of the deep learning network is described in detail, and then the optimization process of the pose graph by the SLAM closed-loop detection and pose graph optimization method based on motion constraints is described in detail.

[0061] 1. Training process of the deep learning network

[0062] A10. Collect continuous video data with one-way motion and no closed-loop as input data;

[0063] In this embodiment, first collect continuous images in a scene, which conform to one-way motion and no closed-loop, and construct a continuous motion model.

[0064] A20. Use the t-th frame image in the input data as the query image, the images in the [t - d, t + d] frames as the similar images, and the images other than the query image and the similar images as the dissimilar images;

[0065] The continuous image sequence is where t is the timestamp when the image is collected. There is a time period [t - d, t + d], and the image x t is similar to each frame image in this time period, and the similarity is inversely proportional to the interval timestamp.

[0066] In this embodiment, use the t-th frame image in the continuous video data as the query image, the images in the [t - d, t + d] frames as the similar images, and the images other than the query image and the similar images as the dissimilar images.

[0067] A30. Extract the global two-dimensional features of the query image, the similar images, and the dissimilar images through a pre-constructed deep learning network, and use them as the first global feature, the second global feature, and the third global feature respectively;

[0068] In this embodiment, extract the query image, the similar images, and the dissimilar images as global binary features through the deep learning network.

[0069] A40. Calculate the distance between the first global feature and the second global feature as the first distance; calculate the distance between the first global feature and the third global feature as the second distance; calculate the distance between the second global feature and the third global feature as the third distance.

[0070] A50. Input the first distance, the second distance, and the third distance into a pre-constructed loss function to obtain a loss value; and in combination with the loss value, update the model parameters of the deep learning network through backpropagation.

[0071] In this embodiment, the pre-constructed loss function is:

[0072]

[0073] where represents the distance between the i-th similar image p i and the dissimilar image n i in the Hamming space, represents the distance between the i-th query image q i and the similar image p i in the Hamming space, and the subscript 1 represents the similarity grading parameter, represents the distance between the i-th query image q i and the dissimilar image n i in the Hamming space, represents the hash code corresponding to the continuous video data, b t is equivalent to b i , p(.) represents the conditional probability, M represents the number of the triplet {q i , p i , n i}, λ represents the set weight, N represents the length of the continuous video data, and L is a positive integer representing an L-dimensional vector.

[0074] Based on the above pre-constructed loss function, calculate the loss value and update the model parameters of the deep learning network through backpropagation.

[0075] A60. Loop and execute steps A30 - A50 until a trained deep learning network is obtained.

[0076] In this embodiment, let the hash code corresponding to the continuous image sequence X be where b t ={0, 1} L represents the hash code of length L corresponding to the image x t . For a pair of hash codes b i and b j , the Hamming distance can be expressed as:

[0077]

[0078] The problem of measuring the image similarity in the Hamming space is transformed into the problem that is consistent with (the similarity between q i and p i ).

[0079] Define the Hamming distance between the hash codes of two frames of images i and j as:

[0080] dist i,j = 2θ ij β (3)

[0081] where θ ij represents the similarity degree (i.e., similarity) between two frames of images, and β is a constant that can control the Hamming distance of the hash codes corresponding to a pair of similar images. Different from the traditional triplet loss function, the proposed method can adjust the Hamming distance between two similar images through the similarity. The designed loss function is based on probability. According to a triplet and a similarity label, the maximum a posteriori estimation p(B|T,Θ) of the hash code can be expressed as:

[0082]

[0083] where B represents the hash code, the triplet and the similarity label represent the similarity between q i and p i , and the conditional probability p(t i ,θ i |b i ) is defined as follows:

[0084]

[0085]

[0086]

[0087]

[0088]

[0089] where d q,p represents the distance between the binary codes q and p in the Hamming space, and σ(x) is the Sigmoid activation function. The above last two formulas control d q,p to achieve the similarity grading of similar images. According to the maximum likelihood estimation, the loss function for learning the hash code proposed by us is as follows:

[0090]

[0091] Because the Sigmoid activation function is used in the last fully connected layer, the value of the tensor b output by the model for the image is in the range of [0, 1]. The output is made to approach 0 or 1 by the constraint of maximizing the sum of squared errors between the output tensor b and 0.5.

[0092] 2. SLAM Closed-loop Detection and Pose Graph Optimization Method Based on Motion Constraints

[0093] S10. Obtain a sequence of historical key frames and the current frame image, and determine whether the current frame image is a key frame through a preset key frame selection method. If it is, then combine the rotation matrix and translation matrix corresponding to the visual-inertial odometer to calculate the relative poses between key frames, and construct a pose graph.

[0094] In this embodiment, during the pose graph optimization process, the visual-inertial odometer system is selected as the front end to obtain the pose of each frame image and perform pose graph optimization simultaneously. To reduce the computational load, closed-loop detection and pose graph optimization are performed on key frames. Therefore, a key frame selection mechanism needs to be added to the visual-inertial odometer.

[0095] The key frame selection method is as follows: If the number of three-dimensional points observable in the current image frame is greater than N, the disparity between the current frame image and the previous key frame image is greater than M, and the time interval between the current frame image and the previous key frame image is greater than the set interval threshold, then the current frame image is a key frame; N and M are positive integers. Among them, N is preferably set to 3.

[0096] If the current frame image is a key frame, then combine the rotation matrix R and translation matrix t corresponding to the visual-inertial odometer to calculate the relative poses between key frames, and construct a pose graph (specifically, construct the edges of the pose graph).

[0097] S20. Extract the global binary feature of the current frame image through a pre-trained deep learning network as the first feature; calculate the distances between the first feature and the global binary features corresponding to each historical key frame, and take the first N historical key frames with the smallest distances as the closed-loop candidate frames.

[0098] When a system running SLAM reaches a closed-loop point, then the system will be in the closed-loop state for a period of time next, as Figure 5 shown. Therefore, the motion states can be distinguished as the non-closed-loop state and the closed-loop state. Assume that the current query frame (i.e., the current frame image or the current key frame) Q 0 is similar to the historical closed-loop frame R 0 Then there exists a period of time of length t during which the upcoming image frame Q t is similar to the historical closed-loop frame R tSimilar. By distinguishing the motion state, fusing global and local features, and using a linear storage structure, a flexible and efficient retrieval strategy is designed.

[0099] For a pair of images from different viewpoints of the same 3D scene, a feature correspondence means that a feature in one image can be projected through a 3D point to a feature in the other image. Assuming a smooth motion process, neighboring features move together. True correspondences are constrained by smoothness, while false correspondences are not. Therefore, true correspondences have more similar neighbors. Divide image I 1 and image I 2 into non-overlapping grids respectively. Assume c i is a correspondence in grids G a and G b . We define the similar neighbors of c i as:

[0100] S i ={c j |c j ∈C ab ,c i ≠c j} (11)

[0101] Here, C ab is the set of correspondences that fall within grids G a and G b . We let |S i |, which is the count of S i , represent the motion support of c i . This motion support is used to distinguish correct and incorrect correspondences. As Figure 4 shown, grid G a has a motion support of |S b | = 2 in grid G i .

[0102] In the non-closed-loop state, the system's requirements for the closed-loop frame are reliable. For each acquired image frame, a deep learning network is used to extract global binary features. This global binary feature is the hash code corresponding to the current image frame Q 0 . It is added to the end of the linear storage structure and a brute-force search is performed in the Hamming space. Specifically, the Hamming distances between the current image frame Q 0 and all historical key frames are calculated, and N image frames with the smallest Hamming distances and less than the threshold δ 1 are selected as closed-loop candidate frames.

[0103] S30. Determine whether the Hamming distances between each closed-loop candidate frame and the current frame image are all greater than a set distance threshold. If not, take the closed-loop candidate frame with the minimum Hamming distance as the closed-loop frame and jump to step S40. Otherwise, extract the local features of each closed-loop candidate frame as the second features. Match the second features with the corresponding local features of the current frame image through an image feature matching algorithm based on grid-based motion statistics and perform closed-loop detection. If the closed-loop detection is successful, take the closed-loop candidate frame with the maximum matching similarity as the closed-loop frame and jump to step S40. Otherwise, jump to step S10;

[0104] In this embodiment, first determine whether the distances between each closed-loop candidate frame and the current frame image are all greater than a set distance threshold. If not, the closed-loop detection is successful. Take the closed-loop candidate frame with the minimum Hamming distance as the closed-loop frame and jump to step S40 to perform pose graph optimization. Otherwise, extract the local features of the closed-loop candidate frames returned in the above process for geometric consistency checking based on grid-based motion statistics (i.e., calculate the similarity between each second feature and the corresponding local feature of the current frame image. If the maximum similarity is greater than the set similarity threshold, take the closed-loop candidate frame corresponding to the maximum similarity as the pending closed-loop frame), as Figure 4 shown. The inlier rate (i.e., similarity) in the matched image pair is the largest and greater than the threshold γ 1 of the closed-loop candidate frame R 0 will continue to perform temporal consistency checking, that is, perform geometric consistency checking again between consecutive frame images Q 1 and R 1 . The assumption of temporal consistency is ideal because the displacement difference between the current two consecutive frames and the displacement difference between their candidate two frames are inconsistent. This means that the similarity between Q 1 and R 1 may be lower than the similarity between Q 0 and R 0 . Therefore, the threshold of the inlier rate of grid-based motion statistics in the temporal consistency link should be less than γ 1 , which is specifically set to γ 2 . If R 0 and R 1 pass through the above process, they will be accepted together as the final closed-loop frame, and the system will also enter the closed-loop state (i.e., determine whether there is a pending closed-loop frame for the next frame of the current frame. If so, take the pending closed-loop frame corresponding to the current frame as the correct closed-loop frame, and the closed-loop detection is successful).

[0105] Within a subsequent period n, the current image frame Q n preferably calculates the Hamming distance with the historical key frame R n . If this Hamming distance is less than the threshold δ 2, the frame is accepted as a closed loop; otherwise, the system performs grid-based motion statistics, and the inlier rate threshold is set to γ 3 . All parameter relationships are summarized as follows:

[0106] 0 < δ 1 < δ 2 < dist(HashCode) (12)

[0107] 0 < γ 3 < γ 2 < γ 1 < 1 (13)

[0108] Among them, dist(HashCode) represents the Hamming distance between the hash codes corresponding to two frames of images.

[0109] If both the global and local feature matching methods fail, the closed-loop state is exited. A very small number of ambiguous positive results appear at the end of the closed-loop sequence. Figure 6 It shows that the proposed invention can continuously retrieve closed loops and can adapt to difficult scenarios such as occlusion. Among them, local feature matching is visualized as color-corresponding points, and global features are visualized by gradient class activation map technology. Figure 7 It is the change curve of the recall rate and average execution time when adjusting the Hamming distance threshold at 100% accuracy. It can be seen that flexible threshold setting can reduce the execution time while increasing the recall rate.

[0110] S40, use the pyramid LK optical flow method to predict the image coordinates of the three-dimensional points observable in the closed-loop frame on the current frame image, and establish 3d-2d matching; based on the matched 3d-2d points, calculate the pose of the current frame image in the world coordinate system through the RANSAC algorithm and the PnP algorithm, and optimize the generated pose graph;

[0111] In this embodiment, the pyramid LK optical flow method is used to predict the image coordinates of the three-dimensional points observable in the closed-loop frame on the current key frame, so as to establish 3d-2d matching. Based on the matched 3d-2d points, the RANSAC+PnP algorithm is used to calculate the pose of the current key frame in the world coordinate system. According to the calculated current key frame pose graph, an edge between the current key frame and the closed-loop frame is established in the pose graph. Optimize the generated pose graph to suppress error drift. The objective function of pose graph optimization is as follows:

[0112]

[0113] Among them, R i and t i respectively represent the rotation matrix and translation vector of the i-th frame relative to the world coordinate system, R ij and t ijrespectively represent the relative rotation and translation between the i-th frame and the j-th frame, ε represents the set of edges in the pose graph, (i, j) represents the edge connecting the i-th frame and the j-th frame, T represents the transpose, SO(3) represents the special orthogonal group, and R 3 represents a 3D vector space, represents the 2-F norm.

[0114] During the pose graph optimization process, the rotation matrix R dominates. Therefore, the second error term in the above equation can be considered first to obtain the following objective function:

[0115]

[0116] This objective function is a linear least squares problem and is very easy to solve. The obtained is very likely not a rotation matrix and needs to be corrected. Perform a singular value decomposition on Finally, take R i = Sdiag[1 1 det(SV T )]V T , det() represents the determinant of a matrix, where S is an m×m matrix, D is an m×n matrix with all elements outside the main diagonal being 0, each element on the main diagonal is called a singular value, V is an n×n matrix, and both S and V are unitary matrices satisfying S T S = I, V T V = I.

[0117] Finally, take the obtained R i as the initial value of the pose graph optimization, and solve the objective function of the pose graph optimization to obtain the camera pose R i and t i .

[0118] In addition, to verify the effectiveness of the method of the present invention, experiments are carried out on various publicly available datasets in China. The experimental results are shown in Table 1, that is, the recall rate at 100% accuracy.

[0119] Table 1

[0120]

[0121] ​Table 2 shows the average execution time test of loop detection for each dataset on the CPU and GPU, including the average execution time of each important link. Among them, global feature extraction and hash code conversion can be performed on the CPU or GPU. The time of TopN brute-force search will increase with the increase of the database, but it has little impact on the global average time test of the system. The average time of local feature extraction, matching, and geometric consistency verification composed of GMS accounts for a relatively small proportion in the global average time of the system. Obviously, the time performance of the proposed loop detection method meets the indicators required by the project.

[0122] Table 2

[0123]

[0124]

[0125] Among them, query represents the query frame.

[0126] Test the loop closure detection and pose graph optimization algorithms in the self-collected scenario. The time consumption and reprojection error of the pose graph optimization algorithm are shown in Table 3 and Figure 8 as follows. The reprojection error refers to predicting the position of the 3D points of the loop closure key frame on the current key frame using the pyramid LK optical flow method to determine the 3D-2D matching, then obtaining the pose of the current key frame according to the PnP algorithm, and at the same time calculating the reprojection error of the 3D points of the loop closure key frame on the current key frame.

[0127] Table 3

[0128]

[0129] Among them, keyframe represents the key frame and pixel represents the pixel.

[0130] A SLAM loop closure detection and pose graph optimization system based on motion constraints according to the second embodiment of the present invention, as Figure 2 shown, includes: a pose graph construction module 100, a global feature matching module 200, a local feature matching module 300, and a pose graph optimization module 400;

[0131] The pose graph construction module 100 is configured to obtain a historical key frame sequence and the current frame image, and determine whether the current frame image is a key frame through a preset key frame selection method. If so, calculate the relative pose between each key frame by combining the rotation matrix and translation matrix corresponding to the visual-inertial odometer, and construct a pose graph;

[0132] The global feature matching module 200 is configured to extract the global binary features of the current frame image through a pre-trained deep learning network as the first feature; calculate the distances between the first feature and the global binary features corresponding to each historical key frame, and use the top N historical key frames with the smallest distances as the closed-loop candidate frames;

[0133] The local feature matching module 300 is configured to determine whether the Hamming distances between each closed-loop candidate frame and the current frame image are all greater than a set distance threshold. If not, it uses the closed-loop candidate frame with the smallest Hamming distance as the closed-loop frame and jumps to the pose graph optimization module 400. Otherwise, it extracts the local features of each closed-loop candidate frame as the second feature; performs matching between each second feature and the local features corresponding to the current frame image through an image feature matching algorithm based on grid-based motion statistics for closed-loop detection. If the closed-loop detection is successful, it uses the closed-loop candidate frame with the largest matching similarity as the closed-loop frame and jumps to the pose graph optimization module 400. Otherwise, it jumps to the pose graph construction module 100;

[0134] The pose graph optimization module 400 is configured to use the pyramid LK optical flow method to predict the image coordinates of the three-dimensional points observable in the closed-loop frame on the current frame image and establish 3d-2d matches; based on the matched 3d-2d points, calculate the pose of the current frame image in the world coordinate system through the RANSAC algorithm and the PnP algorithm to optimize the generated pose graph.

[0135] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and related descriptions of the above-described system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0136] It should be noted that the SLAM closed-loop detection and pose graph optimization system based on motion constraints provided in the above embodiments is only illustrated by dividing the above-mentioned functional modules. In practical applications, the above functions can be allocated to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only used to distinguish each module or step and are not regarded as an improper limitation of the present invention.

[0137] An apparatus according to a third embodiment of the present invention includes at least one processor; and a memory communicatively connected to at least one of the processors; wherein the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned SLAM closed-loop detection and pose graph optimization method based on motion constraints.

[0138] A computer-readable storage medium according to a fourth embodiment of the present invention, wherein the computer-readable storage medium stores computer instructions for being executed by a computer to implement the above-described SLAM closed-loop detection and pose graph optimization method based on motion constraints.

[0139] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and related descriptions of the above-described storage device and processing device can refer to the corresponding processes in the foregoing method examples and will not be elaborated herein.

[0140] Next, refer to Figure 9 , which shows a schematic structural diagram of a computer system of a server suitable for implementing the method, system, and device embodiments of the present application. Figure 9 The shown server is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0141] As Figure 9 shown, the computer system includes a central processing unit (CPU, Central Processing Unit) 901, which can execute various appropriate actions and processes according to a program stored in a read-only memory (ROM, Read Only Memory) 902 or a program loaded from a storage section 908 into a random access memory (RAM, Random Access Memory) 903. In the RAM 903, various programs and data required for system operations are also stored. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O, Input / Output) interface 905 is also connected to the bus 904.

[0142] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including, for example, a cathode ray tube (CRT, Cathode Ray Tube), a liquid crystal display (LCD, Liquid Crystal Display), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage section 908 as needed.

[0143] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 909 and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above functions defined in the method of the present application are performed. It should be noted that the computer-readable medium in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0144] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).

[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0146] The terms "first", "second", etc. are used to distinguish similar objects and not to describe or represent a specific order or sequence.

[0147] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to those processes, methods, articles, or apparatus / device.

[0148] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easily understood by those skilled in the art that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A method for SLAM loop closure detection and pose graph optimization based on motion constraints, characterized in that, the method includes: S10. Obtain a sequence of historical key frames and the current frame image, and determine whether the current frame image is a key frame through a preset key frame selection method. If so, calculate the relative poses between key frames by combining the rotation matrix and translation matrix corresponding to the visual-inertial odometer, and construct a pose graph; S20. Extract the global binary feature of the current frame image through a pre-trained deep learning network as the first feature; calculate the distances between the first feature and the global binary features corresponding to each historical key frame, and use the first N historical key frames with the smallest distances as loop closure candidate frames; S30. Determine whether the Hamming distances between each loop closure candidate frame and the current frame image are all greater than a set distance threshold. If not, use the loop closure candidate frame with the smallest Hamming distance as the loop closure frame and jump to step S40. Otherwise, extract the local features of each loop closure candidate frame as the second feature; perform matching on the second features and the local features corresponding to the current frame image through an image feature matching algorithm based on grid-based motion statistics for loop closure detection. If the loop closure detection is successful, use the loop closure candidate frame with the largest matching similarity as the loop closure frame and jump to step S40. Otherwise, jump to step S10; S40. Use the pyramid LK optical flow method to predict the image coordinates of the three-dimensional points observable by the loop closure frame on the current frame image and establish 3d-2d matching; based on the matched 3d-2d points, calculate the pose of the current frame image in the world coordinate system through the RANSAC algorithm and the PnP algorithm, and optimize the generated pose graph; The optimization objective function corresponding to the pose graph is: where, R i and t i represent the rotation matrix and the translation vector of the i-th frame with respect to the world coordinate system respectively, R ij and t ij represent the relative rotation and translation between the i-th frame and the j-th frame respectively, ε represents the set of edges in the pose graph, (i, j) represents the edge connecting the i-th frame and the j-th frame, T represents the transpose, SO(3) represents the special orthogonal group, and R 3 represents a 3-dimensional vector space.

2. The method for SLAM loop closure detection and pose graph optimization based on motion constraints according to claim 1, characterized in that, the preset key frame selection method is: If the number of three-dimensional points observable in the current image frame is greater than N, the parallax between the current frame image and the previous key frame image is greater than M, and the time interval between the current frame image and the previous key frame image is greater than the set interval threshold, then the current frame image is a key frame; N and M are positive integers.

3. The method for SLAM loop closure detection and pose graph optimization based on motion constraints according to claim 2, characterized in that, the training method of the deep learning network is: A10. Collect continuous video data with one-way motion and no loop closure as input data; A20. Use the t-th frame image in the input data as the query image, use the [t - d, t + d]-th frame images as the similar images, and use the images other than the query image and the similar images as the dissimilar images; A30. Extract the global binary features of the query image, the similar images, and the dissimilar images through a pre-constructed deep learning network as the first global feature, the second global feature, and the third global feature respectively; A40. Calculate the distance between the first global feature and the second global feature as the first distance; calculate the distance between the first global feature and the third global feature as the second distance; calculate the distance between the second global feature and the third global feature as the third distance; A50. Input the first distance, the second distance, and the third distance into a pre-constructed loss function to obtain a loss value; and in combination with the loss value, update the model parameters of the deep learning network through backpropagation. A60. Loop through steps A30 - A50 until a trained deep learning network is obtained.

4. The method for SLAM loop closure detection and pose graph optimization based on motion constraints according to claim 3, wherein, the pre-constructed loss function Loss is: Among them, represents the distance in the Hamming space between the i-th similar image p i and the dissimilar image n i ; represents the distance in the Hamming space between the i-th query image q i and the similar image p i , where the subscript 1 represents the similarity grading parameter; represents the distance in the Hamming space between the i-th query image q i and the dissimilar image n i ; represents the hash code corresponding to the continuous video data, p(.) represents the conditional probability, M represents the number of the triple {q i , p i , n i}, λ represents the set weight, N represents the length of the continuous video data, and L is a positive integer representing the L-dimensional vector.​​​ 5. The method for SLAM loop closure detection and pose graph optimization based on motion constraints according to claim 1, wherein, "Match each second feature with the local feature corresponding to the current frame image through an image feature matching algorithm based on grid-based motion statistics and perform loop closure detection", and the method is: Calculate the similarity between each second feature and the local feature corresponding to the current frame image. If the maximum similarity is greater than a set similarity threshold, then use the loop closure candidate frame corresponding to the maximum similarity as the pending loop closure frame. Judge whether there is a pending loop closure frame in the next frame of the current frame. If so, then use the pending loop closure frame corresponding to the current frame as the correct loop closure frame, and the loop closure detection is successful.

6. The method for SLAM loop closure detection and pose graph optimization based on motion constraints according to claim 4, wherein, the optimization solution process corresponding to the objective function of pose graph optimization is: The second error term of the optimized objective function is solved to obtain the initial rotation matrix of the i-th frame relative to the world coordinate system For perform singular value decomposition to obtain the final rotation matrix R of the i-th frame relative to the world coordinate system i ; use R i as the initial value for pose graph optimization, substitute it into the optimization objective function and solve to obtain the camera pose in the visual-inertial sensor after pose graph optimization.

7. A system for SLAM loop closure detection and pose graph optimization based on motion constraints, wherein, the system includes: a pose graph construction module, a global feature matching module, a local feature matching module, and a pose graph optimization module; The pose graph construction module is configured to obtain a historical key frame sequence and the current frame image, and determine whether the current frame image is a key frame through a preset key frame selection method. If so, then in combination with the rotation matrix and translation matrix corresponding to the visual-inertial odometer, calculate the relative poses between key frames and construct a pose graph. The global feature matching module is configured to extract the global binary feature of the current frame image through a pre-trained deep learning network as the first feature; calculate the distances between the first feature and the global binary features corresponding to each historical key frame, and use the top N historical key frames with the smallest distances as loop closure candidate frames. The local feature matching module is configured to judge whether the Hamming distances between each loop closure candidate frame and the current frame image are all greater than a set distance threshold. If not, then use the loop closure candidate frame with the smallest Hamming distance as the loop closure frame and jump to the pose graph optimization module. Otherwise, extract the local features of each loop closure candidate frame as the second feature; match each second feature with the local feature corresponding to the current frame image through an image feature matching algorithm based on grid-based motion statistics and perform loop closure detection. If the loop closure detection is successful, then use the loop closure candidate frame with the largest matching similarity as the loop closure frame and jump to the pose graph optimization module. Otherwise, jump to the pose graph construction module. The pose graph optimization module is configured to predict the image coordinates of the three-dimensional points observable in the closed-loop frame on the current frame image using the pyramid LK optical flow method and establish 3d-2d matches; based on the matched 3d-2d points, calculate the pose of the current frame image in the world coordinate system through the RANSAC algorithm and the PnP algorithm, and optimize the generated pose graph. The optimization objective function corresponding to the pose graph is: where, R i and t i represent the rotation matrix and translation vector of the i-th frame relative to the world coordinate system, R ij and t ij represent the relative rotation and translation between the i-th frame and the j-th frame, ε represents the set of edges in the pose graph, (i, j) represents the edge connecting the i-th frame and the j-th frame, T represents the transpose, SO(3) represents the special orthogonal group, and R 3 represents a 3-dimensional vector space.

8. An electronic device characterized in that it includes: at least one processor; and a memory communicatively connected to at least one of the processors; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the motion constraint-based SLAM loop closure detection and pose graph optimization method according to any one of claims 1-6.

9. A computer-readable storage medium characterized in that the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the motion constraint-based SLAM loop closure detection and pose graph optimization method according to any one of claims 1-6.