Bag-of-word model loopback detection method based on depth image

Through the bag-of-word model loopback detection method based on depth images, the cumulative error problem of laser vision fusion SLAM system during long run is solved, and the accuracy of loopback detection and the stability of the navigation system are improved.

CN120355756APending Publication Date: 2025-07-22NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510201404.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing laser vision fusion SLAM system has a large cumulative error during long runs, and the point cloud loopback method based on context scanning is less efficient, while the visual bag-of-word model loopback method has limited accuracy and visual blind spots, making it difficult to accurately identify loopback.

Method used

The bag-to-loop detection method based on depth images is adopted. Through the registration of laser point clouds and visual images, the straight edges of the object contour are extracted, depth projection and completion are performed, and the complete depth image is generated by combining local linear fitting, and the bag-to-loop detection is used.

Benefits of technology

The accuracy of loopback detection is improved, the system's cumulative error is corrected, and the stability and positioning accuracy of the navigation system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355756A_ABST
    Figure CN120355756A_ABST
Patent Text Reader

Abstract

The invention discloses a bag-of-word model loopback detection method based on a depth image, and the method specifically comprises the following steps: 1, carrying out the registration of a laser point cloud and a visual image, and calibrating the external parameters of a laser radar and a camera; 2, extracting linear edges of all object contours in the field of view, and converting all point cloud coordinates to a camera coordinate system by aligning edge features in a point cloud image of the laser radar and a visible light image of the camera to realize depth projection to obtain an initial depth image; 3, complementing the initial depth image; generating a complemented complete depth image; and step 4, for the complemented complete depth image, detecting a loop by using a bag-of-word model loop detection mechanism. The loopback detection precision is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision. In particular, it relates to a loop detection method based on a bag-of-words model of depth images. Background Art

[0002] The wide application of intelligent robots poses high requirements for their environmental perception capabilities, that is, to have functions such as accurately recognizing the surrounding environment, simultaneous localization and mapping (SLAM), and visual navigation, etc., to ensure that the robot can complete upstream tasks. Laser-vision fusion SLAM can provide rich environmental information and has strong robustness, but the system will still accumulate large errors during long-term operation, resulting in positioning failure.

[0003] Adding a loop detection module can effectively reduce the accumulated error. For the point cloud loop method based on context scanning, the information is accurate, but the efficiency is low; for the loop method based on the visual bag-of-words model, the detection efficiency is high, but its accuracy is limited, and there is a problem of visual blind spots. For example, it may be difficult to accurately recognize loops at the same intersection in different traveling directions. Summary of the Invention

[0004] Object of the Invention: In order to solve the problems existing in the above-mentioned prior art, the present invention provides a loop detection method based on a bag-of-words model of depth images.

[0005] Technical Solution: The present invention provides a loop detection method based on a bag-of-words model of depth images, specifically as follows:

[0006] Step 1: Register the laser point cloud and the visual image, and calibrate the external parameters of the lidar and the camera;

[0007] Step 2: Extract the straight edges of all object contours in the field of view. By aligning the edge features in the point cloud map of the lidar and the visible light image of the camera, all point cloud coordinates are transformed into the camera coordinate system to achieve depth projection and obtain an initial depth image;

[0008] Step 3: Complete the initial depth image; generate a complete depth image after completion;

[0009] Step 4: For the complete depth image after completion, adopt a bag-of-words model loop detection mechanism to detect loops.

[0010] Further, the depth projection in Step 2 is specifically as follows:

[0011] Step 2.1: Combine the internal parameters and distortion parameters of the camera, project the points in the camera coordinate system into the pixel coordinate system, and correct the distortion of the camera projection model:

[0012] C p i =π(C P i )

[0013] p i = f( C p i )

[0014] where C P i represents the i-th point cloud in the camera coordinate system, C p i represents the i-th point cloud in the pixel coordinate system, C p i represents the i-th point cloud in the camera coordinate system, π(·) represents the projection model of the camera, and f(·) represents the distortion model of the camera; p i represents the i-th point cloud after distortion correction in the pixel coordinate system;

[0015] The error of the optimized reprojection model is used to obtain the Jacobian matrix J from the lidar coordinate system to the image coordinate system:

[0016] J = [F, E]

[0017] where F is the partial derivative matrix of the reprojection model cost function with respect to the pose, the transpose of the expression unit vector, p represents the point cloud after distortion correction, P represents the original point cloud, P ^ represents the skew-symmetric matrix of P, I represents the identity matrix, and E is the partial derivative matrix of the reprojection model cost function with respect to the spatial point, represents the rotation matrix from the radar coordinate system to the camera coordinate system;

[0018] The initial depth image of the point cloud projection in the pixel coordinate system is obtained by using the depth projection function O(·).

[0019] Furthermore, it specifically includes the following steps: Step 3 uses the local linear fitting method to complete the initial depth image, and the objective function is specifically:

[0020]

[0021] where N is the total number of pixels in the image, represents the neighborhood of pixel j, i is the pixel point in, d H,i is the estimated depth value of pixel i, is the linear fitting value of pixel i in N(j), w i,j is the weight of the similarity measure between pixel i and j, D jTo measure the consistency value of the depth data of the completed image and the uncompleted image, E j is the value for measuring the gradient of the pixel;

[0022] The expression of

[0023]

[0024] is as follows: j where α j and β j are linear fitting parameters. α j is a 2×1 vector, and β i,j is a scalar. g j represents the relative coordinate between the coordinates of pixel i minus the coordinates of pixel j;

[0025] The expressions of D j and E

[0026] are as follows: j D L,j = d H,j - d j E L,j = e H,j

[0027] where d L,j is the depth value of pixel j in the initial depth image, d H,j is the estimated depth value of pixel j in the completed depth image, e L,j is the edge value of pixel j in the initial depth image, and e H,j is the estimated edge value of pixel j in the completed depth image;

[0028] The expression of w i,j is as follows:

[0029]

[0030] where s i represents the RGB pixel value of the pixel in the initial depth image, and s j represents the RGB pixel value of pixel j in the fitted image; both σ1 and σ2 are pre-reviewed pixel similarity values, and respectively represent the depth information estimated by the median filter of pixel i and pixel j based on the initial depth image.

[0031] Furthermore, step 4 is specifically as follows:

[0032] Step 4.1: Extract the words in the current key frame image and all historical key frame images using the bag-of-words model,

[0033] Step 4.2: Select historical key frames that share more words with the current key frame image than a preset threshold as loop closure candidates;

[0034] Step 4.3: Calculate the similarity between the x-th key frame in the loop closure candidates and the y key frames adjacent to the x-th key frame in the historical key frames and the current key frame, where x = 1, 2, …, X;

[0035] Step 4.4: Select key frames with similarity greater than a preset similarity threshold as the loop closure key frame set;

[0036] Step 4.5: Perform continuity detection on the loop closure key frame set. If the continuity requirement is not met, directly go to Step 4.1; otherwise, confirm the existence of the loop, use the loop closure key frame set as a global constraint; and go to Step 4.1.

[0037] Furthermore, to calculate the similarity between the current key frame image and the historical key frame image, the L1 norm is used as a metric standard for similarity evaluation, and the specific formula is as follows:

[0038]

[0039] Among them, v1 and v2 represent the feature vectors of two frames of images, and scores represent the similarity function between the two feature vectors.

[0040] An electronic device / system for loop closure detection based on a bag-of-words model of depth images, including a processor and a memory. The memory stores execution instructions for the processor, and the processor is configured to execute the execution instructions to implement the above loop closure detection method.

[0041] A computer-readable storage medium for storing a program, and executing the program to implement the above loop closure detection method.

[0042] Beneficial effects: The depth projection method adopted by the present invention, based on the calibration of the extrinsic parameters of the lidar and the camera, without a fixed target reference, uses the edge features in the natural environment to project the laser point cloud onto the camera coordinate system to generate an initial depth image, which is beneficial to improving the system's environmental perception ability. The method of locally linearly fitting to complete the depth map adopted by the present invention sets a minimization objective function to locally fit the missing depth, combines color change and spatial change, restores the depth discontinuous region, and generates a complete depth image, which is beneficial for the system to implement loop closure detection. The loop closure detection method based on depth images adopted by the present invention has more advantages than simply relying on point cloud context scanning and ordinary image bag-of-words models, has higher loop closure detection accuracy, corrects the global error, reduces the cumulative error of long-term system operation, and is beneficial to improving the stability and positioning accuracy of the navigation system. Description of the Drawings

[0043] Figure 1 This is the flowchart of the present invention.

[0044] Figure 2 It is the effect diagram of depth map completion. Among them, (a) is the original depth image, (b) is the enlarged partial view of the original depth image, (c) is the depth image after completion, and (d) is the enlarged partial view of the depth image after completion.

[0045] Figure 3 It is the key frame selection strategy diagram of the present invention. Detailed implementation manners

[0046] The accompanying drawings that form a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0047] As Figure 1 shown, the present invention provides a loop detection method based on the bag-of-words model of depth images, specifically as follows:

[0048] Step 1: Implement the registration of the laser point cloud and the visual image, perform the external parameter calibration of the lidar and the camera, complete the external parameter calibration in the case of no fixed targets such as checkerboards, and can perform real-time online calibration in natural indoor and outdoor scenes.

[0049] Step 2: Extract the straight edges of all object contours in the field of view. The straight edges are easy to extract and align. By aligning the obvious edge features in the point cloud map of the lidar and the visible light image of the camera; convert all the point cloud coordinates to the camera coordinate system to achieve depth projection and obtain the initial depth image.

[0050] By aligning the obvious natural edge features in the point cloud map of the lidar and the visible light image of the camera, all the point cloud coordinates L P i are converted to the camera coordinate system:

[0051]

[0052] Among them, is the rigid body transformation function of the lidar and the camera, is the rotation matrix, is the translation matrix, and SE(3) is the 4×4 transformation matrix space.

[0053] Combined with the internal parameters and distortion parameters of the camera, project the points in the camera coordinate system to the pixel coordinate system and perform distortion correction on the camera projection model:

[0054] C p i = π(C P i )

[0055] p i = f( C p i )

[0056] where C P i represents the i-th point cloud in the camera coordinate system, C p i represents the i-th point cloud in the pixel coordinate system, C p i represents the i-th point cloud in the camera coordinate system, π(·) represents the projection model of the camera, and f(·) represents the distortion model of the camera; p i represents the i-th point cloud after distortion correction in the pixel coordinate system.

[0057] Optimize the error of the reprojection model to obtain the extrinsic parameters from the lidar coordinate system to the image coordinate system:

[0058]

[0059] J = [F, E]

[0060] M = O(p i )

[0061] where F is the partial derivative matrix of the reprojection model cost function with respect to the pose, the transpose of the expression unit vector, N1 represents the total number of elements in the matrix, p represents the point cloud after distortion correction, P represents the original point cloud, P ^ represents the skew-symmetric matrix of P, I represents the identity matrix, E is the partial derivative matrix of the reprojection model cost function with respect to the spatial point, represents the rotation matrix from the lidar coordinate system to the camera coordinate system; J is the Jacobian matrix. Finally, the depth image M of the point cloud projection in the pixel coordinate system is obtained by the depth projection function O(·) to achieve the registration of the laser point cloud and the visual image.

[0062] Step 3: Complete the sparse and discontinuous initial depth image. Use the local linear fitting method to set the minimization objective function to achieve local fitting of the missing depth, combine the color change and spatial change to restore the depth discontinuous region, and generate a complete depth image.

[0063] First, a minimization of the overall objective function is set up to solve for the best - fit values of the missing depth values. The first term of this function is a prior term corresponding to the weighted deviation of the local linear fitting, and an adaptive weight based on color similarity is used for the local linear fitting. The aim is to construct a linear interpolation only using the pixel points within the same object, which may be similar in both depth and color. The second term is a data - fidelity term, which measures the observed deviation at the available positions of the low - resolution depth map to preserve the edges. The conjugate gradient method is used to iteratively solve a large - scale sparse linear system equation of the form Ax = b for minimizing the objective function.

[0064] The objective function is expressed as

[0065]

[0066] where N is the total number of pixels in the image, denotes the neighborhood of pixel j, i is a pixel point in , d H,i is the estimated depth value of pixel i, is the linear fitting value of pixel i in N(j), w i,j is the weight of the similarity measure between pixel i and j, D j is the value measuring the consistency of the depth data between the completed image and the uncompleted image, E j is the value measuring the gradient of the pixel; λ is a free parameter controlling the relationship between fidelity and smoothness.

[0067] The local linear fitting is defined as:

[0068]

[0069] where α j and β j are the linear fitting parameters, α j is a 2×1 vector, β j is a scalar, g i,j represents the relative coordinate of pixel i with respect to pixel j, which is a 2×1 vector.

[0070] D j = d L,j - d H,j , E j = e L,j - e H,j

[0071] where d L,j is the depth value of pixel j in the initial depth image, d H,j is the estimated depth value of pixel j in the completed depth image, e L,j is the edge value of pixel j in the initial depth image, e H,jis the estimated edge value of pixel j in the completed depth image; the weighted local linear fitting is set according to the local relative spatial position of the neighborhood, and by adding a multiplication factor to the data fidelity term, this method can achieve depth completion.

[0072] The weight function w i,j Combines Gaussian contour information and depth information and is defined as:

[0073]

[0074] where s i represents the RGB pixel value of the pixel in the initial depth image, and s j represents the RGB pixel value of pixel j in the fitted image; both σ1 and σ2 are pre - examined pixel similarity values (set according to the pixel size of the camera), respectively representing the depth information estimated by the median filter of pixel i and pixel j based on the initial depth image.

[0075] Using the weight function w i,j This combination of color similarity and depth similarity is used to avoid inappropriate weighting in the presence of complex color patterns with planar depth spatial variations. The depth images before and after completion are as Figure 2 shown.

[0076] The bag - of - words model loop detection method based on the depth image in step 4 is as follows:

[0077] Combining the term frequency (TF) and the inverse document frequency (IDF), the weight η of the visual word is defined as the product of the two, which fuses the local importance of the feature in a single image (reflected by TF) and the global distinctiveness in the entire image dataset (reflected by IDF), so as to more accurately measure the contribution of each visual word in the image representation.

[0078]

[0079] where n i represents the frequency of occurrence of the i - th feature word in a single - frame image, and n c represents the total number of all feature words in this image, and n sum represents the total frequency of occurrence of all words in the bag - of - words dictionary. When performing loop detection on two frames of images, v1 and v2 are used to represent the feature vectors of the two frames of images, and it is necessary to evaluate the similarity of their feature vectors v1 and v2. The L1 norm is selected as the metric for this evaluation:

[0080]

[0081] After extracting the words of two key-frame images (the current key frame and any historical key frame) using the bag-of-words model, it is necessary to calculate the image similarity. In this embodiment, K-means clustering is used to calculate the similarity, specifically as follows Figure 3 shown. This involves traversing all key frames, finding historical key frames that share words with the current key frame, and recording the maximum number of shared words. Subsequently, according to a preset threshold, historical key frames with a relatively large number of shared words are initially screened out as loop closure candidates. Then, calculate the similarity between these candidate frames and their adjacent 10 key frames (the adjacent 10 key frames refer to the 10 key frames adjacent to the candidate key frame in the historical key frames) and the current frame, perform this operation on all candidate frames, select the frames with similarity exceeding the preset threshold to form a set of loop closure key frames, and perform continuity detection on the key frames in the set. If the similarity of consecutive frames can meet the set threshold and successful matching occurs, it is confirmed that the loop closure exists and a global constraint is formed.

[0082] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present invention does not further describe various possible combination methods.

Claims

1. A loop detection method based on the bag-of-words model of depth images, characterized in that, Specifically, it includes the following steps: Step 1: Register the lidar point cloud with the visual image and calibrate the extrinsic parameters of the lidar and the camera; Step 2: Extract the straight edges of all object contours within the field of view. By aligning the edge features in the point cloud map of the lidar and the visible light image of the camera, convert all point cloud coordinates to the camera coordinate system to achieve depth projection and obtain the initial depth image; Step 3: Complete the initial depth image; generate the complete depth image after completion; Step 4: For the complete depth image after completion, use the bag-of-words model loop detection mechanism to detect loops.

2. The loop closure detection method based on the bag-of-words model of depth images according to claim 1, characterized in that, The depth projection in Step 2 is specifically as follows: Step 2.1: Combine the intrinsic parameters and distortion parameters of the camera to project the points in the camera coordinate system to the pixel coordinate system and correct the distortion of the camera projection model: C p i = π( C P i ) p i = f( C p i ) Among them, C P i represents the i-th point cloud in the camera coordinate system, C p i represents the i-th point cloud in the pixel coordinate system, C p i represents the i-th point cloud in the camera coordinate system, π(·) represents the projection model of the camera, and f(·) represents the distortion model of the camera; p i represents the i-th point cloud after distortion correction in the pixel coordinate system; Optimize the error of the reprojection model to obtain the Jacobian matrix J from the lidar coordinate system to the image coordinate system: J = [F, E] Among them, F is the partial derivative matrix of the reprojection model cost function with respect to the pose, The transpose of the expression unit vector, N1 represents the total number of elements in the matrix, p represents the point cloud after distortion correction, P represents the original point cloud, P ^ Represents the skew-symmetric matrix of P, I represents the identity matrix, E is the partial derivative matrix of the reprojection model cost function with respect to the spatial point, Represents the rotation matrix from the radar coordinate system to the camera coordinate system; Use the depth projection function O(·) to obtain the initial depth image of the point cloud projection in the pixel coordinate system.

3. A bag-of-words model loop detection method based on depth images according to claim 1, characterized in that, Specifically, it includes the following steps: In Step 3, use the local linear fitting method to complete the initial depth image, and the objective function is specifically: Where N is the total number of pixels in the image, denotes the neighborhood of pixel j, and i is the pixel point in H,i the estimated depth value of pixel i, is the linear fitting value of pixel i in N(j), and w i,j is the weight of the similarity measure between pixel i and j, and D j is the consistency value for measuring the depth data of the completed image and the uncompleted image, and E j is the value for measuring the gradient of the pixel; The expression is as follows: where α j and β j are linear fitting parameters, α j is a 2×1 vector, β j is a scalar, and g i,j represents the relative coordinates between the coordinates of pixel i and the coordinates of pixel j; D j and E j The expressions are as follows: D j = d L,j -d H,j ,E j = e L,j -e H,j where d L,j is the depth value of pixel j in the initial depth image, and d H,j is the estimated depth value of pixel j in the completed depth image, e L,j is the edge value of pixel j in the initial depth image, and e H,j is the estimated edge value of pixel j in the completed depth image; w i,j The expression is as follows: Among them, s i represents the RGB pixel value of the pixel in the initial depth image, and s j represents the RGB pixel value of the pixel j in the fitted image; both σ1 and σ2 are pre-reviewed pixel similarity values, respectively representing the depth information estimated by the median filter of the pixel i and the pixel j based on the initial depth image.

4. A bag-of-words model loop detection method based on depth images according to claim 1, characterized in that The specific content of Step 4 is as follows: Step 4.1: Use the bag-of-words model to extract the words in the current key frame image and all historical key frame images; Step 4.2: Select the historical key frames that share more words with the current key frame image than the preset threshold as loop candidates; Step 4.3: Calculate the similarity between the x-th key frame in the loop candidates and the y key frames adjacent to the x-th key frame in the historical key frames, where x = 1, 2,..., X; Step 4.4: Select the key frames with similarity greater than the preset similarity threshold as the loop key frame set; Step 4.5: Perform continuity detection on the loop key frame set. If the continuity requirement is not met, directly go to Step 4.1; otherwise, confirm the existence of the loop, and use the loop key frame set as the global constraint; And go to Step 4.

1.

5. A loop detection method based on a bag-of-words model of depth images according to claim 4, characterized in that Calculate the similarity between the current key frame image and the historical key frame image, and use the L1 norm as the metric standard for similarity evaluation. The specific formula is as follows: Among them, v1 and v2 represent the feature vectors of two frames of images, and scores represents the similarity function between the two feature vectors.

6. An electronic device / system for loop detection of the bag-of-words model based on depth images, characterized in that, It includes a processor and a memory. The memory stores the execution instructions of the processor, and the processor is configured to execute the execution instructions to implement the loop detection method described in any one of claims 1-5.

7. A computer-readable storage medium for storing a program, characterized in that, Execute the program to implement the loop detection method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Simultaneous localization and mapping method based on vision and laser radar

    CN112258600A

  • Unmanned aerial vehicle scene dense reconstruction method based on VI-SLAM and depth estimation network

    CN112435325A

  • Underground coal mine positioning and mapping method based on multi-sensor fusion

    CN118730117A

  • Dynamic calibration method and device for joint calibration of three-dimensional laser radar and camera, and storage medium

    CN119228909A