A loop detection method based on autoencoder

By using an unsupervised autoencoder to extract image features in visual SLAM, the problems of low efficiency and high labor cost of traditional loopback detection methods are solved, and more efficient and accurate loopback detection is achieved.

CN114565671BActive Publication Date: 2025-05-09BEIHANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210158768.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-21
Publication Date
2025-05-09
Estimated Expiration
2042-02-21

AI Technical Summary

Technical Problem

The loopback detection method in traditional visual SLAM relies on the bag of words model, resulting in low feature extraction efficiency and high labor cost of data set annotation, making it difficult to achieve accurate loopback detection.

Method used

The loopback detection method based on an unsupervised autoencoder is adopted to extract image features by training the autoencoder, and use these features to perform similarity measurements in the loopback detection scenario to identify the loopback.

Benefits of technology

It improves the accuracy of loopback detection algorithm, reduces dependence on labeled data, reduces labor costs, and performs well in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565671B_ABST
    Figure CN114565671B_ABST
Patent Text Reader

Abstract

The present application provides a loop detection method based on an autoencoder. The provided loop detection method based on an autoencoder includes: collecting a first plurality of images in a training scenario, generating training samples from the first plurality of images to train the autoencoder; collecting a second plurality of images in a loop detection scenario, extracting ORB feature points and feature point image blocks for each of the second plurality of images; processing the feature point image blocks extracted from each of the second plurality of images with the trained autoencoder, and using the hidden layer output of the autoencoder as a feature vector of the feature point image block provided to the autoencoder; and calculating the similarity between the feature vectors of the first image and the second plurality of images to identify loops.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a loop detection solution based on an autoencoder. Background Art

[0002] GPS can be used for positioning in outdoor environments, which is low-cost and high-precision. However, when there is no GPS signal or the GPS signal is weak indoors, other methods are needed to achieve positioning. For robots and drones, when exploring unknown environments, they need to solve two problems at the same time: self-positioning and external perception. This method of estimating the subject's own position while building a model of the surrounding environment is called Simultaneous Localization and Mapping, or SLAM technology for short. Commonly used sensors in SLAM technology are lidar and cameras, which are divided into laser SLAM and visual SLAM. Classic visual SLAM consists of five modules.

[0003] As part of the SLAM technology framework (see Figure 1 ), loop detection allows the subject to identify the places it has been to, thereby eliminating the accumulated errors generated in the loop process. In traditional visual SLAM, loop detection uses the bag-of-words model to compare the similarity of images to determine whether the subject has reached the same location. However, in the feature extraction based on the bag-of-words model, the extracted features are all artificially designed, the information utilization rate in the image is not high, and the generalization performance is poor. In recent years, with the great success of deep learning in the field of images, researchers have begun to try to use convolutional neural networks to solve the loop detection problem, because the image features extracted from convolutional neural networks have better performance than artificially designed features. However, if a new convolutional neural network needs to be trained, a related data set needs to be constructed. However, the construction of an annotated data set requires a lot of manpower and material resources. Therefore, a loop detection method for indoor environments based on unsupervised autoencoders is proposed. Summary of the invention

[0004] One purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art and propose a new loop detection method based on an autoencoder. Compared with the traditional bag-of-words model method, the features extracted by the autoencoder have better performance and can improve the accuracy of the loop detection algorithm. At the same time, the autoencoder is trained by reconstructing the input, and there is no need to label the data, which can greatly save labor costs.

[0005] According to a first aspect of the present application, a loop detection method based on an unsupervised autoencoder is provided, comprising: collecting a first plurality of images in a training scenario, generating training samples from the first plurality of images to train the autoencoder, wherein the neural network of the autoencoder comprises an input layer, a hidden layer and an output layer; collecting a second plurality of images in a loop detection scenario, extracting ORB feature points and feature point image blocks for each of the second plurality of images; processing the feature point image blocks extracted from each of the second plurality of images with the trained autoencoder, and outputting the hidden layer of the autoencoder as feature vectors provided to the feature point image blocks of the autoencoder; for a first image in the second plurality of images, calculating the similarity of the feature vectors of the first image with one or more images in the second plurality of images; if the similarity of the feature vectors of the first image with images in the second plurality of images other than the first image is greater than a specified threshold, a loop is identified.

[0006] According to the second aspect of the present application, an information processing device is provided, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor implements a loop detection method based on an unsupervised autoencoder according to the first aspect of the present application when executing the program. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0008] Figure 1 Shows the framework diagram of classic visual SLAM technology;

[0009] Figure 2 is a flow chart of a loop detection method according to an embodiment of the present application. DETAILED DESCRIPTION

[0010] The following is a clear and complete description of the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0011] The present application provides a new visual SLAM loop detection method based on an unsupervised autoencoder, the method comprising:

[0012] Step 1: Collect a large number of images offline in the training scene to train the autoencoder. The training scene is similar to the scene where the SLAM technology is actually applied and loop detection is performed, such as indoor scenes, park scenes, etc.

[0013] Step 2: extract ORB feature points from the collected single image, select N feature points, and resize the feature point image blocks corresponding to the feature points to a fixed size of s×s, where N and s are, for example, positive integers.

[0014] Step 3: Use the feature point image blocks as training samples to train the autoencoder, vectorize the feature point image blocks into one dimension, and then feed them into the autoencoder. The autoencoder network structure is divided into three layers: input layer x; hidden layer h; output layer y, where x is the input of h, and h is the input of y.

[0015] h=f(x)=σ(W T x+b)

[0016] W T , b is the weight and bias of the hidden layer

[0017] y=g(h)=σ(W′ T h+b′)=g(f(x))

[0018] W′ T , b′ is the weight and bias of the output layer

[0019] The purpose of the autoencoder is to restore the input, so y≈x. In practice, w, b are trained by minimizing the loss function. Here, cross entropy is used to measure the distance d between the input and output.

[0020]

[0021] i represents the index of the training sample (feature point image block) provided to the autoencoder, and n is the total number of training samples.

[0022] Step 4: Use the optimizer to solve w, b by minimizing the loss function J = KL(x, g(f(x))).

[0023] Step 5: In the loop detection scenario, the output of the previously trained autoencoder is used to measure the similarity of multiple images obtained during the loop detection process to detect whether a loop occurs. During detection, ORB feature point extraction is performed on different images collected in the loop detection scenario. For example, N feature points are selected and the corresponding image block size is adjusted to a fixed size of s×s. The image block is vectorized in one dimension and fed into the neural network. The output of the corresponding hidden layer is extracted for subsequent similarity measurement scoring. If the score exceeds the threshold θ, it is confirmed that a loop has been detected. The threshold θ is a hyperparameter.

[0024] Definition of similarity between two frames of images:

[0025] Two frames of image F (1) , F (2) Contains k 1 , k 2 ORB feature points, after extracting the feature points, can be input into the trained autoencoder to obtain the dense representation of the image blocks corresponding to the feature points in the two frames of images, that is Where h is the hidden layer output of the autoencoder ( The superscript of h refers to the two frames (1) and (2), and the subscript refers to the index of the ORB feature point. 1 =k 2 Compared to the original input, the output of the hidden layer has a lower dimension and is denser, which is a better feature representation and can be used to measure the similarity between two images. The specific similarity measurement is shown in the following table:

[0026] Algorithm Flow

[0027]

[0028]

[0029] It can be understood that the above steps 1 to 4 are used to train the autoencoder and use the images obtained in the training scene. The process of obtaining images in the training scene and the process of training the autoencoder do not need to occur at the same time. The scene of loop detection and the training scene can be similar but different scenes. The process of loop detection and the process of training the autoencoder do not need to occur at the same time.

[0030] In summary, the loop closure detection method disclosed in this example can effectively improve the shortcomings of low efficiency in image feature extraction and high labor cost of data set annotation in traditional loop closure detection methods, and can achieve good results in different application scenarios, and has very broad use value and application prospects.

[0031] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application. Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. Loop detection method based on unsupervised autoencoder, including: Collecting a first plurality of images in a training scene, generating training samples from the first plurality of images to train the autoencoder, wherein the neural network of the autoencoder includes an input layer, a hidden layer, and an output layer; Collecting a second plurality of images in the loop closure detection scene, and extracting ORB feature points and feature point image blocks from each of the second plurality of images; Processing the feature point image blocks extracted from each of the second plurality of images with the trained autoencoder, and using the hidden layer output of the autoencoder as feature vectors of the feature point image blocks provided to the autoencoder; For a first image in the second plurality of images, calculating a similarity between a feature vector of the first image and one or more images in the second plurality of images; If the similarity between the feature vectors of the first image and the image other than the first image in the second plurality of images is greater than a specified threshold, a loop is identified; The step of processing the feature point image blocks extracted from each of the second plurality of images using the trained autoencoder comprises: acquiring the first image and a second image of the second plurality of images; Extracting a plurality of ORB feature points from the first image and the second image respectively, and adjusting the size of a feature point image block of each ORB feature point to s×s, where s is a positive integer; The feature point image blocks of the first image and the second image are respectively one-dimensionally vectorized and fed into the trained autoencoder. The feature vector of the feature point image block of the ORB feature point with index j1 of the first image output by the hidden layer of the neural network of the autoencoder is Where 1≤j1≤k1, k1 is the number of ORB feature points of the first image, and the feature vector of the feature point image block of the ORB feature point of the second image with index j2 output by the hidden layer of the neural network of the autoencoder is Wherein, 1≤j2≤k2, k2 is the number of ORB feature points of the second image; and wherein, The calculating the similarity between the feature vectors of the first image and one or more images of the second plurality of images comprises: Match the ORB feature points of the first image with the ORB feature points of the second image to obtain a set of matching points 1≤j1≤k1,1≤j2≤k2}; For each element mk in the matching point set, calculate the similarity score 1≤j1≤k1,1≤j2≤k2; Accumulate the similarity score sk of each element mk of the matching point set to obtain the similarity between the first image and the second image; wherein, The training of the autoencoder comprises: Extracting multiple ORB feature points from the first multiple images, and adjusting the size of a feature point image block of each ORB feature point to sxs, where s is a positive integer, and the feature point image block is used as a training sample; The feature point image block is one-dimensionally vectorized and fed into the input layer of the autoencoder to train the neural network of the autoencoder, where x is the result of the one-dimensional vectorization of the image block, and the output layer of the neural network of the autoencoder is y. The distance between x and y is measured by cross entropy. Where i represents the index of the feature point image block provided to the autoencoder, x i is the result of one-dimensional vectorization of the i-th image block, y i For the same i The corresponding output layer output of the neural network of the autoencoder, n is the total number of training samples, i and n are positive integers; The neural network of the autoencoder is trained by minimizing the loss function J=KL(x,y).

2. The method according to claim 1, wherein The input of the input layer of the neural network of the autoencoder is x, and the corresponding output of the hidden layer of the neural network of the autoencoder is h, h=f(x)=σ(W T x+b), W T , b is the weight and bias of the hidden layer, y=g(h)=σ(W′ T h+b′)=g(f(x)),W′ T , b′ is the weight and bias of the output layer, The autoencoder is trained by minimizing the loss function J to obtain the weight W of the hidden layer. T , and the bias b of the hidden layer.

3. The method according to claim 2, further comprising: taking each image in the second plurality of images except the first image as the second image, and calculating the similarity between the first image and the second image respectively; If the similarity between the first image and the second image is greater than a specified threshold, the step of calculating the similarity between the feature vectors of the first image and one or more images of the second plurality of images is terminated.

4. The method according to claim 3, further comprising: A plurality of images in the second plurality of images are used as the first images, and similarities between feature vectors of the first image and one or more images in the second plurality of images are calculated to identify loops.

5. An information processing device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the loop detection method based on the unsupervised autoencoder of any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Visual SLAM closed-loop detection method based on depth neural network

    CN107330357A

  • A loopback detection method based on a convolutional neural network and ORB features

    CN109934857A

  • Robot vision SLAM closed-loop detection method based on stack type combined auto-encoder

    CN111753789A