A laser SLAM loop detection method based on semantic information

By spherical projection and semantic segmentation of laser point clouds, the problem of failure to fully utilize environmental semantic information in the SLAM algorithm is solved, efficient loopback detection and error correction are achieved, and real-time and accuracy of the SLAM algorithm are improved.

CN115345932BActive Publication Date: 2025-08-19UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210734017.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-08-19
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

The existing SLAM algorithm fails to fully utilize environmental semantic information in loopback detection, resulting in low detection efficiency.

Method used

By spherical projection and semantic segmentation of the laser point cloud, a global descriptor with rotation invariance is generated, and a geometric verification is used to determine the loopback frame.

Benefits of technology

It improves the efficiency and accuracy of loopback detection, reduces error accumulation, and improves the real-time and accuracy of SLAM algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115345932B_ABST
    Figure CN115345932B_ABST
Patent Text Reader

Abstract

The present invention discloses a laser SLAM loop detection method based on semantic information, comprising: performing spherical projection on the laser point cloud in the current environment and obtaining the depth information of each point in the two-dimensional image generated after the projection; performing semantic segmentation through a fully convolutional neural network and generating a weight matrix using the obtained semantic labels; aggregating and rotating the weight matrix through a sliding window to generate a global descriptor with rotation invariance; constructing a Kd-Tree to search historical frames and obtain candidate frames based on time thresholds and distance thresholds; obtaining the angular difference between two frames of point clouds using the global descriptor, and performing geometric verification through an ICP algorithm with an initial angle to determine the final loop frame and obtain a closed-loop position. The present invention solves the problem that the existing technology does not fully utilize the semantic information of the environment when performing loop detection using pure laser, resulting in low algorithm efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of simultaneous positioning and mapping (SLAM) technology, and in particular to a laser SLAM loop detection method based on semantic information. Background Art

[0002] In recent years, the rapid development of the robotics industry has greatly facilitated people's daily lives and improved their quality of life. The key to autonomous robot mobility lies in simultaneous localization and mapping (SLAM). Starting from an unknown location in an unfamiliar environment, robots use their sensors to collect information such as distance, complete localization, and construct a map of the environment. Compared to GPS, SLAM can be applied to smaller environments and has fewer usage restrictions, making it a research hotspot in the robotics field.

[0003] When using SLAM algorithms to construct an environmental map, the robot matches the images or laser information collected by its sensors at each moment in time to create a complete map. However, as the robot continues to move and is disturbed by environmental factors, errors accumulate constantly during the map-building process. Although many algorithms provide real-time corrections, the accumulation of errors is unavoidable. Therefore, loop closure detection is extremely important. Loop closure detection is used to detect whether the robot has returned to the same location it has previously visited, thereby making large-scale corrections to previously accumulated errors, greatly improving the robot's positioning accuracy and the accuracy of map construction.

[0004] In laser SLAM, the most commonly used loop detection method is to use KD-Tree to perform position search, and then use the ICP algorithm to obtain the matching score between the current frame and the historical frame, so as to confirm the loop frame and add the optimization link, which is relatively time-consuming. In order to realize the loop detection function more efficiently, the method of constructing a global descriptor is often adopted. The global descriptor is a feature vector constructed from a frame of point cloud. Matching the descriptor can save a lot of time. After the matching is completed, geometric verification is required to prevent greater errors caused by incorrect matching. However, the extracted descriptors can often only be used for the loop detection process and cannot further improve the operating efficiency of the SLAM algorithm. Therefore, when the existing SLAM algorithm uses pure laser for loop detection, it does not make full use of the semantic information of the environment, resulting in low algorithm efficiency. Summary of the Invention

[0005] The present invention provides a laser SLAM loop detection method based on semantic information to solve the technical problem that when the current SLAM algorithm uses pure laser for loop detection, it does not fully utilize the semantic information of the environment, resulting in low algorithm efficiency, thereby achieving efficient and rapid loop detection.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In one aspect, the present invention provides a laser SLAM loop detection method based on semantic information, comprising:

[0008] Perform spherical projection on the laser points in the laser point cloud in the current environment to obtain a two-dimensional image generated after projection, and obtain the depth information of each point in the two-dimensional image generated after projection;

[0009] Performing semantic segmentation on the two-dimensional image using a preset fully convolutional neural network to obtain a semantic label for each point in the two-dimensional image, and generating a weight matrix using the obtained semantic label;

[0010] Aggregate and rotate the weight matrix through a sliding window to generate a global descriptor with rotation invariance;

[0011] Construct a Kd-Tree to search historical frames, measure the distance between the global descriptors of the current frame and the historical frames, and the interval time between the current frame and the historical frames, and obtain candidate frames according to the time threshold and distance threshold;

[0012] The global descriptor is used to obtain the angle difference between the two frames of point cloud, and the geometric verification is performed through the ICP algorithm with the initial angle. The final loop frame is determined from the candidate frames to obtain the closed-loop position.

[0013] Furthermore, the performing spherical projection on the laser points in the laser point cloud in the current environment to obtain a two-dimensional image generated after the projection, and obtaining the depth information of each point in the two-dimensional image generated after the projection, includes:

[0014] Using the laser point coordinates of the current frame, each laser point in the laser point cloud in the current environment is spherically projected to obtain a two-dimensional image generated after projection, thereby achieving coordinate dimensionality reduction. The projection formula is:

[0015]

[0016] Where u and v are the horizontal and vertical coordinates of the midpoint of the projected two-dimensional image, x, y, and z are the coordinates of the laser point in the current coordinate system, R is the distance measured by the laser radar, and FOV represents the number of lines of the mechanical laser radar used. up Indicates the number of lines where the LiDAR is located in the upper half, row_size and col_size respectively indicate the number of rows of points that can be obtained by the LiDAR in one scan and the number of laser points collected in each row;

[0017] For the two-dimensional image obtained after projection, the depth of each point is calculated and normalized to obtain the final depth information. The formula for calculating the depth of each point and normalizing it is:

[0018]

[0019] Where Range is the depth information and max_Range is the maximum distance of the point cloud in the current frame.

[0020] Furthermore, the preset fully convolutional neural network uses the Darknet53 framework as the skeleton and adopts an encoding-decoding structure, including residual blocks, convolution blocks and deconvolution blocks; wherein,

[0021] The output features of the encoder are jump-linked to the corresponding output features of the decoder;

[0022] Each residual block consists of two convolution blocks, each of which consists of a convolution layer, batch normalization, and activation function. The convolution kernel size, stride, and padding of the two convolution blocks in each residual block are 1x1, 1x1, 0 and 3x3, 1x1, 1, respectively. The activation function uses the LeakyRelu function. Each individual convolution block is used to achieve horizontal downsampling, with a convolution kernel size of 3x3, a stride of 1x2, and a padding of 1.

[0023] During the encoding process, the input passes through a residual block and then enters the first convolution block, followed by two residual blocks, the second convolution block, eight residual blocks, the third convolution block, eight residual blocks, the fourth convolution block and four residual blocks. After encoding, the horizontal resolution is reduced to one sixteenth of the original; the decoding process uses four deconvolution blocks, and each time it passes through a deconvolution block, the horizontal resolution becomes twice the original; after deconvolving the image to the original resolution, a 1x1 convolution is used to generate an output with M channels, and a Softmax operation is performed on this basis to obtain the probability distribution of each pixel; where M is the number of categories divided; the training process uses stochastic gradient descent and weighted cross entropy as the loss function.

[0024] Furthermore, the expression of the LeakyRelu function is:

[0025] y Activate =max(0,x Activate )+leak*min(0,x Activate )

[0026] Where leak is a constant, x Activate is the activation function input, y Activate Output of the function;

[0027] The expression of the Softmax function is:

[0028]

[0029] Where, Represents classification probability, logits c Represents the output probability value of category c;

[0030] The expression of the loss function is:

[0031]

[0032] Where L represents the loss function value, N is the number of samples, and y c Indicates whether the current classification is correct. If the classification is correct, y c Take 1, if the classification is incorrect, then y c Take 0, i represents the i-th sample.

[0033] Furthermore, semantic segmentation is performed on the two-dimensional image using a preset fully convolutional neural network to obtain a semantic label for each point in the two-dimensional image, and a weight matrix is generated using the obtained semantic label, including:

[0034] The two-dimensional image is semantically segmented using a preset fully convolutional neural network to obtain a semantic label for each point in the two-dimensional image. Weighted processing is performed according to the label probability and label category to generate a weight matrix. The weight generation formula is:

[0035]

[0036] In the formula, w represents the weight value, is the pre-set category weight, p c Represents the probability value of the point category.

[0037] Furthermore, the step of aggregating and rotating the weight matrix through a sliding window to generate a global descriptor with rotation invariance includes:

[0038] Superimposing the weights of each column of the weight matrix to reduce the number of rows to one;

[0039] For the weight matrix after dimensionality reduction, a sliding window is used to reduce the dimensionality again and normalize it to generate the value of each dimension of the descriptor to obtain the feature descriptor;

[0040] The obtained feature descriptor is rotated, the maximum weight is rotated to the initial position of the matrix, and the rotation dimension is recorded; a global descriptor with rotation invariance is generated.

[0041] Furthermore, the calculation formula for generating the value of each dimension of the descriptor is:

[0042]

[0043] In the formula, j represents the current calculation dimension, w j Represents the value of the j-th dimension of the weight matrix after dimensionality reduction, n is the window size, desc max The maximum dimension of the generated descriptor.

[0044] Furthermore, when measuring the distance between the global descriptors of the current frame and the historical frame, the distance metric adopts L1 distance, and the formula is:

[0045]

[0046] In the formula, dis is the calculated distance value, dim is the descriptor dimension, Represents the value of the j-th dimension of the global descriptor of the current frame, Represents the value of the j-th dimension of the global descriptor of the historical frame.

[0047] Furthermore, the global descriptor is used to obtain the angle difference between the two frames of point cloud and geometric verification is performed through the ICP algorithm with the initial angle. The final loop frame is determined from the candidate frames to obtain the closed-loop position, including:

[0048] Geometric verification is performed using the ICP algorithm with an initial angle, and only laser points with the same classification label are matched from high to low according to the weight. When the ICP score corresponding to the candidate frame exceeds the preset threshold, the current candidate frame is confirmed to be a loop frame and the closed-loop position is obtained.

[0049] Furthermore, the estimation of the initial angle is derived from the following formula:

[0050]

[0051] Where A represents the estimated value of the initial angle, Shift_num cur Indicates the descriptor rotation dimension of the current frame, Shift_num his Indicates the descriptor rotation dimension of the historical frame, dim is the descriptor dimension.

[0052] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.

[0053] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, and the instruction is loaded and executed by a processor to implement the above method.

[0054] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0055] 1. Compared with existing laser SLAM loop detection methods, the method of the present invention fully utilizes the depth image after laser point cloud projection and the semantic information in each frame of laser point cloud. On the one hand, the semantics are used to weight the laser point cloud, effectively reducing the interference caused by dynamic obstacles during the registration process. On the other hand, the global descriptor is generated by using the projected matrix, which reduces the redundancy of the intermediate process and eliminates the need for additional methods to generate features and match them, effectively improving the real-time performance of the SLAM algorithm.

[0056] 2. The laser SLAM loop detection method of the present invention uses semantic segmentation to segment laser point clouds, generating rotationally invariant descriptors. This fully utilizes spatial information and further accelerates the search for closed-loop frames using L1 distance. The descriptors themselves contain angular information. Using the descriptors generated from two point cloud frames, they provide an initial angular estimate for the point cloud registration algorithm, further improving loop detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0058] Figure 1 1 is a schematic diagram of the execution flow of the laser SLAM loop detection method based on semantic information provided by an embodiment of the present invention;

[0059] Figure 2 3D laser point cloud projection process according to an embodiment of the present invention;

[0060] Figure 3 This is a diagram of the semantic segmentation network structure provided by an embodiment of the present invention;

[0061] Figure 4 It is a schematic diagram of the descriptor extraction process provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0062] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0063] First embodiment

[0064] With the development of deep learning, more and more image processing methods have been introduced into the processing of laser point clouds, providing new ideas for the loop detection module of laser SLAM. Although the information collected by the laser radar does not have texture features like images, the distance information measured by the laser radar itself can also be converted into depth information, so that a frame of laser point cloud can be approximated as an image for processing. Through depth information, deep learning can be used to extract deeper semantic information, and the semantic labels of laser points can be used to generate globally consistent descriptors for loop detection and point cloud registration processes. Based on this, in view of the relatively low real-time performance of the current laser SLAM loop detection part and the problem that the information contained in the feature descriptor cannot be fully utilized, this embodiment provides a laser SLAM loop detection method based on semantic information to solve the problem of low laser loop detection efficiency and further optimize the real-time performance of the algorithm.

[0065] The method can be implemented by an electronic device. The execution process of the method is as follows Figure 1 As shown, the following steps are included:

[0066] S1, performing spherical projection on the laser points in the laser point cloud in the current environment to obtain a two-dimensional image generated after the projection, and obtaining the depth information of each point in the two-dimensional image generated after the projection;

[0067] Specifically, in this embodiment, the above S1 includes the following steps:

[0068] S11, performing spherical projection on each laser point in the laser point cloud in the current environment using the laser point coordinates of the current frame to obtain a two-dimensional image generated after projection, thereby achieving dimensionality reduction of the coordinates;

[0069] The projection process is as follows: Figure 2 As shown in the figure, the specific process is: spherical projection of the laser points is performed, and the three-dimensional laser point cloud is projected into a two-dimensional image with a size of 64x1024. The specific projection formula is:

[0070]

[0071] Where u and v are the horizontal and vertical coordinates of the midpoint of the projected two-dimensional image, x, y, and z are the coordinates of the laser point in the current coordinate system, R is the distance measured by the laser radar, and FOV represents the number of lines of the mechanical laser radar used. In this embodiment, it is 64. up Indicates the number of lines where the LiDAR is located in the upper half. In this embodiment, it is 32. row_size and col_size respectively indicate the number of rows of points that can be obtained by the LiDAR in one scan and the number of laser points collected in each row. In the worst case, 1028 points can be obtained per row. Here, to facilitate subsequent semantic segmentation, they are set to 64 and 1024, respectively.

[0072] S12, calculating the depth of each point in the two-dimensional image obtained after projection and normalizing it to obtain final depth information; wherein the formula for calculating the depth of each point and normalizing it is:

[0073]

[0074] Where Range is the depth information and max_Range is the maximum distance of the point cloud in the current frame.

[0075] S2, performing semantic segmentation on the two-dimensional image using a preset fully convolutional neural network to obtain a semantic label for each point in the two-dimensional image, and generating a weight matrix using the obtained semantic label;

[0076] Specifically, in this embodiment, the preset full convolutional neural network is as follows: Figure 3 As shown in the figure, In is the input two-dimensional image, which in this example is a 64x1024x4 tensor. Conv is the convolution block, RB is the residual block, DeConv is the deconvolution block, and OUT is the output, storing the pixel classification results. This model uses the Darknet53 framework as its backbone and adopts an encoder-decoder structure, primarily consisting of residual blocks, convolution blocks, and deconvolution blocks. The output features of the encoder are jump-linked to the corresponding output features of the decoder.

[0077] Each residual block consists of two convolution blocks, each of which consists of a convolution layer, batch normalization, and activation function. The convolution kernel size, stride, and padding of the two convolution blocks in each residual block are 1x1, 1x1, 0 and 3x3, 1x1, 1, respectively. The activation function uses the LeakyRelu function. Each individual convolution block is used to achieve horizontal downsampling, with a convolution kernel size of 3x3, a stride of 1x2, and a padding of 1.

[0078] During the encoding process, the input in this embodiment is a 64x1024x4 tensor, and the four channels are depth and three spatial coordinates. After the input passes through a residual block, it enters the first convolution block, and then sequentially connects two residual blocks, the second convolution block, eight residual blocks, the third convolution block, eight residual blocks, the fourth convolution block and four residual blocks. After encoding, the horizontal resolution is reduced to one sixteenth of the original. This embodiment generates an output vector of size 64x64x1024.

[0079] The decoding process uses four deconvolution blocks. Each time the image passes through a deconvolution block, the horizontal resolution is doubled. After deconvolving the image to the original resolution, a 1x1 convolution is used to generate an output with M channels. A Softmax operation is performed on this basis to obtain the probability distribution of each pixel. Here, M is the number of categories divided, which is 14 in this embodiment.

[0080] The training process uses stochastic gradient descent and weighted cross entropy as the loss function.

[0081] Furthermore, the expression of the LeakyRelu function is:

[0082] y Activate =max(0,x Activate )+leak*min(0,x Activate )

[0083] Where, leak represents a very small constant, which is 0.1 in this embodiment, and x Activate is the activation function input, y Activate Output of the function;

[0084] The expression of the Softmax function is:

[0085]

[0086] Where, Represents classification probability, logits c Represents the output probability value of category c;

[0087] The expression of the loss function is:

[0088]

[0089] Where L represents the loss function value, N is the number of samples, and y c Indicates whether the current classification is correct. If the classification is correct, y c Take 1, if the classification is incorrect, then y c Take 0, i represents the i-th sample.

[0090] In the above S2, after obtaining the semantic labels, they are weighted according to their label probabilities and label categories to generate a weight matrix. The weight generation formula is:

[0091]

[0092] In the formula, w represents the weight value, is the pre-set category weight, p c Represents the probability value of the point category.

[0093] Specifically, in this embodiment, the label categories are divided into 14 categories, including: abnormal points, vehicles (bicycles, motorcycles, cars, trucks, buses), pedestrians, roads, buildings, shrubs, trees, railings, traffic signals, and columnar objects. The labels are divided into 5 categories from static to dynamic, from space to the ground, and the weights of the points are divided into 5 categories from high to low, and the weights are set from 0.4 to 0 respectively.

[0094] S3, aggregates and rotates the weight matrix through a sliding window to generate a global descriptor with rotation invariance;

[0095] Specifically, in this embodiment, the above S3 includes the following steps:

[0096] S31, superimposing the weights of each column of the weight matrix to reduce the number of rows to one;

[0097] S32, using a sliding window to reduce the dimension of the weight matrix after dimensionality reduction, and performing normalization processing to generate the value of each dimension of the descriptor to obtain a feature descriptor;

[0098] The calculation formula for generating the value of each dimension of the descriptor is:

[0099]

[0100] In the formula, j represents the current calculation dimension, w j Represents the value of the j-th dimension of the weight matrix after dimensionality reduction, n is the window size, desc max The maximum dimension of the generated descriptor is . In this embodiment, the window size is set to 8 and the step size is set to 8. This example finally generates a 1x128-dimensional descriptor, such as Figure 4 shown.

[0101] S33, rotating the obtained feature descriptor, rotating the maximum weight to the initial position of the matrix, and recording the rotation dimension; generating a global descriptor with rotation invariance.

[0102] S4, constructing a Kd-Tree to search historical frames, measuring the distance between the global descriptors of the current frame and the historical frames and the interval time between the current frame and the historical frames, and obtaining candidate frames according to the time threshold and distance threshold;

[0103] Specifically, in this embodiment, when measuring the distance between the global descriptors of the current frame and the historical frame, the distance metric adopts the L1 distance, and the formula is:

[0104]

[0105] In the formula, dis is the calculated distance value, dim is the descriptor dimension, Represents the value of the j-th dimension of the global descriptor of the current frame, Represents the value of the j-th dimension of the global descriptor of the historical frame.

[0106] Specifically, in this embodiment, the time is set to 30 seconds, and the 20 frames with the closest distance are selected as candidate frames.

[0107] S5, using the global descriptor to obtain the angle difference between the two frames of point cloud, and performing geometric verification through the ICP algorithm with the initial angle, determining the final loop frame from the candidate frames, and obtaining the closed-loop position.

[0108] Specifically, in this embodiment, in S5 above, geometric verification is performed using the ICP algorithm with an initial angle, and only laser points with the same classification label are matched according to the weight from high to low. When the ICP score corresponding to the candidate frame exceeds the preset threshold, the current candidate frame is confirmed to be a loop frame, and the closed-loop position is obtained. In this embodiment, the preset threshold is 0.1, and the initial angle is estimated from the following formula:

[0109]

[0110] Where A represents the estimated value of the initial angle, Shift_num cur Indicates the descriptor rotation dimension of the current frame, Shift_num his Indicates the descriptor rotation dimension of the historical frame, dim is the descriptor dimension.

[0111] In summary, this embodiment makes full use of the depth image after the laser point cloud is projected and the semantic information in each frame of the laser point cloud. On the one hand, semantics is used to weight the laser point cloud, which effectively reduces the interference caused by dynamic obstacles in the registration process. On the other hand, the global descriptor is generated by using the projected matrix to reduce the redundant part of the intermediate process. There is no need to use additional methods to generate features and match them, which effectively improves the real-time performance of the SLAM algorithm. The loop detection method of this embodiment uses semantics to segment the laser point cloud, thereby generating a descriptor with rotation invariance, making full use of spatial information, and using L1 distance to further speed up the search for closed-loop frames. The descriptor itself has angle information. The descriptor generated by two frames of point cloud provides an initial angle estimate for the point cloud registration algorithm, further improving the efficiency of loop detection.

[0112] Second embodiment

[0113] This embodiment provides an electronic device, which includes a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method of the first embodiment.

[0114] The electronic device may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) and one or more memories, wherein the memory stores at least one instruction, which is loaded by the processor to execute the above method.

[0115] Third embodiment

[0116] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment described above. The computer-readable storage medium may be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device. The instructions stored therein can be loaded by a processor in a terminal to execute the method described above.

[0117] Furthermore, it should be noted that the present invention may be provided as a method, apparatus, or computer program product. Thus, embodiments of the present invention may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention may take the form of a computer program product embodied on one or more computer-usable storage media containing computer-usable program code.

[0118] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0119] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0120] It should also be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal device comprising the element.

[0121] Finally, it should be noted that the above is a preferred embodiment of the present invention. It should be noted that although the preferred embodiment of the present invention has been described, it is clear that those skilled in the art, once they understand the basic inventive concept of the present invention, can make various improvements and modifications without departing from the principles of the present invention. Such improvements and modifications should also be considered as within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.

Claims

1. A laser SLAM loop detection method based on semantic information, characterized in that: include: Perform spherical projection on the laser points in the laser point cloud in the current environment to obtain a two-dimensional image generated after projection, and obtain the depth information of each point in the two-dimensional image generated after projection; Performing semantic segmentation on the two-dimensional image using a preset fully convolutional neural network to obtain a semantic label for each point in the two-dimensional image, and generating a weight matrix using the obtained semantic label; Aggregate and rotate the weight matrix through a sliding window to generate a global descriptor with rotation invariance; Construct a Kd-Tree to search historical frames, measure the distance between the global descriptors of the current frame and the historical frames, and the interval time between the current frame and the historical frames, and obtain candidate frames according to the time threshold and distance threshold; The angle difference between the two frames of point cloud is obtained using the global descriptor, and geometric verification is performed using the ICP algorithm with the initial angle. The final loop frame is determined from the candidate frames to obtain the closed-loop position. The method of performing semantic segmentation on the two-dimensional image by using a preset fully convolutional neural network to obtain a semantic label for each point in the two-dimensional image and generating a weight matrix using the obtained semantic label includes: The two-dimensional image is semantically segmented using a preset fully convolutional neural network to obtain a semantic label for each point in the two-dimensional image. Weighted processing is performed according to the label probability and label category to generate a weight matrix. The weight generation formula is: In the formula, w represents the weight value, is the pre-set category weight, p c Represents the probability value of the point category; The weight matrix is aggregated and rotated through a sliding window to generate a global descriptor with rotation invariance, including: Superimposing the weights of each column of the weight matrix to reduce the number of rows to one; For the weight matrix after dimensionality reduction, a sliding window is used to reduce the dimensionality again and normalize it to generate the value of each dimension of the descriptor to obtain the feature descriptor; The obtained feature descriptor is rotated, the maximum weight is rotated to the initial position of the matrix, and the rotation dimension is recorded; a global descriptor with rotation invariance is generated; The calculation formula for generating the value of each dimension of the descriptor is: In the formula, j represents the current calculation dimension, w j Represents the value of the j-th dimension of the weight matrix after dimensionality reduction, n is the window size, desc max The maximum dimension of the generated descriptor.

2. The laser SLAM loop detection method based on semantic information as claimed in claim 1, wherein The spherical projection is performed on the laser points in the laser point cloud in the current environment to obtain a two-dimensional image generated after the projection, and the depth information of each point in the two-dimensional image generated after the projection is obtained, including: Using the laser point coordinates of the current frame, each laser point in the laser point cloud in the current environment is spherically projected to obtain a two-dimensional image generated after projection, thereby achieving coordinate dimensionality reduction. The projection formula is: Where u and v are the horizontal and vertical coordinates of the midpoint of the projected two-dimensional image, x, y, and z are the coordinates of the laser point in the current coordinate system, R is the distance measured by the laser radar, and FOV represents the number of lines of the mechanical laser radar used. up Indicates the number of lines where the LiDAR is located in the upper half, row_size and col_size respectively indicate the number of rows of points that can be obtained by the LiDAR in one scan and the number of laser points collected in each row; For the two-dimensional image obtained after projection, the depth of each point is calculated and normalized to obtain the final depth information. The formula for calculating the depth of each point and normalizing it is: Where Range is the depth information and max_Range is the maximum distance of the point cloud in the current frame.

3. The laser SLAM loop detection method based on semantic information as claimed in claim 1, wherein The preset fully convolutional neural network uses the Darknet53 framework as its skeleton and adopts an encoding-decoding structure, including residual blocks, convolution blocks, and deconvolution blocks; wherein, The output features of the encoder are jump-linked to the corresponding output features of the decoder; Each residual block consists of two convolution blocks, each of which consists of a convolution layer, batch normalization, and activation function. The convolution kernel size, stride, and padding of the two convolution blocks in each residual block are 1x1, 1x1, 0 and 3x3, 1x1, 1, respectively. The activation function uses the LeakyRelu function. Each individual convolution block is used to achieve horizontal downsampling, with a convolution kernel size of 3x3, a stride of 1x2, and a padding of 1. During the encoding process, the input passes through a residual block and then enters the first convolution block, followed by two residual blocks, the second convolution block, eight residual blocks, the third convolution block, eight residual blocks, the fourth convolution block and four residual blocks. After encoding, the horizontal resolution is reduced to one sixteenth of the original; the decoding process uses four deconvolution blocks, and each time it passes through a deconvolution block, the horizontal resolution becomes twice the original; after deconvolving the image to the original resolution, a 1x1 convolution is used to generate an output with M channels, and a Softmax operation is performed on this basis to obtain the probability distribution of each pixel; where M is the number of categories divided; the training process uses stochastic gradient descent and weighted cross entropy as the loss function.

4. The laser SLAM loop detection method based on semantic information as claimed in claim 3, wherein The expression of the LeakyRelu function is: y Activate =max(0,x Activate )+leak*min(0,x Activare ) Where leak is a constant, x Activate is the activation function input, y Activate Output of the function; The expression of the Softmax function is: Where, Represents classification probability, logits c Represents the output probability value of category c; The expression of the loss function is: Where L represents the loss function value, N is the number of samples, and y c Indicates whether the current classification is correct. If the classification is correct, y c Take 1, if the classification is incorrect, then y c Take 0, i represents the i-th sample.

5. The laser SLAM loop detection method based on semantic information as claimed in claim 1, wherein When measuring the distance between the global descriptors of the current frame and the historical frame, the distance metric adopts L1 distance, and the formula is: In the formula, dis is the calculated distance value, dim is the descriptor dimension, Represents the value of the j-th dimension of the global descriptor of the current frame, Represents the value of the j-th dimension of the global descriptor of the historical frame.

6. The laser SLAM loop detection method based on semantic information as claimed in claim 1, wherein The global descriptor is used to obtain the angle difference between the two frames of point cloud, and the geometric verification is performed through the ICP algorithm with the initial angle. The final loop frame is determined from the candidate frames to obtain the closed-loop position, including: Geometric verification is performed using the ICP algorithm with an initial angle, and only laser points with the same classification label are matched from high to low according to the weight. When the ICP score corresponding to the candidate frame exceeds the preset threshold, the current candidate frame is confirmed to be a loop frame and the closed-loop position is obtained.

7. The laser SLAM loop detection method based on semantic information as claimed in claim 6, wherein The initial angle is estimated from the following formula: Where A represents the estimated value of the initial angle, Shift_num cur Indicates the descriptor rotation dimension of the current frame, Shift_num his Indicates the descriptor rotation dimension of the historical frame, dim is the descriptor dimension.

Citation Information

Patent Citations

  • Laser SLAM method based on surface line corner feature extraction

    CN111583369A

  • Instant positioning and map construction system and method with semantic perception

    CN111968129A