Positioning method and device of self-moving robot, self-moving robot and medium

The generation of semantic point cloud data through deep learning feature extraction and semantic segmentation technology is solved, and the problem of inaccurate positioning of self-mobile robots in complex environments is achieved, and the positioning of self-mobile robots with high accuracy and robustness is achieved.

CN120279094APending Publication Date: 2025-07-08SHENZHEN MAMMOTION INNOVATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510400318.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing self-mobile robots have low positioning accuracy and robustness in complex environments, mainly because the manual design feature extraction algorithm is sensitive to environmental changes and occlusion and lacks semantic information.

Method used

The feature extraction model and semantic segmentation technology of deep learning are adopted, combined with point cloud registration algorithms, semantic point cloud data is generated, the accuracy and stability of feature extraction are improved, and the feature matching model of deep learning is positioned.

Benefits of technology

Implement accurate positioning of self-mobile robots in complex environments, improve positioning accuracy and robustness, can stably position within centimeter-level accuracy, and support intelligent obstacle avoidance and path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279094A_ABST
    Figure CN120279094A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the field of positioning, and provides a positioning method and device of a self-moving robot, the self-moving robot and a storage medium. The method comprises the steps of collecting a to-be-matched image, obtaining a first feature point set of the to-be-matched image and first descriptors of feature points in the first feature point set through a deep learning feature extraction model, and generating point cloud data related to the to-be-matched image; performing semantic segmentation on the to-be-matched image to obtain semantic information of the to-be-matched image, and combining the semantic information with the point cloud data to generate a semantic point cloud; and matching the semantic point cloud, the first feature point set and the first descriptors of the feature points in the first feature point set with a historical image through a point cloud registration algorithm, and determining the position of the self-moving robot. According to the technical scheme of the embodiment of the invention, the positioning accuracy and robustness of the self-moving robot in a complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of positioning, and in particular, to a positioning method, device, self-mobile robot, and storage medium for a self-mobile robot. Background Art

[0002] Self-mobile robots mainly use visual Simultaneous Localization and Mapping (SLAM) technology to achieve autonomous navigation and build maps simultaneously. Currently, when visual SLAM locates a self-mobile robot, manually designed feature extraction algorithms are usually used for feature extraction. However, these algorithms are sensitive to environmental changes and occlusions, and cannot guarantee the accuracy of the features extracted in complex environments. Moreover, the positioning of self-mobile robots lacks semantic information, resulting in the inability of self-mobile robots to perform accurate positioning in complex environments, and thus the accuracy and robustness of self-mobile robot positioning are relatively low. Therefore, how to improve the accuracy and robustness of self-mobile robot positioning in complex environments is an urgent problem to be solved currently. Summary of the Invention

[0003] Embodiments of the present invention provide a positioning method, device, self-mobile robot, and storage medium for a self-mobile robot, aiming to improve the accuracy and robustness of self-mobile robot positioning in complex environments.

[0004] In a first aspect, embodiments of the present invention provide a positioning method for a self-mobile robot, including:

[0005] Collect a to-be-matched image, obtain a first set of feature points of the to-be-matched image and first descriptors of the feature points in the first set of feature points through a feature extraction model of deep learning, and generate point cloud data related to the to-be-matched image;

[0006] Perform semantic segmentation on the to-be-matched image, obtain semantic information of the to-be-matched image, and combine the semantic information with the point cloud data to generate semantic point cloud;

[0007] Match the semantic point cloud, the first set of feature points, and the first descriptors of the feature points in the first set of feature points with historical images through a point cloud registration algorithm to determine the position of the self-mobile robot.

[0008] In a second aspect, embodiments of the present invention further provide a positioning device for a self-mobile robot. The positioning device includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection communication between the processor and the memory. When the computer program is executed by the processor, the positioning method for the self-mobile robot as described in the first aspect is implemented.

[0009] In a third aspect, an embodiment of the present invention further provides a self-mobile robot, including:

[0010] A machine body;

[0011] A walking component, provided on the machine body, for driving the self-mobile robot to move forward;

[0012] The positioning device as described in the second aspect, provided on the machine body, for determining the position of the self-mobile robot.

[0013] In a fourth aspect, an embodiment of the present invention further provides a storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the positioning method of the self-mobile robot as described in the first aspect.

[0014] The embodiment of the present invention provides a self-mobile robot positioning method, device, self-mobile robot and storage medium. The feature extraction model of deep learning in the embodiment of the present invention is insensitive to environmental changes and occlusion, and is applicable not only to extracting features in a simple environment, but also to extracting features in a complex environment. Therefore, the features extracted from the image to be matched by the feature extraction model of deep learning are more accurate and stable than the features extracted by using a manually designed feature extraction algorithm. And when positioning the self-mobile robot, not only the accurate features extracted from the image to be matched by the feature extraction model of deep learning are used, but also the semantic point cloud combined with the semantic information of the image to be matched is used, so as to be able to achieve accurate positioning of the self-mobile robot in a complex environment, and improve the accuracy and robustness of the self-mobile robot positioning. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 is a schematic flowchart of a self-mobile robot positioning method provided by an embodiment of the present invention;

[0017] Figure 2 is Figure 1 a schematic sub-step flowchart of the self-mobile robot positioning method in

[0018] Figure 3 is Figure 2 a schematic sub-step flowchart of the self-mobile robot positioning method in

[0019] Figure 4 It is a structural schematic block diagram of a positioning device of a self - moving robot provided by an embodiment of the present invention;

[0020] Figure 5 It is a structural schematic block diagram of a self - moving robot provided by an embodiment of the present invention. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0022] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0023] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0024] Self - moving robots mainly use visual Simultaneous Localization and Mapping (SLAM) technology to achieve autonomous navigation and build maps at the same time. Currently, when visual SLAM locates self - moving robots, it usually uses manually designed feature extraction algorithms for feature extraction. However, these algorithms are sensitive to environmental changes and occlusions, and cannot guarantee the accuracy of the features extracted in complex environments. Moreover, the positioning of self - moving robots lacks semantic information, resulting in the inability of self - moving robots to perform accurate positioning in complex environments, and making the accuracy and robustness of self - moving robot positioning relatively low.

[0025] To solve the above problems, an embodiment of the present invention provides a positioning method, device, self-mobile robot, and storage medium for a self-mobile robot. The feature extraction model of deep learning in the embodiment of the present invention is insensitive to environmental changes and occlusions, and is applicable not only to extracting features in a simple environment but also to extracting features in a complex environment. Therefore, the features extracted from the image to be matched by the feature extraction model of deep learning are more accurate and stable than the features extracted by using a manually designed feature extraction algorithm. Moreover, when positioning the self-mobile robot, not only the accurate features extracted from the image to be matched by the feature extraction model of deep learning are used, but also the semantic point cloud combined with the semantic information of the image to be matched is used, so as to realize the precise positioning of the self-mobile robot in a complex environment, and improve the accuracy and robustness of the positioning of the self-mobile robot.

[0026] Among them, the self-mobile robot may include, but is not limited to, a lawn mowing robot, a cleaning robot, a patrol robot, an indoor mobile robot, or an autonomous vehicle, etc. The cleaning robot may be a floor sweeping robot, a mopping robot, or a robot that combines sweeping and mopping.

[0027] The following will describe in detail some embodiments of the present invention with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0028] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a positioning method for a self-mobile robot provided by an embodiment of the present invention.

[0029] As Figure 1 shown, the positioning method of the self-mobile robot includes steps S101 to S103.

[0030] Step S101, collect an image to be matched, obtain a first set of feature points of the image to be matched and first descriptors of the feature points in the first set of feature points through a feature extraction model of deep learning, and generate point cloud data related to the image to be matched.

[0031] In this embodiment, the feature extraction model of deep learning is insensitive to environmental changes and occlusion, and is applicable not only to feature extraction in a simple environment, but also to feature extraction in a complex environment. Therefore, the features extracted from the image to be matched by the feature extraction model of deep learning are more accurate and stable than the features extracted by using a manually designed feature extraction algorithm. Among them, the feature extraction model of deep learning may include a SuperPoint model, a GCNv2 (Geometric Consistency Network version 2) model, or a DISK (DIScrete Keypoints) model, etc. The generation of the point cloud data related to the image to be matched uses the first feature point set of the image to be matched and the first descriptors of the feature points in the first feature point set.

[0032] In some embodiments, the feature extraction model of deep learning includes an encoder, a probability heatmap generation network, a descriptor generation network, a feature point quality evaluation network, and a descriptor calculation layer. Obtaining the first feature point set of the image to be matched and the first descriptors of the feature points in the first feature point set by the feature extraction model of deep learning may include: encoding the image to be matched through the encoder to obtain a first feature map; processing the first feature map through the probability heatmap generation network to obtain a first key point heatmap, where the first key point heatmap is used to describe the probability that each pixel point in the first feature map is a feature point; obtaining a candidate feature point set of the image to be matched according to the first key point heatmap; determining the texture complexity and geometric features of the feature points in the candidate feature point set based on the first key point heatmap; determining the quality scores of the feature points in the candidate feature point set by the feature point quality evaluation network based on the texture complexity and geometric features of the feature points in the candidate feature point set; removing the feature points with quality scores lower than a preset quality score threshold in the candidate feature point set to obtain a first feature point set; processing the first feature map through the descriptor generation network to obtain a descriptor map corresponding to the first feature map, where the descriptor map corresponding to the first feature map includes descriptors of each pixel point in the first feature map; and determining the first descriptors of the feature points in the first feature point set from the descriptor map corresponding to the first feature map through the descriptor calculation layer. Based on the texture complexity and geometric features of the feature points, this embodiment can accurately reflect the distinctiveness of the feature points and then determine their quality scores, and then remove the feature points with quality scores lower than the set threshold, so that the feature points in the obtained first feature point set have high distinctiveness and high quality, improving the accuracy of the extracted feature points and descriptors.

[0033] In some embodiments, the encoder includes a VGG-style encoder, where the VGG-style encoder consists of convolutional layers, linear layers, and non-linear activation layers. The VGG-style encoder can reduce the image size to extract features. For example, the VGG-style encoder reduces the size of the output feature points to 1 / 8 of the input image to be matched through three linear layers.

[0034] In some embodiments, the probability heatmap generation network includes convolutional layers, a non-linear activation layer (softmax), and a reshape layer (Reshape). Among them, processing the first feature map through the probability heatmap generation network to obtain the first key point heatmap may include: performing convolution on the first feature map through the convolutional layer, performing probability prediction processing on the convolved first feature map through the non-linear activation layer to obtain the first key point heatmap with an image size smaller than that of the image to be matched, and finally performing a reshape process on the first key point heatmap with an image size smaller than that of the image to be matched through the reshape layer to obtain the first key point heatmap with the same image size as the image to be matched.

[0035] In some embodiments, determining the quality scores of the feature points in the candidate feature point set based on the texture complexity and geometric features of the feature points in the first feature point set through the feature point quality evaluation network may include: inputting the texture complexity and geometric features of each feature point in the candidate feature point set into the feature point quality evaluation network, and the feature point quality evaluation network performs quality evaluation processing based on the texture complexity and geometric features of each feature point in the first feature point set and outputs the quality scores of each feature point in the candidate feature point set. Among them, the feature point quality evaluation network is obtained by training a neural network model in advance according to multiple training samples, and the training samples may include the texture complexity, geometric features, and labeled quality scores of the feature points.

[0036] In some embodiments, obtaining the feature point set according to the first key point heatmap may include: obtaining the pixel points with the probability of being a feature point greater than or equal to a preset probability threshold from the first key point heatmap as feature points, thereby obtaining the candidate feature point set. Or, among them, obtaining the pixel points with the probability of being a feature point greater than or equal to a preset probability threshold from the first key point heatmap as candidate feature points to obtain a plurality of candidate feature points; using the non-maximum suppression algorithm to screen the plurality of candidate feature points, and further obtaining the candidate feature point set of the image to be matched. Among them, the preset probability threshold can be set based on the actual situation, and the embodiments of the present invention do not make specific limitations on this.

[0037] In some embodiments, the texture complexity may include an entropy value and / or a gradient magnitude variance, and the geometric features may include the scale and direction of feature points. Determining the texture complexity and geometric features of feature points in the candidate feature point set based on the first key point heat map may include: for each feature point in the candidate feature point set, determining the neighborhood of the feature point in the first key point heat map; determining the entropy value and / or gradient magnitude variance of the neighborhood of the feature point according to the probability of each key point in the neighborhood, and determining the entropy value and / or gradient magnitude variance of the neighborhood as the texture complexity of the feature point; calculating the second moment of the neighborhood according to the probability of each key point in the neighborhood, and determining the covariance matrix according to the second moment; decomposing the covariance matrix to obtain eigenvalues and eigenvectors; determining the scale of the feature point according to the eigenvalues and determining the direction of the feature point according to the eigenvectors.

[0038] In some embodiments, the descriptor generation network includes a convolutional layer, an interpolation layer, and a regularization layer. Processing the first feature map through the descriptor generation network to obtain the corresponding descriptor map of the first feature map may include: performing convolutional processing on the first feature map through the convolutional layer, performing bicubic interpolation processing on the first feature map after convolutional processing through the interpolation layer, and performing regularization processing on the first feature map after interpolation processing through the regularization layer to obtain the corresponding descriptor map of the first feature map. The regularization term adopted by the regularization layer includes the L2 norm.

[0039] In some embodiments, the point cloud data related to the image to be matched is generated according to a plurality of feature point matching pairs, and the plurality of feature point matching pairs are obtained by performing feature extraction and feature matching on the image to be matched using a feature extraction model of deep learning and a feature matching model of deep learning. In this embodiment, the feature extraction model of deep learning and the feature matching model of deep learning are not only suitable for feature extraction and feature matching in a simple environment, but also suitable for feature extraction and feature matching in a complex environment. Therefore, by performing feature extraction and feature matching on the image to be matched using the feature extraction model of deep learning and the feature matching model of deep learning, accurate multiple feature point matching pairs can be obtained, thereby ensuring the accuracy of the point cloud data generated based on the multiple feature point matching pairs.

[0040] Specifically, the images to be matched include a first image to be matched and a second image to be matched, and the first image to be matched and the second image to be matched are acquired by a vision sensor at the same moment. Feature point sets of the first image to be matched and descriptors of the feature points in the feature point sets are extracted through a feature extraction model of deep learning, and feature point sets of the second image to be matched and descriptors of the feature points in the feature point sets are extracted through the feature extraction model of deep learning; a feature matching model of deep learning is used to perform feature point matching on the feature point set of the first image to be matched and the feature point set of the second image to be matched based on the descriptors of the feature points in the feature point set of the first image to be matched and the descriptors of the feature points in the feature point set of the second image to be matched, so as to obtain multiple pairs of matched feature points; a triangulation algorithm is used to generate point cloud data related to the images to be matched based on the multiple pairs of feature points. Among them, the feature matching model of deep learning may include a SuperGlue model or an XFeat model.

[0041] In some embodiments, the feature extraction model of deep learning includes a multi-layer feature extraction network of deep learning. The feature point sets of the first image to be matched and the second image to be matched both include multi-layer feature points extracted through the multi-layer feature extraction network, and the descriptors of the feature points in the feature point set of the first image to be matched and the descriptors of the feature points in the feature point set of the second image to be matched both include the descriptors of each layer of feature points extracted through the multi-layer feature extraction network. The feature matching model of deep learning includes a multi-layer feature fusion network and a feature matching network. The multi-layer feature fusion network is used to fuse the feature points in the feature point set and the corresponding descriptors to obtain feature vectors of the feature points in the feature point set, and the feature matching network is used to perform feature point matching on the feature point set of the first image to be matched and the feature point set of the second image to be matched according to the fused feature vectors to obtain multiple feature point matching pairs. In this embodiment, at least shallow detail features and deep semantic features can be extracted through the multi-layer feature extraction network, so that the feature point matching comprehensively considers the detail information and semantic information of the images to be matched, improves the robustness and accuracy of the feature point matching in complex scenarios such as occlusion, scale change or repetitive texture, and has a higher recall rate.

[0042] Step S102: Perform semantic segmentation on the images to be matched, obtain semantic information of the images to be matched, and combine the semantic information with the point cloud data to generate semantic point clouds.

[0043] In this embodiment, a preset semantic segmentation model can be used to perform semantic segmentation on the image to be matched. Among them, the preset semantic segmentation model is obtained by training a neural network model based on a plurality of sample data including sample images and labeled semantic information. For example, the preset semantic segmentation model can include the DeepLab model, the PSPNet model, the Swin Transformer model, the BiSeNet model, the Fast-SCNN model, etc. Among them, the semantic information of the image to be matched can include the semantic labels of each pixel point in the image to be matched.

[0044] In some embodiments, combining the semantic information with the point cloud data to generate a semantic point cloud may include: obtaining the semantic label corresponding to each point in the point cloud data from the semantic information; associatively storing each point in the point cloud data and the semantic label corresponding to each point to obtain a semantic point cloud. Among them, obtaining the semantic label corresponding to each point in the point cloud data from the semantic information may include projecting each point in the point cloud data back to the image to be matched to determine the pixel point matched by each point in the point cloud data in the image to be matched, and obtaining the semantic label corresponding to the pixel point matched by each point in the point cloud data from the semantic information.

[0045] Step S103, match the semantic point cloud, the first set of feature points, and the first descriptors of the feature points in the first set of feature points with the historical image through a point cloud registration algorithm to determine the position of the mobile robot itself.

[0046] The feature extraction model of deep learning in this embodiment is insensitive to environmental changes and occlusions, and is applicable not only to extracting features in a simple environment, but also to extracting features in a complex environment. Therefore, the features extracted from the image to be matched by the feature extraction model of deep learning are more accurate and stable than the features extracted by using a manually designed feature extraction algorithm. Moreover, when positioning the mobile robot itself, not only the accurate features extracted from the image to be matched by the feature extraction model of deep learning are used, but also the semantic point cloud combined with the semantic information of the image to be matched is used, so as to be able to achieve precise positioning of the mobile robot itself in a complex environment, and improve the accuracy and robustness of the positioning of the mobile robot itself.

[0047] In some embodiments, as Figure 2 shown, step S103 includes sub-steps S1031 to S1033.

[0048] Sub-step S1031, update the semantic point cloud map of the mobile robot itself according to the semantic point cloud through a point cloud registration algorithm.

[0049] In this embodiment, the semantic point cloud map of the self-mobile robot is pre-constructed using a semantic SLAM system based on deep learning. Specifically, when the self-mobile robot starts the mapping function, two images to be matched are collected (when the self-mobile robot starts the mapping function, there is no historical image, so loop detection is not involved); the feature extraction model of deep learning is used to obtain the set of feature points and the descriptors of the feature points in the set of feature points of the two images to be matched; the feature matching model of deep learning is used to perform feature point matching on the set of feature points of the two images to be matched based on the descriptors of the feature points in the set of feature points of the two images to be matched, obtaining multiple pairs of matched feature points, and based on the multiple pairs of feature points, initial point cloud data is generated through the triangulation algorithm; the semantic information of any one of the two images to be matched is determined, and this semantic information is combined with the initial point cloud data to obtain the initial semantic point cloud, and the initial semantic point cloud map is constructed based on the initial semantic point cloud through the point cloud reconstruction algorithm.

[0050] After two new images to be matched are collected, in the same manner as generating the initial point cloud data above, new point cloud data is generated based on the two new images to be matched. Then, based on the new point cloud data through the point cloud registration algorithm, the semantic point cloud map is updated. Then, through the loop detection algorithm, it is detected whether there is a target historical image in the historical images (the previously collected images to be matched are historical images) that forms a loop with the new image to be matched. If there is, then according to the relative pose between the image to be matched and the target historical image, the semantic point cloud map is optimized, and then the self-mobile robot is repositioned to obtain the position of the self-mobile robot. In this way, as the self-mobile robot moves, the self-mobile robot continuously collects two new images to be matched, and thus continuously updates the semantic point cloud map, so that a high-precision global semantic point cloud map can be established. Therefore, through this semantic point cloud map, the self-mobile robot can be accurately positioned, achieving a centimeter-level positioning accuracy, and can be stably and accurately positioned in a complex environment, with higher robustness. And through the semantic point cloud map, different obstacles can be recognized. Therefore, accurate positioning, path planning, and intelligent obstacle avoidance of the self-mobile robot can be realized, which has technological advancement and practicality, and has significant technical advantages in the field of self-mobile robots. Among them, the two images to be matched collected can be the left camera and the right camera in a binocular camera respectively collected at the same time, or can be the single camera collected at two very short intervals respectively.

[0051] In some embodiments, updating the semantic point cloud map of the self - moving robot according to the semantic point cloud through a point cloud registration algorithm may include: determining a transformation matrix between the semantic point cloud and the semantic point cloud map of the self - moving robot through the point cloud registration algorithm; and fusing the semantic point cloud into the semantic point cloud map of the self - moving robot according to the transformation matrix to update the semantic point cloud map. Among them, the point cloud configuration algorithm may include the Iterative Closest Point (ICP) algorithm. Based on the current semantic point cloud to update the semantic point cloud map in this embodiment, it can ensure that the semantic point cloud map established by the self - moving robot is a high - precision global semantic point cloud map.

[0052] In some embodiments, determining a transformation matrix between the semantic point cloud and the semantic point cloud map of the self - moving robot through a point cloud registration algorithm may include: obtaining map points with the same semantic labels in the semantic point cloud map from the semantic point cloud map to obtain a set of map points; performing feature point matching on the semantic point cloud and the set of map points to obtain a plurality of matched feature point pairs, and determining a transformation matrix between the semantic point cloud and the semantic point cloud map based on the plurality of matched feature point pairs through the point cloud registration algorithm. In this embodiment, only the map points with the same semantic labels in the semantic point cloud map and the semantic point cloud are used for feature point matching, which can reduce the amount of calculation while ensuring the accuracy of the transformation matrix between the semantic point cloud and the semantic point cloud map.

[0053] In some embodiments, fusing the semantic point cloud into the semantic point cloud map of the self - moving robot according to the transformation matrix to update the semantic point cloud map may include: projecting the semantic point cloud onto the coordinate system where the semantic point cloud map is located according to the transformation matrix to obtain a set of projected map points; dividing the set of projected map points into a first subset of map points and a second subset of map points, the semantic point cloud map includes the first subset of map points and does not include the second subset of map points; updating the semantic labels of the corresponding map points in the semantic point cloud map according to the semantic labels of each map point in the first subset of map points, and expanding the second subset of map points into the semantic point cloud map to update the semantic point cloud map.

[0054] Sub - step S1032: Through a preset closed - loop detection algorithm based on deep learning, based on the first feature point set and the first descriptors of the feature points in the first feature point set, obtain a target historical image that forms a closed loop with the image to be matched from multiple historical images.

[0055] In this embodiment, the preset closed-loop detection algorithm based on deep learning uses a feature extraction model of deep learning to obtain a set of feature points and descriptors, and then uses a feature matching model of deep learning to perform feature point matching based on the obtained descriptors to obtain multiple pairs of matched feature points. Geometric verification is performed based on the multiple feature points. If the geometric verification passes, it is determined that there is a closed loop. Among them, the feature extraction model of deep learning may include a SuperPoint model, a GCNv2 (Geometric Consistency Network version 2) model, or a DISK (DIScrete Keypoints) model, etc., and the feature matching model of deep learning may include a SuperGlue model or an XFeat model.

[0056] Specifically, through the preset closed-loop detection algorithm based on deep learning, obtaining a target historical image that forms a closed loop with the image to be matched from multiple historical images based on the first set of feature points and the first descriptors of the feature points in the first set of feature points may include: using a feature extraction model of deep learning to obtain a second set of feature points of the historical image and a second descriptor of the feature points in the second set of feature points, and then using a feature matching model of deep learning to perform feature point matching on the first set of feature points and the second set of feature points based on the first descriptor and the second descriptor to obtain multiple pairs of matched feature points; performing geometric verification based on the multiple feature points, and in response to the geometric verification failing, repeating the above process to perform closed-loop detection on the image to be matched and the remaining historical images until the geometric verification is successful or the image to be matched has been subjected to closed-loop detection with each historical image; in response to the geometric verification passing, determining the historical image as the target historical image that forms a closed loop with the image to be matched.

[0057] In some embodiments, as Figure 3 shown, sub-step S1032 may include sub-steps S10321 to S10324.

[0058] Sub-step S10321, screening out closed-loop detection images from multiple historical images according to the consistency index between the semantic information of the image to be matched and the semantic information of the historical images.

[0059] In this embodiment, the consistency index between the semantic information of the image to be matched and the semantic information of the historical images may include semantic distribution similarity and semantic similarity. The semantic distribution similarity refers to the similarity of the semantic distribution between the image to be matched and the historical image, and the semantic similarity refers to the similarity between the semantic information of the image to be matched and the semantic information of the historical image.

[0060] In some embodiments, screening out the closed-loop detection image from multiple historical images according to the consistency index between the semantic information of the image to be matched and the semantic information of the historical images may include: screening out a candidate historical image set from multiple historical images according to the semantic distribution similarity between the semantic information of the image to be matched and the semantic information of the historical images, where the semantic distribution similarity between each candidate historical image in the candidate historical image set and the image to be matched is greater than or equal to a preset semantic distribution similarity; determining the closed-loop detection image from the candidate historical image set according to the semantic similarity between the semantic information of the image to be matched and the semantic information of the historical images. For example, selecting the candidate historical image with the highest semantic similarity from the candidate historical image set as the closed-loop detection image. This embodiment can ensure the semantic consistency between the determined closed-loop detection image and the image to be matched, and improve the screening accuracy of the closed-loop detection image.

[0061] Sub-step S10322, obtaining a second set of feature points of the closed-loop detection image and a second descriptor of the feature points in the second set of feature points through a deep learning-based feature extraction model.

[0062] In this embodiment, the deep learning-based feature extraction model may include a SuperPoint model, a GCNv2 (Geometric Consistency Network version 2) model, or a DISK (DIScrete Keypoints) model, etc.

[0063] In some embodiments, the feature extraction model of deep learning may include an encoder, a probability heatmap generation network, a descriptor generation network, a feature point quality evaluation network, and a descriptor calculation layer. Obtaining the second set of feature points of the closed-loop detection image and the second descriptors of the feature points in the second set of feature points through the feature extraction model of deep learning may include: encoding the closed-loop detection image through the encoder to obtain a second feature map; processing the second feature map through the probability heatmap generation network to obtain a second key point heatmap, where the second key point heatmap is used to describe the probability that each pixel point in the second feature map is a feature point; obtaining a candidate set of feature points of the closed-loop detection image according to the second key point heatmap; determining the texture complexity and geometric features of the feature points in the candidate set of feature points of the closed-loop detection image based on the first key point heatmap; determining the quality scores of the feature points in the candidate set of feature points of the closed-loop detection image based on the texture complexity and geometric features of the feature points in the candidate set of feature points of the closed-loop detection image through the feature point quality evaluation network; removing the feature points with quality scores lower than the preset quality score threshold in the candidate set of feature points of the closed-loop detection image to obtain the second set of feature points; processing the second feature map through the descriptor generation network to obtain a descriptor map corresponding to the second feature map, where the descriptor map corresponding to the second feature map includes descriptors of each pixel point in the second feature map; and determining the second descriptors of the feature points in the second set of feature points from the descriptor map corresponding to the second feature map through the descriptor calculation layer.

[0064] Sub-step S10323: Based on the first descriptor and the second descriptor, perform feature point matching on the first set of feature points and the second set of feature points through the feature matching model of deep learning to obtain multiple pairs of matched feature points.

[0065] In this embodiment, the feature matching model of deep learning may include a SuperGlue model or an XFeat model. For example, the feature matching model of deep learning includes a key point encoding layer, a graph neural network with self-attention and cross-attention, a feature point matching layer, and a feature point screening layer.

[0066] In some embodiments, a feature matching model based on deep learning performs feature point matching on a first set of feature points and a second set of feature points based on a first descriptor and a second descriptor. The obtained multiple pairs of matched feature points may include: encoding the position coordinates and corresponding descriptors of the feature points in the first set of feature points through a key point encoding layer to obtain the feature vectors of the feature points in the first set of feature points; encoding the position coordinates and corresponding descriptors of the feature points in the second set of feature points through the key point encoding layer to obtain the feature vectors of the feature points in the second set of feature points; updating the feature vectors of the feature points in the first set of feature points and the feature vectors of the feature points in the second set of feature points through a graph neural network based on self-attention and cross-attention to obtain the target feature vectors of the feature points in the first set of feature points and the target feature vectors of the feature points in the second set of feature points; determining a matching score matrix between the first set of feature points and the second set of feature points based on the target feature vectors of the feature points in the first set of feature points and the target feature vectors of the feature points in the second set of feature points through a feature point matching layer, and performing iterative normalization on the matching score matrix to obtain a doubly stochastic matrix, where the doubly stochastic matrix is used to describe the matching scores of each feature point in the first set of feature points with each feature point in the second set of feature points; screening out the feature point pairs with the matching scores greater than a second preset threshold through a feature point screening layer to obtain multiple pairs of matched feature points.

[0067] Sub-step S10324: performing geometric verification based on multiple feature points, and in response to passing the geometric verification, determining the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched.

[0068] In this embodiment, the closed-loop detection image is screened out through the consistency index between the semantic information of the image to be matched and the semantic information of the historical image, and only the closed-loop detection image and the image to be matched are subjected to closed-loop detection, which greatly reduces the amount of calculation and improves the efficiency of closed-loop detection. Moreover, the closed-loop detection algorithm based on deep learning uses a feature extraction model of deep learning to extract feature points and a feature matching model of deep learning to perform feature point matching. Compared with using manually designed features for closed-loop detection with historical images, the accuracy of closed-loop detection is higher.

[0069] In some embodiments, the positioning method provided by the embodiments of the present invention further includes: in response to failing the geometric verification, storing the image to be matched as a new historical image.

[0070] In some embodiments, for geometric verification based on multiple feature points, and in response to passing the geometric verification, determining the closed-loop detection image as the target historical image forming a closed loop with the image to be matched may include: screening out multiple target feature point pairs from the multiple feature point pairs according to the consistency index of the semantic information between two feature points in the feature point pair, where the semantic information between the two feature points in the target feature point pair is consistent; performing geometric verification based on the multiple target feature point pairs, and in response to passing the geometric verification, determining the closed-loop detection image as the target historical image forming a closed loop with the image to be matched. In this embodiment, by screening out the target feature point pairs with consistent semantic information between the two included feature points and then performing geometric verification, it is possible to avoid the influence of feature point pairs with inconsistent semantic information between the two included feature points on the geometric verification, and further improve the accuracy of closed-loop detection.

[0071] In some embodiments, for geometric verification based on multiple target feature point pairs, and in response to passing the geometric verification, determining the closed-loop detection image as the target historical image forming a closed loop with the image to be matched may include: determining the relative pose between the closed-loop detection image and the image to be matched according to the multiple target feature points; in response to the relative pose satisfying a preset closed-loop condition, determining the closed-loop detection image as the target historical image forming a closed loop with the image to be matched. Among them, the preset closed-loop condition can be set based on the actual situation, and the embodiments of the present invention do not make specific limitations thereto. For example, the preset closed-loop condition includes a preset rotation angle threshold and a preset displacement threshold, that is, when the rotation angle in the relative pose between the closed-loop detection image and the image to be matched is less than or equal to the preset rotation angle threshold, and the displacement in the relative pose is less than or equal to the preset displacement threshold, it is determined that the relative pose between the closed-loop detection image and the image to be matched satisfies the preset closed-loop condition.

[0072] In some embodiments, determining the relative pose between the closed-loop detection image and the image to be matched according to the multiple target feature point pairs may include: using an iterative optimization algorithm to determine the relative pose between the closed-loop detection image and the image to be matched based on the multiple target feature point pairs. For example, using an iterative optimization algorithm to adjust the relative pose between the closed-loop detection image and the image to be matched so that the reprojection error between the multiple target feature point pairs is minimized, thereby obtaining the relative pose between the closed-loop detection image and the image to be matched.

[0073] In some embodiments, obtaining a target historical image that forms a closed loop with the image to be matched from multiple historical images based on a preset closed-loop detection algorithm based on deep learning, based on the first set of feature points and the first descriptors of the feature points in the first set of feature points, may include: obtaining a second set of feature points of the historical image and the second descriptors of the feature points in the second set of feature points through a feature extraction model of deep learning; performing feature point matching on the first set of feature points and the second set of feature points by a feature matching model of deep learning based on the first descriptors and the second descriptors to obtain multiple pairs of matched feature points; screening out multiple target feature point pairs from the multiple pairs of feature points according to the semantic information of the two feature points in the feature point pairs, where the semantic information of the two feature points in the target feature point pairs is consistent; determining the historical image with the largest number of multiple target feature point pairs as the closed-loop detection image, and performing geometric verification according to the multiple target feature point pairs corresponding to the closed-loop detection image; in response to passing the geometric verification, determining the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched. In this embodiment, by screening out the target feature point pairs with consistent semantic information between the two included feature points and determining the historical image with the largest number of target feature point pairs as the closed-loop detection image, the screening accuracy of the closed-loop detection image can be improved. At the same time, only performing closed-loop detection according to the multiple target feature points corresponding to the closed-loop detection image not only improves the efficiency of closed-loop detection, but also can avoid the feature point pairs with inconsistent semantic information between the two included feature points from affecting the geometric verification, further improving the accuracy of closed-loop detection.

[0074] In some embodiments, obtaining a target historical image that forms a closed loop with the image to be matched from multiple historical images based on a preset closed-loop detection algorithm based on deep learning, based on the first set of feature points and the first descriptors of the feature points in the first set of feature points, may include: screening out multiple candidate historical images according to the semantic distribution similarity between the semantic information of the image to be matched and the semantic information of the historical images; obtaining a second set of feature points of the candidate historical images and second descriptors of the feature points in the second set of feature points through a deep learning-based feature extraction model; performing feature point matching on the first set of feature points and the second set of feature points based on the first descriptors and the second descriptors through a deep learning-based feature matching model to obtain multiple pairs of matched feature points; screening out multiple target feature point pairs from the multiple pairs of feature points according to the semantic information of the two feature points in the feature point pairs, where the semantic information of the two feature points in the target feature point pairs is consistent; determining the candidate historical image with the largest number of multiple target feature point pairs as the closed-loop detection image, and performing geometric verification according to the multiple target feature point pairs corresponding to the closed-loop detection image; in response to passing the geometric verification, determining the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched. In this embodiment, by the semantic distribution similarity between the semantic information of the image to be matched and the semantic information of the historical images, multiple candidate historical images are first screened out, and then the closed-loop detection of the image to be matched and the candidate historical images is performed based on the closed-loop detection algorithm based on deep learning, which not only improves the efficiency of closed-loop detection but also improves the accuracy of closed-loop detection.

[0075] In some embodiments, the first set of feature points includes feature points of multiple layers, the first descriptors of the feature points in the first set of feature points include the descriptors of the feature points of each layer, the second set of feature points includes feature points of multiple layers, the second descriptors of the feature points in the second set of feature points include the descriptors of the feature points of each layer, and the feature matching model by deep learning includes a multi-layer feature fusion network and a feature matching network. The feature matching model by deep learning performs feature point matching on the first set of feature points and the second set of feature points based on the first descriptors and the second descriptors, and the obtained multiple pairs of matching feature points may include: fusing the feature points of multiple layers and the corresponding descriptors in the first set of feature points through the multi-layer feature fusion network to obtain the feature vectors of the feature points in the first set of feature points; fusing the feature points of multiple layers and the corresponding descriptors in the second set of feature points through the multi-layer feature fusion network to obtain the feature vectors of the feature points in the second set of feature points; performing feature point matching through the feature matching network based on the feature vectors of the feature points in the first set of feature points and the feature vectors of the feature points in the second set of feature points to obtain multiple pairs of matching feature points. The feature point matching in this embodiment comprehensively considers the multi-layer features of the image to be matched and the historical image (closed-loop detection image), such as detailed features and semantic features, and can match accurate feature point pairs in complex scenarios such as occlusion, scale change, or repetitive texture, improving the accuracy and robustness of closed-loop detection in complex scenarios such as occlusion, scale change, or repetitive texture.

[0076] In some embodiments, the multi-layer feature fusion network includes a key point encoding sub-network and a feature fusion sub-network. Fusing the feature points of multiple layers and the corresponding descriptors in the first set of feature points through the multi-layer feature fusion network to obtain the feature vectors of the feature points in the first set of feature points may include: encoding the position coordinates and the corresponding descriptors of the feature points of each layer in the first set of feature points through the key point encoding sub-network to obtain the feature vectors of the feature points of each layer in the first set of feature points; fusing the feature vectors of the feature points of each layer in the first set of feature points through the multi-layer feature fusion sub-network to obtain the feature vectors of the feature points in the first set of feature points.

[0077] In some embodiments, fusing the feature points of multiple layers and the corresponding descriptors in the second set of feature points through the multi-layer feature fusion network to obtain the feature vectors of the feature points in the second set of feature points may include: encoding the position coordinates and the corresponding descriptors of the feature points of each layer in the second set of feature points through the key point encoding sub-network to obtain the feature vectors of the feature points of each layer in the second set of feature points; fusing the feature vectors of the feature points of each layer in the second set of feature points through the multi-layer feature fusion sub-network to obtain the feature vectors of the feature points in the second set of feature points.

[0078] In some embodiments, the key-point encoding sub-network may include a first key-point encoding layer and a second key-point encoding layer, or the key-point encoding sub-network may include a first key-point encoding layer, a second key-point encoding layer, and a third key-point encoding layer. The first key-point encoding layer is configured to encode the position coordinates and corresponding descriptors of the shallow feature points into feature vectors suitable for processing by the attention graph neural network. The second key-point encoding layer is configured to encode the position coordinates and corresponding descriptors of the deep feature points into feature vectors suitable for processing by the attention graph neural network. The third key-point encoding layer is configured to encode the position coordinates and corresponding descriptors of the middle feature points into feature vectors suitable for processing by the attention graph neural network.

[0079] In some embodiments, the feature matching model based on deep learning includes a multi-layer feature fusion network and a feature matching network. The feature matching model based on deep learning performs feature point matching on the first set of feature points and the second set of feature points based on the first descriptor and the second descriptor, and obtains multiple pairs of matching feature points, including: fusing the multi-layer feature points and corresponding descriptors in the first set of feature points through the multi-layer feature fusion network to obtain feature vectors of the feature points in the first set of feature points; fusing the multi-layer feature points and corresponding descriptors in the second set of feature points through the multi-layer feature fusion network to obtain feature vectors of the feature points in the second set of feature points; performing feature point matching through the feature matching network based on the feature vectors of the feature points in the first set of feature points and the feature vectors of the feature points in the second set of feature points to obtain multiple pairs of matching feature points. The feature point matching in this embodiment comprehensively considers the multi-layer features of the image to be matched and the historical image (closed-loop detection image), such as detail features and semantic features, and can match accurate feature point pairs in complex scenarios such as occlusion, scale change, or repetitive texture, improving the accuracy and robustness of closed-loop detection in complex scenarios such as occlusion, scale change, or repetitive texture.

[0080] In some embodiments, the multi-layer feature fusion network includes a key-point encoding sub-network and a feature fusion sub-network. Fusing the multi-layer feature points and corresponding descriptors in the first set of feature points through the multi-layer feature fusion network to obtain feature vectors of the feature points in the first set of feature points may include: encoding the position coordinates and corresponding descriptors of the feature points in each layer in the first set of feature points through the key-point encoding sub-network to obtain feature vectors of the feature points in each layer in the first set of feature points; fusing the feature vectors of the feature points in each layer in the first set of feature points through the multi-layer feature fusion sub-network to obtain feature vectors of the feature points in the first set of feature points.

[0081] In some embodiments, the multi-layer feature fusion network is used to fuse the feature points and corresponding descriptors of multiple layers in the second set of feature points, and the feature vectors of the feature points in the second set of feature points obtained may include: encoding the position coordinates and corresponding descriptors of the feature points of each layer in the second set of feature points through the key point encoding sub-network to obtain the feature vectors of the feature points of each layer in the second set of feature points; fusing the feature vectors of the feature points of each layer in the second set of feature points through the multi-layer feature fusion sub-network to obtain the feature vectors of the feature points in the second set of feature points.

[0082] In some embodiments, the key point encoding sub-network may include a first key point encoding layer and a second key point encoding layer. The first key point encoding layer is used to encode the position coordinates and corresponding descriptors of the shallow feature points into feature vectors suitable for processing by the attention graph neural network, and the second key point encoding layer is used to encode the position coordinates and corresponding descriptors of the deep feature points into feature vectors suitable for processing by the attention graph neural network.

[0083] For example, encoding the position coordinates and corresponding descriptors of the feature points of each layer in the first set of feature points through the key point encoding sub-network, and the feature vectors of the feature points of each layer in the first set of feature points obtained may include: encoding the position coordinates and corresponding descriptors of the first shallow feature points in the first set of feature points through the first key point encoding layer to obtain the feature vectors of the first shallow feature points in the first set of feature points; encoding the position coordinates and corresponding descriptors of the first deep feature points in the first set of feature points through the second key point encoding layer to obtain the feature vectors of the first deep feature points in the first set of feature points.

[0084] In this embodiment, encoding the position coordinates of the shallow feature points and the descriptors of the shallow feature points (edges, corners, textures, light intensity features, colors, etc.) to obtain feature vectors can accurately represent the detailed features, and encoding the position coordinates of the deep feature points and the descriptors of the deep feature points (semantic information, object categories, object overall shapes, object overall structures, relationships between components, and global context, etc.) to obtain feature vectors can accurately represent the global features. Therefore, after fusing the feature vectors representing the detailed features of the shallow feature points and the feature vectors representing the global features of the deep feature points through the multi-layer feature fusion sub-network, feature vectors that complement the detailed features and global features can be obtained. In this way, the feature vectors obtained through fusion can stably perform feature point matching in different environments, especially in complex environments such as occlusion or obvious illumination changes, still being able to stably perform feature point matching, improving the robustness and accuracy of feature point matching, and thus improving the accuracy and robustness of loop detection.

[0085] In some embodiments, the key point encoding sub-network may include a first key point encoding layer, a second key point encoding layer, and a third key point encoding layer. The first key point encoding layer is used to encode the position coordinates and corresponding descriptors of the shallow feature points into feature vectors suitable for processing by the attention graph neural network. The second key point encoding layer is used to encode the position coordinates and corresponding descriptors of the deep feature points into feature vectors suitable for processing by the attention graph neural network. The third key point encoding layer is used to encode the position coordinates and corresponding descriptors of the middle feature points into feature vectors suitable for processing by the attention graph neural network.

[0086] For example, by encoding the position coordinates and corresponding descriptors of the feature points at each layer in the first set of feature points through the key point encoding sub-network, the feature vectors of the feature points at each layer in the first set of feature points obtained may include: encoding the position coordinates and corresponding descriptors of the first shallow feature points in the first set of feature points through the first key point encoding layer to obtain the feature vectors of the first shallow feature points in the first set of feature points; encoding the position coordinates and corresponding descriptors of the first deep feature points in the first set of feature points through the second key point encoding layer to obtain the feature vectors of the first deep feature points in the first set of feature points; encoding the position coordinates and corresponding descriptors of the first middle feature points in the first set of feature points through the third key point encoding layer to obtain the feature vectors of the first middle feature points in the first set of feature points. In this embodiment, encoding the position coordinates of the shallow feature points and the descriptors of the shallow feature points (edges, corners, textures, light intensity features, colors, etc.) to obtain feature vectors can accurately represent the detailed features. Encoding the position coordinates of the deep feature points and the descriptors of the deep feature points (semantic information, object categories, overall object shapes, overall object structures, relationships between components, and global context, etc.) to obtain feature vectors can accurately represent the global features. Encoding the position coordinates of the middle feature points and the descriptors of the middle feature points (local structures and global information) to obtain feature vectors can accurately represent the local and global features. Therefore, after fusing the feature vectors representing the detailed features of the shallow feature points, the feature vectors representing the global features of the deep feature points, and the feature vectors representing the local and global features of the middle feature points through the multi-layer feature fusion sub-network, feature vectors that complement the detailed features, local features, and global features can be obtained. In this way, the feature vectors obtained through fusion can stably perform feature point matching in different environments, especially in complex environments such as occlusion or obvious lighting changes, still being able to stably perform feature point matching, improving the robustness and accuracy of feature point matching, and thus improving the accuracy and robustness of loop detection.

[0087] In some embodiments, the feature matching network in the feature matching model of deep learning includes a graph neural network based on self-attention and cross-attention, a feature point matching layer, and a feature point screening layer. The feature point matching is performed based on the feature vectors of the feature points in the first feature point set and the feature vectors of the feature points in the second feature point set through the feature matching network. The obtained multiple matched feature point pairs may include: updating the feature vectors of the feature points in the first feature point set and the feature vectors of the feature points in the second feature point set through the graph neural network based on self-attention and cross-attention to obtain the target feature vectors of the feature points in the first feature point set and the target feature vectors of the feature points in the second feature point set; determining a matching score matrix between the first feature point set and the second feature point set based on the target feature vectors of the feature points in the first feature point set and the target feature vectors of the feature points in the second feature point set through the feature point matching layer, and performing iterative normalization on the matching score matrix to obtain a doubly stochastic matrix, where the doubly stochastic matrix is used to describe the matching scores of each feature point in the first feature point set with each feature point in the second feature point set; screening out the feature point pairs with the matching scores greater than the second preset threshold through the feature point screening layer to obtain multiple matched feature point pairs.

[0088] In some embodiments, the first feature point set includes a plurality of first shallow feature points and a plurality of first deep feature points. The first descriptors of the feature points in the first feature point set include the descriptors of the first shallow feature points and the descriptors of the first deep feature points. The first shallow feature points and the corresponding descriptors are used to describe the detail information of the image to be matched, and the first deep feature points and the corresponding descriptors are used to describe the semantic information of the image to be matched. Among them, the first shallow feature points in the first feature point set are obtained based on the feature map (shallow feature map) of the first resolution of the image to be matched, and the first deep feature points in the first feature point set are obtained based on the feature map (deep feature map) of the second resolution of the image to be matched, where the first resolution is greater than the second resolution. The second feature point set includes a plurality of second shallow feature points and a plurality of second deep feature points. The second descriptors of the feature points in the second feature point set include the descriptors of the second shallow feature points and the descriptors of the second deep feature points. The second shallow feature points and the corresponding descriptors are used to describe the detail information of the historical image, and the second deep feature points and the corresponding descriptors are used to describe the semantic information of the historical image. Among them, the second shallow feature points in the second feature point set are obtained based on the feature map (shallow feature map) of the first resolution of the historical image (closed-loop detection image), and the second deep feature points in the second feature point set are obtained based on the feature map (deep feature map) of the second resolution of the historical image.

[0089] In some embodiments, the multi-layer feature extraction network of deep learning includes a multi-scale encoder, a first feature extraction network (shallow feature extraction network), and a second feature extraction network (deep feature extraction network). The first feature extraction network and the second feature extraction network have the same structure but different model parameters. Both the first feature extraction network and the second feature extraction network include a probability heat map generation network, a descriptor generation network, a feature point quality evaluation network, and a descriptor calculation layer. The multi-scale encoder includes a first encoding layer (for generating a feature map with the first resolution, i.e., a shallow feature map) and a second encoding layer (for generating a feature map with the second resolution, i.e., a deep feature map). The number of convolutional layers used in the first encoding layer and the second encoding layer can be different or the same. In this embodiment, the shallow feature extraction network and the deep feature extraction network can extract shallow detail features (edges, corners, textures, light intensity features, colors, etc.) and deep abstract features (semantic information, object categories, overall object shapes, overall object structures, relationships between components, and global context, etc.). The shallow detail features are sensitive to interferences such as illumination changes and noise, while the deep abstract features are insensitive to interferences such as illumination changes and noise. Thus, in different environments, feature point matching can be stably performed through the extracted shallow detail features and deep abstract features. Especially in complex environments with occlusions or obvious illumination changes, feature point matching can still be stably performed, improving the robustness and accuracy of feature point matching.

[0090] In some embodiments, obtaining the first set of feature points of the image to be matched and the first descriptors of the feature points in the first set of feature points through the multi-layer feature extraction network of deep learning may include: encoding the image to be matched through the first encoding layer to obtain a feature map with the first resolution (shallow feature map); encoding the feature map with the first resolution through the second encoding layer to obtain a feature map with the second resolution (deep feature map); performing feature point extraction and descriptor generation processing on the feature map with the first resolution through the first feature extraction network to obtain a plurality of first shallow feature points and corresponding descriptors; performing feature point extraction and descriptor generation processing on the feature map with the second resolution through the second feature extraction network to obtain a plurality of first deep feature points and corresponding descriptors. The feature map with the first resolution in this embodiment includes low-level visual features such as edges, corners, textures, light intensity features, and colors, which are sensitive to details and can ensure the accurate positioning of key points. The feature map with the second resolution includes abstract high-level features such as semantic information, object categories, overall object shapes, overall object structures, relationships between components, and global context, which are insensitive to interferences such as illumination conditions and noise. Thus, in different environments, feature point matching can be stably performed through the extracted shallow detail features and deep abstract features. Especially in complex environments with occlusions or obvious illumination changes, feature point matching can still be stably performed, improving the robustness and accuracy of feature point matching.

[0091] In some embodiments, the process of extracting feature points and generating descriptors based on the feature map of the first resolution by the first feature extraction network to obtain a plurality of first shallow feature points and corresponding descriptors may include: processing the feature map of the first resolution by the probability heat map generation network in the first feature extraction network to obtain a key point heat map of the first resolution (shallow key point heat map); obtaining a first candidate feature point set based on the key point heat map of the first resolution through the first feature point screening layer in the first feature extraction network; determining the texture information and geometric features of each feature point in the first candidate feature point set according to the key point heat map of the first resolution; obtaining the quality score of each feature point in the first candidate feature point set based on the texture information and geometric features of each feature point in the first candidate feature point set through the feature point quality evaluation network in the first feature extraction network; removing the feature points with quality scores lower than the preset quality score threshold to obtain a plurality of first shallow feature points; processing the feature map of the first resolution by the descriptor generation network in the first feature extraction network to obtain a descriptor map of the first resolution, where the descriptor map of the first resolution includes descriptors of each pixel point in the feature map of the first resolution; and obtaining the descriptor of each first shallow feature point from the descriptor map of the first resolution.

[0092] In some embodiments, the process of extracting feature points and generating descriptors based on the feature map of the second resolution by the second feature extraction network to obtain a plurality of first deep feature points and corresponding descriptors may include: processing the feature map of the second resolution by the probability heat map generation network in the second feature extraction network to obtain a key point heat map of the second resolution (deep key point heat map); obtaining a second candidate feature point set based on the key point heat map of the second resolution through the first feature point screening layer in the second feature extraction network; determining the texture information and geometric features of each feature point in the second candidate feature point set according to the key point heat map of the second resolution; obtaining the quality score of each feature point in the second candidate feature point set based on the texture information and geometric features of each feature point in the second candidate feature point set through the feature point quality evaluation network in the second feature extraction network; removing the feature points with quality scores lower than the preset quality score threshold to obtain a plurality of first deep feature points; processing the feature map of the second resolution by the descriptor generation network in the second feature extraction network to obtain a descriptor map of the second resolution, where the descriptor map of the second resolution includes descriptors of each pixel point in the feature map of the second resolution; and obtaining the descriptor of each first deep feature point from the descriptor map of the second resolution.

[0093] In some embodiments, the first set of feature points includes a plurality of first shallow feature points, a plurality of first middle-layer feature points, and a plurality of first deep feature points. The descriptors of the feature points in the first set of feature points include the descriptors of the first shallow feature points, the descriptors of the first middle-layer feature points, and the descriptors of the first deep feature points. The first middle-layer feature points and the corresponding descriptors are used to balance the detail information and semantic information of the image to be matched. Among them, the first middle-layer feature points in the first set of feature points are obtained based on the feature map (middle-layer feature map) of the third resolution of the image to be matched. The first resolution is greater than the third resolution, and the third resolution is greater than the second resolution. The second set of feature points includes a plurality of second shallow feature points, a plurality of second middle-layer feature points, and a plurality of second deep feature points. The descriptors of the feature points in the second set of feature points include the descriptors of the second shallow feature points, the descriptors of the second middle-layer feature points, and the descriptors of the second deep feature points. The second middle-layer feature points and the corresponding descriptors are used to balance the detail information and semantic information of the image to be detected. Among them, the second middle-layer feature points in the second set of feature points are obtained based on the feature map (middle-layer feature map) of the third resolution of the historical image.

[0094] In some embodiments, the multi-layer feature extraction network includes a multi-scale encoder, a first feature extraction network (shallow feature extraction network), a second feature extraction network (deep feature extraction network), and a third feature extraction network (middle-layer feature extraction network). The first feature extraction network, the second feature extraction network, and the third feature extraction network have the same structure but different model parameters. The first feature extraction network, the second feature extraction network, and the third feature extraction network all include a probability heatmap generation network, a descriptor generation network, a feature point quality assessment network, and a descriptor calculation layer. The multi-scale encoder includes a first encoding layer (for generating features of the first resolution, i.e., shallow feature map), a second encoding layer (for generating features of the second resolution, i.e., deep feature map), and a third encoding layer (for generating features of the third resolution, i.e., middle-layer feature map). The number of convolutional layers used in the first encoding layer, the second encoding layer, and the third encoding layer can be different or the same. In this embodiment, the shallow detail features include edges, corners, textures, light intensity features, colors, etc. They are sensitive to details and can ensure the accurate positioning of key points. The middle-layer coarse-grained features include information such as shapes, edges, local structures, colors, etc., containing more texture and shape information, and still affect feature matching. The deep abstract features include semantic information, object categories, object overall shapes, object overall structures, relationships between components, and global context, etc. They are insensitive to interferences such as illumination changes and noise. In this way, in different environments, feature point matching can be stably performed through the extracted shallow detail features and deep abstract features, especially in complex environments with occlusion or obvious illumination changes, and the robustness and accuracy of feature point matching are improved.

[0095] In some embodiments, obtaining the first set of feature points of the image to be matched and the first descriptors of the feature points in the first set of feature points through a multi-layer feature extraction network of deep learning may include: encoding the image to be matched through a first encoding layer to obtain a feature map of the first resolution (shallow feature map); encoding the feature map of the first resolution through a third encoding layer to obtain a feature map of the third resolution (mid-level feature map); encoding the feature map of the third resolution through a second encoding layer to obtain a feature map of the second resolution (deep feature map); performing feature point extraction and descriptor generation processing on the feature map of the first resolution through a first feature extraction network to obtain a plurality of first shallow feature points and corresponding descriptors; performing feature point extraction and descriptor generation processing on the feature map of the second resolution through a second feature extraction network to obtain a plurality of first deep feature points and corresponding descriptors; performing feature point extraction and descriptor generation processing on the feature map of the third resolution through a third feature extraction network to obtain a plurality of first mid-level feature points and corresponding descriptors.

[0096] In some embodiments, performing feature point extraction and descriptor generation processing on the feature map of the third resolution through a third feature extraction network to obtain a plurality of first mid-level feature points and corresponding descriptors may include: processing the feature map of the third resolution through a probability heat map generation network in the third feature extraction network to obtain a key point heat map of the third resolution (mid-level key point heat map); obtaining a third set of candidate feature points through a first feature point screening layer in the third feature extraction network based on the key point heat map of the third resolution; determining the texture information and geometric features of each feature point in the third set of candidate feature points according to the key point heat map of the third resolution; obtaining the quality score of each feature point in the third set of candidate feature points through a feature point quality evaluation network in the third feature extraction network based on the texture information and geometric features of each feature point in the third set of candidate feature points; removing the feature points with quality scores lower than a preset quality score threshold to obtain a plurality of first mid-level feature points; processing the feature map of the third resolution through a descriptor generation network in the third feature extraction network to obtain a descriptor map of the third resolution, and the descriptor map of the third resolution includes descriptors of each pixel point in the feature map of the third resolution; obtaining the descriptors of each first mid-level feature point from the descriptor map of the third resolution.

[0097] It should be noted that the specific implementation manner of obtaining the second set of feature points of the historical image (closed-loop detection image) and the second descriptors of the feature points in the second set of feature points through a multi-layer feature extraction network of deep learning may refer to the specific implementation manner of obtaining the first set of feature points of the image to be matched and the first descriptors of the feature points in the first set of feature points through a multi-layer feature extraction network of deep learning as described above, and will not be elaborated here.

[0098] Sub-step S1033: Determine the position of the self-mobile robot according to the relative pose between the image to be matched and the target historical image and the updated semantic point cloud map.

[0099] Based on the closed-loop detection algorithm of deep learning in this embodiment, closed-loop detection is performed between the first feature point set and the first descriptors of the feature points in the first feature point set and the historical images, and then the target historical image that forms a closed loop with the image to be matched is obtained from multiple historical images. Compared with performing closed-loop detection using manually designed features and historical images, the accuracy of closed-loop detection is higher. In this way, using the relative pose between the target historical image obtained through closed-loop detection and the image to be matched can more accurately achieve the positioning of the self-mobile robot in a complex environment, further improving the accuracy and robustness of the positioning of the self-mobile robot.

[0100] In some embodiments, determining the position of the self-mobile robot according to the relative pose between the image to be matched and the target historical image and the updated semantic point cloud map may include: optimizing the updated semantic point cloud map according to the relative pose between the image to be matched and the target historical image; determining the position of the self-mobile robot according to the globally optimized semantic point cloud map. Among them, a pose graph optimization algorithm can be used to optimize the updated semantic point cloud map based on the relative pose between the image to be matched and the target historical image.

[0101] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of the structure of a positioning device for a self-mobile robot provided by an embodiment of the present invention.

[0102] As Figure 4 shown, the positioning device 130 of the self-mobile robot includes a processor 131 and a memory 132, and the processor 131 and the memory 132 are connected through a bus 133, and this bus is, for example, an I2C (Inter-integrated Circuit) bus.

[0103] Specifically, the processor 131 is used to provide computing and control capabilities to support the operation of the entire self-mobile robot. The processor 131 can be a Central Processing Unit (CPU), and this processor 131 can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.

[0104] Specifically, the memory 132 can be a Flash chip, Read-Only Memory (ROM), magnetic disk, optical disc, USB flash drive, or mobile hard disk, etc.

[0105] Those skilled in the art can understand that Figure 4 the structure shown in [the figure] is only a block diagram of some structures related to the solution of the embodiment of the present invention, and does not constitute a limitation on the positioning device of the self-mobile robot to which the solution of the embodiment of the present invention is applied. The specific positioning device of the self-mobile robot may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0106] Among them, the processor is used to run the computer program stored in the memory and implement any one of the positioning methods of the self-mobile robot provided by the embodiment of the present invention when executing the computer program.

[0107] In some embodiments, the processor is used to run the computer program stored in the memory and implement the following steps when executing the computer program:

[0108] Collect the image to be matched, obtain the first set of feature points of the image to be matched and the first descriptors of the feature points in the first set of feature points through a feature extraction model of deep learning, and generate point cloud data related to the image to be matched;

[0109] Perform semantic segmentation on the image to be matched, obtain the semantic information of the image to be matched, and combine the semantic information with the point cloud data to generate semantic point cloud;

[0110] Match the semantic point cloud, the first set of feature points, and the first descriptors of the feature points in the first set of feature points with a historical image through a point cloud registration algorithm to determine the position of the self - moving robot.

[0111] In some embodiments, the point cloud data is generated based on multiple feature point matching pairs, and the multiple feature point matching pairs are obtained by performing feature extraction and feature matching on the image to be matched using a feature extraction model of deep learning and a feature matching model of deep learning.

[0112] In some embodiments, when the processor implements matching the semantic point cloud, the first set of feature points, and the first descriptors of the feature points in the first set of feature points with a historical image through a point cloud registration algorithm to determine the position of the self - moving robot, it is used to implement:

[0113] Update the semantic point cloud map of the self - moving robot according to the semantic point cloud through a point cloud registration algorithm;

[0114] Obtain a target historical image that forms a closed loop with the image to be matched from multiple historical images based on the first set of feature points and the first descriptors of the feature points in the first set of feature points through a preset closed - loop detection algorithm based on deep learning;

[0115] Determine the position of the self - moving robot according to the relative pose between the image to be matched and the target historical image and the updated semantic point cloud map.

[0116] In some embodiments, when the processor implements obtaining a target historical image that forms a closed loop with the image to be matched from multiple historical images based on the first set of feature points and the first descriptors of the feature points in the first set of feature points through a preset closed - loop detection algorithm based on deep learning, it is used to implement:

[0117] Screen out closed - loop detection images from multiple historical images according to the consistency index between the semantic information of the image to be matched and the semantic information of the historical images;

[0118] Obtain a second set of feature points of the closed - loop detection image and second descriptors of the feature points in the second set of feature points through the feature extraction model of deep learning;

[0119] Perform feature point matching on the first set of feature points and the second set of feature points based on the first descriptor and the second descriptor through a feature matching model of deep learning to obtain multiple matched feature point pairs;

[0120] Perform geometric verification based on the multiple feature points, and in response to passing the geometric verification, determine the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched.

[0121] In some embodiments, when the processor implements the consistency index between the semantic information of the image to be matched and the semantic information of the historical images and filters out the closed-loop detection images from the multiple historical images, it is used to implement:

[0122] Filter out a set of candidate historical images from the multiple historical images according to the semantic distribution similarity between the semantic information of the image to be matched and the semantic information of the historical images;

[0123] Determine the closed-loop detection image from the set of candidate historical images according to the semantic similarity between the semantic information of the image to be matched and the semantic information of the historical images.

[0124] In some embodiments, when the processor implements geometric verification based on the multiple feature points and, in response to passing the geometric verification, determines the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched, it is used to implement:

[0125] Filter out a plurality of target feature point pairs from the multiple feature point pairs according to the consistency index of the semantic information between two feature points in the feature point pair, and the semantic information between the two feature points in the target feature point pair is consistent;

[0126] Perform geometric verification based on the multiple target feature point pairs, and in response to passing the geometric verification, determine the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched.

[0127] In some embodiments, when the processor implements obtaining a target historical image that forms a closed loop with the image to be matched from multiple historical images through a preset deep learning-based closed-loop detection algorithm based on the first feature point set and the first descriptors of the feature points in the first feature point set, it is used to implement:

[0128] Obtain a second feature point set of the historical image and second descriptors of the feature points in the second feature point set through the feature extraction model of the deep learning;

[0129] Perform feature point matching on the first feature point set and the second feature point set through a deep learning-based feature matching model based on the first descriptors and the second descriptors to obtain a plurality of matched feature point pairs;

[0130] Filter out a plurality of target feature point pairs from the plurality of feature point pairs according to the semantic information of two feature points in the feature point pair, where the semantic information of the two feature points in the target feature point pair is consistent;

[0131] Determine the historical image with the largest number of the plurality of target feature point pairs as the closed-loop detection image, and perform geometric verification according to the plurality of target feature point pairs corresponding to the closed-loop detection image;

[0132] In response to the geometric verification passing, determine the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched.

[0133] In some embodiments, when the processor implements obtaining the second feature point set of the historical image and the second descriptors of the feature points in the second feature point set through the feature extraction model of deep learning, it is used to implement:

[0134] Filter out a plurality of candidate historical images according to the semantic distribution similarity between the semantic information of the image to be matched and the semantic information of the historical image;

[0135] Obtain the second feature point set of the candidate historical image and the second descriptors of the feature points in the second feature point set through the feature extraction model of deep learning;

[0136] When the processor implements determining the historical image with the largest number of the plurality of target feature point pairs as the closed-loop detection image, it is used to implement:

[0137] Determine the candidate historical image with the largest number of the plurality of target feature point pairs as the closed-loop detection image.

[0138] In some embodiments, the feature matching model of deep learning includes a multi-layer feature fusion network and a feature matching network. When the processor implements feature point matching between the first feature point set and the second feature point set based on the first descriptor and the second descriptor through the feature matching model of deep learning to obtain a plurality of matched feature point pairs, it is used to implement:

[0139] Fuse the multi-layer feature points and corresponding descriptors in the first feature point set through the multi-layer feature fusion network to obtain the feature vectors of the feature points in the first feature point set;

[0140] Fuse the multi-layer feature points and corresponding descriptors in the second feature point set through the multi-layer feature fusion network to obtain the feature vectors of the feature points in the second feature point set;

[0141] Based on the feature vectors of the feature points in the first feature point set and the feature vectors of the feature points in the second feature point set, the feature point matching is performed by the feature matching network to obtain multiple pairs of matched feature points.

[0142] In some embodiments, when the processor implements updating the semantic point cloud map of the self-mobile robot according to the semantic point cloud through a point cloud registration algorithm, it is used to implement:

[0143] Determine the transformation matrix between the semantic point cloud and the semantic point cloud map of the self-mobile robot through a point cloud registration algorithm;

[0144] According to the transformation matrix, fuse the semantic point cloud into the semantic point cloud map of the self-mobile robot to update the semantic point cloud map.

[0145] In some embodiments, the feature extraction model of deep learning includes an encoder, a probability heatmap generation network, a descriptor generation network, a feature point quality evaluation network, and a descriptor calculation layer. When the processor implements feature extraction through deep learning, it is used to implement:

[0146] Encode the image to be matched through the encoder to obtain a first feature map;

[0147] Process the first feature map through the probability heatmap generation network to obtain a first key point heatmap, and the first key point heatmap is used to describe the probability that each pixel point in the first feature map is a feature point;

[0148] According to the first key point heatmap, obtain the candidate feature point set of the image to be matched;

[0149] Based on the first key point heatmap, determine the texture complexity and geometric features of the feature points in the candidate feature point set;

[0150] Based on the texture complexity and geometric features of the feature points in the candidate feature point set, determine the quality scores of the feature points in the candidate feature point set through the feature point quality evaluation network;

[0151] Eliminate the feature points in the candidate feature point set whose quality scores are lower than the preset quality score threshold to obtain the first feature point set;

[0152] Process the first feature map through the descriptor generation network to obtain the descriptor map corresponding to the first feature map, and the descriptor map corresponding to the first feature map includes the descriptors of each pixel point in the first feature map;

[0153] Determine the first descriptors of the feature points in the first set of feature points from the corresponding descriptor maps of the first feature maps through the described descriptor calculation layer.

[0154] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the positioning device of the self-mobile robot described above can refer to the corresponding process in the foregoing embodiment of the positioning method of the self-mobile robot, and will not be elaborated here.

[0155] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of the structure of a self-mobile robot provided by an embodiment of the present invention.

[0156] As Figure 5 shown, the self-mobile robot 100 includes a machine body 110, a traveling assembly 120, and a positioning device 130. Among them, the traveling assembly 120 is provided on the machine body 110 and is used to drive the self-mobile robot 100 to travel. The positioning device 130 is provided inside the machine body 110 ( Figure 5 not shown) and is used to implement any one of the positioning methods provided in the specification of the embodiment of the present invention.

[0157] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the self-mobile robot described above can refer to the corresponding process in the foregoing embodiment of the positioning method of the self-mobile robot, and will not be elaborated here.

[0158] An embodiment of the present invention further provides a storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement any one of the positioning methods of the self-mobile robot provided in the specification of the embodiment of the present invention.

[0159] Among them, the storage medium can be the internal storage unit of the self-mobile robot described in the foregoing embodiment, such as the hard disk or memory of the self-mobile robot. The storage medium can also be an external storage device of the self-mobile robot, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the self-mobile robot.

[0160] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0161] It should be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this article, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or system. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article, or system comprising the element.

[0162] The serial numbers of the embodiments of the present invention above are only for description and do not represent the advantages or disadvantages of the embodiments. As described above, the specific embodiments of the present invention are only described, but the protection scope of the present invention is not limited thereto. Any person skilled in the art in the technical field disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A positioning method for a self - moving robot, characterized in that, Including: Collect the image to be matched, obtain the first set of feature points of the image to be matched and the first descriptors of the feature points in the first set of feature points through a feature extraction model of deep learning, and generate point cloud data related to the image to be matched; Perform semantic segmentation on the image to be matched, obtain the semantic information of the image to be matched, and combine the semantic information with the point cloud data to generate a semantic point cloud; Match the semantic point cloud, the first set of feature points, and the first descriptors of the feature points in the first set of feature points with historical images through a point cloud registration algorithm to determine the position of the self-mobile robot.

2. The positioning method according to claim 1, characterized in that, The point cloud data is generated based on multiple feature point matching pairs, and the multiple feature point matching pairs are obtained by performing feature extraction and feature matching on the image to be matched using a feature extraction model of deep learning and a feature matching model of deep learning.

3. The positioning method according to claim 1, wherein The step of matching the feature points of the semantic point cloud, the first set of feature points, and the first descriptors of the feature points in the first set of feature points with historical images through a point cloud registration algorithm to determine the position of the self-mobile robot includes: Update the semantic point cloud map of the self-mobile robot according to the semantic point cloud through a point cloud registration algorithm; Obtain a target historical image that forms a closed loop with the image to be matched from multiple historical images based on the first set of feature points and the first descriptors of the feature points in the first set of feature points through a preset closed-loop detection algorithm based on deep learning; Determine the position of the self-mobile robot according to the relative pose between the image to be matched and the target historical image and the updated semantic point cloud map.

4. The positioning method according to claim 3, wherein, The step of obtaining a target historical image that forms a closed loop with the image to be matched from multiple historical images based on the first set of feature points and the first descriptors of the feature points in the first set of feature points through a preset closed-loop detection algorithm based on deep learning includes: Screen out closed-loop detection images from multiple historical images according to the consistency index between the semantic information of the image to be matched and the semantic information of the historical images; Obtain the second set of feature points of the closed-loop detection image and the second descriptors of the feature points in the second set of feature points through the feature extraction model of deep learning; Perform feature point matching on the first set of feature points and the second set of feature points based on the first descriptor and the second descriptor through a feature matching model of deep learning to obtain multiple matched feature point pairs; Perform geometric verification according to the multiple feature points, and in response to the geometric verification passing, determine the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched.

5. The positioning method according to claim 4, characterized in that, The step of screening out closed-loop detection images from multiple historical images according to the consistency index between the semantic information of the image to be matched and the semantic information of the historical images includes: Screen out a set of candidate historical images from multiple historical images according to the semantic distribution similarity between the semantic information of the image to be matched and the semantic information of the historical images; Determine a closed-loop detection image from the set of candidate historical images according to the semantic similarity between the semantic information of the image to be matched and the semantic information of the historical image.

6. The positioning method according to claim 4, wherein The geometric verification based on the multiple feature points, and in response to the geometric verification passing, determining the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched, includes: According to the consistency index of the semantic information between two feature points in the feature point pairs, filter out multiple target feature point pairs from the multiple feature point pairs, and the semantic information between the two feature points in the target feature point pairs is consistent; Perform geometric verification based on the multiple target feature point pairs, and in response to the geometric verification passing, determine the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched.

7. The positioning method according to claim 3, wherein The method of obtaining a target historical image that forms a closed loop with the image to be matched from multiple historical images through a preset deep learning-based closed-loop detection algorithm, based on the first feature point set and the first descriptors of the feature points in the first feature point set, includes: Obtain a second feature point set of the historical image and second descriptors of the feature points in the second feature point set through the feature extraction model of the deep learning; Perform feature point matching on the first feature point set and the second feature point set based on the first descriptor and the second descriptor through a feature matching model of the deep learning to obtain multiple matched feature point pairs; According to the semantic information of two feature points in the feature point pairs, filter out multiple target feature point pairs from the multiple feature point pairs, and the semantic information between the two feature points in the target feature point pairs is consistent; Determine the historical image with the largest number of the multiple target feature point pairs as the closed-loop detection image, and perform geometric verification according to the multiple target feature point pairs corresponding to the closed-loop detection image; In response to the geometric verification passing, determine the closed-loop detection image as the target historical image that forms a closed loop with the image to be matched.

8. The positioning method according to claim 7, characterized in that, The method of obtaining a second feature point set of the historical image and second descriptors of the feature points in the second feature point set through the feature extraction model of the deep learning includes: Filter out multiple candidate historical images according to the semantic distribution similarity between the semantic information of the image to be matched and the semantic information of the historical image; Obtain a second feature point set of the candidate historical image and second descriptors of the feature points in the second feature point set through the feature extraction model of the deep learning; The method of determining the historical image with the largest number of the multiple target feature point pairs as the closed-loop detection image includes: Determine the candidate historical image with the largest number of the multiple target feature point pairs as the closed-loop detection image.

9. The positioning method according to claim 4 or 7, characterized in that, The feature matching model of the deep learning includes a multi-layer feature fusion network and a feature matching network. The method of performing feature point matching on the first feature point set and the second feature point set based on the first descriptor and the second descriptor through the feature matching model of the deep learning to obtain multiple matched feature point pairs includes: Fuse the multi - layer feature points and corresponding descriptors in the first feature point set through the multi - layer feature fusion network to obtain the feature vectors of the feature points in the first feature point set; Fuse the multi - layer feature points and corresponding descriptors in the second feature point set through the multi - layer feature fusion network to obtain the feature vectors of the feature points in the second feature point set; Based on the feature vectors of the feature points in the first feature point set and the feature vectors of the feature points in the second feature point set, perform feature point matching through the feature matching network to obtain multiple pairs of matched feature points.

10. The positioning method according to claim 3, wherein, The updating of the semantic point cloud map of the self - moving robot according to the semantic point cloud by the point cloud registration algorithm includes: Determine the transformation matrix between the semantic point cloud and the semantic point cloud map of the self - moving robot through the point cloud registration algorithm; According to the transformation matrix, fuse the semantic point cloud into the semantic point cloud map of the self - moving robot to update the semantic point cloud map.

11. The positioning method according to any one of claims 1-8 or claim 10, characterized in that, The feature extraction model of deep learning includes an encoder, a probability heat map generation network, a descriptor generation network, a feature point quality evaluation network, and a descriptor calculation layer. The obtaining of the first feature point set of the image to be matched and the first descriptor of the feature points in the first feature point set through the feature extraction model of deep learning includes: Encode the image to be matched through the encoder to obtain a first feature map; Process the first feature map through the probability heat map generation network to obtain a first key point heat map, where the first key point heat map is used to describe the probability that each pixel point in the first feature map is a feature point; According to the first key point heat map, obtain the candidate feature point set of the image to be matched; Based on the first key point heat map, determine the texture complexity and geometric features of the feature points in the candidate feature point set; Based on the texture complexity and geometric features of the feature points in the candidate feature point set, determine the quality scores of the feature points in the candidate feature point set through the feature point quality evaluation network; Eliminate the feature points with quality scores lower than the preset quality score threshold in the candidate feature point set to obtain the first feature point set; Process the first feature map through the descriptor generation network to obtain the corresponding descriptor map of the first feature map, and the corresponding descriptor map of the first feature map includes the descriptors of each pixel point in the first feature map; Determine the first descriptors of the feature points in the first feature point set from the corresponding descriptor map of the first feature map through the descriptor calculation layer.

12. A positioning device for a self - moving robot, characterized in that, The positioning device includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the computer program is executed by the processor, it realizes the positioning method of the self - moving robot as described in any one of claims 1 to 11.

13. A self - moving robot, characterized in that, Including: Machine body; A walking component, arranged on the machine body, for driving the self - moving robot to move forward; The positioning device described in claim 12 is provided on the machine body and is used to determine the position of the self-mobile robot.

14. A storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the positioning method of the self-mobile robot described in any one of claims 1 to 11.