Simultaneous positioning and mapping method and related device

By extracting the visual data, obtaining multimodal features and matching features, the problem of insufficient robustness of existing SLAM technologies is solved, and a more stable and accurate positioning and mapping effect is achieved.

CN120176645APending Publication Date: 2025-06-20HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311743423.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the process of positioning and drawing, the existing SLAM technology relies on low-level visual features such as simple points, lines and surfaces, resulting in insufficient robustness.

Method used

By extracting the visual data feature, multimodal features are obtained, including visual features and auxiliary features corresponding to the visual data, and feature matching is performed to improve the robustness of SLAM.

Benefits of technology

The use of multimodal features improves the robustness of SLAM and enhances the stability and accuracy of positioning and mapping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120176645A_ABST
    Figure CN120176645A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a simultaneous positioning and mapping method and related device.The simultaneous positioning and mapping method comprises the steps that after visual data of a to-be-positioned target is obtained, feature extraction is conducted on the visual data, multi-modal features corresponding to the visual data are obtained, and the multi-modal features comprise visual features corresponding to the visual data and auxiliary features; the auxiliary features and the visual features are different modal features of the visual data; then feature matching is performed on the multi-modal features to obtain a matching result, simultaneous positioning and mapping are performed according to the matching result, the multi-modal features are richer and more robust compared with point-line-surface features, positioning and mapping are performed based on the multi-modal features, and the robustness of positioning and mapping is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of simultaneous localization and mapping, and particularly to a method for simultaneous localization and mapping and related devices. Background Art

[0002] Simultaneous Localization and Mapping (SLAM) is a technology that locates a sensor or its carrier based on environmental information data obtained by real-time sensors and simultaneously reconstructs an environmental map. As an important research direction in the field of robotics, SLAM is widely used in positioning and navigation.

[0003] However, during the process of localization and mapping through SLAM, due to the characteristics of SLAM being low-level visual features such as points, lines, and planes, the features are relatively simple, resulting in SLAM not being robust. Summary of the Invention

[0004] This application provides a method for simultaneous localization and mapping and related devices. By extracting features from visual data to obtain multi-modal features and performing feature matching on the multi-modal features, the multi-modal features are richer and more robust compared to point, line, and plane features, thereby improving the robustness of SLAM.

[0005] To achieve the above objective, this application adopts the following technical solutions:

[0006] In a first aspect, a method for simultaneous localization and mapping is provided. The method includes: obtaining visual data of a target to be located; extracting features from the visual data to obtain multi-modal features, where the multi-modal features include a first visual feature corresponding to the visual data and an auxiliary feature, the auxiliary feature is aligned with the first visual feature, and the auxiliary feature and the first visual feature are different modal features of the visual data respectively; performing feature matching based on the multi-modal features to obtain a matching result; and performing simultaneous localization and mapping according to the matching result.

[0007] In this way, after obtaining the visual data of the target to be detected, by extracting features from the visual data to obtain multi-modal features, since the visual feature and the auxiliary feature in the multi-modal features are aligned, and feature matching is performed based on the multi-modal features, simultaneous localization and mapping are performed through the matching result; by aligning the multi-modal features of the visual data, feature matching can be performed on multiple modal features of the visual data, and the features of multiple modalities of the same visual data increase the richness of the features of the visual data and improve the robustness of simultaneous localization and mapping. For example, if the construction of the map is based on multi-modal features, the constructed map can be applied in multiple modalities, such as text navigation, voice navigation, etc.

[0008] In a possible implementation of the first aspect, the first visual feature is determined based on the visual data of the target to be located at the first moment; the feature matching based on the multimodal features to obtain a matching result includes:

[0009] Obtain a second visual feature, where the second visual feature is determined based on the visual data of the target to be located at the second moment, the second moment is earlier than the first moment, and the multimodal features corresponding to the first moment and the multimodal features corresponding to the second moment have the same auxiliary feature;

[0010] Perform a first matching process on the auxiliary feature and the first visual feature to obtain a first matching result;

[0011] Perform a first matching process on the auxiliary feature and the second visual feature to obtain a second matching result;

[0012] Perform a second matching process on the first matching result and the second matching result to obtain a matching result.

[0013] First, perform a first matching process on the visual features in the two multimodal features and the same auxiliary feature to obtain two matching results; then, perform a secondary matching based on the two matching results to obtain a matching result. Since the first matching process is a matching between different modal features of the same visual data, and since the first matching process is between two aligned modal features, the matching efficiency can be improved; the second matching process is based on the results of the first matching, that is, the multimodal features are preprocessed through the first matching process, and then the second matching process is performed based on the preprocessed multimodal features, which can improve the overall efficiency and the accuracy of feature matching.

[0014] In a possible implementation of the first aspect, the performing a first matching process on the auxiliary feature and the first visual feature to obtain a first matching result; performing a first matching process on the auxiliary feature and the second visual feature to obtain a second matching result includes: calculating the similarity between the auxiliary feature and the first visual feature to obtain a first matching result; calculating the similarity between the auxiliary feature and the second visual feature to obtain a second matching result. In this way, the matching result of the two features is determined by calculating the similarity between the auxiliary feature and the visual feature.

[0015] In a possible implementation of the first aspect, the extracting features from the visual data to obtain multimodal features includes:

[0016] Input the visual data into a first deep learning model to obtain visual features; obtain auxiliary features aligned with the visual features.

[0017] In this way, the visual features in the multi-modal features are obtained by extracting features from the vision through the deep learning model; the auxiliary features aligned with the visual features can be determined according to the current scene, or can be determined by inputting the visual data into the corresponding deep learning model and extracting features by the deep learning model.

[0018] In a possible implementation manner of the first aspect, the training method of the first deep learning model includes: obtaining historical visual data and inputting the historical visual data into the first deep learning model to obtain visual features; obtaining the auxiliary features corresponding to the historical visual data; determining the loss value between the auxiliary features and the visual features according to a preset loss function; optimizing the first deep learning model according to the loss value; and then repeating the above steps until the loss value between the auxiliary features and the visual features converges. By calculating the loss value between the visual features and the auxiliary features, the model parameters of the first deep learning model are optimized through the loss value; the visual data is re-extracted by the optimized first deep learning model, and the loss value between the extracted visual features and the auxiliary features is calculated again. Through multiple iterations, the first deep learning model is continuously trained to achieve the convergence of the loss value between the auxiliary features and the visual features, that is, the visual features output by the first deep learning model are aligned with the auxiliary features. The objective of the loss function here is to measure the gap between the visual features and the auxiliary features of the model, and the parameters of the first deep learning model are optimized by minimizing the value of the loss function, so as to improve the accuracy and generalization ability of the first deep learning model. That is, the first deep learning model is trained with a large number of visual feature - auxiliary feature pairs to achieve the alignment of the visual features input to the first deep learning model and the corresponding auxiliary features.

[0019] In a possible implementation manner of the first aspect, the obtaining of the auxiliary features corresponding to the visual data includes: obtaining the first auxiliary features aligned with the visual features; and obtaining the second auxiliary features aligned with the visual features according to the first auxiliary features. If the auxiliary features include features of multiple other modalities different from the visual features, after obtaining the auxiliary features of one modality, the auxiliary features of another modality can be determined through the auxiliary features of this modality, so as to realize the diversity of the acquisition methods of the auxiliary features to meet the diverse needs of users.

[0020] In a possible implementation manner of the first aspect, the method further includes: receiving the instruction information of the user; determining the auxiliary features according to the instruction information; determining the multi-modal features according to the auxiliary features; and determining the destination in the built map corresponding to the instruction information according to the multi-modal features. The map is constructed through the multi-modal features. After the map construction is completed, the constructed map can be used according to the auxiliary features in the multi-modal features. To meet the different map usage needs of users.

[0021] Further, since the auxiliary features are features of other modalities different from the visual features, the user can select the corresponding implementation method of the constructed map according to the type or modality of the auxiliary features. For example, if the auxiliary features include voice features, the user can upload a voice file (such as an audio file of a cat's voice) or input the corresponding voice information feature (the location with a cat's voice) through an indication message. After receiving the user's indication message, the corresponding auxiliary features can be determined based on the indication message, and then the multi-modal features corresponding to the auxiliary features can be determined based on the auxiliary features; then the address information that the user is looking for or the address information that the user is going to navigate to can be determined based on the multi-modal features. Thus, the richness of the usage scenarios of the map constructed based on multi-modal features is achieved.

[0022] In a second aspect, a simultaneous localization and mapping device is provided, including: a memory, where the memory includes computer-readable instructions;

[0023] a processor communicatively coupled to the memory, where the processor is configured to execute the computer-readable instructions, so that the simultaneous localization and mapping device executes the simultaneous localization and mapping method according to any one of the first aspect.

[0024] In a third aspect, a computer-readable storage medium is provided, including a program or instructions, which when executed by a processor, implement the simultaneous localization and mapping method according to any one of the first aspect.

[0025] In a fourth aspect, a chip is provided, including a processor, which is configured to call and run instructions stored in a memory from the memory, so that an electronic device installed with the chip executes the simultaneous localization and mapping method according to any one of the first aspect.

[0026] For the beneficial effects brought by each possible implementation manner of the simultaneous localization and mapping device provided in the second aspect, the computer-readable storage medium provided in the third aspect, and the chip provided in the fourth aspect of the embodiments of the present application, reference can be made to the descriptions in various possible implementation manners of the first aspect, and details are not described herein again. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of the architecture of a visual SLAM;

[0028] Figure 2 It is a schematic flowchart of a simultaneous localization and mapping method provided by an embodiment of the present application;

[0029] Figure 3 It is a schematic diagram of training a first deep learning model provided by an embodiment of the present application;

[0030] Figure 4A schematic diagram of feature matching provided by an embodiment of this application;

[0031] Figure 5 A schematic diagram of location recognition provided by an embodiment of this application;

[0032] Figure 6 A flowchart of a simultaneous localization and mapping method provided by an embodiment of this application;

[0033] Figure 7 A schematic diagram of the architecture of visual SLAM provided by an embodiment of this application;

[0034] Figure 8 A schematic diagram of the structure of a simultaneous localization and mapping device provided by an embodiment of this application; Detailed implementation manners

[0035] Next, the technical solutions in this application will be described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments.

[0036] Next, it will be described through Figure 1 the visual SLAM architecture shown. Please refer to Figure 1 , Figure 1 which is a schematic diagram of the architecture of a visual SLAM. Figure 1 The SLAM architecture in

[0037] includes a front end and a back end. The front end includes visual feature extraction, feature matching, and visual odometry, and the back end includes loop detection.

[0038] Among them, the function of the front end is to extract representative feature points or feature descriptors from the video or image obtained by the sensor, and match these feature points or descriptors to adjacent images. These matched feature points or descriptors can be used to estimate the motion of the sensor or its carrier and the scene; among them, image data in the environment is obtained through the sensor, such as image data or video data; visual feature extraction is used to extract feature points or feature descriptors from the image data; feature matching is used to match the feature points in images at different times to estimate the relative motion of the sensor or its carrier between adjacent frames, so as to obtain a preliminary pose estimation of the sensor or its carrier.

[0039] In the above SLAM process, the visual features extracted from the image by visual feature extraction (such as low-level visual features like points, lines, and planes) are relatively simple, resulting in insufficient robustness of SLAM.

[0040] Based on the above problems, the embodiments of the present application provide a simultaneous localization and mapping method. After obtaining the visual data of the target to be localized, the visual data is subjected to feature extraction to obtain multi-modal features corresponding to the visual data. The multi-modal features include the visual features corresponding to the visual data and auxiliary features, where the auxiliary features and the visual features are different modal features of the visual data respectively. For example, the auxiliary features can be text features or voice features corresponding to the visual data. Then, by performing feature matching on the multi-modal features, a matching result is obtained, and simultaneous localization and mapping are performed based on the matching result. By performing feature extraction to obtain multi-modal features and performing feature matching on the multi-modal features, the multi-modal features are richer and more robust compared to point, line, and plane features, thereby improving the robustness of SLAM.

[0041] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a simultaneous localization and mapping method provided by the embodiments of the present application. Figure 2 The simultaneous localization and mapping method in

[0042] S201. Obtain the visual data corresponding to the target to be localized.

[0043] Optionally, the visual data is image data or video data obtained by a visual sensor (such as a camera).

[0044] Optionally, the target to be localized can be a visual sensor or a carrier carrying the visual sensor, such as a mobile device like a robot or a floor sweeper.

[0045] S202. Perform feature extraction on the visual data to obtain multi-modal features. The multi-modal features include the first visual features corresponding to the visual data and auxiliary features. The first visual features and the auxiliary features are different modal features of the visual data respectively, and the auxiliary features are aligned with the first visual features.

[0046] It is easy to understand that multi-modal features refer to data features obtained for the same description object from different fields or perspectives, and each field or perspective for describing these data is called a modality, such as text, image, audio, video, mixed data, etc. Different modal features of data can be different characteristics or attributes of the same data. The multi-modal features in this application are different modal features of the same visual data obtained by a visual sensor. For example, the same visual data has visual features (such as color, texture, shape, etc.), text features (such as vocabulary, grammar, syntax, semantics), and voice features (pitch, timbre, volume). Then the visual features, text features, and voice features are respectively different modal features of the same visual data. The multi-modal features of this application refer to, in addition to the visual features of visual data, other modal features.

[0047] Among them, the alignment of auxiliary features with visual features means associating and synchronizing auxiliary features (such as speech alignment) with visual features to achieve multi-modal information fusion and joint processing of visual data. Here, the visual features of the same visual data are aligned with the features of other modalities. Optionally, after the alignment of auxiliary features with visual features, there are corresponding feature points or feature descriptors. For example, in the visual features of an image, there is an image of a dog, and the corresponding feature point in the text features is the word "dog".

[0048] S203. Perform feature matching on the multi-modal features to obtain a matching result.

[0049] Optionally, the multi-modal feature is to perform feature matching on the multi-modal feature at the current moment with all multi-modal features at the previous moment or the multi-modal features of the target quantity in all multi-modal features, so as to determine the relationship between the two, such as the similarity between the two, or whether the two are the same multi-modal feature obtained at different times.

[0050] Furthermore, because the visual features and auxiliary features in the multi-modal features are aligned, that is, the aligned visual features and auxiliary features have corresponding feature points or feature descriptors, thereby enhancing the richness of the features for feature matching in adjacent frames or adjacent moments.

[0051] S204. Perform simultaneous localization and mapping on the target to be located based on the matching result.

[0052] Optionally, the matching result can be the change in multimodal features between adjacent frames of the target to be detected, such as rotational change or translational change. Then, based on the matching result, the pose change (including the position and angle of the target to be detected) or trajectory change of the target to be detected (the camera or the carrier of the camera) can be determined, so as to achieve the positioning of the target to be located. At the same time, based on the change in its own position and the acquired visual data, map construction is realized. For example, feature matching is performed based on the multimodal features of the visual data acquired in real time and the multimodal features corresponding to the previously constructed map, and the map is updated according to the matching result.

[0053] In this way, after acquiring the visual data of the target to be detected, multi-modal features are obtained by extracting features from the visual data. Since the visual features and auxiliary features in the multi-modal features are aligned, and feature matching is performed based on the multi-modal features, simultaneous localization and mapping are performed through the matching result; by aligning the multi-modal features of the visual data, feature matching can be performed on multiple modal features of the visual data. The features of multiple modalities of the same visual data increase the richness of the visual data and the diversity of feature matching, and improve the robustness of simultaneous localization and mapping.

[0054] Optionally, after acquiring the visual data, the visual data can be input into a first deep learning model (such as a neural network model), that is, the first deep learning model extracts features from the visual data to obtain the visual features of the visual data, and then the auxiliary features of the visual data aligned with the visual features are obtained. The visual features and the auxiliary features of the visual data aligned with the visual features form the multi-modal features of the visual data. Among them, the auxiliary features and the visual features are respectively features of different modalities of the visual data. For example, the auxiliary features can be the text features of the visual data. In this way, the visual features corresponding to the visual data are determined through the deep learning model, and then other modal features (i.e., auxiliary features) aligned with the visual features are obtained. The visual features and the auxiliary features form multi-modal features, thereby improving the richness of the features obtained by feature extraction.

[0055] It is easy to understand that the auxiliary features can also be obtained by feature extraction through a deep learning model. Optionally, the visual data can be input into a second deep learning model, and the second deep learning model extracts features from the visual data to obtain the auxiliary features, and the visual features output by the first deep learning model are aligned with the auxiliary features output by the second deep learning model.

[0056] Exemplarily, the second deep learning model performs image recognition on the visual data to obtain the text features in the visual data.

[0057] Optionally, the second deep learning model can be a neural network model.

[0058] Optionally, if the first deep learning model and the second deep learning model are the same deep learning model, the visual data is input into the deep learning model, and the deep learning model outputs aligned visual features and auxiliary features. By training the deep learning model, after inputting the visual data into the deep learning model, the deep learning model outputs aligned visual features and auxiliary features. Of course, in other embodiments, the first deep learning model and the second deep learning model can also be different deep learning models. By training the first deep learning model and the second deep learning model, after inputting the same visual data into the two deep learning models, the two deep learning models output aligned visual features and auxiliary features.

[0059] It is easy to understand that the auxiliary features can also be obtained in other ways. By presetting the auxiliary features corresponding to each scenario, before obtaining the auxiliary features corresponding to the visual data, first obtain the scenario where the visual data is located, and then determine the corresponding auxiliary features according to the scenario. For example, if the auxiliary features are text features, after determining that the scene where the target to be located is an office, search for the text features corresponding to the office scene from the preset auxiliary features, such as computers, workstations, the numbers of each workstation, water dispensers, meeting rooms, etc. For example, if the auxiliary features are sound features, after determining that the scene where the target to be located is a zoo, search for the sound features corresponding to the office scene from the preset auxiliary features, such as the calls of monkeys, lions, dogs, etc.

[0060] It is easy to understand that the first deep learning model is used to extract features from the visual data to obtain the visual features corresponding to the visual data. In order to align the visual features output by the deep learning model with the auxiliary features, the first deep learning model can be adaptively trained so that the visual features output by the trained first deep learning model are aligned with the auxiliary features.

[0061] Optionally, to align the auxiliary features with the visual features output by the first deep learning model, the training method of the first deep learning model can be optimized. Specifically, first, input the visual data into the first deep learning model, and the first deep learning model outputs the corresponding visual features. Then, obtain the auxiliary features corresponding to the visual data. Then, determine the loss value between the auxiliary features and the visual features according to the preset loss function, and optimize the first deep learning model according to the loss value. Then, input the visual data into the optimized first deep learning model and loop in turn until the loss value between the auxiliary features and the visual features converges, so as to realize the alignment of the visual features extracted from the visual data by the first deep learning model with the auxiliary features corresponding to the visual data. It is easy to understand that the above method calculates the loss value between the visual features and the auxiliary features, and optimizes the model parameters of the first deep learning model through the loss value; the optimized first deep learning model extracts features from the visual data again, and calculates the loss value between the extracted visual features and the auxiliary features again. Through multiple iterations, the first deep learning model is continuously trained to make the loss value between the auxiliary features and the visual features converge, that is, the visual features output by the first deep learning model are aligned with the auxiliary features. The goal of the loss function here is to measure the gap between the visual features and the auxiliary features of the model, and optimize the parameters of the first deep learning model by minimizing the value of the loss function, so as to improve the accuracy and generalization ability of the first deep learning model. That is, the first deep learning model is trained with a large number of visual feature-auxiliary feature pairs to align the visual features input to the first deep learning model with the corresponding auxiliary features.

[0062] Optionally, before calculating the loss value between the visual features and the auxiliary features, in order to meet the calculation requirements of the loss function, the visual features and the auxiliary features can be encoded respectively. For example, if the auxiliary features are text features, the text features can be encoded by a text encoder, and the visual features (such as images) can be encoded by an image encoder.

[0063] Optionally, after calculating the loss value between the visual features and the auxiliary features, the gradient of the loss function with respect to the parameters of the first deep learning model can be calculated by the backpropagation algorithm, and then the gradient descent algorithm can be used to update the parameters of the first deep learning model, so that the loss value between the alignment of the visual features and the auxiliary features corresponding to the visual data gradually becomes smaller.

[0064] Optionally, the loss function can be a mean square error loss function or a cross entropy loss function. The goal of the above loss function is to measure the difference or similarity between the visual features and the auxiliary features.

[0065] Please refer to Figure 3 , Figure 3 which is a schematic diagram of training the first deep learning model provided by an embodiment of the present application. Figure 3In this method, visual data (such as videos or images) are first input into a first deep learning model. The first deep learning model outputs visual features corresponding to the visual data (such as the images in Figure 3 ), and text features corresponding to the visual data are obtained. The text features can be manually annotated. For example, by manually viewing the image or video to determine the objects corresponding to the image or video (such as a dog, a table), and determining the text corresponding to the target object (such as the text corresponding to a dog); then the image features are encoded by an image encoder to obtain an image feature map; the text features are encoded by a text encoder to obtain a text feature map. Then, according to a pre-set loss function, the loss value between the text features and the visual features is calculated, and the first deep learning model is optimized according to the loss value; then the visual data are input into the optimized first deep learning model, and the above steps are repeated until the loss value between the text features and the visual features converges or meets a pre-set number of iterations, so as to realize the alignment of the visual features output by the first deep learning model and the auxiliary features corresponding to the visual data.

[0066] Optionally, if the visual features in the multi-modal features are obtained by the first deep learning model extracting features from visual data; and the auxiliary features are obtained by the second deep learning model extracting features from visual data, then after the training of the first deep learning model is completed, the second deep learning model can be trained according to the same concept: inputting the visual data into the second deep learning model to obtain auxiliary features, calculating the loss value between the auxiliary features and the visual features according to a pre-set loss function; optimizing the second deep learning model according to the loss value; then inputting the visual data into the optimized second deep learning model, and repeating the above steps until the loss value between the text features and the visual features converges or meets a pre-set number of iterations, so as to realize the alignment of the auxiliary features output by the second deep learning model and the visual features output by the first deep learning model.

[0067] Optionally, the auxiliary features may include other modal features of a kind of visual data that are different from the visual features. For example, the auxiliary features are text features, and the text features are aligned with the visual features, and the two modal features of the visual data form the multi-modal features of the visual data; of course, in other embodiments, the auxiliary features may include at least two other modal features of the visual data that are different from the visual features. For example, the auxiliary features include text features and speech features, then the text features are aligned with the visual features and the speech features are aligned with the visual features.

[0068] Further, if the auxiliary features include multiple other modal features different from visual features, after obtaining one of the auxiliary features, another different auxiliary feature can be obtained through the obtained auxiliary feature. Exemplarily, the visual data is image information, and the image information contains a Husky. First, the text features corresponding to the visual data are obtained as "dog" and "Husky". Then, according to the text features, the voice feature (another auxiliary feature) can be at least one of the sound of the Husky "woof woof woof", the pronunciation of "dog", and the pronunciation of "Husky". Among them, the voice feature can be obtained from the network according to the text features. For example, the barking of a dog can be searched from the network through "dog" in the text features.

[0069] It is easy to understand that compared with features such as points, lines, and planes, the multi-modal features contain richer feature content. The multi-modal features of the current frame can be used to perform feature matching on adjacent frames, or the multi-modal features of the current moment can be used to perform feature matching on the multi-modal features of adjacent moments. The target to be located and the map can be updated according to the changes in the multi-modal features. However, since the multi-modal features contain more features, if feature matching is directly performed, the performance requirements for the device are relatively high. To improve the accuracy and efficiency of matching, a feature matching method from coarse to fine can be adopted. Specifically, if the first visual feature is determined based on the visual data of the target to be located at the first moment, S203 includes: obtaining the second visual feature of the target to be located, where the second visual feature is determined based on the visual data of the target to be located at the second moment, and the second moment is earlier than the first moment; performing a first matching process on the auxiliary feature and the first visual feature to obtain a first matching result; performing a first matching process on the auxiliary feature and the second visual feature to obtain a second matching result; performing a second matching process on the first matching result and the second matching result to obtain a matching result. That is, after obtaining the visual data of the data to be located at the first moment, feature extraction is performed on the visual data to obtain the multi-modal features corresponding to the first moment; then the multi-modal features corresponding to the second moment are obtained, and the second moment is earlier than the first moment; since the two multi-modal features for feature matching are respectively determined based on the visual data obtained by the sensor in two adjacent frames or two adjacent moments, it is determined that the auxiliary features in the two multi-modal features for feature matching are the same; first, a first matching process is performed between the visual features in the two multi-modal features and the same auxiliary feature to obtain two matching results; then a second matching is performed based on the two matching results to obtain a matching result. Since the first matching process is a matching between different modal features of the same visual data, and since the first matching process is between two aligned modal features, the matching efficiency can be improved; the second matching process is based on the result of the first matching, that is, the multi-modal features are preprocessed through the first matching process, and then the second matching process is performed based on the preprocessed multi-modal features, which can improve the overall efficiency and the accuracy of feature matching.

[0070] It is easy to understand that if the auxiliary features in the multi-modal features corresponding to the first moment and the multi-modal features corresponding to the second moment are different, when performing feature matching in S203, the visual features corresponding to the two moments can be directly subjected to feature matching to obtain a first matching result, and the auxiliary features corresponding to the two moments can be subjected to feature matching to obtain a second matching result, and the final matching result can be determined based on the first matching result and the second matching result. Optionally, by calculating the similarity of the visual features at the two moments and the similarity of the auxiliary features at the two moments; then the matching result is determined by the two similarities; for example, if the similarity of the visual features at the two moments is calculated to be 0.6 and the similarity of the auxiliary features at the two moments is calculated to be 0.7, then the matching result is a similarity of 0.65.

[0071] Optionally, perform a first matching process on the auxiliary feature and the first visual feature to obtain a first matching result; perform a first matching process on the auxiliary feature and the second visual feature to obtain a second matching result, including: calculating the similarity between the auxiliary feature and the first visual feature to obtain a first matching result; calculating the similarity between the auxiliary feature and the second visual feature to obtain a second matching result. By calculating the similarity between the first visual feature and the auxiliary feature and the similarity between the second visual feature and the auxiliary feature, the matching degree between the visual features and the auxiliary features at the two moments is determined.

[0072] It is easy to understand that the visual features obtained from visual data may include multiple features such as color features, texture features, shape features, and spatial relationship features, amplitude features, histogram features, transform coefficient features, features of points and lines, and gray edge features. Before calculating the similarity, the visual features and the auxiliary features can be encoded respectively to obtain a visual feature map and an auxiliary feature map; then calculate the similarity between the visual feature map and the auxiliary feature map, and the first matching result is the first similarity feature heat map; the second matching result is the second similarity feature heat map.

[0073] Furthermore, after calculating the similarity between the first visual feature and the auxiliary feature and the similarity between the second visual feature and the auxiliary feature, that is, after obtaining the two similarity feature heat maps; perform feature matching on the two similarity feature heat maps to obtain the matching result of the multi-modal features at the two moments, and then perform a second matching process on the first matching result and the second matching result to obtain the matching result, including: performing feature matching on the first similarity feature heat map and the second similarity feature heat map to obtain a second matching result, where the second matching result can be the similarity of the two similarity feature heat maps, and each pixel point or each specific region in the visual feature can be converted into a vector, that is, the similarity feature heat map, and then the similarity calculation method is used to calculate the similarity between the vectors.

[0074] Optionally, after obtaining the first similar feature heat map and the second similar feature heat map, the two similar feature heat maps can be matched pixel by pixel to obtain corresponding matching results.

[0075] In some embodiments, if the number of auxiliary features in the multi-modal features is multiple, then when performing feature matching, the auxiliary features common to the multi-modal features corresponding to the first moment and the second moment can be selected from the multiple auxiliary features for feature matching. For example, if the multi-modal features include three auxiliary features, and the multi-modal features at the first moment and the multi-modal features corresponding to the second moment have the same text feature, then the text feature is used as the auxiliary feature for feature matching. Then, the first matching process is performed on the text feature and the first visual feature to obtain the first matching result; the first matching process is performed on the text feature and the second visual feature to obtain the second matching result.

[0076] Please refer to Figure 4 , Figure 4 which is a schematic diagram of a feature matching provided by an embodiment of the present application. Figure 4 where the auxiliary feature is a text feature. After obtaining the visual features of two adjacent frames (i.e., Figure 4 the image A and the image B in Figure 4 ), the same text feature (i.e., the text in

[0077] It is easy to understand that during the movement of the sensor or its carrier, visual data is collected by a visual sensor, and then a map is constructed based on multi-modal features. For example, based on multi-modal features and corresponding spatial position information, map construction is performed through the SLAM algorithm; and during the movement of the sensor or its carrier, data collection and feature extraction are continuously performed, and the constructed map is updated based on the obtained multi-modal features. At the same time, visual data is collected by the visual sensor, and feature extraction is performed on the visual data to obtain multi-modal features, and then the multi-modal features are stored, for example, stored in a specific database; the currently obtained multi-modal features can also be feature-matched with the previously stored multi-modal features, and based on the matching result, it is determined whether the current position is a position previously visited by the sensor or its carrier, so as to identify the current position, that is, to accurately locate the sensor or its carrier, which is convenient for realizing the positioning and navigation of the sensor or its carrier in an unknown area. Further, place recognition can be applied to the loop closure detection of SLAM and the relocalization stage based on a prior map.

[0078] Please refer to Figure 5 , Figure 5 FIG. [FIG. NUMBER] is a schematic diagram of a place recognition provided by an embodiment of the present application. During the movement of the sensor or its carrier, visual data is collected by a visual sensor, and feature extraction is performed on the visual data to obtain multi-modal features, and then a map is constructed based on the multi-modal features, and the obtained multi-modal features are stored.

[0079] Figure 5 Among them, the multi-modal features corresponding to the key frames are selected from the multi-modal features of the obtained consecutive frames, and the multi-modal features corresponding to the key frames are stored, for example, stored in a preset database.

[0080] After obtaining the visual data corresponding to the currently acquired current frame, feature extraction is performed on the visual data to obtain multi-modal features, and then the multi-modal features corresponding to the current frame are feature-matched with the multi-modal features corresponding to the stored key frames to obtain target key frames. Among them, the target key frames can be a specific number of key frames with a similarity greater than a preset threshold;

[0081] Further, the multi-modal features corresponding to the target key frames and the multi-modal features corresponding to the current frame are feature-matched to obtain a matching result, and based on the matching result, the pose of the sensor or its carrier is determined, and it is judged whether the determined pose is reasonable; if the determined pose of the sensor or its carrier is reasonable, it is determined that the position recognition of the sensor or its carrier is successful; if the determined pose of the sensor or its carrier is unreasonable (for example, an angle or position that the carrier or sensor cannot achieve), it is determined that the position recognition of the sensor or its carrier fails.

[0082] It is easy to understand that the positioning and mapping in the embodiments of the present application are realized based on the multi-modal features of visual data. After the map is constructed, if the auxiliary features include text features, the target to be determined receives a text instruction or the user's voice instruction, and determines the corresponding text features according to the user's instruction; then, the corresponding multi-modal features are found according to the text features, so that navigation can be performed based on the multi-modal features.

[0083] Please refer to Figure 6 , Figure 6 which is a schematic diagram of a simultaneous localization and mapping method provided by the embodiments of the present application. After the mapping is completed, the above method further includes: S601 - S604.

[0084] S601. Receive the instruction information of the user.

[0085] Optionally, the user can input the instruction information to the target to be located through touch screen input or file input, etc. Here, the instruction information can be navigation information for navigating the target to be located to the destination corresponding to the instruction information; of course, the instruction information can also be query information for querying the location in the built map.

[0086] The target to be located here can be a movable device such as a robot;

[0087] S602. Determine the auxiliary features according to the instruction information.

[0088] Optionally, after obtaining the navigation information, the user's navigation information is converted into the corresponding auxiliary features. If the auxiliary features are text features, after determining the navigation text or navigation voice input by the user, the corresponding navigation voice or navigation text is converted into the corresponding text features; if the auxiliary features are voice features, after obtaining the navigation text or navigation voice input by the user, the navigation text or navigation voice can be converted into the corresponding voice features. Exemplarily, if the navigation text input by the user is "Go to the monkey garden in the zoo", the text feature: the monkey garden in the zoo or the voice feature: the monkey garden in the zoo or the sound of monkeys is determined according to the navigation text.

[0089] Optionally, if the auxiliary features include multiple modal features, after determining one modal auxiliary feature according to the navigation information, the features of another modality can be determined according to the features of this modality. Exemplarily, if the navigation text input by the user is "Go to the monkey garden in the zoo", the text feature: the monkey garden in the zoo is determined according to the navigation text; then, the corresponding voice feature: the sound of monkeys is determined according to the monkey garden in the zoo.

[0090] S603. Determine the multi-modal features according to the auxiliary features.

[0091] Optionally, during the map building process, multiple multi-modal features are stored. After obtaining the auxiliary features, the stored multi-modal features can be searched through the auxiliary features to determine the multi-modal features that contain the auxiliary features.

[0092] S604. Determine the destination in the built map according to the multi-modal features.

[0093] It is easy to understand that the construction of the map is based on multi-modal features. Each multi-modal feature corresponds to a specific position in the built map. After determining the multi-modal feature, the position corresponding to the multi-modal feature can be determined by feature matching or other means through the multi-modal feature, that is, the destination that the user wants to navigate to.

[0094] In this way, the map is constructed through multi-modal features. After the map construction is completed, the built map can be used according to the auxiliary features in the multi-modal features to meet the different map usage requirements of users.

[0095] Furthermore, since the auxiliary features are features of other modalities different from visual features, the user can select the implementation method of the corresponding built map according to the type or modality of the auxiliary features. For example, if the auxiliary features include voice features, the user can upload a voice file (such as an audio file of a cat's voice) or input the corresponding voice information through the indication information (the position with the cat's voice). After receiving the user's indication information, the corresponding auxiliary features can be determined based on the indication information, and then the multi-modal features corresponding to the auxiliary features can be determined according to the auxiliary features; then the address information searched by the user or the address information that the user will navigate to can be determined according to the multi-modal features. This enriches the usage scenarios of the map constructed based on multi-modal features.

[0096] Please refer to Figure 7 , Figure 7 For the schematic architecture diagram of a visual SLAM provided by an embodiment of the present application, please refer to Figure 7 ,First, visual data is collected through a visual sensor. The visual data can be image data or video data. Then, feature extraction is performed on the visual data. The visual data can be input into the first deep learning model to obtain the visual features in the multi-modal features; then, auxiliary features aligned with the visual features are obtained to obtain multi-modal features.

[0097] Then, by performing feature matching on the multi-modal features, if the current multi-modal features are determined from the visual data at the first moment, the second visual feature is obtained from the stored multi-modal features. The second visual feature is determined from the visual data of the visual sensor at the second moment, and the second moment is earlier than the first moment. Then, calculate the similarity between the auxiliary feature and the first visual feature to obtain a similarity feature heat map; calculate the similarity between the auxiliary feature and the second visual feature to obtain a similarity feature heat map. Then, perform feature matching on the two similarity feature heat maps to obtain a feature matching matrix, that is, the matching result.

[0098] Then, the pose of the target to be located is determined according to the matching result, and it is judged whether the target to be located has returned to a previously visited position based on the matching result, so as to correct the accumulated positioning errors and improve the accuracy of the built map.

[0099] Furthermore, since the map of the present application is constructed based on the multi-modal features of visual data, the visualization process of each process of positioning and mapping can be realized based on the auxiliary features. For example, if the auxiliary feature is a text feature, in the feature extraction stage, the extracted multi-modal features are displayed in a readable manner according to the text feature, so that the user can read the multi-modal features through the text feature; in the feature matching stage, a visual display is performed according to the similarity feature heat map; in the loop closure stage, the trajectory of the target to be located and the current trajectory can be displayed through the visualization stage, so that it can be visually judged whether the user has been to a previous position. After the map construction is completed, the built map is displayed through a visualization interface, and the characteristics of each position in the map are displayed through text features.

[0100] The simultaneous localization and mapping method provided by the present application, after obtaining the visual data of the target to be located, performs feature extraction on the visual data to obtain the multi-modal features corresponding to the visual data. The multi-modal features include the visual features corresponding to the visual data and the auxiliary features. Among them, the auxiliary feature and the visual feature are respectively different modal features of the visual data. For example, the auxiliary feature can be the text feature or the voice feature corresponding to the visual data. Then, by performing feature matching on the multi-modal features, a matching result is obtained, and simultaneous localization and mapping are performed according to the matching result. By performing feature extraction to obtain multi-modal features and performing feature matching on the multi-modal features, compared with point, line, and plane features, the multi-modal features are richer and more robust, so as to improve the robustness of SLAM.

[0101] It should be understood that the above is only to help those skilled in the art better understand the embodiments of the present application, rather than to limit the scope of the embodiments of the present application. Those skilled in the art can obviously make various equivalent modifications or changes according to the above examples. For example, some steps in the above methods may not be necessary, or some steps may be newly added, etc. Or any combination of any two or any multiple of the above embodiments. The solutions after such modifications, changes or combinations also fall within the scope of the embodiments of the present application.

[0102] It should also be understood that the manners, situations, categories, and the division of embodiments in the embodiments of the present application are only for the convenience of description and should not constitute special limitations. The features in various manners, categories, situations, and embodiments can be combined without contradiction.

[0103] It should also be understood that the various digital numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. The size of the serial numbers of the above processes does not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0104] It should also be understood that the above description of the embodiments of the present application focuses on emphasizing the differences between the various embodiments. The same or similar parts not mentioned can be referred to each other. For the sake of brevity, they are not described here again.

[0105] The above combination Figures 1-7 describes the embodiments of the method and system provided by the embodiments of the present application. Next, the simultaneous localization and mapping device provided by the embodiments of the present application is described.

[0106] In this embodiment, the functional modules of the simultaneous localization and mapping device can be divided according to the above method. For example, corresponding to each function, they can be divided into each functional module, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative and is only a logical function division. There may be other division methods in actual implementation.

[0107] It should be noted that the relevant content of each step involved in the above method embodiment can be cited to the functional description of the corresponding functional module, and will not be repeated here.

[0108] The simultaneous localization and mapping device provided by the embodiments of the present application is used to execute the simultaneous localization and mapping method provided by the above method embodiment, so the same effect as the above implementation method can be achieved.

[0109] In other embodiments, in the case of adopting an integrated unit, the simultaneous localization and mapping device may include a processing module, a storage module, and a communication module. Among them, the processing module may be used to control and manage the operations of the simultaneous localization and mapping device. For example, it may be used to support the simultaneous localization and mapping device in executing the steps performed by the processing unit. The storage module may be used to support the storage of program codes, data, etc. The communication module may be used to support the communication between the simultaneous localization and mapping device and other devices.

[0110] Among them, the processing module may be a processor or a controller. It may be to implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, and so on. The storage module may be a memory. The communication module may specifically be a device that interacts with other simultaneous localization and mapping devices, such as a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, etc.

[0111] See Figure 8 , Figure 8 which shows a schematic structural diagram of an exemplary simultaneous localization and mapping device of the present application. Figure 8 The shown simultaneous localization and mapping device may execute the steps in any of the simultaneous localization and mapping methods performed by the simultaneous localization and mapping devices provided in the embodiments of the present application.

[0112] The simultaneous localization and mapping device 80 includes at least one processor 801, a memory 803, and at least one network interface 804.

[0113] The processor 801 is, for example, a general-purpose CPU, a digital signal processor (DSP), a network processor (NP), a GPU, a neural network processing unit (NPU), a data processing unit (DPU), a microprocessor, or one or more integrated circuits or application-specific integrated circuits (ASICs) for implementing the solution of this application, a programmable logic device (PLD), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The PLD is, for example, a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It can implement or execute various logic blocks, modules, and circuits described in connection with the disclosure of this application. The processor can also be a combination that implements a computing function, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on.

[0114] Optionally, the simultaneous localization and mapping device 80 further includes a bus 802. The bus 802 is used to transfer information between the components of the simultaneous localization and mapping device 80. The bus 802 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 802 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0115] The memory 803 is, for example, a read only memory (ROM) or other type of storage device that can store static information and instructions, or a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read only memory (EEPROM), a compact disc read only memory (CD ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 803 exists independently, for example, and is connected to the processor 801 through the bus 802. The memory 803 can also be integrated with the processor 801.

[0116] The network interface 804 uses any device such as a transceiver to communicate with other devices or communication networks, and the communication network can be an Ethernet, a radio access network (RAN) or a wireless local area network (WLAN), etc. The network interface 804 can include a wired network interface and can also include a wireless network interface. Specifically, the network interface 804 can be an Ethernet interface, such as a fast Ethernet (FE) interface, a gigabit Ethernet (GE) interface, an asynchronous transfer mode (ATM) interface, a WLAN interface, a cellular network interface or a combination thereof. The Ethernet interface can be an optical interface, an electrical interface or a combination thereof. In some embodiments of the present application, the network interface 804 can be used for the simultaneous localization and mapping device 80 to communicate with other devices, for example, to communicate with other simultaneous localization and mapping devices, and multiple simultaneous localization and mapping devices to communicate.

[0117] In a specific implementation, as some embodiments, the processor 801 can include one or more CPUs. Each of these processors can be a single-core processor or a multi-core processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0118] In specific implementations, as some embodiments, the simultaneous localization and mapping device 80 may include multiple processors. Each of these processors may be a single-core processor or a multi-core processor. The processors herein may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0119] In some embodiments, the memory 803 is used to store program instructions for executing the solution of this application, and the processor 801 may execute the program instructions stored in the memory 803. That is, the simultaneous localization and mapping device 80 may implement the method provided by the method embodiment shown in the above embodiment through the program instructions in the processor 801 and the memory 803. The program instructions may include one or more software modules. Optionally, the processor 801 itself may also store program instructions for executing the solution of this application.

[0120] During the specific implementation process, the processor 801 in the simultaneous localization and mapping device 80 of this application reads the instructions in the memory 803, so that Figure 8 the shown simultaneous localization and mapping device 80 can execute all or part of the steps in the simultaneous localization and mapping method executed by the simultaneous localization and mapping device in the above embodiment.

[0121] Among them, each step of the method described in the above embodiment is completed by the integrated logic circuit of the hardware in the processor of the simultaneous localization and mapping device 80 or the instructions in software form. The steps of the method embodiment disclosed in combination with this application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method embodiment. To avoid repetition, it will not be described in detail here.

[0122] It should be understood that the above-mentioned processor can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It is worth noting that the processor can be a processor that supports the advanced RISC machines (ARM) architecture.

[0123] Furthermore, in an optional embodiment, the above-mentioned memory can include a read-only memory and a random access memory, and provide instructions and data to the processor. The memory can also include a non-volatile random access memory. For example, the memory can also store information about the device type.

[0124] The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0125] In an exemplary embodiment, an embodiment of the present application provides a computer program (product), the computer program (product) comprising: computer program code which, when run on a computer, causes the computer to execute steps in any one of the simultaneous localization and mapping methods performed by the simultaneous localization and mapping device provided in the embodiments of the present application.

[0126] In an exemplary embodiment, an embodiment of the present application provides a computer-readable storage medium storing a program or instructions which, when run on a computer, cause the simultaneous localization and mapping method performed by any one of the simultaneous localization and mapping devices provided in the embodiments of the present application to be executed.

[0127] In an exemplary embodiment, an embodiment of the present application provides a chip comprising a processor for calling and running instructions stored in a memory, such that a communication device installed with the chip executes the simultaneous localization and mapping method performed by any one of the simultaneous localization and mapping devices provided in the embodiments of the present application.

[0128] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that integrates one or more available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk).

[0129] In this application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. It should be understood that there is no logical or chronological dependency among "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another.

[0130] It should also be understood that in various embodiments of this application, the magnitude of the serial numbers of each process does not mean the sequence of execution. The execution order of each process should be determined by its function and internal logic, and should not impose any limitation on the implementation process of the embodiments of this application.

[0131] In this application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "a plurality of" refers to two or more. For example, a plurality of second devices means two or more second devices. In this text, the terms "system" and "network" are often used interchangeably.

[0132] It should be understood that the terms used in the description of various examples herein are only for describing specific examples and are not intended to be limiting. As used in the description of various examples and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.

[0133] It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this application generally represents an "or" relationship between the preceding and following associated objects.

[0134] It should also be understood that the terms "if" and "when" can be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined..." or "if [the stated condition or event] is detected" can be interpreted to mean "when it is determined..." or "in response to determining..." or "when [the stated condition or event] is detected" or "in response to detecting [the stated condition or event]".

[0135] The above are only embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included within the protection scope of the present application.

Claims

1. A simultaneous localization and mapping method, characterized in that, The method includes: Obtaining visual data of a target to be located; Performing feature extraction on the visual data to obtain multi-modal features, where the multi-modal features include a first visual feature corresponding to the visual data and an auxiliary feature, the auxiliary feature is aligned with the first visual feature, and the auxiliary feature and the first visual feature are different modal features of the visual data respectively; Performing feature matching based on the multi-modal features to obtain a matching result; Performing simultaneous localization and mapping according to the matching result.

2. The method according to claim 1, characterized in that, The first visual feature is determined based on the visual data of the target to be located at a first moment; The performing feature matching based on the multi-modal features to obtain a matching result includes: Obtaining a second visual feature, where the second visual feature is determined based on the visual data of the target to be located at a second moment, the second moment is earlier than the first moment, and the multi-modal features corresponding to the first moment and the multi-modal features corresponding to the second moment have the same auxiliary feature; Performing a first matching process on the auxiliary feature and the first visual feature to obtain a first matching result; Performing a first matching process on the auxiliary feature and the second visual feature to obtain a second matching result; Performing a second matching process on the first matching result and the second matching result to obtain a matching result.

3. The method according to claim 2, characterized in that, The performing a first matching process on the auxiliary feature and the first visual feature to obtain a first matching result; The performing a first matching process on the auxiliary feature and the second visual feature to obtain a second matching result includes: Calculating the similarity between the auxiliary feature and the first visual feature to obtain a first matching result; Calculating the similarity between the auxiliary feature and the second visual feature to obtain a second matching result.

4. The method according to any one of claims 1 to 3, characterized in that, The performing feature extraction on the visual data to obtain multi-modal features includes: Inputting the visual data into a first deep learning model to obtain visual features; Obtaining an auxiliary feature aligned with the visual features.

5. The method according to claim 4, characterized in that, The training method of the first deep learning model includes: Obtaining historical visual data and inputting the historical visual data into the first deep learning model to obtain visual features; Obtaining the auxiliary feature corresponding to the historical visual data; Determining the loss value between the auxiliary feature and the visual feature according to a preset loss function; Optimizing the first deep learning model according to the loss value; Then repeating the above steps until the convergence of the loss value between the auxiliary feature and the visual feature.

6. The method according to any one of claims 1 to 5, characterized in that, The obtaining the auxiliary feature corresponding to the visual data includes: Obtaining a first auxiliary feature aligned with the visual features; Obtaining a second auxiliary feature aligned with the visual features according to the first auxiliary feature.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Receiving an instruction message from a user; Determining an auxiliary feature according to the instruction message; Determining the multi-modal features according to the auxiliary feature; Determining a destination in the built map corresponding to the instruction message according to the multi-modal features.

8. A simultaneous localization and mapping device, characterized in that, Including: A memory, where the memory includes computer-readable instructions; A processor communicating with the memory, the processor being configured to execute the computer-readable instructions such that the simultaneous localization and mapping device performs the simultaneous localization and mapping method according to any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, Comprising a program or instructions which, when executed by a processor, implement the simultaneous localization and mapping method according to any one of claims 1-7.

10. A chip, characterized in that, Comprising a processor for calling and running instructions stored in a memory from the memory, such that an electronic device installed with the chip executes the simultaneous localization and mapping method according to any one of claims 1-7.