A method for visual relocalization and an electronic device

Through the multi-stage screening model, the problem of insufficient recall and matching accuracy in the existing technology in scenarios with low texture or large environmental changes is solved, and efficient and accurate visual relocation is achieved.

CN120163877BActive Publication Date: 2025-07-29HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510609129.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-29
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

Existing visual relocation technology is difficult to take into account both recall and matching accuracy in scenarios with low texture or large environmental changes, resulting in system positioning failure or mismatch.

Method used

Through multi-stage filtering models, including visual bag of words, local feature filtering models, global image matching models and feature point detection and matching models, the relocation process is dynamically adjusted, and the relocation results are determined layer by layer.

Benefits of technology

It improves the accuracy and stability of visual repositioning, adapts to efficient positioning in different environments, and reduces the positioning error rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163877B_ABST
    Figure CN120163877B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of visual relocalization, and specifically provides a method and an electronic device for visual relocalization. The method may include: obtaining relevant parameters of a first candidate set similar to a query image recalled by a first screening model, where the relevant parameters include: the number of recalled items in the first candidate set, recall item evaluation parameters, and recall similarity; determining a subsequent screening model based on the relevant parameters, where the subsequent screening model includes at least one of a second screening model, a third screening model, and a fourth screening model; and determining a relocalization result of the query image based on the subsequent screening model and the query image. Some embodiments of this application can improve the accuracy of visual relocalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of visual relocalization, and in particular, to a method and an electronic device for visual relocalization. Background Art

[0002] Visual relocalization is a key technology in the SLAM (Simultaneous Localization And Mapping) system, and its main function is to re-determine the position and pose of the target object after tracking is lost. Current visual relocalization technology has been widely used in many application scenarios, such as robot navigation, augmented reality, driverless, etc. Most of the existing visual relocalization schemes are implemented based on feature point matching, and traditional image matching methods mostly rely on a single global matching or local matching algorithm. However, a single matching method often has difficulty in taking into account both the recall rate and the matching accuracy. Especially in scenes with low texture or large environmental changes, it may lead to system visual positioning failure or matching errors.

[0003] Therefore, how to provide a technical solution for a method of visual relocalization with higher accuracy has become a technical problem that needs to be solved urgently. Summary of the Invention

[0004] Some embodiments of this application aim to provide a method and an electronic device for visual relocalization. Through the technical solutions of the embodiments of this application, the accuracy of visual relocalization can be improved, and the positioning error rate can be reduced.

[0005] In a first aspect, some embodiments of this application provide a method for visual relocalization, including: obtaining relevant parameters of a first candidate set similar to a query image recalled by a first screening model, where the relevant parameters include: the number of recall items in the first candidate set, a recall item evaluation parameter, and a recall similarity, and the recall item evaluation parameter is used to measure the consistency of the matching results between images; determining a subsequent screening model based on the relevant parameters, where the subsequent screening model includes at least one of a second screening model, a third screening model, and a fourth screening model; and determining a relocalization result of the query image based on the subsequent screening model and the query image.

[0006] In some embodiments of the present application, after using the first screening model to recall the first candidate set similar to the image to be queried, the subsequent screening model for subsequent relocalization is determined based on the relevant parameters of the first candidate set, and then the relocalization result of the image to be queried is determined through the subsequent screening model. Some embodiments of the present application use different screening methods under different circumstances, and layer by layer screening to determine the relocalization result, improving the accuracy of the relocalization result and overcoming the defect of a relatively high error rate in single matching in the prior art; at the same time, ensuring the stability and efficiency of the relocalization system.

[0007] In some embodiments, the determining the subsequent screening model based on the relevant parameters includes: when the number of recall items is less than the first threshold, the subsequent screening model is the fourth screening model; the determining the relocalization result of the image to be queried based on the subsequent screening model and the image to be queried includes: matching the relocalization result from the first candidate set based on the fourth screening model and the image to be queried.

[0008] Some embodiments of the present application compare the number of recall items with the first threshold, confirm that the fourth screening model is used to process the image to be queried when it is less than the first threshold, and select the relocalization result from the first candidate set, which can achieve efficient and accurate relocalization in different situations.

[0009] In some embodiments, the determining the subsequent screening model based on the relevant parameters includes: when the number of recall items is not less than the first threshold, and the recall item evaluation parameter and the recall similarity meet the first condition, the subsequent screening models are the second screening model, the third screening model, and the fourth screening model; the first condition is that the recall item evaluation parameter is less than the second threshold and the recall similarity is greater than the third threshold; the determining the relocalization result of the image to be queried based on the subsequent screening model and the image to be queried includes: determining a second candidate set from the first candidate set based on the second screening model and the image to be queried; determining a third candidate set based on the third screening model and the image to be queried; matching the relocalization result from the second candidate set and the third candidate set based on the fourth screening model and the image to be queried.

[0010] Some embodiments of the present application analyze the number of recall items, the recall evaluation parameter, and the recall similarity, select the type of the subsequent screening model, and analyze and process the image to be queried and the first candidate set accordingly to determine the relocalization result, improving the positioning accuracy.

[0011] In some embodiments, determining a subsequent screening model based on the relevant parameters includes: when the number of recall items is not less than a first threshold, and the recall item evaluation parameter and the recall similarity meet a second condition, the subsequent screening model is a third screening model and a fourth screening model; the second condition is that the recall item evaluation parameter is not less than a second threshold and the recall similarity is not greater than a third threshold; determining a relocalization result of the image to be queried based on the subsequent screening model and the image to be queried includes: determining a third candidate set based on the third screening model and the image to be queried; and matching the relocalization result from the first candidate set and the third candidate set based on the fourth screening model and the image to be queried.

[0012] In some embodiments of the present application, by analyzing the number of recall items, the recall evaluation parameter, and the recall similarity, the type of the subsequent screening model is selected, and thus the image to be queried and the first candidate set are analyzed and processed to determine the relocalization result, improving the localization accuracy.

[0013] In some embodiments, the relocalization result is obtained by the following method: inputting the image to be queried into the fourth screening model to obtain image feature points; performing feature point matching between the image feature points and the recalled images in the target candidate set to obtain a matching image; and obtaining the relocalization result by verifying the geometric consistency between the matching image and the image to be queried; wherein the target candidate set is the first candidate set, or is composed of the first candidate set and the third candidate set determined by the third screening model, or is composed of the second candidate set determined by the second screening model and the third candidate set determined by the third screening model.

[0014] In some embodiments of the present application, the fourth screening model is used to extract, match, and verify the feature points of the image to be queried to determine the relocalization result, improving the relocalization accuracy.

[0015] In some embodiments, obtaining the relevant parameters of the first candidate set recalled by the first screening model and similar to the image to be queried includes: inputting the image to be queried into the first screening model to obtain a first image feature; matching the first image feature with a visual word item library to obtain the first candidate set; calculating the recall item evaluation parameter between the first image feature and the images in the target image database; and obtaining the recall similarity between the image to be queried and the recall items.

[0016] In some embodiments of the present application, the first candidate set is matched by the first screening model and the visual word item library, and calculations are performed on the first candidate set to obtain relevant parameters, providing a screening basis for subsequent visual relocalization.

[0017] In some embodiments, determining a second candidate set from the first candidate set based on the second screening model and the image to be queried includes: processing the image to be queried using the second screening model to obtain second image features; performing similarity matching between the second image features and the images in the target image database to obtain a matching result; and taking the intersection of the matching result and the first candidate set as the second candidate set.

[0018] In some embodiments of the present application, after processing the image to be queried through the second screening model to obtain a matching result, and then combining it with the first candidate set to obtain the second candidate set, precise screening of candidate items is achieved, laying a foundation for improving the positioning accuracy subsequently.

[0019] In some embodiments, determining a third candidate set based on the third screening model and the image to be queried includes: extracting the image to be queried using the third screening model to obtain third image features; and performing similarity matching between the third image features and the images in the target image database to obtain the third candidate set.

[0020] In some embodiments of the present application, after extracting and matching the image to be queried through the third screening model, the third candidate set is determined, laying a foundation for improving the positioning accuracy subsequently.

[0021] In some embodiments, the first screening model is a bag-of-visual-words model; the second screening model is a local feature screening model; the third screening model is a global image matching model; and the fourth screening model is a feature point detection and matching model.

[0022] In a second aspect, some embodiments of the present application provide a visual relocalization device, including: an acquisition module, configured to acquire relevant parameters of a first candidate set similar to the image to be queried recalled by the first screening model, where the relevant parameters include: the number of recalled items in the first candidate set, the recall item evaluation parameter, and the recall similarity, and the recall item evaluation parameter is used to measure the consistency of the matching results between images; a determination module, configured to determine subsequent screening models based on the relevant parameters, where the subsequent screening models include at least one of the second screening model, the third screening model, and the fourth screening model; and a localization module, configured to determine the relocalization result of the image to be queried based on the subsequent screening models and the image to be queried.

[0023] In a third aspect, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any embodiment of the first aspect can be implemented.

[0024] Fourth aspect, some embodiments of the present application provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the program, the method described in any embodiment of the first aspect can be implemented.

[0025] Fifth aspect, some embodiments of the present application provide a computer program product, the computer program product includes a computer program, wherein, when the computer program is executed by a processor, the method described in any embodiment of the first aspect can be implemented. Description of the Drawings

[0026] To more clearly illustrate the technical solutions of some embodiments of the present application, the following will briefly introduce the drawings required to be used in some embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0027] Figure 1 System diagram of visual relocalization provided for some embodiments of the present application;

[0028] Figure 2 One of the method flowcharts of visual relocalization provided for some embodiments of the present application;

[0029] Figure 3 Schematic diagram of the implementation process of visual relocalization provided for some embodiments of the present application;

[0030] Figure 4 Block diagram of the device composition of visual relocalization provided for some embodiments of the present application;

[0031] Figure 5 Schematic diagram of an electronic device provided for some embodiments of the present application. Detailed Embodiments

[0032] The following will describe the technical solutions in some embodiments of the present application in combination with the drawings in some embodiments of the present application.

[0033] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0034] Most of the visual relocalization schemes in the related art are based on feature point matching or global feature matching. However, in complex environments, especially in scenes with weak texture or drastic changes, the existing schemes often face challenges such as insufficient recall rate or false matching. Traditional image matching methods mostly rely on a single global matching or local matching algorithm, and a single matching method often has difficulty in balancing the recall rate and matching accuracy. Especially in scenes with low texture or large environmental changes, it may lead to system localization failure or matching error.

[0035] As can be seen from the above related art, how to dynamically adjust the relocalization process in different environments to balance performance, recall rate, and accuracy has become a technical problem to be solved.

[0036] In view of this, some embodiments of the present application provide a method for visual relocalization. In this method, a first candidate set can be obtained through primary screening by a first screening model; then, subsequent screening models for subsequent screening and localization are determined based on the relevant parameters of the first candidate set; finally, a relocalization result that matches the query image is selected through the subsequent screening models. Some embodiments of the present application can dynamically adjust the relocalization process in different environments, adapt to different models to achieve relocalization, thereby balancing performance, recall rate, and accuracy, and having a wide range of applications.

[0037] The following will Figure 1 exemplarily elaborate on the overall composition structure of the visual relocalization system provided by some embodiments of the present application.

[0038] As Figure 1 shown, some embodiments of the present application provide a system diagram of visual relocalization. The visual relocalization system may include: a terminal 100 and a server 200. Among them, a first screening model, a second screening model, a third screening model, and a fourth screening model are pre-deployed in the server 200. The terminal 100 can send the query image to the server 200; after receiving the query image, the server 200 first performs primary screening through the first screening model and the visual vocabulary library to determine a first candidate set similar to the query image. Then, at least one of the second screening model, the third screening model, and the fourth screening model for subsequent positioning operations is determined based on the relevant parameters of the first candidate set, and a relocalization result of the query image is obtained.

[0039] In some embodiments of the present application, the terminal 100 can be a mobile terminal or a non-portable computer terminal, and the embodiments of the present application do not make specific limitations here.

[0040] In some embodiments of the present application, the first screening model is a bag-of-visual-words model; the second screening model is a local feature screening model; the third screening model is a global image matching model; the fourth screening model is a feature point detection and matching model.

[0041] For example, in some embodiments of the present application, the bag-of-words model, i.e., the BoW (Bag of Words) model, belongs to a model for image feature representation. By converting image features into a set of discrete term representations, it realizes image feature extraction and matching. The local feature screening model can be EigenPlaces. EigenPlaces is a local feature screening method based on dimensionality reduction, which uses algorithms such as principal component analysis (PCA) or t-SNE to reduce the dimensionality of the candidates recalled by the BoW model, reduce redundant information, and improve the matching accuracy. The global image matching model can be NetVLAD. NetVLAD is a global image matching algorithm based on deep learning, which realizes image matching under different perspectives or environments by calculating the global feature representation of the image. The feature point detection and matching model is SuperPoint+SuperGlue. Among them, SuperPoint is a key point detection and description algorithm based on deep learning, which is used to extract stable feature points in the image. SuperGlue is a deep learning method for feature point matching. Combining the feature points provided by SuperPoint, it generates accurate feature matching pairs and performs geometric consistency checks.

[0042] It should be noted that the types of the first screening model, the second screening model, the third screening model, and the fourth screening model can be adjusted according to the actual application scenario, and the embodiments of the present application are not limited to the above embodiments.

[0043] The following Figure 2 exemplarily elaborates the implementation process of visual relocalization executed by the server 200 provided by some embodiments of the present application.

[0044] Please refer to the Figure 2 , Figure 2 which is a flowchart of a method for visual relocalization provided by some embodiments of the present application. The method for visual relocalization may include: S210, obtaining relevant parameters of a first candidate set similar to the image to be queried recalled by the first screening model, where the relevant parameters include: the number of recalled items, the recall item evaluation parameter, and the recall similarity in the first candidate set, and the recall item evaluation parameter is used to measure the consistency of the matching results between images. S220, determining a subsequent screening model based on the relevant parameters, where the subsequent screening model includes at least one of the second screening model, the third screening model, and the fourth screening model. S230, determining the relocalization result of the image to be queried based on the subsequent screening model and the image to be queried.

[0045] For example, in some embodiments of this application, the Bow model is first used for initial screening to select a first set of candidates similar to the query image from the visual term library. The parameters of the first set of candidates are then dynamically used to determine whether to use the multi-stage screening process of EigenPlaces, NetVLAD, or SuperPoint+SuperGlue to obtain the relocalization results, thereby achieving a balance between precision and recall in visual relocalization.

[0046] During the initial screening phase of the Bow of Objects (BOW), the visual vocabulary library (as a specific example of a visual term library) must first be loaded. This visual vocabulary library contains a large vocabulary of visual word vectors. A visual vocabulary library is a collection used in image feature representation. A "vocabulary" refers to local feature patterns in an image. Each "visual word" represents a group of local image patches with similar characteristics (such as SIFT and ORB features). For example, given a set of street scene images, the visual vocabulary library might include the following feature patterns: "window": composed of a cluster of local feature points with window-like shapes. "wheel": composed of a group of circular-like areas with strong black-white contrast. "sign": composed of a group of feature points with bright colors and sharp edges. These "visual words" are clustered (e.g., using the K-means algorithm) to form a collection of feature vectors that summarize the key visual elements of a specific scene. These feature vectors are pre-trained and can be directly downloaded to the "Universal Scene Bow."

[0047] The above process is described below as an example.

[0048] In some embodiments of the present application, S210 may include: inputting the image to be queried into the first screening model to obtain a first image feature; matching the first image feature with the visual term library to obtain the first candidate set; calculating the recall item evaluation parameter between the first image feature and the image in the target image database; and obtaining the recall similarity between the image to be queried and the recall item.

[0049] For example, in some embodiments of the present application, the BoW model is used to extract the features of the input query image to be queried and convert them into a BoW representation, obtaining the BoW image features (as a specific example of the first image features). Using a similarity algorithm, the BoW image features are matched with the terms in the visual vocabulary library, and the terms with a matching value greater than the set threshold are selected to achieve the purpose of screening out the matching candidates (as a specific example of the first candidate set) from the visual vocabulary library. Then, the relevant parameters related to the candidates are statistically calculated. For example, the number of recalled terms (i.e., the number of recalled items), the standard deviation of the recalled items (as a specific example of the recalled item evaluation parameter), and the recall similarity. Among them, the number of recalled terms can measure the texture complexity of the image. During recall, each query image to be queried will be converted into a "term frequency histogram" by the BoW model, and the "number of recalled terms" refers to the number of non-zero terms in this histogram, indicating which visual words in the visual vocabulary library appear in this query image to be queried.

[0050] In addition, through the candidates screened above, the target image database related to them is screened out from the original database, where the target image database contains multiple image data. The feature similarity between each image data in the target image database and the BoW image of the query image is calculated and the standard deviation is solved to obtain the recalled item evaluation parameter. The recalled item evaluation parameter can measure the consistency of the matching result. By solving the mean value of the BoW feature similarity between the recalled terms and the query image, the recall similarity is obtained. The recall similarity can analyze the similarity distribution of the recalled terms and evaluate the stability of the image quality.

[0051] In some embodiments of the present application, S220 may include: when the number of recalled items is less than the first threshold, the subsequent screening model is the fourth screening model. S230 may include: based on the fourth screening model and the query image, the relocalization result is matched from the first candidate set.

[0052] For example, in some embodiments of the present application, when the number of recalled items is less than the set first threshold, at this time, there are fewer candidates and the image features are scarce, and directly enter the SuperPoint+SuperGlue local matching. The SuperPoint+SuperGlue model can extract and analyze the query image to determine the relocalization result from the candidates.

[0053] In some embodiments of the present application, S220 may include: when the number of recalled items is not less than a first threshold, and the recalled item evaluation parameter and the recall similarity meet a first condition, the subsequent screening models are a second screening model, a third screening model, and a fourth screening model; the first condition is that the recalled item evaluation parameter is less than a second threshold and the recall similarity is greater than a third threshold; S230 may include: determining a second candidate set from the first candidate set based on the second screening model and the image to be queried; determining a third candidate set based on the third screening model and the image to be queried; and matching the relocalization result from the second candidate set and the third candidate set based on the fourth screening model and the image to be queried.

[0054] For example, in some embodiments of the present application, when the number of recalled items is less than a set first threshold, the standard deviation σ of the recalled items is less than a second threshold, and the recall similarity μ is greater than a third threshold, it indicates that the image features are good. At this time, screening can be performed layer by layer in the order of EigenPlaces, NetVLAD, and SuperPoint+SuperGlue to determine the relocalization result.

[0055] In some embodiments of the present application, S220 may include: when the number of recalled items is not less than a first threshold, and the recalled item evaluation parameter and the recall similarity meet a second condition, the subsequent screening models are a third screening model and a fourth screening model; the second condition is that the recalled item evaluation parameter is not less than a second threshold and the recall similarity is not greater than a third threshold; S230 may include: determining a third candidate set based on the third screening model and the image to be queried; and matching the relocalization result from the first candidate set and the third candidate set based on the fourth screening model and the image to be queried.

[0056] For example, in some embodiments of the present application, when the number of recalled items is not less than a set first threshold, the standard deviation σ of the recalled items is not less than a second threshold σ_th, and the recall similarity μ is not greater than a third threshold μ_th, it indicates that the image may have inconsistent candidate items due to low quality or drastic scene changes. Skip EigenPlaces and directly perform screening layer by layer in the order of NetVLAD and SuperPoint+SuperGlue to determine the relocalization result.

[0057] It can be understood that the values of the above first threshold, second threshold, and third threshold can be set according to the actual application scenario, and the embodiments of the present application do not make specific limitations here.

[0058] The processing process of each screening model is described below by way of example.

[0059] In some embodiments of the present application, the processing process of the second screening model includes: processing the image to be queried by using the second screening model to obtain second image features; performing similarity matching between the second image features and the images in the target image database to obtain a matching result; and taking the intersection of the matching result and the first candidate set as the second candidate set.

[0060] For example, in some embodiments of the present application, EigenPlaces uses PCA and t-SNE to perform dimensionality reduction processing on the image to be queried to obtain second image features, so as to reduce redundant information and retain important features. Through similarity matching between the second image features and the images in the target image database, image matching results with similarity exceeding the threshold are screened out. The intersection of the image matching results and the candidates screened by BoW is used as the EigenPlaces candidate set (as a specific example of the second candidate set).

[0061] In some embodiments of the present application, the processing process of the third screening model includes: extracting the image to be queried by using the third screening model to obtain third image features; and performing similarity matching between the third image features and the images in the target image database to obtain the third candidate set.

[0062] For example, in some embodiments of the present application, a pre-trained NetVLAD model is loaded for global matching. The image to be queried is input into the NetVLAD model for NetVLAD feature extraction to obtain NetVLAD features (as a specific example of the third image features). The NetVLAD features are respectively subjected to global matching with the images in the target image database to obtain a matching similarity, and the images with a matching similarity higher than the threshold are used as the third candidate set.

[0063] In some embodiments of the present application, the relocalization result is obtained by the following method: inputting the image to be queried into the fourth screening model to obtain image feature points; performing feature point matching between the image feature points and the recalled images in the target candidate set to obtain matching images; and verifying the geometric consistency between the matching images and the image to be queried to obtain the relocalization result; wherein the target candidate set is the first candidate set, or is composed of the first candidate set and the third candidate set determined by the third screening model, or is composed of the second candidate set determined by the second screening model and the third candidate set determined by the third screening model.

[0064] For example, in some embodiments of the present application, SuperPoint is used to extract the query image to obtain image feature points. SuperGlue is used to perform feature point matching on the query image and the candidate images in the target candidate set obtained through the preliminary screening stage to obtain the matching images. Finally, a geometric verification method, such as Ransac to count the inliers, is used to verify the geometric consistency between the query image and the matching images, exclude the incorrect matches, and determine the relocalization result. That is to say, what remains in this step are the query image and the candidate images in the finally screened database. Using superpoint + superglue for local matching, the local matching result of the query image in the entire target image database can be directly calculated, and then the relocalization result of the query image can be calculated.

[0065] The following will Figure 3 exemplarily elaborate on the specific process of visual relocalization provided by some embodiments of the present application.

[0066] Please refer to the Figure 3 , Figure 3 schematic diagram of the implementation process of a visual relocalization provided by some embodiments of the present application.

[0067] The following will exemplarily elaborate on the above process.

[0068] In the first stage, the BoW model is used to perform preliminary screening on the query image to obtain the first candidate set.

[0069] If there are too few relevant recall items in the first candidate set (i.e., the number of recall items N is less than the first threshold Nmin), then directly enter the superpoint + superglue local geometric matching stage (i.e., the fourth stage). If the number of recall items is greater than or equal to the first threshold, and the recall evaluation parameter and recall similarity satisfy the first condition (i.e., Figure 3 σ < σ_th and μ > μ_th in Figure 3 ), then enter the EigenPlaces screening in the second stage. If the number of recall items is greater than or equal to the first threshold, and the recall evaluation parameter and recall similarity satisfy the second condition (i.e.,

[0070] σ ≥ σ_th and μ ≤ μ_th in

[0071] ), then enter the NetVLAD global matching in the third stage. Figure 3 The second candidate set is the data output by the EigenPlaces screening in

[0072] In the third stage, NetVLAD is used to perform global matching on the query image and the images in the target image database to obtain the third candidate set.

[0073] Among them, the third candidate set is the data input to superpoint+superglue.

[0074] In the fourth stage, superpoint+superglue is used to perform local matching and geometric consistency check on the query image and the above-screened candidates, and the final matching result is output after completing the geometric consistency verification (as a specific example of the relocalization result).

[0075] From the above four stages, it can be seen that the entire process of the embodiment of the present application is a process of gradually screening out the candidate sets to simplify the subsequent quantity scale, and minimizing the chance of incorrect matching during the process to improve the accuracy of the final relocalization.

[0076] In addition, the embodiments of the present application can be applied in the following multiple fields:

[0077] 1) Robot positioning and navigation: In a complex environment, by optimizing visual relocalization in the present application, the positioning accuracy of the robot in a weak texture or dynamic scene is improved.

[0078] 2) Autonomous driving: Through large-scale maps and multi-stage screening, the positioning and relocalization accuracy of vehicles in complex urban environments is improved, and the risk of incorrect matching is reduced.

[0079] 3) Augmented reality (AR): By optimizing the visual relocalization process, the stability and interaction experience of the AR system in various scenarios are improved.

[0080] 4) UAV positioning: By using the present application, the positioning accuracy of UAVs in a dynamic environment can be improved, and the occurrence of incorrect matching and positioning failure can be reduced.

[0081] From some embodiments of the present application above, it can be seen that the present application dynamically determines whether to perform EigenPlaces or NetVLAD optimization through the statistical information in the BoW stage, effectively improving the accuracy and performance of visual relocalization. Through the multi-stage screening strategy combining BoW, EigenPlaces, NetVLAD, and SuperPoint+SuperGlue for multi-stage screening, a high recall rate is ensured while avoiding incorrect matching, meeting the requirements of different scenarios. Through the combination of NetVLAD global matching and SuperPoint+SuperGlue local matching, the final geometric accuracy is ensured. The present application integrates the above multi-stage screening into a complete visual relocalization system, providing an efficient and accurate solution, which is widely applicable to positioning and navigation tasks in complex environments.

[0082] Please refer to Figure 4 , Figure 4 The following is a block diagram illustrating the components of a visual relocalization apparatus provided in some embodiments of the present application. It should be understood that the visual relocalization apparatus corresponds to the aforementioned method embodiments and is capable of executing each of the steps involved in the aforementioned method embodiments. The specific functions of the visual relocalization apparatus can be found in the description above, and a detailed description is omitted here to avoid repetition.

[0083] Figure 4 The visual repositioning device includes at least one software functional module that can be stored in a memory in the form of software or firmware or solidified in the visual repositioning device, and the visual repositioning device includes: an acquisition module 410, used to obtain relevant parameters of a first candidate set similar to the image to be queried and recalled by a first screening model, wherein the relevant parameters include: the number of recalled items in the first candidate set, recall item evaluation parameters and recall similarity, and the recall item evaluation parameters are used to measure the consistency of the matching results between images; a determination module 420, used to determine a subsequent screening model based on the relevant parameters, wherein the subsequent screening model includes at least one of a second screening model, a third screening model and a fourth screening model; a positioning module 430, used to determine the repositioning result of the image to be queried based on the subsequent screening model and the image to be queried.

[0084] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.

[0085] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the operations corresponding to any of the above methods provided in the above embodiments.

[0086] Some embodiments of the present application further provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operations corresponding to any of the above methods provided in the above embodiments.

[0087] like Figure 5 As shown, some embodiments of the present application provide an electronic device 500, which includes: a memory 510, a processor 520, and a computer program stored in the memory 510 and executable on the processor 520, wherein the processor 520 can implement a method as described in any of the above embodiments when reading the program from the memory 510 through the bus 530 and executing the program.

[0088] The processor 520 can process digital signals and can include various computing architectures, such as a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements a combination of multiple instruction sets. In some examples, the processor 520 can be a microprocessor.

[0089] The memory 510 can be used to store instructions executed by the processor 520 or data related to the instruction execution process. These instructions and / or data can include code for implementing some or all of the functions of one or more modules described in the embodiments of the present application. The processor 520 in the embodiments of the present disclosure can be used to execute the instructions in the memory 510 to implement the methods shown above. The memory 510 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.

[0090] The above are only the embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0091] As mentioned above, the above are only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0092] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

Claims

1. A method for visual relocalization, characterized in that, Including: Obtaining relevant parameters of a first candidate set similar to a query image recalled by a first screening model, where the relevant parameters include: the number of recalled items in the first candidate set, a recalled item evaluation parameter, and a recall similarity, and the recalled item evaluation parameter is used to measure the consistency of the matching results between images; the first screening model is used to extract and match features of the query image; Determining a subsequent screening model based on the relevant parameters, where the subsequent screening model includes at least one of a second screening model, a third screening model, and a fourth screening model; the second screening model is used to perform dimensionality reduction local feature screening on the query image; the third screening model is used to extract and match global features of the query image; the fourth screening model is used to perform global feature point matching and geometric consistency checking on the query image; Determining a relocalization result of the query image based on the subsequent screening model and the query image.

2. The method according to claim 1, characterized in that The determining the subsequent screening model based on the relevant parameters includes: In a case where the number of recalled items is less than a first threshold, the subsequent screening model is the fourth screening model; The determining the relocalization result of the query image based on the subsequent screening model and the query image includes: Based on the fourth screening model and the query image, matching to the relocalization result from the first candidate set.

3. The method according to claim 1, characterized in that, The determining the subsequent screening model based on the relevant parameters includes: In a case where the number of recalled items is not less than the first threshold, and the recalled item evaluation parameter and the recall similarity satisfy a first condition, the subsequent screening model is the second screening model, the third screening model, and the fourth screening model; the first condition is that the recalled item evaluation parameter is less than a second threshold and the recall similarity is greater than a third threshold; The determining the relocalization result of the query image based on the subsequent screening model and the query image includes: Based on the second screening model and the query image, determining a second candidate set from the first candidate set; Based on the third screening model and the query image, determining a third candidate set; Based on the fourth screening model and the query image, matching to the relocalization result from the second candidate set and the third candidate set.

4. The method according to claim 1, characterized in that, The determining the subsequent screening model based on the relevant parameters includes: In a case where the number of recalled items is not less than the first threshold, and the recalled item evaluation parameter and the recall similarity satisfy a second condition, the subsequent screening model is the third screening model and the fourth screening model; the second condition is that the recalled item evaluation parameter is not less than the second threshold and the recall similarity is not greater than the third threshold; The determining the relocalization result of the query image based on the subsequent screening model and the query image includes: Based on the third screening model and the query image, determining a third candidate set; Based on the fourth screening model and the query image, matching to the relocalization result from the first candidate set and the third candidate set.

5. The method according to any one of claims 2-4, characterized in that, The relocation result is obtained by the following method: Input the image to be queried into the fourth screening model to obtain image feature points; Match the image feature points with the recalled images in the target candidate set to obtain matching images; Verify the geometric consistency between the matching images and the image to be queried to obtain the relocation result; Wherein, the target candidate set is the first candidate set, or consists of the first candidate set and the third candidate set determined by the third screening model, or consists of the second candidate set determined by the second screening model and the third candidate set determined by the third screening model.

6. The method according to claim 3, wherein Determining the second candidate set from the first candidate set based on the second screening model and the image to be queried includes: Process the image to be queried using the second screening model to obtain second image features; Perform similarity matching between the second image features and the images in the target image database to obtain a matching result; Take the intersection of the matching result and the first candidate set as the second candidate set.

7. The method according to claim 3 or 4, characterized in that, Determining the third candidate set based on the third screening model and the image to be queried includes: Extract the image to be queried using the third screening model to obtain third image features; Perform similarity matching between the third image features and the images in the target image database to obtain the third candidate set.

8. The method according to any one of claims 1-4 and 6, characterized in that, Obtaining the relevant parameters of the first candidate set recalled by the first screening model that is similar to the image to be queried includes: Input the image to be queried into the first screening model to obtain first image features; Match the first image features with the visual word item library to obtain the first candidate set; Calculate the recall item evaluation parameter between the first image features and the images in the target image database; and obtain the recall similarity between the image to be queried and the recall items.

9. The method according to any one of claims 1-4, 6, characterized in that The first screening model is a bag-of-visual-words model; the second screening model is a local feature screening model; the third screening model is a global image matching model; the fourth screening model is a feature point detection and matching model.

10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor. Wherein, when the computer program is run by the processor, it executes the method described in any one of claims 1-9.

Citation Information

Patent Citations

  • Picture query method and device, electronic equipment and storage medium

    CN114860975A

  • Visual repositioning method and device, electronic equipment, storage medium and program product

    CN118747773A

Cited By

  • Visual relocation and pose estimation method based on space-time dual compression and related device

    CN121837372A