Visual repositioning method and electronic equipment
Patent Information
- Application Number
- CN202510609129.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing visual relocation technologies are difficult to take into account recall and matching accuracy in scenarios with low texture or large environmental changes, which may lead to system visual positioning failure or matching errors.
The first filter model obtains the relevant parameters of the first candidate set similar to the image to be query, including the number of recall items, the recall item evaluation parameters and the recall similarity, and dynamically determines the subsequent filter model, including the second, third and fourth filter models, and filters layer by layer to determine the relocation result.
It improves the accuracy of visual repositioning, reduces the positioning error rate, ensures the stability and efficiency of the system, and is suitable for positioning and navigation tasks in complex environments.
Smart Images

Figure CN120163877A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of visual repositioning, and in particular to a method and electronic device for visual repositioning. Background Art
[0002] Visual relocalization is a key technology in SLAM (Simultaneous Localization And Mapping) systems. Its main function is to re-determine the position and posture of the target object after tracking is lost. Currently, visual relocalization technology has been widely used in many application scenarios, such as robot navigation, augmented reality, and unmanned driving. Most existing visual relocalization schemes are based on feature point matching, while traditional image matching methods rely on a single global matching or local matching algorithm. However, a single matching method often fails to balance recall rate and matching accuracy, especially in scenes with low texture or large environmental changes, which may cause system visual positioning failure or matching errors.
[0003] Therefore, how to provide a technical solution for a method of visual relocalization with higher accuracy has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] The purpose of some embodiments of the present application is to provide a method and electronic device for visual repositioning. The technical solutions of the embodiments of the present application can improve the accuracy of visual repositioning and reduce the positioning error rate.
[0005] In a first aspect, some embodiments of the present application provide a method for visual relocalization, comprising: obtaining relevant parameters of a first candidate set similar to an image to be queried and recalled by a first screening model, wherein the relevant parameters include: the number of recalled items in the first candidate set, recall item evaluation parameters and recall similarity, and the recall item evaluation parameters are used to measure the consistency of matching results between images; based on the relevant parameters, determining a subsequent screening model, wherein the subsequent screening model includes at least one of a second screening model, a third screening model and a fourth screening model; based on the subsequent screening model and the image to be queried, determining a relocalization result of the image to be queried.
[0006] In some embodiments of the present application, after using the first screening model to recall the first candidate set similar to the image to be queried, the subsequent screening model for subsequent relocalization is determined based on the relevant parameters of the first candidate set, and then the relocalization result of the image to be queried is determined through the subsequent screening model. Some embodiments of the present application use different screening methods in different situations, and layer by layer screening to determine the relocalization result, improving the accuracy of the relocalization result and overcoming the defect of high error rate in single matching in the prior art; at the same time, ensuring the stability and efficiency of the relocalization system.
[0007] In some embodiments, determining the subsequent screening model based on the relevant parameters includes: when the number of recall items is less than the first threshold, the subsequent screening model is the fourth screening model; determining the relocalization result of the image to be queried based on the subsequent screening model and the image to be queried includes: matching the relocalization result from the first candidate set based on the fourth screening model and the image to be queried.
[0008] Some embodiments of the present application compare the number of recall items with the first threshold, confirm that the fourth screening model is used to process the image to be queried when it is less than the first threshold, and select the relocalization result from the first candidate set, which can achieve efficient and accurate relocalization in different situations.
[0009] In some embodiments, determining the subsequent screening model based on the relevant parameters includes: when the number of recall items is not less than the first threshold, and the recall item evaluation parameter and the recall similarity meet the first condition, the subsequent screening models are the second screening model, the third screening model, and the fourth screening model; the first condition is that the recall item evaluation parameter is less than the second threshold and the recall similarity is greater than the third threshold; determining the relocalization result of the image to be queried based on the subsequent screening model and the image to be queried includes: determining the second candidate set from the first candidate set based on the second screening model and the image to be queried; determining the third candidate set based on the third screening model and the image to be queried; matching the relocalization result from the second candidate set and the third candidate set based on the fourth screening model and the image to be queried.
[0010] Some embodiments of the present application analyze the number of recall items, the recall evaluation parameter, and the recall similarity, select the type of the subsequent screening model, and analyze and process the image to be queried and the first candidate set with this, so as to determine the relocalization result and improve the positioning accuracy.
[0011] In some embodiments, determining a subsequent screening model based on the relevant parameters includes: when the number of recall items is not less than a first threshold, and the recall item evaluation parameter and the recall similarity satisfy a second condition, the subsequent screening model is a third screening model and a fourth screening model; the second condition is that the recall item evaluation parameter is not less than a second threshold and the recall similarity is not greater than a third threshold; determining a relocalization result of the query image based on the subsequent screening model and the query image includes: determining a third candidate set based on the third screening model and the query image; and matching the relocalization result from the first candidate set and the third candidate set based on the fourth screening model and the query image.
[0012] Some embodiments of the present application analyze the number of recall items, the recall evaluation parameter, and the recall similarity to select the type of the subsequent screening model, and then analyze and process the query image and the first candidate set to determine the relocalization result, improving the localization accuracy.
[0013] In some embodiments, the relocalization result is obtained by the following method: inputting the query image into the fourth screening model to obtain image feature points; performing feature point matching between the image feature points and the recalled images in the target candidate set to obtain a matching image; and obtaining the relocalization result by verifying the geometric consistency between the matching image and the query image; where the target candidate set is the first candidate set, or consists of the first candidate set and the third candidate set determined by the third screening model, or consists of the second candidate set determined by the second screening model and the third candidate set determined by the third screening model.
[0014] Some embodiments of the present application extract, match, and verify feature points of the query image through the fourth screening model to determine the relocalization result, improving the relocalization accuracy.
[0015] In some embodiments, obtaining the relevant parameters of the first candidate set recalled by the first screening model and similar to the query image includes: inputting the query image into the first screening model to obtain a first image feature; matching the first image feature with a visual word item library to obtain the first candidate set; calculating the recall item evaluation parameter between the first image feature and the images in the target image database; and obtaining the recall similarity between the query image and the recall items.
[0016] Some embodiments of the present application match the first candidate set through the first screening model and the visual word item library, and calculate the relevant parameters for the first candidate set to provide a screening basis for subsequent visual relocalization.
[0017] In some embodiments, determining a second candidate set from the first candidate set based on the second screening model and the image to be queried includes: processing the image to be queried using the second screening model to obtain second image features; performing similarity matching between the second image features and the images in the target image database to obtain a matching result; and taking the intersection of the matching result and the first candidate set as the second candidate set.
[0018] In some embodiments of the present application, after processing the image to be queried through the second screening model to obtain a matching result, and then combining it with the first candidate set to obtain the second candidate set, which realizes the precise screening of candidates and lays a foundation for improving the positioning accuracy subsequently.
[0019] In some embodiments, determining a third candidate set based on the third screening model and the image to be queried includes: extracting the image to be queried using the third screening model to obtain third image features; performing similarity matching between the third image features and the images in the target image database to obtain the third candidate set.
[0020] In some embodiments of the present application, after extracting and matching the image to be queried through the third screening model to determine the third candidate set, it lays a foundation for improving the positioning accuracy subsequently.
[0021] In some embodiments, the first screening model is a bag-of-visual-words model; the second screening model is a local feature screening model; the third screening model is a global image matching model; the fourth screening model is a feature point detection and matching model.
[0022] In a second aspect, some embodiments of the present application provide a visual relocalization device, including: an acquisition module, configured to acquire relevant parameters of a first candidate set similar to the image to be queried recalled by the first screening model, where the relevant parameters include: the number of recalled items in the first candidate set, a recall item evaluation parameter, and a recall similarity, and the recall item evaluation parameter is used to measure the consistency of the matching results between images; a determination module, configured to determine subsequent screening models based on the relevant parameters, where the subsequent screening models include at least one of a second screening model, a third screening model, and a fourth screening model; and a localization module, configured to determine a relocalization result of the image to be queried based on the subsequent screening models and the image to be queried.
[0023] In a third aspect, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any embodiment of the first aspect can be implemented.
[0024] Fourth aspect, some embodiments of the present application provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any embodiment of the first aspect can be implemented.
[0025] Fifth aspect, some embodiments of the present application provide a computer program product, which includes a computer program. When the computer program is executed by a processor, the method described in any embodiment of the first aspect can be implemented. Description of the Drawings
[0026] To more clearly illustrate the technical solutions of some embodiments of the present application, the drawings required for some embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as a limitation of the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 System diagram of visual relocalization provided for some embodiments of the present application; Figure 2 One of the method flowcharts of visual relocalization provided for some embodiments of the present application; Figure 3 Schematic diagram of the implementation process of visual relocalization provided for some embodiments of the present application; Figure 4 Block diagram of the device composition of visual relocalization provided for some embodiments of the present application; Figure 5 Schematic diagram of an electronic device provided for some embodiments of the present application. Detailed Embodiments
[0028] Next, the technical solutions in some embodiments of the present application will be described in conjunction with the drawings in some embodiments of the present application.
[0029] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0030] Most of the visual relocalization solutions in the related art are based on feature point matching or global feature matching. However, in complex environments, especially in scenes with weak texture or drastic changes, existing solutions often face challenges such as insufficient recall rate or false matching. Traditional image matching methods mostly rely on a single global matching or local matching algorithm, and a single matching method often has difficulty balancing the recall rate and matching accuracy. Especially in scenes with low texture or large environmental changes, it may lead to system localization failure or matching errors.
[0031] As can be seen from the above related art, how to dynamically adjust the relocalization process in different environments to balance performance, recall rate, and accuracy has become a technical problem to be solved.
[0032] In view of this, some embodiments of the present application provide a method for visual relocalization. In this method, a first candidate set can be obtained through primary screening by a first screening model; then, a subsequent screening model for subsequent screening and localization can be determined based on the relevant parameters of the first candidate set; finally, a relocalization result that matches the image to be queried can be selected through the subsequent screening model. Some embodiments of the present application can dynamically adjust the relocalization process in different environments, adapt to different models to achieve relocalization, thereby balancing performance, recall rate, and accuracy, and having a wide range of applications.
[0033] The following combines the attached Figure 1 Exemplarily elaborates the overall composition structure of the visual relocalization system provided by some embodiments of the present application.
[0034] As Figure 1 shown, some embodiments of the present application provide a system diagram of visual relocalization. The visual relocalization system may include: a terminal 100 and a server 200. Among them, a first screening model, a second screening model, a third screening model, and a fourth screening model are pre-deployed in the server 200. The terminal 100 can send the image to be queried to the server 200; after receiving the image to be queried, the server 200 first performs primary screening through the first screening model and the visual vocabulary library to determine a first candidate set similar to the image to be queried. Then, at least one of the second screening model, the third screening model, and the fourth screening model for subsequent localization operations is determined based on the relevant parameters of the first candidate set to obtain the relocalization result of the image to be queried.
[0035] In some embodiments of the present application, the terminal 100 can be a mobile terminal or a non-portable computer terminal, and the embodiments of the present application do not make specific limitations here.
[0036] In some embodiments of the present application, the first screening model is a bag-of-visual-words model; the second screening model is a local feature screening model; the third screening model is a global image matching model; the fourth screening model is a feature point detection and matching model.
[0037] For example, in some embodiments of the present application, the bag-of-words model, i.e., the BoW (Bag of Words) model, belongs to a model for image feature representation. By converting image features into a set of discrete term representations, it realizes image feature extraction and matching. The local feature screening model can be EigenPlaces, which is a local feature screening method based on dimensionality reduction. It uses algorithms such as principal component analysis (PCA) or t-SNE to reduce the dimensionality of the candidates recalled by the BoW model, reduce redundant information, and improve the matching accuracy. The global image matching model can be NetVLAD, which is a global image matching algorithm based on deep learning. By calculating the global feature representation of the image, it realizes image matching under different perspectives or environments. The feature point detection and matching model is SuperPoint+SuperGlue. Among them, SuperPoint is a key point detection and description algorithm based on deep learning, used to extract stable feature points in the image. SuperGlue is a deep learning method for feature point matching. Combining the feature points provided by SuperPoint, it generates accurate feature matching pairs and performs geometric consistency checks.
[0038] It should be noted that the types of the first screening model, the second screening model, the third screening model, and the fourth screening model can be adjusted according to the actual application scenario, and the embodiments of the present application are not limited to the above embodiments.
[0039] The following combines the attached Figure 2 Exemplarily illustrate the implementation process of visual relocalization executed by the server 200 provided by some embodiments of the present application.
[0040] Please refer to the attached Figure 2 , Figure 2 which is a flowchart of a method for visual relocalization provided by some embodiments of the present application. The method for visual relocalization may include: S210, obtaining relevant parameters of a first candidate set similar to the query image recalled by the first screening model, where the relevant parameters include: the number of recalled items in the first candidate set, the recall item evaluation parameter, and the recall similarity, and the recall item evaluation parameter is used to measure the consistency of the matching results between images. S220, determining subsequent screening models based on the relevant parameters, where the subsequent screening models include at least one of the second screening model, the third screening model, and the fourth screening model. S230, determining the relocalization result of the query image based on the subsequent screening models and the query image.
[0041] For example, in some embodiments of the present application, the Bow model is first used for initial screening to select a first candidate set similar to the query image from the visual term library. The relevant parameters of the first candidate set are used to dynamically select whether to use the multi-stage screening process of EigenPlaces, NetVLAD or SuperPoint+SuperGlue to obtain the relocalization result, thereby achieving a balance between precision and recall in visual relocalization.
[0042] In the initial screening stage of BoW, you first need to load the visual vocabulary library (as a specific example of a visual word library), which contains a large number of visual word vectors. The visual vocabulary library is a collection used in image feature representation. "Vocabulary" refers to the local feature pattern in the image. Each "visual word" represents a group of local image blocks with similar features (such as SIFT, ORB and other features). Assuming there is a set of street scene pictures, the visual vocabulary library may contain the following feature patterns: "Window": It is composed of a group of local feature points with window shapes. "Wheel": It is composed of a group of circular-like areas with strong black and white contrast. "Signboard": It is composed of a group of feature points with bright colors and clear edges. These "visual words" are clustered (such as K-means algorithm) to form a set of feature vectors that can summarize the key visual elements in a specific scene. It is obtained through pre-training and can be directly downloaded to the "General Scene Bow".
[0043] The above process is explained below as an example.
[0044] In some embodiments of the present application, S210 may include: inputting the image to be queried into the first screening model to obtain a first image feature; matching the first image feature with a visual term library to obtain the first candidate set; calculating the recall item evaluation parameter between the first image feature and the image in the target image database; and obtaining the recall similarity between the image to be queried and the recall item.
[0045] For example, in some embodiments of the present application, the BoW model is used to extract the features of the input query image to be queried and convert them into a BoW representation, obtaining the BoW image features (as a specific example of the first image features). Using a similarity algorithm, the BoW image features are matched with the terms in the visual vocabulary library, and the terms with a matching value greater than the set threshold are selected, achieving the purpose of screening out the matching candidates (as a specific example of the first candidate set) from the visual vocabulary library. Then, the relevant parameters related to the candidates are statistically calculated. For example, the number of recalled terms (i.e., the number of recalled items), the standard deviation of the recalled items (as a specific example of the recalled item evaluation parameter), and the recall similarity. Among them, the number of recalled terms can measure the texture complexity of the image. During recall, each query image to be queried will be converted by the BoW model into a "term frequency histogram", and the "number of recalled terms" refers to the number of non-zero terms in this histogram, indicating which visual words in the visual vocabulary library appear in the query image to be queried.
[0046] In addition, through the candidates screened above, the target image database related to them is screened out from the original database, where the target image database contains multiple image data. The feature similarity between each image data in the target image database and the BoW image of the query image is calculated and the standard deviation is solved to obtain the recalled item evaluation parameter. The recalled item evaluation parameter can measure the consistency of the matching result. By solving the mean value of the BoW feature similarity between the recalled terms and the query image, the recall similarity is obtained. The recall similarity can analyze the similarity distribution of the recalled terms and evaluate the image quality stability.
[0047] In some embodiments of the present application, S220 may include: when the number of recalled items is less than the first threshold, the subsequent screening model is the fourth screening model. S230 may include: based on the fourth screening model and the query image, the relocalization result is matched from the first candidate set.
[0048] For example, in some embodiments of the present application, when the number of recalled items is less than the set first threshold, at this time, there are fewer candidates and the image features are scarce, and directly enter the SuperPoint+SuperGlue local matching. The SuperPoint+SuperGlue model can extract and analyze the query image and determine the relocalization result from the candidates.
[0049] In some embodiments of the present application, S220 may include: when the number of recalled items is not less than a first threshold, and the recalled item evaluation parameter and the recall similarity meet a first condition, the subsequent screening models are a second screening model, a third screening model, and a fourth screening model; the first condition is that the recalled item evaluation parameter is less than a second threshold and the recall similarity is greater than a third threshold; S230 may include: determining a second candidate set from the first candidate set based on the second screening model and the query image; determining a third candidate set based on the third screening model and the query image; and matching the relocalization result from the second candidate set and the third candidate set based on the fourth screening model and the query image.
[0050] For example, in some embodiments of the present application, when the number of recalled items is less than a set first threshold, the standard deviation σ of the recalled items is less than a second threshold, and the recall similarity μ is greater than a third threshold, it indicates that the image features are good. At this time, screening can be performed layer by layer in the order of EigenPlaces, NetVLAD, and SuperPoint+SuperGlue to determine the relocalization result.
[0051] In some embodiments of the present application, S220 may include: when the number of recalled items is not less than a first threshold, and the recalled item evaluation parameter and the recall similarity meet a second condition, the subsequent screening models are a third screening model and a fourth screening model; the second condition is that the recalled item evaluation parameter is not less than a second threshold and the recall similarity is not greater than a third threshold; S230 may include: determining a third candidate set based on the third screening model and the query image; and matching the relocalization result from the first candidate set and the third candidate set based on the fourth screening model and the query image.
[0052] For example, in some embodiments of the present application, when the number of recalled items is not less than a set first threshold, the standard deviation σ of the recalled items is not less than a second threshold σ_th, and the recall similarity μ is not greater than a third threshold μ_th, it indicates that the image may have inconsistent candidate items due to low quality or drastic scene changes. Skip EigenPlaces and directly perform layer-by-layer screening in the order of NetVLAD and SuperPoint+SuperGlue to determine the relocalization result.
[0053] It can be understood that the values of the above first threshold, second threshold, and third threshold can be set according to the actual application scenario, and the embodiments of the present application do not make specific limitations here.
[0054] The processing process of each screening model is described below by way of example.
[0055] In some embodiments of the present application, the processing process of the second screening model includes: processing the to-be-query image by using the second screening model to obtain second image features; performing similarity matching between the second image features and the images in the target image database to obtain a matching result; taking the intersection of the matching result and the first candidate set as the second candidate set.
[0056] For example, in some embodiments of the present application, EigenPlaces uses PCA and t-SNE to perform dimensionality reduction processing on the to-be-query image to obtain second image features, so as to reduce redundant information and retain important features. Through similarity matching between the second image features and the images in the target image database, image matching results with similarity exceeding the threshold are screened out. Taking the intersection of the image matching results and the candidates screened by BoW as the EigenPlaces candidate set (as a specific example of the second candidate set).
[0057] In some embodiments of the present application, the processing process of the third screening model includes: extracting the to-be-query image by using the third screening model to obtain third image features; performing similarity matching between the third image features and the images in the target image database to obtain the third candidate set.
[0058] For example, in some embodiments of the present application, a pre-trained NetVLAD model is loaded for global matching. The to-be-query image is input into the NetVLAD model for NetVLAD feature extraction to obtain NetVLAD features (as a specific example of the third image features). The NetVLAD features are respectively subjected to global matching with the images in the target image database to obtain matching similarities, and the images with matching similarities higher than the threshold are used as the third candidate set.
[0059] In some embodiments of the present application, the repositioning result is obtained by the following method: inputting the to-be-query image into the fourth screening model to obtain image feature points; performing feature point matching between the image feature points and the recalled images in the target candidate set to obtain matching images; verifying the geometric consistency between the matching images and the to-be-query image to obtain the repositioning result; wherein, the target candidate set is the first candidate set, or is composed of the first candidate set and the third candidate set determined by the third screening model, or is composed of the second candidate set determined by the second screening model and the third candidate set determined by the third screening model.
[0060] For example, in some embodiments of the present application, SuperPoint is used to extract the query image to obtain image feature points. SuperGlue is used to perform feature point matching on the query image and the candidate images in the target candidate set obtained through the previous screening stage to obtain matching images. Finally, a geometric verification method, such as Ransac to count inliers, is used to verify the geometric consistency between the query image and the matching image, exclude incorrect matches, and determine the relocalization result. That is to say, what remains in this step are the query image and the candidate images in the finally screened database. Using superpoint+superglue for local matching can directly calculate the local matching result of the query image in the entire target image database, and then the relocalization result of the query image can be calculated.
[0061] The following will Figure 3 exemplarily illustrate the specific process of visual relocalization provided by some embodiments of the present application.
[0062] Please refer to the Figure 3 , Figure 3 which is a schematic diagram of the implementation process of a visual relocalization provided by some embodiments of the present application.
[0063] The following will exemplarily illustrate the above process.
[0064] In the first stage, the BoW model is used to perform a preliminary screening on the query image to obtain the first candidate set.
[0065] If there are too few relevant recall items in the first candidate set (i.e., the number of recall items N is less than the first threshold Nmin), then directly enter the superpoint+superglue local geometric matching stage (i.e., the fourth stage). If the number of recall items is greater than or equal to the first threshold, and the recall evaluation parameter and recall similarity meet the first condition (i.e., Figure 3 the σ in is less than σ_th and μ is greater than μ_th), then enter the EigenPlaces screening in the second stage. If the number of recall items is greater than or equal to the first threshold, and the recall evaluation parameter and recall similarity meet the second condition (i.e., Figure 3 the σ in is greater than or equal to σ_th and μ is less than or equal to μ_th), then enter the NetVLAD global matching in the third stage.
[0066] In the second stage, the EigenPlaces model is used to perform global feature dimensionality reduction screening to optimize the first candidate set and obtain the second candidate set.
[0067] Among them, the second candidate set is the Figure 3 data output by the EigenPlaces screening in and is used as the input to the NetVLAD global feature matching.
[0068] In the third stage, NetVLAD is used to perform global matching on the query image and the images in the target image database to obtain the third candidate set.
[0069] Among them, the third candidate set is the data input to superpoint+superglue.
[0070] In the fourth stage, superpoint+superglue is used to perform local matching and geometric consistency check on the query image and the above-screened candidates, and the final matching result (as a specific example of the relocalization result) is output after completing the geometric consistency verification.
[0071] As can be seen from the above four stages, the entire process of the embodiment of the present application is a process of gradually screening out the candidate sets to simplify the subsequent quantity scale, and minimizing the chance of incorrect matching during the process to improve the accuracy of the final relocalization.
[0072] In addition, the embodiments of the present application can be applied in the following multiple fields: 1) Robot positioning and navigation: In a complex environment, by optimizing visual relocalization in the present application, the positioning accuracy of the robot in a weak texture or dynamic scene can be improved.
[0073] 2) Autonomous driving: Through large-scale maps and multi-stage screening, the positioning and relocalization accuracy of vehicles in complex urban environments can be improved, and the risk of incorrect matching can be reduced.
[0074] 3) Augmented reality (AR): By optimizing the visual relocalization process, the stability and interaction experience of the AR system in various scenarios can be improved.
[0075] 4) UAV positioning: By using the present application, the positioning accuracy of UAVs in a dynamic environment can be improved, and the occurrence of incorrect matching and positioning failure can be reduced.
[0076] As can be seen from some embodiments of the present application above, the present application dynamically determines whether to perform EigenPlaces or NetVLAD optimization through the statistical information in the BoW stage, effectively improving the accuracy and performance of visual relocalization. Through the multi-stage screening strategy combined with the multi-stage screening of BoW, EigenPlaces, NetVLAD, and SuperPoint+SuperGlue, while ensuring a high recall rate, incorrect matching is avoided, meeting the requirements of different scenarios. Through the combination of NetVLAD global matching and SuperPoint+SuperGlue local matching, the final geometric accuracy is ensured. The present application integrates the above multi-stage screening into a complete visual relocalization system, providing an efficient and accurate solution, which is widely applicable to positioning and navigation tasks in complex environments.
[0077] Please refer toFigure 4 , Figure 4 The block diagram of the visual relocalization device provided by some embodiments of the present application is shown. It should be understood that this visual relocalization device corresponds to the above method embodiments and can execute each step involved in the above method embodiments. The specific functions of this visual relocalization device can be referred to the descriptions above. To avoid repetition, the detailed descriptions are appropriately omitted here.
[0078] Figure 4 The visual relocalization device includes at least one software functional module that can be stored in the memory in the form of software or firmware or solidified in the visual relocalization device. The visual relocalization device includes: an acquisition module 410, configured to acquire relevant parameters of a first candidate set similar to the query image recalled by a first screening model, where the relevant parameters include: the number of recalled items in the first candidate set, the recall item evaluation parameter, and the recall similarity, and the recall item evaluation parameter is used to measure the consistency of the matching results between images; a determination module 420, configured to determine a subsequent screening model based on the relevant parameters, where the subsequent screening model includes at least one of a second screening model, a third screening model, and a fourth screening model; and a localization module 430, configured to determine the relocalization result of the query image based on the subsequent screening model and the query image.
[0079] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method, and will not be elaborated here too much.
[0080] Some embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the operations corresponding to any of the methods in the above methods provided by the above embodiments can be implemented.
[0081] Some embodiments of the present application also provide a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the operations corresponding to any of the methods in the above methods provided by the above embodiments can be implemented.
[0082] As Figure 5 shown, some embodiments of the present application provide an electronic device 500, and the electronic device 500 includes: a memory 510, a processor 520, and a computer program stored on the memory 510 and executable on the processor 520. When the processor 520 reads the program from the memory 510 through a bus 530 and executes the program, the methods of any of the above embodiments can be implemented.
[0083] The processor 520 can process digital signals and can include various computing architectures. For example, a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements a combination of multiple instruction sets. In some examples, the processor 520 can be a microprocessor.
[0084] The memory 510 can be used to store instructions executed by the processor 520 or data related to the instruction execution process. These instructions and / or data can include code for implementing some or all of the functions of one or more modules described in the embodiments of the present application. The processor 520 of the embodiments of the present disclosure can be used to execute the instructions in the memory 510 to implement the methods shown above. The memory 510 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.
[0085] The above are only the embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0086] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0087] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
Claims
1. A method of visual relocalization, characterized in that: include: Obtaining relevant parameters of a first candidate set similar to the query image recalled by the first screening model, wherein the relevant parameters include: the number of recalled items in the first candidate set, a recalled item evaluation parameter, and a recalled similarity, wherein the recalled item evaluation parameter is used to measure the consistency of the matching results between images; Determining a subsequent screening model based on the relevant parameters, wherein the subsequent screening model includes at least one of a second screening model, a third screening model, and a fourth screening model; Based on the subsequent screening model and the image to be queried, a relocation result of the image to be queried is determined.
2. The method according to claim 1, characterized in that Determining a subsequent screening model based on the relevant parameters includes: When the number of recalled items is less than the first threshold, the subsequent screening model is the fourth screening model; The step of determining a relocation result of the image to be queried based on the subsequent screening model and the image to be queried includes: Based on the fourth screening model and the query image, the relocation result is matched from the first candidate set.
3. The method according to claim 1, characterized in that Determining a subsequent screening model based on the relevant parameters includes: When the number of recalled items is not less than the first threshold, and the recalled item evaluation parameter and the recalled similarity meet the first condition, the subsequent screening models are the second screening model, the third screening model and the fourth screening model; the first condition is that the recalled item evaluation parameter is less than the second threshold and the recalled similarity is greater than the third threshold; The step of determining a relocation result of the image to be queried based on the subsequent screening model and the image to be queried includes: Determine a second candidate set from the first candidate set based on the second screening model and the query image; Determine a third candidate set based on the third screening model and the image to be queried; Based on the fourth screening model and the query image, the relocation result is matched from the second candidate set and the third candidate set.
4. The method according to claim 1, characterized in that Determining a subsequent screening model based on the relevant parameters includes: When the number of recalled items is not less than the first threshold, and the recalled item evaluation parameter and the recalled similarity meet the second condition, the subsequent screening model is the third screening model and the fourth screening model; the second condition is that the recalled item evaluation parameter is not less than the second threshold and the recalled similarity is not greater than the third threshold; The step of determining a relocation result of the image to be queried based on the subsequent screening model and the image to be queried includes: Determine a third candidate set based on the third screening model and the image to be queried; Based on the fourth screening model and the query image, the relocation result is matched from the first candidate set and the third candidate set.
5. The method according to any one of claims 2 to 4, characterized in that: The relocation result is obtained by the following method: Inputting the query image into the fourth screening model to obtain image feature points; Matching the image feature points with the recalled images in the target candidate set to obtain a matching image; Obtaining the relocation result by verifying the geometric consistency between the matching image and the image to be queried; Among them, the target candidate set is the first candidate set, or is composed of the first candidate set and the third candidate set determined by the third screening model, or is composed of the second candidate set determined by the second screening model and the third candidate set determined by the third screening model.
6. The method according to claim 3, characterized in that The step of determining a second candidate set from the first candidate set based on the second screening model and the query image includes: Processing the query image using the second screening model to obtain a second image feature; Performing similarity matching between the second image feature and an image in a target image database to obtain a matching result; The intersection of the matching result and the first candidate set is used as the second candidate set.
7. The method according to claim 3 or 4, characterized in that The step of determining a third candidate set based on the third screening model and the image to be queried includes: Extracting the query image using the third screening model to obtain a third image feature; The third image feature is matched with images in a target image database to obtain the third candidate set.
8. The method according to any one of claims 1 to 4 and 6, characterized in that: The step of obtaining relevant parameters of a first candidate set similar to the query image and recalled by the first screening model includes: Inputting the query image into the first screening model to obtain a first image feature; Matching the first image feature with the visual term library to obtain the first candidate set; The recall item evaluation parameter between the first image feature and an image in a target image database is calculated; and the recall similarity between the image to be queried and the recall item is obtained.
9. The method according to any one of claims 1 to 4 and 6, characterized in that: The first screening model is a visual word bag model; the second screening model is a local feature screening model; the third screening model is a global image matching model; and the fourth screening model is a feature point detection and matching model.
10. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the computer program executes the method according to any one of claims 1 to 9 when being run by the processor.
Citation Information
Patent Citations
Picture query method and device, electronic equipment and storage medium
CN114860975A
Visual repositioning method and device, electronic equipment, storage medium and program product
CN118747773A
Identifying content related to a visual search query
US11354349B1
Method and apparatus for retrieving image, device, and medium
US20210209408A1
Positioning method and apparatus, and electronic device
WO2023179523A1