Two-stage location identification method based on similarity constraint deep hash
By adopting the two-stage recognition paradigm of similarity-constrained deep hashing in the visual position recognition method, using binary hash features for rapid search and floating-point features for precise reordering, the shortcomings of existing methods in terms of search delay and storage consumption are solved, and efficient and accurate visual position recognition is achieved.
Patent Information
- Application Number
- CN202510115667.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing visual position recognition methods have shortcomings in retrieval delay and storage consumption, especially as models and training data sets continue to expand.
Using a deep hashing method based on similarity constraints, triple training samples are constructed through a triple training framework, image features are extracted and dimension reduction and dimension increase processing are performed to generate binary hashing features and floating-point features. The first phase is quickly searched using binary hash features, the second phase is precisely reordered using floating point features, and the model parameters are optimized through weak supervision methods.
It significantly reduces the search delay and storage space of the visual position recognition method, improves Rank-1 accuracy and average accuracy (mAP), and improves the model's retrieval efficiency and memory usage efficiency.
Smart Images

Figure CN120045737A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and deep learning, and particularly to a two-stage place recognition method based on similarity-constrained deep hashing. Background Art
[0002] Visual place recognition (VPR), also known as image localization or visual geolocation, aims to roughly estimate the location of a query location image by searching for the best match of the query location image from a database of geotagged images. VPR has a wide range of applications in fields such as mobile robot localization and augmented reality.
[0003] VPR tasks usually use global or local descriptors to represent image features. Aggregation algorithms such as VLAD are usually used to aggregate local features into a vector as a global feature. This compact global feature helps with fast location retrieval and is robust to viewpoint changes. However, these global features ignore spatial information, making VPR methods based on them prone to perceptual aliasing. The current solution is two-stage VPR, that is, using global features to retrieve the top k in the database as candidate results, and then performing spatial verification to reorder these candidate results by matching local features. The training of VPR models follows the "pre-train then fine-tune" pattern. Most VPR models are initialized with model parameters pre-trained on the ImageNet dataset and fine-tuned on VPR datasets (such as Mapillary street-level sequences, MSLS). As the models and training datasets continue to expand, two-stage training becomes more expensive in terms of computation and memory usage.
[0004] It should be noted that the information disclosed in the above background art section is only for understanding the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the defects existing in the above background art, and to provide a two-stage place recognition method based on similarity-constrained deep hashing.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A two-stage place recognition method based on similarity-constrained deep hashing, comprising the following steps:
[0008] S1. Based on a triplet training framework, construct triplet training samples including query images, positive example images, and negative example images; input the training samples into a feature extraction network to extract image features for two-stage retrieval;
[0009] S2. Perform dimensionality reduction and dimensionality increase processing on the high-dimensional semantic feature map output by the feature extraction network; convert the dimensionality-reduced features into binary hash features through a deep hashing method with similarity constraints, and simultaneously obtain the floating-point features after dimensionality increase.
[0010] S3. Use the binary hash features to perform the first-stage fast retrieval in the image database, and screen out the top k candidate images with features similar to the query image; use the floating-point features to perform an accurate re-ranking of the candidate images in the second stage, and determine the final top k images that are most similar to the query image.
[0011] S4. Adopt a weakly supervised method to calculate the overall loss function that includes the quantization loss and the metric loss with similarity constraints; iteratively optimize the model parameters until the preset convergence condition is met to achieve the end-to-end training of the model.
[0012] Furthermore, in step S1, the construction of the triplet training samples is based on geographical location tags. Each query image corresponds to several positive example images and negative example images, and the model training is optimized through an online hard mining strategy.
[0013] Furthermore, in step S1, the feature extraction network uses a pre-trained vision large model to extract a high-dimensional semantic feature map, and aggregates the high-dimensional semantic feature map into a one-dimensional feature vector through generalized mean pooling (Gempooling).
[0014] Furthermore, in step S2, the high-dimensional semantic feature map is processed for dimensionality reduction and dimensionality increase through a linear layer and a fully connected layer, and then binary hash features and floating-point features are generated respectively.
[0015] Furthermore, in step S2, the deep hashing method with similarity constraints constrains the similarity between the binary hash features and the floating-point features through the quantization loss, and realizes end-to-end training in combination with the straight-through estimator (STE).
[0016] Furthermore, in step S3, the first-stage retrieval screens out the top k candidate images with features closest to the query image through the fast matching of binary hash features, and the second-stage re-ranking determines the final top k images that are most similar to the query image by calculating the similarity metric between the floating-point features.
[0017] Furthermore, in step S4, the overall loss function combines the quantization loss and the metric loss through weighted summation, and iteratively optimizes the model parameters through the backpropagation algorithm. The convergence condition for model optimization is that the error value calculated by the loss function is less than the preset threshold, or the number of iterations reaches the predetermined value. The weakly supervised method improves the retrieval accuracy and efficiency of the model through the joint optimization of the metric loss and the quantization loss.
[0018] Further, in step S4, the metric loss includes a multi-similarity loss (MS loss) for optimizing the similarity metric between image features.
[0019] Further, in step S4, the metric loss of the similarity constraint is calculated through the following steps:
[0020] Calculate the quantization loss: Calculate the quantization loss based on the similarity difference between the floating-point features and the binary hash features;
[0021] Calculate the metric loss: Calculate the metric loss based on the multi-similarity loss, and use a straight-through estimator to calculate the gradient during backpropagation;
[0022] Calculate the overall loss: Combine the quantization loss and the metric loss by weighted summation to obtain the overall loss function.
[0023] A computer program product includes a computer program which, when executed by a processor, implements the two-stage place recognition method based on similarity constraint deep hashing described above.
[0024] The present invention has the following beneficial effects:
[0025] The present invention provides an efficient two-stage visual place recognition method, which innovatively solves the key problems of retrieval latency and storage in the field of visual place recognition applications, and significantly reduces the consumption of the model in these two aspects. First, by introducing a two-stage recognition location recognition paradigm of similarity-constrained deep hashing, image features for the first-stage retrieval are obtained through hashing, and then higher-dimensional floating-point features are used for the second-stage re-ranking. The use of compact binary features during the first-stage retrieval can significantly reduce the memory consumption of VPR and improve the retrieval efficiency of the model. The present invention proposes a similarity-constrained quantization loss, which can effectively resist the problems of quantization loss and metric loss resistance existing during the training of deep hashing, and preferably integrates the binary hashing operation into the neural network by combining with the straight-through estimator (STE), realizing end-to-end training containing quantization operations. The quantization operation is performed during forward propagation, and the situation where the gradient is 0 and the parameters cannot be updated is avoided during backpropagation. Compared with the traditional two-stage visual place recognition method, the method of the present invention significantly reduces the retrieval latency and storage space on multiple datasets such as Pittsburgh, Tokyo24 / 7, Mapillary Street-Level Sequences (MSLS), and improves the Rank-1 accuracy and mean average precision (mAP). The present invention can directly match local features without spatial verification, making the re-ranking much faster and requiring less memory than the mainstream two-stage VPR method. The present invention is applicable to the visual place recognition scenario and has important application value in the fields of mobile robot positioning and augmented reality.
[0026] Other beneficial effects in the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flowchart of the two-stage place recognition method based on similarity-constrained deep hashing according to an embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of the model of the two-stage place recognition method based on similarity-constrained deep hashing according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The following makes a detailed description of the embodiments of the present invention. It should be emphasized that the following description is merely exemplary and not intended to limit the scope of the present invention and its applications.
[0030] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more such features. In the description of the embodiments of the present invention, "a plurality" means two or more, unless otherwise specifically defined.
[0031] Referring to Figure 1 and Figure 2 , an embodiment of the present invention provides a two-stage location recognition method based on similarity-constrained deep hashing, including the following steps:
[0032] Step S1, data preparation and feature extraction: Based on a triplet training framework, construct triplet training samples including query images, positive example images, and negative example images; input the training samples into a feature extraction network to extract image features for two-stage retrieval.
[0033] The triplet training framework can adopt, for example, the training framework of the GSV-Cities dataset.
[0034] In a preferred embodiment, in step S1, the construction of the triplet training samples is based on geographical location tags, each query image corresponds to a number of positive example images and negative example images, and the model training is optimized through an online hard mining strategy.
[0035] In a preferred embodiment, the feature extraction network adopts a pre-trained vision large model (such as the DINOv2 model) to extract a high-dimensional semantic feature map, and aggregates the high-dimensional semantic feature map into a one-dimensional feature vector through generalized mean pooling (Gempooling).
[0036] Step S2, feature processing and transformation: Perform dimensionality reduction and dimensionality increase processing on the high-dimensional semantic feature map output by the feature extraction network; convert the dimensionality-reduced features into binary hash features through a similarity-constrained deep hashing method, and at the same time obtain the floating-point features after dimensionality increase.
[0037] In a preferred embodiment, in step S2, the high-dimensional semantic feature map is processed for dimensionality reduction and dimensionality increase through a linear layer and a fully connected layer, and then binary hash features and floating-point features are generated respectively.
[0038] In a preferred embodiment, the similarity-constrained deep hashing method constrains the similarity between the binary hash features and the floating-point features through a quantization loss, and realizes end-to-end training in combination with a straight-through estimator (STE).
[0039] Step S3, Two-stage Retrieval: Use binary hash features to perform a fast first-stage retrieval in the image database to filter out the top k candidate images with features similar to the query image; use floating-point features to perform an accurate re-ranking of the candidate images in the second stage to determine the final top k images that are most similar to the query image.
[0040] Specifically, in the first-stage retrieval, the top k candidate images with features most similar to the query image are filtered out through the fast matching of binary hash features, and in the second-stage re-ranking, the top k images that are finally most similar to the query image are determined by calculating the similarity metric between floating-point features.
[0041] Step S4, Model Optimization: Adopt a weakly-supervised method to calculate the overall loss function that includes quantization loss and metric loss with similarity constraints; iteratively optimize the model parameters until the preset convergence condition is met to achieve end-to-end training of the model.
[0042] In a preferred embodiment, in step S4, the overall loss function combines the quantization loss and the metric loss by weighted summation, and iteratively optimizes the model parameters through the backpropagation algorithm. The convergence condition for model optimization is that the error value calculated by the loss function is less than the preset threshold, or the number of iterations reaches a predetermined value. The weakly-supervised method improves the retrieval accuracy and efficiency of the model through the joint optimization of the metric loss and the quantization loss.
[0043] In a preferred embodiment, the metric loss includes multi-similarity loss (MS loss) for optimizing the similarity metric between image features.
[0044] In a further preferred embodiment, the metric loss with similarity constraints is calculated through the following steps: calculate the quantization loss based on the similarity difference between floating-point features and binary hash features; calculate the metric loss based on the multi-similarity loss, and use a straight-through estimator to calculate the gradient during backpropagation; combine the quantization loss and the metric loss by weighted summation to obtain the overall loss function.
[0045] The present invention proposes an efficient two-stage visual location recognition method, effectively solving the deficiencies of traditional methods in terms of retrieval latency and storage consumption. In the first stage, compact binary hash features are used for fast retrieval, significantly reducing memory occupancy and improving retrieval efficiency; in the second stage, higher-dimensional floating-point features are used for precise reordering to ensure the accuracy of the recognition results. Further, through the combination of similarity-constrained quantization loss and the Straight-Through Estimator (STE), end-to-end training with quantization operations is achieved, avoiding the problem of parameter update caused by zero gradients and further improving the training effect of the model. Compared with the prior art, the present invention verifies its superiority on multiple datasets, not only significantly reducing the retrieval latency and storage space, but also remarkably improving the Rank-1 accuracy and mean average precision (mAP), providing a more efficient and accurate solution for applications in fields such as mobile robot positioning and augmented reality, and having important practical application value.
[0046] The following further describes specific embodiments of the present invention and their algorithm examples.
[0047] An efficient method for two-stage visual location recognition and its model training solves the problem of low efficiency in memory usage during the training process of the two-stage visual location recognition method, and reduces the expensive consumption of retrieval latency and storage during reordering in the second stage. The method includes: in the first stage, the Deep Hashing method is used to generate low-dimensional binary global hash features to select the top K candidate images that are most similar to the query image, and then higher-dimensional floating-point features are used to reorder the K candidate images; the similarity-constrained deep hashing method is adopted, and the similarity-constrained quantization loss L_Q is used to replace the traditional deep hashing quantization loss L1 / L2, and the binary hashing operation is integrated into the neural network by combining the Straight-Through Estimator (STE) to achieve end-to-end training of the network with quantization operations.
[0048] Refer to Figure 1 and Figure 2 , a two-stage place recognition method based on similarity-constrained deep hashing, includes the following steps:
[0049] Step S1: Adopt the training framework of the GSV-Cities dataset, and in one batch, each query image corresponds to several positive and negative example images to form a triplet for training. Input the mini-batch of training image triplets with geographical location labels into the feature extraction network to extract the image features for two-stage retrieval.
[0050] Step S2: The high-dimensional semantic feature map obtained through the feature extraction network is successively passed through a linear layer, an aggregation layer, and a fully connected layer for dimensionality reduction and dimensionality increase respectively. The features after dimensionality reduction are subjected to deep hashing with similarity constraints to obtain binary hashing features, and the features after dimensionality increase are floating-point features.
[0051] Step S3: First, use the binary hashing features to perform the first-stage retrieval of the query image in the image database to obtain the top k candidate images with the most similar features to the query image. Then, use the floating-point features to perform a more accurate re-ranking among the k candidate images, so as to obtain the top k images that are finally most similar to the query image.
[0052] Step S4: Adopt a weakly supervised method to calculate the overall loss function, including the metric loss and quantization loss of similarity constraints, and iteratively optimize the model parameters until the training meets the preset convergence conditions.
[0053] Specifically, in a preferred embodiment, the training method for two-stage visual position recognition further includes the following steps:
[0054] Further, in step S2, the feature extraction network uses a pre-trained vision large model DINOv2 model, uses the patch tokens output by DINO v2 as the high-dimensional feature map, and then uses generalized mean pooling (Gem pooling) to aggregate into a one-dimensional feature vector in the aggregation layer.
[0055] Further, in step S4, the metric loss of the similarity constraint is specifically calculated through the following steps:
[0056] S41: Calculate the quantization loss of the similarity constraint:
[0057]
[0058] where f i and f j are the floating-point features of images i and j, b i and b j are the hashing features of images i and j, and K represents the number of pairs of images in a batch.
[0059]
[0060] S42: Calculate the metric loss L M (b) calculated using the multi-similarity loss (MS loss), and adopt the straight-through estimator (STE) to facilitate the calculation of the gradient during backpropagation. (b), and adopt the straight-through estimator (STE) to facilitate the calculation of the gradient during backpropagation.
[0060] S43: Calculate the total loss L = L M (b) + λL Q(f, b) controls the weights of the quantization loss and the metric loss.
[0061] During the training process, the training data will be continuously input into the model from step S1 to step S4, calculate the overall loss function, and iteratively optimize the model parameters until the training meets the preset convergence condition. Specifically, use the metric loss function and the quantization loss function to calculate the error of the model, and update the model parameters according to the error through the backpropagation algorithm to optimize the performance of the linear layer, aggregation layer, and fully connected layer. The training result meeting the preset convergence condition can be that the error value calculated by the loss function is less than a preset threshold, or the number of iterations reaches a predetermined value.
[0062] Compared with the traditional two-stage visual place recognition method, the method of the present invention significantly reduces the retrieval latency and storage space on multiple datasets such as Pittsburgh, Tokyo24 / 7, Mapillary Street-Level Sequences (MSLS), and improves the Rank-1 accuracy and mean average precision (mAP). The present invention can directly match local features without spatial verification, making the re-ranking much faster than the mainstream two-stage VPR method and requiring less memory.
[0063] In summary, the two-stage visual place recognition method of the present invention significantly reduces the retrieval latency and storage space compared with the traditional method, and improves the Rank-1 accuracy and mean average precision (mAP). The present invention is applicable to the visual place recognition scenario and has important application value in the fields of mobile robot positioning and augmented reality.
[0064] The embodiment of the present invention also provides a storage medium for storing a computer program, and when the computer program is executed, it at least executes the method described above.
[0065] The embodiment of the present invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein, the processor is used to execute the computer program and at least execute the method described above.
[0066] The embodiment of the present invention also provides a processor, and the processor executes a computer program and at least executes the method described above.
[0067] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memories.
[0068] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the various components shown or discussed can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0069] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0070] In addition, in each embodiment of the present invention, the various functional units can all be integrated in one processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0071] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media that can store program codes such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0072] Alternatively, if the above integrated unit is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: various media that can store program codes such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0073] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0074] The features disclosed in the several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0075] The features disclosed in the several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0076] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention belongs, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and as long as the performance or use is the same, they should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A two-stage location recognition method based on similarity-constrained deep hashing, characterized in that: The following steps are involved: S1. Based on the triple training framework, construct triple training samples including query image, positive image and negative image; input the training samples into the feature extraction network to extract image features for two-stage retrieval; S2. Perform dimensionality reduction and dimensionality increase processing on the high-dimensional semantic feature graph output by the feature extraction network; convert the dimensionality reduction features into binary hash features through the similarity-constrained deep hashing method, and obtain the floating-point features after dimensionality increase; S3, using binary hash features to perform the first stage of rapid retrieval in the image database, screening out the top k candidate images with features similar to the query image; using floating-point features to perform the second stage of precise reordering of the candidate images, and finally determining the top k images most similar to the query image; S4. Use weak supervision method to calculate the overall loss function including quantization loss and metric loss of similarity constraint; optimize the model parameters by iteration until the preset convergence conditions are met to achieve end-to-end training of the model.
2. The two-stage location recognition method according to claim 1, characterized in that: In step S1, the construction of the triplet training samples is based on geographic location tags, each query image corresponds to a number of positive and negative images, and the model training is optimized through an online hard mining strategy.
3. The two-stage location recognition method according to claim 1 or 2, characterized in that: In step S1, the feature extraction network uses a pre-trained visual large model to extract a high-dimensional semantic feature map, and aggregates the high-dimensional semantic feature map into a one-dimensional feature vector through generalized mean pooling (Gempooling).
4. The two-stage location recognition method according to any one of claims 1 to 3, characterized in that: In step S2, the high-dimensional semantic feature map is processed by dimensionality reduction and dimensionality increase through a linear layer and a fully connected layer, and then binary hash features and floating-point features are generated respectively.
5. The two-stage location recognition method according to any one of claims 1 to 4, characterized in that: In step S2, the similarity-constrained deep hashing method constrains the similarity between binary hash features and floating-point features through quantization loss, and implements end-to-end training in combination with a straight-through estimator (STE).
6. The two-stage location recognition method according to any one of claims 1 to 5, characterized in that: In step S3, the first stage retrieval screens out the top k candidate images that are closest to the query image features through fast matching of binary hash features, and the second stage re-ranking determines the top k images that are most similar to the query image by calculating the similarity measure between floating-point features.
7. The two-stage location recognition method according to any one of claims 1 to 6, characterized in that: In step S4, the overall loss function combines the quantization loss and the measurement loss by weighted summation, and iteratively optimizes the model parameters by back propagation algorithm. The convergence condition of model optimization is that the error value calculated by the loss function is less than a preset threshold, or the number of iterations reaches a predetermined value. The weak supervision method improves the retrieval accuracy and efficiency of the model by jointly optimizing the measurement loss and the quantization loss.
8. The two-stage location recognition method according to any one of claims 1 to 7, characterized in that: In step S4, the metric loss includes a multi-similarity loss (MS loss), which is used to optimize the similarity metric between image features.
9. The two-stage location recognition method according to claim 8, characterized in that: In step S4, the metric loss of the similarity constraint is calculated by the following steps: Calculate quantization loss: Calculate quantization loss based on the similarity difference between floating-point features and binary hash features; Calculate metric loss: Calculate metric loss based on multi-similarity loss, and use the straight-through estimator to calculate the gradient during backpropagation; Calculate the overall loss: Combine the quantization loss and the measurement loss by weighted summation to get the overall loss function.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the two-stage location recognition method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Multi-level retrieval pedestrian re-identification method based on deep hash
CN112818859A
Model training method, image feature extraction method, system, device and medium
CN114549907A
Training a convolutional neural network for image retrieval with a listwise ranking loss function
EP3731154A1