Visual location identification method based on frequency domain-space domain double-domain aggregation
Through the dual-domain aggregation method of frequency domain and spatial domain, using the multi-scale contextual attention module and triple fusion strategy, the problem of existing technology that only processes local features in the spatial domain is solved, and more discriminative and robust global features are generated, thereby improving the accuracy of visual place recognition.
Patent Information
- Application Number
- CN202510745824.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-12
AI Technical Summary
Existing visual place recognition methods are limited to processing local features in the spatial domain and ignore information in the frequency domain, resulting in poor feature aggregation effects.
A dual-domain aggregation method based on frequency domain and spatial domain is adopted. The spatial domain feature map is converted to the frequency domain through two-dimensional fast Fourier transform. The feature map is optimized by combining the multi-scale contextual attention module. Finally, the spatial domain and frequency domain feature maps are aggregated into a unified global feature through a triple fusion strategy.
More discriminative and robust global features are generated, which improves the accuracy and reliability of visual place recognition, especially the excellent recognition effect on multiple challenging datasets.
Smart Images

Figure CN120635704A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual place recognition, and in particular to a visual place recognition method based on dual-domain aggregation of frequency domain and spatial domain. Background Art
[0002] The goal of visual place recognition is to identify the location of a desired image, given a database of geotagged images. Visual place recognition is crucial in many real-world applications, including autonomous driving, mobile robotics, and augmented reality. However, due to factors such as lighting, occlusion, seasonality, and perspective variations, the appearance of the same location can vary significantly across different images, making this task extremely challenging.
[0003] The task of visual place recognition involves identifying the specific geographic location of a query image from a database of geotagged images. This task is typically addressed as an image retrieval task. Existing visual place recognition methods generally aggregate local image features into global features and perform image retrieval based on similarity matching between the query image and the global features of the database images. Therefore, the feature aggregation strategy is crucial to the final recognition performance. The entire process plays a key role in the effectiveness of visual place recognition, and much existing research has focused on improving aggregation strategies. For example, VLAD aggregates features by calculating the difference between local features and cluster centers and has been widely adopted and expanded in visual place recognition methods. The Conv–AP method uses convolutional layers followed by average pooling to generate global features. SALAD utilizes the Sinkhorn algorithm for optimal transfer to aggregate features. Despite these advances, existing aggregation methods are limited to processing local features in the spatial domain. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned technical deficiencies and provide a visual place recognition method based on dual-domain aggregation of frequency domain and spatial domain, so as to solve the technical problem that the existing technology is limited to processing local features in the spatial domain and ignores the processing in the frequency domain that can extract richer information from the image.
[0005] To achieve the above technical objectives, in a first aspect, the technical solution of the present invention provides a visual place recognition method based on frequency domain-spatial domain dual-domain aggregation, comprising the steps of:
[0006] Build a geotagged image database;
[0007] Given an input image, extracting a local spatial domain feature map of the input image using a visual feature extractor;
[0008] Converting the spatial domain feature map to the frequency domain by two-dimensional fast Fourier transform to obtain a frequency domain feature map;
[0009] Optimizing the spatial domain feature map and the frequency domain feature map in parallel through a multi-scale contextual attention module to obtain optimized spatial domain feature map and frequency domain feature map;
[0010] Aggregating the optimized spatial domain feature map and the frequency domain feature map into a unified global feature through triple fusion;
[0011] A reference image that is most similar to the input image is retrieved from the image database based on the global features, thereby identifying the location of the input image based on the geographic tag of the reference image.
[0012] Compared with the prior art, the present invention has the following beneficial effects:
[0013] This application proposes a novel visual place recognition method - a dual-domain aggregation recognition method. This method integrates spatial domain and frequency domain information to enhance the aggregation effect. Its core motivation is that the frequency domain can more easily capture the global and structural patterns of the image, providing information that is difficult to obtain in the spatial domain. This information is discriminative and can complement the information of the spatial domain features. In order to better optimize features in the frequency domain and spatial domain, this application designs a multi-scale contextual attention module to effectively utilize multi-scale information in the downsampling process of feature aggregation and retain key details. In addition, in order to bridge the gap between spatial and frequency domain features, this application introduces a triple fusion strategy to promote cross-domain interaction, deeply fuse spatial domain and frequency domain features, and thus generate unified and robust global features. In this way, the global features finally aggregated by the dual-domain aggregation recognition method are more discriminative than the features aggregated only in the spatial domain by existing methods. Experimental results on multiple challenging datasets show that the dual-domain aggregation recognition method of this application has excellent performance and achieves the best recognition effect to date.
[0014] According to some embodiments of the present invention, a visual feature extractor is used to extract a local spatial domain feature map from the input image, and the spatial domain feature map is converted into a frequency domain by a two-dimensional fast Fourier transform to obtain a frequency domain feature map, including the steps of:
[0015] Use the visual feature extractor to extract local spatial domain feature maps:
[0016] F s =F(I)
[0017] in Represents the local spatial domain feature map; H and W represent the height and width respectively, and D is the feature dimension; F sInformation used to describe the visual content of the input image in the spatial domain;
[0018] The F is transformed by two-dimensional fast Fourier transform (FFT) s Convert to the frequency domain:
[0019] F f =FFT(F s )
[0020] in It is the frequency domain feature map.
[0021] According to some embodiments of the present invention, optimizing the spatial domain feature map and the frequency domain feature map in parallel by a multi-scale context attention module comprises the steps of:
[0022]
[0023] Among them, IFFT(·) converts the frequency domain feature map back to the spatial domain, and MSC-ATTN(·) is a multi-scale context attention module. The multi-scale context attention module converts F s and F f Downsample respectively to obtain the optimized spatial domain feature map And the frequency domain feature map Than F s and F f More differentiated and compact; Where S = (H / 4) × (W / 4).
[0024] According to some embodiments of the present invention, the multi-scale context attention module optimizes the spatial domain feature map and the frequency domain feature map by:
[0025] The local spatial domain feature map F s Project into query, key, and value matrices Q, K, and V with the same resolution H×W;
[0026] Q=W q F s ,K=W k F s ,V=W v F s
[0027] The matrix is stratified downsampled:
[0028]
[0029] Where D(·) represents a 3×3 convolution with a stride of 2, which can achieve a 2×2 downsampling rate. In order to integrate multi-scale information, the keys and values of different scales are flattened and concatenated to obtain multi-scale key and value features:
[0030]
[0031] Query after downsampling Operate with multi-scale keys and values to obtain optimized features:
[0032]
[0033] The optimization process of the frequency domain feature map is consistent with the optimization process of the spatial domain feature map.
[0034] According to some embodiments of the present invention, the optimized spatial domain feature map and the frequency domain feature map are aggregated into a unified global feature by triple fusion, including the steps of:
[0035] According to the optimized frequency domain feature map The spatial domain feature map is in the form of To reorganize:
[0036]
[0037] where q s (·), k s (·) and v s (·) are linear projection functions used to generate query, key and value respectively, F s′ Contains spatial domain information, but With similar signal organization structure, similarly, F f′ Obtained through:
[0038]
[0039] F s′ and F f′ The signal forms have been reorganized to be compatible with each other, reducing and Domain gap between features; based on implicit alignment of F s′ and F f′ , using the self-attention mechanism to fuse into a unified global feature:
[0040]
[0041] in cat(·;·) represents the concatenation operation of two features, and the aggregated unified global feature F contains feature information from both spatial and frequency domains.
[0042] According to some embodiments of the present invention, a reference image that is most similar to the input image is retrieved from the image database based on the global features. The retrieval process is formalized as follows:
[0043]
[0044] where s(·,·) represents the similarity measure, F(·) represents the visual feature extractor used to extract local features from the input image, and f agg (·) is an aggregation function that aggregates local features into compact global features.
[0045] According to some embodiments of the present invention, the geographic tagged image database includes at least any one of the following: Pitts30k dataset, MSLS dataset, SPED dataset, Tokyo24 / 7 dataset, and NordLand dataset.
[0046] In a second aspect, the technical solution of the present invention provides a visual place recognition system based on frequency domain-spatial domain dual-domain aggregation, comprising:
[0047] Data construction module for building a database of geotagged images;
[0048] A spatial domain feature extraction module, given an input image, uses a visual feature extractor to extract a local spatial domain feature map of the input image;
[0049] A frequency domain conversion module is in communication with the spatial domain feature extraction module and converts the spatial domain feature map into the frequency domain by a two-dimensional fast Fourier transform to obtain a frequency domain feature map;
[0050] a feature optimization module, which is in communication with the spatial domain feature extraction module and the frequency domain transformation module, and optimizes the spatial domain feature map and the frequency domain feature map in parallel through a multi-scale context attention module to obtain the optimized spatial domain feature map and the frequency domain feature map;
[0051] A feature fusion module is connected to the feature optimization module and aggregates the optimized spatial domain feature map and the frequency domain feature map into a unified global feature through triple fusion;
[0052] An image retrieval module is in communication with the data construction module and the feature fusion module, and retrieves a reference image that is most similar to the input image from the image database based on the global features, thereby identifying the location of the input image based on the geographic tag of the reference image.
[0053] In a third aspect, the technical solution of the present invention provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the visual place recognition method based on dual-domain aggregation of frequency domain and spatial domain as described in any one of the first aspects.
[0054] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, wherein the abstract drawing is identical to one of the drawings in the specification:
[0056] Figure 1 A flow chart of a visual place recognition method based on frequency-spatial dual-domain aggregation provided by one embodiment of the present invention;
[0057] Figure 2 A flow chart of a visual place recognition method based on frequency-spatial dual-domain aggregation provided by one embodiment of the present invention;
[0058] Figure 3 A diagram showing the structure of a multi-scale contextual attention module according to one embodiment of the present invention;
[0059] Figure 4 A triple fusion structure diagram provided for one embodiment of the present invention. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0061] It should be noted that although the system diagrams illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the system or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0062] The goal of visual place recognition is to identify the location of a desired image, given a database of geotagged images. Visual place recognition is crucial in many real-world applications, including autonomous driving, mobile robotics, and augmented reality. However, due to factors such as lighting, occlusion, seasonality, and perspective variations, the appearance of the same location can vary significantly across different images, making this task extremely challenging.
[0063] Visual place recognition is often addressed as an image retrieval problem. Existing methods use feature extractors to extract local features from images and then aggregate these local features into global features. This approach then performs a similarity match between the global features of the query image and database images, selecting the location in the database image with the highest similarity as the query location. Therefore, the feature aggregation strategy throughout this process plays a key role in the effectiveness of visual place recognition, and much existing research has focused on improving this aggregation strategy. For example, VLAD aggregates features by calculating the difference between local features and cluster centers and has been widely adopted and extended in visual place recognition methods. The Conv-AP method utilizes convolutional layers followed by average pooling to generate global features. SALAD utilizes the Sinkhorn algorithm for optimal transfer to aggregate features. Despite these advances, existing aggregation methods are limited to processing local features in the spatial domain, neglecting the potential of processing in the frequency domain to extract richer information from the image. Frequency domain information has a global receptive field and can encode global features, effectively capturing the overall pattern and structure of the image. In visual place recognition, images often contain repetitive textures and shapes, such as the neat and regular arrangement of building windows. These patterns can be concentrated in the frequency domain, which is more conducive to efficient processing and feature extraction than the spatial domain.
[0064] Therefore, the present application proposes a dual-domain aggregation recognition method for visual place recognition, which aims to jointly enhance global features by using key information in the spatial and frequency domains. The dual-domain aggregation recognition method adopts a dual-branch architecture, and each branch is responsible for processing the features of a specific domain. The features obtained by processing the frequency domain branch contain global and structural information that is complementary to the spatial domain. In each branch, the present application designs a multi-scale context attention module that can retain key local details when downsampling features to extract information. By utilizing multi-scale context, the module ensures comprehensive information extraction during feature aggregation. Finally, a triple fusion mechanism is designed to integrate features from two different domains and compensate for the differences between features in different domains. The triple fusion mechanism can fuse the complementary information of spatial and frequency domain features and ultimately generate robust global features.
[0065] Compared to existing methods that only aggregate features in the spatial domain, our dual-domain aggregation recognition method can generate more informative and discriminative global features. We evaluate this method on several challenging public datasets, demonstrating its superior performance in visual place recognition tasks. The main contributions of this work are as follows:
[0066] 1. This application proposes a dual-domain aggregation recognition method for visual place recognition, which aggregates features through complementary clues in the spatial and frequency domains to obtain robust global features.
[0067] 2. A novel multi-scale contextual attention module is designed to explore multi-scale information and preserve key details during downsampling. In addition, a triple fusion mechanism is designed to compensate for the difference between spatial and frequency domain features to generate discriminative global features.
[0068] 3. Extensive experiments conducted on multiple benchmark datasets demonstrate the effectiveness and generalization ability of the dual-domain aggregation recognition method, achieving the current optimal performance.
[0069] Reference Figures 1 to 4 , Figure 1 A flow chart of a visual place recognition method based on frequency-spatial dual-domain aggregation provided by one embodiment of the present invention; Figure 2 A flow chart of a visual place recognition method based on frequency-spatial dual-domain aggregation provided by one embodiment of the present invention; Figure 3 A diagram showing the structure of a multi-scale contextual attention module according to one embodiment of the present invention; Figure 4 A triple fusion structure diagram provided for one embodiment of the present invention.
[0070] The visual place recognition method based on frequency domain and spatial domain dual-domain aggregation includes but is not limited to the following steps:
[0071] Step S110, constructing a geographically tagged image database;
[0072] Step S120: Given an input image, extract a local spatial domain feature map from the input image using a visual feature extractor, and convert the spatial domain feature map into a frequency domain by a two-dimensional fast Fourier transform to obtain a frequency domain feature map;
[0073] Step S130, optimizing the spatial domain feature map and the frequency domain feature map in parallel through a multi-scale contextual attention module to obtain optimized spatial domain feature map and frequency domain feature map;
[0074] Step S140, aggregating the optimized spatial domain feature map and frequency domain feature map into a unified global feature through triple fusion;
[0075] Step S150 , based on the global features, a reference image that is most similar to the input image is retrieved from the image database, thereby identifying the location of the input image based on the geographic tag of the reference image.
[0076] In one embodiment, a visual place recognition method based on dual-domain aggregation of frequency domain and spatial domain includes the following steps: constructing an image database with geographic tags; given an input image, using a visual feature extractor to extract a local spatial domain feature map of the input image; converting the spatial domain feature map to the frequency domain through a two-dimensional fast Fourier transform to obtain a frequency domain feature map; optimizing the spatial domain feature map and the frequency domain feature map in parallel through a multi-scale contextual attention module to obtain optimized spatial domain feature map and frequency domain feature map; aggregating the optimized spatial domain feature map and frequency domain feature map into a unified global feature through a triple fusion method; retrieving a reference image that is most similar to the input image from the image database based on the global feature, thereby identifying the location of the input image based on the geographic tag of the reference image.
[0077] This application is no longer limited to information in a single domain by utilizing features in both the spatial domain and the frequency domain. Spatial domain features can capture intuitive visual information such as the local structure, shape, and texture of objects in the image, while frequency domain features can reflect the distribution of different frequency components of the image from a frequency perspective, revealing some periodicity, directionality, and other features that are difficult to directly observe in the spatial domain. The combination of the two makes the extracted features more comprehensive and can cover richer image content, which helps to improve the ability to describe images in different scenes and under different shooting conditions, thereby improving the accuracy of location recognition.
[0078] This application uses a multi-scale contextual attention module in the process of processing features, which means that both spatial domain feature maps and frequency domain feature maps can be optimized at multiple scales. Different scales correspond to different ranges of image information. For example, a small scale can focus on the details of the image, while a large scale can grasp the overall scene layout. Comprehensive multi-scale optimization enables the extracted features to take into account both local subtle differences and overall scene features, which is more advantageous for visual place recognition in complex and changeable actual environments, and reduces the possibility of recognition errors due to scale changes (such as shooting the same place from a long distance and close up).
[0079] This application uses a multi-scale contextual attention module to optimize the spatial domain feature map and the frequency domain feature map in parallel, and can simultaneously make targeted improvements to the characteristics of the features of the two domains. Compared with processing the features of a certain domain in sequence or processing the features of a certain domain separately, parallel operations can make full use of computing resources, efficiently mine more valuable information in the features of the two domains, and avoid ignoring the relevance and complementarity of the features of another domain when optimizing the features of one domain, so that the final optimized features are of higher quality and can better highlight the key features of the image in terms of location recognition. The multi-scale contextual attention module can automatically assign attention weights according to the image content, focusing on more important areas and feature components. For example, in the spatial domain, higher weights can be given to areas containing key location recognition elements such as landmark buildings, and in the frequency domain, focus is placed on parts that can reflect the unique frequency patterns of the scene. Through this targeted optimization and enhancement, the distinguishability of the features is further enhanced, making the differences in image features at different locations more obvious, which helps to more accurately match similar images in subsequent retrievals.
[0080] This application uses a triple fusion approach to aggregate the optimized spatial domain feature map and frequency domain feature map into a unified global feature. This fusion method can fully integrate the optimized feature information of the two domains, avoiding the problems of information loss or insufficient fusion that may be caused by simple and crude splicing or a single fusion method. Through a reasonable triple fusion strategy, the aggregated global features can comprehensively reflect the advantages of spatial domain and frequency domain features, forming a more representative, more compact feature representation that contains rich discriminant information, which facilitates subsequent efficient image similarity retrieval based on this feature.
[0081] This application uses aggregated global features to retrieve the reference image most similar to the input image from an image database. Because global features incorporate key information from both the spatial and frequency domains, they can more comprehensively and accurately measure the similarity between images when performing similarity measurement, compared to retrieval based on features in a single domain. It can better cope with various complex situations, such as changes in image appearance caused by different lighting or photographing the same location from different angles, while still accurately finding the corresponding reference image. Furthermore, it can accurately identify the location of the input image based on the geographic markers in the reference image, improving the overall accuracy and reliability of visual place recognition.
[0082] The method uses a visual feature extractor to extract a local spatial domain feature map from the input image, and converts the spatial domain feature map into a frequency domain feature map through a two-dimensional fast Fourier transform, including the following steps:
[0083] Use the visual feature extractor to extract local spatial domain feature maps:
[0084] F s =F(I)
[0085] in Represents the local spatial domain feature map; H and W represent the height and width respectively, and D is the feature dimension; F s Information used to describe the visual content of the input image in the spatial domain;
[0086] The F is transformed by two-dimensional fast Fourier transform (FFT) s Convert to the frequency domain:
[0087] F f =FFT(F s )
[0088] in It is the frequency domain feature map.
[0089] Among them, the spatial domain feature map and the frequency domain feature map are optimized in parallel through the multi-scale context attention module, including the following steps:
[0090]
[0091] Among them, IFFT(·) converts the frequency domain feature map back to the spatial domain, and MSC-ATTN(·) is a multi-scale context attention module. The multi-scale context attention module converts F s and F f Downsample separately to obtain the optimized spatial domain feature map and frequency domain feature maps Than F s and F f More differentiated and compact;
[0092] Where S = (H / 4) × (W / 4).
[0093] Among them, feature aggregation involves downsampling operations to optimize and compress features. However, traditional downsampling operations, such as average pooling, directly aggregate adjacent features, which will result in the discarding of key details and limit the adaptability of the network to different scales. To address these problems, this application designs a novel multi-scale contextual attention (MSC-Attn) module, which can effectively integrate multi-scale contextual information and retain key fine-grained details during the downsampling process in the spatial and frequency domains. The multi-scale contextual attention module optimizes the spatial domain feature map and the frequency domain feature map in the following steps:
[0094] The local spatial domain feature map F s Project into query, key, and value matrices Q, K, and V with the same resolution H×W;
[0095] Q=W q F s ,K=W kF s ,V=W v F s
[0096] The matrix is stratified downsampled:
[0097]
[0098] Where D(·) represents a 3×3 convolution with a stride of 2, which can achieve a 2×2 downsampling rate. In order to integrate multi-scale information, the keys and values of different scales are flattened and concatenated to obtain multi-scale key and value features:
[0099]
[0100] Query after downsampling Operate with multi-scale keys and values to obtain optimized features:
[0101]
[0102] The optimization process of the frequency domain feature map is consistent with that of the spatial domain feature map.
[0103] Here, this application utilizes multi-scale information to improve the accuracy of feature extraction.
[0104] Compared with general downsampling operations (such as average pooling or generalized average pooling), MSC-Attn can effectively integrate contextual information at multiple scales in spatial and frequency domains, ensuring that discriminative details at different scales are preserved.
[0105] The optimized spatial domain feature map and frequency domain feature map are aggregated into a unified global feature through triple fusion, including the following steps:
[0106] According to the optimized frequency domain feature map The spatial domain feature map To reorganize:
[0107]
[0108] where q s (·), k s (·) and v s (·) are linear projection functions used to generate query, key and value respectively, F s′ Contains spatial domain information, but With similar signal organization structure, similarly, F f′ Obtained through:
[0109]
[0110] F s′ and F f′ The signal forms have been reorganized to be compatible with each other, reducing and Domain gap between features; based on implicit alignment of F s′ and F f′ , using the self-attention mechanism to fuse into a unified global feature:
[0111]
[0112] in cat(·;·) represents the concatenation operation of two features, and the aggregated unified global feature F contains feature information from both spatial and frequency domains.
[0113] Among them, visual place recognition aims to retrieve and input images from database B Most similar reference image Thus, the position of the input image is identified. Based on the global features, the reference image that is most similar to the input image is retrieved from the image database. The retrieval process is formalized as follows:
[0114]
[0115] where s(·,·) represents the similarity measure, F(·) represents the visual feature extractor used to extract local features from the input image, and f agg (·) is an aggregation function that aggregates local features into compact global features.
[0116] Furthermore, the geotagged image database includes at least one of the following: the Pitts30k dataset, the MSLS dataset, the SPED dataset, the Tokyo24 / 7 dataset, and the NordLand dataset. This application validates the method on the following multiple visual place recognition benchmark datasets: Pitts30k, MSLS, SPED, Tokyo24 / 7, and NordLand. The Pitts30k dataset contains over 6,000 query images and 10,000 database images, with significant perspective changes between images. The MSLS dataset contains over 1.6 million images from 30 major cities on six continents. The SPED dataset contains 607 query images and 607 database reference images, with seasonal and diurnal variations. The Tokyo24 / 7 dataset contains over 75,000 database images and 315 query images, showing diurnal variations. The NordLand dataset contains over 2,700 query images and over 27,000 database images, which are taken from the perspective of a train driver and include changes across four seasons.
[0117] After 40 epochs of training, our network achieved state-of-the-art results on multiple widely used benchmark datasets. Using the Recall@1 metric, a commonly used retrieval metric, our approach surpassed existing methods on all datasets. On the Pitts30k dataset, our approach achieved a 0.2% improvement over the existing best method. We also achieved a 0.6% improvement on the MSLSval dataset, a 0.4% improvement on the SPED dataset, a 2.2% improvement on the Tokyo24 / 7 dataset, a 0.8% improvement on the MSLS challenge dataset, and a 2.6% improvement on the Nordland dataset.
[0118] In one embodiment, a visual place recognition system based on frequency-domain and spatial-domain dual-domain aggregation includes: a data construction module for constructing an image database with geographic tags; a spatial domain feature extraction module for extracting a local spatial domain feature map from the input image using a visual feature extractor given an input image; a frequency domain transformation module for communicating with the spatial domain feature extraction module and converting the spatial domain feature map to the frequency domain through a two-dimensional fast Fourier transform to obtain a frequency domain feature map; a feature optimization module for communicating with the spatial domain feature extraction module and the frequency domain transformation module and optimizing the spatial domain feature map and the frequency domain feature map in parallel through a multi-scale contextual attention module to obtain optimized spatial domain feature map and frequency domain feature map; a feature fusion module for communicating with the feature optimization module and aggregating the optimized spatial domain feature map and the frequency domain feature map into a unified global feature through a triple fusion method; and an image retrieval module for communicating with the data construction module and the feature fusion module and retrieving a reference image that is most similar to the input image from the image database based on the global feature, thereby identifying the location of the input image based on the geographic tag of the reference image.
[0119] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate and may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0121] In addition, an embodiment of the present invention also provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by a processor or controller, for example, by a processor in the above-mentioned terminal embodiment, so that the above-mentioned processor can execute the visual place recognition method based on dual-domain aggregation of frequency domain and spatial domain in the above-mentioned embodiment.
[0122] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media is included in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data) and is volatile and non-volatile, removable, and non-removable. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage, or other magnetic storage devices, or can be used to store desired information and any other medium that can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0123] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above implementation. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.
[0124] The specific embodiments of the present invention described above do not limit the scope of protection of the present invention. Any other corresponding changes and modifications made based on the technical concept of the present invention should be included in the scope of protection of the claims of the present invention.
Claims
1. A visual place recognition method based on frequency domain and spatial domain dual-domain aggregation, characterized in that: Including steps: Build a geotagged image database; Given an input image, extracting a local spatial domain feature map of the input image using a visual feature extractor; Converting the spatial domain feature map to the frequency domain by two-dimensional fast Fourier transform to obtain a frequency domain feature map; Optimizing the spatial domain feature map and the frequency domain feature map in parallel through a multi-scale contextual attention module to obtain optimized spatial domain feature map and frequency domain feature map; Aggregating the optimized spatial domain feature map and the frequency domain feature map into a unified global feature through triple fusion; A reference image that is most similar to the input image is retrieved from the image database based on the global features, thereby identifying the location of the input image based on the geographic tag of the reference image.
2. The visual place recognition method based on frequency domain and spatial domain dual-domain aggregation according to claim 1 is characterized in that: The method comprises the following steps: extracting a local spatial domain feature map from the input image using a visual feature extractor, and converting the spatial domain feature map into a frequency domain by a two-dimensional fast Fourier transform to obtain a frequency domain feature map. Use the visual feature extractor to extract local spatial domain feature maps: F s =F(I) in Represents a local spatial domain feature map; H and W represent height and width respectively, and D is the feature dimension; F s Information used to describe the visual content of the input image in the spatial domain; The F is transformed by two-dimensional fast Fourier transform (FFT) s Convert to the frequency domain: F f =FFT(F s ) in It is the frequency domain feature map.
3. The visual place recognition method based on frequency domain and spatial domain dual-domain aggregation according to claim 2 is characterized in that: The spatial domain feature map and the frequency domain feature map are optimized in parallel by a multi-scale context attention module, comprising the steps of: Among them, IFFT(·) converts the frequency domain feature map back to the spatial domain, and MSC-ATTN(·) is a multi-scale context attention module. The multi-scale context attention module converts F s and F f Downsample respectively to obtain the optimized spatial domain feature map And the frequency domain feature map Than F s and F f More differentiated and compact; Where S = (H / 4) × (W / 4).
4. The visual place recognition method based on frequency domain and spatial domain dual-domain aggregation according to claim 3 is characterized in that: The multi-scale context attention module optimizes the spatial domain feature map and the frequency domain feature map in the following steps: The local spatial domain feature map F s Project into query, key, and value matrices Q, K, and V with the same resolution H×W; Q=W q F s ,K=W k F s ,V=W v F s The matrix is stratified downsampled: Where D(·) represents a 3×3 convolution with a stride of 2, which can achieve a 2×2 downsampling rate. In order to integrate multi-scale information, the keys and values of different scales are flattened and concatenated to obtain multi-scale key and value features: Query after downsampling Operate with multi-scale keys and values to obtain optimized features: The optimization process of the frequency domain feature map is consistent with the optimization process of the spatial domain feature map.
5. The visual place recognition method based on frequency domain and spatial domain dual-domain aggregation according to claim 4 is characterized in that: The optimized spatial domain feature map and the frequency domain feature map are aggregated into a unified global feature by triple fusion, including the following steps: According to the optimized frequency domain feature map The spatial domain feature map is in the form of To reorganize: where q s (·), k s (·) and v s (·) are linear projection functions used to generate query, key and value respectively, F s′ Contains spatial domain information, but With similar signal organization structure, similarly, F f′ Obtained through: F s′ and F f′ The signal forms have been reorganized to be compatible with each other, reducing and Domain gap between features; based on implicit alignment of F s′ and F f′ , using the self-attention mechanism to fuse into a unified global feature: in cat(·;·) represents the concatenation operation of two features, and the aggregated unified global feature F contains feature information from both spatial and frequency domains.
6. The visual place recognition method based on frequency domain and spatial domain dual-domain aggregation according to claim 5 is characterized in that: Based on the global features, a reference image that is most similar to the input image is retrieved from the image database. The retrieval process is formalized as follows: where s(·,·) represents the similarity measure, F(·) represents the visual feature extractor used to extract local features from the input image, and f agg (·) is an aggregation function that aggregates local features into compact global features.
7. The visual place recognition method based on frequency domain and spatial domain dual-domain aggregation according to claim 1 is characterized in that: The geotagged image database includes at least one of the following: Pitts30k dataset, MSLS dataset, SPED dataset, Tokyo24 / 7 dataset, and NordLand dataset.
8. A visual place recognition system based on frequency domain and spatial domain dual-domain aggregation, characterized in that: include: Data construction module for building a database of geotagged images; A spatial domain feature extraction module, given an input image, uses a visual feature extractor to extract a local spatial domain feature map of the input image; A frequency domain conversion module is in communication with the spatial domain feature extraction module and converts the spatial domain feature map into the frequency domain by a two-dimensional fast Fourier transform to obtain a frequency domain feature map; a feature optimization module, which is in communication with the spatial domain feature extraction module and the frequency domain transformation module, and optimizes the spatial domain feature map and the frequency domain feature map in parallel through a multi-scale context attention module to obtain the optimized spatial domain feature map and the frequency domain feature map; A feature fusion module is connected to the feature optimization module and aggregates the optimized spatial domain feature map and the frequency domain feature map into a unified global feature through triple fusion; An image retrieval module is in communication with the data construction module and the feature fusion module, and retrieves a reference image that is most similar to the input image from the image database based on the global features, thereby identifying the location of the input image based on the geographic tag of the reference image.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the visual place recognition method based on frequency domain-spatial domain dual-domain aggregation as described in any one of claims 1 to 7.