Visual position identification method
Through global feature extraction network and relationship perception global attention technology, multi-scale features of images are extracted and enhanced, and the problem of poor recognition effect in complex appearance environments is solved, and more robust and generalized visual position recognition is achieved.
Patent Information
- Application Number
- CN202311463849.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-13
AI Technical Summary
The existing visual position recognition methods are poor in the face of complex appearance environment changes such as image viewing angle, lighting, and seasons, resulting in poor recognition results.
A global feature extraction network is adopted to extract features of different scales and fusion, and the fusion features are enhanced and redundant noise is suppressed by using relationship perception global attention. Finally, MixVPR is used for feature encoding to generate a robust image global descriptor.
It improves the robustness and generalization of image descriptors to changes in view angles and complex appearance environments, and achieves better visual position recognition effect.
Smart Images

Figure CN119992432A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network space surveying and mapping, and in particular relates to a visual position recognition method. Background Art
[0002] With the rapid development of the Internet, cyberspace, as a carrier of information technology, communication technology, and network science and technology, has become the fifth largest space after "land, sea, air, and space". In recent years, the new research field of "cyberspace mapping" combines cyberspace with geographic space to mine the geographic attributes in network data, realize the tracking and precise positioning of cyberspace resources, events, and relationships, and solve the core problem of "where is the target".
[0003] Image data, as one of the most common forms of big data analysis, contains rich visual information and is known as "a picture is worth a thousand words". In recent years, the widespread use of mobile devices such as smartphones and tablets has made it possible for us to obtain more and more image data, such as tourist photos taken during travel and street photos taken casually on the street. More and more images with information such as buildings, landmarks, and text notes appear on social media, but they do not directly contain the geographic location information we need, which limits further image information mining. Therefore, using Internet image positioning technology to geolocate such street view images and obtain the geographic location of the shooting location is the key and prerequisite for information mining and comprehensive analysis. At present, using images and related information posted by Internet social users to geolocate the entity targets of interest in the image and realize the geographic mapping of network virtual users is one of the key technologies for cyberspace mapping and is of great significance to national security.
[0004] Given a target image, the task of roughly estimating the geographic location of the currently captured target image based on the geographic locations of a set of previously visited images is called Visual Place Recognition (VPR). The VPR task has long been regarded as an image retrieval problem, and the geographic location of the target image is determined based on the known geographic location tags of the most relevant image in the reference library. Currently, scene recognition is difficult due to complex factors such as short-term or long-term illumination changes, camera shooting angle changes, moving object occlusion, building appearance changes, and repeated structures between target images and reference images, which in turn leads to poor robustness and generalization of existing VPR methods. The key technology is to design an image descriptor with strong robustness and good generalization, so that it has good adaptability to complex appearance environment changes such as image perspective, illumination, season, and interfering target occlusion.
[0005] In the early traditional VPR research, most of the classic VPR methods used rule-based manual local features such as SIFT, SURF, and ORB combined with the visual word bag model (Bag of Words, BoW), Fisher Vector (FisherVector, FV) or local aggregation descriptor vector (Vector of Locally Aggregated Descriptors, VLAD) to construct image descriptors to characterize images. The image description constructed by the traditional VPR method has a certain robustness to changes in image perspective, but is more sensitive to scenes such as long-term complex environmental changes such as lighting, weather and seasonal changes, and has poor robustness and generalization.
[0006] In recent years, with the rapid development of deep learning, various data-driven convolutional neural networks (CNNs) such as AlexNet, VGG, ResNet, and Transformer have achieved great success in image classification, face recognition, pedestrian re-identification, and scene recognition tasks with their powerful feature learning and image representation capabilities. Similarly, CNN has also been applied to VPR tasks. Currently, optimizing the VPR algorithm based on CNN has become the preferred method for VPR tasks. The reason is that the deep features extracted based on CNN have better robustness to image perspective and appearance changes, the features extracted from the middle layer of CNN as image descriptors have good robustness to appearance changes, and the high-level fully connected layer as an image descriptor has a certain robustness to perspective changes.
[0007] At present, existing VPR methods usually use multi-scale information and attention mechanisms to enhance the performance of image descriptors, so that they can better cope with complex appearance environment changes such as image perspective, lighting, season, and interference target occlusion. However, the existing VPR methods use multi-scale information either by using convolution kernels of different sizes and expansion coefficients on the last layer of the network to extract multi-scale features, or by processing multiple image resolutions before and after training. They often ignore the use of multi-scale information generated in the feature extraction process and the loss of target feature information caused by the continuous downsampling of small-sized key targets. The use of attention mechanisms is often through learning through convolution within a limited receptive field, ignoring the mining of knowledge from global structural patterns, lacking the overall effectiveness enhancement of local features, and usually performing well on data sets with single image perspective changes. They often fail in cases with complex appearance environment changes and cannot cope with complex appearance environment changes such as image perspective, lighting, and seasons, resulting in poor VPR recognition results. Summary of the invention
[0008] The purpose of the present invention is to provide a visual position recognition method to solve the problem that the existing methods cannot simultaneously cope with the poor recognition effect caused by changes in complex appearance environments such as image viewing angle, lighting, and seasons.
[0009] To solve the above technical problems, the present invention provides a visual location recognition method, which uses a trained global feature extraction network to perform feature extraction on a query image and each reference image in a reference library to obtain respective global descriptors, calculates the similarity between the query image and each reference image according to the respective global descriptors, and identifies the location of the query image according to the geographical location of the reference image with the highest similarity; wherein the global feature extraction network is used to generate a global descriptor in the following manner: first, extract features of different scales of the input image, and fuse the features of different scales to obtain a fused feature; then, use relation-aware global attention to enhance the fused features and suppress redundant noise to obtain an enhanced feature; finally, use MixVPR to feature encode the enhanced features to generate a global descriptor.
[0010] The beneficial effects of the above technical solution are as follows: the global feature extraction network used in the present invention first needs to extract features of different scales of the input image and perform fusion processing to compensate for the problem of insufficient expression ability of single-scale feature information, so as to better cope with complex appearance environment changes, and then use relationship-aware global attention to enhance the fused features and suppress redundant noise processing to avoid repetition and redundancy between features of different scales, and mine the spatial dependency relationship between features from the global structure to obtain the corresponding attention weights, so as to further enhance the robustness and generalization of the image descriptor in response to changes in image perspective and complex appearance environment changes, and finally use MixVPR for feature encoding to construct a more compact and robust image global descriptor, so as to achieve an overall improvement in the robustness and generalization in response to factors such as changes in image perspective and complex appearance environment.
[0011] Furthermore, the relationship-aware global attention is used to perceive global attention through spatial relationships and channel relationships.
[0012] The beneficial effects of the above technical solution are as follows: relation-aware global attention (i.e., RGA) provides two types of relation-aware global attention, namely RGA-S (realizing spatial relationship-aware global attention) and RGA-C (realizing channel relationship-aware global attention). By using RGA to enhance the fused features, one is that redundant noise information can be further suppressed to achieve the purpose of focusing on learning key targets in scene images; the second is that the fused features can be modeled for structural information from a global spatial range, which can better mine the semantic information of static objects such as buildings to effectively deal with occlusion caused by dynamic objects such as pedestrians and cars, further improving the expression ability of the fused features.
[0013] Furthermore, features of different scales include shallow-level features, mid-level features, and high-level features.
[0014] The beneficial effects of the above technical solution are: features of three different scales contain feature information at different levels, and making full use of features of different scales facilitates obtaining features that are robust to multiple factors such as image perspective and lighting, thereby achieving feature complementarity and helping to improve the global feature extraction effect of the global feature extraction network.
[0015] Furthermore, the method of fusing features of different scales is to use adaptive mean pooling to unify the sizes of shallow-level features, medium-level features and high-level features, and then connect the unified-sized shallow-level features, medium-level features and high-level features in series along the channel direction.
[0016] The beneficial effects of the above technical solution are: using adaptive mean pooling to unify the sizes of features at different levels, improving processing quality and efficiency, and facilitating subsequent fusion by a simple and effective way of serial merging along the channel direction.
[0017] Furthermore, the network for extracting features of different scales of the input image includes a preliminary feature extraction module and three groups of residual modules connected in sequence, and the outputs of the first group of residual modules, the second group of residual modules, and the third group of residual modules are the shallow-level features, the medium-level features, and the high-level features, respectively.
[0018] The beneficial effect of the above technical solution is: on the basis of preliminary feature extraction, the residual module is used to extract richer features, which serve as the basis for subsequent feature fusion.
[0019] Furthermore, the loss function used to train the global feature extraction network is multiple similarity loss.
[0020] The beneficial effects of the above technical solution are: multiple similarity loss alleviates the problems of too large inter-class distance and too small intra-class distance in metric learning by considering multiple similarities. It no longer uses absolute distance in space as the only metric, but uses the overall distance distribution of other sample pairs in a BatchSize to weight the loss. This calculation method can effectively promote model convergence in the early stage.
[0021] Furthermore, the normalized Euclidean distance is used to calculate the similarity between the query image and each reference image, and the smallest Euclidean distance indicates that the query image has the highest similarity with the reference image.
[0022] The beneficial effect of the above technical solution is: using Euclidean distance to calculate the similarity between two images is simple and easy to use. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flow chart of the VPR technology based on image retrieval according to an embodiment of the present invention;
[0024] Figure 2 is an overall technical framework diagram of MRGA-Mix of an embodiment of the present invention;
[0025] Figure 3 Schematic diagram of the principle of the deep residual network used in the embodiment of the present invention;
[0026] Figure 4 Schematic diagram of the principle of multi-level feature fusion and RGA feature enhancement in an embodiment of the present invention;
[0027] Figure 5 is a schematic diagram of spatial relationship perception global attention using RGA in an embodiment of the present invention;
[0028] Figure 6 is a schematic diagram of channel relationship perception global attention using RGA in an embodiment of the present invention;
[0029] Figure 7 is a detailed network structure diagram of a local aggregation vector network MixVPR in an embodiment of the present invention;
[0030] Figure 8 This is a thermal comparison diagram of several VPR methods;
[0031] Fig. 9 This is a comparison chart of global feature retrieval and positioning experiments of several VPR methods;
[0032] Fig.10 This is a comparison chart of the recall rates of several VPR methods on the commonly used Pittsburgh250k (with perspective changes) dataset;
[0033] Fig.11 This is a comparison chart of the recall rates of several VPR methods on the commonly used Tokyo 24 / 7 (with viewpoint changes and illumination changes) dataset. DETAILED DESCRIPTION
[0034] The core innovation of the present invention is that, based on MixVPR, the multi-scale features and attention mechanism generated by the feature extraction process are fully utilized to propose a VPR method that integrates multi-level features and relation-aware global attention (MRGA-Mix). This method requires the use of a global feature extraction network to extract global descriptors. The specific extraction process is as follows: first, features of different scales of the input image are extracted, and the features of different scales are integrated to obtain integrated features to make up for the defect of insufficient information expression ability of single-scale features; then, the integrated features are enhanced and redundant noise is suppressed by relation-aware global attention to obtain enhanced features, so as to achieve the purpose of focusing on learning key targets in scene images; finally, MixVPR is used to feature encode the enhanced features to generate global descriptors. Furthermore, the specific process of the visual location recognition method implemented based on the global feature extraction network is as follows: the global feature extraction network after training is used to extract global features of the query image and each reference image in the reference library to obtain their respective global descriptors, and then the similarity between the query image and each reference image is calculated according to their respective global descriptors, and the reference image with the highest similarity is selected as the matching image, and the location of the query image is identified according to the geographical location of the matching image. The entire solution can better cope with factors such as changes in image perspective and complex appearance environment, and has good robustness and generalization performance.
[0035] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings. Figure 1 As shown, the specific process is as follows:
[0036] Step 1: construct a global feature extraction network. The global feature extraction network constructed in this embodiment is as follows: Figure 2 As shown, the process of generating a global descriptor is as follows:
[0037] 1) Use deep convolutional neural networks to extract features at different levels (i.e., features at different scales) of the input image.
[0038] The background of street view images is relatively complex, and there are differences in target sizes, as well as problems such as perspective, lighting, and occlusion. Using only the final feature map of CNN for recognition is likely to ignore small targets on the image, the spatial position dependencies between targets, and the inability to obtain features that are robust to image perspective, lighting, seasonal changes, etc. The experiment used ResNet to extract multi-scale features. First, the commonly used ResNet18, ResNet34, and ResNet50 were directly tested on the public street view dataset. Among them, ResNet18 and ResNet34 achieved good results in terms of comprehensive performance. Therefore, ResNet18 and ResNet34 were selected as the backbone network to extract feature maps at different levels of the image. Cutting out the final fully connected layer, adaptive mean pooling layer, and the last residual module layer4 of ResNet18 and ResNet34 can reduce the number of model parameters by more than half. Figure 3 As shown, in this embodiment, ResNet undergoes convolutional layer, batch normalization layer, activation function and maximum pooling layer ( Figure 3 The batch normalization layer and activation function are omitted in the algorithm to complete the preliminary feature extraction, and then three groups of residual modules layer1, layer2 and layer3 are used to extract deeper features.
[0039] according to Figure 3 The ResNet feature extraction process in , ResNet is divided into 5 stages. After the input image or feature map is processed by each stage, the number of channels is often expanded, while the width and height are reduced to 1 / 2 of the original. For example: when the input image size is 320×320×3, the feature maps output at different stages are recorded as C 1 , C 2 , C 3 , C 4 and C 5 , the feature map sizes are 160×160×n 1 ,80×80×n 2 ,80×80×n 3 ,40×40×n 4 and 20×20×n 5 , n i Represents the number of channels, i∈{1, 2, 3, 4, 5}. The feature maps of the first two stages of ResNet (C 1 and C 2 ) information is not rich enough and is generally not used for VPR. Therefore, the feature maps (also called features) C of the last three stages are 3 , C 4 , C 5 The features are taken out and recorded as low-level features, medium-level features and high-level features in turn for subsequent feature fusion processing.
[0040] 2) The multi-scale features generated in the feature extraction process are fused to make up for the defect of insufficient information expression ability of single-scale features.
[0041] Making full use of the multi-scale features generated during feature extraction and fusing CNN multi-layer features is an effective way to improve scene recognition. The feature maps at the shallow level of CNN are large in size, high in resolution, small in receptive field, rich in spatial detail information, but weak in high-level semantic information; the feature maps at the deep level of CNN are small in size, low in resolution, large in receptive field, and severely lose spatial detail information, but rich in high-level semantic information. Most methods only use high-level CNN features and ignore the characteristics of CNN features at different levels, and cannot extract features that are robust to multiple factors such as viewing angle and illumination. Therefore, when the scene appearance changes dramatically, the robustness and generalization are poor. Compared with only using high-level features, multi-level fusion can integrate the multi-level information of the scene image, realize feature complementarity, and help achieve better recognition results.
[0042] like Figure 4 As shown, first C 3 , C 4 and C 5 Take out one by one, among which the middle-level feature C 4 Contains more geometric information and is more robust to changes in illumination, while high-level features C 5 Contains more semantic information, which can effectively overcome the change of perspective; secondly, in order to make full use of C 3 , C 4 , C 5 The global context information of C is unified using adaptive mean pooling 3 , C 4 and C 5 To avoid C 3 , C 4 and C 5 The feature map information is lost too much during the downsampling process and considering C 5 The semantic information of C is rich enough, so 3 and C 4 Adjust the size to C 5 The same size, i.e. 20×20; finally, the obtained C′ 3 , C′ 4 , C′ 5 The size of the series connection along the channel is (n 3 +n 4 +n 5 )×20×20C′ 6 , called fusion features.
[0043] 3) Relation-aware global attention (RGA) is used to enhance the fused features and suppress redundant noise information, so as to focus on learning the key targets in the scene image.
[0044] The relative position relationship of visual features is crucial to the location recognition results. In order to maintain the robustness and generalization of location recognition in different environmental changes, the network model needs to have the ability to screen important environmental features and obtain the dependency relationship between visual features at any two locations. 6 In order to address the redundant noise information in the image and the insensitivity to the spatial position dependency of image features, we use relation-aware global attention RGA to 6 The enhancement is performed by capturing the dependency between any two positions of the feature graph through the relationship affinity matrix to globally learn the attention weight of each feature node. Specifically, for each feature node, by superimposing its paired relationship with all feature nodes as a vector, a compact relationship representation is established and the corresponding attention weight is mined and learned from it. Specifically, it can be RGA-S for learning spatial attention weights and RGA-C for learning channel attention weights.
[0045] RGA-S is used to learn the spatial attention weight map. For the input feature tensor C′ 6 ∈R C×H×W , taking the C-dimensional feature vector at each spatial position as the feature node, it contains a total of N = H × W nodes, such as Figure 5 As shown, N feature nodes are represented as x i ∈R C , where i = 1, ... N, the affinity relationship between node i and node j is represented by r i,j It is expressed as, calculated by formula (1).
[0046] r i,j =θ S (x i ) T φ S (x j ) (1)
[0047] In the formula, θ S and φ S represents a 1×1 convolutional layer with a BN layer and ReLU activation.
[0048] Similarly, the affinity relationship r from node j to node i is calculated by formula (1): j,i And use the affinity matrix R S ∈R N×N Represents the pairwise relationship of all nodes, such as Figure 5 R S The 7th row and 7th column of is used as the relationship feature, i.e. r7 =[R s (7,:),R s (:,7)], where R s (7,:) represents the affinity relationship between the 7th feature node and all nodes. Similarly, R s (:,7) represents the affinity relationship between all nodes and the 7th node, which are derived as the 7th spatial position attention.
[0049] In order to calculate the attention of the i-th feature node, in addition to the pairwise relationship r i In addition, the feature C′ needs to be input 6 The spatial relationship perception feature y is calculated by formula (2) i .
[0050]
[0051] In the formula, ψ s and They represent the input features and the global affinity embedding function, respectively, which consists of a 1×1 convolutional layer with a BN layer and a ReLU activation. c Represents global average pooling along the channel dimension.
[0052] According to formula (2), the spatial relationship perception feature y is obtained i Finally, the attention value of the spatial position node is obtained by mining the associated semantic information from this information through a learnable function. The calculation method is shown in formula (3).
[0053] a i =Sigmoid(w 2 ReLU(w 1 y i ) (3)
[0054] In the formula, w 1 and w 2 They represent the weights of the 1×1 convolutional layer, and Sigmoid() represents the activation function.
[0055] RGA-C is used to learn the C-dimensional channel attention weight map, taking the d = H × W dimension feature map of each channel as the feature node, including C nodes, and using Represents, where i = 1, ..c. Similar to the RGA-S calculation, the pairwise affinity relationship r between node i and node j can be obtained. i,j With r j,i , the obtained pairwise affinity relationships are represented by the affinity matrix R c ∈R C×C Indicates that Figure 6 As shown, the relationship vector r obtained for the sixth feature node6 =[R c (6,:),R c (:,6)] to represent the global structural information of the 6th node. In order to obtain the attention weight of the i-th feature node, a similar calculation as RGA-S can be performed to obtain the attention value of the i-th node at the channel position.
[0056] 4) The local aggregation vector network MixVPR is used to encode local features to construct a robust and compact image descriptor, namely the global descriptor.
[0057] The purpose of feature encoding is to integrate local features in a holistic manner to generate a more compact and robust image descriptor. In previous work, most VPR methods use NetVLAD to encode and construct image descriptors, but this method is easily affected by interfering objects, usually treating visual components as independent entities and ignoring their spatial relationship with the surroundings. At the same time, the dimension of the image descriptor encoded by this method is often too high, resulting in a large number of NetVLAD parameters. Therefore, the present invention uses a relatively lightweight local aggregation vector network MixVPR that considers feature spatial relationships for feature encoding to construct a more compact and robust global image descriptor. The MixVPR network structure is as follows: Figure 7 shown.
[0058] For the input C′ c ∈R C×H×W , MixVPR first transforms the three-dimensional tensor C′ 7 ∈R C×H×C It is regarded as a set of C two-dimensional features of size H×W, as shown in formula (4):
[0059] F={X i},i={1,…,C} (4)
[0060] In the formula, X i Corresponding to the i-th activation map in the feature map F. Secondly, each two-dimensional feature Figure X i Expand it into a one-dimensional vector to represent it, thus obtaining the flattened feature map C′ 7 ∈R C×N , where N = H × W. Then, the C flattened feature maps C′ 7 It is sent to the feature mixer Feature-Mixer, which is composed of L multi-layer perceptrons (MLP) with the same structure, such as Figure 7 As shown in Figure 2, Feature-Mixer takes the flattened feature map set as input and sequentially incorporates the spatial global relationship into each X i ∈C′ 7 As shown in formula (5):
[0061] Xi ←W 2 (σ(W 1 X i ))+X i ,i={1,…,C} (5)
[0062] Where W 1 and W 2 are the weights of the two fully connected layers that make up the MLP, and σ is the nonlinear activation function ReLU.
[0063] For C′ 7 ∈R C×N , Feature-Mixer produces an output Z∈R with the same shape due to its isotropic architecture C×N , which is then fed into the second Feature-Mixer block, and so on, until L consecutive blocks are reached, as shown in Formula (6):
[0064] Z=C′ 7 M L (C′ 7 M L-1 (…C′ 7 M 1 (C′ 7 ))) (6)
[0065] In the formula, Z and feature map C′ 7 have the same dimensions.
[0066] To control the dimensions of the final global descriptor, two fully connected layers are used to transform the channel and row dimensions in turn. First, a depth projection is used to transform Z from R C×N Mapping to R d×N , as shown in formula (7)
[0067] Z′=W d (Transpose(Z)) (7)
[0068] Where W d is the weight of the fully connected layer. Then a row-by-row projection is used to transform the output Z′ from R d×N The resulting output is mapped to R d×r , as shown in formula (8):
[0069] O=W r (Transpose(Z′)) (8)
[0070] Where W r is the weight of another fully connected layer. The final output O has the dimension d×r, which is flattened and L 2After normalization, a global feature vector, namely the global descriptor, is formed.
[0071] Step 2: Train the global feature extraction network constructed in step 1.
[0072] In existing VPR research, most of them use a triplet loss function based on weak supervision to train the network. The disadvantage of triplet loss training is that it involves a large amount of GPU memory usage and computational overhead. The present invention uses a multiple similarity loss function for training, which has been proven to perform best in VPR tasks. Multiple similarity loss is proposed on the basis of improving structural loss and triplet loss. It alleviates the problems of too large inter-class distance and too small intra-class distance in metric learning by considering multiple similarities. It no longer uses absolute distance in space as the only metric, but uses the overall distance distribution of other sample pairs in a BatchSize to weight the loss. This calculation method can effectively promote model convergence in the early stage. The calculation of multiple similarity loss function is shown in formula (9):
[0073]
[0074] In the formula, Indicates the loss value; P i Represents the set of positive sample pairs for each instance in each BatchSize; N i Represents the set of negative sample pairs for each instance in each BatchSize; S ij and S ik Represents the similarity between two images; α, β and m are hyperparameters, where m represents the minimum distance between positive and negative sample pairs, and its value is generally 0.1.
[0075] Step three, using the trained global feature extraction network obtained in step two to perform feature extraction on the query image and each benchmark image in the benchmark library to obtain their respective global descriptors, calculating the similarity between the query image and each benchmark image based on their respective global descriptors, selecting the benchmark image with the highest similarity as the matching image, and identifying the location of the query image based on the geographic location of the matching image.
[0076] The output result of the VPR network after feature extraction and encoding of the reference and query images is a vector of fixed dimension, so the similarity calculation between images is converted into the similarity calculation between feature vectors. Commonly used feature vector similarity calculation methods include: Euclidean distance, cosine distance, and hash distance. This embodiment uses the normalized (specifically, L2 norm normalization) Euclidean distance to calculate the similarity between the contents of two images. Let X i =[X 1 ,X 2 ,...,XN ] and Y i =[Y 1 ,Y 2 ,...,Y N ] are two N-dimensional feature vectors, then the Euclidean distance between the two feature vectors is calculated as shown in formula (10):
[0077]
[0078] The following experiments are conducted to verify the effect of the method of the present invention.
[0079] The experiment will be divided into two groups. The first group will be tested on the Pittsburgh250k, Pittsburgh30k and SF-XL-Val datasets, where the view angle changes significantly but the image appearance does not change significantly. The second group will be tested on the Tokyo 24 / 7, Nordland and SF-XL-Testv1 datasets, where the view angle changes and the image appearance changes significantly. Both groups of experiments use Gsv-cities as the training set, and finally test the test set in turn according to the trained model. The experiment uses the most commonly used Recall@N in VPR to evaluate the recognition accuracy of different methods, that is, if there is at least one image matching the query image in the first N returned retrieval results of a query image within 25 meters of the actual geographic space distance, then the query image is considered to be successfully located; the length of time for feature extraction of a single image on the CPU is used as the basis for evaluating the real-time performance of the algorithm; and FLOPs is used to evaluate the computational complexity of the CNN model.
[0080] (1) Performance comparison between MRGA-Mix and other VPR methods.
[0081] To test the performance of MRGA-Mix, Recall@N and feature dimension are used as evaluation indicators to compare MRGA-Mix with other VPR methods on 6 public datasets. The experimental results are shown in Table 1.
[0082] Table 1
[0083]
[0084] In Table 1, the bold part is the method of the present invention, and the bold and single underlined part is the optimal value obtained by the method of the present invention on Recall@N. Among the different VPR methods, this experiment uses the two main methods mentioned above to construct the global image descriptor, one is the feature map spatial pooling method, and the other is NetVLAD and its variants that emphasize local aggregation. From the results in Table 1, it can be seen that the method of the present invention, MRGA-Mix, is significantly better than the method of constructing image description by feature map spatial pooling. Compared with other VPR methods with additional multi-scale information, the Recall@1 (R@1 in Table 1) of the proposed method on Pittsburgh250k reached 94.20%, which is 4.3% and 7.5% higher than SPE-NetVLAD and MultiRes-NetVLAD, respectively; the Recall@1 of the proposed method on Tokyo24 / 7 reached 83.81%, which is 14.01% higher than MultiRes-NetVLAD; compared with other VPR methods with added attention, the Recall@1 of the proposed method on Pittsburgh250k was 3.6% and 4.5% higher than CRN and Gated NetVLAD, respectively; On 24 / 7, it is 13.65% higher than CRN; compared with the VPR method without adding multi-scale information and attention mechanism, it is 6.20%, 3.80%, 4.31% and 2.68% higher than SARE, SFRS, CospPlace and ConvAP respectively on Pittsburgh250k; on Tokyo 24 / 7, it is 9.01%, 3.51%, 20.64% and 11.75% higher than SARE, SFRS, CospPlace and ConvAP respectively. The experimental results show that compared with other VPR methods that add multi-scale information and attention mechanism to enhance image descriptors, the image descriptor constructed by the method of the present invention has better robustness and generalization.
[0085] In addition, from the Recall@N of different VPR methods in Table 1, it can be concluded that the image descriptors constructed by the feature map spatial pooling method represented by the GeM method and the image descriptors constructed by the method represented by NetVLAD have certain robustness to data sets with single image perspective changes such as Pittsburgh250k, but have poor robustness to complex appearance environment changes such as Tokyo 24 / 7 and Nordland. Both methods often fail under severe lighting and seasonal changes, especially on the Nordland data set with drastic seasonal changes. The Recall@1 of both methods is below 50%, especially for NetVLAD trained on Pittsburgh30k, whose Recall@1 on Nordland is only 7.86%, while the method MRGA-Mix of the present invention reaches the optimal value of 75%. At the same time, the performance of NetVLAD can be further improved by training on Gsv-cities with complex environmental changes, especially on the Nordland data set, where its Recall@1 is improved by 28.39% compared with the original.
[0086] In summary, other VPR methods cannot take into account both image perspective and complex appearance environment changes at the same time. They usually perform well on datasets with single image perspective changes, but often fail in cases with complex appearance environment changes. The method MRGA-Mix of the present invention can always maintain high recall accuracy in low dimensions, whether on the Pittsburgh250k, Pittsburgh30k and SF-XL-Val datasets that focus on perspective changes, or on the Tokyo 24 / 7, Nordland and SF-XL-Testv1 datasets with perspective changes, illumination changes and seasonal changes. The experimental results show that the method of the present invention has certain robustness and generalization to image perspective changes and complex appearance environment changes.
[0087] (2) Test of multi-level feature fusion and RGA feature enhancement effect.
[0088] The effects of feature fusion and RGA enhancement are tested by comparing the effectiveness index Recall@N, efficiency index T, FLOPs and model size of the multi-level feature fusion model Multi-level MixVPR (hereinafter referred to as M-Mix), the multi-level feature fusion and relationship-aware global attention model Multi-level RGA MixVPR (hereinafter referred to as MRGA-Mix), and the Baseline (MixVPR(Res18), MixVPR(Res34)) on 6 public datasets. The experimental results are shown in Table 2.
[0089] Table 2
[0090]
[0091] In Table 2, the bold part is the method of the present invention, and the bold and single underline indicate that the method of the present invention obtains the optimal value in Recall@N.
[0092] From the comparison of Recall@N of ResNet18 and ResNet34 on the test set, we can see that the image descriptor directly obtained by the fully connected layer is more robust to changes in perspective, but less robust to changes in appearance environment, especially on the Nordland dataset with obvious seasonal changes. From the comparison of Recall@N of ResNet18, ResNet34, MixVPR (Res18) and MixVPR (Res34) on the test set, we can see that the image descriptor obtained by the local aggregation vector network specially designed by VPR has better performance, whether on datasets with changes in perspective or complex appearance environment.
[0093] From the comparison of Recall@N of MixVPR and M-Mix on the test set, it can be seen that compared with the Baseline (MixVPR(Res18) and MixVPR(Res34)), the recall rates of M-Mix(Res18) and M-Mix(Res34) on the six data sets are higher than those of MixVPR(Res18) and MixVPR(Res34), and the feature extraction time is only extended by 4ms. The Recall@1 of M-Mix(Res18) is 1.03, 1.1, 2.38, 2.22, 8.62 and 5.1 percentage points higher than that of MixVPR(Res18), respectively, and the Recall@1 of M-Mix(Res34) is 0.35, 0.69, 0.86, 3.17, 2.15 and 3.0 percentage points higher than that of MixVPR(Res34), respectively. From the comparative test results of M-Mix and MixVPR, it can be seen that single-layer CNN features cannot extract features that are robust to multiple factors such as lighting and viewing angle. Making full use of the multi-scale features generated in the feature extraction process can effectively make up for the insufficient expression ability of single-scale features.
[0094] As shown in Table 2, Fig.10 and Fig.11As shown in the figure, from the comparison of Recall@N of MRGA-Mix(Res18), MRGA-Mix(Res34) and M-Mix(Res18), M-Mix(Res34) on the test set, it can be seen that the scene recognition ability of the model is further improved after using RGA, and the recall rate is improved to a certain extent on each data set, but the improvement is smaller on the data set with obvious perspective changes, and it is improved more on the data set with complex appearance environment changes. Among them, MRGA-Mix(Res18) has improved Recall@1 by 2.54, 2.06 and 0.9 percentage points on Tokyo 24 / 7, Nordland and SF-XL-Testv1 data sets compared with M-Mix(Res18), and MRGA-Mix(Res34) has improved Recall@1 by 1.27, 2.79 and 1.6 percentage points respectively compared with M-Mix(Res34). The experimental results show that spatial position dependency is more sensitive to complex appearance changes.
[0095] In summary, compared with ResNet18 and ResNet34, MRGA-Mix (Res18) and MRGA-Mix (Res34) of the present invention have a smaller model, lower computational complexity, and a slightly improved feature extraction time for a single image on the CPU while greatly improving the recall rate. Therefore, the method of the present invention is more practical and can achieve a good balance between effectiveness and efficiency.
[0096] (3) Analysis of the superiority of MRGA-Mix.
[0097] ① Grad-CAM visualization. In order to more intuitively verify the impact of MRGA-Mix on feature extraction, Grad-CAM is used to visualize the class activation map of two images at the same location in the Tokyo 24 / 7 dataset under different viewing angles and lighting conditions. The original image examples and experimental results are shown in Figure 2. Figure 8 As shown in the figure. The red area in the class activation map represents the feature area that the neural network pays more attention to during image recognition, and the blue area represents the area that the neural network pays less attention to during image recognition. The closer to the center of the red area, the greater the contribution of the area to the recognition result. Figure 8The test results show that most of the GeM area is blue and unresponsive, the red areas are scattered and not concentrated enough, and the red area focus is too small; NetVLAD, CRN, CosPlace and ConvAP are also in a state of unresponsiveness, and the red areas with large responses are not concentrated on the buildings. When the viewing angle and appearance change, the red areas change greatly, which is not robust enough; MixVPR can keep the red areas focused on the buildings during the day, but when the viewing angle and appearance environment change, the red areas also change greatly, spreading to the surroundings, and some red areas are not concentrated on the buildings; the method of the present invention, MRGA-Mix, can have a high response in the key areas after the image is processed by feature fusion and RGA, and the red areas are relatively concentrated. When the viewing angle and illumination change, the red areas also change but are always well concentrated on the roads and buildings, so they can show good performance in recognition.
[0098] ②MRGA-Mix global feature retrieval and positioning test. The present invention uses image retrieval and positioning to test the image representation performance of MRGA-Mix global features. The test results are as follows: Fig. 9 As shown. Fig. 9 In the experiment, the image to be located is an image from Tokyo 24 / 7, which has drastic changes in illumination and is difficult to identify. In this experiment, the first five retrieval and positioning results returned by the five VPR methods were compared with the method of the present invention, MRGA-Mix. The green retrieval results indicate that the actual geographic location distance to the image to be located is within 25 meters, and the red retrieval results indicate that the actual geographic location distance to the image to be located is beyond 25 meters. The experimental results show that under conditions of severe illumination changes, only the method of the present invention and MixVPR can successfully retrieve and locate the query image, and compared with MixVPR, the method of the present invention performs better, with more and more successful recall results.
Claims
1. A visual position recognition method, characterized in that: The trained global feature extraction network is used to extract features from the query image and each reference image in the reference library to obtain their respective global descriptors. The similarities between the query image and each reference image are calculated based on their respective global descriptors. The location of the query image is identified based on the geographical location of the reference image with the highest similarity. The global feature extraction network is used to generate a global descriptor in the following way: first, features of different scales of the input image are extracted, and the features of different scales are fused to obtain fused features; then, the relation-aware global attention is used to enhance the fused features and suppress redundant noise to obtain enhanced features; finally, MixVPR is used to encode the enhanced features to generate a global descriptor.
2. The visual position recognition method according to claim 1, characterized in that: The relationship-aware global attention is used to perceive global attention through spatial relationships and channel relationships.
3. The visual position recognition method according to claim 1, characterized in that: Features of different scales include shallow-level features, mid-level features, and high-level features.
4. The visual position recognition method according to claim 3, characterized in that: The way to fuse features of different scales is to use adaptive mean pooling to unify the sizes of shallow, medium and high-level features, and then connect the unified shallow, medium and high-level features in series along the channel direction.
5. The visual position recognition method according to claim 3, characterized in that: The network for extracting features of different scales of the input image includes a preliminary feature extraction module and three groups of residual modules connected in sequence, and the outputs of the first group of residual modules, the second group of residual modules, and the third group of residual modules are the shallow-level features, the middle-level features, and the high-level features, respectively.
6. The visual position recognition method according to any one of claims 1 to 5, characterized in that: The loss function used to train the global feature extraction network is multiple similarity loss.
7. The visual position recognition method according to any one of claims 1 to 5, characterized in that: The normalized Euclidean distance is used to calculate the similarity between the query image and each reference image, and the smallest Euclidean distance indicates that the query image has the highest similarity with the reference image.
Citation Information
Cited By
Unstructured environment significance semantic segmentation method and system
CN120495669A
A method and system for unstructured environment saliency semantic segmentation
CN120495669B
Visual identification method and device for mixed placement of articles in hazardous chemical substance storehouse
CN121170433A