Feature extraction method for network space image positioning
By using convolutional multi-layer perceptron and enhanced spatial attention modules in the geolocation of cyberspace images, the problem of insufficient robustness in the existing technology is solved, and more robust and generalized image positioning feature extraction is achieved.
Patent Information
- Application Number
- CN202311463850.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is not robust enough in the geographic positioning of images in cyberspace, especially in different seasons, lighting, occlusion and viewing angle changes, and it is difficult to obtain feature descriptors with strong robustness and generalization of image changes.
By acquiring semantic features of different scales of network space images, a stacked multi-layer convolution multi-layer perceptron is used for feature aggregation, and feature dimensionality reduction is combined with adaptive average pooling to obtain robust global features. At the same time, the enhanced spatial attention module is used to process semantic features of different scales, eliminate noise and enhance the robustness of multi-scale features. Finally, the global features and multi-scale features are fused through orthogonal projection method to realize image positioning feature extraction.
This method can effectively extract the key deep semantic feature information in the query image, enhance the robustness and generalization of image descriptors, and is suitable for image geolocation tasks, especially in the case of environmental changes and perspective changes, and exhibit high positioning accuracy.
Smart Images

Figure CN119992264A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cyberspace surveying and mapping, and in particular relates to a feature extraction method for cyberspace image positioning. Background Art
[0002] Image geolocation refers to the process of identifying and obtaining the geographic location of a query image in a pre-built image database for a given query image. Image geolocation in cyberspace is usually solved as an image retrieval task. Image geolocation plays an important role in many robotics and computer vision tasks, such as autonomous driving, SLAM, 3D scene reconstruction, and control point-free positioning of drones in GNSS-denied environments. The challenges of image geolocation tasks mainly come from the constantly changing external environment such as different seasons, different lighting, occlusions, moving objects, environments with high appearance similarity such as trees and buildings, and differences in camera perspectives. Therefore, how to obtain feature descriptors that are robust to image changes and have good generalization is one of the current researchers' concerns.
[0003] Traditional methods usually use hand-designed feature extraction operators such as SIFT, SURF, and HOG to obtain local features of the image, and then further use Fisher Vectors (FV), Bag of Words (BoW), and Vector of Locally Aggregated Descriptor (VLAD) to aggregate them into global descriptors representing the entire image to reduce the computational and storage overhead caused by the high dimension of the descriptor. With the rapid development of artificial intelligence and deep learning, deep learning methods represented by convolutional neural networks (CNN) and transformers have achieved good results in many computer vision tasks such as image classification, object detection, and semantic segmentation. Many researchers have also used CNN and Vision Transformer to solve image geolocation tasks.
[0004] Most studies directly train CNNs to complete location recognition and positioning by designing end-to-end trainable layers, and insert these trainable layers into pre-trained feature extraction networks to obtain rich and robust descriptor representations. These methods have been successful in large-scale image geolocation test datasets. Subsequent studies have enhanced the robustness of image descriptors by adding semantic and contextual information, using attention mechanisms, and leveraging multi-scale features to achieve better positioning results.
[0005] Recent studies have focused on improving the robustness of image-based location descriptions by introducing multi-scale information into global descriptors. Efficient use of multi-scale information is the key to enhancing the robustness of global descriptors. However, these CNN-based methods either process multiple image resolutions before and after training, or use convolution kernels of different sizes and dilation coefficients to extract multi-scale features on the last convolutional layer of the model, ignoring the information loss problem caused by continuous downsampling during multi-scale feature extraction, which affects the robustness of the descriptor. Summary of the invention
[0006] The purpose of the present invention is to provide a feature extraction method for cyberspace image positioning to solve the problem of insufficient robustness of the methods in the prior art.
[0007] In order to solve the above technical problems, the present invention provides a feature extraction method for cyberspace image positioning, comprising the following steps:
[0008] 1) Obtain semantic features of cyberspace images at different scales;
[0009] 2) The deepest semantic features are aggregated through a multi-layer convolutional multi-layer perceptron. After aggregation, the feature dimension is reduced using adaptive average pooling. After dimension reduction, the features are processed into global features for representing the entire image.
[0010] 3) Processing semantic features of different scales except the deepest layer to obtain multi-scale features;
[0011] 4) Global features and multi-scale features are integrated to realize image positioning feature extraction.
[0012] The beneficial effects of the above technical solution are as follows: considering that the convolutional multi-layer perceptron has strong feature extraction and expression capabilities, the present invention uses stacked multi-layer convolutional multi-layer perceptrons to perform feature aggregation to highlight channel information, and then uses adaptive average pooling to enhance spatial information, so as to obtain robust global features, highlight the content of the query image, especially the key features have stronger expression capabilities, and can effectively extract more critical deep semantic feature information in the query image; further, the global features and multi-scale features are integrated to realize feature extraction for cyberspace image positioning, which is more suitable for image geographic positioning tasks.
[0013] Furthermore, the convolutional multi-layer perceptron is a 1×1 convolutional multi-layer perceptron.
[0014] The beneficial effect of the above technical solution is that using a 1×1 convolutional multilayer perceptron can avoid excessive calculation and parameter amounts.
[0015] Furthermore, the processing process in step 3) is: the semantic features of different scales except the deepest layer are respectively subjected to the enhanced spatial attention module to embed spatial layout information and eliminate noise, so as to obtain the spatial attention weighted features corresponding to each semantic feature, and each spatial attention weighted feature is fused to obtain a multi-scale feature.
[0016] The beneficial effect of the above technical solution is: using the enhanced spatial attention module to eliminate the noise of multi-scale information generated in the feature extraction process can enhance the robustness of multi-scale features.
[0017] Furthermore, the processing process of the enhanced spatial attention module is: first use maximum pooling along the channel direction to reduce the dimension of the input features, use convolution kernels with different receptive fields to extract the reduced features to obtain intermediate features of different scales, and then perform tensor splicing on the intermediate features of different scales along the channel dimension. After splicing, the attention scores are processed to obtain the attention scores of the features, and the attention scores are broadcast to the input features to obtain the corresponding spatial attention weighted features.
[0018] The beneficial effects of the above technical solution are: the above enhanced spatial attention module can enable the network model to focus on areas that are valuable for image geolocation, suppress other irrelevant objects and noise, and improve the robustness of the spatial attention weighted feature expression.
[0019] Furthermore, in step 4), the orthogonal projection method is used to fuse the global features and the multi-scale features, and the processing process of the orthogonal projection method is: first, the projection of each multi-scale feature on the global feature is calculated pixel by pixel, and then the difference between the vector of the multi-scale feature and its projection vector is calculated and used as the corresponding orthogonal component, and finally the features of the orthogonal components are aggregated into an orthogonal vector descriptor, and spliced and fused with the global feature.
[0020] The beneficial effect of the above technical solution is that redundant scale information can be eliminated through the above method, so that the obtained multi-scale information and global information can enhance each other to generate a more compact descriptor.
[0021] Furthermore, ResNet50 with the last convolutional layer and classification layer removed is used to extract semantic features.
[0022] The beneficial effect of the above technical solution is that the above processing is more suitable for geographic positioning feature extraction.
[0023] Furthermore, the processing process after dimensionality reduction in step 2) is: flattening the features and performing L2 regularization.
[0024] The beneficial effect of the above technical solution is that the above processing can prevent the model from overfitting while ensuring the lightweight of the model.
[0025] Furthermore, the process of aggregating into orthogonal vector descriptors is as follows: first, adaptive average pooling is used to reduce the dimension of the orthogonal components, and then L2 regularization is performed after the dimension reduction.
[0026] The beneficial effects of the above technical solution are: first using adaptive average pooling can realize the aggregation of spatial features, and then performing L2 regularization processing can prevent the model from overfitting. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is the overall technical roadmap of the network architecture of the embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of a convolutional multilayer perceptron according to an embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of enhancing spatial attention proposed in an embodiment of the present invention;
[0030] Figure 4 It is a schematic diagram of the basic principle of feature orthogonal projection fusion adopted in an embodiment of the present invention;
[0031] Figure 5 This is a graph showing the result of the present invention's efficient use of multi-scale information proof;
[0032] Figure 6 These are the search result diagrams of the method of the present invention in some extreme cases;
[0033] Figure 7 It is a thermal diagram comparison between the method of the present invention and several other methods. DETAILED DESCRIPTION
[0034] Due to the ever-changing external environment such as different seasons, different lighting, occlusion, moving objects, images with highly similar appearance such as trees and buildings, and differences in camera perspectives, there are widespread problems in image geolocation such as insufficient robustness of image descriptors and feature redundancy in the use of multi-scale information, which makes the geolocation of cyberspace images still challenging. Based on this, the core innovation of the present invention lies in using a convolutional multi-layer perceptron (ConvMLP) for feature aggregation to obtain a robust global descriptor. The process of the feature extraction method for cyberspace image positioning of the present invention implemented based on this is: first, semantic features of different scales of cyberspace images are extracted; then the extracted deepest semantic features are subjected to feature aggregation by ConvMLP, and after aggregation, they are subjected to adaptive average pooling processing for dimensionality reduction, and then processed into global features for representing the entire image; and, the semantic features of different scales except the deepest layer are respectively subjected to the enhanced spatial attention module for spatial layout information embedding and noise elimination, so as to obtain spatial attention weighted features corresponding to each semantic feature, and each spatial attention weighted feature is fused to obtain multi-scale features to eliminate the noise of multi-scale information generated in the feature extraction process; finally, the global features and multi-scale features are fused to enhance the robustness of the image descriptor to environmental changes and perspective changes.
[0035] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings.
[0036] The network architecture used in this embodiment is as follows Figure 1 As shown, hereinafter referred to as ConvMLP-OFME, the process of a feature extraction method for network space image positioning implemented by using this network architecture is as follows:
[0037] Step 1: Obtain cyberspace images and extract semantic features of different scales. Specifically:
[0038] For a given input image That is, the network space image, first use the ResNet50 pre-trained on ImageNet with the last convolutional layer and classification layer removed as the feature extraction network to obtain semantic features of different scales. The three-dimensional tensor of a certain scale is F∈R c×h×w , h×w is the spatial resolution, and c is the number of channels. Figure 1 In this paper, semantic features of four scales are extracted.
[0039] Step 2: The deepest semantic features are aggregated through ConvMLP, and then adaptive average pooling is used to reduce the dimension of the features. The reduced features are flattened and L2 regularized to obtain a global descriptor for representing the entire image, i.e., the global feature. Specifically:
[0040] In this step, Figure 2 As shown, we first use a set of 1×1 convolutions to build a convolutional multilayer perceptron. For the feature F∈R obtained by the backbone network c×h×w , by stacking multiple layers of ConvMLP to achieve feature aggregation, then using adaptive average pooling to transform the spatial dimension of the feature from h×w to 1×1, and finally flattening the reduced feature and performing L2 regularization to obtain a global descriptor f that represents the entire image g ∈R c×1 Part of the process can be expressed by the following formula:
[0041] f g =AAP(ConvMLP D (ConvMLP D-1 (···ConvMLP1(F))))
[0042] Where AAP() represents adaptive average pooling; ConvMLP represents convolutional multi-layer perceptron; F∈R c×h×w represents the deepest semantic features obtained in step 1; D represents the depth of ConvMLP.
[0043] Here, during the processing, this method neither focuses on local features nor uses the attention mechanism. Instead, it uses the characteristics of 1×1 convolution to aggregate channel features and uses adaptive average pooling to aggregate spatial features, so as to achieve full feature aggregation while avoiding excessive parameters and calculations as much as possible.
[0044] Step 3: The semantic features of different scales except the deepest layer are respectively subjected to the enhanced spatial attention module to embed spatial layout information and eliminate noise, so as to obtain the spatial attention weighted features corresponding to each semantic feature, and each spatial attention weighted feature is fused (i.e., spliced) to obtain multi-scale features. Specifically:
[0045] In this embodiment, Figure 1 In addition to the deepest semantic features, there are three scales of semantic features, and three enhanced spatial attention modules are set in the corresponding network.
[0046] A common method used in image geolocation is to use Spatial Pyramid Pooling or similar structures to obtain spatial features of different scales. However, this method is only performed on the features extracted by the last convolutional layer of the model, and does not take into account the multi-scale information generated in the feature extraction process. Therefore, the present invention uses the features of different scales generated in the feature extraction process to perform feature fusion operations. Since the features generated in the feature extraction process are often shallow features and contain more noise, the present invention proposes an enhanced spatial attention module (ESA Module), which can embed spatial layout information into the feature representation, so that the network focuses on areas that are valuable for image geolocation and suppresses other irrelevant objects and noise, such as Figure 3 The calculation process of ESA Module can be expressed as:
[0047]
[0048] That is, for a given feature F∈R c×h×w , first use the maximum pooling along the channel direction To highlight the more valuable features to obtain S(F), and then use convolution kernels with different receptive fields to focus on information of different scales to obtain intermediate features of different scales. Here, it is specifically f in the formula 3×3 (), f 5×5 (), f 7×7 () processing, which are 3×3 convolution, 5×5 convolution, and 7×7 convolution. Then, the three features of different scales are concatenated along the channel dimension (that is, the ∪ operation in the formula), and then a 1×1 convolution f is used. 1×1 () and sigmoid function δ() to obtain the attention score about the input feature, and finally broadcast the attention score to the input feature, that is, Get the spatial attention weighted feature F'.
[0049] Step 4: Fusion the global features and multi-scale features to extract geographic positioning features.
[0050] Traditional multi-scale feature fusion is usually implemented by tensor splicing, but this method does not take into account the duplication and redundancy between information of different scales, which will affect the retrieval accuracy. Therefore, the present invention adopts a feature orthogonal projection fusion method, which can eliminate redundant scale information through this feature projection decomposition method, so that the obtained scale information and global information can enhance each other to generate a more compact image descriptor.
[0051] The basic principle of feature orthogonal projection fusion used in the present invention is as follows Figure 4As shown, it requires the global feature f g and multi-scale features f s As input, each multi-scale feature is calculated pixel by pixel In the global feature f g Projection on
[0052]
[0053] Then, by calculating the multi-scale feature vector Its projection vector The difference between them gives the corresponding orthogonal components
[0054]
[0055] Finally, for the orthogonal components The method of obtaining global features by imitating the face uses adaptive average pooling and L2 regularization to aggregate the features of the orthogonal components into a c×1 orthogonal vector descriptor and compares it with the global descriptor f of the same dimension. g Perform splicing and fusion, and use the fused descriptor vector for image retrieval.
[0056] The following experimental comparison demonstrates the effect of the method of the present invention. Specifically, several image geolocation methods based on global descriptor retrieval, including AVG, GeM, NetVLAD, SPE-NetVLAD, GatedNetVLAD, CosPlace, and ConvAP, are compared with the network architecture ConvMLP-OFME proposed in the present invention. All methods use ResNet50 pre-trained on ImageNet as the feature extraction network and are trained on GSV-Cities. The results are shown in Table 1 and Figure 6 shown.
[0057] Table 1
[0058]
[0059] As can be seen from Table 1, the method proposed in the present invention outperforms other methods on several co-open datasets. On the Pitts250k dataset, the method of the present invention achieves 92.5% Recall@1 (i.e., R@1 in Table 1), which is slightly improved compared with the previous methods. On the MSLS dataset, the method of the present invention obtains 86.5% Recall@1, which is 2% and 3.1% higher than CosPlace and ConvAP, respectively. This fully demonstrates that the network architecture proposed in the present invention can effectively cope with perspective changes and illumination changes in image geolocation tasks. On the SPED dataset and NordLand dataset with extreme illumination changes and appearance changes, the method of the present invention achieves the best performance of 80.6% and 43.2%, respectively. In addition, it can be seen from Table 1 that after orthogonally fusing multi-scale features on the basis of ConvMLP, the Recall@1 on each dataset is improved, with the maximum reaching 3%, indicating the effectiveness of the orthogonal fusion multi-scale information strategy adopted by the present invention.
[0060] like Figure 5 As shown in the figure, compared with other image geolocation methods with additional multi-scale information, the network architecture proposed in the present invention not only has a good retrieval accuracy performance, but also achieves a Recall@1 of 86.5% on the MSLS dataset, which is 8.3% and 3.7% higher than SPE-NetVLAD and MultiRes-NetVLAD respectively, and the required Floating Point Operations (FLOPs) are only 77.4% of the former and 50.7% of the latter.
[0061] In order to reflect the role of ConvMLP, 4 groups of experiments are conducted below by stacking only the number of ConvMLPs D. When D = 0, it is the BaseLine network used in the present invention, indicating that only adaptive average pooling is used to obtain the global descriptor for the features obtained by the backbone network. As can be seen from Table 2, when only one layer of ConvMLP is stacked, the Recall@1 performance of Pitts30k_val can be improved by 7.73 percentage points, from 83.94% to 91.67%, and by 12.56 percentage points, from 71.49% to 84.05% on MSLS. When the number of ConvMLPs is further stacked, the performance improvement achieved on the two data sets of Pitts30k_val and MSLS is limited, and even the problem of performance degradation occurs. Taking into account the increase in the number of parameters and calculations caused by stacking ConvMLP, stacking 1 ConvMLP is selected as the benchmark.
[0062] Table 2
[0063]
[0064] In order to more intuitively illustrate the feature extraction and expression capabilities of ConvMLP, Figure 7 The heat maps of the input images of several methods are listed. It can be seen that compared with the ResNet50, CosPlace, and ConvAP methods, the ConvMLP feature aggregation method proposed in the present invention can more accurately highlight the content of the query image, indicating that this method has a stronger ability to express key features and can effectively extract more critical semantic feature information in the query image, thereby achieving better performance.
[0065] As shown in Table 3, after adding ESA, the Recall@1 on the Pitts30k_val and MSLS datasets increased by 1.6% and 3.65%, respectively, indicating that adding ESA can effectively remove the noise of the shallow network and make the network focus on more valuable information for image geolocation. And from this experiment, it can be seen that if the multi-scale information generated in the feature extraction process is used alone to construct the image descriptor for retrieval and recognition, its Recall@1 on the Pitts30k_val and MSLS datasets is 8.31% and 13.78% lower than the global image descriptor obtained by ConvMLP, respectively, indicating that the image descriptor constructed using the multi-scale information generated in the feature extraction process alone is not suitable for solving the image geolocation task. This is because the multi-scale information generated in the feature extraction process contains more shallow features and the deep semantic information is insufficiently represented. Combined with the experimental results in Table 1, it can be seen that using multi-scale information to enhance the global descriptor obtained by ConvMLP can effectively increase the robustness and generalization of the descriptor.
[0066] Table 3
[0067]
[0068] In order to prove the effectiveness of the orthogonal fusion module (which implements the orthogonal projection method), we remove Figure 1 In the orthogonal fusion module, the multi-scale features f s With the global feature f g Direct splicing and fusion are performed for comparative experiments. Here, the use of Hadamard product to fuse two vectors is also explored, which is also a common method for fusing two vectors. The experimental results are shown in Table 4. Compared with the commonly used tensor splicing fusion method, the Recall@1 of the present invention on the Pitts30k_val and MSLS datasets is improved by 3.3% and 5.54%, respectively. It shows that through the orthogonal projection process, redundant information and repeated information in multi-scale features can be eliminated, making the output multi-scale information richer and more informative, so that a large number of shallow features in the multi-scale information will not affect the performance of the global descriptor, thereby achieving the purpose of complementary enhancement.
[0069] Table 4
[0070]
Claims
1. A feature extraction method for cyberspace image positioning, characterized in that: The steps include: 1) Obtain semantic features of cyberspace images at different scales; 2) The deepest semantic features are aggregated through a multi-layer convolutional multi-layer perceptron. After aggregation, the feature dimension is reduced using adaptive average pooling. After dimension reduction, the features are processed into global features for representing the entire image. 3) Processing semantic features of different scales except the deepest layer to obtain multi-scale features; 4) Global features and multi-scale features are integrated to realize image positioning feature extraction.
2. The feature extraction method for cyberspace image positioning according to claim 1, characterized in that: The convolutional multi-layer perceptron is a 1×1 convolutional multi-layer perceptron.
3. The feature extraction method for cyberspace image positioning according to claim 1, characterized in that: The processing process in step 3) is as follows: the semantic features of different scales except the deepest layer are respectively subjected to the enhanced spatial attention module to embed spatial layout information and eliminate noise, so as to obtain the spatial attention weighted features corresponding to each semantic feature, and each spatial attention weighted feature is fused to obtain a multi-scale feature.
4. The feature extraction method for cyberspace image positioning according to claim 3 is characterized in that: The processing process of the enhanced spatial attention module is as follows: first, the input features are reduced in dimension using the maximum pooling along the channel direction, and the convolution kernels with different receptive fields are used to extract the reduced features to obtain intermediate features of different scales, and then the intermediate features of different scales are tensor-joined along the channel dimension. After the joints, they are processed to obtain the attention scores about the features, and the attention scores are broadcasted to the input features to obtain the corresponding spatial attention weighted features.
5. The feature extraction method for cyberspace image positioning according to claim 1, characterized in that: In step 4), the orthogonal projection method is used to fuse the global features and multi-scale features, and the processing process of the orthogonal projection method is: first, the projection of each multi-scale feature on the global feature is calculated pixel by pixel, and then the difference between the vector of the multi-scale feature and its projection vector is calculated and used as the corresponding orthogonal component, and finally the features of the orthogonal components are aggregated into an orthogonal vector descriptor, and spliced and fused with the global feature.
6. The feature extraction method for cyberspace image positioning according to any one of claims 1 to 5, characterized in that: Semantic features are extracted using ResNet50 without the last convolutional layer and classification layer.
7. The feature extraction method for cyberspace image positioning according to claim 1, characterized in that: The processing process after dimensionality reduction in step 2) is: flattening the features and performing L2 regularization.
8. The feature extraction method for cyberspace image positioning according to claim 5, characterized in that: The process of aggregating into orthogonal vector descriptors is as follows: first, adaptive average pooling is used to reduce the dimension of the orthogonal components, and then L2 regularization is performed after dimension reduction.