Streetscape image matching and positioning method and system based on structural feature enhancement and scene alignment
By using panoramic image and structural feature enhancement modules in visual geolocation of street scene images, the problems of weak supervision of GPS and the flood of false positive samples are solved, and a more accurate and stable matching positioning effect is achieved in complex urban environments.
Patent Information
- Application Number
- CN202510280344.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-13
AI Technical Summary
The existing visual geolocation method for street scene images based on deep learning is difficult to achieve stable matching positioning in complex urban environments, and is affected by the problems of weak GPS supervision and the flood of false positive samples.
A method of matching and positioning of street scene images based on structural feature enhancement and scene alignment is proposed. By selecting panoramic images as reference images, an adaptive similar scene alignment module and structural feature enhancement module are designed to eliminate redundant information and focus on stable structural features.
It effectively avoids the problem of weak GPS supervision, improves the accuracy and stability of matching positioning in dynamic urban environments, and achieves more accurate perspective query image matching positioning.
Smart Images

Figure CN120147423A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of street view image visual geolocation, and particularly to a method and system for matching and positioning street view images based on structural feature enhancement and scene alignment. Background Art
[0002] In recent years, the Global Navigation Satellite System (GNSS) has become the main tool for obtaining geographical location information and plays a key role in people's daily production and life. However, the stability of GNSS signals is challenged in complex urban environments. Especially when facing dense building blockages and various electronic signal interferences, its positioning accuracy often fails to meet the requirements of high-precision positioning. The image-based visual geolocation method can assist in solving the geolocation task when GNSS signals are blocked in cities.
[0003] Visual geolocation (VG), also known as Visual Place Recognition or Image-based Localization, aims to estimate the geographical location where the query image is located without relying on additional information such as GPS. The specific process is as Figure 1As shown, the user uses a handheld device to obtain the image to be queried (upper left), and then identifies the top-K reference images that are most similar to the query image from a city-scale database with GPS tags (lower left), so as to estimate the geographical location of the query image through the geographical location (GPS) of the reference images (right). The blue pentagram represents the GPS coordinates of the correctly retrieved image, and the red pentagram represents the GPS coordinates of the incorrectly retrieved image. This task is usually transformed into an image retrieval problem, so the key to solving this problem lies in extracting the key features of the image. In recent years, deep learning technology has achieved great success in the field of computer vision with its powerful feature extraction ability, and has also promoted the rapid development of the visual geolocation task. However, there are still two major problems to be solved in the research on visual geolocation of street view images based on deep learning. (1) The problem of the proliferation of false positive samples caused by weak GPS tag supervision. Affected by the limited shooting perspective of perspective images, slight differences in the shooting direction and angle of images under the same GPS tag will lead to great differences in the scene content they contain. And the existing deep learning-based algorithms usually screen positive and negative samples according to GPS tags, which makes a large number of false positive samples exist in the selected positive samples. There is very little or no overlapping area between these samples and the query image, and a large amount of misleading information will inevitably be introduced during the training process, ultimately leading to model collapse. (2) The lack of the ability to perform stable matching and positioning in complex and dynamic urban environments. The dynamic changes in urban scenes are reflected in many aspects, such as the changes in natural landscapes such as lighting, weather and seasons, and the movement of objects (such as the movement of pedestrians, vehicles, obstacles), etc. Under the combined action of these factors, there are certain content differences in street view images taken at different times and from different perspectives, which puts high requirements on the ability of the matching and positioning algorithm to extract stable features.
[0004] To alleviate the problem of weak GPS tag supervision, Torii et al. cropped panoramic images into multiple images at a certain ratio to construct a reference image database, but the uncertainty of the hard cropping process not only cannot guarantee the integrity of the matching content, but also significantly increases the data volume. And to achieve stable geolocation in dynamic urban environments, some scholars are committed to constructing street view image datasets containing multiple time nodes and multiple conditions to enhance the robustness of the model to the dynamic changes in urban environments, but this will undoubtedly increase the complexity of data collection and preprocessing. Another part of the research focuses on the transformation of network models and the integration of attention mechanisms, aiming to improve the ability of the model to learn discriminative features by increasing network complexity. However, the above methods all focus on the extraction of object visual appearance information and ignore the enhancement of object structure information. Summary of the Invention
[0005] The present invention aims to solve the problems of weak GPS supervision and ignoring street view structure information, and proposes a street view image matching and positioning method and system based on structural feature enhancement and scene alignment. By selecting panoramic images as reference images, the problem of weak GPS supervision is effectively avoided, and an adaptive similar scene alignment module is designed to eliminate redundant scene information; a structural feature enhancement module is designed to guide the model to focus on the relatively stable structural information in the dynamically changing urban environment, making the perspective query image matching and positioning more accurate.
[0006] To achieve the above object, the technical solutions adopted are as follows:
[0007] The present invention provides a street view image matching and positioning method based on structural feature enhancement and scene alignment, including:
[0008] Select the LskNet operator with the fourth block and the fully connected layer removed as the backbone network for feature extraction, and use the first and second blocks of this backbone network to extract the shallow features of the perspective image and the panoramic image; then input the extracted shallow feature images into the structural feature enhancement module SFE to guide the model to focus on the stable structural information; then use the third block of the backbone network to extract the deep features of the perspective image and the panoramic image;
[0009] Input the deep features of the perspective image and the panoramic image into the adaptive similar scene alignment module ASAM to calculate the similarity of each potential matching region of the perspective feature image on the panoramic feature image, and cut down the region with the maximum similarity;
[0010] Introduce the feature aggregation module MixVPR to aggregate global features;
[0011] Use the weighted soft margin triplet loss function to train the model;
[0012] Use the trained model to match the most similar panoramic reference image from the panoramic image library for the perspective query image to estimate the geographical location of the perspective query image.
[0013] According to the street view image matching and positioning method based on structural feature enhancement and scene alignment of the present invention, further, the structural feature enhancement module SFE includes feature aggregation, and the feature aggregation specifically includes:
[0014] Given the input feature image F l , perform sum pooling operation along the channel axis to aggregate the feature map, and obtain the aggregated feature map F a .
[0015] According to the street view image matching and positioning method based on structural feature enhancement and scene alignment of the present invention, further, the structural feature enhancement module SFE includes line feature mask generation, and the line feature mask generation specifically includes:
[0016] The aggregated feature map F a is processed by two convolutional kernels, and these two convolutional kernels are respectively designed to produce the maximum response to vertical edges and horizontal edges; during the processing, for each pixel F a in the feature map F a (i,j), its local receptive field F a (i-u,j-v) is convolved with these two convolutional kernels respectively; subsequently, the sum of the squares of these two convolution results is calculated, and the square root is taken, and the obtained value is used as the intensity value F a (i,j) on the output mask a '(i,j);
[0017] Finally, after normalization, the final line feature attention mask M is obtained.
[0018] According to the street view image matching and positioning method based on structural feature enhancement and scene alignment of the present invention, further, the structural feature enhancement module SFE includes residual attention fusion, and the residual attention fusion specifically includes:
[0019] First, the number of channels of the line feature attention mask M is expanded to be the same as that of the input feature image F l to obtain the mask M' with expanded channels; M' is multiplied element-wise with the input feature image F l and then added to the input feature image F l to output the feature map F l '.
[0020] According to the street view image matching and positioning method based on structural feature enhancement and scene alignment of the present invention, further, the working process of the adaptive similar scene alignment module ASAM is as follows:
[0021] First, the left and right boundaries of the panoramic feature image are circularly spliced to obtain a seamless panoramic feature image The perspective query feature image is used as a sliding window to slide and search along the equatorial direction on the panoramic feature image; during the sliding search process, the similarity between the perspective feature image and the panoramic feature image in each direction is calculated according to the following formula:
[0022]
[0023] where C and H respectively represent the number of channels and the height of the feature image, and W m ,W prespectively represent the widths of the perspective feature image and the panoramic feature image; after calculating the similarity score, the region corresponding to the maximum similarity score is the scene alignment region of the panoramic image relative to the perspective image; subsequently, this region is cropped from the panoramic image.
[0024] According to the street view image matching and positioning method based on structural feature enhancement and scene alignment of the present invention, further, the working process of the feature aggregation module MixVPR is as follows:
[0025] For the feature image F extracted from the backbone network, first perform a flattening process, and then use a feature mixer to generate an output Z with the same shape as the feature image F 1 , and then send it to the second feature mixer, and so on, until it is sent to the L-th feature mixer. This process is expressed by the formula:
[0026] Z = FM L (FM L-1 (...FM 1 (F)))
[0027] wherein, the shape of Z is the same as that of the feature image F, and FM represents the feature mixer; two fully connected layers are added after the feature mixer to reduce the channel and row dimensions of Z; finally, perform flattening and L2 normalization to output the global feature descriptor.
[0028] According to the street view image matching and positioning method based on structural feature enhancement and scene alignment of the present invention, further, the mathematical formula of the weighted soft margin triplet loss function is as follows:
[0029]
[0030] wherein, F m represents the perspective query image feature descriptor, F p' , F p* represent the feature descriptors of the panoramic image that matches the perspective query image and the panoramic image that does not match after being cropped by ASAM, d(,) represents the Euclidean distance, and the parameter γ controls the convergence speed of the training process.
[0031] Further, the present invention also provides a street view image matching and positioning system based on structural feature enhancement and scene alignment, which is used to implement the street view image matching and positioning method based on structural feature enhancement and scene alignment as described above. This system includes:
[0032] The feature extraction module is used to select the LskNet operator after removing the fourth block and the fully connected layer as the backbone network for feature extraction, and use the first and second blocks of the backbone network to extract the shallow features of the perspective image and the panoramic image. Then, the extracted shallow feature images are input into the Structural Feature Enhancement module (SFE) to guide the model to focus on stable structural information. Next, the third block of the backbone network is used to extract the deep features of the perspective image and the panoramic image.
[0033] The alignment module is used to input the deep features of the perspective image and the panoramic image into the Adaptive Similar Scene Alignment module (ASAM) to calculate the similarity of each potential matching region of the perspective feature image on the panoramic feature image, and crop the region with the maximum similarity.
[0034] The feature aggregation module is used to introduce the feature aggregation module MixVPR to aggregate global features.
[0035] The model training module is used to train the model using the weighted soft margin triplet loss function.
[0036] The localization module is used to use the trained model to match the most similar panoramic reference image from the panoramic image library for the perspective query image to estimate the geographical location of the perspective query image.
[0037] Adopting the above technical solution, the beneficial effects obtained are:
[0038] The task of visual geolocation based on street view images is to estimate the geographic location of the query image by matching the most similar reference image from an image library with a geographic location (Global Positioning System, GPS) tag. The research on this technology has important application value in fields such as unmanned driving and robot navigation. Current visual geolocation methods generally use perspective street view images as reference images. The lack of scene content in such images due to the limited shooting angle is the main reason why the query image and the reference image under the same GPS tag cannot be correctly matched and positioned due to the mismatch in scene content. In response to the above problems, the present invention proposes a new architecture of visual geolocation based on perspective-panoramic street view images, which uses 360° panoramic street views as reference images, thereby avoiding the problem of mismatch in scene content caused by limited shooting angles, and greatly simplifies the data clipping preparation process of reference images. An improved feature extraction backbone network based on a structural feature enhancement module is used to mine the stable structural features of images. An adaptive similar scene alignment module is designed to eliminate feature redundancy and data asymmetry between perspective and panoramic images, and achieve preliminary scene alignment. Finally, a lightweight feature aggregation module MixVPR that considers spatial structural relationships is used to aggregate the features of the alignment area into a robust global feature descriptor to complete matching and positioning. After experiments, the R@1 index of the model proposed in this paper reached 72.5% and 58.4% on the Pitts250k-P2E and YQ360 datasets, respectively, achieving the most advanced performance compared with similar algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings of the embodiments of the present invention, wherein the drawings are only used to illustrate some embodiments of the present invention, but not to limit all embodiments of the present invention thereto.
[0040] Figure 1 It is a flow chart of the existing technology of visual geo-positioning based on street view imagery;
[0041] Figure 2 is a framework diagram of a street view image matching and positioning method based on structural feature enhancement and scene alignment according to an embodiment of the present invention, wherein the blue dotted box is a structural feature enhancement module;
[0042] Figure 3 is a flow chart of line feature mask generation according to an embodiment of the present invention;
[0043] Figure 4It is a flowchart of the adaptive similar scene alignment module according to an embodiment of the present invention; the yellow curve on the panoramic feature image represents the similarity score between the perspective and the panoramic feature along the equatorial direction, and the area corresponding to the maximum similarity score is the scene alignment area (red window) of the panoramic image relative to the perspective image; note that for the sake of saving space, the circular stitching process of the panoramic feature image is not shown in the figure;
[0044] Figure 5 It is an architecture diagram of the feature aggregation module MixVPR according to an embodiment of the present invention;
[0045] Figure 6 They are two examples of street view datasets according to an embodiment of the present invention. The left side is the perspective query image, and the right side is the panoramic reference image;
[0046] Figure 7 It is a performance comparison of the visual geolocation algorithm according to an embodiment of the present invention on the Pitts250k - P2E dataset;
[0047] Figure 8 It is a performance comparison of the visual geolocation algorithm according to an embodiment of the present invention on the YQ360 dataset;
[0048] Figure 9 It is a validation of the effectiveness of ASAM and SFE according to an embodiment of the present invention on the Pitts250k - P2E and YQ360 datasets;
[0049] Figure 10 It is a model visualization heat map according to an embodiment of the present invention;
[0050] Figure 11 They are the qualitative results according to an embodiment of the present invention on the Pitts250k - P2E and YQ360 datasets. The left side is the query image, and the right side is the Top - 3 images retrieved from the image database, where the green box and the red box represent successful retrieval and retrieval failure respectively. Detailed implementation manners
[0051] In the following, the exemplary solutions of the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the specific embodiments of the present invention. Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the art.
[0052] This embodiment discloses a street view image matching and positioning method based on structural feature enhancement and scene alignment, as Figure 2 shown, which includes the following steps:
[0053] Step S101: Select the LskNet operator with the fourth block and the fully connected layer removed as the backbone network for feature extraction, and use the first and second blocks of this backbone network to extract the shallow features of the perspective image and the panoramic image. Then, input the extracted shallow feature images into the Structure Feature Enhancement (SFE) to guide the model to focus on stable structural information, and use the third block of the backbone network to extract the deep features of the perspective image and the panoramic image. This solution avoids the problem of inconsistent scene content caused by limited shooting angles by matching the perspective image and the panoramic image. At the same time, in order to obtain stable features in a dynamic urban environment, a structure feature enhancement module is designed based on the principle of linear feature extraction. This module automatically identifies and enhances continuous and clear linear features in the image, guiding the model to focus on areas rich in structural information to extract more robust feature representations.
[0054] This model uses LskNet as the basic feature extraction network, which can dynamically adjust the convolution kernel size for different target objects in the image. This feature not only enables the model to extract the feature information of the target itself, but also prompts the model to capture the background information related to the foreground target information. On this basis, a structure feature enhancement module is designed and integrated into the LskNet network, aiming to further strengthen the model's learning ability of stable structural features in the image, thereby helping to generate more comprehensive and robust feature representations.
[0055] The main application scope of the visual geolocation task based on street view images is in large-scale urban environments. By studying the urban scene content, it can be found that there are a large number of ground objects with stable structural features in the city, such as buildings, roads, street lights, etc. Compared with visual appearance features, geometric structural features often remain stable in a dynamically changing urban environment and are not easily affected by factors such as the passage of time and seasonal changes and undergo large appearance changes. This characteristic determines that it can be used as a key feature for matching urban street view images. Generally speaking, ground objects with stable structural features usually have clear and continuous linear features, while messy dynamic objects such as pedestrians and vehicles lack obvious linear features and usually appear as irregular textures and shapes in the image. Based on the above understanding, this solution designs a structure feature enhancement module based on linear feature extraction. This module completes the modeling of structural features by identifying and capturing linear features in the image through a linear feature extractor, and guides the model to focus its attention on areas with stable structural features in the image, finally generating robust and reliable feature representations. This module can be flexibly embedded between different stages of the network, and its structure diagram is as Figure 2 shown by the blue dotted box. This module consists of three parts: feature aggregation, line feature mask generation, and residual attention fusion.
[0056] Given the input feature image Fl , in order to effectively highlight the significant information regions in the feature map, a sum pooling operation is performed along the channel axis to aggregate the feature map, and the aggregated feature map F a can be expressed as:
[0057] F a = SumPool(F l )
[0058] where SumPool represents the sum pooling operation.
[0059] The aggregated feature map F a is processed through two 3×3 convolutional kernels, as Figure 3 shown. These two convolutional kernels are respectively designed to produce the maximum response to vertical edges and horizontal edges. During the processing, for each pixel F a in the feature map F a (i,j), the local receptive field F a (i - u,j - v) is convolved with these two convolutional kernels respectively. Subsequently, the sum of the squares of these two convolution results is calculated, and the square root is taken. The obtained value is used as the intensity value F a (i,j) of the pixel on the output mask, which directly reflects the saliency of this pixel as a line feature. Finally, after normalization processing, the final line feature attention mask M is obtained. The specific operation of this process can be described by the mathematical expression: a M = ReLU(BN(F
[0060]
[0061] ')) a '))
[0062] where BN is batch normalization, ReLU is the rectified linear unit activation function, i,j represent the position coordinates of the element in the feature image, represents two different convolutional kernels, and u,v represent the position coordinates in the local receptive field area A = {-1,0,1}.
[0063] After obtaining the line feature attention mask M through the above processing, in order to further use it to highlight the important regions in the feature map, the number of channels of the line feature attention mask M is expanded to be the same as that of the input feature image F l . Assuming the number of channels of the input feature image is c, the mask M' after expanding the channels is:
[0064] M' = Expand(M,c)
[0065] where c is the number of channels of F l .
[0066] In this way, a mask M' with the same size as the input feature image can be obtained. The mask is equivalent to the weight of F. Since directly superimposing the line feature mask M' on the original feature image may lead to a decline in the model performance, therefore, a method similar to residual learning is selected. The obtained attention feature map and the original feature image F l are multiplied and added element-wise, and the output feature map F l ' can be expressed as: l ' can be expressed as:
[0067] F l ' = F l *M' + F l
[0068] Step S102: To address the significant data volume and information content asymmetry between the perspective query image and the panoramic reference image, an Adaptive Similar Scene Alignment Module (ASAM) is designed. The deep features of the perspective image and the panoramic image are input into the adaptive similar scene alignment module to calculate the similarity of each potential matching region of the perspective feature image on the panoramic feature image. The region with the maximum similarity is cropped for further refined matching. This process can discard most of the redundant information in the panoramic image, achieve scene alignment, and improve the accuracy and computational efficiency of the model.
[0069] In this solution, the panoramic image is selected as the reference image. The panoramic image is a collection of scene contents at a certain GPS coordinate, while the query perspective image is taken with a standard limited field of view camera, with a fixed and single scene. There are significant differences in the perspective coverage and content information between the two. Therefore, directly comparing at the feature level will introduce a large amount of redundant features into the matching process, thereby reducing the accuracy and efficiency of the matching. Inspired by the human visual system's recognition of scenes, that is, when humans try to find a specific scene, the human visual system will first determine the general visual appearance similarity, and then perform more refined scene element comparison and matching. To enable the neural network to imitate this process, this solution designs an adaptive scene alignment module. After extracting the deep features of the image, this module calculates the correlation between the query features and the panoramic features along the equatorial direction, and then uses this correlation to adaptively guide the network to crop the region with the highest response. This strategy aims to first screen out the candidate regions with similar scenes, thereby providing strong initial conditions for the subsequent more refined matching process. The specific process is as Figure 4 shown.
[0070] After the backbone network completes feature extraction, the perspective query feature image can be used as a sliding window on the panoramic feature image Perform a sliding search on it. However, since the panoramic image adopts an imaging method of cylindrical equidistant projection, there is discontinuity after its left and right boundaries are cropped and unfolded. In order to simulate the continuity and integrity in the real scene, before the sliding search window, the left and right boundaries of the panoramic image are first circularly stitched to obtain a seamless panoramic feature image Considering that the vertical viewing angles of the perspective feature image and the panoramic feature image are the same, but there are differences in the horizontal viewing angles, the sliding path will be along the equator direction of the panoramic feature image. During the sliding search process, the similarity between the perspective feature image and the panoramic feature image in each direction is calculated according to the following formula to quantitatively evaluate the correlation between the two. This strategy helps to more accurately locate the scene alignment area of the panoramic image relative to the query image
[0071]
[0072] Among them, C and H respectively represent the number of channels and the height of the feature image, and W m ,W p respectively represent the widths of the perspective feature image and the panoramic feature image
[0073] After calculating the similarity score, the region corresponding to the maximum similarity score is the scene alignment area of the panoramic image relative to the perspective image. Subsequently, the cropping technique is used to accurately crop this region from the panoramic image for subsequent processing
[0074] Step S103: Introduce a feature aggregation module MixVPR to aggregate global features
[0075] To achieve efficient and accurate visual geolocation tasks in an urban-scale database, it is necessary to adopt a suitable feature aggregation strategy to semantically express the feature image pairs that have achieved scene alignment after ASAM processing to generate a compact and robust feature representation. Based on this, this solution introduces a feature aggregation strategy MixVPR. This strategy uses multiple feature mixers with the same structure and completely composed of multi-layer perceptrons (MLPs) to iteratively incorporate the spatial structure relationships extracted by the structure feature enhancement module into each individual feature map. And, compared with traditional feature aggregation techniques CosPlace and NetVLAD, its number of parameters is less than half. MixVPR has proven its high performance and lightweight characteristics through some qualitative and quantitative results, and its architecture is as Figure 5 shown
[0076] For the feature image F∈R extracted from the backbone network H×W×c , it can be regarded as a set of c two-dimensional features of size H×W:
[0077] F = {X 1 ,X2 , ···, X c}
[0078] Among them, X i represents the i-th two-dimensional feature along the channel dimension in F.
[0079] Next, perform flattening on each X in F i , that is, use the Flatten() function to flatten X i ∈ R H×W into a one-dimensional vector, and finally obtain the flattened feature map F ∈ R c×n , where n = H × W.
[0080] Then apply the feature mixer FM 1 to F ∈ R c×n to produce an output Z of the same shape 1 ∈ R c×n . This process can be expressed by the mathematical formula as:
[0081] Z 1 = FM 1 (F) = W 2 (σ(W 1 X i )) + X i , i = {1, ··· c}
[0082] Among them, W 1 , W 2 represent the weights of two fully connected layers, and σ represents the non-linear activation function.
[0083] Then send Z 1 to the second feature mixer FM 2 , and so on, until it is sent to the L-th feature mixer FM L . Through the way of iterative fusion, the cross-level feature relationship fusion is effectively realized. This process can be expressed by the mathematical formula as:
[0084] Z = FM L (FM L-1 (...FM 1 (F)))
[0085] Among them, the shape of Z is the same as that of the feature map F; in order to further reduce its dimension, two fully connected layers are added after the feature mixer to reduce the depth (channel) and row dimension of Z. This operation can be regarded as a weighted pooling operation to control the size of the final global descriptor. First, use depth projection to map Z from R c×n to R d×n , as shown in the following formula:
[0086] Z' = Wd (Transpose(Z))
[0087] where W d is a fully connected layer. Next, apply a row-by-row projection to map the output Z' from R d×n to R d×r , as shown in the following formula:
[0088] O = W r (Transpose(Z'))
[0089] where W r is a fully connected layer. The dimension of the final output O is d×r. Finally, perform flattening and L2 normalization to output the global feature descriptor.
[0090] Step S104: Train the model using the weighted soft margin triplet loss function.
[0091] During the training process, ASAM acts on all pairs of perspective images and panoramic images. For the matching image pairs, the training process focuses on maximizing the similarity score between the perspective image feature descriptor and the feature descriptor of the scene alignment region (the region with the highest similarity to the perspective image) cropped from the panoramic image by ASAM, so as to enhance the model's ability to distinguish perspective differences. For the non-matching image pairs, although there is no region on the panoramic image with the same perspective as the perspective image, there is a region with the highest similarity to it, which is the most challenging part for distinguishing whether the image pair is a match. Therefore, during the training process, by minimizing the similarity score between the perspective image feature descriptor and the feature descriptor of this region, the model's ability to distinguish difficult features can be enhanced. Therefore, in this solution, the weighted soft margin triplet loss is used to train the model network, and the formula is as follows:
[0092]
[0093] where F m represents the perspective query image feature descriptor, F p' , F p* represent the feature descriptors of the panoramic image that matches the perspective query image and the panoramic image that does not match after being cropped by ASAM, d(,) is used to calculate the Euclidean distance between the two, and the parameter γ controls the convergence speed of the training process.
[0094] Step S105: Use the trained model to match the most similar panoramic reference image from the panoramic image library for the perspective query image to estimate the geographical location of the perspective query image.
[0095] Correspondingly to the above method, this embodiment also discloses a street view image matching and positioning system based on structural feature enhancement and scene alignment, including:
[0096] The feature extraction module is used to select the LskNet operator after removing the fourth block and the fully connected layer as the backbone network for feature extraction, and use the first and second blocks of this backbone network to extract the shallow features of the perspective image and the panoramic image. Then, the extracted shallow feature images are input into the Structure Feature Enhancement module (SFE) to guide the model to focus on stable structural information. Next, the third block of the backbone network is used to extract the deep features of the perspective image and the panoramic image.
[0097] The alignment module is used to input the deep features of the perspective image and the panoramic image into the Adaptive Similar Scene Alignment module (ASAM) to calculate the similarity of each potential matching region of the perspective feature image on the panoramic feature image, and cut down the region with the maximum similarity.
[0098] The feature aggregation module is used to introduce the feature aggregation module MixVPR to aggregate global features.
[0099] The model training module is used to train the model using the weighted soft margin triplet loss function.
[0100] The positioning module is used to use the trained model to match the most similar panoramic reference image from the panoramic image library for the perspective query image to estimate the geographical location of the perspective query image.
[0101] To verify the effectiveness of this solution, further explanations will be given below in combination with experimental data.
[0102] (1) Datasets
[0103] Relevant experiments were carried out on two public datasets, Pitts250k-P2E and YQ360. Figure 6 Some examples of these two datasets are shown, and Table 1 shows the sizes and partitions of these two datasets.
[0104] Pitts250k-P2E is constructed based on the original Pitts250k dataset. The reference image set in this dataset comes from Google Street View panoramic images of the Pittsburgh area downloaded from the Internet, and the query image set comes from 12 perspective images with relatively small yaw angles processed from each panoramic image. These query images overlap with the field of view of the panoramic images but are taken at different times.
[0105] The YQ360 dataset was collected by Zhejiang University. The captured content mainly includes roads with repetitive texture structures and distinctive open street views. Among them, the panoramic images were captured by a panoramic annular lens (PAL) camera mounted on the top of the vehicle, and the query images were captured by a pinhole camera. Different from the Pitts250k-P2E dataset, in order to simulate a more realistic test scenario, the vertical fields of view of the query images and panoramic images in this dataset do not completely overlap.
[0106] Table 1 Size and Division of the Dataset
[0107]
[0108] (2) Experimental Details
[0109] This algorithm is based on the PyTorch architecture and is trained and tested on an NVIDIA GeForce RTX 3090 graphics card. The number of training epochs is 100, the training optimizer is the Adam optimizer, and the learning rate selection adopts a stepwise learning rate adjustment strategy, that is, the entire training process of the network is divided into three stages. In the first 30 epochs of training, the learning rate is set to 1×e -3 ; in the 30th to 50th epochs of training, the learning rate is set to 1×e -4 ; in the subsequent training epochs, the learning rate is set to 1×e -5 . For the loss function, the most commonly used triplet loss function in the field of visual geolocation is selected to train the model. Each triplet consists of a query perspective image, a positive sample panoramic image, and multiple negative sample panoramic images. The definition of positive and negative samples is based on the distance between the reference image and the query image. During the training stage, 10 meters is used as the threshold, that is, image pairs with a distance less than or equal to 10 meters are regarded as positive sample pairs, and vice versa as negative sample pairs. In the test stage, in order to more comprehensively evaluate the performance of the model, this threshold is adjusted to 25 meters. In addition, for ease of processing, the sizes of the perspective street view images and panoramic street view images input into the network are adjusted to 128×128 pixels and 128×768 pixels respectively.
[0110] (3) Evaluation Metrics
[0111] Similar to the most widely used evaluation metrics for visual geolocation algorithms, this experiment uses R@K (Recall@K) as the metric to measure the performance of the model in the experiment. Specifically, if at least one of the top K database images retrieved is within 25 meters of the ground truth position of the query image, then the query image is considered to be correctly located, and the percentage of all query images that are successfully matched and located is R@K.
[0112] (4) Experimental Results
[0113] (4.1) Comparison of Results with Other Related Methods
[0114] To prove the effectiveness and superiority of the method proposed in the present invention, comparative experiments were carried out with other algorithms of the same type on two datasets, Pitts250k-P2E and YQ360. Among them, the NetVLAD and Berton et al. algorithms are relatively basic VG algorithms; the Dino_Mix and LskNet algorithms have powerful image feature extraction capabilities; the SFRS algorithm has the same research goal as the present invention, and also explores the weakly supervised problem in VG and trains the model in a self-supervised manner; while the Orhan et al. and PanoVPR algorithms adopt the same perspective-panoramic image VG method as the present invention.
[0115] ① Experimental Results on the Pitts250k-P2E Dataset
[0116] The experimental results on the Pitts250k-P2E dataset are shown in Table 2. Figure 7 This is a complete performance comparison chart for the experiments of these VG algorithms. It can be seen from the results that the accuracy of this algorithm leads comprehensively in four indicators and achieves the current optimal accuracy. In addition, while achieving the above excellent performance, the number of parameters required by this algorithm is only 5.67, and it also has obvious advantages in terms of lightweight. Generally speaking, the method proposed in the present invention achieves the best effect.
[0117] Table 2 Comparison of the Effects of Various Methods on the Pitts250k-P2E Dataset
[0118]
[0119] Note: The bold font is the optimal value in each column.
[0120] ② Experimental Results on the YQ360 Dataset
[0121] Deep learning is a technology highly dependent on data driving. However, the data volume of the YQ360 dataset is difficult to meet the requirements of the deep learning network for the number of samples during the training process. Therefore, to address the problem of insufficient data volume in the YQ360 dataset, this experiment adopted the same training method as PanoVPR, that is, using the model trained on the Pitts250k-P2E dataset as the pre-training weights for the training of the YQ360 dataset. The experimental results are shown in Table 3. Figure 8This is the complete performance comparison chart for the experiments of these several VG algorithms. It can be seen from the results that the accuracy of this algorithm is superior to the PanoVPR algorithm in terms of the two indicators of R@1 and R@10, achieving the highest accuracy at present. It is on par with the PanoVPR algorithm in terms of the two indicators of R@20, and is only slightly lower than the PanoVPR algorithm in terms of the R@5 indicator. At the same time, while achieving the above excellent performance, the number of parameters required by this algorithm is also the least. Therefore, generally speaking, the method proposed in the present invention has the best effect.
[0122] Table 3 Comparison of the effects of various methods on the YQ360 dataset
[0123]
[0124] Note: The bold font is the optimal value in each column.
[0125] From the comparison results in Table 2 and Table 3, the method of the present invention has achieved a large degree of leading compared with the algorithms of Orhan et al. and PanoVPR. This shows that the algorithm that pre-uses the sliding window strategy to accurately match regions on panoramic images before image matching is advanced. The reason for this result is that panoramic images inevitably contain a large amount of redundant information that is not relevant to perspective images due to their wider field of view. In the case of whole-image matching, this redundant information makes it impossible for the model to accurately extract image features related to the matching task from panoramic images, increasing the difficulty of key feature extraction. After precise matching of the regions, these complex redundant information can be significantly removed, thereby improving the algorithm accuracy. Secondly, compared with the same type of algorithms of Orhan et al. and PanoVPR, the method of the present invention still leads. The main reason is that the algorithms of Orhan et al. and PanoVPR need to generate feature descriptors to calculate the similarity within the sliding window during the sliding process, but the dataset does not provide region-level similarity labels, so it is impossible to supervise the model to learn feature descriptors that can distinguish perspective differences. In addition, the process of generating feature descriptors will inevitably lose structural and spatial information, which is very important for finding regions with consistent perspectives. In contrast, during the process of sliding the search window, this algorithm evaluates the similarity based on the overall pixel values within the window. This strategy does not rely on region-level similarity label supervision and maximally retains the structural and spatial information of the image. Therefore, this algorithm can more accurately locate the regions in panoramic images that have the same perspective as perspective images, providing a solid foundation for subsequent refined matching.
[0126] It should be noted that although this algorithm performs excellently on the Pitts250k-P2E dataset, the improvement in accuracy on the YQ360 dataset is relatively limited. This is mainly attributed to the unique challenges of the YQ360 dataset: (1) The YQ360 dataset is collected from areas with severe green coverage, resulting in serious occlusion of key matching features such as urban buildings and roads, and fewer effective matching features. (2) In order to simulate a more realistic test scenario, the vertical fields of view of the perspective images and panoramic images collected in the YQ dataset do not completely overlap, resulting in the loss of image content, a sharp increase in the matching ambiguity, and thus seriously affecting the matching accuracy of the model.
[0127] In addition, this algorithm has achieved a significant reduction in the number of parameters (from 50.22M to 5.67M), which is mainly due to the following two points: (1) The present invention innovatively introduces a feature extraction operator LskNet with lightweight design and a feature aggregation module MixVPR. LskNet obtains the context features of the image by sequentially decomposing the method to replace the traditional large convolution kernel, effectively reducing the number of parameters. MixVPR combines the global relationship between elements in each feature map in the cascade of feature mixing, avoiding complex local or pyramid aggregation operations, and further reducing the number of parameters. (2) The designed structural feature enhancement module and similarity scene alignment module of the present invention only contain very few learning parameters, and these parameters are the learnable affine parameters in the BN layer. Therefore, only negligible computational costs are required.
[0128] (4.2), Ablation experiments
[0129] (4.2.1), Hyperparameter ablation experiments
[0130] In this section, a series of ablation experiments are carried out to explore and quantify the impact of different hyperparameter configurations inside the adaptive similarity scene alignment module and the structural feature enhancement module on the model performance.
[0131] ① Hyperparameter ablation experiment of the adaptive similarity scene alignment module
[0132] During the process of the perspective feature image sliding and searching on the panoramic feature image as a sliding window, the set sliding step size and whether to perform circular stitching processing on the panoramic image are both key parameters affecting the model performance. Because the choice of the step size is directly related to the fineness during the sliding process, although performing circular stitching processing on the panoramic image ensures the continuity of the field of view, it also increases the additional computational amount and complexity. To verify the impact of different parameter combinations on the model performance, multiple ablation experiments are carried out on the Pitts250k-P2E and YQ360 datasets respectively, and the results are shown in Table 4.
[0133] Table 4 Influence of different sliding step sizes and whether to use circular stitching technology on ASAM
[0134]
[0135]
[0136] Note: The bold font represents the optimal value for each column. The parameter stride represents the step size during the sliding process, and the parameter cycle represents
[0137] whether to perform circular stitching processing on the panoramic image.
[0138] First, by comparing the effects of different stride settings {stride = 1, 2, 3, 4} in Table 4 on the model accuracy, it can be found that as the stride decreases, the recall rate of the model will increase accordingly (on the Pitts250k-P2E dataset, when the stride decreases from 4 to 1, R@1 increases from 68.7% to 73.9%; on the YQ360 dataset, when the stride decreases from 4 to 1, R@1 increases from 49.6% to 56.0%). The reason for this is that: the smaller the stride means the shorter the step size of the window moving on the feature map, so more local detailed information can be captured, and this local detailed information helps the model to more accurately judge the image content. In addition, the sliding process simulates the way the human visual system searches for specific scenes. In actual situations, the human visual system shows seamless and smooth characteristics when performing such searches. The smaller the stride means a smoother transition, which is closer to this natural and smooth search mode, so it can more effectively imitate the working mode of the human visual system.
[0139] Second, by comparing the effects of whether to perform circular stitching processing on the panoramic image on the model accuracy in Table 4 {cycle = true, false}, it can be found that performing circular stitching processing on the panoramic image can increase R@1 by 4.2% and 3.2% on the Pitts250k-P2E dataset and the YQ360 dataset respectively, and other evaluation indicators have also been widely improved. The reason for this result is that: the 360° field of view in the real scene is three-dimensional and circular. In order to project the three-dimensional 360° field of view onto a two-dimensional plane, it needs to be unfolded into a planar image. This unfolding process will cause regions that are continuous in three-dimensional space to be cut off and distributed at both ends in the two-dimensional plane. When the field of view of the perspective image is located in the cut-off region, it will cause an inconsistent field of view, resulting in insufficient matching features, which in turn affects the model accuracy.
[0140] ② Hyperparameter ablation experiment of the structural feature enhancement module
[0141] In deep learning, there are depth differences in the features extracted at different network stages. To explore the performance of the SFE proposed in this invention at each stage of the model, in this section, the three blocks of the baseline model are regarded as three stages, and the impact of injecting SFE at the three stages on the model accuracy is studied. Specifically, first, the impact of embedding SFE only after one of the three stages on the model performance is evaluated, and then the comprehensive improvement effect of jointly embedding SFE after multiple stages on the model performance is evaluated. The results are shown in Table 5.
[0142] Table 5 Comparative experiments on injecting the structural feature enhancement module at different stages
[0143]
[0144] Note: The bold font is the optimal value in each column, and the parameter Stage indicates injecting SFE after the {1, 2, 3}rd stage.
[0145] First of all, by comparing the impact of injecting SFE after a single stage {stage = 1, 2, 3} in Table 5 on the model accuracy, it can be found that compared with injecting SFE after the deep stage (stage 3), injecting SFE after the shallow stages (stage 1 and stage 2) can achieve higher performance. Especially injecting SFE after stage 2 makes the R@1 accuracy reach the maximum value on both datasets (on the Pitts250k - P2E dataset, the R@1 accuracy reaches 74.4%. On the YQ360 dataset, the R@1 accuracy reaches 58.4%). In addition, by comparing the results of injecting SFE after a single stage {stage = 1, 2, 3} with injecting SFE after multiple stages {stage = 1 + 2, 1 + 2 + 3} in Table 5, it can be found that embedding SFE only after a single stage has a better effect than jointly injecting SFE in multiple stages.
[0146] The experimental results show that injecting the structural feature enhancement module after relatively shallow stages (i.e., stage1, 2) can effectively improve the model accuracy, but injecting the structural feature enhancement module after deeper stages (i.e., stage3) leads to a decrease in the model accuracy. This is because: the shallow stages mainly reflect the basic attributes of the data, such as color attributes, texture attributes, and structural attributes. Emphasizing the extraction of linear features at this stage can effectively highlight the stable structural information in the image and lay a foundation for subsequent image matching. While the deep stages contain more complex and abstract high-level semantic information. Therefore, continuing to emphasize the extraction of linear features in the deep stages will interfere with the model's effective capture and processing of high-level semantic information, resulting in a decline in the overall performance of the model.
[0147] (4.2.2), Ablation experiment on the effectiveness of the module
[0148] In this section, relevant experiments are designed to verify the effectiveness of the proposed ASAM and SFE modules of the present invention. In the experiments, the lightweight LskNet network is selected as the baseline network, with the query image and the panoramic image as the network inputs. Fixing the optimal parameters determined in the experiments in Section 4.2.1, SFE and ASAM are respectively injected after the second block and the third block of the baseline network, and training and testing are carried out on two datasets, Pitts250k-P2E and YQ360. The comparison results are shown in Table 6.
[0149] Table 6 Verification of the effectiveness of ASAM and SFE on the Pitts250k-P2E and YQ360 datasets
[0150]
[0151] Note: The bold font is the optimal value in each column. The parameter Stage indicates that SFE is injected after the {1, 2, 3}rd stage.
[0152] Table 6 compares the influence of adding ASAM and SFE to the baseline model on the algorithm performance on the two datasets, Pitts250k-P2E and YQ360. First, adding the ASAM designed in the present invention to the BaseLine improves the R@1 accuracy by 35.6% and 16% respectively on the two datasets. In addition, the SFE proposed in the present invention also plays a role in the experiments, further improving the accuracy of the algorithm on the two general datasets (on the Pitts250k-P2E dataset, R@1 increases from 73.9% to 74.4%, and on the YQ360 dataset, R@1 increases from 56.0% to 58.4%). The complete R@K curve graph is as Figure 9 shown.
[0153] From the above experimental results, it can be concluded that the ASAM module and the structural feature enhancement module designed in the present invention play an important role in the model. This is because: (1) There is a large field of view difference between the panoramic image and the perspective image. When matching these two images by generating compact tensors of the same dimension, the panoramic feature vector needs to capture and encode the global information of the entire scene, while the perspective feature vector only needs to encode the global information of a partial scene. This results in a significant difference in the amount of information and the focus points contained in the feature tensors of the same dimension. Before generating the compact feature vectors in this study, the ASAM module was used to pre-determine the approximate orientation of the perspective image on the panoramic image, and then the cropping operation was used to extract the area consistent with the field of view range of the perspective image, effectively eliminating the influence of irrelevant areas on the panoramic image on image matching. (2) The traditional VG algorithm mainly focuses on the extraction of visual features of images. The model designed in this study expands the extraction and enhancement of structural features in images on the basis of traditional methods. As a more stable feature type in images that is rich in spatial layout information, structural features can better retain highly stable information such as buildings, street lights, and roads in a complex urban environment with dynamic changes. By extracting and enhancing the structural features in images, an accurate and reliable reference basis can be provided for visual geolocation tasks, thereby improving the model accuracy.
[0154] (4.3), Heatmap Visualization
[0155] To more intuitively demonstrate the role of the ASAM module designed in the present invention, Grad-CAM is used to generate a visualization feature heatmap to visually interpret the embedded feature map. As Figure 10 shown, a total of three groups of street view images (group1, group2, group3) are selected. In each group, on the left side of the first row is the perspective image, and on the right side is the panoramic image. The second row is the feature heatmap of the perspective image and the panoramic image after being processed by the baseline model and the structural feature enhancement module. The depth of color in the heatmap is proportional to the degree of network attention. The lighter the yellow, the lower the degree of network attention to this area, and vice versa, it is the key attention area. The yellow curve on the panoramic feature heatmap is the similarity curve between the perspective image and each area on the panoramic image generated by the ASAM module designed in the present invention. The area where the maximum value on the similarity curve is located (the red dashed box) is the area consistent with the field of view angle of the perspective image.
[0156] From Figure 10It can be seen that the region (red box) where the peak of the similarity curve (yellow curve) in the panoramic feature heat map is located is the region that exhibits the highest similarity to the perspective feature heat map in terms of hue distribution and structural distribution. This indicates that after the ASAM module designed in the present invention completes image feature extraction, it can efficiently and accurately locate the region on the panoramic feature image that has the same field of view angle as the perspective feature image, and this region is the region for subsequent cropping.
[0157] (4.4), Visualization of Retrieval Results
[0158] To visually demonstrate the retrieval and positioning efficiency of this algorithm, two pairs of representative image retrieval cases (query1, query2) are selected from the Pitts250k-P2E and YQ360 datasets respectively for visualization. The results are as Figure 11 shown. The results in the figure show that this algorithm can still exhibit accurate and robust retrieval and positioning capabilities in complex urban environments with varying lighting and viewpoints.
[0159] In summary, to solve the problem of the proliferation of false positive samples caused by weak GPS supervision, the present invention constructs a new framework for visual geographical positioning of perspective-panoramic street view images. This framework uses panoramic images to construct a reference dataset. At the same time, in order to eliminate the asymmetry of data capacity and information content existing in the matching process of perspective-panoramic images, an adaptive similar scene alignment module is designed. This module can adaptively determine the region in the panoramic image that is most similar to the perspective image without relying on the supervision of region-level similarity labels, thereby eliminating redundant information irrelevant to the matching. In addition, in order to improve the model's ability to perform stable matching and positioning in complex and dynamic urban environments, a structural feature enhancement module is designed to guide the model to focus on the more stable building structure information relative to visual appearance information. The experimental results show that the street view image matching and positioning method proposed in the present invention has a higher recall rate and fewer parameters than the baseline method and the state-of-the-art method.
[0160] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent substitution on some of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A street view image matching and positioning method based on structural feature enhancement and scene alignment, characterized in that: Include: The LskNet operator with the fourth block and the fully connected layer removed is selected as the backbone network for feature extraction, and the first and second blocks of the backbone network are used to extract shallow features of perspective images and panoramic images; then the extracted shallow feature images are input into the structural feature enhancement module SFE to guide the model to focus on stable structural information; and the third block of the backbone network is used to extract deep features of perspective images and panoramic images; The deep features of the perspective image and the panoramic image are input into the adaptive similar scene alignment module ASAM to calculate the similarity of each potential matching area of the perspective feature image on the panoramic feature image, and the area with the greatest similarity is cropped; Introduce the feature aggregation module MixVPR to aggregate global features; The model is trained using a weighted soft-margin triplet loss function; The trained model is used to match the most similar panoramic reference image from the panoramic image library to the perspective query image to estimate the geographic location of the perspective query image.
2. The street view image matching and positioning method based on structural feature enhancement and scene alignment according to claim 1, characterized in that: The structural feature enhancement module SFE includes feature aggregation, which specifically includes: Given an input feature image F l , perform summing and pooling operations along the channel axis to aggregate the feature map and obtain the aggregated feature map F a .
3. The street view image matching and positioning method based on structural feature enhancement and scene alignment according to claim 2, characterized in that: The structural feature enhancement module SFE includes line feature mask generation, which specifically includes: The aggregated feature map F a It is processed by two convolution kernels. are designed to produce the maximum response to vertical edges and horizontal edges respectively; during the processing, the feature map F a Each pixel F a The local receptive field F of (i,j) a (iu, jv) is convolved with the two convolution kernels respectively; then, the sum of the squares of the two convolution results is calculated and the square root is taken, and the obtained value is used as the pixel F a The intensity value F of (i,j) on the output mask a '(i,j); Finally, after normalization, the final line feature attention mask M is obtained.
4. The street view image matching and positioning method based on structural feature enhancement and scene alignment according to claim 3, characterized in that: The structural feature enhancement module SFE includes residual attention fusion, which specifically includes: First, the number of channels of the line feature attention mask M is expanded to match the input feature image F l The number of channels is the same as that of the original image, and the mask M' after the expanded channel is obtained; M' is combined with the input feature image F l After element-wise multiplication, it is then multiplied with the input feature image F l Add and output feature map F l '.
5. The street view image matching and positioning method based on structural feature enhancement and scene alignment according to claim 1, characterized in that: The working process of the adaptive similar scene alignment module ASAM is as follows: First, the panoramic feature image The left and right borders are stitched in a circle to obtain a seamless panoramic feature image Perspective query feature image As a sliding window, a sliding search is performed along the equatorial direction on the panoramic feature image. During the sliding search, the similarity between the perspective feature image and the panoramic feature image in each direction is calculated according to the following formula: Among them, C, H represent the number of channels and height of the feature image respectively, and W m ,W p Respectively represent the width of the perspective feature image and the panoramic feature image; after the similarity score is calculated, the area corresponding to the maximum similarity score is the alignment area of the panoramic image relative to the perspective image scene; then the area is cropped out from the panoramic image.
6. The street view image matching and positioning method based on structural feature enhancement and scene alignment according to claim 1, characterized in that: The working process of the feature aggregation module MixVPR is as follows: For the feature image F extracted from the backbone network, it is first flattened, and then the feature mixer is used to generate an output Z1 of the same shape as the feature image F, which is then sent to the second feature mixer, and so on, until it is sent to the Lth feature mixer. The process is expressed by the formula: Z=FM L (FM L-1 (...FM1(F))) Among them, the shape of Z is the same as the feature image F, and FM represents the feature mixer; two fully connected layers are added after the feature mixer to reduce the channel and row dimensions of Z; finally, flattening and L2 normalization are performed to output the global feature descriptor.
7. The street view image matching and positioning method based on structural feature enhancement and scene alignment according to claim 1, characterized in that: The mathematical formula of the weighted soft margin triplet loss function is as follows: Among them, F m represents the perspective query image feature descriptor, F p' , F p* represents the feature descriptors of the panoramic images that match the perspective query image and the panoramic images that do not match the ASAM cropped image, d(,) represents the Euclidean distance, and the parameter γ controls the convergence speed of the training process.
8. A street view image matching and positioning system based on structural feature enhancement and scene alignment, characterized in that: The system is used to implement the street view image matching and positioning method based on structural feature enhancement and scene alignment as described in any one of claims 1 to 7, comprising: The feature extraction module is used to select the LskNet operator with the fourth block and the fully connected layer removed as the backbone network for feature extraction, and use the first and second blocks of the backbone network to extract shallow features of perspective images and panoramic images; then the extracted shallow feature images are input into the structural feature enhancement module SFE to guide the model to focus on stable structural information; and then the third block of the backbone network is used to extract deep features of perspective images and panoramic images; An alignment module is used to input the deep features of the perspective image and the panoramic image into the adaptive similar scene alignment module ASAM to calculate the similarity of each potential matching area of the perspective feature image on the panoramic feature image, and cut out the area with the greatest similarity; Feature aggregation module, used to introduce feature aggregation module MixVPR to aggregate global features; A model training module for training models using a weighted soft-margin triplet loss function; The positioning module is used to estimate the geographic location of the perspective query image by matching the most similar panoramic reference image from the panoramic image library with the perspective query image using the trained model.
9. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.