A method for training a cross-view feature extraction model, a geolocation method and apparatus
By integrating visual and semantic features in the cross-view feature extraction model, and using the Transformer encoder to extract feature embeddings of street and aerial images, the problem of insufficient matching accuracy and robustness of cross-view image in the prior art is solved, and more efficient cross-view geolocation is achieved.
Patent Information
- Application Number
- CN202510455319.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing cross-view image geolocation methods rely on visual features, making it difficult to achieve accurate matching and poor robustness, especially when there are differences in view angles, geometric structures and resolutions between street images and aerial images.
A cross-view feature extraction model training method is adopted. By constructing a training sample pair set including positive and negative sample pairs, using a Transformer encoder model including visual branches and semantic branches, visual and semantic features of street and aerial images are extracted, feature embeddings are generated, and the model is optimized through iterative training.
By integrating visual and semantic features, the cross-view feature extraction model can capture higher-level information, such as geographical indications and road topology, significantly improving the performance of cross-view matching and improving the accuracy of geolocation.
Smart Images

Figure CN119992117B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of geolocation, and particularly to a method for training a cross-view feature extraction model, a geolocation method and a device. Background Art
[0002] The cross-view image geolocation method based on the CNN convolutional neural network usually relies on polar coordinate transformation to reduce the geometric difference between ground images and aerial images. However, this method requires prior knowledge of geometric bodies and may lead to matching failures when street view images and aerial images are not aligned. In addition, since the CNN convolutional neural network cannot effectively model global correlation and explicit position information, there is a large domain gap between the cross-view retrieval system for street view and aerial images. In view of these deficiencies, methods based on Transformer encoders have emerged in recent years, which utilize their advantages in global information modeling and explicit position information encoding to improve the performance of cross-view matching.
[0003] The core objective of the cross-view image geolocation method is to match the same geographical location from two different perspectives, namely the ground street view image (i.e., the street photo) and the aerial image. This technology has important application values in fields such as autonomous driving and navigation. Existing methods usually rely on visual features and extract features through deep learning models for matching. However, there are differences in the perspectives, geometric structures, and resolutions of street photos and aerial images, and it is difficult to achieve accurate matching and poor robustness only relying on visual features. Summary of the Invention
[0004] This application aims to at least solve the technical problems existing in the prior art, and provides a method for training a cross-view feature extraction model, a geolocation method and a device.
[0005] In a first aspect, the present application provides a method for training a cross-view feature extraction model, including: constructing a training sample pair set including one or more positive sample pairs and one or more negative sample pairs, each sample pair including a street view image and an aerial view image; constructing a cross-view feature extraction model, the cross-view feature extraction model including: a street view feature extraction module for obtaining a street view feature embedding of a street view image, the street view feature extraction module including a street view visual branch and a street view semantic branch, as well as a street view token fusion unit, a street view Transformer encoder, and a street view embedding feature generation unit connected in sequence, the input end of the street view token fusion unit being respectively connected to the output ends of the street view visual branch and the street view semantic branch; an aerial view feature extraction module for obtaining an aerial view feature embedding of an aerial view image, the aerial view feature extraction module including an aerial view visual branch and an aerial view semantic branch, as well as an aerial view token fusion unit, an aerial view Transformer encoder, and an aerial view embedding feature generation unit connected in sequence, the input end of the aerial view token fusion unit being respectively connected to the output ends of the aerial view visual branch and the aerial view semantic branch; using the training sample pair set to iteratively train the constructed cross-view feature extraction model in batches, and obtaining a cross-view feature extraction model after reaching the training stop condition.
[0006] In a second aspect, the present application provides a cross-view geolocation method based on the fusion of visual and semantic features, including: obtaining a street view image to be located; forming multiple image pairs by respectively combining the street view image to be located with multiple aerial view images in a pre-constructed aerial view image set;
[0007] inputting each image pair into a trained cross-view feature extraction model to obtain a pair of street view feature embeddings and aerial view feature embeddings, the trained cross-view feature extraction model being trained according to the method provided in the first aspect of the present application; calculating the distance between the street view feature embedding and the aerial view feature embedding in each image pair, and selecting the image pair with the smallest distance between the street view feature embedding and the aerial view feature embedding as the matching image pair; using the geographical location information associated with the aerial view image in the matching image pair as the geographical location information of the street view image to be located, and outputting the obtained geographical location information.
[0008] In a third aspect, the present application provides a cross-view geolocation device based on the fusion of visual and semantic features for implementing the cross-view geolocation method based on the fusion of visual and semantic features provided in the second aspect of the present application, including: an image input module for acquiring a street view image to be located; an image pair acquisition module for forming multiple image pairs by respectively combining the street view image to be located with multiple aerial images in a pre-constructed aerial image set; a feature embedding acquisition module for inputting each image pair into a trained cross-view feature extraction model to obtain a pair of street view feature embeddings and aerial image feature embeddings, where the trained cross-view feature extraction model is obtained by training according to the method provided in the first aspect of the present application; a matching image pair selection module for calculating the distance between the street view feature embedding and the aerial image feature embedding in each image pair, and selecting the image pair with the smallest distance between the street view feature embedding and the aerial image feature embedding as the matching image pair; and an output module for using the geographical location information associated with the aerial image in the matching image pair as the geographical location information of the street view image to be located and outputting the acquired geographical location information.
[0009] In a fourth aspect, the present application provides a computer program product including a computer program which, when executed by a processor, implements the steps of the method provided in the first aspect or the second aspect of the present application.
[0010] In a fifth aspect, the present application provides an electronic device, where the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the steps of the method provided in the first aspect or the second aspect of the present application.
[0011] The beneficial technical effects of this application are as follows: In the street photography feature extraction module and the aerial photography feature extraction module of the cross-view feature extraction model, corresponding visual branches and semantic branches are both set at the same time. Street photography visual tokens containing the visual features of street photography images and street photography semantic tokens containing the semantic features of street photography images are extracted simultaneously. After splicing the street photography visual tokens and the street photography semantic tokens, the spliced tokens are encoded and mapped to generate street photography feature embeddings. And aerial photography visual tokens containing the visual features of aerial photography images and aerial photography semantic tokens containing the semantic features of aerial photography images are extracted simultaneously. After splicing the aerial photography visual tokens and the aerial photography semantic tokens, the spliced tokens are encoded and mapped to generate aerial photography feature embeddings. This makes the obtained street photography feature embeddings and aerial photography feature embeddings both contain semantic features, supplementing the expressive ability of visual features, enabling the cross-view feature extraction model to capture higher-level information, such as geographical landmarks, road topologies, etc., significantly improving the performance of cross-view matching, and enabling the cross-view feature extraction model to better adapt to the perspective differences and geometric structure differences between ground images and aerial photography images, thereby improving the accuracy of cross-view geographical positioning. Description of the Drawings
[0012] Figure 1 is a schematic flowchart of a method for training a cross-view feature extraction model in a preferred embodiment of the present invention;
[0013] Figure 2 is a street photography image in a positive sample pair in an example;
[0014] Figure 3 is Figure 2 the aerial photography image in the positive sample pair in the example;
[0015] Figure 4 is a schematic structural diagram of a cross-view feature extraction model in a preferred embodiment of the present invention;
[0016] Figure 5 is a schematic diagram of the aerial photography image cropping process in a preferred embodiment of the present invention;
[0017] Figure 6 is a schematic flowchart of a cross-view geographical positioning method based on the fusion of visual and semantic features in a preferred embodiment of the present invention;
[0018] Figure 7 is a system block diagram of a cross-view geographical positioning device based on the fusion of visual and semantic features in a preferred embodiment of the present invention;
[0019] Figure 8 is a schematic structural diagram of an electronic device in a preferred embodiment of the present invention;
[0020] Reference Numerals:
[0021] 10 Processor; 11 Memory; 12 Communication Bus; 13 Communication Interface; 30 Street Photography Visual Branch; 301 First Linear Projection Unit; 302 First Position Embedding Unit; 31 Street Photography Semantic Branch; 311 First Semantic Feature Extraction Unit; 312 First MLP Unit; 32 Street Photography Token Fusion Unit; 33 Street Photography Transformer Encoder; 34 Street Photography Embedded Feature Generation Unit; 35 Street Photography Feature Embedding; 20 Aerial Photography Visual Branch; 201 Second Linear Projection Unit; 202 Second Position Embedding Unit; 21 Aerial Photography Semantic Branch; 211 Second Semantic Feature Extraction Unit; 212 Second MLP Unit; 22 Aerial Photography Token Fusion Unit; 23 Aerial Photography Transformer Encoder; 24 Aerial Photography Embedded Feature Generation Unit; 25 Aerial Photography Feature Embedding. Detailed Implementation Manner
[0022] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0023] In the description of the present invention, it should be understood that the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0024] In the description of the present invention, unless otherwise specified and defined, it should be noted that the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal communication of two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0025] The execution subject of the cross-view feature extraction model training method or the cross-view geolocation method based on the fusion of visual and semantic features provided by the present invention includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, the cross-view feature extraction model training method or the cross-view geolocation method based on the fusion of visual and semantic features can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0026] The present invention provides a cross-view feature extraction model training method. In a preferred embodiment, as Figure 1 shown, it includes:
[0027] Step A1, constructing a training sample pair set including more than one positive sample pair and more than one negative sample pair, and each sample pair includes a street view image and an aerial view image.
[0028] In this embodiment, the positive sample pair is composed of a street view image and an aerial view image from the same geographical location, and the negative sample pair is composed of a street view image and an aerial view image from different geographical locations. Figure 2 And Figure 3 shows an example of a positive sample pair. Among them, Figure 2 shows the street view image in the positive sample pair example, Figure 3 shows the aerial view image in the positive sample pair example. It can be intuitively observed that the street view image is an image obtained by a person taking a picture on the ground at a certain geographical location point with an image device, and the aerial view image is an image taken from a high altitude downward at a certain geographical location point. Preferably, in order to more clearly display the geographical location detail information, the aerial view image can be a bird's-eye view.
[0029] In this embodiment, the method for constructing a training sample pair set is as follows: Obtain an existing aerial image set. Each aerial image in the aerial image set has geographical location information of the shooting point. Street view images can be taken at the geographical location corresponding to each aerial image, so as to obtain a set of street view images that can correspond one-to-one with the aerial images in the aerial image set. Based on the aerial image set and the corresponding street view image set, the aerial image and the street view image at the same geographical location are formed into a positive sample pair, and the aerial image and the street view image at different geographical locations are formed into a negative sample pair. Multiple positive sample pairs and multiple negative sample pairs are generated in the above manner, and the multiple positive sample pairs and multiple negative sample pairs form a training sample pair set. Similarly, a validation sample pair set can be constructed according to the above method.
[0030] Step A2, construct a cross-view feature extraction model.
[0031] Please refer to the appendix Figure 4 , the cross-view feature extraction model includes:
[0032] A street view feature extraction module, which is used to obtain the street view feature embedding 35 of the street view image. The street view feature extraction module includes a street view visual branch 30 and a street view semantic branch 31, as well as a street view token fusion unit 32, a street view Transformer encoder 33, and a street view embedding feature generation unit 34 connected in sequence. The input end of the street view token fusion unit 32 is respectively connected to the output ends of the street view visual branch 30 and the street view semantic branch 31.
[0033] An aerial image feature extraction module, which is used to obtain the aerial image feature embedding 25 of the aerial image. The aerial image feature extraction module includes an aerial image visual branch 20 and an aerial image semantic branch 21, as well as an aerial image token fusion unit 22, an aerial image Transformer encoder 23, and an aerial image embedding feature generation unit 24 connected in sequence. The input end of the aerial image token fusion unit 22 is respectively connected to the output ends of the aerial image visual branch 20 and the aerial image semantic branch 21.
[0034] Exemplarily, taking one sample as the input, the process of the cross-view feature extraction model extracting the street view feature embedding 35 of the street view image and the aerial image feature embedding 25 of the aerial image in the sample pair is introduced below. It includes:
[0035] Both the street view image and the aerial image of the sample pair are subjected to block processing, and each image block is flattened.
[0036] The image patches of the street view images are input into the street view feature extraction module. The street view vision branch 30 linearly projects the flattened image patches into a fixed dimension to obtain a visual token of the street view image. A learnable position embedding is added to the visual token to retain the spatial information of the street view image. All the visual tokens of the street view image that have incorporated the position information are used as street view visual tokens. The street view semantic branch 31 extracts the semantic features of the street view image, flattens the semantic features, and projects them into the same feature dimension as the street view visual tokens to obtain street view semantic tokens. The street view token fusion unit 32 is used to concatenate the street view visual tokens and the street view semantic tokens, and the concatenated street view tokens are obtained after concatenation. The street view Transformer encoder 33 encodes the concatenated street view tokens to obtain the street view encoded features. The street view embedding feature generation unit 34 is used to perform a mapping transformation on the street view encoded features to obtain the street view feature embedding 35. The street view embedding feature generation unit 34 is preferably but not limited to an existing MLP Head unit or one or more cascaded fully connected layers.
[0037] The image patches of the aerial view images are input into the aerial view feature extraction module. The aerial view vision branch 20 linearly projects the flattened image patches into a fixed dimension to obtain a visual token of the aerial view image. A learnable position embedding is added to the visual token to retain the spatial information of the aerial view image. All the visual tokens of the aerial view image that have incorporated the position information are used as aerial view visual tokens. The aerial view semantic branch 21 extracts the semantic features of the aerial view image, flattens the semantic features, and projects them into the same feature dimension as the aerial view visual tokens to obtain aerial view semantic tokens. The aerial view token fusion unit 22 is used to concatenate the aerial view visual tokens and the aerial view semantic tokens, and the concatenated aerial view tokens are obtained after concatenation. The aerial view Transformer encoder 23 encodes the concatenated aerial view tokens to obtain the aerial view encoded features. The aerial view embedding feature generation unit 24 is used to perform a mapping transformation on the aerial view encoded features to obtain the aerial view feature embedding 25. The aerial view embedding feature generation unit 24 is preferably but not limited to an existing MLP Head unit or one or more cascaded fully connected layers.
[0038] Step A3: Use the training sample set to iteratively train the constructed cross-view feature extraction model batch by batch. After reaching the training stop condition, the cross-view feature extraction model is obtained.
[0039] In this embodiment, the training stop condition is preferably but not limited to that the number of training times reaches a preset maximum number of training times or the total batch loss converges. The total batch loss is the loss calculated after each batch is completed, and the total batch loss can be the batch triplet loss. Preferably, after step A3 is completed, step A4 is further included, and the cross-view feature extraction model obtained in step A3 is verified by using a pre-constructed set of verification samples. If the verification is passed, the cross-view feature extraction model obtained in step A3 is output. If the verification fails, the training parameters, such as the learning rate, are adjusted, and then step A3 is returned to be executed.
[0040] In this embodiment, preferably, the batch triplet loss is calculated by using the street view feature embeddings 35 and the aerial view feature embeddings 25 of multiple sample pairs in the subset of the training sample pairs of the current batch output by the cross-view feature extraction model. Exemplarily, the batch triplet loss is calculated by using the street view feature embeddings 35 and the aerial view feature embeddings 25 of all sample pairs included in each batch. :
[0041]
[0042] Among them, represents the first hyperparameter, which is used to control the margin range of the batch triplet loss, and the default value can be 0.5; represents the average value of the squares of the distances or L2 distances (i.e., Euclidean distances) between the street view feature embeddings 35 and the aerial view feature embeddings 25 of all positive sample pairs obtained from the subset of the training sample pairs of the current batch, , represents the number of positive sample pairs in the subset of the training sample pairs of the current batch, which is a positive integer, represents the index of the positive sample pairs in the subset of the training sample pairs of the current batch, represents the positive sample pairs in the subset of the training sample pairs of the current batch obtained from the squares of the distances or L2 distances between the street view feature embeddings 35 and the aerial view feature embeddings 25; represents the average value of the squares of the distances or L2 distances between the street view feature embeddings 35 and the aerial view feature embeddings 25 in the negative sample pairs in the subset of the training sample pairs of the current batch, , represents the number of negative sample pairs in the subset of the training sample pairs of the current training batch, which is a positive integer, represents the index of the negative sample pairs in the subset of the training sample pairs of the current batch, represents the negative sample pairs in the subset of the training sample pairs of the current batch obtained from the squares of the distances or L2 distances between the street view feature embeddings 35 and the aerial view feature embeddings 25.
[0043] In this embodiment, the batch triplet loss is used to increase the difference between the street view feature embedding 35 and the aerial view feature embedding 25 of the negative sample pair, while reducing the difference between the street view feature embedding 35 and the aerial view feature embedding 25 of the positive sample pair, thereby optimizing the global matching of the street view image and the aerial view image.
[0044] In this embodiment, preferably, the street view visual branch 30 includes a first linear projection unit 301 and a first position embedding unit 302 connected in sequence, and the street view semantic branch 31 includes a first semantic feature extraction unit 311 and a first MLP unit 312 connected in sequence; alternatively, the aerial view visual branch 20 includes a second linear projection unit 201 and a second position embedding unit 202 connected in sequence, and the aerial view semantic branch 21 includes a second semantic feature extraction unit 211 and a second MLP unit 212 connected in sequence.
[0045] In this embodiment, preferably, referring to Figure 4 As shown, the street view visual branch 30 includes a first linear projection unit 301 and a first position embedding unit 302 connected in sequence, and the street view semantic branch 31 includes a first semantic feature extraction unit 311 and a first MLP unit 312 connected in sequence; and the aerial view visual branch 20 includes a second linear projection unit 201 and a second position embedding unit 202 connected in sequence, and the aerial view semantic branch 21 includes a second semantic feature extraction unit 211 and a second MLP unit 212 connected in sequence.
[0046] In this embodiment, the first linear projection unit 301 and the second linear projection unit 201 are preferably but not limited to cascaded fully connected layers with more than one layer. The first linear projection unit 301 is used to linearly project the image blocks after flattening the street view image into a fixed dimension to obtain a visual token of the street view image. The first position embedding unit 302 is used to add learnable position embeddings to the visual tokens of the street view image to retain the spatial information of the street view image. The second linear projection unit 201 is used to linearly project the image blocks after flattening the aerial view image into a fixed dimension to obtain a visual token of the aerial view image. The second position embedding unit 202 is used to add learnable position embeddings to the visual tokens of the aerial view image to retain the spatial information of the aerial view image.
[0047] In this embodiment, the first semantic feature extraction unit 311 and the second semantic feature extraction unit 211 are not limited to a pre-trained DINOv2 model or a MAE (Masked AutoEncoders) model. Preferably, the DINOv2 model is selected because the DINOv2 model can learn high-quality feature representations from unlabeled data. Through a self-supervised learning method, it can extract deep semantic features of images without relying on a large amount of manually labeled data. In addition, the DINOv2 model can capture the global context information in images, is suitable for processing complex image scenes, and can effectively understand the global semantic structure of images. The features learned by the DINOv2 model have good robustness and adaptability when facing images in different environments and conditions, and can handle changes such as illumination changes and perspective changes.
[0048] This application can also adopt Training Method 1 or Training Method 2 to train the constructed cross-view feature extraction model.
[0049] In a preferred embodiment, Training Method 1 is adopted for training. Training Method 1 is as follows: In step A3, during the process of iteratively training the constructed cross-view feature extraction model batch by batch using the training sample pair set, each batch includes only the first training stage, and the first training stage of each batch includes:
[0050] Step A311: Input the subset of the training sample pairs of the current batch into the cross-view feature extraction model; wherein, the subset of the training sample pairs of the current batch includes more than one positive sample pair and more than one negative sample pair. After each sample pair is processed by the cross-view feature extraction model, the street photography feature embedding 35, street photography visual tokens, and street photography semantic tokens of the street photography image in the sample pair are obtained, and the aerial photography feature embedding 25, aerial photography visual tokens, and aerial photography semantic tokens of the aerial photography image in the sample pair are obtained.
[0051] Step A312: Calculate the batch triplet loss using the street photography feature embeddings 35 and aerial photography feature embeddings 25 of multiple sample pairs .
[0052] Step A313: Calculate the batch semantic alignment loss using the street photography visual tokens and street photography semantic tokens of the street photography images in the subset of the training sample pairs of the current batch output by the cross-view feature extraction model, and the aerial photography visual tokens and aerial photography semantic tokens of the aerial photography images; wherein, the street photography image is input into the street photography visual branch 30 to obtain the street photography visual tokens, the street photography image is input into the street photography semantic branch 31 to obtain the street photography semantic tokens, the aerial photography image is input into the aerial photography visual branch 20 to obtain the aerial photography visual tokens, and the aerial photography image is input into the aerial photography semantic branch 21 to obtain the aerial photography semantic tokens.
[0053] Step A314: Calculate the batch total loss based on the batch triplet loss and the batch semantic alignment loss.
[0054] Step A315: Update the network parameters of the cross-view feature extraction model based on the batch total loss. Specifically, in step A315, first determine whether the training stop condition is reached. If the iterative training stop condition is reached, output the network parameters of the current cross-view feature extraction model, that is, obtain the trained cross-view feature extraction model. If the training stop condition is not reached, after updating the network parameters of the cross-view feature extraction model based on the batch total loss, enter the next batch. Preferably but not limited to, update the network parameters of the cross-view feature extraction model by the gradient descent method based on the batch total loss. The network parameters of the cross-view feature extraction model include the network parameters of the first linear projection unit 301, the second linear projection unit 201, the first MLP unit 312, the second MLP unit 212, the street view Transformer encoder 33, the aerial view Transformer encoder 23, the street view embedding feature generation unit 34, and the aerial view embedding feature generation unit 24.
[0055] In this embodiment, in training method one, the batch triplet loss and the batch semantic alignment loss are simultaneously incorporated into the batch total loss. The batch semantic alignment loss enhances the consistency between visual features and semantic features, alleviates the feature offset problem in cross-view matching. The batch triplet loss optimizes the global matching of ground and aerial images, while the batch semantic alignment loss ensures the collaborative consistency of visual and semantic features, improving the robustness and generalization ability of the model in complex scenarios.
[0056] In this embodiment, to facilitate adjusting the fusion degree of the batch triplet loss and the batch semantic alignment loss, preferably, calculate the batch total loss according to the following formula:
[0057]
[0058] where, represents the batch triplet loss, represents the batch semantic alignment loss; represents the adjustment coefficient, The value range of is preferably but not limited to between 0.01 and 10, and the Bayesian optimization method can be used to tune it.
[0059] In this embodiment, to accurately calculate the batch semantic alignment loss, further preferably, the street view image and the aerial view image of each sample pair are both divided into blocks and then input into the cross-view feature extraction model. Suppose a street view image is divided into image blocks, and an aerial view image is divided into image blocks, and are all positive integers. The calculation method of the batch semantic alignment loss includes:
[0060] Calculate the two-norm of the difference between the street photography visual token and the street photography semantic token of each street photography image block in the current batch, denoted as the difference two-norm of the street photography image block. For the th sample pair in the current batch, the difference two-norm of the th image block of the street photography image is: . represents the visual token of the th image block of the street photography image in the th sample pair, represents the semantic token of the th image block of the street photography image in the th sample pair, . is the two-norm operator.
[0061] Calculate the two-norm of the difference between the aerial photography visual token and the aerial photography semantic token of each aerial photography image block in the current batch, denoted as the difference two-norm of the aerial photography image block; for the th sample pair in the current batch, the difference two-norm of the th image block of the aerial photography image is: . represents the visual token of the th image block of the aerial photography image in the th sample pair, represents the semantic token of the th image block of the aerial photography image in the th sample pair, .
[0062] Obtain the average value of the difference two-norms of all street photography image blocks in the current batch, denoted as the first average value; the first average value , represents the number of sample pairs in the subset of training sample pairs in the current batch, is the sample pair index, and are all positive integers, .
[0063] Obtain the average value of the difference two-norms of all aerial photography image blocks in the current batch, denoted as the second average value. The second average value .
[0064] Obtain the average value of the first average value and the second average value to obtain the batch semantic alignment loss , .
[0065] The batch semantic alignment loss can effectively utilize semantic information to enhance the ability of cross-view feature modeling, thereby improving the robustness of the model in complex scenarios.
[0066] In a preferred embodiment, the cross-view feature extraction model further includes a cropping module, which is trained using training method 2. Training method 2 is as follows: In step A3, during the iterative training of the constructed cross-view feature extraction model in batches using the training sample pair set, each batch includes a first training stage and a second training stage, specifically including:
[0067] The first training stage includes:
[0068] Step A3110, input the subset of training sample pairs of the current batch into the cross-view feature extraction model; wherein, the subset of training sample pairs of the current batch includes more than one positive sample pair and more than one negative sample pair. After each sample pair is processed by the cross-view feature extraction model, the street photography feature embedding 35, street photography visual tokens, and street photography semantic tokens of the street photography image in the sample pair are obtained, the aerial photography feature embedding 25, aerial photography visual tokens, and aerial photography semantic tokens of the aerial photography image in the sample pair are obtained. And for each aerial photography image in each sample pair, an attention map output by the last multi-head attention unit of the aerial photography Transformer encoder 23 is also obtained.
[0069] Step A3120, calculate the batch triplet loss using the street photography feature embeddings 35 and aerial photography feature embeddings 25 of multiple sample pairs .
[0070] Step A3130, calculate the batch semantic alignment loss using the street photography visual tokens and street photography semantic tokens of the street photography images in the subset of training sample pairs of the current batch, and the aerial photography visual tokens and aerial photography semantic tokens of the aerial photography images.
[0071] Step A3140, calculate the total batch loss based on the batch triplet loss and the batch semantic alignment loss.
[0072] Step A3150, update the network parameters of the cross-view feature extraction model based on the total batch loss.
[0073] After the first training stage is completed, enter the second training stage. The second training stage of each batch includes:
[0074] Step A321, obtain the attention map extracted by the aerial photography Transformer encoder 23 for each aerial photography image in the subset of training sample pairs of the current batch during the first training stage.
[0075] Step A322: The cropping module crops each aerial image according to the attention map of each aerial image to obtain a cropped aerial image. The cropped aerial image and the street view image in the sample pair to which each aerial image belongs form a new sample pair.
[0076] Step A323: Input all the new sample pairs into the cross-view feature extraction model obtained in the first training stage to obtain new street view feature embeddings 35 and new aerial image feature embeddings 25 of multiple new sample pairs, the new street view visual tokens and new street view semantic tokens of the street view images, and the new aerial image visual tokens and new aerial image semantic tokens of the cropped aerial images.
[0077] Step A324: Calculate the new batch triplet loss and the new batch semantic alignment loss, and calculate the new batch total loss based on the new batch triplet loss and the new batch semantic alignment loss. The calculation processes of the new batch triplet loss, the new batch semantic alignment loss, and the new batch total loss respectively refer to the calculation processes of the batch triplet loss, the batch semantic alignment loss, and the batch total loss in the first training stage, which will not be elaborated here.
[0078] Step A325: Update the network parameters of the cross-view feature extraction model again based on the new batch total loss.
[0079] Specifically, in Step A325, first determine whether the training stop condition is reached. If the iterative training stop condition is reached, output the network parameters of the current cross-view feature extraction model, that is, obtain the trained cross-view feature extraction model. If the training stop condition is not reached, after updating the network parameters of the cross-view feature extraction model based on the new batch total loss, enter the next batch. The specific method of updating the network parameters can refer to the implementation manner including only the first training stage, which will not be elaborated here.
[0080] In this embodiment, in Step A322, the step of the cropping module cropping each aerial image according to the attention map of each aerial image to obtain a cropped aerial image refers to Figure 5 , and specifically includes:
[0081] Step a: Perform magnification and binarization processing on the attention map of the aerial image. First, perform magnification processing on the attention map to make its size the same as that of the aerial image. Specifically, the resolution of the attention map can be increased (for example, increased by times) to increase the details of the key area, thereby increasing the number of patches ( times). When the image resolution is increased by times, the number of patches increases by times. Denote the magnification factor. Then, perform binarization on the magnified attention map. According to the pre-determined proportion β of the patches to be retained after cropping, for example, β is 64%, that is, 64% of the area is retained. Combine the numerical range and numerical distribution of the attention scores of the image regions in the magnified attention map to determine the attention score threshold. The higher the attention score, the greater the contribution of the image region to the final output result. For example, for aerial images, the attention map will show which regions (such as streets) are more important for geolocation, while regions that are not visible from another perspective (such as the rooftops of high-rise buildings) contribute less. Therefore, perform binarization according to the comparison result between the attention score of the image region in the attention map and the attention score threshold. For example, set the pixel value of the image region with an attention score greater than the attention score threshold to "1", and conversely, set the pixel value to "0", so as to obtain the binarized attention map. "1" represents the first pixel value, and "0" represents the second pixel value.
[0082] Step b, correspond the binarized attention map to the pixel positions of the aerial image, and crop the image regions with pixel value "0" in the aerial image, and retain the image regions with pixel value "1", so as to crop the regions with lower attention scores and retain the regions that make important contributions to the geolocation task, and obtain the cropped aerial image.
[0083] In this embodiment, by setting and β, the number of patches retained after the final cropping of the aerial image is . Select β × = 1, the amount of calculation remains unchanged, the resolution is improved, and the performance is improved. By adjusting these parameters, the model can reduce the amount of calculation.
[0084] Use = 1, the resolution is not increased, but due to cropping, the amount of calculation is reduced.
[0085] In the cross-view geolocation task, only some regions are commonly visible between the two perspectives. For example, many regions in the aerial image (such as the rooftops of high-rise buildings) are not visible from the street view perspective and contribute little to the positioning task. By cropping these unimportant regions, the model can process the parts that are important for matching (such as the street area). Retain the key regions, that is, the regions with high attention scores, which are also the regions that make important contributions to the geolocation task (such as streets).
[0086] This embodiment uses training method two for training, which includes a first training stage and a second training stage. In the second training stage, an image cropping mechanism is adopted. Specifically, by using the attention map generated by the Transformer encoder, regions with low contribution to geolocation in the aerial image are removed, and regions with high contribution are retained. Then, the detail information is enhanced through resolution improvement to generate a high-resolution image, reducing the computational burden while improving the performance.
[0087] The present invention also discloses a cross-view geolocation method based on the fusion of visual and semantic features. In a preferred embodiment, please refer to Figure 6 , this geolocation method includes:
[0088] Step B1, obtaining the street view image to be located. Specifically, it is not limited to the street view images taken by users on the ground through mobile terminals, cameras or robots.
[0089] Step B2, forming multiple image pairs by respectively combining the street view image to be located with multiple aerial images in the pre-constructed aerial image set. Each image pair includes the street view image to be located and one aerial image in the aerial image set. Each aerial image in the aerial image set is associated with an accurate geographical location information.
[0090] Step B3, inputting each image pair into the trained cross-view feature extraction model to obtain a pair of street view feature embeddings 35 and aerial view feature embeddings 25. The trained cross-view feature extraction model is obtained by training according to the above-mentioned cross-view feature extraction model training method provided by the present invention. The street view image of each image pair is input into the street view feature extraction module of the cross-view feature extraction model to obtain the street view feature embedding 35, and the aerial image of each image pair is input into the aerial view feature extraction module of the cross-view feature extraction model to obtain the aerial view feature embedding 25.
[0091] Step B4, calculating the distance between the street view feature embedding 35 and the aerial view feature embedding 25 in each image pair. Preferably but not limited to the Euclidean distance, and selecting the image pair with the smallest distance between the street view feature embedding 35 and the aerial view feature embedding 25 as the matching image pair;
[0092] Step B5, using the geographical location information associated with the aerial image in the matching image pair as the geographical location information of the street view image to be located, and outputting the obtained geographical location information.
[0093] The present invention also discloses a cross-view geolocation device based on the fusion of visual and semantic features, which is used to implement the above-mentioned cross-view geolocation method based on the fusion of visual and semantic features provided by the present invention. In a preferred embodiment, please refer to Figure 7 , this device includes:
[0094] An image input module, which obtains the street view image to be located;
[0095] An image pair acquisition module that forms multiple image pairs by respectively combining the street view images to be located with multiple aerial images in a pre-constructed aerial image set;
[0096] A feature embedding acquisition module that inputs each image pair into a trained cross-view feature extraction model to obtain a pair of street view feature embeddings 35 and aerial image feature embeddings 25, and the trained cross-view feature extraction model is obtained by training according to the above-mentioned cross-view feature extraction model training method provided by the present invention;
[0097] A matching image pair selection module that calculates the distance between the street view feature embedding 35 and the aerial image feature embedding 25 in each image pair, and selects the image pair with the smallest distance between the street view feature embedding 35 and the aerial image feature embedding 25 as the matching image pair;
[0098] An output module that uses the geographical location information associated with the aerial image in the matching image pair as the geographical location information of the street view image to be located, and outputs the obtained geographical location information.
[0099] In this embodiment, the image input module, the image pair acquisition module, the feature embedding acquisition module, the matching image pair selection module, and the output module respectively correspond to step B1, step B2, step B3, step B4, and step B5 in the above-mentioned cross-view geographical location method based on the fusion of visual and semantic features, and will not be elaborated here.
[0100] In a preferred embodiment, the above-mentioned cross-view geographical location device further includes an aerial image storage module for storing the aerial image set and the geographical location information associated with each aerial image in the aerial image set.
[0101] The present invention also discloses a computer program product, including a computer program, which when executed by a processor implements the steps of the above-mentioned cross-view feature extraction model training method or the cross-view geographical location method based on the fusion of visual and semantic features provided by the present invention. The computer program product should be understood as a software product that mainly implements its solution through a computer program, such as a program product integrated in the cloud or a software library.
[0102] The present invention also discloses an electronic device. In one embodiment, the electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the cross-view feature extraction model training method or the cross-view geographical location method based on the fusion of visual and semantic features provided by the present invention.
[0103] Such as Figure 8As shown in the figure, it is a schematic structural diagram of an electronic device for the cross-view feature extraction model training method or the cross-view geolocation method based on the fusion of visual and semantic features provided by an embodiment of the present invention. The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a cross-view feature extraction model training method or a cross-view geolocation method program based on the fusion of visual and semantic features.
[0104] Among them, the processor 10 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as executing the cross-view feature extraction model training method or the cross-view geolocation method based on the fusion of visual and semantic features, etc.), and calling the data stored in the memory 11, to execute various functions of the electronic device and process data.
[0105] The memory 11 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical discs, etc. The memory 11 may be an internal storage unit of the electronic device in some embodiments, such as the mobile hard disk of the electronic device. The memory 11 may also be an external storage device of the electronic device in other embodiments, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may include both an internal storage unit and an external storage device of the electronic device. The memory 11 can not only be used to store application software installed on the electronic device and various types of data, such as the code of the cross-view feature extraction model training method or the cross-view geolocation method program based on the fusion of visual and semantic features, but also be used to temporarily store data that has been output or will be output.
[0106] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to implement connection communication between the memory 11 and at least one processor 10, etc.
[0107] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.
[0108] Figure 8 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 8 the shown structure does not constitute a limitation on the electronic device, and it can include fewer or more components than shown, or combine certain components, or have a different component arrangement.
[0109] For example, although not shown, the electronic device can also include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source can also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device can also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0110] It should be understood that the embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0111] Furthermore, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).
[0112] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one implementation manner", "one preferred implementation manner" or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0113] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A cross-view feature extraction model training method, characterized in that: include: Construct a training sample pair set including one or more positive sample pairs and one or more negative sample pairs, each sample pair including one street photography image and one aerial photography image; Build a cross-view feature extraction model, which includes: A street photography feature extraction module, used for obtaining street photography feature embedding of a street photography image, wherein the street photography feature extraction module comprises a street photography visual branch and a street photography semantic branch, and a street photography token fusion unit, a street photography Transformer encoder and a street photography embedding feature generation unit connected in sequence, wherein an input end of the street photography token fusion unit is respectively connected to an output end of the street photography visual branch and an output end of the street photography semantic branch; An aerial feature extraction module is used to obtain aerial feature embedding of an aerial image. The aerial feature extraction module includes an aerial vision branch and an aerial semantic branch, and an aerial token fusion unit, an aerial Transformer encoder, and an aerial embedding feature generation unit connected in sequence. The input end of the aerial token fusion unit is respectively connected to the output end of the aerial vision branch and the output end of the aerial semantic branch. The constructed cross-view feature extraction model is iteratively trained in batches using the training sample pair set. When the training stop condition is reached, the cross-view feature extraction model is obtained. The iterative training steps include: Calculate the total batch loss based on batch triplet loss and batch semantic alignment loss; Update the network parameters of the cross-view feature extraction model based on the total batch loss; The street-level images and aerial images of each sample pair are processed in blocks and then input into the cross-view feature extraction model. The batch semantic alignment loss is calculated by: Calculate the binary norm of the difference between the street photography visual token and the street photography semantic token of each street photography image block in each street photography image in the current batch, and record it as the binary norm of the difference of the street photography image block; Calculate the binary norm of the difference between the aerial visual token and the aerial semantic token of each aerial image block in each aerial image in the current batch, and record it as the binary norm of the difference of the aerial image block; The average value of the difference two norms of all street photography image blocks in the current batch is calculated, which is recorded as the first average value. The average value of the difference two norms of all aerial photography image blocks in the current batch is calculated, which is recorded as the second average value. The average of the first average value and the second average value is calculated to obtain the batch semantic alignment loss.
2. A cross-view feature extraction model training method as claimed in claim 1, characterized in that: The street photography visual branch includes a first linear projection unit and a first position embedding unit connected in sequence, and the street photography semantic branch includes a first semantic feature extraction unit and a first MLP unit connected in sequence; and / or, The aerial photography vision branch includes a second linear projection unit and a second position embedding unit connected in sequence, and the aerial photography semantic branch includes a second semantic feature extraction unit and a second MLP unit connected in sequence.
3. A cross-view feature extraction model training method as claimed in claim 1 or 2, characterized in that: In the step of iteratively training the constructed cross-view feature extraction model in batches using the training sample pair set, each batch includes a first training stage, and the first training stage of each batch includes: Inputting the subset of training sample pairs of the current batch into the cross-view feature extraction model; wherein the subset of training sample pairs of the current batch includes one or more positive sample pairs and one or more negative sample pairs; The batch triplet loss is calculated using the street photography feature embeddings and aerial photography feature embeddings of multiple sample pairs in the current batch of training sample pair subsets output by the cross-view feature extraction model; The batch semantic alignment loss is calculated using the street visual tokens and street semantic tokens of the street images in the subset, as well as the aerial visual tokens and aerial semantic tokens of the aerial images, using the training samples of the current batch output by the cross-view feature extraction model; wherein, the street images are input into the street vision branch to obtain street visual tokens, the street images are input into the street semantic branch to obtain street semantic tokens, the aerial images are input into the aerial vision branch to obtain aerial visual tokens, and the aerial images are input into the aerial semantic branch to obtain aerial semantic tokens.
4. A cross-view feature extraction model training method as claimed in claim 1, characterized in that: The total batch loss is calculated according to the following formula: , in, represents the batch triplet loss, represents the batch semantic alignment loss, Represents the adjustment factor.
5. A cross-view feature extraction model training method as claimed in claim 3, characterized in that: The cross-view feature extraction model also includes a cropping module; Each batch also includes a second training phase, which includes: Get the attention map extracted by the aerial Transformer encoder during the first training phase for each aerial image in the current batch of training sample pairs; The cropping module crops each aerial image according to the attention map of each aerial image to obtain a cropped aerial image, and the cropped aerial image and the street image in the sample pair to which each aerial image belongs form a new sample pair; Input all new sample pairs into the cross-view feature extraction model obtained in the first training stage to obtain new street feature embeddings and new aerial feature embeddings for multiple new sample pairs, new street visual tokens and new street semantic tokens for street images, and new aerial visual tokens and new aerial semantic tokens for cropped aerial images; Calculate the triplet loss of the new batch and the semantic alignment loss of the new batch, and calculate the total loss of the new batch based on the triplet loss of the new batch and the semantic alignment loss of the new batch; The network parameters of the cross-view feature extraction model are updated again based on the total loss of the new batch.
6. A cross-view geolocation method based on the fusion of visual and semantic features, characterized in that: include: Obtain the street photography image to be located; The street-photographed image to be located is respectively combined with a plurality of aerial images in a pre-constructed aerial image set to form a plurality of image pairs; Input each image pair into a trained cross-view feature extraction model to obtain a pair of street photography feature embedding and aerial photography feature embedding, wherein the trained cross-view feature extraction model is trained according to the method of any one of claims 1 to 5; Calculate the distance between the street photography feature embedding and the aerial photography feature embedding in each image pair, and select the image pair with the smallest distance between the street photography feature embedding and the aerial photography feature embedding as the matching image pair; The geographical location information associated with the aerial image in the matching image pair is used as the geographical location information of the street-shot image to be located, and the obtained geographical location information is output.
7. A cross-view geolocation device based on the fusion of visual and semantic features, used to implement the cross-view geolocation method based on the fusion of visual and semantic features as claimed in claim 6, characterized in that: include: Image input module, obtaining the street-shot image to be located; An image pair acquisition module combines the street-photographed image to be located with multiple aerial images in a pre-built aerial image set to form multiple image pairs; A feature embedding acquisition module, inputting each image pair into a trained cross-view feature extraction model to obtain a pair of street photography feature embedding and aerial photography feature embedding, wherein the trained cross-view feature extraction model is trained according to the method of any one of claims 1 to 5; The matching image pair selection module calculates the distance between the street photography feature embedding and the aerial photography feature embedding in each image pair, and selects the image pair with the smallest distance between the street photography feature embedding and the aerial photography feature embedding as the matching image pair; The output module uses the geographic location information associated with the aerial image in the matching image pair as the geographic location information of the street-shot image to be located, and outputs the obtained geographic location information.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.
9. An electronic device, characterized in that: The electronic device comprises: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor so that the at least one processor can perform the method as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Unsupervised cross-modal hash retrieval method based on CLIP and attention fusion mechanism
CN118861327A
Cross-view image geo-localization
US20240303770A1