Cross-view feature extraction model training method, geographic positioning method and device

By adopting a Transformer encoder model that integrates visual and semantic features in cross-view image geolocation, the problem of insufficient accuracy and robustness of cross-view matching in the prior art is solved, and a more efficient geolocation effect is achieved.

CN119992117AActive Publication Date: 2025-05-13SEVNCE ROBOTICS CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510455319.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

Existing cross-view image geolocation methods rely on visual features, making it difficult to achieve accurate matching and poor robustness, especially when there are differences in view angles, geometric structures and resolutions of street and aerial images.

Method used

A cross-view feature extraction model training method is adopted. By constructing a training sample pair set including positive and negative sample pairs, using a Transformer encoder model including visual branches and semantic branches, visual and semantic features of street and aerial images are extracted, feature embeddings are generated, and the model is optimized through iterative training.

Benefits of technology

By fusing visual and semantic features, the cross-view feature extraction model can capture higher-level information, such as geographical indications and road topology, significantly improving the performance of cross-view matching and improving the accuracy of geolocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992117A_ABST
    Figure CN119992117A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of geographic positioning, and provides a cross-view feature extraction model training method and device and a geographic positioning method and device.The training method comprises the steps that a training sample pair set is constructed; the constructed cross-view feature extraction model comprises a street beat feature extraction module which comprises a street beat visual branch and a street beat semantic branch, and a street beat token fusion unit, a street beat Transform encoder and a street beat embedded feature generation unit which are connected in sequence; the aerial photography feature extraction module comprises an aerial photography visual branch and an aerial photography semantic branch, and an aerial photography token fusion unit, an aerial photography Transform encoder and an aerial photography embedded feature generation unit which are connected in sequence; constructing a cross-view feature extraction model in batches; the invention further provides a cross-view geographic positioning method based on visual and semantic feature fusion, a cross-view geographic positioning device based on visual and semantic feature fusion, a computer program product and electronic equipment. According to the invention, the accuracy of cross-view geographic positioning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of geographic positioning technology, and in particular to a cross-view feature extraction model training method, a geographic positioning method and a device. Background Art

[0002] Cross-view image geolocation methods based on CNN usually rely on polar coordinate transformation to reduce the geometric differences between ground images and aerial images. However, this method requires prior knowledge of geometry and may lead to matching failure when street view images and aerial images are not aligned. In addition, since CNN convolutional neural networks cannot effectively model global correlation and explicit location information, there is a large domain gap between street view images and aerial images in cross-view retrieval systems. To address these shortcomings, methods based on Transformer encoders have emerged in recent years, which take advantage of their advantages in global information modeling and explicit location information encoding to improve the performance of cross-view matching.

[0003] The core goal of the cross-view image geolocation method is to match the same geographic location from two different perspectives: a ground street view image (i.e., street photography) and an aerial image. This technology has important application value in the fields of autonomous driving and navigation. Existing methods usually rely on visual features and extract features for matching through deep learning models. However, there are differences in the perspective, geometric structure, and resolution of street photography and aerial images. It is difficult to achieve accurate matching based on visual features alone, and the robustness is poor. Summary of the invention

[0004] The present application aims to at least solve the technical problems existing in the prior art and provide a cross-view feature extraction model training method, a geolocation method and a device.

[0005] In a first aspect, the present application provides a cross-view feature extraction model training method, comprising: constructing a training sample pair set including one or more positive sample pairs and one or more negative sample pairs, each sample pair including a street shot image and an aerial photo; constructing a cross-view feature extraction model, the cross-view feature extraction model including: a street shot feature extraction module, for obtaining street shot feature embedding of a street shot image, the street shot feature extraction module including a street shot visual branch and a street shot semantic branch, and a street shot token fusion unit, a street shot Transformer encoder and a street shot embedding feature generation unit connected in sequence, the input end of the street shot token fusion unit is divided into: The aerial photography feature extraction module is used to obtain the aerial feature embedding of the aerial image, and the aerial photography feature extraction module includes an aerial photography visual branch and an aerial photography semantic branch, as well as an aerial photography token fusion unit, an aerial photography Transformer encoder and an aerial photography embedding feature generation unit connected in sequence, and the input end of the aerial photography token fusion unit is respectively connected to the output end of the aerial photography visual branch and the output end of the aerial photography semantic branch; the constructed cross-view feature extraction model is iteratively trained in batches using a training sample set, and the cross-view feature extraction model is obtained when the training stop condition is reached.

[0006] In a second aspect, the present application provides a cross-view geolocation method based on the fusion of visual and semantic features, comprising: obtaining a street-photographed image to be located; combining the street-photographed image to be located with a plurality of aerial images in a pre-constructed aerial image set to form a plurality of image pairs; Input each image pair into a trained cross-view feature extraction model to obtain a pair of street photography feature embedding and aerial photography feature embedding, wherein the trained cross-view feature extraction model is trained according to the method provided in the first aspect of the present application; calculate the distance between the street photography feature embedding and the aerial photography feature embedding in each image pair, and select the image pair with the smallest distance between the street photography feature embedding and the aerial photography feature embedding as the matching image pair; use the geographic location information associated with the aerial photography image in the matching image pair as the geographic location information of the street photography image to be located, and output the obtained geographic location information.

[0007] In a third aspect, the present application provides a cross-view geolocation device based on the fusion of visual and semantic features, which is used to implement the cross-view geolocation method based on the fusion of visual and semantic features provided in the second aspect of the present application, including: an image input module, which obtains street-shot images to be located; an image pair acquisition module, which combines the street-shot images to be located with multiple aerial images in a pre-constructed aerial image set to form multiple image pairs; a feature embedding acquisition module, which inputs each image pair into a trained cross-view feature extraction model to obtain a pair of street-shot feature embedding and aerial feature embedding, and the trained cross-view feature extraction model is trained according to the method provided in the first aspect of the present application; a matching image pair selection module, which calculates the distance between the street-shot feature embedding and the aerial feature embedding in each image pair, and selects the image pair with the smallest distance between the street-shot feature embedding and the aerial feature embedding as the matching image pair; an output module, which uses the geographic location information associated with the aerial image in the matching image pair as the geographic location information of the street-shot image to be located, and outputs the acquired geographic location information.

[0008] In a fourth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method provided in the first aspect or the second aspect of the present application.

[0009] In a fifth aspect, the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the steps of the method provided in the first aspect or the second aspect of the present application.

[0010] The beneficial technical effect of the present application is as follows: corresponding visual branches and semantic branches are simultaneously provided in the street photography feature extraction module and the aerial photography feature extraction module of the cross-view feature extraction model, and street photography visual tokens containing the visual features of the street photography images and street photography semantic tokens containing the semantic features of the street photography images are simultaneously extracted, and after splicing the street photography visual tokens and the street photography semantic tokens, the spliced ​​tokens are encoded and mapped to generate street photography feature embedding, and the aerial photography visual tokens containing the visual features of the aerial photography images and the aerial photography semantic tokens containing the semantic features of the aerial photography images are simultaneously extracted. After concatenating aerial visual tokens and aerial semantic tokens, the concatenated tokens are encoded and mapped to generate aerial feature embeddings, so that the obtained street photography feature embeddings and aerial photography feature embeddings both contain semantic features, which supplements the expression ability of visual features and enables the cross-view feature extraction model to capture higher-level information, such as geographical landmarks, road topology, etc., significantly improving the performance of cross-view matching, and enabling the cross-view feature extraction model to better adapt to the perspective differences and geometric structure differences between ground images and aerial images, thereby improving the accuracy of cross-view geographic positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flowchart of a cross-view feature extraction model training method in a preferred embodiment of the present invention; Figure 2 is a street-photographed image in a positive pair in an example; Figure 3 yes Figure 2 Aerial images in the positive sample pairs in the example; Figure 4 It is a schematic diagram of the structure of a cross-view feature extraction model in a preferred embodiment of the present invention; Figure 5 This is a schematic diagram of an aerial image cropping process in a preferred embodiment of the present invention; Figure 6 It is a flow chart of a cross-view geolocation method based on the fusion of visual and semantic features in a preferred embodiment of the present invention; Figure 7 It is a system block diagram of a cross-view geographic positioning device based on the fusion of visual and semantic features in a preferred embodiment of the present invention; Figure 8 is a schematic structural diagram of an electronic device in a preferred embodiment of the present invention; Reference numerals: 10 processor;11 memory;12 communication bus;13 communication interface;30 street photography vision branch;301 first linear projection unit;302 first position embedding unit;31 street photography semantic branch;311 first semantic feature extraction unit;312 first MLP unit;32 street photography token fusion unit;33 street photography Transformer encoder;34 street photography embedding feature generation unit;35 street photography feature embedding;20 aerial photography vision branch;201 second linear projection unit;202 second position embedding unit;21 aerial photography semantic branch;211 second semantic feature extraction unit;212 second MLP unit;22 aerial photography token fusion unit;23 aerial photography Transformer encoder;24 aerial photography embedding feature generation unit;25 aerial photography feature embedding. DETAILED DESCRIPTION

[0012] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0013] In the description of the present invention, it is necessary to understand that the terms "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0014] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal connection of two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.

[0015] The execution subject of the cross-view feature extraction model training method or the cross-view geolocation method based on the fusion of visual and semantic features provided by the present invention includes but is not limited to at least one of the electronic devices such as the server and the terminal that can be configured to execute the method provided in the embodiment of the present application. In other words, the cross-view feature extraction model training method or the cross-view geolocation method based on the fusion of visual and semantic features can be executed by software or hardware installed on the terminal device or the server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.

[0016] The present invention provides a cross-view feature extraction model training method. In a preferred embodiment, Figure 1 As shown, including: Step A1: construct a training sample pair set including one or more positive sample pairs and one or more negative sample pairs, each sample pair including one street-photographed image and one aerial-photographed image.

[0017] In this embodiment, the positive sample pair consists of a street image and an aerial image from the same geographical location, and the negative sample pair consists of a street image and an aerial image from different geographical locations. Figure 2 and Figure 3 An example of a positive pair is shown, where Figure 2 Shows street photography images from a positive pair example. Figure 3 The aerial images in the positive sample pair example are shown. It can be intuitively observed that the street photography images are images taken by an imaging device on the ground of a person at a certain geographical location, and the aerial images are images taken from a high altitude at a certain geographical location. Preferably, in order to more clearly show the detailed information of the geographical location, the aerial image can be a bird's-eye view.

[0018] In this embodiment, the method for constructing a training sample pair set is: obtain an existing aerial image set, each aerial image in the aerial image set has geographical location information of the shooting point, and a street image can be taken at the geographical location corresponding to each aerial image, so as to obtain a set of street image sets that can correspond one-to-one with the aerial images in the aerial image set. Based on the aerial image set and the corresponding street image set, the aerial images and street images in the same geographical location are combined into positive sample pairs, and the aerial images and street images in different geographical locations are combined into negative sample pairs. According to the above method, multiple positive sample pairs and multiple negative sample pairs are generated, and the multiple positive sample pairs and multiple negative sample pairs constitute a training sample pair set. Similarly, a verification sample pair set can be constructed according to the above method.

[0019] Step A2: construct a cross-view feature extraction model.

[0020] Please refer to the attached Figure 4 , the cross-view feature extraction model includes: A street photography feature extraction module is used to obtain street photography feature embedding 35 of a street photography image. The street photography feature extraction module includes a street photography visual branch 30 and a street photography semantic branch 31, as well as a street photography token fusion unit 32, a street photography Transformer encoder 33 and a street photography embedding feature generation unit 34 connected in sequence. The input end of the street photography token fusion unit 32 is respectively connected to the output end of the street photography visual branch 30 and the output end of the street photography semantic branch 31.

[0021] The aerial photography feature extraction module is used to obtain the aerial photography feature embedding 25 of the aerial photography image. The aerial photography feature extraction module includes an aerial photography vision branch 20 and an aerial photography semantic branch 21, as well as an aerial photography token fusion unit 22, an aerial photography Transformer encoder 23 and an aerial photography embedding feature generation unit 24 connected in sequence. The input end of the aerial photography token fusion unit 22 is respectively connected to the output end of the aerial photography vision branch 20 and the output end of the aerial photography semantic branch 21.

[0022] Exemplarily, taking a sample as input, the following describes the process of extracting street photography feature embedding 35 of the street photography image in the sample pair and extracting aerial photography feature embedding 25 of the aerial photography image by the cross-view feature extraction model. It includes: The street photography images and aerial photography images of the sample pairs are divided into blocks, and each image block is flattened.

[0023] The image block of the street photography image is input into the street photography feature extraction module, the street photography visual branch 30 linearly projects the flattened image block to a fixed dimension to obtain a visual token of the street photography image, adds a learnable position embedding to the visual token to retain the spatial information of the street photography image, and takes all the visual tokens of the street photography image that incorporate the position information as street photography visual tokens; the street photography semantic branch 31 extracts the semantic features of the street photography image, and flattens the semantic features and projects them to the same feature dimension as the street photography visual tokens to obtain street photography semantic tokens; the street photography token fusion unit 32 is used to splice the street photography visual tokens and the street photography semantic tokens to obtain street photography spliced ​​tokens after splicing; the street photography Transformer encoder 33 encodes the street photography spliced ​​tokens to obtain street photography encoding features, and the street photography embedding feature generation unit 34 is used to map the street photography encoding features to obtain street photography feature embedding 35, and the street photography embedding feature generation unit 34 is preferably but not limited to an existing MLP Head unit or one or more cascaded fully connected layers.

[0024] The image block of the aerial image is input into the aerial feature extraction module, the aerial vision branch 20 linearly projects the flattened image block to a fixed dimension to obtain a visual token of the aerial image, adds a learnable position embedding to the visual token to retain the spatial information of the aerial image, and takes all the visual tokens of the aerial image that are integrated with the position information as aerial visual tokens; the aerial semantic branch 21 extracts the semantic features of the aerial image, and flattens the semantic features and projects them to the same feature dimension as the aerial visual tokens to obtain aerial semantic tokens; the aerial token fusion unit 22 is used to splice the aerial visual tokens and the aerial semantic tokens to obtain aerial splicing tokens after splicing; the aerial Transformer encoder 23 encodes the aerial splicing tokens to obtain aerial coding features, and the aerial embedding feature generation unit 24 is used to map the aerial coding features to obtain aerial feature embedding 25, and the aerial embedding feature generation unit 24 is preferably but not limited to an existing MLP Head unit or one or more cascaded fully connected layers.

[0025] In step A3, the constructed cross-view feature extraction model is iteratively trained in batches using the training sample pair set, and a cross-view feature extraction model is obtained when the training stop condition is reached.

[0026] In this embodiment, the training stop condition is preferably but not limited to the number of training times reaching the preset maximum number of training times or the total batch loss convergence. The total batch loss is the loss calculated after each batch is completed, and the total batch loss can be the batch triplet loss. Preferably, after completing step A3, step A4 is also included, using a pre-constructed verification sample pair set to verify the cross-view feature extraction model obtained in step A3. If the verification passes, the cross-view feature extraction model obtained in step A3 is output. If the verification fails, the training parameters, such as the learning rate, are adjusted, and then step A3 is returned to execute.

[0027] In this embodiment, preferably, the batch triplet loss is calculated using the street photography feature embedding 35 and the aerial photography feature embedding 25 of multiple sample pairs in the current batch of training sample pair subsets output by the cross-view feature extraction model. , for example, the batch triplet loss is calculated using the street photography feature embedding 35 and the aerial photography feature embedding 25 of all sample pairs included in each batch : in, Represents the first hyperparameter, which is used to control the marginal range of batch triple loss. The default value can be 0.5; It represents the average of the square of the distance or L2 distance (i.e., Euclidean distance) between the street photography feature embedding 35 and the aerial photography feature embedding 25 obtained by all positive sample pairs in the current batch of training sample pair subsets. , Indicates the number of positive sample pairs in the current batch of training sample pair subsets, which is a positive integer. Represents the index of the positive sample pair in the subset of training sample pairs in the current batch, Represents the positive sample pairs in the current batch of training sample pairs The distance or the square of the L2 distance between the obtained street photography feature embedding 35 and the aerial photography feature embedding 25; It represents the average of the square of the L2 distance between the street photography feature embedding 35 and the aerial photography feature embedding 25 in the negative sample pairs in the current batch of training sample pairs. , Represents the number of negative sample pairs in the training sample pair subset of the current training batch, which is a positive integer. Represents the index of the negative sample pair in the subset of training sample pairs in the current batch, Represents the negative sample pairs in the current batch of training sample pairs The distance or the square of the L2 distance between the obtained street photography feature embedding 35 and the aerial photography feature embedding 25.

[0028] In this embodiment, batch triplet loss is used to increase the difference between the street photography feature embedding 35 and the aerial photography feature embedding 25 of the negative sample pair, while the difference between the street photography feature embedding 35 and the aerial photography feature embedding 25 of the positive sample pair is reduced, thereby optimizing the global matching of street photography images and aerial photography images.

[0029] In this embodiment, preferably, the street photography visual branch 30 includes a first linear projection unit 301 and a first position embedding unit 302 connected in sequence, and the street photography semantic branch 31 includes a first semantic feature extraction unit 311 and a first MLP unit 312 connected in sequence; or, the aerial photography visual branch 20 includes a second linear projection unit 201 and a second position embedding unit 202 connected in sequence, and the aerial photography semantic branch 21 includes a second semantic feature extraction unit 211 and a second MLP unit 212 connected in sequence.

[0030] In this embodiment, preferably, referring to Figure 4 As shown, the street photography visual branch 30 includes a first linear projection unit 301 and a first position embedding unit 302 connected in sequence, and the street photography semantic branch 31 includes a first semantic feature extraction unit 311 and a first MLP unit 312 connected in sequence; and the aerial photography visual branch 20 includes a second linear projection unit 201 and a second position embedding unit 202 connected in sequence, and the aerial photography semantic branch 21 includes a second semantic feature extraction unit 211 and a second MLP unit 212 connected in sequence.

[0031] In this embodiment, the first linear projection unit 301 and the second linear projection unit 201 are preferably, but not limited to, fully connected layers cascaded at more than one layer. The first linear projection unit 301 is used to linearly project the image block after flattening the street image to a fixed dimension to obtain a visual token of the street image. The first position embedding unit 302 is used to add a learnable position embedding to the visual token of the street image to retain the spatial information of the street image. The second linear projection unit 201 is used to linearly project the image block after flattening the aerial image to a fixed dimension to obtain a visual token of the aerial image. The second position embedding unit 202 is used to add a learnable position embedding to the visual token of the aerial image to retain the spatial information of the aerial image.

[0032] In this embodiment, the first semantic feature extraction unit 311 and the second semantic feature extraction unit 211 are not limited to the pre-trained DINOv2 model or MAE (Masked AutoEncoders) model. Preferably, the DINOv2 model is selected because the DINOv2 model can learn high-quality feature representations from unlabeled data. It can extract deep semantic features of images through self-supervised learning methods without relying on a large amount of manually annotated data. In addition, the DINOv2 model can capture global contextual information in the image, is suitable for processing complex image scenes, and can effectively understand the global semantic structure of the image. The features learned by the DINOv2 model have good robustness and adaptability when facing images in different environments and conditions, and can handle changes such as changes in illumination and perspective.

[0033] The present application may also use training method one or training method two to train the constructed cross-view feature extraction model.

[0034] In a preferred embodiment, training is performed using training method 1, which is: in step A3, in the process of iteratively training the constructed cross-view feature extraction model in batches using the training sample pair set, each batch includes only the first training stage, and the first training stage of each batch includes: Step A311, input the current batch of training sample pair subsets into the cross-view feature extraction model; wherein the current batch of training sample pair subsets includes one or more positive sample pairs and one or more negative sample pairs. After each sample pair is processed by the cross-view feature extraction model, the street photography feature embedding 35, street photography visual tokens and street photography semantic tokens of the street photography image in the sample pair are obtained, and the aerial photography feature embedding 25, aerial photography visual tokens and aerial photography semantic tokens of the aerial photography image in the sample pair are obtained.

[0035] Step A312, using the street photography feature embedding 35 and the aerial photography feature embedding 25 of multiple sample pairs to calculate the batch triplet loss .

[0036] Step A313, using the current batch of training samples output by the cross-view feature extraction model to calculate the batch semantic alignment loss for the street visual tokens and street semantic tokens of the street images in the subset, as well as the aerial visual tokens and aerial semantic tokens of the aerial images; wherein, the street images are input into the street visual branch 30 to obtain street visual tokens, the street images are input into the street semantic branch 31 to obtain street semantic tokens, the aerial images are input into the aerial visual branch 20 to obtain aerial visual tokens, and the aerial images are input into the aerial semantic branch 21 to obtain aerial semantic tokens.

[0037] Step A314, calculate the batch total loss based on the batch triplet loss and the batch semantic alignment loss.

[0038] Step A315, update the network parameters of the cross-view feature extraction model based on the total batch loss. Specifically, in step A315, first determine whether the training stop condition is met. If the iterative training stop condition is met, then output the network parameters of the current cross-view feature extraction model, that is, obtain the trained cross-view feature extraction model. If the training stop condition is not met, then update the network parameters of the cross-view feature extraction model based on the total batch loss, and then enter the next batch. Preferably, but not limited to, update the network parameters of the cross-view feature extraction model by gradient descent based on the total batch loss. The network parameters of the cross-view feature extraction model include the network parameters of the first linear projection unit 301, the second linear projection unit 201, the first MLP unit 312, the second MLP unit 212, the street photography Transformer encoder 33, the aerial photography Transformer encoder 23, the street photography embedding feature generation unit 34, and the aerial photography embedding feature generation unit 24.

[0039] In this implementation, batch triplet loss and batch semantic alignment loss are simultaneously incorporated into the batch total loss in training method one. The batch semantic alignment loss enhances the consistency between visual features and semantic features, and alleviates the problem of feature offset in cross-view matching. The batch triplet loss optimizes the global matching of ground and aerial images, while the batch semantic alignment loss ensures the coordinated consistency of visual and semantic features, and improves the robustness and generalization ability of the model in complex scenarios.

[0040] In this embodiment, in order to facilitate the adjustment of the fusion degree of the optimal batch triple loss and the batch semantic alignment loss, preferably, the total batch loss is calculated according to the following formula: in, represents the batch triplet loss, represents batch semantic alignment loss; represents the adjustment factor, The value range of is preferably but not limited to 0.01 to 10, and can be tuned using the Bayesian optimization method.

[0041] In this embodiment, in order to accurately calculate the batch semantic alignment loss, it is further preferred that the street photography image and the aerial photography image of each sample pair are divided into blocks and then input into the cross-view feature extraction model. Suppose a street photography image is divided into image blocks, an aerial image is divided into image blocks, and are all positive integers. Then the calculation method of batch semantic alignment loss includes: Calculate the binary norm of the difference between the street photography visual token and the street photography semantic token of each street photography image block in each street photography image in the current batch, and record it as the binary norm of the difference of the street photography image block. The first street image in the sample pair The difference bi-norm of the image blocks is: . Indicates The first street image in the sample pair The visual token of the image patch, Indicates The first street image in the sample pair The semantic token of the image block, . is the two-norm operator.

[0042] Calculate the binary norm of the difference between the aerial visual token and the aerial semantic token of each aerial image block in each aerial image in the current batch, recorded as the binary norm of the difference between the aerial image blocks; The first sample pair of aerial images The difference bi-norm of the image blocks is: . Indicates The first sample pair of aerial images The visual token of the image patch, Indicates The first sample pair of aerial images The semantic token of the image block, .

[0043] Calculate the average value of the difference two norms of all street photography image blocks in the current batch, which is recorded as the first average value; the first average value , Indicates the number of sample pairs in the training sample pair subset in the current batch, is the sample pair index, and are all positive integers, .

[0044] Calculate the average value of the difference two norms of all aerial image blocks in the current batch, and record it as the second average value. .

[0045] Take the average of the first and second averages to get the batch semantic alignment loss , .

[0046] The batch semantic alignment loss can effectively utilize semantic information to enhance the ability of cross-view feature modeling, thereby improving the robustness of the model in complex scenes.

[0047] In a preferred embodiment, the cross-view feature extraction model further includes a cropping module, and is trained using the second training method. The second training method is: in step A3, the constructed cross-view feature extraction model is iteratively trained in batches using the training sample set, and each batch includes a first training stage and a second training stage, specifically including: The first training phase includes: Step A3110, input the subset of training sample pairs of the current batch into the cross-view feature extraction model; wherein the subset of training sample pairs of the current batch includes one or more positive sample pairs and one or more negative sample pairs. After each sample pair is processed by the cross-view feature extraction model, the street photography feature embedding 35, street photography visual tokens and street photography semantic tokens of the street photography image in the sample pair are obtained, and the aerial photography feature embedding 25, aerial photography visual tokens and aerial photography semantic tokens of the aerial photography image in the sample pair are obtained. And the aerial image of each sample pair also obtains the attention map output by the last layer of multi-head attention unit of the aerial photography Transformer encoder 23.

[0048] Step A3120, using the street photography feature embedding 35 and the aerial photography feature embedding 25 of multiple sample pairs to calculate the batch triplet loss .

[0049] Step A3130, using the training samples of the current batch to calculate the batch semantic alignment loss for the street visual tokens and street semantic tokens of the street images in the subset, and the aerial visual tokens and aerial semantic tokens of the aerial images.

[0050] Step A3140, calculate the batch total loss based on the batch triplet loss and the batch semantic alignment loss.

[0051] Step A3150, updates the network parameters of the cross-view feature extraction model based on the total batch loss.

[0052] After the first training phase is completed, the second training phase begins. The second training phase for each batch includes: Step A321, obtain the attention map extracted by the aerial Transformer encoder 23 for each aerial image in the subset of training samples of the current batch during the first training stage.

[0053] In step A322, the cropping module crops each aerial image according to the attention map of each aerial image to obtain a cropped aerial image. The cropped aerial image and the street image in the sample pair to which each aerial image belongs form a new sample pair.

[0054] Step A323, input all new sample pairs into the cross-view feature extraction model obtained in the first training stage, obtain new street photography feature embeddings 35 and new aerial photography feature embeddings 25 for multiple new sample pairs, new street photography visual tokens and new street photography semantic tokens for street photography images, and new aerial visual tokens and new aerial semantic tokens for cropped aerial images.

[0055] Step A324, calculate the new batch triple loss and the new batch semantic alignment loss, and calculate the new batch total loss based on the new batch triple loss and the new batch semantic alignment loss. The new batch triple loss, the new batch semantic alignment loss and the new batch total loss are all calculated by referring to the calculation process of the batch triple loss, the batch semantic alignment loss and the batch total loss in the first training stage, respectively, and will not be repeated here.

[0056] Step A325, update the network parameters of the cross-view feature extraction model again based on the new batch total loss.

[0057] Specifically, in step A325, it is first determined whether the training stop condition is met. If the iterative training stop condition is met, the network parameters of the current cross-view feature extraction model are output, that is, the trained cross-view feature extraction model is obtained. If the training stop condition is not met, the network parameters of the cross-view feature extraction model are updated based on the total loss of the new batch, and then the next batch is entered. For the specific method of updating the network parameters, please refer to the implementation method that only includes the first training stage, which will not be repeated here.

[0058] In this embodiment, in step A322, the cropping module crops each aerial image according to the attention map of each aerial image to obtain the step of cropping the aerial image. Figure 5 , specifically including: Step a: enlarge and binarize the attention map of the aerial image. First, enlarge the attention map to make it the same size as the aerial image. Specifically, the resolution of the attention map can be increased (for example, by increasing times) to increase the details of key areas, thereby increasing the number of patches ( times). times the image resolution, the number of patches increases times. Represents the multiplication factor. After that, the attention map after the magnification processing is binarized. According to the predetermined patch ratio β to be retained after cropping, for example, β is 64%, that is, 64% of the area is retained, combined with the numerical range and numerical distribution of the attention score of the image area in the attention map after the magnification processing, the attention score threshold is determined. The higher the attention score, the greater the contribution of the image area to the final output result. For example, for aerial images, the attention map will show which areas (such as streets) are more important for geographic positioning, and which areas that are not visible from another perspective (such as the roof of a high-rise building) contribute less. Therefore, binarization is performed based on the comparison result between the attention score of the image area in the attention map and the attention score threshold, such as setting the pixel value of the image area with an attention score greater than the attention score threshold to "1", and vice versa, the pixel value is set to "0", so as to obtain a binary attention map. "1" represents the first pixel value, and "0" represents the second pixel value.

[0059] In step b, the binary attention map is matched with the pixel position of the aerial image, and the image area with the pixel value of "0" is cropped out in the aerial image, and the image area with the pixel value of "1" is retained. In this way, the area with low attention score is cropped out, and the area that makes important contribution to the geolocation task is retained to obtain the cropped aerial image.

[0060] In this embodiment, by setting and β, the number of patches retained after the final cropping of the aerial image is . Select β× =1, the amount of calculation remains unchanged, the resolution is increased, and the performance is improved. By adjusting these parameters, the model can reduce the amount of calculation.

[0061] use =1, the resolution is not increased, but the amount of computation is reduced due to cropping.

[0062] In the cross-view geolocation task, only some areas are visible between the two views. For example, many areas in aerial images (such as rooftops of tall buildings) are invisible in street view and contribute little to the positioning task. By cropping out these unimportant areas, the model can handle parts that are important for matching (such as street areas). The key areas are retained, that is, areas with high attention scores, which are also areas that contribute significantly to the geolocation task (such as streets).

[0063] This implementation adopts training method 2 for training, which includes a first training stage and a second training stage. The second training stage adopts an image cropping mechanism, which specifically uses the attention map generated by the Transformer encoder to remove areas with low contribution to geographic positioning in the aerial image, retain high contribution areas, and enhance detail information through resolution improvement to generate high-resolution images, thereby reducing the computational burden and improving performance.

[0064] The present invention also discloses a cross-view geolocation method based on the fusion of visual and semantic features. In a preferred embodiment, see Figure 6 , the geolocation method includes: Step B1, obtaining a street-shot image to be located, which is not limited to street-shot images taken by a user on the ground through a mobile phone terminal, a camera or a robot.

[0065] Step B2, the street image to be located is combined with multiple aerial images in the pre-constructed aerial image set to form multiple image pairs. Each image pair includes the street image to be located and an aerial image in the aerial image set. Each aerial image in the aerial image set is associated with an accurate geographic location information.

[0066] Step B3, input each image pair into the trained cross-view feature extraction model to obtain a pair of street photography feature embedding 35 and aerial photography feature embedding 25, and the trained cross-view feature extraction model is obtained by training according to the cross-view feature extraction model training method provided by the present invention. The street photography image to be located of each image pair is input into the street photography feature extraction module of the cross-view feature extraction model to obtain the street photography feature embedding 35, and the aerial photography image of each image pair is input into the aerial photography feature extraction module of the cross-view feature extraction model to obtain the aerial photography feature embedding 25.

[0067] Step B4, calculating the distance between the street photography feature embedding 35 and the aerial photography feature embedding 25 in each image pair, preferably but not limited to the Euclidean distance, and selecting the image pair with the smallest distance between the street photography feature embedding 35 and the aerial photography feature embedding 25 as the matching image pair; Step B5: taking the geographical location information associated with the aerial image in the matching image pair as the geographical location information of the street-photographed image to be located, and outputting the acquired geographical location information.

[0068] The present invention also discloses a cross-view geolocation device based on the fusion of visual and semantic features, which is used to implement the cross-view geolocation method based on the fusion of visual and semantic features provided by the present invention. In a preferred embodiment, see Figure 7 , the device comprises: Image input module, obtaining street photography images to be located; An image pair acquisition module combines the street-photographed image to be located with multiple aerial images in a pre-built aerial image set to form multiple image pairs; The feature embedding acquisition module inputs each image pair into a trained cross-view feature extraction model to obtain a pair of street photography feature embedding 35 and aerial photography feature embedding 25, wherein the trained cross-view feature extraction model is obtained by training according to the cross-view feature extraction model training method provided by the present invention; A matching image pair selection module calculates the distance between the street photography feature embedding 35 and the aerial photography feature embedding 25 in each image pair, and selects the image pair with the smallest distance between the street photography feature embedding 35 and the aerial photography feature embedding 25 as the matching image pair; The output module uses the geographic location information associated with the aerial image in the matching image pair as the geographic location information of the street-shot image to be located, and outputs the obtained geographic location information.

[0069] In this embodiment, the image input module, the image pair acquisition module, the feature embedding acquisition module, the matching image pair selection module and the output module respectively correspond to step B1, step B2, step B3, step B4 and step B5 in the above-mentioned cross-view geolocation method based on the fusion of visual and semantic features, and will not be repeated here.

[0070] In a preferred embodiment, the cross-view geographic positioning device further includes an aerial image storage module for storing an aerial image set and geographic location information associated with each aerial image in the aerial image set.

[0071] The present invention also discloses a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the above-mentioned cross-view feature extraction model training method or the cross-view geolocation method based on the fusion of visual and semantic features provided by the present invention. The computer program product should be understood as a software product that mainly implements its solution through a computer program, such as a program product integrated in the cloud or a software library.

[0072] The present invention also discloses an electronic device. In one embodiment, the electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the cross-view feature extraction model training method or the cross-view geo-positioning method based on the fusion of visual and semantic features provided by the present invention.

[0073] like Figure 8, which is a schematic diagram of the structure of an electronic device for a cross-view feature extraction model training method or a cross-view geo-positioning method based on the fusion of visual and semantic features provided by an embodiment of the present invention. The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a cross-view feature extraction model training method or a cross-view geo-positioning method based on the fusion of visual and semantic features.

[0074] Among them, the processor 10 may be composed of an integrated circuit in some embodiments, for example, it may be composed of a single packaged integrated circuit, or it may be composed of multiple integrated circuits packaged with the same function or different functions, including one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, and a combination of various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes or executes programs or modules stored in the memory 11 (for example, executing a cross-view feature extraction model training method or a cross-view geolocation method based on visual and semantic feature fusion, etc.), and calls the data stored in the memory 11 to execute various functions of the electronic device and process data.

[0075] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 11 may also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (SmartMediaCard, SMC), a secure digital (SecureDigital, SD) card, a flash card (FlashCard), etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit of the electronic device and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device, such as a cross-view feature extraction model training method or a cross-view geolocation method program based on visual and semantic feature fusion, but also can be used to temporarily store data that has been output or is to be output.

[0076] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.

[0077] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device and to display a visual user interface.

[0078] Figure 8 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the electronic device, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0079] For example, although not shown, the electronic device may also include a power source (such as a battery) for supplying power to various components. Preferably, the power source may be logically connected to at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, and power status indicators. The electronic device may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.

[0080] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited by this structure.

[0081] Furthermore, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory).

[0082] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", "an implementation", "a preferred implementation" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0083] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A cross-view feature extraction model training method, characterized in that: include: Construct a training sample pair set including one or more positive sample pairs and one or more negative sample pairs, each sample pair including one street photography image and one aerial photography image; Build a cross-view feature extraction model, which includes: A street photography feature extraction module, used for obtaining street photography feature embedding of a street photography image, wherein the street photography feature extraction module comprises a street photography visual branch and a street photography semantic branch, and a street photography token fusion unit, a street photography Transformer encoder and a street photography embedding feature generation unit connected in sequence, wherein an input end of the street photography token fusion unit is respectively connected to an output end of the street photography visual branch and an output end of the street photography semantic branch; An aerial feature extraction module is used to obtain aerial feature embedding of an aerial image. The aerial feature extraction module includes an aerial vision branch and an aerial semantic branch, and an aerial token fusion unit, an aerial Transformer encoder, and an aerial embedding feature generation unit connected in sequence. The input end of the aerial token fusion unit is respectively connected to the output end of the aerial vision branch and the output end of the aerial semantic branch. The constructed cross-view feature extraction model is iteratively trained in batches using a set of training sample pairs, and a cross-view feature extraction model is obtained when the training stop condition is reached.

2. A cross-view feature extraction model training method as claimed in claim 1, characterized in that: The street photography visual branch includes a first linear projection unit and a first position embedding unit connected in sequence, and the street photography semantic branch includes a first semantic feature extraction unit and a first MLP unit connected in sequence; and / or, The aerial photography vision branch includes a second linear projection unit and a second position embedding unit connected in sequence, and the aerial photography semantic branch includes a second semantic feature extraction unit and a second MLP unit connected in sequence.

3. A cross-view feature extraction model training method as claimed in claim 1 or 2, characterized in that: In the step of iteratively training the constructed cross-view feature extraction model in batches using the training sample pair set, each batch includes a first training stage, and the first training stage of each batch includes: Inputting the subset of training sample pairs of the current batch into the cross-view feature extraction model; wherein the subset of training sample pairs of the current batch includes one or more positive sample pairs and one or more negative sample pairs; The batch triplet loss is calculated using the street photography feature embeddings and aerial photography feature embeddings of multiple sample pairs in the current batch of training sample pair subsets output by the cross-view feature extraction model; The batch semantic alignment loss is calculated using the street visual tokens and street semantic tokens of the street images in the subset, as well as the aerial visual tokens and aerial semantic tokens of the aerial images, using the current batch of training samples output by the cross-view feature extraction model; wherein the street images are input into the street visual branch to obtain street visual tokens, the street images are input into the street semantic branch to obtain street semantic tokens, the aerial images are input into the aerial visual branch to obtain aerial visual tokens, and the aerial images are input into the aerial semantic branch to obtain aerial semantic tokens; Calculate the total batch loss based on batch triplet loss and batch semantic alignment loss; Update the network parameters of the cross-view feature extraction model based on the total batch loss.

4. A cross-view feature extraction model training method as claimed in claim 3, characterized in that: The street-level images and aerial images of each sample pair are processed in blocks and then input into the cross-view feature extraction model. The batch semantic alignment loss is calculated by: Calculate the binary norm of the difference between the street photography visual token and the street photography semantic token of each street photography image block in each street photography image in the current batch, and record it as the binary norm of the difference of the street photography image block; Calculate the binary norm of the difference between the aerial visual token and the aerial semantic token of each aerial image block in each aerial image in the current batch, and record it as the binary norm of the difference of the aerial image block; The average value of the difference two norms of all street photography image blocks in the current batch is calculated, which is recorded as the first average value. The average value of the difference two norms of all aerial photography image blocks in the current batch is calculated, which is recorded as the second average value. The average of the first average value and the second average value is calculated to obtain the batch semantic alignment loss.

5. A cross-view feature extraction model training method as claimed in claim 3, characterized in that: The total batch loss is calculated according to the following formula: ,in, represents the batch triplet loss, represents the batch semantic alignment loss, Represents the adjustment factor.

6. A cross-view feature extraction model training method as claimed in claim 3, characterized in that: The cross-view feature extraction model also includes a cropping module; Each batch also includes a second training phase, which includes: Get the attention map extracted by the aerial Transformer encoder during the first training phase for each aerial image in the current batch of training sample pairs; The cropping module crops each aerial image according to the attention map of each aerial image to obtain a cropped aerial image, and the cropped aerial image and the street image in the sample pair to which each aerial image belongs form a new sample pair; Input all new sample pairs into the cross-view feature extraction model obtained in the first training stage to obtain new street feature embeddings and new aerial feature embeddings for multiple new sample pairs, new street visual tokens and new street semantic tokens for street images, and new aerial visual tokens and new aerial semantic tokens for cropped aerial images; Calculate the triplet loss of the new batch and the semantic alignment loss of the new batch, and calculate the total loss of the new batch based on the triplet loss of the new batch and the semantic alignment loss of the new batch; The network parameters of the cross-view feature extraction model are updated again based on the total loss of the new batch.

7. A cross-view geolocation method based on the fusion of visual and semantic features, characterized in that: include: Obtain the street photography image to be located; The street-photographed image to be located is respectively combined with a plurality of aerial images in a pre-constructed aerial image set to form a plurality of image pairs; Input each image pair into a trained cross-view feature extraction model to obtain a pair of street photography feature embedding and aerial photography feature embedding, wherein the trained cross-view feature extraction model is trained according to the method of any one of claims 1 to 6; Calculate the distance between the street photography feature embedding and the aerial photography feature embedding in each image pair, and select the image pair with the smallest distance between the street photography feature embedding and the aerial photography feature embedding as the matching image pair; The geographical location information associated with the aerial image in the matching image pair is used as the geographical location information of the street-shot image to be located, and the obtained geographical location information is output.

8. A cross-view geolocation device based on the fusion of visual and semantic features, used to implement the cross-view geolocation method based on the fusion of visual and semantic features as claimed in claim 7, characterized in that: include: Image input module, obtaining the street-shot image to be located; An image pair acquisition module combines the street-photographed image to be located with multiple aerial images in a pre-built aerial image set to form multiple image pairs; A feature embedding acquisition module, inputting each image pair into a trained cross-view feature extraction model to obtain a pair of street photography feature embedding and aerial photography feature embedding, wherein the trained cross-view feature extraction model is trained according to the method of any one of claims 1 to 6; The matching image pair selection module calculates the distance between the street photography feature embedding and the aerial photography feature embedding in each image pair, and selects the image pair with the smallest distance between the street photography feature embedding and the aerial photography feature embedding as the matching image pair; The output module uses the geographic location information associated with the aerial image in the matching image pair as the geographic location information of the street-shot image to be located, and outputs the obtained geographic location information.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: The electronic device comprises: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor so that the at least one processor can perform the method as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Natural language pedestrian retrieval method and system combining token and feature alignment

    CN115311687A

  • Semantic segmentation method and device for aerial image, equipment and storage medium

    CN115471765A

  • Image classification method based on generalized zero sample learning and endangered animal identification method

    CN118470421A

  • Multi-modal semantic segmentation method, system and device in industrial scene and storage medium

    CN118657942A

  • Unsupervised cross-modal hash retrieval method based on CLIP and attention fusion mechanism

    CN118861327A