A pixel-level cross-view image positioning method and system based on deep learning

Through a deep learning-based method, the features of the image to be located and the candidate images for shooting are extracted, and the feature similarity is calculated to achieve the probability distribution of the target location, which solves the problem of insufficient positioning accuracy and generalization ability of cross-view angle images in the prior art, and achieves pixel-level positioning with high precision and high generalization ability.

CN115203460BActive Publication Date: 2025-05-09SUN YAT SEN UNIVERSITY SHENZHEN +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210782818.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-05-09
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

Existing cross-view image positioning technology is difficult to achieve high-precision and high generalization capabilities of pixel-level positioning, especially when the viewing angles of the images to be queried and the database images change dramatically and the time span is uncertain.

Method used

Using a deep learning-based method, image feature extraction is performed through a convolutional neural network to extract the set of image to be positioned and top-shot candidate images, and feature similarity between the ground feature map and top-shot feature map is calculated, and the probability distribution of the target location is calculated based on this.

Benefits of technology

The pixel-level cross-view image positioning with high flexibility, high precision and high generalization capabilities can accurately locate and obtain the accurate geographical location of the shooting point without requiring the image to be positioned to be aligned with the center of the database image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115203460B_ABST
    Figure CN115203460B_ABST
Patent Text Reader

Abstract

The present invention discloses a pixel-level cross-view image positioning method and system based on deep learning, the method comprising: obtaining an image to be positioned of a target to be positioned and a set of overhead candidate images corresponding to the image to be positioned; extracting image features of the image to be positioned and the set of overhead candidate images through a convolutional neural network to obtain a ground feature map and an overhead feature map; calculating the target location probability distribution of the target to be positioned based on feature similarity between features, and then calculating the pixel-level positioning coordinates; according to the pixel-level positioning coordinates, combined with the shooting parameter information of the set of overhead candidate images, determining the positioning information of the target to be positioned. The present invention has high flexibility, high precision and high generalization ability, calculates the positioning probability map through high-resolution overhead features and ground global features, and then obtains the pixel coordinates of the ground image, which are finally converted into actual geographic coordinates, and can be widely used in the field of image processing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a pixel-level cross-viewing image positioning method and system based on deep learning. Background Art

[0002] The goal of cross-view image localization is to obtain the location information of the ground query image using the overhead database image with geographic coordinates. Due to the drastic change in the perspective of the query image and the database image and the uncertain time span, there is a huge apparent difference between the query image and the database image, making the cross-view image localization task extremely challenging.

[0003] The existing approximate implementation solutions mainly include the following two:

[0004] (1) Cross-view image retrieval scheme

[0005] This type of solution simplifies the cross-view localization task into a cross-view image retrieval task, requiring the center of the query image and the database image to be aligned. The main challenge faced by this type of solution is the apparent difference caused by changes in perspective. First, the shape and appearance of objects under different perspectives, especially objects with significant height differences such as buildings, mountains, and vegetation, will change significantly. Second, the visible range of images from different perspectives is significantly different, resulting in greatly different information contained in the images. Finally, due to the severe distortion of ground panoramic images, the geometric relationships between objects are also completely different.

[0006] To alleviate the impact of the above problems, such solutions usually use twin neural networks to extract global features of ground and database images, and use the powerful representation ability of deep neural networks to achieve Robin's image feature representation to cope with the apparent changes of the target under different viewing angles. Secondly, such solutions use technologies such as attention mechanisms to achieve effective feature extraction and enhancement to cope with the impact of changes in the visible range. Finally, such solutions usually use polar coordinate transformation to transform the overhead image into a pseudo-panoramic image, reducing the apparent difference between the image to be located and the overhead image.

[0007] (2) Direct regression cross-view image localization solution

[0008] This type of scheme proposes a coarse-to-fine joint retrieval and calibration geolocation framework, models the coordinate calculation as a regression problem, takes the global feature vectors of the image to be located and the database image as input, and uses MLP (multi-layer perceptron) to predict the pixel coordinates of the image to be located. The query image can be arbitrary in the region of interest, and the reference image is captured before the query appears. This scheme breaks the one-to-one retrieval setting of most datasets because the query and reference images are not perfectly aligned, and there may be multiple reference images covering a query location.

[0009] The scheme first uses the twin network to output its representation embedded in high-dimensional space to compare the similarity between the two samples, thereby screening out the image with the highest similarity to the image to be located from the overhead view images in the database to complete image retrieval, and then fuses the screened overhead view feature vector with the feature vector of the image to be located, and then inputs it into a multi-layer perceptron (MLP) to predict the offset of the query position relative to the center of the retrieved overhead view image.

[0010] Disadvantages of the above two related technical solutions:

[0011] (1) Disadvantages of cross-view image retrieval schemes

[0012] a. Cross-view image retrieval schemes are mostly one-to-one retrieval, which simply assumes that each query ground view image has a corresponding reference top view image, and the top view image is aligned with the center of the query image. This is not practical in real-world applications because the query image may appear at any location in the region of interest (top view image). In this case, the cross-view image retrieval scheme cannot obtain the precise geographic location of the ground image.

[0013] b. The existing cross-view localization network is basically a twin network. The twin neural network takes two samples as input and outputs their representation embedded in a high-dimensional space to compare the similarity of the two samples, thereby selecting the image with the highest similarity to the image to be located from the overhead image in the database to complete image retrieval, but cannot further achieve pixel-level image localization.

[0014] (2) Disadvantages of direct regression cross-view image localization solutions

[0015] The cross-view geolocation scheme of direct regression models the coordinate calculation as a regression problem, takes the global feature vectors of the image to be located and the database image as input, and uses MLP (multi-layer perceptron) to predict the pixel coordinates of the image to be located. This type of method has limited performance and poor generalization, making it difficult to be applied in practice. Summary of the invention

[0016] In view of this, an embodiment of the present invention provides a deep learning-based pixel-level cross-view image positioning method and system with high flexibility, high precision and high generalization capability.

[0017] An aspect of an embodiment of the present invention provides a pixel-level cross-view image positioning method based on deep learning, comprising:

[0018] Acquire an image to be located of the target to be located and a set of overhead candidate images corresponding to the image to be located;

[0019] Extracting image features from the image to be located and the set of candidate overhead images by using a convolutional neural network to obtain a ground feature map and an overhead feature map;

[0020] Calculating a target location probability distribution of a target to be located according to feature similarity between the ground feature map and the overhead feature map;

[0021] Calculate pixel-level positioning coordinates according to the target location probability distribution;

[0022] The positioning information of the target to be positioned is determined according to the pixel-level positioning coordinates and in combination with the shooting parameter information of the overhead candidate image set.

[0023] Optionally, the step of extracting image features from the image to be located and the set of candidate overhead images by using a convolutional neural network to obtain a ground feature map and an overhead feature map includes:

[0024] Extracting image features from the image to be located by using a ground image feature extraction network to obtain a ground feature map;

[0025] Extracting image features from the candidate overhead image set using an overhead image feature extraction network to obtain an overhead feature map;

[0026] Wherein, the ground image feature extraction network is used to map the ground image into a high-dimensional feature vector;

[0027] The overhead image feature extraction network is used to aggregate image information while maintaining image resolution, and generate an overhead feature map that maintains spatial structure and spatial resolution and has specificity.

[0028] Optionally, the ground image feature extraction network adopts an "encoder-decoder" network structure; the overhead image feature extraction network adopts an "encoder-decoder" network structure;

[0029] Optionally, the encoder of the ground image feature extraction network is based on a VGG16 network for parsing image information; the decoder of the ground image feature extraction network uses a shallow convolutional neural network for compressing the spatial size of the feature map to obtain a feature vector;

[0030] The encoder of the ground image feature extraction network uses the first thirteen layers of the VGG16 network. The pooling layer of the encoder of the ground image feature extraction network uses a size of 2x2. After each pooling layer processing, the length and width of the image are reduced by half. After the 13 convolutional layers and pooling layers of the encoder of the ground image feature extraction network, the number of channels of the original image is 512;

[0031] The decoder of the ground image feature extraction network uses a shallow convolutional neural network. The first two layers of the network are used to reduce the size and number of channels of the feature image. The third layer of the network performs global average pooling along the spatial direction to generate a 1x1x128 feature vector, which is used to perform pixel-level similarity calculation with the feature map of high-resolution dense features of the overhead image.

[0032] Optionally, the overhead image feature extraction network is based on a U-net network, and the processing process of the overhead image feature extraction network includes a downsampling process and an upsampling process, wherein the downsampling process is used to extract image features, and the upsampling process is used to convert a low-resolution image containing high-level abstract features into a high-resolution image while retaining the high-level abstract features, and then perform a feature fusion operation with a high-resolution image of low-level surface features, so as to obtain a feature map that maintains the original resolution;

[0033] The downsampling process of the overhead image feature extraction network is implemented by a convolution block and two downsampling modules of the encoder. Each downsampling module contains two 3x3 convolution layers and a 2x2 pooling layer. The downsampling module is used to extract features, thereby obtaining local features and performing image-level classification to obtain abstract semantic features. After downsampling, the length and width of the image are reduced to 1 / 4 of the original, and the number of channels is 512.

[0034] The upsampling process of the overhead image feature extraction network is implemented by a layer of deconvolution, feature concatenation and two 3x3 convolution layers of the decoder. During each upsampling operation, the length and width of the image are doubled.

[0035] After the image obtained by the upsampling operation is spliced ​​with the downsampled image, a 1×1 convolution layer is used to reduce the dimension, and the number of channels is reduced to 128 to obtain a top-view image feature map at the original resolution.

[0036] Optionally, the calculating the target location probability distribution of the target to be located according to the feature similarity between the ground feature map and the overhead feature map includes:

[0037] The similarity of each pixel between the ground feature map and the top view feature map is calculated one by one by using a cosine similarity calculation method to obtain an initial response map;

[0038] After multiplying the initial response map by a preset temperature coefficient, the map is processed by a softmax function to obtain a probability map of each location, thereby determining a target location probability distribution of the target to be located;

[0039] The number of channels of the ground feature map and the top view feature map is the same.

[0040] Optionally, the method further comprises: after obtaining the pixel-level positioning coordinates, calculating the loss value of each coordinate by using a loss function, and when the loss value meets a preset condition, determining that the network training is completed;

[0041] The calculation formula of the loss value is:

[0042]

[0043] Among them, loss(x,y) represents a function related to the (x,y) coordinates; x1 represents the x-axis coordinate of the actual positioning coordinate; x2 represents the x-axis coordinate of the predicted positioning coordinate; y1 represents the y-axis coordinate of the actual positioning coordinate; y2 represents the y-axis coordinate of the predicted positioning coordinate.

[0044] Optionally, in the step of calculating the pixel-level positioning coordinates according to the target location probability distribution, the calculation formula of the pixel-level positioning coordinates is:

[0045]

[0046] Where r is the radius of the earth; (lat1, lon1) represents the longitude and latitude of the center point of the overhead view; (lat2, lon2) represents the longitude and latitude of the network predicted position, and d represents the actual distance between the two points (in meters). By making the longitude and latitude equal, the actual distances of the x and y axis offsets of the pixel coordinates of the overhead view and the predicted position can be calculated respectively. The actual distance of each pixel in the overhead view is known, so the pixel coordinates can be calculated.

[0047] In the network inference process, after obtaining the pixel coordinates of the predicted position, the pixel coordinates are converted to actual geographic coordinates through the inverse transformation of the above formula. The conversion formula is:

[0048]

[0049]

[0050] where d y is the vertical distance, d x is the horizontal distance between the two points, (lat1, lon1) represents the longitude and latitude of the center point of the top view; (lat2, lon2) represents the longitude and latitude of the network predicted position.

[0051] Another aspect of the embodiment of the present invention further provides a pixel-level cross-view image positioning system based on deep learning, including:

[0052] The first module is used to obtain an image to be located of the target to be located and a set of candidate overhead images corresponding to the image to be located;

[0053] The second module is used to extract image features of the image to be located and the set of candidate overhead images through a convolutional neural network to obtain a ground feature map and an overhead feature map;

[0054] The third module is used to calculate the target location probability distribution of the target to be located according to the feature similarity between the ground feature map and the overhead feature map;

[0055] A fourth module is used to calculate pixel-level positioning coordinates according to the probability distribution of the target location;

[0056] The fifth module is used to determine the positioning information of the target to be located according to the pixel-level positioning coordinates in combination with the shooting parameter information of the overhead candidate image set.

[0057] Another aspect of an embodiment of the present invention further provides an electronic device, including a processor and a memory;

[0058] The memory is used to store programs;

[0059] The processor executes the program to implement the method described above.

[0060] Another aspect of the embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a program.

[0061] The program is executed by a processor to implement the above-mentioned method.

[0062] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.

[0063] The embodiments of the present invention obtain an image to be located of a target to be located and a set of overhead candidate images corresponding to the image to be located; perform image feature extraction on the image to be located and the set of overhead candidate images through a convolutional neural network to obtain a ground feature map and an overhead feature map; calculate the target location probability distribution of the target to be located based on the feature similarity between the ground feature map and the overhead feature map; calculate pixel-level positioning coordinates based on the target location probability distribution; determine the positioning information of the target to be located based on the pixel-level positioning coordinates and the shooting parameter information of the set of overhead candidate images. The present invention has high flexibility, high precision and high generalization ability, and calculates the positioning probability map through high-resolution overhead features and ground global features, thereby obtaining the pixel coordinates of the ground image, which are finally converted into actual geographic coordinates. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0065] Figure 1 An overall step flow chart provided for an embodiment of the present invention;

[0066] Figure 2 A schematic diagram of the structure of a ground image feature extraction network provided by an embodiment of the present invention;

[0067] Figure 3 A schematic diagram of the structure of a feature extraction network for overhead images provided in an embodiment of the present invention;

[0068] Figure 4 A flowchart of calculating similarity provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0069] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0070] Since the cross-view image retrieval solution in the prior art cannot obtain the accurate geographical location of the ground image, the present invention does not need to consider the problem of image center alignment to address this shortcoming;

[0071] Since the existing cross-view positioning network is basically a twin network, the twin neural network takes two samples as input and outputs its representation embedded in a high-dimensional space to compare the similarity of the two samples, thereby screening out the image with the highest similarity to the image to be positioned from the overhead image in the database to complete the image retrieval, and cannot further achieve pixel-level image positioning. In response to this shortcoming, the present invention uses a heterogeneous network design, uses a high-resolution network to generate a high-resolution dense feature map of the overhead image, uses a pyramid network to extract a global feature vector with query, and finally achieves pixel-level positioning by measuring feature similarity and converting it into corresponding geographic coordinates.

[0072] Since the cross-view geolocation solution of direct regression in the prior art models the coordinate calculation as a regression problem, the global feature vectors of the image to be located and the database image are used as input, and the pixel coordinates of the image to be located are predicted using MLP (multi-layer perceptron). This type of method has limited performance and poor generalization, making it difficult to be applied in practice. In view of this shortcoming, the present invention realizes coordinate positioning by representation learning, calculates the positioning probability map through high-resolution overhead features and global ground features, and then obtains the pixel coordinates of the ground image, which are finally converted into actual geographic coordinates.

[0073] Specifically, an aspect of an embodiment of the present invention provides a pixel-level cross-view image positioning method based on deep learning, including:

[0074] Acquire an image to be located of the target to be located and a set of overhead candidate images corresponding to the image to be located;

[0075] Extracting image features from the image to be located and the set of candidate overhead images by using a convolutional neural network to obtain a ground feature map and an overhead feature map;

[0076] Calculating a target location probability distribution of a target to be located according to feature similarity between the ground feature map and the overhead feature map;

[0077] Calculate pixel-level positioning coordinates according to the target location probability distribution;

[0078] The positioning information of the target to be positioned is determined according to the pixel-level positioning coordinates and in combination with the shooting parameter information of the overhead candidate image set.

[0079] Optionally, the step of extracting image features from the image to be located and the set of candidate overhead images by using a convolutional neural network to obtain a ground feature map and an overhead feature map includes:

[0080] Extracting image features from the image to be located by using a ground image feature extraction network to obtain a ground feature map;

[0081] Extracting image features from the candidate overhead image set using an overhead image feature extraction network to obtain an overhead feature map;

[0082] Wherein, the ground image feature extraction network is used to map the ground image into a high-dimensional feature vector;

[0083] The overhead image feature extraction network is used to aggregate image information while maintaining image resolution, and generate an overhead feature map that maintains spatial structure and spatial resolution and has specificity.

[0084] Optionally, the ground image feature extraction network adopts an "encoder-decoder" network structure; the overhead image feature extraction network adopts an "encoder-decoder" network structure;

[0085] Optionally, the encoder of the ground image feature extraction network is based on a VGG16 network for parsing image information; the decoder of the ground image feature extraction network uses a shallow convolutional neural network for compressing the spatial size of the feature map to obtain a feature vector;

[0086] The encoder of the ground image feature extraction network uses the first thirteen layers of the VGG16 network. The pooling layer of the encoder of the ground image feature extraction network uses a size of 2x2. After each pooling layer processing, the length and width of the image are reduced by half. After the 13 convolutional layers and pooling layers of the encoder of the ground image feature extraction network, the number of channels of the original image is 512;

[0087] The decoder of the ground image feature extraction network uses a shallow convolutional neural network. The first two layers of the network are used to reduce the size and number of channels of the feature image. The third layer of the network performs global average pooling along the spatial direction to generate a 1x1x128 feature vector, which is used to perform pixel-level similarity calculation with the feature map of high-resolution dense features of the overhead image.

[0088] Optionally, the overhead image feature extraction network is based on U-net, and the processing process of the overhead image feature extraction network includes a downsampling process and an upsampling process, wherein the downsampling process is used to extract image features, and the upsampling process is used to convert a low-resolution image containing high-level abstract features into a high-resolution image while retaining the high-level abstract features, and then perform a feature fusion operation with a high-resolution image of low-level surface features, so as to obtain a feature map that maintains the original resolution;

[0089] The downsampling process of the overhead image feature extraction network is implemented by a convolution block and two downsampling modules of the encoder. Each downsampling module contains two 3x3 convolution layers and a 2x2 pooling layer. The downsampling module is used to extract features, thereby obtaining local features and performing image-level classification to obtain abstract semantic features. After downsampling, the length and width of the image are reduced to 1 / 4 of the original, and the number of channels is 512.

[0090] The upsampling process of the overhead image feature extraction network is implemented by a layer of deconvolution, feature concatenation and two 3x3 convolution layers of the decoder. During each upsampling operation, the length and width of the image are doubled.

[0091] After the image obtained by the upsampling operation is spliced ​​with the downsampled image, a 1×1 convolution layer is used to reduce the dimension, and the number of channels is reduced to 128 to obtain a top-view image feature map at the original resolution.

[0092] Optionally, the calculating the target location probability distribution of the target to be located according to the feature similarity between the ground feature map and the overhead feature map includes:

[0093] The similarity of each pixel between the ground feature map and the top view feature map is calculated one by one by using a cosine similarity calculation method to obtain an initial response map;

[0094] After multiplying the initial response map by a preset temperature coefficient, the map is processed by a softmax function to obtain a probability map of each location, thereby determining a target location probability distribution of the target to be located;

[0095] The number of channels of the ground feature map and the top view feature map is the same.

[0096] Optionally, the method further comprises: after obtaining the pixel-level positioning coordinates, calculating the loss value of each coordinate by using a loss function, and when the loss value meets a preset condition, determining that the network training is completed;

[0097] The calculation formula of the loss value is:

[0098]

[0099] Among them, loss(x,y) represents a function related to the (x,y) coordinates, which is the loss value between the actual position and the network predicted position; (x1,y1) represents the pixel coordinates of the actual position; (x2,y2) represents the pixel coordinates of the network predicted position.

[0100] Optionally, in the step of calculating the pixel-level positioning coordinates according to the target location probability distribution, the calculation formula of the pixel-level positioning coordinates is:

[0101]

[0102] Where r is the radius of the earth; (lat1, lon1) represents the longitude and latitude of the center point of the overhead view; (lat2, lon2) represents the longitude and latitude of the network predicted position, and d represents the actual distance between the two points (in meters). By making the longitude and latitude equal, the actual distances of the x and y axis offsets of the pixel coordinates of the overhead view and the predicted position can be calculated respectively. The actual distance of each pixel in the overhead view is known, so the pixel coordinates can be calculated.

[0103] In the network inference process, after obtaining the pixel coordinates of the predicted position, the pixel coordinates are converted to actual geographic coordinates through the inverse transformation of the above formula. The conversion formula is:

[0104]

[0105]

[0106] where d y is the vertical distance, d x is the horizontal distance between the two points, (lat1, lon1) represents the longitude and latitude of the center point of the top view; (lat2, lon2) represents the longitude and latitude of the network predicted position.

[0107] Another aspect of the embodiment of the present invention further provides a pixel-level cross-view image positioning system based on deep learning, including:

[0108] The first module is used to obtain an image to be located of the target to be located and a set of candidate overhead images corresponding to the image to be located;

[0109] The second module is used to extract image features of the image to be located and the set of candidate overhead images through a convolutional neural network to obtain a ground feature map and an overhead feature map;

[0110] The third module is used to calculate the target location probability distribution of the target to be located according to the feature similarity between the ground feature map and the overhead feature map;

[0111] A fourth module is used to calculate pixel-level positioning coordinates according to the probability distribution of the target location;

[0112] The fifth module is used to determine the positioning information of the target to be located according to the pixel-level positioning coordinates in combination with the shooting parameter information of the overhead candidate image set.

[0113] Another aspect of an embodiment of the present invention further provides an electronic device, including a processor and a memory;

[0114] The memory is used to store programs;

[0115] The processor executes the program to implement the method described above.

[0116] Another aspect of the embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a program,

[0117] The program is executed by a processor to implement the above-mentioned method.

[0118] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.

[0119] The specific implementation process of the present invention is described in detail below in conjunction with the accompanying drawings:

[0120] The purpose of the present invention is to realize cross-view image positioning, that is, given a panoramic image taken on the ground, the positioning prediction of the shooting location of the ground image is realized by combining the overhead panoramic image collected by satellite or drone. The algorithm input of the present invention is an image to be positioned and several overhead candidate images. The image to be positioned is a panoramic image taken on the ground, and the overhead candidate image is a non-center-aligned database image retrieved from the database that may contain the shooting location of the image to be positioned.

[0121] Taking an image to be located and an overhead candidate image as an example, the core idea of ​​the present invention is to use a convolutional neural network to extract features from the overhead image of the image to be located, calculate the probability distribution of the shooting location based on the similarity between the features, and obtain the pixel-level positioning coordinates based on the probability distribution. Finally, the geographic coordinates of the shooting location are obtained by combining the geographic tags of the overhead image and the camera parameters. The overall process is as follows: Figure 1 shown.

[0122] 1. Image feature extraction process:

[0123] In order to ensure the generalization performance of the algorithm, the present invention models the cross-view positioning problem as a representation learning problem, that is, the position coordinates of the image to be positioned are calculated based on the feature similarity between the overhead image and the image to be positioned. In view of the powerful function fitting ability of the convolutional neural network, the present invention uses a convolutional neural network as a feature extractor to extract features from the overhead image and the image to be positioned. Since there are different algorithm requirements for the overhead feature map and the ground feature map, the present invention uses a heterogeneous feature extraction network to process the overhead map and the ground panoramic map respectively. Hereinafter, the ground image feature extraction network and the overhead image feature extraction network are referred to as the ground network and the overhead network respectively.

[0124] Ground image feature extraction network: The purpose of the ground network is to map the ground image into a high-dimensional feature representation (feature vector). The present invention uses a convolutional neural network with an "encoder-decoder" structure as the ground network. The encoder is based on the VGG16 network to parse the image information, and the decoder uses a shallow convolutional neural network to compress the spatial size of the feature map to obtain a feature vector.

[0125] The network structure of the ground image feature extraction network is as follows: Figure 2 As shown in the figure. The encoder uses the first thirteen layers of the VGG16 network. The VGG16 network has a simple structure and uses multiple convolution layers with smaller convolution kernels (3x3) instead of a convolution layer with a larger convolution kernel. On the one hand, it can reduce parameters, and on the other hand, it is equivalent to performing more nonlinear mappings, which can increase the network's expressive power. The pooling layer uses a size of 2x2. After each pooling layer, the length and width of the image are reduced to half of the original. After 13 convolutional layers and pooling layers, the number of channels of the original image reaches 512, so that more information of the original image can be extracted.

[0126] Overhead image feature extraction network: The purpose of the overhead network is to aggregate image information while maintaining image resolution, generating an overhead feature map with specificity and spatial structure and resolution. This paper uses UNet as the basic framework of the overhead network, and its structure is as follows: Figure 3 shown.

[0127] The left half is the downsampling process, which is used to extract the features of the image, while the right half is the upsampling process, which allows the low-resolution image containing high-level abstract features to become high-resolution while retaining the high-level abstract features, and then performs feature fusion operation with the high-resolution image with low-level surface features on the left to obtain a feature map that maintains the original resolution.

[0128] Encoder of the overhead image feature extraction network: The left half is the downsampling process, which consists of a convolution block and two downsampling modules. Each downsampling module contains two 3x3 convolution layers and a 2x2 pooling layer. Its function is to extract features (obtain local features and perform image-level classification) to obtain abstract semantic features. After downsampling, the length and width of the image are reduced to 1 / 4 of the original, and the number of channels is 512, that is, more information of the original image is extracted.

[0129] Decoder of the overhead image feature extraction network: The right half (up-sampling) consists of a layer of deconvolution, feature concatenation, and two 3x3 convolution layers, which are repeated 4 times, corresponding to the feature extraction network. After each up-sampling operation, the length and width of the image are doubled, and then concatenated with the down-sampled image. Finally, a layer of 1x1 convolution is performed to reduce the number of channels to a specific number of 128, matching the number of channels of the ground feature map, thereby obtaining the overhead image feature map at the original resolution.

[0130] 2. Network model training process:

[0131] The feature maps of the ground image and the overhead image can be obtained through the aforementioned image feature extraction network, where the ground image is a feature vector of 1x1xC, and the overhead view is a feature map of HxWxC, with the same number of channels, so the similarity of the pixels of the two image feature maps can be calculated one by one. This embodiment uses cosine similarity to measure the similarity between images. The closer the cosine value is to 1, the closer the angle is to 0 degrees, that is, the more similar the two vectors are. The calculation formula for the cosine of the angle is as follows.

[0132]

[0133] Among them, a is the ground feature vector, b is the feature vector of a main pixel in the top view feature map, and after the cosine similarity calculation, an initial response map of size HxW can be obtained. The initial response map is multiplied by the temperature coefficient (to avoid the value obtained after softmax being too small) and then the probability map is obtained through softmax, so that the output value range of each position of the response map is mapped to [0,1], and the sum of the output values ​​of each position is constrained to be 1. Let each output value of the probability map be multiplied by the two-dimensional coordinates of the corresponding position and then summed to obtain the predicted coordinates of the network. Specifically, the process of calculating the similarity is as follows Figure 4 shown.

[0134] After obtaining the network prediction coordinates, the present invention uses the L1 loss function as the loss function for network training, as shown in the following formula, to calculate the loss value of the network prediction coordinates.

[0135] The calculation formula of the loss value is:

[0136]

[0137] Among them, loss(x,y) represents a function related to the (x,y) coordinates, which is the loss value between the actual position and the network predicted position; (x1,y1) represents the pixel coordinates of the actual position; (x2,y2) represents the pixel coordinates of the network predicted position.

[0138] The present invention uses multi-level network supervision, calculates the pixel correlation between each layer of feature map in the Unet upsampling stage and the ground feature map to obtain the corresponding corresponding map and probability map, and obtains their respective loss values. The sum of the loss values ​​of the three layers is used as the loss value of the entire network training, thereby supervising the learning of each upsampling process and improving the accuracy of training.

[0139] 3. Geographic coordinate calculation process:

[0140] After the aforementioned network training, the predicted coordinates of a certain overhead image can be obtained by placing the new image to be located and the database image into the present invention. The coordinates are pixel coordinates and need to be converted into actual geographic coordinates through a formula. The present invention uses the haversine formula to calculate the actual geographic coordinates. The calculation formula of the actual geographic coordinate d is as follows:

[0141]

[0142]

[0143]

[0144] Among them, d y The vertical distance between the prediction point and the center point of the top view, d x is the horizontal distance between the two points, (lat1, lon1) represents the longitude and latitude of the center point of the top view; (lat2, lon2) represents the longitude and latitude of the network predicted position, r is the radius of the earth, lat1 and lat2 are the latitude coordinates of the two points, lon1 and lon2 are the longitude coordinates of the two points. The actual geographic location coordinates of the center point of the top view in the database image can be obtained according to the camera's internal parameters to obtain the actual distance offset of each pixel point, thereby calculating the actual geographic coordinates.

[0145] In summary, the present invention adopts a pixel-level cross-view image positioning overall solution of metric learning, while the existing cross-view image positioning solutions are mostly image retrieval solutions, which require that the shooting point of the image to be positioned is located at the center of the database image, and the image to be positioned and the database image are matched by image retrieval, and finally the geographic tag carried by the matched database image is used as the positioning result of the image to be positioned. The present invention does not require the center of the image to be positioned to be aligned with the database image, and can accurately position the image in the overhead image containing the shooting point of the image to be positioned and obtain the accurate geographic location of the shooting point.

[0146] The present invention proposes a special feature extraction network training method. Since the existing pixel-level cross-view image localization scheme uses the global feature vectors of the image to be located and the database image as input, the MLP (multi-layer perceptron) is used to predict the pixel coordinates of the image to be located. The present invention uses a method of measuring feature similarity to obtain a response map to achieve pixel-level localization. In order to obtain a high-resolution overhead feature map, a heterogeneous network design is used to extract image features.

[0147] The present invention adopts a heterogeneous image feature extraction network design. Since the existing cross-view positioning network is basically a twin network, the twin neural network takes two samples as input and outputs its representation embedded in a high-dimensional space to compare the similarity of the two samples, thereby screening out the image with the highest similarity to the image to be positioned from the overhead image in the database to complete the image retrieval. In order to achieve pixel-level positioning, the present invention uses a heterogeneous network design, uses a high-resolution network to generate a high-resolution dense feature map of the overhead view, uses a pyramid network to extract a global feature vector with query, and finally achieves pixel-level positioning by measuring feature similarity and converting it into corresponding geographic coordinates.

[0148] The present invention adopts a pixel-level cross-view positioning algorithm based on feature similarity. The existing cross-view positioning coordinate calculation scheme models the coordinate calculation as a regression problem, takes the global feature vectors of the image to be positioned and the database image as input, and uses MLP (multi-layer perceptron) to predict the pixel coordinates of the image to be positioned. This type of method has limited performance and poor generalization, making it difficult to be applied in practice. The present invention realizes coordinate positioning in the form of representation learning, calculates the positioning probability map through the similarity of high-resolution overhead features and global ground features, and then obtains the pixel coordinates of the ground image, which is finally converted into actual geographic coordinates.

[0149] Compared with the prior art, the present invention has the following advantages:

[0150] 1. High flexibility: There is no need to align the center of the image to be located with the database image.

[0151] 2. High positioning accuracy: It can accurately locate in the overhead image containing the shooting point of the image to be located, and obtain the exact geographical location of the shooting point.

[0152] 3. High generalization ability: The pixel coordinates are predicted using the similarity of the feature maps of the ground image and the overhead image, which has strong generalization ability.

[0153] 4. Multi-level supervised learning: The system will correct the model according to the loss values ​​​​at multiple stages, so that the system will become more and more accurate as the training time progresses.

[0154] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.

[0155] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0156] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0157] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0158] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0159] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0160] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0161] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

[0162] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A pixel-level cross-view image localization method based on deep learning, characterized in that: include: Acquire an image to be located of the target to be located and a set of overhead candidate images corresponding to the image to be located; Extracting image features from the image to be located and the set of candidate overhead images by using a convolutional neural network to obtain a ground feature map and an overhead feature map; Calculating a target location probability distribution of a target to be located according to feature similarity between the ground feature map and the overhead feature map; Calculate pixel-level positioning coordinates according to the target location probability distribution; Determine the positioning information of the target to be positioned according to the pixel-level positioning coordinates and in combination with the shooting parameter information of the overhead candidate image set; The calculating the target location probability distribution of the target to be located according to the feature similarity between the ground feature map and the overhead feature map includes: The similarity of each pixel between the ground feature map and the top view feature map is calculated one by one by using a cosine similarity calculation method to obtain an initial response map; After multiplying the initial response map by a preset temperature coefficient, the map is processed by a softmax function to obtain a probability map of each location, thereby determining a target location probability distribution of the target to be located; The ground feature map and the top view feature map have the same number of channels; In the step of calculating the pixel-level positioning coordinates according to the probability distribution of the target location, the calculation formula of the pixel-level positioning coordinates is: Among them, r is the radius of the earth; (lat1, lon1) represents the longitude and latitude of the center point of the bird's-eye view; (lat2, lon2) represents the longitude and latitude of the network predicted location.

2. The pixel-level cross-view image positioning method based on deep learning according to claim 1, characterized in that: The step of extracting image features from the image to be located and the set of candidate overhead images by using a convolutional neural network to obtain a ground feature map and an overhead feature map includes: Extracting image features from the image to be located by using a ground image feature extraction network to obtain a ground feature map; Extracting image features from the candidate overhead image set using an overhead image feature extraction network to obtain an overhead feature map; Wherein, the ground image feature extraction network is used to map the ground image into a high-dimensional feature vector; The overhead image feature extraction network is used to aggregate image information while maintaining image resolution, and generate an overhead feature map that maintains spatial structure and spatial resolution and has specificity.

3. The pixel-level cross-view image positioning method based on deep learning according to claim 2, characterized in that: The ground image feature extraction network adopts an "encoder-decoder" network structure; the overhead image feature extraction network adopts an "encoder-decoder" network structure; The encoder of the ground image feature extraction network is based on the VGG16 network and is used to parse the image information; the decoder of the ground image feature extraction network uses a shallow convolutional neural network to compress the spatial size of the feature map to obtain a feature vector; The encoder of the ground image feature extraction network uses the first thirteen layers of the VGG16 network. The pooling layer of the encoder of the ground image feature extraction network uses a size of 2x2. After each pooling layer processing, the length and width of the image are reduced by half. After the 13 convolutional layers and pooling layers of the encoder of the ground image feature extraction network, the number of channels of the original image is 512; The decoder of the ground image feature extraction network uses a shallow convolutional neural network. The first two layers of the network are used to reduce the size and number of channels of the feature image. The third layer of the network performs global average pooling along the spatial direction to generate a 1x1x128 feature vector, which is used to perform pixel-level similarity calculation with the feature map of high-resolution dense features of the overhead image. The overhead image feature extraction network is based on the U-net network. The processing process of the overhead image feature extraction network includes a downsampling process and an upsampling process, wherein the downsampling process is used to extract image features, and the upsampling process is used to convert a low-resolution image containing high-level abstract features into a high-resolution image while retaining the high-level abstract features, and then perform a feature fusion operation with a high-resolution image of low-level surface features, thereby obtaining a feature map that maintains the original resolution; The downsampling process of the overhead image feature extraction network is implemented by a convolution block and two downsampling modules of the encoder. Each downsampling module contains two 3x3 convolution layers and a 2x2 pooling layer. The downsampling module is used to extract features, thereby obtaining local features and performing image-level classification to obtain abstract semantic features. After downsampling, the length and width of the image are reduced to 1 / 4 of the original, and the number of channels is 512. The upsampling process of the overhead image feature extraction network is implemented by a layer of deconvolution, feature concatenation and two 3x3 convolution layers of the decoder. During each upsampling operation, the length and width of the image are doubled. After the image obtained by the upsampling operation is spliced ​​with the downsampled image, a 1×1 convolution layer is used to reduce the dimension, and the number of channels is reduced to 128 to obtain a top-view image feature map at the original resolution.

4. The pixel-level cross-view image positioning method based on deep learning according to claim 1, characterized in that: The method further comprises: after obtaining the pixel-level positioning coordinates, calculating the loss value of each coordinate by using a loss function, and when the loss value meets a preset condition, determining that the network training is completed; The calculation formula of the loss value is: Among them, loss(x,y) represents a function related to the (x,y) coordinates; x1 represents the x-axis coordinate of the actual positioning coordinate; x2 represents the x-axis coordinate of the predicted positioning coordinate; y1 represents the y-axis coordinate of the actual positioning coordinate; y2 represents the y-axis coordinate of the predicted positioning coordinate.

5. A system using the deep learning-based pixel-level cross-view image positioning method as described in any one of claims 1 to 4, characterized in that: include: The first module is used to obtain an image to be located of the target to be located and a set of candidate overhead images corresponding to the image to be located; The second module is used to extract image features of the image to be located and the set of candidate overhead images through a convolutional neural network to obtain a ground feature map and an overhead feature map; The third module is used to calculate the target location probability distribution of the target to be located according to the feature similarity between the ground feature map and the overhead feature map; A fourth module is used to calculate pixel-level positioning coordinates according to the probability distribution of the target location; The fifth module is used to determine the positioning information of the target to be located according to the pixel-level positioning coordinates in combination with the shooting parameter information of the overhead candidate image set.

6. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 4.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Cross-view-angle image real-time matching geographic positioning method and system based on deep learning

    CN114241464A