Computer-implemented method and server
By introducing an angle loss objective function and a specific coordinate system during model training, the azimuth estimation and geolocation accuracy of street scene images are improved, and the problem of insufficient accuracy in the prior art is solved.
Patent Information
- Application Number
- CN202380068439.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-26
- Filing Date
- 2023-08-23
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to improve the accuracy of azimuth and geographic location matching of street scene images, especially in cross-view matching.
By defining the azimuth alignment coordinates and absolute angle error coordinates for the south-aligned azimuth alignment as part of the objective function during model training, the weighted soft margin triple loss function and absolute angle error loss function are used to improve the azimuth estimation granularity of street scene images.
It significantly improves the accuracy of orientation estimation and geographic positioning accuracy, improves the ground reality definition and coordinate system, and reduces absolute angle errors.
Smart Images

Figure CN119948535A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of machine learning. One aspect of the present invention relates to a computer-implemented method for training a machine learning model. Another aspect of the present invention relates to a server that uses a machine learning model in an inference phase. Another aspect of the present invention relates to a computer-implemented method for inference using a machine learning model. Yet another aspect of the present invention relates to a mobile communication device that uses a machine learning model in an inference phase. Background Art
[0002] Photos not only contain memories, but also provide us with a way to learn and perceive the world through the eyes of others, discover details that people may have overlooked before, and share emotions and knowledge with the community. With the advancement of hardware, personal high-quality cameras have become more affordable. Many creators are keen to share photos on the Internet. The captured images may not be carefully calibrated like those taken by dedicated multi-sensor systems (e.g., Google Street View vehicles), but a large number of crowdsourced images may provide rich information. If we can efficiently estimate the missing meta-information (e.g., geolocation, camera orientation) in these images and calibrate them to make them "ready", this huge hidden treasure can help complete various downstream tasks, such as map information extraction, car navigation and tracking, UAV positioning, hazard detection, social research.
[0003] To achieve this goal, we can perform three or more tasks: (a) straighten the image, (b) find the position, and / or (c) estimate the camera’s viewpoint.
[0004] Image-based geolocation is the study of camera locations aimed at inferring street view imagery. Among various geolocation methods, cross-view geolocation uses georeferenced aerial imagery (primarily satellite imagery). Given a street view query image ( Figure 1 (a)), the system finds the most similar match in the pool of satellite images ( Figure 1 (b)), and then the satellite image center is taken as the localization result. Due to its image retrieval property, cross-view matching with satellite images can be applied to large-scale search and achieve promising results.
[0005] For example, a paper by Y. Shi, X. Yu, D. Campbell, and H. Li, entitled "Where Am I Looking At? Joint Location and Orientation Estimation by Cross-View Matching," published in the 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) (Seattle, WA, USA, 2020, pp. 4064-4072, the contents of which are incorporated herein by reference), discloses a method for estimating the orientation of a query image (such as a street view image) relative to a candidate image (such as a satellite image). This document is referred to as the "DSM" paper.
[0006] One technical problem that may exist in the art is how to improve the accuracy of the position estimation of the query image, the geographic location of the query image, and / or the matching accuracy of any points of interest or other features within the query image. Summary of the invention
[0007] The embodiments may be implemented as set out in the independent claim. Some optional features are defined in the dependent claims.
[0008] Implementations of the technology disclosed herein may provide significant technical advantages. One or more advantages may include:
[0009] Significantly improve the accuracy of position estimation.
[0010] Improve the ground truth definition and / or coordinate system.
[0011] The absolute angle error and angle loss are defined as the ratio of the absolute angle error to the maximum value.
[0012] During model training, angle loss is included as part of the objective function.
[0013] Defines south-aligned azimuth alignment coordinates and continuous absolute angle error coordinates for azimuth estimation in cross-view matching with satellite imagery.
[0014] FOV unchanged due to south-aligned azimuth-aligned coordinates and angle error coordinate system.
[0015] We propose two methods to improve the granularity of orientation estimation for street view imagery without introducing any additional learnable parameters.
[0016] Propose a set of metrics for orientation estimation that is easier to understand and provides better clarity for real-world use cases.
[0017] Significantly improve the accuracy of geolocation estimates.
[0018] In an exemplary implementation, the functionality of the technology disclosed herein may be implemented in software running on a server communication device (such as a server cluster or a cloud computing platform), which communicates with an application running on a terminal (such as a mobile phone). The software that implements the functionality of the technology disclosed herein may be included in a computer program or a computer program product. The server communication device establishes a secure communication channel with the user terminal to receive queries from the user and present search ranking results to the user. The process may also include: training a machine learning model, using the model in the inference phase, and / or identifying points of interest with estimated locations and orientations. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The present invention will now be described by way of example only and with reference to the accompanying drawings, in which:
[0020] Figure 1 (a) is a street view image query with 0° and 315° clockwise misalignment.
[0021] Figure 1 (b) is the georeferenced satellite image used as a reference for cross-view matching.
[0022] Figure 1 (c) is a visualization of the angular misalignment on a sample street view image.
[0023] Figure 2 is a schematic block diagram illustrating an exemplary delivery / transportation service.
[0024] Figure 3 is a schematic block diagram illustrating an exemplary communication server for delivery / transportation services.
[0025] Figure 4 (a) is a schematic diagram of the camera axis.
[0026] Figure 4 (b) is a georeferenced satellite image of the south-aligned reference coordinates used for azimuth offset.
[0027] Figure 4 (c) is the rotation applied to the image to create a visualization of the orientation misalignment.
[0028] Figure 5a is a schematic block diagram of the overall architecture for geolocation and orientation estimation. The example shows a 360° FOV.
[0029] Figure 5b is a schematic block diagram of an inference architecture for geolocation and position estimation.
[0030] Figure 6 is the histogram of angle errors of the best instance of each model in the 360° image test. Each group has 8884 images.
[0031] Figure 7 is a histogram of the differences between the matching cases and all cases for the same / cross-dataset test of CVACT. Most of the removed results have good orientation estimates.
[0032] Figure 8 is the histogram of angular error (last 10 degrees) for CVUSA and CVACT tested on the same dataset.
[0033] Fig. 9 is the histogram of angle errors (first 10 degrees) for CFUSA and CVACT in cross-dataset testing.
[0034] Fig.10 is the histogram of angular error (last 10 degrees) for CVUSA and CVACT in cross-dataset testing. DETAILED DESCRIPTION
[0035] The techniques described in this article are primarily described with reference to their use in cross-view matching of street view imagery with satellite imagery. This may be useful in map creation, augmented reality, navigation, and the like.
[0036] Figure 2 An exemplary architecture of a system 100 is shown, wherein a plurality of users each have a communication device 104, a plurality of merchants each have a communication device 109, a plurality of drivers each have a user interface communication device 106, a server 102 (or geographically distributed servers), and a communication link 108 connecting each component. Each user uses a user software application (app) on a communication device 104 to contact the server 102. Similarly, drivers and merchants can use apps on their devices 106, 109.
[0037] For transactions based on delivery or e-commerce, user device 104 can allow the user to enter a query containing keywords of the goods of interest and the delivery address. The user can see a list of goods provided by the merchant and / or by the merchant, and order goods from the merchant. The merchant can use the merchant device 109 to contact the server 102 to provide information about its goods and receive orders for each confirmation transaction. The driver uses the driver device 106 to contact the server 102. The driver device 106 allows the driver to indicate whether they can accept the delivery job, information about their vehicle, and their location. The server 102 can then match the driver with the delivery based on, for example, the following: the geographic location of the merchant and the driver, income maximization, user or driver feedback rating, weather, driving conditions, traffic level / accidents, relative demand, environmental impact and / or supply level. Specific delivery costs and approximate delivery ETA can be provided to the user. If the user accepts the quotation, the system may pass the payment authorization process. If the authorization is approved, the merchant will receive a notification and be instructed to provide goods for the driver to pick up. Then, the selected driver will receive a notification and be guided to the pickup location to pick up the goods. During the delivery, the user device 104, the driver device 106, the merchant device 109, and the server 102 may be updated with real-time trip information, including the real-time location of the driver's vehicle, the destination, the driver's fare, and / or other trip-related information. At the end of the trip, the driver device 106 may send a confirmation to the server 102 that the trip has ended. Once the transaction is approved and / or the delivery is complete, the user device 104, the driver device 106, the merchant device 109, and the server 102 may be updated with the details of the completed financial transaction. This allows for efficient allocation of resources because the available driver fleet is optimized for user demand in each geographic area.
[0038] For transportation, the user device 104 may allow the user to enter their pickup location, destination address, one or more service parameters and / or post-ride information, such as ratings. One or more service parameters may include the number of seats in the vehicle, the style of the vehicle, the degree of environmental impact, and / or the desired type of transportation service. Each driver uses the driver app on the communication device 106 to contact the server 102. The driver app allows the driver to indicate whether they can participate in the ride job, information about the vehicle, their location, and / or post-ride information, such as ratings. The server 102 can then match the user with the driver based on, for example, the following: the user and driver's geographic location, income maximization, user or driver feedback rating, weather, driving conditions, traffic level / accidents, relative demand, environmental impact, and / or supply level. Specific transportation costs or ranges can be provided to users based on different types of vehicles, as well as approximate ETAs. If the user accepts the quote, the system may go through the payment authorization process. If the authorization is approved, the selected driver will be notified and directed to the pickup location to pick up the user / passenger. During the trip, the user device 104, the driver device 106, the merchant device 109, and the server 102 may be updated with real-time trip information, including the real-time location of the driver's vehicle, the destination, the driver's fare, and / or other trip-related information. At the end of the trip, the driver device 106 may send a confirmation to the server 102 that the trip has ended. Once the transaction is approved and / or the trip is completed, the user device 104, the driver device 106, the merchant device 109, and the server 102 may be updated with the details of the completed financial transaction. This allows for efficient allocation of resources because the available driver fleet is optimized for user demand in each geographic area.
[0039] refer to Figure 3 , we will now describe Figure 2 The communication device 100 includes a communication server 102, and it may include a user communication device 104, a merchant communication device 109, and a driver communication device 106. These devices are connected in a communication network 108 (e.g., the Internet) through corresponding communication links 110, 111, 112, 114 that implement, for example, Internet communication protocols. The communication devices 104, 106, and 109 are capable of communicating through communication networks and / or protocols, including cellular communication networks, LANs, WANs, dedicated data networks, VPNs, fiber optic connections, laser communications, microwave communications, satellite communications, Bluetooth, Wifi, NFC, etc., but for clarity, Figure 3 These are not specified in .
[0040] The communication server device 102 may be Figure 3Alternatively, the functions performed by server device 102 may be distributed across multiple physically or logically separate server components. Figure 3 In the example shown in , the communication server device 102 may include multiple separate components, including but not limited to one or more microprocessors 116, a memory 118 (e.g., volatile memory (such as RAM) and / or a long-term storage device, such as an SSD (solid state or hard disk drive (HDD))) for loading executable instructions 120, the executable instructions defining the functions performed by the server device 102 under the control of the microprocessor 116. The communication server device 102 also includes an input / output module 122 that allows the server to communicate over the communication network 108. A user interface 124 is provided for administrator control and may include, for example, computing peripherals such as a display monitor, a computer keyboard, etc.
[0041] The server device 102 may also include a database 126 stored in the memory 118 for storing data, which may include data about geographic information, images, products, points of interest, users, drivers, merchants, transactions, and other relevant data. The data may be stored in a data structure as required by the application or as described in more detail below. The database 126 may be replicated, distributed, sharded, or otherwise optimized as required by the application or as described in more detail below.
[0042] The user communication device 104 may include a plurality of individual components, including but not limited to one or more microprocessors 128, memory 130 (e.g., volatile memory, such as RAM) and / or a long-term storage device, such as flash memory or SSD (solid state drive), for loading executable instructions 132, which define the functions performed by the user communication device 104 under the control of the microprocessor 128. The user communication device 104 also includes an input / output module 134 that allows the user communication device 104 to communicate via the communication network 108. A user interface 136 is provided for user control. If the user communication device 104 is a smart phone or tablet device, the user interface 136 will have a touch panel display, which is common in many smart phones and other handheld devices. Alternatively, if the user communication device 104 is, for example, a desktop or laptop computer, the user interface 136 may have, for example, computing peripherals, such as a display monitor, a computer keyboard, etc.
[0043] The merchant communication device 109 may be, for example, a smartphone or tablet device having the same or similar hardware architecture as the user communication device 104 .
[0044] The driver communication device 106 may be, for example, a smartphone or tablet device having the same or similar hardware architecture as the user communication device 104. Alternatively, the functionality may be integrated into a custom device, such as a taxi fleet management terminal.
[0045] It can be used as part of a delivery, e-commerce, ride-hailing, mapping, Street View or enterprise mapping solution to provide accurate cross-view matching of professionally sourced Street View imagery, crowdsourced photos and point-of-interest imagery with satellite imagery in a meaningful way. While Google Street View does provide adequate Street View imagery for some locations, it may be outdated or not available at all in some more remote locations. For example, in Southeast Asia, there are many locations without Street View imagery.
[0046] For example, each of the user communication device 104, the driver communication device 106, and / or the merchant communication device 109 may include a camera. In the case of the user communication device 104, the camera 138 may be integrated as part of a smart phone or tablet device. In the case of the driver communication device 106, the camera 140 may be mounted on the driver's vehicle, or mounted on the driver, such as on a helmet. In the case of the merchant communication device 109, the camera 142 may be integrated as part of a smart phone or tablet device. The user communication device 104, the driver communication device 106, and / or the merchant communication device 109 may be individually or collectively referred to as a mobile communication device.
[0047] Cameras 138, 140, 142 can be used to collect street view images suitable for query images in a training dataset for a machine learning model. The cameras can capture 360° geo-tagged still images, or they can capture a reduced field of view (FOV). They can include azimuth estimates (relative to the south heading). Alternatively, for example, camera 140 can include multiple cameras mounted together, each camera having a limited FOV and a different azimuth axis, and the images from each camera are stitched together to form a 360° geo-tagged still image. Images can be captured periodically, for example: every 1 second, they can be captured at pre-specified GPS coordinates, or they can be captured depending on vehicle movement, for example, every 5 meters of travel. The process can be automated, or the user can capture images of specific points of interest and annotate them at the time.
[0048] The geo-tag may be in the form of an estimated longitude and latitude from an onboard GPS module in the corresponding mobile communication device. This may include an estimated compass bearing relative to the camera axis.
[0049] Cameras 138 , 140 , 142 may be used to collect street view images suitable for use as query images during inference using a machine learning model.
[0050] Therefore, it may be desirable to provide a robust cross-view matching machine learning model that can be used in Southeast Asia or other locations where Google Street View is outdated or non-existent.
[0051] In addition to accurately estimating the geographic location of street view images, the inventors also attempt to estimate the fine camera orientation of street view images for three reasons:
[0052] 1) With accurate camera orientation, information extracted from street view imagery, especially using single-image algorithms (e.g., depth estimation, object detection), enables a wider range of real-world applications, such as map creation, augmented reality, navigation. In practice, even very small misalignments in orientation can propagate to large shifts in the physical locations of objects detected in the image. For example, Figure 1 (c) shows that an orientation error of 15° is enough to mislocate the exit to the other lane; an error of 30° is enough to incorrectly assign the attribute to the reverse direction of the road.
[0053] 2) Nowadays, crowd-sourced high-quality 360° or wide-angle images can be taken with semi-professional 360° cameras or even mobile phones. These images are usually not carefully calibrated. It is more likely that the orientation information is missing but the rough location is marked, rather than the other way around. Therefore, the problem of finding the location of a street view image, assuming the orientation is known, is no longer a practical problem.
[0054] 3) Based on our experiments, we observed that the performance of geolocalization can be further improved by introducing finer granularity in the orientation estimation. We hypothesize that finding fine-grained orientations may also improve the performance of geolocalization of crowdsourced images.
[0055] Therefore, it will be understood that Figure 2 , Figure 3 and Figure 5b And the foregoing description shows and describes a system comprising:
[0056] Communication server 102;
[0057] at least one mobile communication device 104, 106, 109; and
[0058] The communication network device 108 is configured to establish communication with the communication server 102 and at least one mobile communication device 104, 106, 109;
[0059] The mobile communication devices 104, 106, and 109 include a first processor and a first memory. The mobile communication devices 104, 106, and 109 are configured to execute a first instruction stored in the first memory under the control of the first processor to:
[0060] capturing a query image;
[0061] transmitting the query image to the communication server 102;
[0062] And wherein the communication server 102 includes a second processor and a second memory, and the communication server 102 is configured to execute a second instruction stored in the second memory under the control of the second processor to:
[0063] In the inference phase, the azimuth rotation and / or geographic location of the query image is estimated using a machine learning model trained based on the weighted soft margin triplet loss function and the absolute angle error loss function.
[0064] Furthermore, it will be understood that Figure 5a A method performed in a communication server device 102 is shown and described, the method comprising: under the control of a microprocessor 116 of the server device 102:
[0065] storing a training data set including a plurality of geo-tagged candidate images and a plurality of query images, each query image having at least one corresponding candidate image having the same geographic location;
[0066] applying a quasi-random or random azimuth rotation to each of the plurality of query images, and storing the azimuth rotation of each of the plurality of rotated query images;
[0067] Training machine learning models, including:
[0068] Extracting features from multiple rotated query images;
[0069] estimating an azimuthal rotation of the rotated query image based on the extracted features of the rotated query image and the inference of features extracted from the candidate images, and
[0070] An objective function is used, the objective function comprising a first loss function based on a weighted soft margin triplet loss, and a second loss function based on an absolute angular error between a stored azimuth rotation and an estimated azimuth rotation of the stored data set.
[0071] Those skilled in the art can adjust the specific methods of selecting a machine learning model, selecting a data set, using the data set to train the model, validating the model, testing the model, and using the model for inference according to the requirements of the desired application. An exemplary implementation will be given below.
[0072] Coordinate system
[0073] Besides geolocation, finding the accurate camera orientation is another critical task in preparing Street View imagery to a “ready-to-use” state. Figure 4 (a) shows the three camera angles 402 required for calibration. Pitch and roll angles and other camera distortions can be corrected to provide upright rectified images. After such correction, the Street View image is upright and only the yaw angle, "azimuth rotation" 404 or orientation needs to be estimated. In some training data sets, the Street View image is oriented north-aligned, meaning that the center column of the image points to the geographic North Pole (denoted by Figure 4 (c) marked by arrow 406 at the top) to ensure that it is aligned with the north direction of the georeferenced satellite image (marked by Figure 4 (b) with arrow 408 marking the north direction).
[0074] If the orientation of the street view image is known, geolocation can achieve better performance. The prior art may have different definitions of orientation misalignment and error.
[0075] Given an upright street view image I g ( Figure 4 (c) top) or “query image” and a set of georeferenced satellite images ( Figure 4 (b) “Candidate images”: the system should identify the candidate images that match the image I from a set of satellite image candidates. g Satellite image I at the same location s I s The center position of the g location.
[0076] Given a set of upright street view images I g = {I g} and a set of georeferenced satellite images I s = {I s}, which are paired and cropped at the same location in the paired street view images, creating an orientation misalignment θ for each street view and satellite pair gt (“quasi-random or random azimuth rotation”). For each street view query image I g , estimate query I g with I S The similarity and orientation θ between each satellite image candidate est The satellite candidates are ranked according to their similarity. The center position of the first satellite image and the estimated orientation of the correct match are extracted as the query image I g Our goal may be to reduce the estimated position θ est With θ gt while maintaining or improving the recall of geolocation.
[0077] Machine Learning Models
[0078] Our model is trained with unknown orientations or "quasi-random or random azimuth rotations" (even though the actual orientation is stored for use in the loss function, as explained later). Specific details of the machine learning model are provided in the exemplary embodiment below.
[0079] Dataset
[0080] The training, validation, and testing datasets can be selected based on the requirements of the application.
[0081] The training dataset may contain two types of images, namely, orthorectified street view images and orthorectified aerial images. This may be captured by cameras 138, 140, and / or 142 using a structured process, or an existing dataset may be used if applicable to the application.
[0082] After the structural process of image acquisition, street view images should ideally include one or more of the following conditions:
[0083] After preprocessing, the images are upright rectified, which means that only the azimuth is not calibrated (in Figure 4 In (a), the yaw angle is the azimuth angle).
[0084] • The image can have a full field of view (FOV), such as a 360° image, or a limited FOV (less than 360 degrees).
[0085] The image can be a single image taken by a phone, a standard camera, or an image stitched together from multiple images.
[0086] The image may have small angular uncertainties in roll and pitch ( Figure 4 (a)).
[0087] Aerial images can be extracted from existing satellite image libraries or acquired through structured acquisition. After preprocessing, aerial images should ideally include one or more of the following conditions:
[0088] · Aerial imagery is orthorectified and georeferenced.
[0089] Images can be taken by different platforms, satellites, aircraft, UAVs, etc.
[0090] Imagery should be of high resolution, e.g. image ground sampling distance (GSD) less than 1 meter per pixel.
[0091] Aerial images can be image chips with standard sizes cropped from larger aerial image tiles.
[0092] In order to use the dataset for supervised training, there must be some pre-existing relationship between the street view images and the aerial images. For the ground truth pairing:
[0093] For each query street view image, its location may be covered by one or more matching aerial images.
[0094] • In the case of creating one-to-one pairings, each query street view image has one positive matching aerial image.
[0095] In the case of creating one-to-many pairings, each query street view image can have multiple positive matching aerial images. However, among all the matching aerial images, there should be a scoring system / ranking to indicate the best to worst matches. For example, this can be calculated by the distance between the query image location and the center location of the aerial image.
[0096] Aerial imagery may also not match any query street view imagery.
[0097] For cropping of aerial images, there are two ways to prepare the dataset:
[0098] oAerial imagery is cropped around the location of the Street View imagery.
[0099] o Aerial imagery is cropped uniformly over the region of interest (e.g., city, state) with / without overlap between image chips.
[0100] The locations of street view images can be used as additional information for dataset creation or evaluation.
[0101] For example, existing datasets include CVUSA and CVACT. Both datasets contain 35,532 training street-satellite matching pairs and 8,884 testing pairs. All images are angularly aligned. During training, random offsets and crops (if FOV is limited) are applied to street view images. During testing, we follow the azimuth offset given for each matching pair. Note that these two datasets were collected in the United States and Australia, respectively, and have a non-negligible domain offset. CVUSA contains a mixture of commercial, residential, suburban, and rural areas, and CVACT tends to be urban / suburban in style. In addition, the satellite images in CVUSA have higher ground coverage but lower resolution than those in CVACT, which introduces another substantial domain offset.
[0102] Note that these two datasets were collected in the United States and Australia, respectively, and have a non-negligible domain shift. CVUSA contains a mixture of commercial, residential, suburban, and rural areas, and CVACT tends to be more urban / suburban in style. In addition, the satellite images in CVUSA have higher ground coverage but lower resolution than CVACT, which introduces another substantial domain shift.
[0103] Training Methods
[0104] The training method can be selected based on the requirements of the specific application.
[0105] For example, the training method 500 is shown in Figure 5(a). It includes random azimuth rotation 502 for street view images, polar coordinate transformation 504 for satellite images, feature extraction 506, fine-grained orientation extraction 508, orientation estimation 510, angle loss based on absolute angle error 512 and triplet loss function 514.
[0106] Training methods can be used Figure 2 and Figure 3 It is implemented by the server 102 in.
[0107] Preprocessing
[0108] Before passing the input images to the feature extractor, the following preprocessing is applied to the street view images and satellite images respectively:
[0109] During training, street view images are randomly rotated to introduce orientation misalignment 502. The ground truth offset w in feature space is recorded gt , to calculate the angle loss and orientation estimation accuracy, where θ gt =w gt / width(Fs)*360°, and Fs is the extracted feature from the satellite image.
[0110] • The satellite image is polar transformed 504 to a similar viewpoint and size as the street view image to physically reduce the gap between the street view and satellite view imagery. Figure 5a The upper left corner shows the effect of polarity reversal.
[0111] All Street View images and polar transformed satellite images are resized to [128, 512] pixels in height and width. [128, 512] is used for polar transformed satellite images and 360° Street View images. For Street View images with limited FOV, the width is reduced proportional to the FOV. For example, 180° images are resized to [128, 256]. The output of this step should be proportional to the FOV.
[0112] To create a misalignment between the Street View image and the georeferenced satellite image, we randomly shift the Street View image clockwise and record this angular offset θ in the south-aligned reference coordinates. gt . Figure 4(c) shows an example of shifting the street view image 315° clockwise. Semantically, the augmentation crops the rightmost portion 410 outside the offset angle, stitches it to the leftmost column 412 of the street view image, and then crops 414 the output of the desired field of view (FOV). Figure 4 In (c), box 416 shows a 180° FOV and box 418 shows a 360° FOV. Although the original image pair is north-aligned, we choose south alignment as a reference, i.e., the first column of the image is aligned with the south direction of the satellite image (e.g. Figure 4 The change is made for images with limited FOV. After cropping, the center column is offset to the left and no longer matches the ground truth value θ gt For example, in a north-oriented coordinate system, if a 360° image is offset 90 degrees clockwise and cropped to a 180° FOV, the angular offset between the center column and the north direction is 45 degrees; if cropped to 120°, the angular offset becomes 30 degrees.
[0113] Table 1 shows the different methods used to create misalignment. Compared with the DSM paper, our method is FOV invariant, θ gt Independent of FOV. Compared to Sijie Zhu, Taojiannan Yang, and Chen Chen. 2021. Revisitingstreet-to-aerial view image geo-location and orientation estimation. In Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, pages 756–765. Rotating street view images is easier to implement, and our method avoids losing the corners of satellite images due to distortion caused by rotation and interpolation.
[0114] Table 1: Different methods of rotating and creating misalignments.
[0115] method Rotate to open Alignment FOV invariance DSM
[36] Street View north Zhu et al.
[61] satellite South √ Our Street View South √
[0116] We propose to calculate the estimated angular offset θ est Angular deviation θ from the ground truth gt The absolute angle error between gt Centered on the error, the error counts up to 180 degrees in the opposite direction ( Figure 4 (b)). Absolute angle error coordinates may have the following advantages:
[0117] Avoid 360° clockwise systems and [-180, 180] system The abrupt change of 360 degrees around the boundary. Our absolute error coordinates are continuous everywhere, which is more natural for defining the angular loss function.
[0118] · Calculate the angle error more fairly. For example, when the ground truth θ gt = 0°, the two estimated values are θ est1 =170° and θ est2 =240°. In a 360° system, θ est1 The angular error is 170°, and θ est2 The angular error is 240°, so even if θ est2 Closer to θ gt , the absolute error is 120°, and θ should also be given priority est1 Our system should prioritize θ est2 Instead of θ est1 .
[0119] The angular error can be easily calculated from the two angular offsets given in the south-aligned coordinate system. Given θ1 and θ2 in the south-aligned coordinate system, their angular difference can be calculated as:
[0120]
[0121] For unknown orientations, the image can be randomly rotated to 360°. The unit of rotation is about 0.70°, corresponding to a shift of one pixel of the input image. For example, 360 degrees divided by 512 pixels = 0.7° in the above example.
[0122] Algorithm 3 shows the augmentation of rotating the street view image in feature space and creating an orientation-shifted ground truth. It allows sub-pixel localization of the ground truth for refined orientation estimation.
[0123]
[0124] We use polar coordinate transformation to preprocess satellite images. Given a size of (S s , S s ) satellite image and convert it into a street view image of size (H g , W g ) plane image, satellite coordinates With street view coordinates The pixel relationship between (both coordinates are based on the upper left corner):
[0125]
[0126] Feature extraction
[0127] The preprocessed satellite and street view images are sent to the Siamese feature extractor 506, namely VGG16. The extractor consists of the first 10 layers of VGG16 and 3 additional convolutional layers to further compress the feature maps on the vertical axis. The output feature depth of the last three layers on the vertical and horizontal axes is [256, 64, 16] and the stride is [(2, 1), (2, 1), (1, 1)]. For the full 360° input of size [128, 512, 3], two feature maps Fs and Fg of size [4, 64, 16] in [H, W, C] are extracted respectively. The two branches have the same architecture and no weight sharing.
[0128] Fine-grained orientation estimation
[0129] The extracted features Fs and Fg are processed by a fine-grained orientation extractor 508 to infer 510 a sub-pixel level angular offset θ est .
[0130] To find fine-grained orientation (at the sub-degree level) 508, we propose two methods to increase the granularity of the estimate without increasing the number of learnable parameters.
[0131] After obtaining the features Fs and Fg, the cross-correlation is calculated as follows:
[0132]
[0133] where Fg / s[m] is the feature slice at horizontal position m across all heights and channels. Ws and Wg are the widths of Fs and Fg. The position with the highest cross-correlation value is considered as the estimated angular offset w in the feature space. est , to estimate the position 508.
[0134] Since our input image has a width of 512 pixels, the offset unit is about 0.7 degrees (360° / 512 pixels). However, Fs and the cross-correlation result (Equation 2) have a width of only 64 pixels (the satellite image has the same width, but the extracted features have a shorter width and are compressed by the model), which makes the orientation extractor have a maximum resolution of 5.625 degrees (360° / 64 pixels). For some applications, it may be desirable to increase the resolution. To refine the granularity of the estimate, we propose two methods.
[0135]
[0136]
[0137] Feature Interpolation (FI): According to Algorithm 1, Fs and Fg are interpolated by a scaling factor S before computing the fine-grained cross-correlation curve. In our implementation, we increase the granularity by a factor of 10. The bin number with the maximum cross-correlation curve is extracted and divided by S to obtain the pixel-level estimate. The resolution of the model is refined to 0.5625 degrees (360° / 640 pixels).
[0138] Curve smoothing (CS): According to Algorithm 2, the cross-correlation curve ((Fg★Fs)[w]) is calculated at the original resolution (64 bins). In order to smooth the curve with a scaling factor S = 10, the rough cross-correlation curve is transformed into the frequency domain using a fast Fourier transform (FFT) and zero-filled by (S-1) times in the middle of the curve:
[0139] ZP(F[w],S)=Concat(F[0:int(W / 2)],
[0140] zeros[0:(S-1)W],F[int(w / 2)+1:]),
[0141] Where W is the width of F[w]. The zero-padded output is again transformed back by inverse FFT (IFFT). The fine-grained orientation extractor with CS has a resolution of 0.5625 degrees.
[0142] Both CS and FI provide the flexibility to adjust the granularity of orientation extraction via a variable zoom factor. For example, if the street view image does not have a full FOV, it may be more difficult to find the correct result. Therefore, in this case, the user may want to reduce the difficulty of the task by using a smaller zoom factor.
[0143] Objective Function
[0144] Based on the estimated orientation, the satellite features Fs can be azimuthally offset (to match the estimated orientation) and cropped (F's) to align with the street view features Fg in orientation and FOV. The output features F's and Fg are used to compute the triplet loss. In addition, the angular loss is used to provide direct supervision on the orientation estimation.
[0145] To provide direct supervision for angle estimation, we propose an angle loss based on absolute angle error 512. Given the ground truth orientation w in the feature space gt , estimated position w est and the width W of the feature map space, the angular loss is given by:
[0146] L 角度 =(0.5W-||w gt -w est |-0.5W|) / 0.5W, (3)
[0147] This corresponds to the ratio of the angular error to the maximum error of 180°.
[0148] Note that this loss only applies to matched pairs.
[0149] For the geolocation task, we use a weighted soft margin triplet loss 514. Given a triplet consisting of an anchor query image A in one view, a positive sample P (correct match) in another view, and a negative sample N, the distance between the feature extracted from A (FA) and the offset and cropped feature in P (F'P) should be smaller than the distance between the offset and cropped feature in N (F'N). We use the cosine distance between features in our implementation. The loss function is given by:
[0150] L 匹配 =In(1+exp(α(D(F A ,F′ P )-D(F A ,F′ N ))))
[0151] D(F1,F2)=2(1-cos(F1,F2)), (4)
[0152] Where α = 10. For the case of batch size B, each query can form (B-1) triples. In each matching direction (street → satellite or satellite → street), B(B-1) triples are constructed. We enforce matching in both directions and there are a total of 2B(B-1) triples in each mini-batch. The overall objective function is:
[0153] L=L 匹配 (A,P,N)+βL 角度 (A,P). (5)
[0154] The loss function weight of the angle loss can be set to β = 0.3. However, depending on the application, the weight of the angle loss can vary between 0.1 and 0.5.
[0155] Implementation details
[0156] Our models are trained with unknown orientation. The first 10 layers of the VGG16 based feature extractor use pre-trained weights on ImageNet, and the last three layers are randomly initialized. Note that in our model, all parameters are learnable. The batch size B is set to 32. We use the Adam optimizer with an initial learning rate between 5 and 15, or for example 11e-5, with a learning rate reduction factor of 0.5 on the platform and patience 8. The maximum training time is set to 200 epochs, and the early stopping threshold is 30 epochs. The models are trained using 4 NVIDIA Tesla V100 GPUs.
[0157] Inference method
[0158] Figure 5b is an example of the inference phase. The query image could be a crowdsourced street view image, 360°, or limited FOV.
[0159] The inference method can be selected based on the requirements of a specific application.
[0160] The dataset used for inference may include, for example, a set of satellite images of a given geographic area. This may allow query street view images within that area to be submitted for inference. Once the query images have an estimated orientation and / or geographic location, they can be incorporated into the dataset for later use.
[0161] For example, the inference method 550 is shown in Figure 5(b). It includes polar transformation 554 of satellite images (as detailed above during training), feature extraction 556, fine-grained orientation extraction 558 (FI or CS as detailed above during training), orientation estimation 560, and ranking by similarity score 562. The similarity score can use cos(Fs, Fg).
[0162] Inference method 550 may use Figure 2 and Figure 3 The server 102 and the mobile devices 104, 106, 109 in the embodiment of the present invention are implemented. The mobile devices 104, 106, 109 can be used to transmit the query image. The data set including the satellite image, the query image (with the orientation and / or geographic location estimate) and / or the extracted and located POI can be stored in the memory 118 or the database 126 for later use in the desired application, such as navigation, VR / AR, 3D map visualization, etc.
[0163] During inference, if the bearing between the satellite and the street view is known (θ gt=0), then the orientation extractor is not used and only cropping Fs is applied to obtain the same FOV; if the orientation is unknown, the fine-grained orientation extractor finds the alignment. Fs is offset and cropped. Given a query street view image, feature similarities and orientation estimates are computed between the query and all possible satellite candidates.
[0164] POI Extraction from Street View Images
[0165] Given a street view image without location / coarse location and orientation information, the location and orientation are first estimated by cross-matching with satellite image candidates in ROI (Region of Interest). The region of interest can be: Large ROI - anywhere in the world; Urban ROI - a specific city or region; Local ROI - selected streets / neighborhoods; Point ROI - an area within 100 meters. The location and orientation of the image with the highest similarity ranking is used as the result.
[0166] Secondly, based on the extracted geolocation and orientation, the street view image is rotated to the correct orientation and given a 2D location on the map. For human viewing applications, this will be an interactive system for the current user's point of view. For downstream applications, only the orientation south of the input image is stored.
[0167] Third, points of interest or objects (buildings, signs, etc.) are extracted from the rotated street view image and the locations of the objects in image coordinates are estimated.
[0168] Fourth, based on the image's geographic location and orientation, object information from the image coordinate system is transferred to the world coordinate system. Downstream applications (such as object detection) can extract the object segmentation (its location in the image) and obtain the relative position of the object to the center of the camera. Then, when we know the orientation and position of the image, we can convert the extracted information to the world coordinate system. Extraction and translation can be part of the downstream task function.
[0169] Fifth, the object's information is placed on the map. This information may include what the object is, a cropped image of the object, the object's location, etc.
[0170] In another case, given a street view image with accurate location but no orientation information, the orientation is first estimated by matching with satellite image candidates cropped at the given location. Satellite images usually come in the form of image patches. Usually, each patch covers a large area, perhaps several square kilometers. Each patch can be cropped into smaller satellite chips as input to the feature extractor.
[0171] Secondly, based on the position and extracted orientation, the street view image is rotated and given a 2D position on the map.
[0172] Third, objects of interest (buildings, signs, etc.) are extracted from the rotated street view image and their locations in image coordinates are estimated.
[0173] Fourth, based on the geographic location and orientation of the image, the object information is transferred from the image coordinate system to the world coordinate system.
[0174] Fifth, the object's information is placed on the map.
[0175] Since the information of the object is on the map, it can be used for navigation viewing or information viewing.
[0176] Alternatively, these steps can include a human validation step instead of obtaining the previous result. The human annotator can select the best result from a set of previous results of the ML model.
[0177] Experimental Results
[0178] Calculate the 1° granularity histogram H(θ) of the absolute angle error. For each image in the test set, given the ground truth θ gt and the estimated value θ est :
[0179] θ err =180°-||θ gt -θ est |-180°|,
[0180] H(θ i )=H(θ i )+1, where θ i -1≤θ err <θ i (6)
[0181] Using fine-grained histograms, one can easily retrieve a 1-to-1 visual comparison between models and the cumulative accuracy curve for any specific degree. It shows the distribution and reliability of the orientation estimates, which is critical for downstream tasks.
[0182] The angular error is averaged over all test images to independently evaluate the orientation performance for geolocation for two reasons: 1) There are alternative sources for obtaining location labels, e.g., social media, and orientation estimation remains a bottleneck for downstream tasks. Any missing orientation information introduces 180° uncertainty to the test image. 2) A 180° error may not affect the geolocation results, but it may be worse for many downstream tasks (e.g., navigation). Therefore, the average is used to linearly compute the error over the entire test set.
[0183] The accuracy of the test image with an estimation error below a certain threshold is calculated. For a given fine-grained histogram H(θ) (Equation 6), the rate below x° is given by:
[0184]
[0185] With these metrics, users can decide whether the estimated orientation is suitable for downstream tasks, which are usually tolerant of orientation errors.
[0186] Table 2 shows the performance of our model on orientation estimation on 360° images. Since only a few works reported their results on orientation estimation and used different metrics, we use the proposed metric for comprehensive comparison.
[0187] The prior art (labeled as DSM
[36] *) is used as the baseline to compute the improvements shown in parentheses. Our models FI and CS achieve significant improvements in the mean errors r@2° and r@5° across all test cases.
[0188] Table 2: Orientation estimates for two datasets.
[0189]
[0190]
[0191] Both datasets have an average error improvement of about 1.38° to 1.52° from the original 5.29° and 6.26°. However, CVUSA observes a higher absolute improvement (about 35%) than CVACT (about 28%) at r@2°. Figure 6 From the histograms shown in , the error distribution of CVUSA is pushed further into the angular error region than the CVACT result distribution. About 57% of the test datasets for CVUSA obtained azimuth estimates with errors below 1°, and about 45% for CVACT. For CVACT, some test cases that were not pushed below the 2° error region were still successfully reduced to within 5°. We assume that the CVUSA dataset contains larger suburban and rural areas than CVACT, which causes the images in CVUSA to naturally have less obvious features, such as buildings. This makes accurate azimuth estimates have a greater impact on CVUSA than CVACT.
[0192] Between FI and CS, FI interpolates coarse feature maps with large scale numbers to generate fine-grained correlation curves with new values, while CS obtains sub-pixel correlation curve values by smoothing the original curves. From the experimental results in Table 1, CS slightly outperforms FI in orientation extraction. However, FI provides a more basic fine-grained orientation curve generation. It may be useful when prior knowledge of coarse orientation is available, which can be added to the street view features before generating the orientation curve.
[0193] Geolocation
[0194] Finding fine-grained orientation not only provides additional orientation information but also improves the performance of geolocation. We evaluate the geolocation results of our model on known / unknown orientation tests. The r@1 of CVUSA and CVACT is shown in Table 3. Note that the performance of our Fl, CS and the best examples of the existing technologies are reported to be fairly compared. Compared with the prior art (labeled as DSM
[36] *), the r@1 of the known / unknown orientation test on CVUSA is improved by 1.93%, FI by 4.70%, and CS by 1.80% and 4.64%; on CVACT, FI is improved by 2.91% and 5.17%, and CS is improved by 2.34% and 5.55%. Additional results on cross-dataset and mixed dataset tests are shown in the supplementary material, as well as visualizations of the top 5 best matches and worst mismatch cases.
[0195] We achieve better r@1 than all existing methods; in particular, we obtain 1.66% and 2.66% absolute improvements on CVACT for known and unknown orientation tests without implementing additional sampling strategies or computationally expensive architectures.
[0196] Table 3: Geolocalization evaluation for two datasets.
[0197]
[0198]
[0199]
[0200] Ablation studies
[0201] In Table 4, the best examples for FI CVACT and CS CVUSA are shown, ‘all’ refers to the results for all test images; ‘match’ refers to the results for only matched images; ‘match all’ refers to the results for the matched case divided by the number of images in the full test set (8884 images). After removing the cross-dataset test. For all tests, r@2° also increases. Both show that images with good geolocation results generally have better orientation estimates. However, if orientation estimation is treated as a separate problem, any missing estimates in the case of test images with 180° uncertainty, the filtering results based on position correctness may end up with a very low percentage of correctly estimated images in the entire dataset. For cross-dataset testing, r@2° (match all) drops to 9.02% and 12.69%, although these models are able to provide high-quality orientations for 57.09% and 55.94% of the full test images. Figure 7 It shows that most of the removed cases obtain high to medium quality orientation estimates. Misplaced images do not necessarily have low quality orientation estimates. Additionally, evaluating only on images with matching positions may lead to unfair comparisons, e.g., the model may fool the evaluation by having only one correctly positioned image with a perfectly estimated orientation. Therefore, we believe that orientation estimation can be evaluated independently of geolocation, unless it is targeted for a specific use case.
[0202] Table 4: Evaluation of orientation estimation on all test data or location matching data.
[0203]
[0204]
[0205] Limited FOV
[0206] We test images using 180° (fisheye camera) and 90° (wide-angle camera). The first test uses a model trained on 360° images to simulate the case where the model is trained on well-collected 360° images, but the test images are crowdsourced and have limited FOV. Table 5 shows the average performance of the FI model on CVACT. The absolute improvement over the state of the art (labeled DSM
[36] *) is given in parentheses. For unknown r@1 geolocation, our results have improvements of 9.60% and 6.83% in 180° and 90° tests. For orientation estimation, we observe improvements in all metrics compared to the state of the art.
[0207] Table 5: Performance of models trained on 360° VACT images and tested on 180°, 90°.
[0208] FOV r@1(%) Average error↓ r@2°(%) r@5°(%) 180° 56.47(9.60) 9.48°(2.79°) 51.12(19.56) 81.85(8.91) 90° 21.88(6.83) 21.57°(3.62°) 30.90(10.92) 55.35(6.64)
[0209] Second, we test on FI models trained on limited FOV to understand whether images with limited FOV still contain enough information to learn fine-grained orientation estimates. Table 6 shows the results for models trained on 180° and 90°. Compared to the state-of-the-art (labeled DSM
[36] *), our model FI improves on most metrics, demonstrating that our approach is applicable to limited FOV. The improvement over the state-of-the-art and the performance itself are much more significant on 180° compared to 90°. This reduction caused by the reduction in FOV is more drastic in the model trained on limited FOV than what we observe in the model trained on full FOV (Table 5). In addition, the unknown r@1 except for 180° achieves higher performance than the model trained on 360°.
[0210] Table 6: CVACT models trained and tested under limited field of view.
[0211]
[0212] Our method improves the orientation estimation performance on both datasets. It has a larger impact on CVUSA because the images in CVUSA have less complex scenes and distinct features. In general, CS obtains slightly better results in our test cases, however, FI provides more basic fine-grained orientation curve generation, which may be useful when prior knowledge of coarse orientation is available. Geolocalization: By incorporating fine-grained orientation estimation, the trained model achieves higher performance compared to the baseline model. Better r@1 scores are achieved than existing methods on both datasets. Evaluation: When street view images are correctly geolocated, the orientation estimation has higher accuracy. However, most of the incorrectly located images can still obtain high to medium quality orientation estimates. It is also fairer to evaluate the orientation estimation independently of geolocalization. Limited FOV: Our model also obtains relative improvements compared to the state-of-the-art models.
[0213] Histograms for testing on the same / across datasets
[0214] Figure 6 , Figure 8 , Fig. 9 and Fig.10 Shown are histograms of the best model of the prior art system (labeled DSM
[36] *) compared to our system, labeled FI and CS, for testing on the same dataset (last 10°) and across datasets (first and last 10°).
[0215] Performance across datasets
[0216] For cross-dataset testing, we train on one dataset and test on another dataset. For example, in Tables 7 and 8, CVACT→CVUSA indicates a model trained on the CVACT training set and tested on the CVUSA test set.
[0217] Table 7: Orientation extraction results in cross-dataset testing.
[0218]
[0219]
[0220] Table 8: Geolocalization results in cross-dataset testing.
[0221]
[0222] The orientation extraction performance is shown in Table 7. Due to the domain shift between the two datasets, both test cases have larger mean errors at the beginning compared to the same dataset test. For both cases, r@2° has an absolute improvement of about 20% to 23% over the original accuracy (about 32%). CVACT→CVUSA achieves better improvements in mean error (about 3.6°) and r@5° (about 9.7%). Combined with the fact that CVACT has worse initial mean error in cross-dataset testing, we believe that the higher proportion of urban images in CVACT causes the model to exploit obvious features, such as buildings. When these features are missing in CVUSA, its performance is worse than the opposite case (CVUSA→CVACT). By integrating fine-grained orientation estimates in training, the model improves generalization and transferability and is less dependent on simple features.
[0223] For geolocation, by learning fine-grained orientation estimates, the trained model achieves higher generalization and transferability in cross-dataset testing without implementing additional sampling strategies or computationally expensive architectures. Compared with our implementation of the prior art (labeled DSM
[36] *), CVUSA→CVACT improves r@1 of known / unknown orientation by 5.11% / 2.43% for FI and 3.57% / 1.01% for CS; CVACT→CVUSA improves by 8.61% / 4.83% for FI and 7.38% / 4.46% for CS. Hongji Yang, Xiufan Lu, and Yingying Zhu. 2021. Cross-view Geo-location with Layer-to-Layer Transformer. Advances in Neural Information Processing Systems 34 (2021), “L2LTR” achieves higher performance on the known orientation test, using a ResNet backbone with 12 layers of visual transformers on each view to bridge the gap between the two views. However, L2LTR does not address orientation uncertainty and can only be applied to images whose orientation is known. This makes it unsuitable for extracting both orientation and location information to prepare street view images into a ready-to-use state.
[0224] Performance on mixed dataset tests
[0225] Tables 9 and 10 show the performance of the best instance of each model trained on CVACT in testing on a mixture of the test sets of both datasets. In addition to the metrics used in the main session, few new metrics are added to understand the improvements in the set and their contribution to the overall performance. For r@1, r@2°, r@5°, and mean error, we show the breakdown on the same dataset (data in the CVACT test set) and across datasets (data in the CVUSA test set). In addition, the hit rate in its own dataset is calculated, which shows the percentage of queries for which each dataset selects satellite candidates from its own satellite pool.
[0226] Table 9: Performance of orientation extraction on the mixed dataset tested on models trained on CVACT.
[0227]
[0228] Table 10: Geolocalization performance on mixed dataset tests for models trained on CVACT.
[0229]
[0230]
[0231] Compared with the existing techniques: (1) Our FI model not only improves the performance on the same test dataset, but also propagates the improvement to cross-test datasets.
[0232] For r@5° mean error, known r@1, and all hit rates, the gains across test sets are even higher than on the same test set. (2) The hit rates for both unknown and known orientation tests achieve more balanced performance and narrow the gap between same and across test sets. Our model gives less unfair treatment to the training dataset. Both results suggest that learning orientation extraction not only improves the performance of the trained model, but also improves the model’s generalization ability to other geographic locations / slightly different acquisition settings.
[0233] Large position offset test
[0234] We consider the pairs in CVUSA and CVACT to be perfectly positionally aligned. However, due to GPS errors, the dataset does contain small positional translation offsets, where the camera positions should be on main roads rather than secondary roads, and the top 1 prediction is actually closer to the actual position. For some extreme cases, the camera positions of street view images are the satellite images they match. The actual camera positions deviate from the matched images (ground truth) and are actually around the top 1 prediction given by our model. Therefore, our model tolerates small positional offsets.
[0235] Table 11: Model performance on the VIGOR dataset using our CS method.
[0236] S10 Known S1 unknown S10 unknown Known r@1 16.99% 20.81% 20.81% Known Hit Rate 22.86% 28.60% 28.56% Unknown r@1 11.47% 16.67% 16.70% Unknown hit rate 16.34% 24.54% 24.46% Average error↓ 36.46° 34.53° 34.51° r@2° 10.54% 9.80% 11.52% r@5° 24.21% 27.23% 26.43%
[0237] We also tested on the VIGOR dataset (where the camera position has a large offset from the center of the satellite image). The results are shown in Table 11. Larger offsets do degrade the overall performance. However, we found that adding fine-grained orientation extraction can still improve performance:
[0238] Compared to models trained with known orientations, training with unknown orientations improves the performance of r@1, r@2°, r@5°, and mean orientation error.
[0239] Training fine-grained orientation with high scaling factors can improve fine-grained orientation extraction compared to models trained with lower scaling factors.
[0240] Performance of geolocation and orientation on two datasets with different configurations.
[0241] Table 12 shows the performance with different configurations. The average performance over three instances of angle weights is reported.
[0242] Table 12: Performance of geolocation and orientation on two datasets with different configurations.
[0243]
Claims
1. A computer-assisted method comprising: storing a training data set including a plurality of geo-tagged candidate images and a plurality of query images, each query image having at least one corresponding candidate image having the same geographic location; applying a quasi-random or random azimuth rotation to each of the plurality of query images, and storing the azimuth rotation of each of the plurality of rotated query images; Training machine learning models, including: extracting features from the plurality of rotated query images; estimating an azimuthal rotation of the rotated query image based on the extracted features of the rotated query image and inferences from the extracted features of the candidate images, and An objective function is used, the objective function comprising a first loss function based on a weighted soft margin triplet loss, and a second loss function based on an error loss between a stored azimuth rotation and an estimated azimuth rotation of the stored data set.
2. The method according to claim 1, further comprising: Each of the plurality of rotated query images is cropped to a restricted field of view.
3. The method according to claim 1 or 2, wherein: The training machine learning model also includes: ranking the correlation between the extracted features of multiple candidate images and the extracted features of the rotated query image, selecting the highest ranked candidate image, and estimating the geographic location of the query image through the geographic location of the highest ranked candidate image.
4. The method according to claim 3, wherein: The training machine learning model further comprises adjusting the orientation of the plurality of candidate images based on the estimated azimuth rotation of the query image and / or cropping the field of view of the plurality of candidate images depending on the field of view of the query image.
5. The method according to claim 3 or 4, further comprising: An approximate geographic location of one or more of the plurality of query images is stored, and a subset of candidate images is selected to be associated with the query image based on proximity to the approximate geographic location of the query image.
6. A method according to any preceding claim, wherein: The objective function is defined as L=L 匹配 (A,P,N)+βL 角度 (A,P)。 7. The method according to claim 6, wherein: The first loss function is defined as 50 匹配 =In(1+exp(α(D(F A ,F P ′ )-D(F A ,F N ′ )))) D(F1,F2)=2(1-cos(F1,F2)).
8. The method according to claim 6 or 7, wherein: The second loss function is defined as <h2 style=";text-align:left;direction:ltr">L<h2 style=";text-align:left;direction:ltr"> 角度 <h2 style=";text-align:left;direction:ltr"> (0.5W-|W<h2 style=";text-align:left;direction:ltr"> gt <h2 style=";text-align:left;direction:ltr"> -W<h2 style=";text-align:left;direction:ltr"> est <h2 style=";text-align:left;direction:ltr"> |-0.5W) / 0.5W.
9. The method according to claim 8, wherein: The error loss is determined using the absolute angle error.
10. The method according to claim 9, wherein: The absolute angular error is calculated using θ err =180°-||θ gt -θ est |-180°| is calculated.
11. The method according to any one of claims 6 to 10, wherein: β is between 0.1 and 0.5 or substantially similar to 0.
3.
12. The method according to any preceding claim, further comprising: One or more metrics are applied to the machine learning model, the metrics selected from the group consisting of: fine-grained histogram, average angular error, accuracy below a certain threshold, and any combination thereof.
13. The method according to claim 12, wherein: The fine-grained histogram is calculated using: i err =180°-||θ gt -θ est |-180°|, H(θ i ) = H(θ i ) + 1, where θ i -1 ≤ θ err < θ i .
14. The method according to claim 12, wherein: The accuracy below a certain threshold is calculated using:
15. The method according to any preceding claim, further comprising: A polar coordinate transform is applied to each of the plurality of candidate images.
16. A method according to any preceding claim, wherein: The applying the random azimuth rotation includes cropping a portion of one side of the image and appending the cropped portion of one side of the image to another side of the image.
17. A method according to any preceding claim, wherein: The training dataset is based on a south-aligned coordinate system, the plurality of query images correspond to street view images, and the plurality of candidate images correspond to aerial images.
18. A method according to any preceding claim, wherein: The training machine learning model also includes: interpolating the extracted features of the rotated query image and the extracted features from the candidate images by a scaling factor, and associating the interpolated extracted features of the rotated query image with the interpolated extracted features from the multiple candidate images using the first loss function.
19. The method according to any one of claims 1 to 17, wherein: The training machine learning model also includes associating the interpolated extracted features of the rotated query image with the interpolated extracted features from the plurality of candidate images using the first loss function, and smoothing a curve associated with the correlation using a scaling factor.
20. The method according to claim 19, wherein: Smoothing the curve includes: Fast Fourier transform (FFT) the correlation curve into the frequency domain; zero-filling a predetermined number of times into the middle of the transformed curve; and Inverse Fast Fourier Transform (FFT) zero-filled curve.
21. A method comprising: A trained machine learning model is used in the inference phase, wherein the machine learning model is trained using the method according to any one of claims 1 to 20.
22. A system comprising: Communication server; at least one mobile communication device; as well as a communication network device configured to establish communication with the communication server and the at least one mobile communication device; The mobile communication device includes a first processor and a first memory, and the mobile communication device is configured to execute a first instruction stored in the first memory under the control of the first processor to: capturing a query image; transmitting the query image to the communication server; And wherein the communication server includes a second processor and a second memory, and the communication server is configured to execute a second instruction stored in the second memory under the control of the second processor to: In the inference phase, a machine learning model trained according to any one of claims 1 to 20 is used to estimate the azimuthal rotation and / or geographic location of the query image.
23. A mobile communication device according to at least one mobile communication device of claim 22.
24. A computer-aided method for position estimation and / or geographic location estimation using a machine learning model, comprising: Extract features from the query image; interpolating the extracted features of the query image and the extracted features from a plurality of candidate images by a scaling factor; estimating an azimuthal rotation of the interpolated query image based on inference of the extracted features of the interpolated query image and the extracted features from the plurality of interpolated candidate images; shifting the plurality of interpolation candidate images based on the estimated azimuth rotation; determining a similarity score between the interpolated query image and the plurality of interpolation candidate images; as well as A geographic location of the query image is inferred based on the similarity score.
25. A computer-aided method for position estimation and / or geographic location estimation using a machine learning model, comprising: Extract features from the query image; estimating an azimuthal rotation of the query image based on the extracted features of the query image and inferences of the extracted features from a plurality of candidate images; shifting the interpolated candidate image based on the estimated azimuth rotation; associating the extracted features of the query image with extracted features from a plurality of offset candidate images; Smooth the curves associated with the correlation; as well as The geographic location of the query image is inferred based on the smoothed correlation curve using the scaling factor.
26. The method according to claim 25, wherein: Smoothing the curve includes: Fast Fourier transform (FFT) the correlation curve into frequency domain; zero-filling a predetermined number of times into the middle of the transformed curve; and Inverse Fast Fourier Transform (FFT) zero-filled curve.
27. The method according to any one of claims 24 to 26, further comprising: The user selects the scaling factor.